Indoor three-dimensional scene generation method and device, equipment and medium
In the indoor three-dimensional scene generation, the initial scene is generated based on room information and object data sets, and the problem of unstable generation quality in the existing technology is solved through the replacement object model and evaluation and screening process, and the generation of high-quality and diverse three-dimensional scenes is achieved.
Patent Information
- Application Number
- CN202411941321.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems of unstable quality when generating indoor three-dimensional scenes.
By determining the three-dimensional spatial information and functional information of the target room based on the room information, combining the preset object data set, an initial three-dimensional scene is generated, and the scene is transformed by replacing the object model. Finally, multiple transformation scenarios are evaluated and filtered according to the preset evaluation dimensions to obtain the target three-dimensional scenario.
Improve the quality and diversity of three-dimensional scene generation, ensuring that the generated scene meets design requirements and has high visual quality.
Smart Images

Figure CN120070729A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of three-dimensional scene technology, and in particular to a method, device, equipment and medium for generating an indoor three-dimensional scene. Background Art
[0002] In recent years, the development of modeling technology has enabled practitioners to efficiently construct high-quality three-dimensional indoor scenes through technical art means, meeting the data requirements of scientific research and practical applications.
[0003] The prior art can use ATISS to generate indoor scenes. ATISS is based on the transform architecture and synthesizes the indoor environment through the target room type and floor plan input by the user. This technical solution regards scene synthesis as an unordered set generation problem and uses the transform architecture to simulate this process.
[0004] However, the prior art has the problem of unstable generation quality. Summary of the Invention
[0005] This application provides a method, device, equipment and medium for generating an indoor three-dimensional scene to solve the problem of unstable generation quality existing in the prior art.
[0006] In a first aspect, this application provides a method for generating an indoor three-dimensional scene, including:
[0007] Determine the three-dimensional space information and function information of the target room according to the room information, where the room information includes the room type and the number of rooms;
[0008] Determine the initial three-dimensional scene of the target room according to the three-dimensional space information, function information and a preset object data set, where the initial three-dimensional scene includes object models and the object positions of each object model;
[0009] Perform transformation processing on the three-dimensional scene according to the replacement object models of each object model in the initial three-dimensional scene to obtain the transformed three-dimensional scene of the target room;
[0010] Evaluate and screen each transformed three-dimensional scene of the target room according to a preset evaluation dimension to obtain the target three-dimensional scene of the target room.
[0011] In this application, evaluating and screening each transformed three-dimensional scene of the target room according to a preset evaluation dimension to obtain the target three-dimensional scene of the target room includes:
[0012] For each transformed three-dimensional scene of the target room, evaluate the transformed three-dimensional scene according to the physical constraint dimension to obtain a first evaluation result, where the physical constraint dimension includes collision constraint, floor projection, wall penetration constraint, wall attachment constraint and passage constraint;
[0013] Evaluate the transformed three-dimensional scene according to the FID metric and the color histogram to obtain a second evaluation result;
[0014] Evaluate the transformed three-dimensional scene according to the occurrence times of each object in the transformed three-dimensional scene and the occurrence rate of each object in the preset three-dimensional scene dataset to obtain a third evaluation result;
[0015] Evaluate the transformed three-dimensional scene according to the two-dimensional projection of the forward positive direction vector of each object in the transformed three-dimensional scene on the floor plane to obtain a fourth evaluation result;
[0016] Evaluate the transformed three-dimensional scene according to the evaluated three-dimensional scene similar to the transformed three-dimensional scene in the three-dimensional scene dataset to obtain a fifth evaluation result;
[0017] According to the first evaluation result, the second evaluation result, the third evaluation result, the fourth evaluation result and the fifth evaluation result of each transformed three-dimensional scene, screen each transformed three-dimensional scene of the target room to obtain the target three-dimensional scene of the target room.
[0018] In this application, according to the room information, determine the three-dimensional space information and functional information of the target room, including:
[0019] According to the room information, determine the target empty scene corresponding to the room information from the empty scene dataset;
[0020] According to the target empty scene, label the type of the target room in the target empty scene to determine the functional information of the target room;
[0021] According to the size information of the target empty scene, determine the three-dimensional space information of the target room.
[0022] In this application, according to the room information, determine the target empty scene corresponding to the room information from the empty scene dataset, including:
[0023] Perform discretization processing on the room information to obtain a first high-dimensional vector;
[0024] Construct a first KDTree according to the first high-dimensional vector and the second high-dimensional vectors of each empty scene in the empty scene dataset;
[0025] According to the first KDTree, determine the target second high-dimensional vector adjacent to the first high-dimensional vector in the first KDTree;
[0026] According to the target second high-dimensional vector, determine the target empty scene corresponding to the target second high-dimensional vector.
[0027] In this application, according to the three-dimensional space information, function information, and a preset object dataset, an initial three-dimensional scene of the target room is determined, including:
[0028] According to the function information, the room configuration corresponding to the function information is determined from the three-dimensional scene dataset, and the room configuration includes the object type and the number of objects of the object.
[0029] According to the room configuration, an object model combination corresponding to the room configuration is obtained from the three-dimensional scene dataset.
[0030] According to the object model combination, the initial three-dimensional scene of the target room is obtained.
[0031] In this application, according to the function information, the room configuration corresponding to the function information is determined from the three-dimensional scene dataset, including:
[0032] According to the function information, a filtered three-dimensional scene that meets the function information is extracted from the three-dimensional scene dataset.
[0033] For each filtered three-dimensional scene, the object type and the number of objects of each object in the filtered three-dimensional scene are counted to obtain a statistical result.
[0034] According to the statistical results of each filtered three-dimensional scene, the object type and the number of objects with an appearance rate meeting the preset appearance rate are used as the room configuration corresponding to the function information.
[0035] In this application, according to the room configuration, an object model combination corresponding to the room configuration is obtained from the object dataset, including:
[0036] The room configuration is discretized to obtain a third high-dimensional vector.
[0037] According to the third high-dimensional vector and the fourth high-dimensional vector of the same type of three-dimensional scenes in the three-dimensional scene dataset, a second KDTree is constructed, and the same type of three-dimensional scenes are three-dimensional scenes with the same room type as the target room.
[0038] According to the second KDTree, a target fourth high-dimensional vector adjacent to the third high-dimensional vector in the second KDTree is determined.
[0039] According to the target fourth high-dimensional vector, a target same-type three-dimensional scene corresponding to the target fourth high-dimensional vector is determined.
[0040] An object model combination corresponding to the room configuration is obtained from the target same-type three-dimensional scene.
[0041] In this application, according to the replacement object model of each object model in the initial three-dimensional scene, the three-dimensional scene is transformed to obtain a transformed three-dimensional scene of the target room, including:
[0042] For each object model, determine the size information and function information of the object model;
[0043] According to the size information and function information, determine a replacement object model for the object model from the object dataset;
[0044] According to the replacement object model of each object model, perform transformation processing on the object models in the initial three-dimensional scene to obtain multiple transformed three-dimensional scenes of the target room.
[0045] In a second aspect, the present application provides a device for generating an indoor three-dimensional scene, including:
[0046] A room determination module, configured to determine the three-dimensional space information and function information of the target room according to the room information, where the room information includes the room type and the number of rooms;
[0047] An initial three-dimensional scene module, configured to determine the initial three-dimensional scene of the target room according to the three-dimensional space information, function information, and a preset object dataset, where the initial three-dimensional scene includes object models and the object positions of each object model;
[0048] A transformed three-dimensional scene module, configured to perform transformation processing on the three-dimensional scene according to the replacement object models of each object model in the initial three-dimensional scene to obtain a transformed three-dimensional scene of the target room;
[0049] An evaluation and screening module, configured to evaluate and screen each transformed three-dimensional scene of the target room according to a preset evaluation dimension to obtain a target three-dimensional scene of the target room.
[0050] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0051] The memory stores computer execution instructions;
[0052] The processor executes the computer execution instructions stored in the memory to implement the method of the first aspect.
[0053] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method of the first aspect.
[0054] The method, device, equipment and medium for generating an indoor three-dimensional scene provided by this application determine the three-dimensional space information and functional information of a target room according to room information, where the room information includes room type and number of rooms; then determine the initial three-dimensional scene of the target room according to the three-dimensional space information, functional information and a preset object data set, and the initial three-dimensional scene includes object models and the object positions of each object model; then perform transformation processing on the three-dimensional scene according to the replacement object models of each object model in the initial three-dimensional scene to obtain the transformed three-dimensional scene of the target room, increasing the diversity of three-dimensional scene generation; finally, evaluate and screen each transformed three-dimensional scene of the target room according to a preset evaluation dimension to obtain the target three-dimensional scene of the target room, improving the quality of the generated scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application and used together with the specification to explain the principles of this application.
[0056] Figure 1 It is a schematic flowchart of a method for generating an indoor three-dimensional scene provided by an embodiment of this application;
[0057] Figure 2 It is a schematic flowchart of an empty scene matching process provided by an embodiment of this application;
[0058] Figure 3 It is a schematic diagram of a furniture combination feature vector provided by an embodiment of this application;
[0059] Figure 4 It is a schematic flowchart of a furniture combination matching process provided by an embodiment of this application;
[0060] Figure 5 It is a schematic flowchart of a furniture combination layout process provided by an embodiment of this application;
[0061] Figure 6 It is an effect schematic diagram of a combination layout generation module provided by an embodiment of this application;
[0062] Figure 7 It is a schematic flowchart of a semantic metadata replacement process provided by an embodiment of this application;
[0063] Figure 8 It is a schematic flowchart of a spatial distribution replacement process provided by an embodiment of this application;
[0064] Figure 9 It is a schematic diagram of furniture update provided by an embodiment of this application;
[0065] Figure 10 It is a schematic diagram of the steps for filling small furniture provided by an embodiment of this application;
[0066] Figure 11 Schematic diagram of resetting the motion state of furniture components provided by an embodiment of the present application;
[0067] Figure 12 Schematic diagram of resetting the physical metadata of furniture provided by an embodiment of the present application;
[0068] Figure 13 Schematic diagram of the process of the scene screening module provided by an embodiment of the present application;
[0069] Figure 14 Schematic diagram of the structure of the physical constraint module provided by an embodiment of the present application;
[0070] Figure 15 Schematic diagram of the structure of bounding box expansion and isolated islands provided by an embodiment of the present application;
[0071] Figure 16 Schematic diagram of comparing the style differences of rendering results provided by an embodiment of the present application;
[0072] Figure 17 Schematic diagram of the vector representation of furniture distribution provided by an embodiment of the present application;
[0073] Figure 18 Schematic diagram of the COPT effect provided by an embodiment of the present application;
[0074] Figure 19 Schematic diagram of the layout specification effect provided by an embodiment of the present application;
[0075] Figure 20 Schematic diagram of the structure of a device for generating an indoor three-dimensional scene provided by an embodiment of the present application;
[0076] Figure 21 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application.
[0077] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0078] Here, exemplary embodiments will be described in detail, and their examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0079] To clearly understand the technical solution of this application, the solutions of the prior art will be introduced in detail first.
[0080] With the continuous development of high-performance computers and modeling tools, a model library of three-dimensional indoor scenes has been established and continuously expanded, which has promoted the development of three-dimensional perception algorithms and generation algorithms, and also driven the rise of virtual platforms that support the iteration of simulation robot models. The refinement, standardization, diversification, and parameterization of these scenes have become important foundations for virtual perception and interaction, making how to efficiently obtain three-dimensional indoor scenes that meet these requirements a research hotspot in the fields of computer graphics, virtual reality, artificial intelligence, and human-computer interaction, and showing broad application potential in multiple fields such as robotics, intelligent manufacturing, and digital entertainment. In recent years, the development of modeling technology has enabled practitioners to efficiently construct high-quality three-dimensional indoor scenes through technical art means, meeting the data requirements of scientific research and practical applications. At the same time, the development of various automated generation systems has provided new ways to obtain these data.
[0081] However, the existing three-dimensional indoor scene generation algorithms face two main problems: one is the low algorithm efficiency, which cannot quickly generate a large number of scenes that meet the requirements, and usually requires high-computing-power hardware support, restricting its application scope; the other is that scenes with low quality and poor effectiveness are often generated during the generation process, and the existing quality assessment algorithms cannot effectively quantify and evaluate a single scene, resulting in the inability to screen and eliminate low-quality scenes.
[0082] In response to any of the above-mentioned multiple technical problems, the inventor found in the research that the quality of the generated scenes can be improved by evaluating and screening the generated scenes.
[0083] Figure 1 The flowchart of a method for generating an indoor three-dimensional scene provided by an embodiment of this application is as Figure 1 shown, and this method includes:
[0084] S101. According to the room information, determine the three-dimensional space information and function information of the target room, where the room information includes the room type and the number of rooms.
[0085] Among them, the room information may refer to the description information of the rooms in the three-dimensional scene input by the user, such as two bedrooms and one living room.
[0086] The three-dimensional space information may refer to the dimension description information of the room, which is used to determine the space size of the room and facilitate the placement of furniture and other objects.
[0087] The function information may refer to the function description of the target room determined according to the room type, which is used to determine the use of the room. For example, if the room type of the target room is a bedroom, the function information may be a children's bedroom.
[0088] The room type can refer to the classification of rooms, and the room types can include mixed - function rooms, balconies, kitchens, bathrooms, living rooms, master bedrooms, bedrooms, second bedrooms, living halls, corridors, libraries, master bathrooms, children's rooms, secondary bathrooms, foyers, dining rooms, passages, cloakrooms, equipment rooms, storage rooms, rooms for the elderly, laundries, platforms, connecting corridors, stairwells, courtyards, nanny rooms, auditoriums, unfurnished rooms, garages, free spaces, etc.
[0089] In some embodiments, determining the three - dimensional space information and functional information of a target room according to the room information may include:
[0090] Determining a target empty scene corresponding to the room information from the empty scene dataset according to the room information;
[0091] Labeling the type of the target room in the target empty scene according to the target empty scene to determine the functional information of the target room;
[0092] Determining the three - dimensional space information of the target room according to the size information of the target empty scene.
[0093] Among them, the empty scene can refer to a space with a certain volume for placing objects.
[0094] In this application, an empty scene dataset can be constructed according to the existing empty scenes. The existing empty scenes can be collected from the network or generated by a trained mathematical model, and this application does not limit this. First, according to the room information input by the user, search for the empty scene corresponding to the room information in the empty scene dataset, which is called the target empty scene here. After obtaining the target empty scene, label each room in the empty scene to obtain the functional information of each room, and then according to the size information of the target empty scene, obtain the three - dimensional space information of each room. The target room is any room in the target empty scene.
[0095] Furthermore, determining the target empty scene corresponding to the room information from the empty scene dataset may include:
[0096] Performing discretization processing on the room information to obtain a first high - dimensional vector;
[0097] Constructing a first KDTree according to the first high - dimensional vector and the second high - dimensional vectors of each empty scene in the empty scene dataset;
[0098] Determining the target second high - dimensional vector adjacent to the first high - dimensional vector in the first KDTree according to the first KDTree;
[0099] Determining the target empty scene corresponding to the target second high - dimensional vector according to the target second high - dimensional vector.
[0100] Among them, the discretization process may refer to converting room information into a vector representation. For example, if the room information is two bedrooms and one living room, all room types are represented by a vector. The position of each sequence represents a different room type, and the number in each sequence represents the number of rooms. Assuming there are only three room types: bedroom, living room, and dining room, the vector can be represented as (X 1 , X 2 , X 3 ), where X 1 represents that the number of bedrooms is X 1 , X 2 represents that the number of living rooms is X 2 , X 3 represents that the number of dining rooms is X 3 , and the discretization result of two bedrooms and one living room is (2, 1, 0).
[0101] KDTree (K-D tree) may refer to a data structure used to organize points in a k-dimensional space, mainly for searching key data in a multi-dimensional space, such as range search and nearest neighbor search.
[0102] Exemplarily, for three-dimensional scene generation, first obtain the constructed empty scene, and assign corresponding semantics according to the room types and quantities input by the user. Use the three-dimensional indoor room semantics generation module to search and fill the appropriate empty scene in the empty scene dataset. To process complex three-dimensional elements and data structures, a general similar element search framework is introduced, which involves mapping elements to a high-dimensional vector space and using KDTree to find nearest neighbors. Figure 2 This is a schematic diagram of the empty scene matching process provided by the embodiment of the present application. The specific steps include: first discretize the user requirements into high-dimensional vectors, and then map the attributes of the empty scene to the same vector space. After constructing the KDTree, use the tree to search for nearest neighbor vectors within the error range from the empty scene set, and extract the corresponding scene from the metadata set. Finally, fine-tune the semantic annotation of the empty scene. For example, modify the study in the empty scene to a bedroom to ensure meeting the user requirements.
[0103] S102. Determine the initial three-dimensional scene of the target room according to the three-dimensional space information, function information, and a preset object dataset. The initial three-dimensional scene includes object models and the object positions of each object model.
[0104] Among them, the object dataset may refer to a dataset constructed based on existing objects, and the types of objects may include cabinets, tables, chairs, lamps, sofas, ornaments, beds, shelves, storage units, home appliances, pillars, potted plants, tools, toys, statues, wall cabinets, outdoor items, electronic devices, toilets, trash cans, water bottles, building materials, multimedia units, cabinets, daily chemicals, and musical instruments, etc.
[0105] The object model may refer to a three-dimensional model used to display an object.
[0106] The object position may refer to the position of the object in the target room, which can be represented by the three-dimensional coordinate system constructed for the target room.
[0107] Specifically, according to the three-dimensional space information, function information, and a preset object data set, an initial three-dimensional scene of the target room can be determined, which may include:
[0108] According to the function information, determine the room configuration corresponding to the function information from the three-dimensional scene data set. The room configuration includes the object type and the number of objects of the object.
[0109] According to the room configuration, obtain the object model combination corresponding to the room configuration from the three-dimensional scene data set.
[0110] According to the object model combination, obtain the initial three-dimensional scene of the target room.
[0111] Among them, the three-dimensional scene data set may refer to a data set constructed based on existing three-dimensional scenes. The existing three-dimensional scenes can be obtained by acquiring various three-dimensional scene-related images from the network and then converting the images into three-dimensional scenes, or by generating three-dimensional scenes through a trained mathematical model. This application does not limit this.
[0112] The room configuration can determine the use of the target room according to the function information, and then search for three-dimensional scenes with the same use in the three-dimensional scene data set. Refer to the room configuration in the searched three-dimensional scenes to obtain the room configuration of the target room. For example, the function information of the target house is a children's bedroom. Search for three-dimensional scenes related to the children's bedroom in the three-dimensional scene, and then statistically find from the searched three-dimensional scenes that a combination of one bed and two cabinets has a relatively high appearance rate. Therefore, one bed and two cabinets can be used as the room configuration of the children's bedroom.
[0113] After obtaining the room configuration, a set of rooms of the same room type can be searched from the three-dimensional scene data set, and then the optimal object model combination can be obtained from the set of rooms. Place the object model combination in the target room to obtain the initial three-dimensional scene of the target room.
[0114] Furthermore, according to the function information, determining the room configuration corresponding to the function information from the three-dimensional scene data set may include:
[0115] According to the function information, extract the screened three-dimensional scenes that meet the function information from the three-dimensional scene data set.
[0116] For each screened three-dimensional scene, count the object type and the number of objects of each object in the screened three-dimensional scene to obtain a statistical result.
[0117] According to the statistical results of each filtered 3D scene, the object types and object quantities with occurrence rates meeting the preset occurrence rate are used as the room configurations corresponding to the function information.
[0118] In this application, a filtered 3D scene corresponding to the function information can be selected from the 3D scene dataset. For example, if the function information is a children's bedroom, then a 3D scene with a description of a children's bedroom is selected from the 3D scene dataset as the filtered 3D scene. Then, the quantity of each object is counted in the filtered 3D scene. For example, taking each filtered 3D scene as the statistical unit, the type and quantity of each object in each filtered 3D scene are counted. Suppose the statistical results show that the proportion of a bed in each scene is 90%, the proportion of a wardrobe is 85%, and the proportion of a desk is 80%, and the preset occurrence rate is 80%. Then, a bed, a wardrobe, and a desk are used as the room configuration of the children's bedroom.
[0119] Furthermore, according to the room configuration, an object model combination corresponding to the room configuration can be obtained from the object dataset, which may include:
[0120] Perform discretization processing on the room configuration to obtain a third high-dimensional vector;
[0121] According to the third high-dimensional vector and the fourth high-dimensional vector of the same type of 3D scenes in the 3D scene dataset, a second KDTree is constructed. The same type of 3D scenes are 3D scenes with the same room type as the target room;
[0122] According to the second KDTree, determine the target fourth high-dimensional vector adjacent to the third high-dimensional vector in the second KDTree;
[0123] According to the target fourth high-dimensional vector, determine the target same type of 3D scene corresponding to the target fourth high-dimensional vector;
[0124] Obtain the object model combination corresponding to the room configuration from the target same type of 3D scenes.
[0125] In this application, in order to match the object model combination most relevant to the room configuration from the three-dimensional scene dataset, first, the room configuration is discretized into a high-dimensional vector, herein referred to as the third high-dimensional vector, and then each three-dimensional scene in the three-dimensional scene dataset is also discretized into a high-dimensional vector, herein referred to as the fourth high-dimensional vector. A KDTree is constructed based on the third high-dimensional vector and the fourth high-dimensional vector, herein referred to as the second KDTree. Based on the second KDTree, the fourth high-dimensional vector closest to the third high-dimensional vector is matched for the third high-dimensional vector, herein referred to as the target fourth high-dimensional vector. Then, the combination method of the objects corresponding to the target fourth high-dimensional vector is used as the object model combination corresponding to the room configuration. For example, if the room configuration is a bed, a wardrobe, and a desk, after obtaining the target fourth high-dimensional vector from the three-dimensional scene dataset, the three-dimensional scene corresponding to the target fourth high-dimensional vector is determined, and the three-dimensional models of a bed, a wardrobe, and a desk in this three-dimensional scene are extracted to obtain an object model combination of a bed, a wardrobe, and a desk.
[0126] Exemplarily, first, the room semantic generation module of the empty scene determines the types, quantities, and layout positions of the furniture required for the room. Next, the room layout generation includes two main components: a furniture combination screening component and a furniture combination layout component.
[0127] The furniture combination screening component screens out the compliant furniture combinations from the set of the same room types by vectorizing the furniture features and using the KDTree to search for the matching furniture combinations. Figure 3 It is a schematic diagram of the furniture combination feature vector provided by the embodiment of this application. The vectorization involves four groups of features: G1 describes the three-dimensional spatial dimensions of the furniture combination, the quantity and average volume of the included furniture, G2 represents the quantity of specific types of furniture, G3 is about the average volume of the furniture, and G4 calculates the average distance between the furniture. In the screening process, the type and quantity features (G2) of the furniture are first used for preliminary screening, and then other features (G1, G3, G4) are used for fine screening, so as to improve the search efficiency and obtain a more natural layout result. Figure 4 It is a schematic diagram of the furniture combination matching process provided by the embodiment of this application. Among them, the three-dimensional indoor scene meta-dataset is the three-dimensional scene dataset. Both rounds of screening are completed relying on the KDTree. The first round uses the type_tree to screen based on the G2 feature, and the second round uses the item_tree to conduct in-depth screening based on the combination features. This method effectively combines the functionality and spatial adaptability of the furniture combination, ensuring that the generated layout meets the actual requirements and has a high-quality visual performance.
[0128] After obtaining the furniture combination solutions that have passed two rounds of screening, the furniture combination layout component reasonably places these furniture items in the room to be laid out. This process has something in common with the spatial deformation task, where the prior room is regarded as the local space to be deformed, and the room to be laid out is the target space. The furniture combination is equivalent to points in space, and through spatial deformation, the prior room is adapted to the shape of the target room. This involves establishing a two-way mapping of the positions of points in the space before and after deformation to determine the positions of the furniture in the new room. In addition, to reduce the complexity, the concept of a two-dimensional room boundary is introduced, simplifying the task to spatial deformation in a two-dimensional space, similar to a two-dimensional image deformation problem. Figure 5 It is a schematic diagram of the furniture combination layout process provided by an embodiment of this application. By selecting control points and using the properties of generalized barycentric coordinates, a region is deformed, and the new positions of the furniture in the target room are determined. In this way, the structural characteristics of the space can be maintained, and any layout problems caused by spatial differences can be solved through post-processing.
[0129] Figure 6 It is a schematic diagram of the effect of the combination layout generation module provided by an embodiment of this application. Figure 6 (a) shows the positions of the furniture combination to be laid out in the original room. Figure 6 (b) shows the original positions of the furniture combination in the room after directly sampling the target points using generalized barycentric coordinates. Finally, Figure 6 (c) shows the room layout after adjustment through post-processing, which corrects the positions of the furniture to ensure the rationality and aesthetics of the layout.
[0130] S103. According to the replacement object models of each object model in the initial three-dimensional scene, perform transformation processing on the three-dimensional scene to obtain the transformed three-dimensional scene of the target room.
[0131] Among them, the transformation processing may refer to replacing the object models in the initial three-dimensional scene to obtain a new three-dimensional scene. For example, if there are other styles of beds in the dataset for the bed in the initial three-dimensional scene, the bed in the initial three-dimensional scene can be replaced with the bed in the dataset to obtain a new three-dimensional scene, which is called the transformed three-dimensional scene here, so that the styles of the three-dimensional scene are not limited to a few. As long as the number of datasets is rich enough, more transformed three-dimensional scenes can be obtained.
[0132] In some embodiments, according to the replacement object models of each object model in the initial three-dimensional scene, performing transformation processing on the three-dimensional scene to obtain the transformed three-dimensional scene of the target room may include:
[0133] For each object model, determine the size information and functional information of the object model;
[0134] Determine the replacement object model of the object model from the object dataset according to the size information and functional information;
[0135] According to the replacement object model of each object model, perform transformation processing on the object model in the initial three-dimensional scene to obtain multiple transformed three-dimensional scenes of the target room.
[0136] In this application, in order to make the size of the obtained transformed three-dimensional scene meet the requirements, it is necessary to search for and replace the replacement object model from the object dataset according to the size information and functional information of each object model in the initial three-dimensional scene. For example, if the size information of the bed in the initial three-dimensional scene is 1.8m * 2m and the functional information is a storage bed, then a storage bed with a size of 1.8m * 2m is selected from the object dataset.
[0137] Exemplarily, in three-dimensional scene generation, to increase room diversity, a furniture replacement module is developed, including two strategies: semantic metadata replacement and spatial distribution replacement. Semantic metadata replacement is based on the detailed classification annotations (first-level and second-level labels) of furniture. For example, although the "wardrobe" in the bedroom and the "sideboard" in the kitchen are both "cabinets", they have different second-level labels. Figure 7 This is a schematic diagram of the semantic metadata replacement process provided for the embodiments of this application. This strategy selects replacement furniture by matching the second-level labels and pays attention to the consistency of the volume of the replacement furniture with the original furniture to avoid scene distortion.
[0138] The spatial distribution replacement strategy is based on the similarity of the three-dimensional grid spatial distribution of furniture. First, the spatial occupancy features of furniture are vectorized through a feature extraction algorithm, and then the KDTree and KNN (K-Nearest Neighbor) methods are used to quickly find furniture with similar spatial distributions in the large dataset as replacement candidates. This method involves normalizing the furniture to the standard space, dividing the space, and counting the proportion of the cube area to define the spatial distribution of the furniture. Figure 8 This is a schematic diagram of the spatial distribution replacement process provided for the embodiments of this application. By comparing the cosine similarity and KL divergence (relative entropy) of the furniture spatial distribution, replacement options highly similar to the original furniture are accurately selected. The combined action of these two strategies effectively improves the diversity and fidelity of the generated scenes.
[0139] Furthermore, during the generation of 3D indoor scenes, a randomization module is introduced to enhance the diversity of the generated scenes. This module hierarchically randomizes the layout and metadata in the generated scenes, aiming to maximize diversity while maintaining rationality. Randomization is mainly applied to two levels: furniture combinations and furniture, using different strategies to enrich the results. The randomization at the furniture combination level includes fine-tuning the local layout between furniture. This is achieved by updating the layout relationships in the furniture combination, especially the parts with lower importance. The importance of furniture is judged by the vector angle between its positive direction and the center of other furniture. Furniture with a smaller angle is considered to point to other furniture, thus affecting their importance ranking. This information helps to preferentially update the furniture with lower importance. For example, Figure 9 is a schematic diagram of furniture update provided for the embodiments of this application, Figure 9 (a) shows furniture combinations identified by three different colored rectangular boxes in the living room. In Figure 9 (b), the positive directions of the furniture are represented by green arrows, and the red numbers mark the counts of the furniture, indicating how many furniture point to this piece of furniture. Based on the pointing and counting of the furniture, the importance degree of the furniture is determined. The update of the furniture includes deletion, replacement, and addition. The deletion operation calculates the update probability according to the importance of the furniture and randomly executes it under this probability. Figure 9 (c) shows the result of the deletion operation. The replacement and addition operations consider the remaining space in the room and evaluate the space occupancy through the occupy map (occupancy grid map) projected onto the floor plane, as shown in Figure 9 (d). When replacing, neighboring furniture is selected from the object dataset according to the furniture feature vector for adjustment. The addition operation selects furniture within the free grid and allows scaling within a certain range. Figure 9 (e) and Figure 9 (f) respectively show the updated room layout and the updated occupy map, and the latter uses shapes of different colors to identify the positions of the replaced and added furniture. Through these operations, the randomization module effectively enhances the diversity and realism of the scene layout.
[0140] Furthermore, in 3D scene generation, the furniture-level randomization component updates furniture features using two methods. First, small furniture is used to fill the empty space on the surface of large furniture to increase scene diversity. In this process, by projecting the large furniture onto the ground and cutting out horizontal layers with the same height, the space suitable for placing small items is found. Figure 10 is a schematic diagram of the small furniture filling steps provided for the embodiments of this application, including space selection and the final filling effect. Second, the metadata of the furniture is randomized, including visual features, motion states, and physical properties. This involves randomly selecting new materials and adjusting the states of furniture components to refresh the visual and physical performance of the furniture. Figure 11Schematic diagram of the effect of resetting the motion state of furniture components provided by the embodiments of the present application Figure 12 Schematic diagram of the effect of resetting the physical metadata of furniture provided by the embodiments of the present application. The above figure shows the randomization results of motion metadata and physical metadata, highlighting the new appearance after the change of furniture state and materials.
[0141] S104. According to the preset evaluation dimensions, evaluate and screen each transformed three-dimensional scene of the target room to obtain the target three-dimensional scene of the target room.
[0142] In the present application, after obtaining multiple transformed three-dimensional scenes, the transformed three-dimensional scenes can be evaluated multi-dimensionally, the sequence of all transformed three-dimensional scenes can be determined according to the evaluation results, and the transformed three-dimensional scenes that meet the sequence requirements are determined as the target three-dimensional scenes.
[0143] In some embodiments, according to the preset evaluation dimensions, evaluating and screening each transformed three-dimensional scene of the target room to obtain the target three-dimensional scene of the target room may include:
[0144] For each transformed three-dimensional scene of the target room, evaluate the transformed three-dimensional scene according to the physical constraint dimension to obtain the first evaluation result. The physical constraint dimension includes collision constraint, floor projection, wall penetration constraint, wall attachment constraint, and passage constraint;
[0145] Evaluate the transformed three-dimensional scene according to the FID index and color histogram to obtain the second evaluation result;
[0146] Evaluate the transformed three-dimensional scene according to the appearance times of each object in the transformed three-dimensional scene and the appearance rate of each object in the preset three-dimensional scene dataset to obtain the third evaluation result;
[0147] Evaluate the transformed three-dimensional scene according to the two-dimensional projection of the forward positive direction vector of each object in the transformed three-dimensional scene on the floor plane to obtain the fourth evaluation result;
[0148] Evaluate the transformed three-dimensional scene according to the evaluation three-dimensional scene similar to the transformed three-dimensional scene in the three-dimensional scene dataset to obtain the fifth evaluation result;
[0149] According to the first evaluation result, the second evaluation result, the third evaluation result, the fourth evaluation result, and the fifth evaluation result of each transformed three-dimensional scene, screen each transformed three-dimensional scene of the target room to obtain the target three-dimensional scene of the target room.
[0150] In this application, five dimensions can be used to evaluate the transformed three-dimensional scene, obtaining five evaluation results. Then, according to the weights of the parameters in each evaluation result, the final evaluation result of the transformed three-dimensional scene is obtained. Based on the final evaluation result, each transformed three-dimensional scene is screened to obtain the target three-dimensional scene of the target room.
[0151] Exemplarily, this application employs five modules for scene screening and how to adopt a combined screening criterion based on the results of these modules to select the optimal scene. This process aims to ensure that the final scene meets the design requirements and has high visual quality. Figure 13 It is a schematic flow chart of the scene screening module provided by the embodiment of this application.
[0152] First, Figure 14 It is a schematic structural diagram of the physical constraint module provided by the embodiment of this application. The screening module based on physical constraints comprehensively evaluates the quality of the generated scene, including collision constraints, floor projection, wall-penetration constraints, wall-adjacent constraints, and passage constraints components. The collision constraint component uses bounding volumes and precise collision detection methods to identify overlaps between indoor objects, and the output is the collision object ratio (C_Ratio). The wall-penetration constraint component detects the collision between the two-dimensional projection of the three-dimensional indoor object and the room boundary, and its output is the wall-penetration object ratio (W_Ratio). The wall-adjacent constraint component evaluates whether the furniture is close to the wall, with a boundary of 0.3 m, and the output is the wall-adjacent object ratio (N_Ratio). The passage constraint component evaluates the passability of the generated three-dimensional indoor scene when used as a virtual environment (such as a robot model iteration task). This component ensures that all objects are reachable and requires at least a 0.3 m wide passage to be reserved for agents such as robots. By projecting the three-dimensional indoor objects in the room onto the ground and expanding their bounding box sizes, the component detects whether there are impassable island areas caused by indoor objects in the room. It calculates the ratio (A_Ratio) of the largest island to the total area of all islands as the output. Figure 15 It is a schematic structural diagram of the bounding box expansion and islands provided by the embodiment of this application, where green represents the area outside the room, white shows the idle area inside the room, black indicates the area occupied by indoor objects, and red is the island area. The left figure shows the case without islands, while in the right figure, the ratio of the island area to the idle area is 42.93%. These modules work together to evaluate the rationality and real-world adaptability of the scene layout room by room.
[0153] Second, as an important way for humans to obtain information, vision makes the style similarity under the viewing angle an effective means to evaluate scene similarity. In the screening module based on style constraints, the scene is rendered from a top-down perspective and a unified monochromatic material is applied to exclude the interference of material differences. Figure 16Schematic diagram for comparing the style differences of rendering results provided by embodiments of this application. This module uses two methods, namely Fréchet Inception Distance (FID) and color histogram, to evaluate the style similarity. FID measures the difference in probability distribution between the generated image and the real image, and the lower the ideal value, the higher the image quality. The color histogram evaluates the distribution of color changes in the image and calculates the color gradient consistency between different scenes through cosine similarity. This module selects the five scenes most similar to the current scene from the 3D scene dataset through KDTree for comparison. When selecting the current scene to compare with the other five scenes, the optimal score of each index of the current scene is used as the final score, and then the average value of the two indexes is calculated as the output to further ensure the visual consistency and high quality of the generated scene.
[0154] Third, in the generated 3D indoor scene, each room is given a specific semantic annotation according to its function and contains furniture types that conform to this semantics. This module aims to evaluate the consistency between the furniture distribution in the generated room and the average furniture distribution of the same room type in the existing dataset. This process is achieved by comparing the differences in furniture distribution. First, the module calculates the probability of the occurrence of specific furniture in each room type to form an overall probability distribution vector, where each dimension represents a different furniture category and its occurrence probability. Figure 17 Schematic diagram of the vector representation of the furniture distribution provided by embodiments of this application. Then, the furniture occurrence in the current room is counted to form a sample probability distribution vector, where each dimension represents a furniture category and the value is 1 when the furniture exists. Finally, the cosine similarity (denoted as F_Ratio) between these two vectors is calculated to measure the similarity between the furniture distribution in the generated room and the expected distribution, and this value is used as the output of the module to evaluate the quality and accuracy of the generated scene.
[0155] Fourth, in 3D indoor scene generation, the screening module evaluates the difference in furniture distribution within a room by comparing the distances between the undirected graphs corresponding to the rooms. Through KDTree, this module selects the five rooms with the most similar furniture types and quantities to the current room and constructs a fully connected undirected graph. Using the COPT (Shanshu Solver) function, based on the optimal transport theory, it calculates the difference between the graphs and then measures the similarity of the rooms. The COPT method regards the transformation of the graph as a graph transport problem and simplifies the calculation using the Laplacian matrix of the graph. The Laplacian matrix, which reflects the structure of the graph, is defined by the adjacency matrix and the degree matrix. The calculation of COPT involves the number of nodes in the graph, the trace of the Laplacian matrix, and the operation of its pseudo-inverse. This method measures the cost of the probability distribution of the node mapping from one graph to another. The result output is G_Ratio, which is a normalized graph distance value. If the COPT measurement is greater than or equal to 6, then G_Ratio is 0; if it is less than 6, then it is calculated as (6 - COPT(X,Y)) / 6, reflecting the similarity degree of the 3D object distribution in the two rooms. Figure 18 This is a schematic diagram of the COPT effect provided by the embodiment of the present application. This method allows for the precise measurement of the structural differences between the generated room and the reference room, ensuring the rationality and consistency of the furniture layout in the generated scene.
[0156] Fifth, Figure 19 This is a schematic diagram of the layout specification effect provided by the embodiment of the present application. In the generated 3D indoor scene, although the furniture layout in some rooms conforms to the type average distribution, there are no collision or blocking areas, and it is similar to the style of similar rooms, the layout may appear chaotic and does not meet the standards of high-quality rooms. To evaluate the degree of messiness of the room layout, the module analyzes the two-dimensional projection of the forward positive direction vectors of the furniture on the floor plane. The inconsistent directions of these vectors may indicate the chaos of the layout. By regarding the geometry of the forward positive direction vectors of the furniture as samples of a Gaussian distribution in a two-dimensional plane, kernel density estimation (KDE) is used to estimate the parameters of this distribution, mainly the variance. KDE is a non-parametric method that estimates the probability density function by placing Gaussian kernel functions at each data point and superimposing them, thereby obtaining the degree of messiness of the overall layout. The larger the variance, the more chaotic the layout. By calculating the variance, Sigma_Ratio is further obtained, which is a quantitative index used to evaluate the standardization degree of the room furniture layout, so as to screen out the rooms that meet the high-quality standards. This method not only considers the quantity and types of furniture, but also considers their relative positions and directions in the room, ensuring that the generated scene is both practical and beautiful.
[0157] Finally, in the 3D indoor scene screening, a combined screening criterion is required. Based on the evaluation results of the above five modules, it quantifies and scores the results room by room in the current scene. The quantitative scoring result of the single-room quality is shown in formula (1), and the quantitative scoring result of the whole-scene quality is shown in formula (2).
[0158] R(room) = w 1 *(1 - C_Ratio)+w 2 *(1 - W_Ratio)+w 3 *
[0159] N_Ratio+w 4 *(1 - A_Ratio)+w 5 *FID_Ratio+w 6 *CGH_Ratio+w 7 *
[0160] F_Ratio+w 8 *G_Ratio+w 9 *Sigma_Ratio(1)
[0161]
[0162] In formula (1), R(room) represents the quantization evaluation result of the transformed three - dimensional scene, and w i i ∈ [1, 9] is the weight adjustment coefficient. The weight adjustment coefficient is, for example, {9, 9, 9, 9, 8, 8, 16, 16, 16}. S(scene) is the quantization evaluation result of the combined scene obtained by combining multiple rooms, where N is the number of rooms in the combined scene.
[0163] A method for generating an indoor three - dimensional scene provided by an embodiment of the present application determines the three - dimensional space information and functional information of a target room according to room information, where the room information includes room type and the number of rooms; then determines the initial three - dimensional scene of the target room according to the three - dimensional space information, functional information, and a preset object data set, where the initial three - dimensional scene includes object models and the object positions of each object model; then performs a transformation process on the three - dimensional scene according to the replacement object models of each object model in the initial three - dimensional scene to obtain the transformed three - dimensional scene of the target room; and finally evaluates and filters each transformed three - dimensional scene of the target room according to a preset evaluation dimension to obtain the target three - dimensional scene of the target room, improving the quality of the generated scene.
[0164] Figure 20 This is a schematic structural diagram of a device for generating an indoor three - dimensional scene provided by an embodiment of the present application, as Figure 20 shown. The device 200 includes:
[0165] A room determination module 201, configured to determine the three - dimensional space information and functional information of a target room according to room information, where the room information includes room type and the number of rooms;
[0166] The initial 3D scene module 202 is used to determine the initial 3D scene of the target room according to the 3D space information, function information, and a preset object dataset. The initial 3D scene includes object models and the object positions of each object model;
[0167] The transformed 3D scene module 203 is used to perform transformation processing on the 3D scene according to the replacement object models of each object model in the initial 3D scene to obtain the transformed 3D scene of the target room;
[0168] The evaluation and screening module 204 is used to evaluate and screen each transformed 3D scene of the target room according to a preset evaluation dimension to obtain the target 3D scene of the target room.
[0169] Among them, in some embodiments, the evaluation and screening module 204 is further used to:
[0170] For each transformed 3D scene of the target room, evaluate the transformed 3D scene according to the physical constraint dimension to obtain a first evaluation result. The physical constraint dimension includes collision constraint, floor projection, wall penetration constraint, wall attachment constraint, and passage constraint;
[0171] Evaluate the transformed 3D scene according to the FID metric and color histogram to obtain a second evaluation result;
[0172] Evaluate the transformed 3D scene according to the occurrence times of each object in the transformed 3D scene and the occurrence rate of each object in a preset 3D scene dataset to obtain a third evaluation result;
[0173] Evaluate the transformed 3D scene according to the two-dimensional projection of the forward positive direction vector of each object in the transformed 3D scene on the floor plane to obtain a fourth evaluation result;
[0174] Evaluate the transformed 3D scene according to the evaluation 3D scenes similar to the transformed 3D scene in the 3D scene dataset to obtain a fifth evaluation result;
[0175] According to the first evaluation result, second evaluation result, third evaluation result, fourth evaluation result, and fifth evaluation result of each transformed 3D scene, screen each transformed 3D scene of the target room to obtain the target 3D scene of the target room.
[0176] Among them, in some embodiments, the room determination module 201 is further used to:
[0177] Determine the target empty scene corresponding to the room information from the empty scene dataset according to the room information;
[0178] Label the type of the target room in the target empty scene according to the target empty scene to determine the function information of the target room;
[0179] Determine the three-dimensional spatial information of the target room according to the size information of the target empty scene.
[0180] Among them, in some embodiments, the room determination module 201 is further configured to:
[0181] Discretize the room information to obtain a first high-dimensional vector;
[0182] Construct a first KDTree according to the first high-dimensional vector and the second high-dimensional vectors of each empty scene in the empty scene dataset;
[0183] Determine the target second high-dimensional vector adjacent to the first high-dimensional vector in the first KDTree according to the first KDTree;
[0184] Determine the target empty scene corresponding to the target second high-dimensional vector according to the target second high-dimensional vector.
[0185] Among them, in some embodiments, the initial three-dimensional scene module 202 is further configured to:
[0186] Determine the room configuration corresponding to the function information from the three-dimensional scene dataset according to the function information, where the room configuration includes the object type and the number of objects of the object;
[0187] Obtain the object model combination corresponding to the room configuration from the three-dimensional scene dataset according to the room configuration;
[0188] Obtain the initial three-dimensional scene of the target room according to the object model combination.
[0189] Among them, in some embodiments, the initial three-dimensional scene module 202 is further configured to:
[0190] Extract the filtered three-dimensional scenes that meet the function information from the three-dimensional scene dataset according to the function information;
[0191] For each filtered three-dimensional scene, count the object type and the number of objects of each object in the filtered three-dimensional scene to obtain a statistical result;
[0192] According to the statistical results of each filtered three-dimensional scene, use the object types and the number of objects whose appearance rate meets the preset appearance rate as the room configuration corresponding to the function information.
[0193] Among them, in some embodiments, the initial three-dimensional scene module 202 is further configured to:
[0194] Discretize the room configuration to obtain a third high-dimensional vector;
[0195] Construct a second KDTree based on the third high-dimensional vector and the fourth high-dimensional vector of the same type of 3D scenes in the 3D scene dataset, where the same type of 3D scenes are 3D scenes with the same room type as the target room;
[0196] Based on the second KDTree, determine the target fourth high-dimensional vector adjacent to the third high-dimensional vector in the second KDTree;
[0197] Based on the target fourth high-dimensional vector, determine the target 3D scene of the same type corresponding to the target fourth high-dimensional vector;
[0198] Obtain the object model combination corresponding to the room configuration from the target 3D scenes of the same type.
[0199] Among them, in some embodiments, the 3D scene transformation module 203 is further configured to:
[0200] For each object model, determine the size information and function information of the object model;
[0201] Based on the size information and function information, determine the replacement object model of the object model from the object dataset;
[0202] Based on the replacement object model of each object model, perform transformation processing on the object models in the initial 3D scene to obtain multiple transformed 3D scenes of the target room.
[0203] Figure 21 This is a schematic structural diagram of the electronic device provided by the embodiments of the present application. As Figure 21 shown, the electronic device 210 includes:
[0204] The electronic device 210 may include a processor 211 with one or more processing cores, a memory 212 with one or more computer-readable storage media, a communication component 213, and other components. Among them, the processor 211, the memory 212, and the communication component 213 are connected through a bus 214.
[0205] In a specific implementation process, at least one processor 211 executes the computer execution instructions stored in the memory 212, so that at least one processor 211 executes the above method for generating an indoor 3D scene.
[0206] The specific implementation process of the processor 211 can refer to the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0207] In the above Figure 21In the illustrated embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.
[0208] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0209] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0210] In some embodiments, a computer program product is also proposed, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps in any of the above-mentioned methods for generating an indoor three-dimensional scene are implemented.
[0211] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated here.
[0212] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0213] Therefore, an embodiment of the present application provides a computer-readable storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any of the methods for generating an indoor three-dimensional scene provided by the embodiments of the present application.
[0214] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM), magnetic disk, optical disc, etc.
[0215] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium.
[0216] Since the instructions stored in the storage medium can execute the steps in any of the indoor three-dimensional scene generation methods provided by the embodiments of the present application, the beneficial effects achievable by any of the indoor three-dimensional scene generation methods provided by the embodiments of the present application can be realized. For details, please refer to the previous embodiments and will not be elaborated here.
[0217] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0218] It should be understood that the present application is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A method for generating an indoor three-dimensional scene, characterized in that: include: Determine the three-dimensional space information and functional information of the target room according to the room information, wherein the room information includes the room type and the number of rooms; Determine an initial three-dimensional scene of the target room according to the three-dimensional space information, the functional information and a preset object data set, wherein the initial three-dimensional scene includes object models and an object position of each object model; According to the replacement object model of each object model in the initial three-dimensional scene, transforming the three-dimensional scene to obtain a transformed three-dimensional scene of the target room; According to a preset evaluation dimension, each transformed three-dimensional scene of the target room is evaluated and screened to obtain a target three-dimensional scene of the target room.
2. The method according to claim 1, characterized in that The step of evaluating and screening each transformed three-dimensional scene of the target room according to a preset evaluation dimension to obtain a target three-dimensional scene of the target room includes: For each transformed three-dimensional scene of the target room, the transformed three-dimensional scene is evaluated according to a physical constraint dimension to obtain a first evaluation result, where the physical constraint dimension includes a collision constraint, a floor projection, a wall penetration constraint, a wall-to-wall constraint, and a passage constraint; Evaluating the transformed three-dimensional scene according to the FID index and the color histogram to obtain a second evaluation result; evaluating the transformed three-dimensional scene according to the number of occurrences of each object in the transformed three-dimensional scene and the appearance rate of each object in a preset three-dimensional scene data set to obtain a third evaluation result; evaluating the transformed three-dimensional scene according to a two-dimensional projection of a forward positive direction vector of each object in the transformed three-dimensional scene on a floor plane to obtain a fourth evaluation result; evaluating the transformed three-dimensional scene according to the evaluation three-dimensional scenes in the three-dimensional scene data set that are similar to the transformed three-dimensional scene, to obtain a fifth evaluation result; Each transformed three-dimensional scene of the target room is screened according to the first evaluation result, the second evaluation result, the third evaluation result, the fourth evaluation result and the fifth evaluation result of each transformed three-dimensional scene to obtain a target three-dimensional scene of the target room.
3. The method according to any one of claims 1-2, characterized in that: Determining the three-dimensional space information and functional information of the target room according to the room information includes: According to the room information, determining a target empty scene corresponding to the room information from an empty scene dataset; According to the target empty scene, marking the type of the target room in the target empty scene, and determining the functional information of the target room; The three-dimensional space information of the target room is determined according to the size information of the target empty scene.
4. The method according to claim 3, characterized in that The step of determining, according to the room information, a target empty scene corresponding to the room information from an empty scene dataset comprises: Discretize the room information to obtain a first high-dimensional vector; Constructing a first KDTree according to the first high-dimensional vector and a second high-dimensional vector of each empty scene in the empty scene dataset; Determine, according to the first KDTree, a target second high-dimensional vector adjacent to the first high-dimensional vector in the first KDTree; According to the target second high-dimensional vector, a target empty scene corresponding to the target second high-dimensional vector is determined.
5. The method according to any one of claims 1-2, characterized in that: The determining the initial three-dimensional scene of the target room according to the three-dimensional space information, the function information and the preset object data set includes: According to the functional information, determining a room configuration corresponding to the functional information from a three-dimensional scene data set, wherein the room configuration includes an object type and an object quantity of an object; According to the room configuration, acquiring a combination of object models corresponding to the room configuration from the three-dimensional scene data set; According to the object model combination, an initial three-dimensional scene of the target room is obtained.
6. The method according to claim 5, characterized in that The step of determining, according to the functional information, a room configuration corresponding to the functional information from a three-dimensional scene data set includes: According to the functional information, extracting a screening three-dimensional scene that meets the functional information from the three-dimensional scene data set; For each screened three-dimensional scene, counting the object type and the object quantity of each object in the screened three-dimensional scene to obtain a statistical result; According to the statistical result of each filtered three-dimensional scene, the type and quantity of objects whose appearance rate meets the preset appearance rate are used as the room configuration corresponding to the functional information.
7. The method according to claim 5, characterized in that The acquiring, according to the room configuration, from the object data set a combination of object models corresponding to the room configuration, comprises: Discretizing the room configuration to obtain a third high-dimensional vector; constructing a second KDTree according to the third high-dimensional vector and a fourth high-dimensional vector of a similar three-dimensional scene in the three-dimensional scene dataset, wherein the similar three-dimensional scene is a three-dimensional scene of the same room type as the target room; Determine, according to the second KDTree, a target fourth high-dimensional vector adjacent to the third high-dimensional vector in the second KDTree; Determining, according to the target fourth high-dimensional vector, a target similar three-dimensional scene corresponding to the target fourth high-dimensional vector; A combination of object models corresponding to the room configuration is obtained from the target homogeneous three-dimensional scene.
8. The method according to any one of claims 1-2, characterized in that: The step of transforming the three-dimensional scene according to the replacement object model of each object model in the initial three-dimensional scene to obtain the transformed three-dimensional scene of the target room includes: For each object model, determining size information and function information of the object model; determining a replacement object model of the object model from the object data set according to the size information and the function information; According to the replacement object model of each object model, the object model in the initial three-dimensional scene is transformed to obtain a plurality of transformed three-dimensional scenes of the target room.
9. A device for generating an indoor three-dimensional scene, characterized in that: include: A room determination module, used to determine the three-dimensional space information and functional information of the target room according to the room information, wherein the room information includes the room type and the number of rooms; An initial three-dimensional scene module, used to determine an initial three-dimensional scene of the target room according to the three-dimensional space information, the functional information and a preset object data set, wherein the initial three-dimensional scene includes object models and an object position of each object model; A three-dimensional scene transformation module, configured to transform the three-dimensional scene according to a replacement object model of each object model in the initial three-dimensional scene to obtain a transformed three-dimensional scene of the target room; The evaluation and screening module is used to evaluate and screen each transformed three-dimensional scene of the target room according to a preset evaluation dimension to obtain a target three-dimensional scene of the target room.
10. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.