A method and device for acquiring indoor data sets based on real-time renderer
By converting the three-dimensional indoor scene data into USD format and generating designated renderings using real-time renderers, the problem of time-consuming and costly generation of synthetic data sets in the prior art is solved, and the effect of quickly obtaining high-quality training data sets is achieved.
Patent Information
- Application Number
- CN202210486435.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-06
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-05-06
AI Technical Summary
The prior art is costly and time-consuming when generating synthetic data sets, and it is difficult to quickly obtain large-scale data sets in offline rendering.
Using a real-time renderer-based method, the three-dimensional indoor scene data is converted into USD format, labels and camera track information are added, and designated renderings are generated and post-processed to obtain the indoor data set.
This greatly accelerates the generation speed of synthetic data sets, ensures rendering effect, reduces costs, and can quickly obtain high-quality training data sets.
Smart Images

Figure CN114926574B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of synthetic image data acquisition, and in particular relates to a method and device for acquiring an indoor data set based on a real-time renderer. Background Art
[0002] With the rapid development of deep learning, data has played a vital role in machine learning, especially in computer vision tasks. However, creating datasets requires significant manual annotation costs. Furthermore, most datasets are designed for a single task, and the accuracy of annotations is difficult to guarantee. Large-scale synthetic datasets, however, offer advantages such as low cost, high accuracy, and a wide range of applications, significantly boosting the data-driven field of machine learning. Currently, most methods for obtaining synthetic datasets rely on offline rendering, which is costly and time-consuming. Rapidly acquiring large datasets requires significant computational resources and requires very long synthesis times, with synthesis of a single photo taking minutes.
[0003] Generating 3D synthetic datasets significantly improves the training process of machine learning. Classic synthetic datasets such as SUNCG, InteriorNet, and House3D provide rich data for network training. Kujiale's Minverse system also demonstrates the improved training results achieved with synthetic datasets. Existing methods for generating 3D synthetic datasets, such as Chinese patent application publication number CN112950760A, disclose a system and method for generating 3D scene data. However, these methods are all offline, requiring defined rules and lengthy offline rendering, resulting in high computational costs and time consumption. Summary of the Invention
[0004] One of the objectives of the present invention is to provide an indoor dataset acquisition method based on a real-time renderer, which has the advantages of short rendering time, unimpaired rendering effect, and low cost, helps to speed up the acquisition of synthetic datasets, and greatly promotes the rapid acquisition of training datasets.
[0005] To achieve the above object, the technical solution adopted by the present invention is:
[0006] A method for acquiring an indoor data set based on a real-time renderer, the method comprising:
[0007] Convert 3D indoor scene data into USD format data;
[0008] Add labels to each object in the 3D indoor scene based on USD format data;
[0009] Add a predefined camera to the 3D indoor scene and obtain the camera's trajectory information in the 3D indoor scene;
[0010] According to the trajectory information of the camera in the 3D indoor scene, a real-time renderer is used to obtain the rendering image with specified rendering attributes under the camera's perspective in real time, and the rendering image is post-processed to obtain the indoor dataset.
[0011] Several optional methods are also provided below, but they are not intended to be additional limitations on the above-mentioned overall solution. They are merely further supplements or optimizations. Under the premise that there are no technical or logical contradictions, each optional method can be combined separately for the above-mentioned overall solution, or multiple optional methods can be combined.
[0012] Preferably, the converting of the three-dimensional indoor scene data into USD format data includes:
[0013] Convert the 3D indoor scene data into the Uasset scene description format commonly used by the Unreal Engine engine, and then convert the Uasset scene description format data into standard USD format data through the Unreal Engine engine. The format conversion objects include scenes, materials, and lights.
[0014] Preferably, the real name of each object is used as the label.
[0015] Preferably, the predefined information of the camera includes: name, focal length, resolution, imaging type, rotation parameters and translation parameters.
[0016] Preferably, the trajectory information includes a plurality of preset trajectory points, and each trajectory point includes the following information: the name of the camera, the position of the camera in the three-dimensional indoor scene, and the direction of the camera.
[0017] Preferably, the rendering images of the specified rendering attributes include a normal map, a depth map, an instance map, a segmentation map, a color map RGB, a two-dimensional box Box2d of the object, and a three-dimensional box Box3d of the object.
[0018] Preferably, the indoor data set is verified for correctness using a visualization program.
[0019] The indoor dataset acquisition method based on a real-time renderer provided by the present invention converts three-dimensional indoor scene data into the universal scene description format USD (Universe Scene Description) format, uses a real-time renderer for real-time rendering, and then adds camera and trajectory route synthesis depth map, color map, instance instance map, semantic segmentation map, object Box2D map, object Box3D map and camera information, greatly accelerating the synthesis speed of the synthetic dataset.
[0020] The second purpose of the present invention is to provide an indoor dataset acquisition device based on a real-time renderer, which has the advantages of short rendering time, unimpaired rendering effect, and low cost, helps to speed up the acquisition of synthetic datasets, and greatly promotes the rapid acquisition of training datasets.
[0021] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is: an indoor dataset acquisition device based on a real-time renderer, comprising a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, implement the steps of the indoor dataset acquisition method based on a real-time renderer. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 Flowchart of the indoor data set acquisition method based on real-time renderer of the present invention;
[0023] Figure 2 A schematic diagram of an embodiment of the color image RGB of the present invention;
[0024] Figure 3 For the present invention and Figure 2 A schematic diagram of an embodiment of a two-dimensional box Box2d for objects in the same three-dimensional indoor scene and viewing angle;
[0025] Figure 4 For the present invention and Figure 2 A schematic diagram of an embodiment of a three-dimensional box Box3d for objects in the same three-dimensional indoor scene and viewing angle;
[0026] Figure 5 For the present invention and Figure 2 A schematic diagram of an embodiment of a depth map for the same three-dimensional indoor scene and viewing angle;
[0027] Figure 6 For the present invention and Figure 2 A schematic diagram of an embodiment of an instance graph Instance under the same three-dimensional indoor scene and perspective;
[0028] Figure 7 For the present invention and Figure 2A schematic diagram of an embodiment of a segmentation map Semantic for the same 3D indoor scene and viewing angle;
[0029] Figure 8 For the present invention and Figure 2 A schematic diagram of an embodiment of a normal map under the same three-dimensional indoor scene and viewing angle. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0032] Currently, real-time rendering engines are constantly developing and improving, such as Unreal Engine, Unity, CRYENGINE, KoolEngine, and the real-time rendering and simulation engine Isaac Sim. Their rendering effects are already very realistic and can meet the requirements for dataset acquisition. Therefore, this embodiment proposes a method for acquiring massive indoor datasets based on a real-time rendering engine. The method greatly reduces the time consumption and can achieve the same effect as offline rendering. It also greatly accelerates the speed and quantity of obtaining synthetic datasets.
[0033] like Figure 1 As shown, the indoor dataset acquisition method based on the real-time renderer of this embodiment converts the three-dimensional indoor scene data into the universal scene description format USD data, adds the physical properties of the scene and the label of each object in the scene based on the USD scene, and then renders it through the real-time renderer, and verifies the rendered result through a separate display program. The results show that this embodiment can greatly speed up the generation of synthetic datasets and ensure that there is no loss of effect.
[0034] Specifically, the indoor dataset acquisition method based on the real-time renderer in this embodiment includes the following steps:
[0035] 1) Convert 3D indoor scene data into USD format data.
[0036] USD (Universal Scene Description) is a highly extensible and open-source 3D scene description format developed by Pixar. Its powerful functions are widely used in architecture, design, rendering and other fields, and most well-known 3D modeling software supports the direct import of USD 3D scenes.
[0037] This embodiment uses an online home design platform to acquire a large amount of 3D indoor scene data, reducing the difficulty of acquiring original data. For example, the Cool Home online home design platform has accumulated a large amount of 3D indoor scene data, and the 3D indoor scene data is in a standard middle platform format.
[0038] When performing format conversion in this embodiment, the three-dimensional indoor scene data is first converted into data in the Uasset scene description format commonly used by the Unreal Engine engine, and then the Uasset scene description format data is converted into standard USD format data to form a USD file through the Unreal Engine engine. The format conversion objects include scenes, materials, and lights.
[0039] The final generated USD format scene file is as follows:
[0040] Materials: usd material system. The value of this parameter is used to describe the material information, including Textures and Mdl.
[0041] Textures: The value of this parameter contains all the texture image information of the scene;
[0042] Mdl: The value of this parameter is used to describe the MDL (Material Definition Language) material information.
[0043] Props: The basic unit of the scene, the basic properties of each object in the scene, mainly including geometric and physical properties.
[0044] Start.usd: The entry file of the entire scene, containing global information such as perspective, lighting, camera, etc.
[0045] After the standard universal scene description format USD is formed, it can be opened and viewed through many large 3D software such as Isaac Sim, Maya, Blender, etc.
[0046] 2) Add labels to each object in the 3D indoor scene based on USD format data.
[0047] To subsequently obtain the Semantic segmentation map and Instance map, it is necessary to add unique labels to each object in the 3D indoor scene USD file. The objects here are each object (including soft furnishings and hard furnishings) as well as walls, ceilings, and floors in the 3D indoor scene. In this embodiment, the real names of each object are added as labels. For example, the labels "ceilings," "walls," and "floors" are added to the roof, floor, and walls, respectively.
[0048] When adding tags to each object, a tag mapping table can be preset, and the corresponding object tag number can be queried by id, and the corresponding tag can be found by object tag number. In actual operation, each object can be tagged directly based on the tag system of the existing online home design platform.
[0049] For example, using the tagging system of the Cool Home online home design platform, you can query the corresponding object tag number using the Mesh ID, and then use the number to find the corresponding tag mapping table. USD files can be read and modified using the Pxr open source library. Its file format uses Prim as a node to connect all object information, and uses the tagging system to assign corresponding tags to each piece of furniture.
[0050] 3) Add a predefined camera to the 3D indoor scene and obtain the trajectory information of the camera in the 3D indoor scene.
[0051] Synthetic datasets require obtaining continuous image information in the scene, so it is necessary to customize the corresponding camera at any position in the scene and add camera trajectory information for continuous rendering.
[0052] 31) Add predefined cameras.
[0053] USD files are scene file description formats that can be modified in real-time programming. By reading the current scene and obtaining the prim of the root node, a predefined camera is added to the 3D indoor scene for shooting. In this embodiment, the information for adding a predefined camera mainly includes:
[0054] Name: For example, / world / camera.
[0055] Focal Length: Used to describe the focal length of the camera.
[0056] Resolution (image size): for example 1024X800, can be customized.
[0057] Projection type: For example, perspective camera, fisheye camera, or orthographic camera.
[0058] Rotation parameter (Rotatezyx): used to describe the rotation of the camera.
[0059] Translation parameters (Translate): used to describe the translation of the camera.
[0060] In other embodiments, there are other optional parameters such as near and far planes, fStop, scale, etc., which can be customized according to actual needs. One or more cameras can be predefined in advance, and the specified camera can be added to the 3D indoor scene when acquiring the data set.
[0061] 32) Add the trajectory information corresponding to the camera.
[0062] The camera trajectory mainly includes how the camera continuously shoots in indoor scenes. Its motion information mainly includes the camera's position and orientation. The position information is represented by the three coordinates X, Y, and Z in the world coordinate system, and its orientation is determined by lookat.
[0063] The camera's trajectory information is set by pre-adding multiple trajectory points in the USD file. Each trajectory point is described in the USD file by the MovementComponent keyword. The name of the camera is specified by adding prim_path; the position of the camera in the three-dimensional indoor scene is specified by adding target_points, and the position information is represented by the X, Y, and Z coordinates in the world coordinate system; the lookat direction of the camera is specified by adding lookat_target_points.
[0064] The same camera can have one or more trajectory information. During data collection, the specific trajectory information for this data collection is obtained, and multiple track points pre-set for this trajectory information are added to the USD file. Subsequently, the track points are sequentially executed to cause the corresponding camera to move. It should be noted that in this embodiment, the track points can be manually set or obtained by discretizing the planned path based on the path planning algorithm based on the 3D indoor scene. This is not a limitation in this embodiment.
[0065] 4) According to the trajectory information of the camera in the three-dimensional indoor scene, a real-time renderer is used to obtain a rendering image with specified rendering attributes under the camera's perspective in real time, and the rendering image is post-processed to obtain an indoor dataset.
[0066] 41) The real-time renderer renders the scene.
[0067] USD files can be opened and operated by many real-time rendering software. This article conducts experiments based on Isaac Sim, which supports real-time ray tracing algorithms, MDL materials, and physical-based rendering. The real-time rendering effect is very good and meets the requirements of synthetic data.
[0068] For example, if you use Isaac Sim 2021.1.1 to directly open the USD file and run it, you can browse and verify the scene data by rotating the various perspectives. The current real-time renderer effect is already very good. The test shows that it takes less than 200ms to render a very realistic 1024X800 picture, which is about 50 to 100 times faster than offline rendering. At the same time, since the three-dimensional indoor scene is known, you can customize the rendering properties to render the three-dimensional indoor scene, obtain a rendering image, and post-process the rendering image based on the post-processing instructions to obtain a high-quality data set that meets the requirements.
[0069] 42) Acquisition and preservation of indoor datasets.
[0070] After defining the label system, camera, and camera trajectory information, a real-time renderer such as Isaac Sim can obtain a rendering image with specified rendering attributes from the camera's perspective in real time. Define the type of rendering image to be obtained. In this embodiment, the normal map, depth map, instance map, segmentation map, RGB color map, two-dimensional box Box2d of the object, and three-dimensional box Box3d of the object are selected. The camera information of the camera at the corresponding moment (i.e., the camera information in step 31) is recorded in the output data. Each type of data corresponds to a folder, and the data is saved as .npy data. The color map is saved as a four-channel png image. The output format is as follows:
[0071] Viewport1
[0072] Semantic:0.npy 1.npy....100.npy
[0073] Instance:0.npy 1.npy....100.npy
[0074] Depth:0.npy 1.npy....100.npy
[0075] Rgb:0.png 1.png....100.png ...
[0077] Camera:0.npy 1.npy....100.npy
[0078] The saved Python.npy data stores data of the corresponding data type in the form of a structured array or a NumPy array. For example, the information contained in Box3d is as follows: dtype([('uniqueId','<i4'),('name','O'),(semanticLabel,'O'),('metadata','O'),('instanceIds','O'),('semanticId','<u4'),('x_min','<f4'),('y_min','<f4'),('z_min','<f4'),('x_max','<f4'),('y_max','<f4'),('z_max','<f4'),('transform','<f4',(4,4)),('corners','<f4',(8,3))]). The saving results of other data types are similar. The saved data is guaranteed to be readable by a visualization program and the results of this type of data can be fully displayed, enabling quick verification of data correctness.
[0079] To ensure the accuracy of the generated dataset, an additional visualization program is required to verify the correctness of the generated data. In this paper, the Python program reads the.npy data of each type of data and converts the data into the form of a visual image for display. The two-dimensional bounding box information of the object can be displayed as a line in the color image. The segmentation map is visualized based on different or the same colors of the objects, etc. The depth map is converted into pixel values from 0 to 255, in the form of nearly white and far black, etc. Finally, seven types of synthetic data can be fully visualized. This method can quickly verify the correctness of the generated data. The visualization images of the seven types of synthetic data are as Figures 2 to 8 shown.
[0080] The seven types of data and the corresponding camera parameters in this embodiment are all synthetically obtained, and their accuracy has a natural advantage over traditional manual annotation, and basically cover the data types required for current image-related deep learning. The generated data can be formatted to be the same as that of traditional well-known datasets, such as ImageNet, COCO, etc. Their quantity can be generated on a large scale, and the quality has also been greatly improved compared to traditional datasets.
[0081] The indoor data set acquisition method based on the real-time renderer provided in this embodiment obtains high-quality three-dimensional indoor scene data with rich simulation scene details based on the online home decoration design platform. It can restore the real scene and has universality after being converted into USD format. The entire conversion process is fully automatic, with a rendering speed of about 200ms (for 1024X800 images), which is 50 to 100 times faster than offline rendering, and can quickly obtain large-scale training data sets including Ground Truth. It is highly scalable, and the universal scene description format USD is not limited to the Isaac Sim renderer, and other real-time renderers still support it, such as Kool Engine. The visualization program can conveniently verify the correctness of the results in real time.
[0082] In another embodiment, the present application also provides an indoor dataset acquisition device based on a real-time renderer, comprising a processor and a memory storing a plurality of computer instructions, wherein the computer instructions, when executed by the processor, implement the steps of the indoor dataset acquisition method based on a real-time renderer.
[0083] The specific definition of the indoor dataset acquisition device based on the real-time renderer can be found in the definition of the indoor dataset acquisition method based on the real-time renderer above, which will not be repeated here.
[0084] The memory and processor are electrically connected, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected via one or more communication buses or signal lines. The memory stores a computer program executable on the processor. The processor executes the computer program stored in the memory to implement the indoor dataset acquisition method based on a real-time renderer according to an embodiment of the present invention.
[0085] The memory may be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. The memory is used to store a program, and the processor executes the program after receiving an execution instruction.
[0086] The processor may be an integrated circuit chip with data processing capabilities. The processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor.
[0087] It should be noted that Figures 2 to 8 The following are schematic diagrams of an embodiment for displaying the color image RGB, the two-dimensional box Box2d of the object, the three-dimensional box Box3d of the object, the depth map, the instance map Instance, the segmentation map Semantic, and the normal map obtained in this embodiment. The graphics and characters in the figure are only elements of the running interface when the software is running, and do not involve the focus of the improvement of this application. The clarity of the running interface is related to the pixels and the scaling ratio, so the presentation effect is relatively limited. Due to the requirement of removing the color information of the above images (except the depth map) for display, the actual software running output has obvious color distinction.
[0088] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0089] The above-described embodiments merely illustrate several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for acquiring indoor data sets based on a real-time renderer, characterized in that: The indoor data set acquisition method based on a real-time renderer includes: Converting 3D indoor scene data into USD format data, including: converting the 3D indoor scene data into data in the Uasset scene description format commonly used by the Unreal Engine engine, and then converting the Uasset scene description format data into standard USD format data through the Unreal Engine engine, and the format conversion objects include scenes, materials and lights; Add labels to each object in the 3D indoor scene based on USD format data, and use the real name of each object as the label; Add a predefined camera to the 3D indoor scene and obtain the camera's trajectory information in the 3D indoor scene; According to the trajectory information of the camera in the 3D indoor scene, a real-time renderer is used to obtain the rendering image with specified rendering attributes under the camera's perspective in real time, and the rendering image is post-processed to obtain the indoor dataset.
2. The indoor dataset acquisition method based on real-time renderer according to claim 1, characterized in that: The predefined information of the camera includes: name, focal length, resolution, imaging type, rotation parameters and translation parameters.
3. The indoor dataset acquisition method based on real-time renderer according to claim 1, characterized in that: The trajectory information includes a plurality of preset trajectory points, and each trajectory point includes the following content: the name of the camera, the position of the camera in the three-dimensional indoor scene, and the direction of the camera.
4. The indoor dataset acquisition method based on real-time renderer according to claim 1, characterized in that: The rendering images of the specified rendering attributes include a normal map, a depth map, an instance map, a segmentation map, a color map RGB, a two-dimensional box Box2d of the object, and a three-dimensional box Box3d of the object.
5. The indoor dataset acquisition method based on real-time renderer according to claim 1, characterized in that: The indoor dataset is verified to be correct using a visualization program.
6. A device for acquiring indoor data sets based on a real-time renderer, comprising a processor and a memory storing a plurality of computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the indoor data set acquisition method based on a real-time renderer according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
USD-based 3D software efficient hardware rendering preview method
CN111383306A
Three-dimensional synthetic scene data generation system and method
CN112950760A