A scene reconstruction method, device, equipment, storage medium and product
By using a hierarchical unified Gaussian meta-method to initialize and optimize images, the problem of simulating complex dynamic objects in large-scale dynamic scenes is solved, achieving highly accurate scene reconstruction and real-time rendering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CO WHEELS TECH CO LTD
- Filing Date
- 2024-11-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to effectively simulate complex dynamic objects in large-scale dynamic scenes, leading to a decrease in the accuracy of scene reconstruction.
The Hierarchy Unified Gaussian Particle (HGP) method is adopted to establish a target spatial scene by initializing and optimizing the sequence of images to be processed. This includes scene division, extraction of device posture information and image information, generation of point cloud information of multi-dimensional scene structure, addition of it to spatial blocks, and merging scene elements to construct the target spatial scene.
It improves the accuracy of scene reconstruction, effectively simulates complex dynamic and static objects in large-scale dynamic scenes, and achieves high-fidelity large-scale dynamic scene reconstruction and real-time rendering.
Smart Images

Figure CN122115700A_ABST
Abstract
Description
Technical Field
[0001] This application relates to image processing technology, and more particularly to a scene reconstruction method, apparatus, device, storage medium, and product. Background Technology
[0002] With the development of electronic technology, more and more researchers are focusing on scene reconstruction technology. Scene reconstruction technology can reconstruct two-dimensional image information into corresponding three-dimensional scenes, which is of great significance for applications in fields such as autonomous driving, virtual reality, and smart cities.
[0003] In related technologies, 3DGS can achieve scene reconstruction in large-scale scenarios by using point-based differentiable rendering techniques, or city-scale reconstruction using block Gaussian, and large-scale scene reconstruction and rendering using hierarchical tree Gaussian primitives. However, these methods cannot effectively simulate complex dynamic objects in large-scale dynamic scenes. When scene reconstruction of dynamic objects is required, the accuracy of scene reconstruction using existing technologies is reduced. Summary of the Invention
[0004] This application provides a scene reconstruction method, apparatus, device, storage medium, and product, which can improve the accuracy of scene reconstruction.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] This application provides a scene reconstruction method, the method comprising:
[0007] Obtain the sequence of images to be processed;
[0008] The target spatial scene corresponding to the sequence of images to be processed is established using hierarchical unified Gaussian primitives.
[0009] In the above scheme, the step of establishing the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian units includes:
[0010] The sequence of images to be processed is initialized to obtain the scene frame of the sequence of images to be processed;
[0011] The sequence of images to be processed is optimized to obtain multiple sets of scene elements in the sequence of images to be processed.
[0012] Based on the scene framework and the multiple sets of scene elements, a target spatial scene corresponding to the sequence of images to be processed is established.
[0013] In the above scheme, the optimization processing of the sequence of images to be processed to obtain multiple sets of scene elements in the sequence of images to be processed includes:
[0014] The sequence of images to be processed is divided into scenes to obtain multiple sub-scenes;
[0015] Scene elements are extracted from the multiple sub-scenes to obtain the multiple sets of scene elements.
[0016] In the above scheme, the step of dividing the sequence of images to be processed into multiple sub-scenes includes:
[0017] The sequence of images to be processed is divided into multiple spatial blocks according to a preset spatial size;
[0018] Extract device pose information from the sequence of images to be processed; the device pose information is the pose information of the acquisition device when acquiring the sequence of images to be processed.
[0019] Obtain the image information corresponding to the sequence of images to be processed;
[0020] Based on the multiple spatial blocks, the device posture information, and the image information, the multiple sub-scenes are determined, and the multiple spatial blocks correspond one-to-one with the multiple scenes.
[0021] In the above scheme, determining the multiple sub-scenes based on the multiple spatial blocks, the device posture information, and the image information includes:
[0022] Generate point cloud information of a multi-dimensional scene structure based on the device posture information and the image information;
[0023] Based on the point cloud information, the device posture information, and the relationship between the multiple spatial blocks, the point cloud information and the sequence of images to be processed are added to the multiple spatial blocks to obtain the multiple sub-scenes.
[0024] In the above scheme, the step of adding the point cloud information and the sequence of images to be processed to the multiple spatial blocks based on the point cloud information, the device pose information, and the relationship between the multiple spatial blocks to obtain the multiple sub-scenes includes:
[0025] The density of the point cloud in the multi-dimensional scene structure is increased in the point cloud information to obtain the increased point cloud information.
[0026] Based on the added point cloud information, the device posture information, and the relationship between the multiple spatial blocks, the added point cloud information and the sequence of images to be processed are added to the multiple spatial blocks to obtain the multiple sub-scenes.
[0027] In the above scheme, the multiple sets of scene elements include at least:
[0028] Sky, background, rigid body objects, and non-rigid body objects are one or more of the following: sky, background, rigid body objects, and non-rigid body objects.
[0029] In the above scheme, establishing the target spatial scene corresponding to the sequence of images to be processed based on the scene framework and the multiple sets of scene elements includes:
[0030] The multiple sets of scene elements are merged into the scene frame according to the relationship between the multiple sub-scenes and the scene frame to obtain the target space scene.
[0031] In the above scheme, the initialization process of the sequence of images to be processed to obtain the scene framework of the sequence of images to be processed includes:
[0032] Establish a multi-layer structure for the sequence of images to be processed;
[0033] The multi-layered structure is used as the scene framework.
[0034] This application provides a scene reconstruction apparatus, including:
[0035] The acquisition unit is used to acquire a sequence of images to be processed.
[0036] A unit is established to construct the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian units.
[0037] In the above scheme, the device further includes a processing unit;
[0038] The processing unit is used to initialize the sequence of images to be processed to obtain the scene framework of the sequence of images to be processed; and to optimize the sequence of images to be processed to obtain multiple sets of scene elements in the sequence of images to be processed.
[0039] The establishment unit is used to establish the target spatial scene corresponding to the sequence of images to be processed based on the scene framework and the multiple sets of scene elements.
[0040] In the above scheme, the device further includes a division unit and an extraction unit;
[0041] The segmentation unit is used to segment the sequence of images to be processed into multiple sub-scenes;
[0042] The extraction unit is used to extract scene elements from the multiple sub-scenes respectively to obtain the multiple sets of scene elements.
[0043] In the above scheme, the device further includes a determining unit;
[0044] The partitioning unit is used to partition the sequence of images to be processed according to a preset spatial size to obtain multiple spatial blocks;
[0045] The extraction unit is used to extract device posture information from the sequence of images to be processed; the device posture information is the posture information of the acquisition device when acquiring the sequence of images to be processed.
[0046] The acquisition unit is used to acquire image information corresponding to the sequence of images to be processed;
[0047] The determining unit is used to determine the multiple sub-scenes based on the multiple spatial blocks, the device posture information, and the image information, wherein the multiple spatial blocks correspond one-to-one with the multiple scenes.
[0048] In the above scheme, the device further includes a generation unit and an addition unit;
[0049] The generation unit is used to generate point cloud information of a multi-dimensional scene structure based on the device posture information and the image information;
[0050] The adding unit is used to add the point cloud information and the sequence of images to be processed to the multiple spatial blocks according to the point cloud information, the device posture information and the relationship between the multiple spatial blocks, so as to obtain the multiple sub-scenes.
[0051] In the above scheme, the device further includes an increasing unit;
[0052] The adding unit is used to increase the density of the point cloud in the multi-dimensional scene structure in the point cloud information to obtain the increased point cloud information.
[0053] The adding unit is used to add the added point cloud information and the sequence of images to be processed to the multiple spatial blocks according to the added point cloud information, the device posture information and the relationship between the multiple spatial blocks, so as to obtain the multiple sub-scenes.
[0054] In the above scheme, the multiple sets of scene elements include at least:
[0055] Sky, background, rigid body objects, and non-rigid body objects are one or more of the following: sky, background, rigid body objects, and non-rigid body objects.
[0056] In the above scheme, the device further includes a merging unit;
[0057] The merging unit is used to merge the multiple sets of scene elements into the scene frame according to the relationship between the multiple sub-scenes and the scene frame, so as to obtain the target space scene.
[0058] In the above scheme, the establishment unit is used to establish a multi-layer structure of the sequence of images to be processed;
[0059] The determining unit is used to use the multi-layer structure as the scene frame.
[0060] This application embodiment further provides a scene reconstruction device, the scene reconstruction device comprising:
[0061] Memory is used to store executable instructions for a computer;
[0062] The processor, when executing computer-executable instructions stored in the memory, implements the scene reconstruction method provided in the embodiments of this application.
[0063] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the scene reconstruction method provided in this application.
[0064] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the scene reconstruction method provided in this application.
[0065] The embodiments of this application have the following beneficial effects: when the scene reconstruction device acquires a sequence of images to be processed, it uses hierarchical unified Gaussian units to establish the target space scene corresponding to the sequence of images to be processed. The hierarchical unified Gaussian units can effectively simulate complex dynamic objects and static objects in large-scale dynamic scenes, thereby improving the accuracy of scene reconstruction. Attached Figure Description
[0066] Figure 1 This application provides a scene reconstruction method flow. Figure 1 ;
[0067] Figure 2 This is an exemplary image scene reconstruction result comparison diagram provided in an embodiment of this application;
[0068] Figure 3 This is a schematic diagram of an exemplary scene reconstruction structure provided in an embodiment of this application;
[0069] Figure 4 This is a comparative diagram of an exemplary dynamic city dataset provided in an embodiment of this application;
[0070] Figure 5 This is a qualitative comparison diagram of an exemplary Waymo dataset provided in an embodiment of this application;
[0071] Figure 6 This is a schematic diagram of an exemplary segmented object ablation training strategy provided in an embodiment of this application;
[0072] Figure 7 This is an exemplary schematic diagram of non-rigid object ablation optimization provided in an embodiment of this application;
[0073] Figure 8 This is a schematic diagram of the composition structure of a scene reconstruction device provided in an embodiment of this application;
[0074] Figure 9 This is a schematic diagram of the composition structure of a scene reconstruction device provided in an embodiment of this application.
[0075] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0076] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0077] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0078] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0079] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0080] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0081] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0082] This application provides a scene reconstruction method, which is applied to a scene reconstruction device. Figure 1 A flowchart of a scene reconstruction method provided in this application embodiment is shown below. Figure 1 As shown, scene reconstruction methods may include:
[0083] S101. Obtain the sequence of images to be processed.
[0084] The scene reconstruction method provided in this application embodiment is applicable to scenarios where an image scene corresponding to the image to be processed is established.
[0085] In the embodiments of this application, the scene reconstruction device can be implemented in various forms. For example, the scene reconstruction device described in this application may include devices such as servers, computers, and cloud computing. The specific scene reconstruction device can be determined according to the actual situation, and the embodiments of this application do not limit it in this regard.
[0086] In this embodiment of the application, the number of images to be processed in the sequence of images to be processed is multiple. The specific number of images to be processed in the sequence of images to be processed can be determined according to the actual situation, and this embodiment of the application does not limit this.
[0087] In this embodiment of the application, the sequence of images to be processed can be images to be processed within a period of time that are continuously acquired, such as images to be processed within 5 minutes.
[0088] It should be noted that a period of time can be 5 minutes, 20 minutes, or other durations. The specific duration of a period of time can be determined according to the actual situation, and this application does not limit it.
[0089] In this embodiment, the scene reconstruction device can acquire images to be processed within a certain period of time from the acquisition device (i.e., image acquisition device), that is, acquire a sequence of images to be processed from the acquisition device; the sequence of images to be processed can also be images acquired by the image acquisition unit in the scene reconstruction device, and the scene reconstruction device can also acquire the sequence of images to be processed from other devices; the specific way in which the scene reconstruction device acquires the sequence of images to be processed can be determined according to the actual situation, and this embodiment does not limit it.
[0090] In this embodiment, the sequence of images to be processed includes dynamic objects. These dynamic objects include moving cars, walking people, or animals, and the specific dynamic objects can be determined based on the actual situation; this embodiment does not limit this.
[0091] In the embodiments of this application, each image in the sequence of images to be processed can be a two-dimensional image.
[0092] S102. Establish the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian primitives.
[0093] In this embodiment of the application, after the scene reconstruction device acquires the sequence of images to be processed, it uses hierarchical unified Gaussian units to establish the target spatial scene corresponding to the sequence of images to be processed.
[0094] In this embodiment of the application, a Hierarchy UGP is used.
[0095] It should be noted that the target space scene can be a three-dimensional space scene, a thought space scene, or a space scene of other dimensions. The specific dimensional information of the target space scene can be determined according to the actual situation, and this application embodiment does not limit it in this way.
[0096] In this embodiment of the application, the process of the scene reconstruction device establishing the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian elements includes: initializing the sequence of images to be processed to obtain a scene frame of the sequence of images to be processed; optimizing the sequence of images to be processed to obtain multiple sets of scene elements in the sequence of images to be processed; and establishing the target spatial scene corresponding to the sequence of images to be processed based on the scene frame and the multiple sets of scene elements.
[0097] In this embodiment of the application, the plurality of scene elements includes at least:
[0098] Sky, background, rigid body objects, and non-rigid body objects are one or more of the following: sky, background, rigid body objects, and non-rigid body objects.
[0099] It should be noted that rigid bodies can be objects like cars, whose size and shape do not change after being subjected to forces or movement. Non-rigid bodies can be objects like people or animals, whose size and shape change after being subjected to forces or movement.
[0100] In this embodiment, the SfM (Structure from Motion) and LiDAR point clouds in a large-scale dynamic scene can be used to initialize the UGP (Unstructured Geometry Processing) model. That is, the SfM and LiDAR point clouds in the large-scale dynamic scene are used to initialize the sequence of images to be processed to obtain the scene frame of the sequence of images to be processed. Other methods can also be used to initialize the sequence of images to be processed to obtain the scene frame of the sequence of images to be processed. The specific method of initializing the sequence of images to obtain the scene frame of the sequence of images to be processed can be determined according to the actual situation, and this embodiment does not limit it.
[0101] In this embodiment, through coarse training, a global UGP framework is constructed, providing powerful global geometric priors for sub-scenes and effectively improving rendering fidelity. Due to the UGP timescale s... t To control its temporal influence range, the non-rigid object's s can be initialized according to the data acquisition frame rate using formula (1). t :
[0102]
[0103] It should be noted that λ is a parameter used for fine-tuning the initial s. t The hyperparameters are: υ is the image acquisition frequency, dt is the time duration corresponding to each frame, and o th It is the opacity threshold used for pruning.
[0104] In this embodiment of the application, the process by which the scene reconstruction device optimizes the sequence of images to be processed to obtain multiple sets of scene elements in the sequence of images to be processed includes: dividing the sequence of images to be processed into scenes to obtain multiple sub-scenes; and extracting scene elements from the multiple sub-scenes to obtain the multiple sets of scene elements.
[0105] In the embodiments of this application, the sky, background, rigid body object or non-rigid body object are extracted in each of the multiple sub-scenes to obtain the scene elements corresponding to each sub-scene, and thus multiple sets of scene elements are obtained for the multiple sub-scenes.
[0106] It should be noted that a sub-scene corresponds to a set of scene elements.
[0107] In this embodiment of the application, the method of extracting multiple sets of scene elements from the multiple sub-scenes can be determined according to the actual situation, and this embodiment of the application does not limit it.
[0108] In this embodiment of the application, the process of the scene reconstruction device dividing the sequence of images to be processed into multiple sub-scenes includes: dividing the sequence of images to be processed into multiple spatial blocks according to a preset spatial size; extracting device posture information from the sequence of images to be processed; obtaining image information corresponding to the sequence of images to be processed; and determining the multiple sub-scenes based on the multiple spatial blocks, the device posture information, and the image information.
[0109] It should be noted that the multiple spatial blocks correspond one-to-one with the multiple scenes, that is, one spatial block corresponds to one scene.
[0110] It should be noted that the device posture information refers to the posture information of the acquisition device when acquiring the sequence of images to be processed.
[0111] In this embodiment, the preset space size can be the space size configured in the scene reconstruction device, the size information transmitted to the scene reconstruction device from other devices, or the size information obtained by the scene reconstruction device through other means. The specific way in which the scene reconstruction device obtains the preset space size can be determined according to the actual situation, and this embodiment does not limit it.
[0112] It should be noted that the size information of the preset space can be determined according to the actual situation, and this application embodiment does not limit it.
[0113] In this embodiment of the application, SfM technology can be used to extract camera pose information (the camera pose information is the device pose information) from a sequence of images to be processed. The camera pose information includes the pose (position and orientation) of each image.
[0114] It should be noted that other methods can also be used to extract device posture information from the sequence of images to be processed. The specific method for extracting device posture information from the sequence of images to be processed can be determined according to the actual situation, and this application embodiment does not limit it.
[0115] In this application embodiment, the method of obtaining the image information corresponding to the sequence of images to be processed can be determined according to the actual situation, and this application embodiment does not limit it.
[0116] In this embodiment of the application, the process of determining the multiple sub-scenes based on the multiple spatial blocks, the device posture information, and the image information includes: generating point cloud information of a multi-dimensional scene structure based on the device posture information and the image information; and adding the point cloud information and the sequence of images to be processed to the multiple spatial blocks according to the relationship between the point cloud information, the device posture information, and the multiple spatial blocks to obtain the multiple sub-scenes.
[0117] It should be noted that the point cloud information of the multi-dimensional scene structure generated based on the device posture information and the image information is sparse point cloud information.
[0118] In this embodiment, the process by which the scene reconstruction device adds the point cloud information and the sequence of images to be processed to the multiple spatial blocks to obtain the multiple sub-scenes, based on the point cloud information, the device posture information, and the relationship between the multiple spatial blocks, includes: increasing the density of the point cloud in the multi-dimensional scene structure to obtain increased point cloud information; and adding the increased point cloud information and the sequence of images to be processed to the multiple spatial blocks to obtain the multiple sub-scenes based on the increased point cloud information, the device posture information, and the relationship between the multiple spatial blocks.
[0119] In this embodiment, the reconstruction quality of non-rigid objects can be improved by increasing the density of the initial point cloud according to execution time and spatial interpolation (i.e., increasing the density of the point cloud in the multi-dimensional scene structure in the point cloud information according to execution time and spatial interpolation to obtain the increased point cloud information); or by increasing the density of the point cloud in the multi-dimensional scene structure in the point cloud information in other ways to obtain the increased point cloud information; specifically, the method of increasing the density of the point cloud in the multi-dimensional scene structure in the point cloud information to obtain the increased point cloud information can be determined according to the actual situation, and this embodiment does not limit it.
[0120] In this embodiment of the application, the density of the point cloud in the multi-dimensional scene structure is increased in the point cloud information. After obtaining the increased point cloud information, a specific adjustment strategy can be introduced to enhance the fitting degree in areas of large-scale motion.
[0121] In this embodiment, the objective function in formula (2) can be used to optimize the sub-scene:
[0122] L=λ rgb L rgb +λ depth L depth +λ warp L warp (2)
[0123] In the objective function, Lrgb combines L1 loss and SSIM (structural similarity) loss. Ldepth is the L1 loss between the rendered depth map and the depth map generated by projecting sparse LiDAR points onto the camera plane. Lwarp is the L1 loss between the rendered image and the virtual deformable view.
[0124] In the embodiments of this application, the relationship between multiple spatial blocks is specifically the spatial relationship between multiple blocks.
[0125] In this embodiment of the application, the process of the scene reconstruction device establishing the target spatial scene corresponding to the sequence of images to be processed based on the scene frame and the multiple sets of scene elements includes: merging the multiple sets of scene elements into the scene frame according to the relationship between multiple sub-scenes and the scene frame to obtain the target spatial scene.
[0126] In this embodiment, Unstructured Geometry Primitives (UGPs) can be organized into a Bounding Volume Hierarchy (BVH) tree, and UGP attributes can be calculated for intermediate nodes. Specifically, the UGPs are recursively divided by spatial median until each UGP is assigned to a leaf node. Once the BVH tree is constructed, starting from the leaf nodes of the first level, the UGP attributes of child nodes are recursively merged in a bottom-up manner.
[0127] For example, such as Figure 2 As shown, this is an overview of the improvements made by the hierarchical unified Gaussian primitives in reconstruction under large-scale and high-motion scenes. In large-scale dynamic scenes, the reconstruction quality of the hierarchical unified Gaussian primitives is significantly better than previous methods, specifically the Street GS and OmniRe methods in existing technologies. The hierarchical unified Gaussian primitives also demonstrate significant improvements in the reconstruction of pedestrians with large movements. Figure 3 As shown, large-scale dynamic scenes are constructed as hierarchical tree structures, where both static and dynamic elements are represented using unified Gaussian primitives, thus achieving efficient large-scale dynamic scene reconstruction. As shown in (a), the scene hierarchy consists of a root level, sub-scene levels, and primitive levels. The root level is the entry point managing the entire structure. At the sub-scene level, the scene is spatially divided into multiple sub-scenes, which are further categorized into sky, background, rigid objects, and non-rigid objects. As shown in (b), at the primitive level, each element is modeled using unified Gaussian primitives with different properties.
[0128] In this embodiment, higher-level UGPs exhibit larger 3D volumes, resulting in coarser scene representations. In contrast, lower-level UGPs have smaller 3D volumes, enabling finer representations. The attributes of the first-level UGPs are obtained by interpolating the attributes of the l-1 level UGPs according to carefully designed weights w, as shown in formula (3):
[0129]
[0130] Since objects typically occupy limited space in a scene, it is empirically recommended to set them as leaf nodes to achieve a better balance between rendering quality and speed.
[0131] In this embodiment, after constructing a hierarchy for all sub-scenes, they are merged to the root level. For each sub-scene, UGPs are loaded, and unnecessary UGPs are filtered out based on their distance from the sub-scene center. The UGPs are then sequentially merged to the root level. Sub-scenes are seamlessly integrated by leveraging the strong geometric priors of the global UGP framework and a segmentation strategy that allows for 50% spatial overlap between adjacent sub-scenes. The merged scene maintains both visual continuity and geometric consistency, without noticeable boundary artifacts.
[0132] In this embodiment, the method of this application is implemented using PyTorch and a custom CUDA kernel. Experiments on the dynamic city dataset were conducted on H20 GPUs, where parallel training of multiple sub-scenes across multiple GPUs was completed within 3 hours. For the Waymo dataset, experiments were conducted on a single 4090 GPU, and training was completed within 2 hours.
[0133] Currently, large-scale dynamic street scene datasets are not publicly available. To address this gap, this application introduces a dynamic city dataset. This dataset comprises eight image and radar data sequences captured at a frequency of 10 Hz, covering street scenes ranging from 600 meters to over 1 kilometer. Compared to publicly available datasets such as Waymo and PandaSet, the dynamic city dataset includes a wider range of street scenes. This application intends to release this dataset as an open resource to advance research on large-scale dynamic street scene reconstruction. Comparing the method in this application with existing methods on the dynamic city dataset is challenging because no current method can handle the complexity of large-scale dynamic scenes. To demonstrate the effectiveness of the algorithm in this application and ensure a fair comparison, two experiments were conducted: one with a large-scale dynamic scene and the other with a sub-scene extracted from a larger scene.
[0134] It's important to note that Waymo uses a real-world dataset, comprising thousands of driving segments collected on actual roads. Each segment spans tens to hundreds of meters and contains 20 seconds of sensor data sampled at 10Hz.
[0135] In the embodiments of this application, the quantitative results in Table 1 demonstrate that this application outperforms other methods in both large-scale dynamic scenes and smaller sub-scenes, showcasing its strong capabilities in large-scale scene representation and dynamic element modeling. Furthermore, by employing gsplat-based LOD (Level of Detail) technology and engineering enhancements, this application achieves real-time rendering of large-scale dynamic scenes. While the rendering speed of this application is slightly slower than 4DGS, the reconstruction quality far surpasses that of 4DGS.
[0136] Table 1
[0137]
[0138]
[0139] like Figure 4 As shown, this application performs best in interpolation and extrapolation tasks, especially exhibiting superior geometric properties when compared with the new, shifted view. This is due to the supervision of lidar depth and virtual warped views. Additional extrapolation results and ablation studies are provided in the supplementary materials.
[0140] In the embodiments of this application, the quantitative results in Table 2 and Figure 5 Qualitative results demonstrate that the proposed method outperforms other methods in reconstruction. Specifically, the proposed method achieves higher visual metrics for human objects, proving its powerful ability to model dynamic elements, even in large motion regions. In contrast, the proposed method and OmniRe exhibit more severe visual quality degradation on test frames. This is due to the inability to capture the motion of non-rigid elements on test frames, resulting in the loss of temporal information and a significant performance degradation on new frames. Nevertheless, the proposed method remains competitive with state-of-the-art methods.
[0141] Table 2
[0142]
[0143] It should be noted that Table 2 shows Waymo's quantitative comparisons. Every 10 frames were selected as test frames, and visual quality metrics were calculated, where * indicates the metric for pedestrian areas. Each cell is colored to indicate best or second best.
[0144] In the embodiments of this application, such as Figure 5 As shown: Qualitative comparison of the Waymo dataset. For typical scenes in the Waymo dataset, the sky is modeled using a sky cubemap. Compared with other methods, the method in this application significantly improves performance for large dynamic regions such as pedestrian feet and legs.
[0145] This application conducts an ablation study on a segmented object training strategy and calculates the visual metrics of vehicles most likely to span multiple sub-scenes. Table 3 and Figure 6 This paper presents the qualitative and quantitative results of the ablation study of the segmented object training strategy presented in this application. Further details regarding this strategy are provided in the supplementary materials.
[0146] Table 3
[0147] method PSNR PSNR+ W / o block-wise objects 28.58 21.51 W-block W-wise object 28.64 23.15
[0148] In the embodiments of this application, Table 3 shows the training strategy for ablation block smart objects. The "+" in Table 3 indicates the metrics for the vehicle region. The results show that the block-based object training strategy of this application can significantly improve the visual metrics of moving objects.
[0149] In the embodiments of this application, such as Figure 6 As shown, a segmented object ablation training strategy is employed. This application highlights the impact of the block-based object training strategy on vehicles to demonstrate its effectiveness.
[0150] Table 4 shows the ablation study results comparing various strategies in this application. Specific t-scale initialization has the greatest impact, indicating that UGP is highly sensitive to the effects of time on initialization. This application finds that the best practice is to align it with the data sampling frequency. Specific adjustment strategies also have significant effects, such as... Figure 7 As shown, this allows UGP to densify and add more primitives in large motion regions. The initial point densification strategy enhances the fit by increasing the initial point cloud, while global covariance enables UGP to better capture the geometry of the elements.
[0151] Table 4
[0152] method PSNR SSIM W / o init point densification 34.28 0.939 W / o specific t size 23.22 0.791 W / o specific adjustments 27.15 0.831 W / o Global Covariance 34.88 0.943 Complete method 34.93 0.943
[0153] As shown in Table 4: Ablation optimization for non-rigid objects. This application conducted an ablation study on non-rigid optimization strategies in Waymo and calculated human visual metrics.
[0154] The contribution of this application lies in proposing a novel hierarchical model, which not only improves the representation of large-scale dynamic scenes but also enables real-time rendering, which is of great significance for promoting technological progress in related fields.
[0155] In this embodiment of the application, the process by which the scene reconstruction device initializes the sequence of images to be processed to obtain a scene frame of the sequence of images to be processed includes: establishing a multi-layer structure of the sequence of images to be processed; and using the multi-layer structure as the scene frame.
[0156] In this embodiment, the Hierarchy Unified Gaussian Primitives (UGP) is based on the idea of constructing a hierarchical structure containing a root layer, sub-scene layers, and primitive layers, using Unified Gaussian Primitives (UGP) defined in 4D space as representations. Establishing a multi-layered structure for a sequence of images to be processed is equivalent to establishing a hierarchical structure containing a root layer, sub-scene layers, and primitive layers.
[0157] It should be noted that the root layer serves as the entry point for the hierarchical structure. In the sub-scene layer, the scene is spatially divided into multiple sub-scenes, from which various elements are extracted. In the primitive layer, each element is modeled using UGP, and its global pose is controlled by time-dependent action priors. This hierarchical design significantly enhances the model's capacity, enabling this application to simulate large-scale scenes. Furthermore, the UGP in this application allows for the reconstruction of arbitrary dynamic elements. Experiments were conducted on Dynamic City (an internally proprietary large-scale dynamic street scene dataset) and the public Waymo dataset. Experimental results show that this application achieves state-of-the-art performance. The accompanying code and the Dynamic City dataset can be released as open resources to further promote research within the community.
[0158] In this application, given a series of images captured from a dynamic large-scale scene over a long period, the goal is to achieve high-fidelity reconstruction and real-time rendering of large-scale scenes containing arbitrary dynamic elements. To this end, this application proposes a novel scene representation method—Hierarchy Unified Gaussian Primitives (HGP)—aimed at efficiently representing large-scale dynamic scenes. The core of this method is to construct a hierarchical UGP tree structure, providing an efficient representation for complex mixtures of static and dynamic environments.
[0159] Previous methods have limitations in modeling large-scale dynamic scenes: methods capable of handling large-scale environments struggle with dynamic elements, while methods focused on dynamic elements are often limited to typical or fragmented scenes, failing to capture comprehensive street scenes. This application proposes a hierarchical UGP, which models large-scale scenes containing arbitrary dynamic elements through hierarchically structured Unified Gaussian Primitives (UGPs).
[0160] The hierarchical structure of this application consists of three levels: the root level, the sub-scene level, and the primitive level. The root level abstracts the entire scene and serves as the entry point for managing the hierarchy. At the sub-scene level, the large-scale scene is divided into multiple sub-scenes based on spatial information and various scene elements. Each sub-scene is reconstructed independently and ultimately merged at the root level. At the primitive level, various scene elements are modeled using UGPs with different attributes.
[0161] Large-scale scene reconstruction at the sub-scene level faces challenges in terms of memory usage, computational speed, and model capacity, while typical scenes are easier to manage. To address this issue, a divide-and-conquer approach is adopted, dividing the large-scale scene into multiple sub-scenes for parallel reconstruction.
[0162] Given a sequence of images captured from a large-scale scene over a long period, the camera pose can be estimated using the Structure from Motion (SfM) method, generating a sparse point cloud. This application predefines the size of each chunk and uses this size to divide the large-scale scene into multiple chunks. Subsequently, based on the camera pose, the point cloud, and the spatial relationships between each chunk, the image sequence and the point cloud are assigned to the corresponding chunks, with each chunk corresponding to a sub-scene in the sub-scene layer.
[0163] In the sub-scene, elements are assembled within a set S. Elements in the sub-scene are categorized into four types: background, sky, rigid body objects, and non-rigid body objects, denoted as Ne. During reconstruction, each element is modeled in its respective local coordinate system. During rendering, elements are transformed to the global coordinate system and then jointly optimized. The transformation of elements from the local coordinate system to the global coordinate system (as shown in Equation (4)) is simulated as a time-dependent motion prior M(t) = (R... t |T t ):
[0164]
[0165] It should be noted that Rt is the orientation (rotation) of the element in the global coordinate system, and Tt is the position (translation) of the element in the global coordinate system. For local coordinates, These are global coordinates.
[0166] For any time point t, the motion prior M(t) = (I|O) is used for the background and sky, while for rigid and non-rigid objects, Rt and Tt can be easily obtained using off-the-shelf trackers. This application's method employs a divide-and-conquer strategy at the sub-scene level to reduce the complexity of large-scale scenes and alleviate reconstruction challenges. Furthermore, this application decouples various elements within the sub-scene, enabling more refined modeling at the primitive level.
[0167] Primitive Layer: Dynamic reconstruction of large-scale scenes is an important and challenging task. Existing methods have demonstrated that Gaussian primitives offer significant advantages in dynamic street reconstruction. Recent methods have shown the great potential of Gaussian primitives in large-scale and dynamic reconstruction.
[0168] In this application, elements in the sub-scene are modeled as Unified Gaussian Primitives (UGPs), denoted as G(μ, ∑, o, SH), which are defined in 4D space. These UGPs are parameterized by the mean μ, the covariance matrix ∑, the opacity o, and the spherical harmonic coefficients SH. Specifically, the mean of the UGP is μ = (μx, μy, μz, μt). The covariance matrix ∑ is parameterized as a configuration of a 4D ellipsoid as shown in Equation (5):
[0169] ∑=RSS T R T (5)
[0170] Here, S = diag(sx, sy, sz, st) is a diagonal matrix representing the scale of UGP in 4D space, while the 4D rotation is parameterized by two isotropic quaternions ql = (a, b, c, d) and qr = (p, q, r, s) and constructed in accordance with Equation (6):
[0171] R = R(q) l )R(q r (6)
[0172] Here, R(q1) and R(qr) are symmetric extensions of two isotropic quaternions. Based on their complexity, this application uses UGPs with different properties to represent various elements in the sub-scene. Specifically, since the mean of the UGPs of local static elements remains constant over time, there is no rotation on the time-dependent plane, and they have infinite temporal extension, this application sets μt = 0, qr = (1, 0, 0, 0), and st = ∞ for the background and rigid objects.
[0173] The rendering process of Hierarchy Unified Gaussian Primitives (Hierarchy UGP) consists of three main stages: element aggregation, primitive selection, and image rendering.
[0174] Element aggregation is a key step in the Hierarchy UGP rendering process, involving summing the contributions of all elements to generate the final image. Given a time point t, the conditional 3D mean and covariance of UGPs at that time point can be obtained through the following steps as shown in Equation (7):
[0175]
[0176] After slicing UGP into 3D space, each element e of the current time step in the sub-scene is transformed into the global coordinate system, and then they are combined as shown in Equation (8):
[0177]
[0178] During this process, the pose of elements can be manually adjusted to facilitate the editing of dynamic scenes. Since 4D rotation lacks geometric meaning and cannot be directly applied to coordinate transformation, the transformation is applied to the covariance matrix. The specific coordinate transformation method is shown in formula (9):
[0179]
[0180] Primitive selection occurs after element aggregation. This step aims to improve rendering efficiency by selecting only the necessary UGPs for rendering from a given viewpoint, rather than rendering all UGPs directly. Specifically, starting from the root node of the Hierarchy UGP, UGPs are selected level by level until a level sufficient for accurate image rendering is reached.
[0181] To achieve efficient primitive selection, a threshold τ is set, and the selection process begins from the root node of the hierarchy. For each UGP, its diameter on the image plane is calculated. Then, it is checked whether this diameter is less than τ, or whether the primitive has child nodes. If neither of these conditions is met, the child nodes of the current UGP are evaluated, and this evaluation process is repeated. This iterative selection process continues until all UGPs satisfy the condition that their diameter is less than the threshold τ, or no more child nodes are available.
[0182] To calculate the diameter of the UGP on the image plane, its covariance matrix must first be projected onto the 2D image plane, as shown in Equation (10):
[0183] ∑′=JW∑ globai W T J T (10)
[0184] In this context, W is the view transformation matrix, which contains the camera's position and orientation information, used to transform 3D world coordinates into the camera coordinate system. J is the Jacobian matrix, which represents the affine approximation of the projection transformation. The projection covariance matrix ∑′ defines an ellipse on the image plane, and the diameter of the UGP is determined by the length of the principal axis of this ellipse.
[0185] During the rendering process, the probability of each UGP appearing at time t is obtained through the slicing process. This probability is determined by... The opacity of UGP at that time is given by the probability, and then weighted by this probability, as shown in Equation (11):
[0186] αt=p(t)α (11)
[0187] The color c of each UGP is calculated using spherical harmonics (SH). Finally, the image I is rendered at time t using differentiable alpha blending, as shown in Equation (12):
[0188]
[0189] The process of building a hierarchical unified Gaussian primitive (Hierarchy UGP) consists of three stages: initialization of the global UGP framework, optimization of the sub-scene layer, and the final merging process.
[0190] The optimization process involves training each sub-scene individually. Using the motion priors of each element, elements in each sub-scene are modeled in the local space, and then transformed into the global space for joint optimization.
[0191] During the optimization phase, a block-wise object training strategy is employed to handle dynamic objects spanning multiple sub-scenes, thus avoiding interference between blocks. To improve the reconstruction quality of non-rigid objects, temporal and spatial interpolation is performed to increase the density of the initial point cloud, and specific adjustment strategies are introduced to enhance the fitting in regions of large-scale motion.
[0192] Understandably, once the scene reconstruction device acquires a sequence of images to be processed, it uses hierarchical unified Gaussian units to establish the target spatial scene corresponding to the sequence of images to be processed. Hierarchical unified Gaussian units can effectively simulate complex dynamic objects and static objects in large-scale dynamic scenes, thereby improving the accuracy of scene reconstruction.
[0193] Based on the same inventive concept as the above-mentioned scene reconstruction method, this application provides a scene reconstruction device 1, corresponding to a scene reconstruction method; Figure 8 This is a schematic diagram of the composition structure of a scene reconstruction device provided in an embodiment of this application. The scene reconstruction device 1 may include:
[0194] Acquisition unit 11 is used to acquire a sequence of images to be processed;
[0195] Establishment unit 12 is used to establish the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian units.
[0196] In some embodiments of this application, the apparatus further includes a processing unit;
[0197] The processing unit is used to initialize the sequence of images to be processed to obtain the scene framework of the sequence of images to be processed; and to optimize the sequence of images to be processed to obtain multiple sets of scene elements in the sequence of images to be processed.
[0198] The establishment unit 12 is used to establish the target space scene corresponding to the sequence of images to be processed based on the scene framework and the multiple sets of scene elements.
[0199] In some embodiments of this application, the apparatus further includes a division unit and an extraction unit;
[0200] The segmentation unit is used to segment the sequence of images to be processed into multiple sub-scenes;
[0201] The extraction unit is used to extract scene elements from the multiple sub-scenes respectively to obtain the multiple sets of scene elements.
[0202] In some embodiments of this application, the apparatus further includes a determining unit;
[0203] The partitioning unit is used to partition the sequence of images to be processed according to a preset spatial size to obtain multiple spatial blocks;
[0204] The extraction unit is used to extract device posture information from the sequence of images to be processed; the device posture information is the posture information of the acquisition device when acquiring the sequence of images to be processed.
[0205] The acquisition unit 11 is used to acquire image information corresponding to the sequence of images to be processed;
[0206] The determining unit is used to determine the multiple sub-scenes based on the multiple spatial blocks, the device posture information, and the image information, wherein the multiple spatial blocks correspond one-to-one with the multiple scenes.
[0207] In some embodiments of this application, the apparatus further includes a generating unit and an adding unit;
[0208] The generation unit is used to generate point cloud information of a multi-dimensional scene structure based on the device posture information and the image information;
[0209] The adding unit is used to add the point cloud information and the sequence of images to be processed to the multiple spatial blocks according to the point cloud information, the device posture information and the relationship between the multiple spatial blocks, so as to obtain the multiple sub-scenes.
[0210] In some embodiments of this application, the apparatus further includes an adding unit;
[0211] The adding unit is used to increase the density of the point cloud in the multi-dimensional scene structure in the point cloud information to obtain the increased point cloud information.
[0212] The adding unit is used to add the added point cloud information and the sequence of images to be processed to the multiple spatial blocks according to the added point cloud information, the device posture information and the relationship between the multiple spatial blocks, so as to obtain the multiple sub-scenes.
[0213] In some embodiments of this application, the plurality of scene elements includes at least:
[0214] Sky, background, rigid body objects, and non-rigid body objects are one or more of the following: sky, background, rigid body objects, and non-rigid body objects.
[0215] In some embodiments of this application, the apparatus further includes a merging unit;
[0216] The merging unit is used to merge the multiple sets of scene elements into the scene frame according to the relationship between the multiple sub-scenes and the scene frame, so as to obtain the target space scene.
[0217] In some embodiments of this application, the establishment unit 12 is used to establish a multi-layer structure of the sequence of images to be processed;
[0218] The determining unit is used to use the multi-layer structure as the scene frame.
[0219] It should be noted that, in practical applications, the acquisition unit 11 and the establishment unit 12 mentioned above can be implemented by the processor 13 on the scene reconstruction device, specifically by a CPU (Central Processing Unit), MPU (Microprocessor Unit), DSP (Digital Signal Processor), or FPGA (Field Programmable Gate Array), etc.; the data storage mentioned above can be implemented by the memory 14 on the scene reconstruction device.
[0220] This application also provides a scene reconstruction device, such as... Figure 9 As shown, the scene reconstruction device includes a processor 13, a memory 14, and a communication bus 15. The memory 14 communicates with the processor 13 through the communication bus 15. The memory 14 stores programs executable by the processor 13. When the program is executed, the scene reconstruction method described above is executed by the processor 13.
[0221] In practical applications, the aforementioned memory 14 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 13.
[0222] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the scene reconstruction apparatus to perform the scene reconstruction method described above in this application.
[0223] This application provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by a processor, they cause the processor to execute the scene reconstruction method provided in this application. For example, ... Figure 1 The scene reconstruction method is shown.
[0224] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEP ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0225] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0226] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0227] Understandably, once the scene reconstruction device acquires a sequence of images to be processed, it uses hierarchical unified Gaussian units to establish the target spatial scene corresponding to the sequence of images to be processed. Hierarchical unified Gaussian units can effectively simulate complex dynamic objects and static objects in large-scale dynamic scenes, thereby improving the accuracy of scene reconstruction.
Claims
1. A scene reconstruction method, characterized in that, The method includes: Obtain the sequence of images to be processed; The target spatial scene corresponding to the sequence of images to be processed is established using hierarchical unified Gaussian primitives.
2. The method according to claim 1, characterized in that, The step of establishing the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian units includes: The sequence of images to be processed is initialized to obtain the scene frame of the sequence of images to be processed; The sequence of images to be processed is optimized to obtain multiple sets of scene elements in the sequence of images to be processed. Based on the scene framework and the multiple sets of scene elements, a target spatial scene corresponding to the sequence of images to be processed is established.
3. The method according to claim 2, characterized in that, The optimization process of the sequence of images to be processed yields multiple sets of scene elements in the sequence of images to be processed, including: The sequence of images to be processed is divided into scenes to obtain multiple sub-scenes; Scene elements are extracted from the multiple sub-scenes to obtain the multiple sets of scene elements.
4. The method according to claim 3, characterized in that, The process of dividing the sequence of images to be processed into multiple sub-scenes includes: The sequence of images to be processed is divided into multiple spatial blocks according to a preset spatial size; Extract device pose information from the sequence of images to be processed; the device pose information is the pose information of the acquisition device when acquiring the sequence of images to be processed. Obtain the image information corresponding to the sequence of images to be processed; Based on the multiple spatial blocks, the device posture information, and the image information, the multiple sub-scenes are determined, and the multiple spatial blocks correspond one-to-one with the multiple scenes.
5. The method according to claim 4, characterized in that, The step of determining the multiple sub-scenes based on the multiple spatial blocks, the device posture information, and the image information includes: Generate point cloud information of a multi-dimensional scene structure based on the device posture information and the image information; Based on the point cloud information, the device posture information, and the relationship between the multiple spatial blocks, the point cloud information and the sequence of images to be processed are added to the multiple spatial blocks to obtain the multiple sub-scenes.
6. The method according to claim 5, characterized in that, The step of adding the point cloud information and the sequence of images to be processed to the multiple spatial blocks based on the point cloud information, the device pose information, and the relationship between the multiple spatial blocks to obtain the multiple sub-scenes includes: The density of the point cloud in the multi-dimensional scene structure is increased in the point cloud information to obtain the increased point cloud information. Based on the added point cloud information, the device posture information, and the relationship between the multiple spatial blocks, the added point cloud information and the sequence of images to be processed are added to the multiple spatial blocks to obtain the multiple sub-scenes.
7. The method according to any one of claims 2-6, characterized in that, The multiple sets of scene elements include at least: Sky, background, rigid body objects, and non-rigid body objects are one or more of the following: sky, background, rigid body objects, and non-rigid body objects.
8. The method according to claim 2, characterized in that, The step of establishing the target spatial scene corresponding to the sequence of images to be processed based on the scene framework and the multiple sets of scene elements includes: The multiple sets of scene elements are merged into the scene frame according to the relationship between the multiple sub-scenes and the scene frame to obtain the target space scene.
9. The method according to claim 2, characterized in that, The initialization process of the sequence of images to be processed to obtain the scene framework of the sequence of images to be processed includes: Establish a multi-layer structure for the sequence of images to be processed; The multi-layered structure is used as the scene framework.
10. A scene reconstruction device, characterized in that, The device includes: The acquisition unit is used to acquire a sequence of images to be processed. A unit is established to construct the target spatial scene corresponding to the sequence of images to be processed using hierarchical unified Gaussian units.
11. A scene reconstruction device, characterized in that, The device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, they implement the method described in any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method according to any one of claims 1 to 9.