Three-dimensional scene reconstruction method, electronic equipment and vehicle
By acquiring images from the vehicle's camera to construct an initial four-dimensional Gaussian plane and motion field, and adjusting parameters to construct the target three-dimensional scene, the problem of dependence on manual annotation and LiDAR in existing technologies is solved, and efficient and high-precision autonomous driving scene reconstruction is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU AUTOMOBILE GROUP CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-05-26
AI Technical Summary
Existing 3D reconstruction methods for autonomous driving scenarios rely on manually labeled 3D bounding boxes or LiDAR point clouds, resulting in high labor and hardware costs, and the reconstruction accuracy is difficult to guarantee in complex environments.
Images are acquired by the vehicle's camera, and an initial four-dimensional Gaussian plane and motion field are constructed using the camera's intrinsic parameters and pose. The rendering results are obtained by combining the image processing model, and the parameters are adjusted to construct the target three-dimensional scene, thus avoiding dependence on point clouds and 3D bounding boxes.
It improves the accuracy and efficiency of 3D scene reconstruction, reduces labor and hardware costs, and achieves high-fidelity reconstruction in dynamic scenes.
Smart Images

Figure CN122089907A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically, to a method for reconstructing a three-dimensional scene, an electronic device, and a vehicle. Background Technology
[0002] With the development of science and technology, vehicles with autonomous driving capabilities are becoming increasingly common. In optimizing the performance of autonomous driving systems, it is often necessary to synthesize virtual scenes using LiDAR point clouds or manually annotated 3D bounding boxes to test the system's behavior in various scenarios, thereby improving its safety. However, related technologies still face the challenge of improving the accuracy of 3D scene reconstruction. Summary of the Invention
[0003] In view of this, embodiments of this application propose a method for reconstructing a three-dimensional scene, an electronic device, and a vehicle to improve the above-mentioned problems.
[0004] In a first aspect, embodiments of this application provide a method for reconstructing a three-dimensional scene. The method includes: obtaining an initial four-dimensional Gaussian image based on an initial surround view image captured by a camera of a vehicle, the camera pose corresponding to the initial surround view image, and camera intrinsic parameters; obtaining an initial motion field based on the vehicle's position information, the camera pose, the camera intrinsic parameters, and the acquisition time of the initial surround view image, wherein the spatial range of the initial motion field is determined based on the position information, the camera pose, and the camera intrinsic parameters, and the temporal range of the initial motion field is determined based on the acquisition time; and obtaining an initial motion field based on the initial four-dimensional Gaussian image. The initial motion field is used to obtain the rendering result corresponding to the acquisition time. The rendering result includes a rendered panoramic image, a rendered depth map, a rendered semantic map, and a rendered motion flow. The rendered motion flow is used to characterize the displacement of pixels in the rendered panoramic image in the initial motion field at the acquisition time. Based on the rendering result and the initial panoramic image, the parameters of the initial four-dimensional Gaussian and the initial motion field are adjusted to obtain the target four-dimensional Gaussian and the target motion field. The target three-dimensional scene is constructed based on the target four-dimensional Gaussian and the target motion field.
[0005] Secondly, embodiments of this application provide a three-dimensional scene reconstruction device, the device comprising: a four-dimensional Gaussian initialization module, a motion field initialization module, a rendering result acquisition module, a four-dimensional Gaussian and motion field adjustment module, and a three-dimensional scene reconstruction module. The four-dimensional Gaussian initialization module is used to acquire an initial four-dimensional Gaussian based on an initial surround view image captured by a camera of a vehicle, the camera pose corresponding to the initial surround view image, and camera intrinsic parameters; the motion field initialization module is used to acquire an initial motion field based on the vehicle's position information, the camera pose, the camera intrinsic parameters, and the acquisition time of the initial surround view image, wherein the spatial range of the initial motion field is determined based on the position information, the camera pose, and the camera intrinsic parameters, and the temporal range of the initial motion field is determined based on the acquisition time; the rendering result acquisition module is used to acquire the rendering result based on the initial four-dimensional Gaussian and the initial motion field. The rendering result corresponding to the acquisition time includes a rendered surround view image, a rendered depth map, a rendered semantic map, and a rendered motion flow. The rendered motion flow is used to characterize the displacement of pixels in the rendered surround view image in the initial motion field at the acquisition time. The four-dimensional Gaussian and motion field adjustment module is used to adjust the parameters in the initial four-dimensional Gaussian and the initial motion field according to the rendering result and the initial surround view image to obtain the target four-dimensional Gaussian and the target motion field. The three-dimensional scene reconstruction module is used to construct the target three-dimensional scene according to the target four-dimensional Gaussian and the target motion field.
[0006] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor, wherein the memory is coupled to the processor, the memory stores instructions, and when the instructions are executed by the processor, the processor executes the three-dimensional scene reconstruction method provided in the first aspect above.
[0007] Fourthly, embodiments of this application provide a vehicle that includes the electronic equipment provided in the third aspect above.
[0008] Fifthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the three-dimensional scene reconstruction method provided in the first aspect.
[0009] In this application's solution, an initial four-dimensional Gaussian plane is obtained based on the initial surround view image captured by the vehicle's camera, the corresponding camera pose, and camera intrinsic parameters. An initial motion field is obtained based on the vehicle's position information, camera pose, camera intrinsic parameters, and the acquisition time of the initial surround view image. The spatial range of the initial motion field is determined based on the position information, camera pose, and camera intrinsic parameters, while the temporal range is determined based on the acquisition time. Based on the initial four-dimensional Gaussian plane and the initial motion field, the rendered surround view image, rendered depth map, and rendered semantic map corresponding to the acquisition time are obtained. The image and the rendering result of the rendered motion flow are shown. The rendered motion flow is used to represent the displacement of pixels in the rendered panoramic image at the acquisition time in the initial motion field. Based on the rendering result and the initial panoramic image, the parameters in the initial four-dimensional Gaussian and the initial motion field are adjusted to obtain the target four-dimensional Gaussian and the target motion field. The target three-dimensional scene is constructed based on the target four-dimensional Gaussian and the target motion field. Thus, without the need for point clouds and three-dimensional bounding boxes, the three-dimensional reconstruction of the dynamic scene is performed only through the images acquired by the vehicle's camera and the calibration parameters of the camera, which improves the efficiency and accuracy of the three-dimensional scene reconstruction. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a three-dimensional scene reconstruction method provided in an embodiment of this application is shown. Figure 2 A flowchart illustrating a three-dimensional scene reconstruction method provided in an embodiment of this application is shown. Figure 3 A flowchart illustrating a three-dimensional scene reconstruction method provided in an embodiment of this application is shown. Figure 4 An initial surround view image provided in an embodiment of this application is shown; Figure 5 This illustration shows a rendered surround view image provided in one embodiment of this application; Figure 6 An initial depth map provided in one embodiment of this application is shown; Figure 7 This application shows a rendered depth map provided in one embodiment. Figure 8 An initial semantic graph provided in one embodiment of this application is shown; Figure 9This application shows a rendered semantic map provided in one embodiment. Figure 10 A schematic diagram of the initial optical flow provided in an embodiment of this application is shown; Figure 11 A schematic diagram of the predicted optical flow provided in an embodiment of this application is shown; Figure 12 A schematic diagram of a target three-dimensional scene provided in an embodiment of this application is shown; Figure 13 A block diagram of a three-dimensional scene reconstruction apparatus provided in one embodiment of this application is shown; Figure 14 A block diagram of an electronic device for performing a method for reconstructing a three-dimensional scene according to an embodiment of the present application is shown. Detailed Implementation
[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0013] To better understand the solutions of the embodiments of this application, the technical terms used in the embodiments of this application will be explained below.
[0014] A multi-scale hash map is a data structure that combines multi-scale analysis with hash tables. It is mainly used for the efficient storage and retrieval of complex geometric or large-scale data.
[0015] CARLA, Car Learning to Act, is an open-source autonomous driving simulation platform based on Unreal Engine 4, primarily used for closed-loop evaluation and testing of end-to-end autonomous driving (E2E-AD) systems.
[0016] AirSim, an open-source cross-platform simulation environment from Microsoft, is designed to provide high-fidelity physical modeling, climate simulation, and perception sensor support for unmanned aerial vehicles (UAVs) and autonomous vehicles.
[0017] Neural Radiance Fields (NeRF).
[0018] Neural Implicit Surfaces (NeuS)
[0019] 3D Gaussian Splatting (3DGS) is a 3D reconstruction and rendering technique based on Gaussian distribution.
[0020] The Waymo Open Dataset is the largest publicly available dataset in the field of autonomous driving, released by Waymo, the self-driving car company owned by Alphabet (Google's parent company).
[0021] DBSCAN, Density-Based Spatial Clustering of Applications with Noise, is a density-based clustering algorithm that defines a cluster as the largest set of density-connected points and can discover clusters of arbitrary shapes in noisy data.
[0022] Peak signal-to-noise ratio (PSNR).
[0023] Structural dissimilarity is an image quality assessment method that measures perceptual differences between images by calculating the inverse of structural similarity (SSIM) to define an error value.
[0024] LPIPS (Learned Perceptual Image Patch Similarity), also known as "perceptual loss," is used to measure the difference between two images.
[0025] S-NeRF is a Neural Radiation Field (NeRF) algorithm optimized for street view data. It mainly solves the problems of small view overlap and no scene boundary in traditional NeRF in street view scenes.
[0026] SUDS, Scalable Urban Dynamic Scenes, is primarily used for the decomposition and reconstruction of dynamic scenes, especially in the fields of computer vision and 3D reconstruction.
[0027] Multivariate Adaptive Regression Splines (MARS) are a nonparametric regression technique.
[0028] EmerNeRF is an autonomous driving simulation framework based on NeRF (Neural Radiation Field). It achieves high-fidelity reconstruction of dynamic driving scenarios through self-supervised learning, without the need for manual annotation or pre-trained model segmentation.
[0029] StreetSurf is a 3D reconstruction and sensor simulation framework for the field of autonomous driving, primarily used to generate high-precision virtual scenes for testing autonomous driving systems (SDV). Its core functions include scene reconstruction, dynamic participant simulation, and closed-loop testing.
[0030] The PVG algorithm, proposed by OpenAI, is an interactive proof system based on game theory. It improves the readability, accuracy, and security of AI model output by training two models (proverer and verifier) to engage in collaborative game.
[0031] StreetGaussians is a novel algorithm for street scene reconstruction in autonomous driving. It enables efficient scene modeling and editing by representing dynamic urban street scenes as a set of 3D Gaussian point clouds.
[0032] SplatFlow is a real-time rendering algorithm based on 3D Gaussian splashing (3DGS), primarily used for the fusion of LiDAR and camera data in autonomous driving scenarios.
[0033] The implementation details of the technical solutions in the embodiments of this application are described in detail below: With the application of end-to-end autonomous driving models, the importance of closed-loop evaluation of autonomous driving systems is becoming increasingly prominent. However, traditional closed-loop evaluation systems (such as CARLA and AirSim) rely on manually constructing virtual scenes, which is not only costly but also deviates from the real world. In contrast, 3D reconstruction technology can utilize images or point cloud data collected by autonomous vehicles to achieve high-precision digital reproduction of scenes. By editing the reconstructed scene, the behavior of the autonomous driving model under different conditions can be systematically tested, thereby allowing for targeted model optimization and ultimately improving its safety.
[0034] Currently, 3D reconstruction methods for autonomous driving scenarios can be mainly divided into three categories: (1) NeRF-based methods; (2) NeuS-based methods; and (3) 3DGS-based methods. Table 1 uses the Waymo public dataset as an example to compare and show the reconstruction accuracy (PSNR, SSIM, LPIPS) and rendering speed (FPS) of representative algorithms in these methods.
[0035] Table 1 As shown in Table 1, 3DGS-based methods are increasingly widely used in autonomous driving scene reconstruction due to their superior performance. However, the reconstruction of dynamic objects remains a key challenge. Existing methods mostly rely on manually annotated 3D bounding boxes or LiDAR point clouds to learn the motion information of dynamic objects. This dependence not only introduces additional labor and hardware costs but also limits the applicable scenarios of 3D reconstruction algorithms.
[0036] For example, the patent "An Unsupervised 4D Automatic Annotation Method for Street Scenes in Autonomous Driving Based on 3D Reconstruction" proposes an unsupervised 4D automatic annotation process. This method first extracts pixel-level semantic and depth information from the input frame-by-frame image using a general pre-trained image segmentation model and a depth estimation model, respectively, and obtains instance segmentation results accordingly. Then, based on the instance segmentation results, it calculates the 2D bounding box and its feature vector for each instance, and achieves cross-frame matching and tracking of instances by comparing the Euclidean distance between feature vectors in different frames. Next, it uses depth information to convert the image into a pseudo-point cloud. Based on this, it employs the implicit surface reconstruction method StreetSurf based on the signed distance field to perform instance-level reconstruction of the entire scene. Finally, it samples the reconstructed scene frame by frame to obtain instance point cloud data for each frame, and after denoising using the DBSCAN algorithm, it finally calculates the required 3D annotation boxes. However, this patented method has several limitations. First, the 2D annotation boxes it relies on are obtained through semantic segmentation and feature matching, which are prone to matching failures in complex situations such as occlusion or foreground changes, directly affecting the quality of subsequent reconstruction. Secondly, the StreetSurf method used is based on the NeuS framework, which suffers from slow training convergence and low rendering efficiency (see Table 1). Finally, the DBSCAN algorithm used in the denoising stage is extremely sensitive to parameter settings, and the variable size of dynamic targets and uneven density of pseudo-point clouds in autonomous driving scenarios further affect the stability and reliability of the denoising effect.
[0037] Therefore, autonomous driving 3D reconstruction technology can generate high-fidelity 3D scenes using images or point cloud data collected by vehicles. Combined with scene editing, this technology can not only synthesize data to improve system performance, but also test system behavior in diverse virtual environments, which is key to achieving closed-loop evaluation. However, most existing methods rely heavily on LiDAR point clouds or manually annotated 3D bounding boxes, which not only leads to high labor and hardware costs, limiting the scope of application, but also makes it difficult to guarantee the accuracy of the created 3D scenes.
[0038] To address the aforementioned problems, the inventors, through extensive research, have developed a method, electronic device, and vehicle for reconstructing a 3D scene, as provided in this application. This method achieves dynamic 3D scene reconstruction using only images captured by the vehicle's camera and the camera's calibration parameters, without requiring point clouds or 3D bounding boxes, thus improving the accuracy of 3D scene reconstruction. The specific 3D scene reconstruction method will be described in detail in subsequent embodiments.
[0039] The embodiments involved in this application will now be described with reference to the accompanying drawings.
[0040] Please see Figure 1 , Figure 1A flowchart illustrating a three-dimensional scene reconstruction method according to an embodiment of this application is shown. In a specific embodiment, this three-dimensional scene reconstruction method can be applied to, for example... Figure 13 The three-dimensional scene reconstruction device 200 shown and the electronic device 100 equipped with the three-dimensional scene reconstruction device 200 are shown. Figure 14 The following will use an electronic device as an example to illustrate the specific process of this embodiment. Of course, it is understood that the electronic device used in this embodiment may include vehicles, in-vehicle terminals, computers, etc., and is not limited thereto. The following will focus on... Figure 1 The process shown is described in detail. The method for reconstructing the 3D scene may specifically include the following steps: Step S110: Obtain the initial four-dimensional Gaussian based on the initial surround view image captured by the vehicle's camera, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters.
[0041] In some implementations, the electronic device can be a vehicle (in this embodiment, it can be understood as a self-driving car). The self-driving car can be equipped with multi-view cameras, which can capture multiple frames of multi-view images. For example, the multi-view cameras of the self-driving car may include cameras in the front, rear, left, and right directions. These four cameras can be understood as the self-driving car's surround-view cameras, enabling 360-degree environmental monitoring.
[0042] The vehicle can acquire multiple consecutive temporal images via surround-view cameras, and these consecutive temporal images can be determined as the initial surround-view image captured by the vehicle's cameras. This initial surround-view image can include images captured by cameras from multiple directions of the vehicle, or it can include an image fused from images captured by cameras from multiple directions of the vehicle. For example, the vehicle can acquire several consecutive temporal images captured by the surround-view cameras and determine these temporal images as the initial surround-view image. The temporal images can be characterized as... , ={ |c=1... k=1... } Where H can represent the height of the time series image, W can represent the width of the time series image, and N... C N can represent the serial number of the surround-view camera. K This can represent the sequence number of the acquired time-series images. Specifically, during the acquisition of time-series images, the vehicle can obtain the timestamp t={ corresponding to the acquired time-series images. }, where k can represent the sequence number of the time series image.
[0043] After acquiring the initial surround-view image from the camera, the vehicle can obtain the camera's intrinsic parameters and, based on its position information at the time of acquiring the initial surround-view image, determine the camera pose in the world coordinate system. For example, the vehicle can use the time-series image acquired by the surround-view camera as the initial surround-view image and, based on the vehicle's position information, obtain the camera pose P={ in the world coordinate system of this time-series image. } and camera intrinsic parameters K={ } Where c can represent the camera sequence number, and k can represent the sequence number of the time-series image. The vehicle can pre-store the intrinsic parameters of the cameras corresponding to the surround-view cameras. The vehicle can be equipped with a positioning device (e.g., a global navigation system, inertial navigation system, etc.), through which the vehicle can obtain its position information in real time. During the process of acquiring initial surround-view images through the cameras, the vehicle can obtain the camera pose corresponding to the camera based on its position information at the time of acquisition of the initial surround-view images.
[0044] In some implementations, after the vehicle acquires the initial surround view image captured by the vehicle's camera, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters, an initial four-dimensional Gaussian can be obtained based on the initial surround view image, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters.
[0045] The initial 4D Gaussian coordinate system can be determined by the vehicle based on a pseudo-point cloud in the camera coordinate system. This pseudo-point cloud can be determined by the vehicle based on an initial surround view image, an initial depth map corresponding to the initial surround view image, and an initial semantic map corresponding to the initial surround view image. The camera coordinate system can be determined by the vehicle based on the camera pose and camera intrinsic parameters.
[0046] The vehicle can be pre-configured with image processing models (e.g., depth estimation models, semantic segmentation models, optical flow estimation models, etc.). After acquiring an initial surround-view image, the vehicle can input this image processing model to obtain the initial depth map, initial semantic map, and initial optical flow corresponding to each pixel in the initial surround-view image. This image processing model can be used to obtain the depth information of each pixel in the image, converting the image into a pseudo-point cloud. It can also be used to obtain the category ID of each pixel in the image, distinguishing whether a pixel belongs to a dynamic or static object based on this category ID. Furthermore, the image processing model can be used to calculate the optical flow of each pixel based on two consecutive frames captured by the same camera on the vehicle, improving the accuracy of 3D reconstruction of dynamic objects in the scene. Optical flow characterizes the displacement of a pixel across two consecutive frames.
[0047] For example, the vehicle is pre-configured with an image processing model, which may include a depth estimation model, a semantic segmentation model, and an optical flow estimation model. The vehicle can then process the acquired initial surround view image... The input is fed into the depth estimation model to obtain the initial depth map D={ The vehicle can collect the initial surround view image. The input is fed into the semantic segmentation model to obtain the initial semantic graph M={ }, where C can represent the number of object categories (e.g., pedestrians, vehicles, lane lines, etc.) included in the initial surround view image output by the semantic segmentation model. The vehicle can then collect the initial surround view image. The input is given to the optical flow estimation model to obtain the initial optical flow F={ }
[0048] In some implementations, after the vehicle acquires the initial semantic map corresponding to the initial surround view image, it can generate a mask based on the initial semantic map, and subsequently define an initial four-dimensional Gaussian mask based on this mask. For example, the vehicle can generate dynamic and static masks based on the classification of dynamic and static objects in the initial semantic map. Specifically, the vehicle can generate a dynamic mask MD={ based on the category ID of the dynamic object in the initial semantic map. }, where the categories of dynamic objects can include: cars, pedestrians, bicycles, etc., where dynamic objects can be represented as 1 on the dynamic mask MD, and other objects can be represented as 0. Specifically, a static mask MS={ can be generated based on the category IDs of static objects in the initial semantic graph. In this context, the categories of static objects can include: roads, buildings, trees, etc. On the static mask MS, static objects can be represented as 1, and other objects can be represented as 0.
[0049] In some implementations, after the vehicle acquires the initial depth map corresponding to the initial surround view image, it can generate a pseudo-point cloud based on the initial depth map, initialize a 3D Gaussian, and then further expand the 3D Gaussian into an initial 4D Gaussian. Specifically, the vehicle can obtain the initial pseudo-point cloud based on the pixel coordinates of the initial surround view image, the initial depth map, and camera intrinsic parameters, and then expand the initialized 3D Gaussian based on the initial pseudo-point cloud, a dynamic mask, and a static mask to obtain the initial 4D Gaussian.
[0050] For example, the vehicle can back-project the pixel coordinates X of the initial surround view image based on the camera intrinsic parameter K to obtain the pseudo point cloud P in the camera coordinate system. C =K -1 X. Among them, the vehicle can also determine the pseudo point cloud P in the camera coordinate system based on the camera pose W. C Transform to world coordinates to obtain the initial pseudo-point cloud P=WP C .
[0051] Optionally, after the vehicle acquires the initial pseudo-point cloud P, in order to reduce the computational load of creating the 3D scene, the initial pseudo-point cloud P can be uniformly sampled, and the initialized 3D Gaussian can be expanded based on the uniformly sampled pseudo-point cloud, dynamic mask, and static mask to obtain the initial 4D Gaussian.
[0052] For example, the vehicle can uniformly sample the pseudo-point cloud P based on a preset sampling interval ts to obtain the initial pseudo-point cloud P after sampling. k ={P k |k=1, ts+1, ...N s ×ts+1}, where N s N can represent the total number of frames sampled from the pseudo-point cloud P. s =[(N K -1) / ts].
[0053] In some implementations, to improve the accuracy of subsequent 3D scene construction by combining the four-dimensional Gaussian with the motion field, dynamic objects can be filtered during the definition of the four-dimensional Gaussian in this embodiment. Specifically, the vehicle can generate a static pseudo-point cloud PS based on the static mask MS1 corresponding to the pseudo-point cloud P1 of the first frame in the initial pseudo-point cloud. 1, And a dynamic pseudo-point cloud PD1 is generated based on the dynamic mask MD1 corresponding to the pseudo-point cloud P1. Additionally, for the pseudo-point clouds {P} of the frames other than the first frame in the initial pseudo-point cloud... k |k=ts+1,...N s ×ts+1}, the vehicle can use the pseudo point cloud {P} from other frames. k |k=ts+1,...N s The static mask corresponding to ×ts+1 generates a static pseudo-point cloud {PS} k |k=ts+1,...N s ×ts+1}. Among them, the vehicle can be based on the static pseudo-point cloud PS1 and the static pseudo-point cloud {PS... k |k=ts+1,...N s ×ts+1}, obtain the overall static pseudo-point cloud PS={PS k |k=1, ts+1, ...N s×ts+1}.
[0054] In some implementations, the vehicle can assemble an initial three-dimensional Gaussian distribution based on multiple anisotropic three-dimensional Gaussian distributions. The anisotropic three-dimensional Gaussian can be characterized by a series of learnable parameters, which can be obtained from third-party experimental data and are not limited here. In order to represent dynamic scenes, in this embodiment, the vehicle can extend the three-dimensional Gaussian to a four-dimensional Gaussian G(t) = {μ(t), s, q, c, α, m}. Here, μ can be used to represent the center point in the four-dimensional Gaussian, s can be used to represent the scale of the four-dimensional Gaussian, q can be used to represent the quaternion of the rotation vector, c can be used to represent color, α can be used to represent opacity, and m can be used to represent the semantics of motion and statics.
[0055] After the vehicle obtains the overall static pseudo-point cloud PS and the first frame's dynamic pseudo-point cloud PD1, it can initialize the four-dimensional Gaussian obtained by expanding the three-dimensional Gaussian based on PS and PD1 to obtain the initial four-dimensional Gaussian. The center point μ in the four-dimensional Gaussian can be initialized based on the position information of PS and PD1, and the color c in the four-dimensional Gaussian can be initialized based on the colors of PS and PD1. The semantic m of the static pseudo-point cloud PS can be initialized to 0.95, 0.9, etc., without limitation. The semantic m of the dynamic pseudo-point cloud PD1 can be initialized to 0.05, 0.1, etc., without limitation. The opacity α can be initialized to 0.95, 0.9, etc., without limitation. The rotation vector q can be initialized to [1,0,0,0], etc., without limitation. The scale s can be initialized to [0.1,0.1,0.1], etc., without limitation.
[0056] Step S120: Obtain the initial motion field based on the vehicle's position information, the camera pose, the camera intrinsic parameters, and the acquisition time of the initial surround view image. The spatial range of the initial motion field is determined based on the position information, the camera pose, and the camera intrinsic parameters, and the temporal range of the initial motion field is determined based on the acquisition time.
[0057] In some implementations, the vehicle can define a motion field Φ based on a deep learning model, where the motion field Φ can include an encoder and a decoder. The encoder can consist of a multi-resolution hash grid and is used to encode the input to obtain features f; the decoder can consist of several fully connected layers and is responsible for decoding the features f into displacements. μ. The input to the motion field Φ can be the three-dimensional coordinates x and time t. k Correspondingly, the output of the motion field Φ can be the position x at time t. kDisplacement relative to the initial time t1 μ.
[0058] in, μ can be characterized by formula (1): μ k =Φ(x,t) k ), Φ(·)=decoder(encoder(·)) (1) In some implementations, a motion field can be predefined in the vehicle, and the motion field can be initialized based on the vehicle's position information, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters to obtain the initial motion field.
[0059] Among them, after acquiring the acquisition time of the initial surround view image, the vehicle's position information at the acquisition time of the initial surround view image, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters, the vehicle can acquire the initial motion field based on the position information, the camera pose, the camera intrinsic parameters, and the acquisition time.
[0060] The spatial range of the initial motion field can be determined by the vehicle based on the location information, the camera pose, and the camera's intrinsic parameters. The temporal range of the initial motion field can be determined by the vehicle based on the data acquisition time.
[0061] For example, the vehicle can obtain the given pixel coordinates [[0, 0], [W, 0], [0, H], [W, H]] and the maximum camera depth dmax based on the camera intrinsic parameters, where H represents the image height and W represents the image width. The vehicle can also back-project the given pixel coordinates and the maximum camera depth dmax based on the camera pose W and camera intrinsic parameters K to obtain the coordinates of the farthest point of the camera frustum, the minimum value xyz_min of the frustum coordinates, and the maximum value xyz_max of the frustum coordinates, thus obtaining the spatial range of the camera frustum. The vehicle can obtain the time range by obtaining the minimum value t_min and the maximum value t_max of the timestamp corresponding to the initial panoramic image. Finally, the vehicle can initialize the range of the motion field Φ based on its position information, [xyz_min, t_min], and [xyz_max, t_max] to obtain the initial motion field.
[0062] Step S130: Based on the initial four-dimensional Gaussian and the initial motion field, obtain the rendering result corresponding to the acquisition time, wherein the rendering result includes the rendered panoramic image, the rendered depth map, the rendered semantic map, and the rendered motion flow, and the rendered motion flow is used to characterize the displacement of the pixels in the rendered panoramic image in the initial motion field at the acquisition time.
[0063] In some implementations, after obtaining the initial four-dimensional Gaussian surface and the initial motion field, the vehicle can obtain the rendering result corresponding to the acquisition time based on the initial four-dimensional Gaussian surface and the initial motion field. Specifically, the vehicle can target time t... k The motion flow corresponding to the pixel in the initial motion field is calculated using formula (1). μ, and this motion flow can be substituted into the initial four-dimensional Gaussian to obtain the four-dimensional Gaussian G(t) for rendering. k ) = {μ(t1) + The expression is defined as follows: μ, s, q, c, α, m. Here, μ(t1) represents the position information of the pixel in the initial motion field at time t1, s represents the scale of the four-dimensional Gaussian used for rendering, q represents the quaternion of the rotation vector of the four-dimensional Gaussian used for rendering, c represents the color of the pixel rendered in the four-dimensional Gaussian, α represents the opacity of the pixel rendered in the four-dimensional Gaussian, and m represents the static / dynamic semantics of the pixel rendered in the four-dimensional Gaussian.
[0064] Among them, the four-dimensional Gaussian G(t) obtained by the vehicle for rendering k ) = {μ(t1) + After μ, s, q, c, α, m}, the four-dimensional Gaussian image to be rendered can be rendered to t. k From the camera's perspective at any given moment, obtain the rendered panoramic image I´, the rendered depth map D´, the rendered semantic map MD´, and the rendered motion flow. F k .
[0065] Among them, the vehicle can obtain the rendered surround view image I´ according to formula (2).
[0066] I´= (1-α) j (2) Among them, the vehicle can obtain the rendered depth map D´ according to formula (3).
[0067] D´= (1-α) j (3) Among them, the vehicle can obtain the rendered semantic graph MD´ according to formula (4).
[0068] MD´= (1-α) j (4) Among them, the vehicle can obtain the rendered motion flow according to formula (5). F k。 F k = (1-α) j (5) In formulas (2) to (5), i can represent the i-th Gaussian distribution, z can represent the distance from the center point of the Gaussian distribution to the camera plane, or the distance from the Gaussian point to the camera, and N can represent the number of Gaussian distributions.
[0069] Step S140: Based on the rendering result and the initial panoramic image, adjust the parameters in the initial four-dimensional Gaussian and the initial motion field to obtain the target four-dimensional Gaussian and the target motion field.
[0070] In some implementations, after the vehicle obtains the rendering results, it can adjust the parameters in the initial four-dimensional Gaussian and the initial motion field based on the rendering results and the initial surround view image to obtain the target four-dimensional Gaussian and the target motion field.
[0071] Optionally, the vehicle can obtain the difference between the rendered result and the initial surround view image, adjust the parameters in the initial four-dimensional Gaussian and the initial motion field, and obtain the target four-dimensional Gaussian and the target motion field.
[0072] As an feasible approach, the vehicle can obtain the first error between the rendered surround view image and the initial surround view image in the rendering result, and optimize the parameters in the initial four-dimensional Gaussian and the initial motion field by gradient descent of the first error to obtain the target four-dimensional Gaussian and the target motion field.
[0073] As another feasible approach, the vehicle can obtain the second error between the rendered depth map in the rendering result and the initial depth map corresponding to the initial surround view image, and optimize the parameters in the initial four-dimensional Gaussian and the initial motion field by gradient descent of the second error to obtain the target four-dimensional Gaussian and the target motion field.
[0074] As another feasible approach, the vehicle can obtain the third error between the rendered semantic map in the rendering result and the initial semantic map corresponding to the initial toroidal image, and optimize the parameters in the initial four-dimensional Gaussian and the initial motion field by gradient descent of the third error to obtain the target four-dimensional Gaussian and the target motion field.
[0075] As another feasible approach, the vehicle can obtain the predicted optical flow of adjacent time points from the rendered motion flow in the rendering result, and can obtain the fifth error between the predicted optical flow and the initial optical flow corresponding to the initial panoramic image. The vehicle can then optimize the parameters in the initial four-dimensional Gaussian and the initial motion field by gradient descent of the fifth error to obtain the target four-dimensional Gaussian and the target motion field.
[0076] As another feasible approach, the vehicle can obtain the first and second motion flows of adjacent moments from the rendered motion flow in the rendering result, and obtain regularization loss and smoothing loss based on the first motion flow, the second motion flow and the second error; and optimize the parameters in the initial four-dimensional Gaussian and the initial motion field by gradient descent of the regularization loss and smoothing loss to obtain the target four-dimensional Gaussian and the target motion field.
[0077] Step S150: Construct a target three-dimensional scene based on the target four-dimensional Gaussian and the target motion field.
[0078] It is understandable that 4D Gaussian is a representation of a 3D scene. Based on this, after the vehicle acquires the target 4D Gaussian and the target motion field, it can construct the target 3D scene based on the target 4D Gaussian and the target motion field. Thus, high-precision 3D reconstruction of the vehicle's driving scene can be completed using only images acquired by the vehicle's camera and calibration parameters, without the need for LiDAR point clouds and 3D bounding boxes, significantly reducing labor and hardware costs. In addition, the 3D reconstruction method based on 3D Gaussian improves the accuracy and rendering speed of 3D scene reconstruction compared to other methods based on NeRF or NeuS.
[0079] An embodiment of this application provides a method for reconstructing a 3D scene. This method obtains an initial four-dimensional Gaussian plane based on an initial surround view image captured by a vehicle's camera, the corresponding camera pose, and camera intrinsic parameters. It then obtains an initial motion field based on the vehicle's position information, camera pose, camera intrinsic parameters, and the acquisition time of the initial surround view image. The spatial range of the initial motion field is determined based on the position information, camera pose, and camera intrinsic parameters, while the temporal range is determined based on the acquisition time. Finally, based on the initial four-dimensional Gaussian plane and the initial motion field, it obtains a rendered surround view image and a rendered depth map corresponding to the acquisition time. The rendered semantic map and the rendered motion flow are used to represent the displacement of pixels in the rendered panoramic image at the acquisition time in the initial motion field. Based on the rendering results and the initial panoramic image, the parameters in the initial four-dimensional Gaussian and the initial motion field are adjusted to obtain the target four-dimensional Gaussian and the target motion field. The target three-dimensional scene is constructed based on the target four-dimensional Gaussian and the target motion field. Thus, without the need for point clouds and three-dimensional bounding boxes, dynamic scene three-dimensional reconstruction can be performed only through images acquired by the vehicle's camera and the camera's calibration parameters, improving the efficiency and accuracy of three-dimensional scene reconstruction.
[0080] Please see Figure 2 , Figure 2 A flowchart illustrating a three-dimensional scene reconstruction method according to an embodiment of this application is shown. This method is applied to the aforementioned electronic device, and will be discussed below. Figure 2 The process shown is described in detail. The method for reconstructing the 3D scene may specifically include the following steps: Step S201: Based on the pixel categories in the initial semantic map corresponding to the initial surround view image, obtain the first mask corresponding to the dynamic object in the initial surround view image and the second mask corresponding to the static object in the initial surround view image.
[0081] In some implementations, the vehicle may have a pre-set semantic segmentation model. This model can be used to obtain the category of each pixel in an image, whereby the category can distinguish whether a pixel is a dynamic object or a static object. After acquiring an initial surround-view image, the vehicle can input this image into the semantic segmentation model to obtain an initial semantic map output by the model. Based on the pixel categories in this initial semantic map, the vehicle can obtain a dynamic mask corresponding to dynamic objects as a first mask and a static mask corresponding to static objects as a second mask.
[0082] Step S202: Obtain the initial pseudo point cloud based on the pixel coordinates, camera pose, camera intrinsic parameters, and the initial depth map corresponding to the initial panoramic image.
[0083] In some implementations, the vehicle is pre-configured with a depth estimation model, which can be used to obtain depth information of each pixel in an image. After acquiring an initial surround-view image, the vehicle can input this image into the depth estimation model to obtain an initial depth map corresponding to the initial surround-view image, output by the depth estimation model.
[0084] In some implementations, after acquiring the camera pose, camera intrinsics, and initial depth map corresponding to the initial surround view image, the vehicle can obtain an initial pseudo point cloud based on the pixel coordinates, camera pose, camera intrinsics, and initial depth map in the initial surround view image.
[0085] Specifically, the vehicle can back-project the pixel coordinates and initial depth map from the initial surround view image based on the camera's intrinsic parameters to obtain a pseudo-point cloud in the camera coordinate system. It can then transform this pseudo-point cloud from the camera coordinate system to the world coordinate system based on the camera pose to obtain a pseudo-point cloud in the world coordinate system. Furthermore, it can sample this pseudo-point cloud in the world coordinate system at a preset sampling interval to obtain an initial pseudo-point cloud. The vehicle can pre-set a preset sampling interval, such as 2 frames or 3 frames, which is not limited here.
[0086] Step S203: Based on the initial pseudo point cloud, the first mask, and the second mask, obtain the second pseudo point cloud and the third pseudo point cloud, wherein the second pseudo point cloud is determined based on the first frame pseudo point cloud in the initial pseudo point cloud and the first mask corresponding to the first frame pseudo point cloud, and the third pseudo point cloud is determined based on the second mask corresponding to the initial pseudo point cloud.
[0087] In some implementations, after the vehicle acquires the initial pseudo point cloud, the first mask, and the second mask, it can acquire the second pseudo point cloud and the third pseudo point cloud based on the initial pseudo point cloud, the first mask, and the second mask.
[0088] The second pseudo-point cloud can be determined by the vehicle based on the first frame pseudo-point cloud in the initial pseudo-point cloud and the first mask corresponding to that first frame pseudo-point cloud. The first frame pseudo-point cloud can be the first pseudo-point cloud sampled in the initial pseudo-point cloud; for example, the initial pseudo-point cloud P... k ={P k |k=1, ts+1, ...N s The first frame of pseudo-point cloud in {×ts+1} is P1. After the vehicle acquires the first frame of pseudo-point cloud, it can obtain the second pseudo-point cloud PD1 based on the first mask MD1 corresponding to the first frame of pseudo-point cloud and the first frame of pseudo-point cloud.
[0089] The third pseudo-point cloud can be determined by the vehicle based on the second mask corresponding to the initial pseudo-point cloud. For example, the initial pseudo-point cloud P... k ={Pk |k=1,ts+1,...N s ×ts+1}, where the vehicle acquires the initial pseudo-point cloud P k ={P k |k=1,ts+1,...N s After ×ts+1}, the third pseudo-point cloud PS={PS} can be obtained based on the second mask corresponding to each initial pseudo-point cloud and the initial pseudo-point cloud. k |k=1, ts+1,...N s ×ts+1}. Here, the third pseudo-point cloud can be understood as the overall static pseudo-point cloud corresponding to the initial pseudo-point cloud.
[0090] Step S204: Obtain the initial four-dimensional Gaussian based on the second pseudo-point cloud and the third pseudo-point cloud.
[0091] In some implementations, after the vehicle acquires the second and third pseudo-point clouds, it can obtain an initial four-dimensional Gaussian based on these pseudo-point clouds. The vehicle can predefine a four-dimensional Gaussian, which can be obtained by combining a three-dimensional Gaussian with a representation of the dynamic scene. This three-dimensional Gaussian can be obtained through third-party experimental data, can be characterized by a series of learnable parameters, and can be composed of multiple anisotropic three-dimensional Gaussians.
[0092] For example, the four-dimensional Gaussian can be G(t) = {μ(t), s, q, c, α, m}. Here, μ can represent the center point in the four-dimensional Gaussian, s can represent the scale of the four-dimensional Gaussian, q can represent the quaternion of the rotation vector, c can represent color, α can represent opacity, and m can represent the semantics of motion and statics. After the vehicle acquires the second and third pseudo-point clouds, it can initialize the four-dimensional Gaussian based on these pseudo-point clouds to obtain an initial four-dimensional Gaussian. Among them, μ and c in the initial four-dimensional Gaussian can be obtained by initializing based on the position and color of the second pseudo point cloud and the third pseudo point cloud, respectively. The semantic m of the third pseudo point cloud PS in the initial four-dimensional Gaussian can be initialized to 0.95. The semantic m of the second pseudo point cloud PD1 in the initial four-dimensional Gaussian can be initialized to 0.05. The opacity α in the initial four-dimensional Gaussian can be initialized to 0.95. The rotation vector q in the initial four-dimensional Gaussian can be initialized to [1,0,0,0]. The scale s in the initial four-dimensional Gaussian can be initialized to [0.1, 0.1, 0.1].
[0093] Step S205: Obtain the initial motion field based on the vehicle's position information, the camera pose, the camera intrinsic parameters, and the acquisition time of the initial surround view image. The spatial range of the initial motion field is determined based on the position information, the camera pose, and the camera intrinsic parameters, and the temporal range of the initial motion field is determined based on the acquisition time.
[0094] For a description of step S205, please refer to the previous description of step S120, which will not be repeated here.
[0095] Step S206: Determine the initial motion flow corresponding to the acquisition time based on the initial motion field.
[0096] In some implementations, after the vehicle acquires the initial motion field, it can determine the initial motion flow corresponding to the acquisition time of the initial surround view image based on the initial motion field. Specifically, the vehicle can determine the initial motion flow based on the acquisition time t. k The initial motion flow is obtained using formula (1). The vehicle can collect data at time t. k Data collection time t k Substituting the coordinates x within the camera's field of view of the vehicle into formula (1), the initial motion flow is obtained. μ k =Φ(x,t) k ).
[0097] Step S207: Determine the four-dimensional Gaussian to be rendered corresponding to the acquisition time based on the initial motion flow and the initial four-dimensional Gaussian.
[0098] In some implementations, after the vehicle acquires the initial motion stream and the initial four-dimensional Gaussian surface, it can determine the four-dimensional Gaussian surface to be rendered corresponding to the acquisition time based on the initial motion stream and the initial four-dimensional Gaussian surface. For example, the vehicle acquires the initial motion stream... μ k After obtaining the initial four-dimensional Gaussian G(t) = {μ(t), s, q, c, α, m}, the four-dimensional Gaussian G(t) to be rendered corresponding to the acquisition time can be determined based on the initial motion flow and the initial four-dimensional Gaussian. k ) = {μ(t1) + μ, s, q, c, α, m}. Here, μ(t1) can be used to characterize the position information of the vehicle's coordinate x within the camera's field of view at time t1 in the initial motion field.
[0099] Step S208: Render the four-dimensional Gaussian image to be rendered to obtain the rendered panoramic image, the rendered depth map, the rendered semantic map, and the rendered motion flow.
[0100] In some implementations, after obtaining the four-dimensional Gaussian image to be rendered, the vehicle can render the four-dimensional Gaussian image to be rendered to obtain the rendered surround view image, the rendered depth map, the rendered semantic map, and the rendered motion flow. Among them, the vehicle can obtain the rendering results (e.g., the rendered surround view image, the rendered depth map, the rendered semantic map, and the rendered motion flow) based on the four-dimensional Gaussian image to be rendered according to formulas (2) to (5).
[0101] Step S209: Obtain the predicted optical flow based on the rendered depth map and the rendered motion flow.
[0102] In some implementations, after the vehicle acquires the rendered depth map and the rendered motion flow, it can obtain the predicted optical flow based on the rendered depth map and the rendered motion flow.
[0103] The vehicle can backproject the rendered depth map onto the world coordinate system to obtain a first pseudo-point cloud. Based on the rendered motion flow and the first pseudo-point cloud, it can obtain the predicted optical flow for the next moment corresponding to the acquisition time. For example, the vehicle can backproject the rendered depth map D´ onto the world coordinate system based on camera intrinsic parameters to obtain the first pseudo-point cloud P´=KW. -1 D´. Among them, the vehicle can be based on the first pseudo-point cloud and the rendered motion flow. F, obtain the coordinates of pixel X´ in the rendered toroidal image: X´=KW -1 (P´± F). The vehicle can obtain the predicted optical flow F' for the next time step corresponding to the acquisition time step based on the coordinates X' of the pixels in the rendered surround view image and the coordinates X' of the pixels in the initial surround view image, where [F', D]=X´-X. Where, D can characterize the difference between the predicted depth of a pixel in the rendered depth map and the actual depth of a pixel in the initial panorama image.
[0104] Step S210: Obtain the target loss based on the predicted optical flow, the rendering result, and the initial panoramic image.
[0105] In some implementations, after the vehicle acquires the predicted optical flow and rendering results, it can obtain the target loss based on the predicted optical flow, rendering results, and the initial surround view image.
[0106] As a feasible approach, the vehicle can obtain a first loss based on the rendered surround view image and the initial surround view image. Specifically, the vehicle can substitute the rendered surround view image I´ and the initial surround view image I into the first loss calculation formula to obtain the first loss. The first loss calculation formula is as follows: L rgb= ·||I´-I||1+ ·SSIM(I,I´) Among them, SSIM can characterize structural similarity loss. It can represent the first weight. The second weight can be used to represent the first weight. The first and second weights can be pre-stored in the vehicle, obtained through third-party experimental data, or set by the user; no restrictions are imposed here.
[0107] As an feasible approach, the vehicle can obtain a second loss based on the rendered depth map and the initial depth map corresponding to the initial surround view image. Specifically, the vehicle can substitute the rendered depth map D´ and the initial depth map D corresponding to the initial surround view image into the second loss calculation formula to obtain the second loss. The second loss calculation formula is as follows: L depth = ·||D´ / D-1||1 in, This can represent the third weight. The third weight can be pre-stored in the vehicle, obtained through third-party experimental data, or set by the user; no restrictions are imposed here.
[0108] As an feasible approach, the vehicle can also obtain a third loss based on the rendered semantic map and the initial semantic map corresponding to the initial surround view image. Specifically, the vehicle can substitute the rendered semantic map MD' and the initial semantic map MD corresponding to the initial surround view image into the third loss calculation formula to obtain the third loss. The third loss calculation formula is as follows: L sem = CrossEntropy (MD, MD´) in, The fourth weight can be represented, and CrossEntropy can represent the cross-entropy loss. The fourth weight can be pre-stored in the vehicle, obtained through third-party experimental data, or set by the user; no restrictions are imposed here.
[0109] As an feasible approach, the vehicle can also obtain a fourth loss based on the predicted optical flow and the initial optical flow corresponding to the initial surround view image. Specifically, the vehicle can substitute the predicted optical flow F' and the initial optical flow F corresponding to the initial surround view image into the fourth loss calculation formula to obtain the fourth loss. The fourth loss calculation formula is as follows: L flow = ·||FF´||1 in, It can represent the fifth weight; the fifth weight can be pre-stored in the vehicle, obtained through third-party experimental data, or set by the user, and is not limited here.
[0110] As an feasible approach, the vehicle can also obtain a fifth loss based on the rendered motion flow, the rendered depth map, and the initial depth map. This fifth loss can be understood as a combination of regularization loss and smoothing loss, thereby improving the stability of 3D reconstruction by combining the two.
[0111] Among them, the vehicle can obtain the data collection time t k The rendered motion flow corresponding to adjacent time moments F k+1 as well as F k-1 It can also distinguish the difference between the rendered depth map, the predicted depth corresponding to the rendered depth map, and the true depth corresponding to the initial depth map. D、 F k+1 as well as F k-1 Substituting into the formula for calculating the fifth loss, we obtain the fifth loss. The formula for calculating the fifth loss is as follows: L reg = ·|| F k+1 + F k-1 ||1+ ·|| D||1 in, It can represent the sixth weight. The seventh weight can be represented; the sixth and seventh weights can be pre-stored in the vehicle, obtained through third-party experimental data, or set by the user, which is not limited here.
[0112] In some implementations, after the vehicle acquires the first loss, second loss, third loss, fourth loss, and fifth loss, it can determine the target loss based on these losses. Optionally, the vehicle can obtain the minimum, maximum, average, sum, or arithmetic mean of the first, second, third, fourth, and fifth losses as the target loss. For example, the target loss L = L... rgb +L depth +L sem +L flow +L reg .
[0113] Step S211: Based on the target loss, iteratively update the parameters in the initial four-dimensional Gaussian and the initial motion field to obtain the target four-dimensional Gaussian and the target motion field.
[0114] In some implementations, after acquiring the target loss, the vehicle can iteratively update the parameters in the initial four-dimensional Gaussian and the initial motion field based on the target loss to obtain the target four-dimensional Gaussian and the target motion field. Specifically, the vehicle can optimize the initial four-dimensional Gaussian and the initial motion field by minimizing the target loss through gradient descent to obtain the target four-dimensional Gaussian and the motion field. This allows for the introduction of different loss functions to characterize the error between the rendered result and the true value.
[0115] For example, please refer to the following: Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 as well as Figure 12 .in, Figure 3 A flowchart illustrating a three-dimensional scene reconstruction method provided in an embodiment of this application is shown; Figure 4 An initial surround view image provided in an embodiment of this application is shown; Figure 5 This illustration shows a rendered surround view image provided in one embodiment of this application; Figure 6 An initial depth map provided in one embodiment of this application is shown; Figure 7 This application shows a rendered depth map provided in one embodiment. Figure 8 An initial semantic graph provided in one embodiment of this application is shown; Figure 9 This application shows a rendered semantic map provided in one embodiment. Figure 10 A schematic diagram of the initial optical flow provided in an embodiment of this application is shown; Figure 11 A schematic diagram of the predicted optical flow provided in an embodiment of this application is shown; Figure 12 This illustration shows a schematic diagram of a target 3D scene represented by point cloud form according to an embodiment of this application.
[0116] After the vehicle acquires an initial surround view image (which may include multiple frames and multiple viewpoints), it can input the initial surround view image into a depth estimation model to obtain an initial depth map corresponding to the initial surround view image output by the depth estimation model. Alternatively, it can output the initial surround view image into a semantic segmentation model to obtain an initial semantic map corresponding to the initial surround view image output by the semantic segmentation model. It can also input the initial surround view image into an optical flow estimation model to obtain the initial optical flow corresponding to the initial surround view image output by the optical flow estimation model.
[0117] The vehicle can acquire an initial pseudo-point cloud based on the acquisition time of the initial surround view image, its position information at that acquisition time, and the corresponding camera pose and camera intrinsics. After acquiring the initial pseudo-point cloud, the vehicle can acquire an initial four-dimensional Gaussian plane based on the initial pseudo-point cloud and the initial semantic map. The vehicle can also acquire an initial motion field based on its position information, camera pose, camera intrinsics, and the acquisition time of the initial surround view image. Furthermore, the vehicle can acquire the initial motion flow corresponding to the acquisition time based on the initial motion field, and obtain rendering results based on the initial four-dimensional Gaussian plane and the initial motion flow, such as the rendered surround view image, rendered depth map, rendered semantic map, and rendered motion flow. Finally, the vehicle can acquire predicted optical flow based on the rendered motion flow, rendered depth map, camera pose, and camera intrinsics.
[0118] Among them, the vehicle can obtain the target loss based on the initial surround view image, the rendered surround view image, the rendered depth map, the rendered semantic map, the rendered motion flow, and the predicted optical flow; and can optimize the initial four-dimensional Gaussian and the initial motion field through gradient descent target loss to obtain the target four-dimensional Gaussian and the target motion field, and can reconstruct the target three-dimensional scene based on the target four-dimensional Gaussian and the target motion field.
[0119] The vehicle can acquire multiple frames of multi-view images I (in this embodiment, this can be understood as the initial surround view image, such as...) through data acquisition. Figure 4 As shown in the figure), and the corresponding camera pose W, camera intrinsic parameters K, and acquisition time timestamp t; then the acquired images are input into the depth estimation model, semantic segmentation model, and optical flow estimation model respectively to obtain the initial depth map D (as shown in the figure). Figure 6 As shown), the initial semantic graph M (as shown) Figure 8 (as shown) and initial optical flow F (as shown) Figure 10(As shown). Then, the vehicle can generate dynamic masks MD and static masks MS based on the category IDs of dynamic objects (cars, pedestrians, bicycles, etc.) and static objects (roads, buildings, trees, etc.). It can also combine the pixel coordinates X of the initial panoramic image and the initial depth map D, and use the camera pose W and camera intrinsic parameters K to perform back projection to obtain a pseudo-point cloud in the world coordinate system. The pseudo-point cloud can be uniformly sampled at intervals ts to obtain the sampled initial pseudo-point cloud P. Furthermore, for the pseudo-point cloud P1 of the first frame, the vehicle can use the static mask MS1 and the dynamic mask MD1 to generate static pseudo-point cloud PS1 and dynamic pseudo-point cloud PD1, respectively. For the pseudo-point clouds of other frames obtained by sampling, the corresponding static masks are used to generate the corresponding static pseudo-point clouds to obtain the overall static pseudo-point cloud PS. In the process of expanding 3D Gaussian to 4D Gaussian, the pseudo-point clouds PS and PD1 can be used to initialize the center point, color, and semantics of 4D Gaussian G to obtain the initial four-dimensional Gaussian. Furthermore, the vehicle can use a deep learning model to define the motion field. Based on the maximum camera depth dmax, the vehicle's position information, camera pose W, and camera intrinsic parameters K, it can calculate the minimum and maximum values of the camera's view frustum space range, xyz_min and xyz_max. This range can be initialized by combining the minimum and maximum values t_min and t_max of the initial panoramic image acquisition time t, thus obtaining the initial motion field Φ. Then, the vehicle can obtain the acquisition time t based on the initial motion field. k Corresponding initial motion flow μ, and t can be obtained from the initial motion flow and the initial four-dimensional Gaussian. k 4D Gaussian G(t) to be rendered at any given moment k ), and can be used to generate 4D Gaussian G(t) k ) Render to the camera viewpoint corresponding to the acquisition time to obtain the rendered panoramic image I´ (e.g. Figure 5 As shown), the rendered depth map D´ (as shown) Figure 7 As shown), the rendered semantic graph MD' (such as Figure 9 (as shown) and motion flow F. Among these, the vehicle can also obtain the predicted optical flow F´ (e.g., based on the rendered motion flow and the rendered depth map) Figure 11 As shown). The vehicle can also determine the target loss based on the rendering results, predicted optical flow, and initial surround view image. By minimizing the target loss, it optimizes the initial 4D Gaussian G and the initial motion field Φ to obtain the target 4D Gaussian and the target motion field for 3D scene reconstruction, thus obtaining the target 3D scene (e.g., Figure 12(As shown). This allows the use of a general deep learning model to generate pseudo-labels, reducing reliance on manual annotation and lowering labor costs. Furthermore, it enables dynamic 3D reconstruction of autonomous driving scenarios without the need for LiDAR point clouds, reducing hardware costs and expanding application scenarios. Additionally, the 3D Gaussian-based method for 3D reconstruction improves the accuracy and rendering speed of 3D scene reconstruction, enhancing the practicality of 3D scene creation.
[0120] Step S212: Construct a target three-dimensional scene based on the target four-dimensional Gaussian and the target motion field.
[0121] For a description of step S212, please refer to the previous description of step S150, which will not be repeated here.
[0122] The three-dimensional scene reconstruction method provided in one embodiment of this application is compared with... Figure 1 The 3D scene reconstruction method shown in this embodiment can also obtain the predicted optical flow based on the rendered depth map and the rendered motion flow; obtain the target loss based on the predicted optical flow, the rendering result and the initial panoramic image; and iteratively update the parameters in the initial four-dimensional Gaussian and the initial motion field based on the target loss to obtain the target four-dimensional Gaussian and the target motion field, thereby combining the dynamic scene to reconstruct the 3D scene and expanding the application scenarios of 3D scene reconstruction.
[0123] Meanwhile, this embodiment can also determine the initial motion flow corresponding to the acquisition time based on the initial motion field; determine the four-dimensional Gaussian to be rendered corresponding to the acquisition time based on the initial motion flow and the initial four-dimensional Gaussian; render the four-dimensional Gaussian to be rendered to obtain the rendered panoramic image, the rendered depth map, the rendered semantic map, and the rendered motion flow, thereby calculating the motion flow based on the motion field, and in the case of introducing the motion flow, performing four-dimensional Gaussian rendering to obtain multiple rendering results including dynamic scenes, so as to perform three-dimensional reconstruction based on multiple rendering results, thereby improving the accuracy of three-dimensional reconstruction.
[0124] Furthermore, this embodiment can also obtain a first mask corresponding to dynamic objects in the initial surround view image and a second mask corresponding to static objects in the initial surround view image based on the pixel categories in the initial semantic map corresponding to the initial surround view image; obtain an initial pseudo-point cloud based on the pixel coordinates, camera pose, camera intrinsic parameters, and the initial depth map corresponding to the initial surround view image; obtain a second pseudo-point cloud and a third pseudo-point cloud based on the initial pseudo-point cloud, the first mask, and the second mask, wherein the second pseudo-point cloud is determined based on the first frame pseudo-point cloud in the initial pseudo-point cloud and the first mask corresponding to the first frame pseudo-point cloud, and the third pseudo-point cloud is determined based on the second mask corresponding to the initial pseudo-point cloud; obtain an initial four-dimensional Gaussian based on the second pseudo-point cloud and the third pseudo-point cloud, thereby filtering dynamic objects during the definition of the four-dimensional Gaussian, avoiding the influence of dynamic objects on the creation of the four-dimensional Gaussian, and improving the accuracy of subsequent three-dimensional scene creation by combining the motion field and the four-dimensional Gaussian.
[0125] Please see Figure 13 , Figure 13 A block diagram of a three-dimensional scene reconstruction apparatus according to an embodiment of this application is shown. This three-dimensional scene reconstruction apparatus 200 is applied to the aforementioned electronic device, and will be discussed below. Figure 7 The process is described in detail below. The 3D scene reconstruction device 200 includes: a four-dimensional Gaussian initialization module 210, a motion field initialization module 220, a rendering result acquisition module 230, a four-dimensional Gaussian and motion field adjustment module 240, and a 3D scene reconstruction module 250, wherein: The four-dimensional Gaussian initialization module 210 is used to obtain an initial four-dimensional Gaussian based on the initial surround view image captured by the vehicle's camera, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters.
[0126] The motion field initialization module 220 is used to obtain an initial motion field based on the vehicle's position information, the camera pose, the camera intrinsic parameters, and the acquisition time of the initial surround view image. The spatial range of the initial motion field is determined based on the position information, the camera pose, and the camera intrinsic parameters, and the temporal range of the initial motion field is determined based on the acquisition time.
[0127] The rendering result acquisition module 230 is used to acquire the rendering result corresponding to the acquisition time based on the initial four-dimensional Gaussian and the initial motion field. The rendering result includes a rendered panoramic image, a rendered depth map, a rendered semantic map, and a rendered motion flow. The rendered motion flow is used to characterize the displacement of pixels in the rendered panoramic image in the initial motion field at the acquisition time.
[0128] The 4D Gaussian and motion field adjustment module 240 is used to adjust the parameters in the initial 4D Gaussian and the initial motion field according to the rendering result and the initial panoramic image to obtain the target 4D Gaussian and the target motion field.
[0129] The 3D scene reconstruction module 250 is used to construct a target 3D scene based on the target 4D Gaussian and the target motion field.
[0130] Furthermore, the four-dimensional Gaussian and motion field adjustment module 240 may include: a predicted optical flow acquisition unit, a target loss acquisition unit, and a four-dimensional Gaussian and motion field adjustment subunit, wherein: The predicted optical flow acquisition unit is used to acquire the predicted optical flow based on the rendered depth map and the rendered motion flow.
[0131] The target loss acquisition unit is used to acquire the target loss based on the predicted optical flow, the rendering result, and the initial panoramic image.
[0132] The four-dimensional Gaussian and motion field adjustment subunit is used to iteratively update the parameters in the initial four-dimensional Gaussian and the initial motion field according to the target loss, so as to obtain the target four-dimensional Gaussian and the target motion field.
[0133] Further, the target loss acquisition unit may include: a first loss acquisition unit, a second loss acquisition unit, a third loss acquisition unit, a fourth loss acquisition unit, a fifth loss acquisition unit, and a target loss acquisition subunit, wherein: The first loss acquisition unit is used to acquire a first loss based on the rendered panoramic image and the initial panoramic image.
[0134] The second loss acquisition unit is used to acquire a second loss based on the rendered depth map and the initial depth map corresponding to the initial surround view image.
[0135] The third loss acquisition unit is used to acquire the third loss based on the rendered semantic graph and the initial semantic graph corresponding to the initial loop image.
[0136] The fourth loss acquisition unit is used to acquire the fourth loss based on the predicted optical flow and the initial optical flow corresponding to the initial toroidal image.
[0137] The fifth loss acquisition unit is used to acquire the fifth loss based on the rendered motion flow, the rendered depth map, and the initial depth map.
[0138] The target loss acquisition subunit is used to determine the target loss based on the first loss, the second loss, the third loss, the fourth loss, and the fifth loss.
[0139] Furthermore, the predicted optical flow acquisition unit may include: a first pseudo-point cloud acquisition unit and a predicted optical flow acquisition subunit, wherein: The first pseudo-point cloud acquisition unit is used to back-project the rendered depth map to the world coordinate system to obtain the first pseudo-point cloud.
[0140] The predicted optical flow acquisition subunit is used to acquire the predicted optical flow of the next moment corresponding to the acquisition moment based on the rendered motion flow and the first pseudo point cloud.
[0141] Furthermore, the rendering result acquisition module 230 may include: an initial motion flow determination unit, a four-dimensional Gaussian determination unit to be rendered, and a rendering result acquisition subunit, wherein: An initial motion flow determination unit is used to determine the initial motion flow corresponding to the acquisition time based on the initial motion field.
[0142] The four-dimensional Gaussian to be rendered determination unit is used to determine the four-dimensional Gaussian to be rendered corresponding to the acquisition time based on the initial motion flow and the initial four-dimensional Gaussian.
[0143] The rendering result acquisition sub-unit is used to render the four-dimensional Gaussian to be rendered, and obtain the rendered panoramic image, the rendered depth map, the rendered semantic map, and the rendered motion flow.
[0144] Further, the four-dimensional Gaussian initialization module 210 may include: a mask generation unit, an initial pseudo-point cloud acquisition unit, a second pseudo-point cloud acquisition unit, and a four-dimensional Gaussian initialization subunit, wherein: The mask generation unit is used to obtain a first mask corresponding to dynamic objects in the initial surround view image and a second mask corresponding to static objects in the initial surround view image based on the pixel categories in the initial semantic map corresponding to the initial surround view image.
[0145] The initial pseudo-point cloud acquisition unit is used to acquire an initial pseudo-point cloud based on the pixel coordinates in the initial panoramic image, the camera pose, the camera intrinsic parameters, and the initial depth map corresponding to the initial panoramic image.
[0146] The second pseudo-point cloud acquisition unit is used to acquire a second pseudo-point cloud and a third pseudo-point cloud based on the initial pseudo-point cloud, the first mask, and the second mask. The second pseudo-point cloud is determined based on the first frame pseudo-point cloud in the initial pseudo-point cloud and the first mask corresponding to the first frame pseudo-point cloud. The third pseudo-point cloud is determined based on the second mask corresponding to the initial pseudo-point cloud.
[0147] The four-dimensional Gaussian initialization subunit is used to obtain the initial four-dimensional Gaussian based on the second pseudo-point cloud and the third pseudo-point cloud.
[0148] Furthermore, the initial pseudo-point cloud acquisition unit may include: a pseudo-point cloud acquisition unit in the camera coordinate system, a pseudo-point cloud acquisition unit in the world coordinate system, and a pseudo-point cloud interval sampling unit, wherein: The pseudo-point cloud acquisition unit in the camera coordinate system is used to back-project the pixel coordinates in the initial panoramic image and the initial depth map based on the camera intrinsic parameters to obtain the pseudo-point cloud in the camera coordinate system.
[0149] The pseudo-point cloud acquisition unit in the world coordinate system is used to transform the pseudo-point cloud in the camera coordinate system to the world coordinate system based on the camera pose, so as to obtain the pseudo-point cloud in the world coordinate system.
[0150] The pseudo-point cloud interval sampling unit is used to sample the pseudo-point cloud in the world coordinate system based on a preset sampling interval to obtain the initial pseudo-point cloud.
[0151] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0152] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0153] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0154] Please see Figure 14 This document illustrates a structural block diagram of an electronic device according to an embodiment of this application. The electronic device 100 can be a vehicle, in-vehicle terminal, server, computer, or other device with processing capabilities. The electronic device 100 in this application may include one or more of the following components: a processor 110, a memory 120, and one or more application programs. The one or more application programs may be stored in the memory 120 and configured to be executed by one or more processors 110. The one or more programs are configured to perform the methods described in the foregoing method embodiments.
[0155] The processor 110 may include one or more processing cores. The processor 110 connects to various parts of the vehicle 100 via various interfaces and lines, and performs various functions and processes data of the vehicle 100 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 110 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 110 and may be implemented separately through a communication chip.
[0156] The memory 120 may include random access memory (RAM) or read-only memory (ROM). The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).
[0157] In this embodiment, a computer-readable medium stores program code, which can be called by a processor to execute the methods described in the above method embodiments.
[0158] Computer-readable storage media can be electronic storage devices such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, computer-readable storage media includes non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for program code that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code can be compressed, for example, in a suitable form.
[0159] In this application, "multiple" refers to two or more.
[0160] In this application, unless otherwise expressly defined, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0161] The terms “first,” “second,” “third,” “fourth,” etc., in this application (if present) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0162] In this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, in this application, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0163] Unless otherwise specified, all steps in this application may be performed sequentially or randomly. For example, if the method includes steps A and B, it means that the method may include steps A and B performed sequentially, or it may include steps B and A performed sequentially. For example, if the method may also include step C, it means that step C may be added to the method in any order. For example, the method may include steps A, B, and C, or it may include steps A, C, and B, or it may include steps C, A, and B, etc.
[0164] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for reconstructing a three-dimensional scene, characterized in that, The method includes: Based on the initial surround view image captured by the vehicle's camera, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters, an initial four-dimensional Gaussian is obtained. Based on the vehicle's location information, the camera pose, the camera intrinsic parameters, and the acquisition time of the initial surround view image, an initial motion field is obtained. The spatial range of the initial motion field is determined based on the location information, the camera pose, and the camera intrinsic parameters, and the temporal range of the initial motion field is determined based on the acquisition time. Based on the initial four-dimensional Gaussian and the initial motion field, the rendering result corresponding to the acquisition time is obtained. The rendering result includes a rendered panoramic image, a rendered depth map, a rendered semantic map, and a rendered motion flow. The rendered motion flow is used to characterize the displacement of pixels in the rendered panoramic image in the initial motion field at the acquisition time. Based on the rendering results and the initial toroidal image, the parameters in the initial four-dimensional Gaussian and the initial motion field are adjusted to obtain the target four-dimensional Gaussian and the target motion field. Construct a target three-dimensional scene based on the target four-dimensional Gaussian and the target motion field.
2. The method according to claim 1, characterized in that, The step of adjusting the parameters in the initial four-dimensional Gaussian and the initial motion field based on the rendering result and the initial toroidal image to obtain the target four-dimensional Gaussian and the target motion field includes: Based on the rendered depth map and the rendered motion flow, the predicted optical flow is obtained; Based on the predicted optical flow, the rendering result, and the initial surround view image, the target loss is obtained; Based on the target loss, the parameters in the initial four-dimensional Gaussian and the initial motion field are iteratively updated to obtain the target four-dimensional Gaussian and the target motion field.
3. The method according to claim 2, characterized in that, The step of obtaining the target loss based on the predicted optical flow, the rendering result, and the initial surround view image includes: Based on the rendered toroidal image and the initial toroidal image, obtain the first loss; Based on the rendered depth map and the initial depth map corresponding to the initial ring view image, the second loss is obtained; Based on the rendered semantic map and the initial semantic map corresponding to the initial loop image, a third loss is obtained; Based on the predicted optical flow and the initial optical flow corresponding to the initial toroidal image, a fourth loss is obtained; Based on the rendered motion flow, the rendered depth map, and the initial depth map, a fifth loss is obtained; The target loss is determined based on the first loss, the second loss, the third loss, the fourth loss, and the fifth loss.
4. The method according to claim 2, characterized in that, The step of obtaining the predicted optical flow based on the rendered depth map and the rendered motion flow includes: The rendered depth map is back-projected onto the world coordinate system to obtain the first pseudo-point cloud; Based on the rendered motion flow and the first pseudo-point cloud, the predicted optical flow for the next moment corresponding to the acquisition moment is obtained.
5. The method according to claim 1, characterized in that, The step of obtaining the rendering result corresponding to the acquisition time based on the initial four-dimensional Gaussian and the initial motion field includes: The initial motion flow corresponding to the acquisition time is determined based on the initial motion field; Based on the initial motion flow and the initial four-dimensional Gaussian, determine the four-dimensional Gaussian to be rendered corresponding to the acquisition time; The four-dimensional Gaussian image to be rendered is rendered to obtain the rendered panoramic image, the rendered depth map, the rendered semantic map, and the rendered motion flow.
6. The method according to any one of claims 1-5, characterized in that, The process of obtaining an initial four-dimensional Gaussian image based on the initial surround view image captured by the vehicle's camera, the camera pose corresponding to the initial surround view image, and the camera intrinsic parameters includes: Based on the pixel categories in the initial semantic map corresponding to the initial panoramic image, obtain the first mask corresponding to the dynamic object in the initial panoramic image and the second mask corresponding to the static object in the initial panoramic image; An initial pseudo-point cloud is obtained based on the pixel coordinates in the initial panoramic image, the camera pose, the camera intrinsic parameters, and the initial depth map corresponding to the initial panoramic image; Based on the initial pseudo point cloud, the first mask, and the second mask, a second pseudo point cloud and a third pseudo point cloud are obtained. The second pseudo point cloud is determined based on the first frame pseudo point cloud in the initial pseudo point cloud and the first mask corresponding to the first frame pseudo point cloud. The third pseudo point cloud is determined based on the second mask corresponding to the initial pseudo point cloud. The initial four-dimensional Gaussian is obtained based on the second pseudo-point cloud and the third pseudo-point cloud.
7. The method according to claim 6, characterized in that, The step of obtaining the initial pseudo-point cloud based on the pixel coordinates in the initial panoramic image, the camera pose, the camera intrinsic parameters, and the initial depth map corresponding to the initial panoramic image includes: Based on the camera intrinsic parameters, the pixel coordinates in the initial panoramic image and the initial depth map are back-projected to obtain a pseudo point cloud in the camera coordinate system. Based on the camera pose, the pseudo point cloud in the camera coordinate system is transformed to the world coordinate system to obtain the pseudo point cloud in the world coordinate system. The initial pseudo-point cloud is obtained by sampling the pseudo-point cloud in the world coordinate system based on a preset sampling interval.
8. An electronic device, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the method as described in any one of claims 1-7.
9. A vehicle, characterized in that, It includes the electronic device as described in claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium contains program code that can be invoked by a processor to execute the method as described in any one of claims 1-7.