Automatic driving scene reconstruction method, system and device based on space-time consistency constraint and storage medium
By constructing a Gaussian ellipsoid model based on spatiotemporal consistency constraints, the problems of high-frequency temporal consistency and spatial high-frequency structural degradation in autonomous driving scenario reconstruction were solved, achieving high-precision reconstruction of dynamic urban scenes and improving the accuracy and robustness of reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-03
- Publication Date
- 2026-05-12
AI Technical Summary
Existing autonomous driving scene reconstruction methods suffer from insufficient high-frequency temporal consistency and spatial high-frequency structure degradation in dynamic urban scenes. They are unable to stably characterize rapidly changing behaviors such as vehicle light flashing and vehicle trajectory, and are difficult to accurately preserve fine-grained geometric structures under sparse perspectives or long-distance observation.
By constructing a Gaussian ellipsoid model based on spatiotemporal consistency constraints, utilizing four-dimensional spatiotemporal coordinates and multi-resolution hash grid encoding, and combining high-frequency temporal embedding and spatial embedding, the properties of the Gaussian ellipsoid are optimized to achieve high-precision reconstruction of dynamic scenes.
It improves the accuracy and robustness of autonomous driving scene reconstruction, solves the problems of light flickering and discontinuous motion trajectory in dynamic scenes, and preserves fine-grained geometric structure under sparse viewpoints, thereby improving the quality and consistency of rendered images.
Smart Images

Figure CN122023670A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision and autonomous driving scene reconstruction technology, specifically to autonomous driving scene reconstruction methods, systems, devices, and storage media based on spatiotemporal consistency constraints. Background Technology
[0002] 3D scene reconstruction and novel view synthesis are core technologies for data generation and scene modeling in the field of autonomous driving. Their goal is to construct 3D representations with geometric consistency and visual realism under limited multi-view observation conditions, and to support the generation of high-quality rendered images at arbitrary camera poses. This serves downstream tasks of autonomous driving systems such as high-precision map maintenance, battery-electric perception (BEV perception), overall scene understanding, and 3D object detection. In recent years, Neural Radiance Fields (NeRF) and its derivatives have achieved good results in novel view synthesis of static scenes, generating relatively realistic rendered images from different perspectives. However, NeRF-based methods typically have high rendering overhead, making it difficult to meet the real-time or near-real-time rendering requirements of autonomous driving scenarios, especially in complex urban environments where deployment costs are high. To improve rendering efficiency, 3D Gaussian Splatting (3DGS) has been proposed as an explicit 3D representation method. It achieves real-time rendering by rasterizing the projection of anisotropic Gaussians, significantly reducing rendering costs and demonstrating application potential in real-time novel view synthesis and reconstruction tasks for autonomous driving scenarios.
[0003] Despite the progress made by 3DGS-based methods in terms of rendering speed and scene representation, existing technologies still have the following key problems when dealing with realistic dynamic urban scenes, especially sudden motion scenes such as lane changes and vehicle headlight flashing: High-frequency temporal consistency issues: Existing 3D scene reconstruction and new view synthesis methods typically lack the ability to explicitly model high-frequency temporal signals in dynamic driving scenarios. This makes it difficult to stably characterize periodic flashing of turn signals, sudden on / off of brake lights, and rapid changes in vehicle behavior such as acceleration, deceleration, or lane changes, leading to the continuous accumulation of frame-by-frame errors over time. The result is inconsistent reconstruction results for dynamic regions between adjacent frames, exhibiting problems such as difficulty in consistently reproducing lighting states and jittering or breaking motion trajectories. This disrupts the overall temporal continuity, making it difficult to meet the temporal consistency requirements of autonomous driving simulation and evaluation applications.
[0004] The degradation of high-frequency spatial structures: Existing methods, under sparse viewpoints or long-distance observation conditions, generally rely on pixel-level loss to supervise the scene, lacking effective constraints on high-frequency spatial structures. This results in the inaccurate preservation of fine-grained geometric structures such as lane lines, curb boundaries, distant lane divisions, and building facade textures. The result is blurred fine lines, blunted edge contours, and missing texture details in the rendered image, weakening the geometric representation and structural recognition of the scene, and thus affecting the reliability and practicality of downstream autonomous driving perception tasks. Summary of the Invention
[0005] The purpose of this invention is to provide a method, system, device and storage medium for autonomous driving scene reconstruction based on spatiotemporal consistency constraints, so as to solve the technical problems of insufficient high-frequency temporal consistency and spatial high-frequency structural degradation in existing autonomous driving 3D scene reconstruction results.
[0006] In a first aspect, the present invention provides an autonomous driving scene reconstruction method based on spatiotemporal consistency constraints, comprising the following steps: Acquire multi-source data of dynamic urban scenes, establish multiple Gaussian ellipsoids based on the multi-source data of dynamic urban scenes, and construct a Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints; The three-dimensional spatial coordinates of each Gaussian ellipsoid are combined with the time dimension to form a four-dimensional spatiotemporal coordinate. Based on the four-dimensional spatiotemporal coordinate, a spatial embedding vector and a high-frequency temporal embedding are constructed, and an instance embedding vector is constructed based on multi-source data of dynamic urban scenes. A joint feature vector is obtained based on the spatial embedding vector, high-frequency temporal embedding, and instance embedding vector; a temporal high-frequency consistency constraint is constructed based on the joint feature vector, and the properties of the Gaussian ellipsoid are optimized based on the temporal high-frequency consistency constraint. Construct spatial high-frequency consistency constraints and optimize the properties of static Gaussian ellipsoids; A Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints is trained using the optimized Gaussian ellipsoid. The trained Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints is used to render multi-source data of dynamic urban scenes.
[0007] The significant advantages of this invention are as follows: By constructing a spatial embedding vector of a Gaussian ellipsoid, this invention effectively solves the technical problem in existing technologies that Gaussian points rely solely on static spatial location encoding and lack dynamic spatiotemporal context modeling capabilities. Compared to the limitations of existing technologies that can only represent static geometric positions, the spatial embedding vector generated by this method possesses stronger spatiotemporal context representation capabilities, providing more accurate geometric and dynamic priors for subsequent scene reconstruction, and significantly improving the accuracy and robustness of autonomous driving scene reconstruction.
[0008] Further, in step S2, the step of constructing a spatial embedding vector based on four-dimensional spatiotemporal coordinates includes: Construct four-dimensional spatiotemporal coordinates based on the spatial and temporal positions of each Gaussian point in the scene containing the Gaussian ellipsoid; Calculate the four-dimensional resolution vector of the four-dimensional spatiotemporal coordinates at each resolution level; Calculate the fractional offset of each resolution coordinate in the four-dimensional resolution vector within a grid point in the hash grid; Spatiotemporal features of each resolution level are generated based on the fractional offset of each resolution coordinate in the four-dimensional resolution vector. The spatial embedding vector of the Gaussian ellipsoid is obtained by concatenating the spatiotemporal features at all resolution levels.
[0009] By constructing four-dimensional spatiotemporal coordinates, the basic modeling of the dynamic properties of Gaussian points is realized, breaking the limitation of traditional static 3D Gaussian Splatting which is based solely on spatial coordinates. By calculating the four-dimensional resolution vector and the fractional offset within the grid at multiple resolution levels, and combining it with a hash grid, the fine-grained encoding of multi-resolution features was achieved, thus achieving the technical goal of improving feature representation ability and generalization. Based on fractional offset, spatiotemporal features of various resolutions are generated and stitched together, realizing the step-by-step fusion from low-dimensional smooth features to high-dimensional fine features. The resulting Gaussian ellipsoid spatial embedding vector fully covers spatial location and time-varying information.
[0010] Furthermore, when constructing the high-frequency time embedding based on the four-dimensional spatiotemporal coordinates, one or more time frequency components are generated based on the four-dimensional spatiotemporal coordinates, and all time frequency components are accumulated to obtain the high-frequency time embedding.
[0011] By constructing high-frequency temporal embeddings through the accumulation of multiple frequency components, and covering the entire dynamic spectrum with an exponentially spaced frequency base, high-frequency dynamic details in autonomous driving scenarios are accurately captured, solving the problem that traditional temporal coding cannot adapt to complex temporal changes. The generated high-frequency temporal embeddings provide accurate temporal priors for subsequent feature fusion and deformation prediction, effectively suppressing flickering and drift issues in dynamic reconstruction, and improving the temporal consistency and visual realism of scene reconstruction. While ensuring feature representation capabilities, computational complexity is controlled, achieving a balance between reconstruction accuracy and real-time performance, thus meeting the engineering requirements of autonomous driving scenarios.
[0012] Further, the steps of obtaining a joint feature vector based on the spatial embedding vector, high-frequency temporal embedding, and instance embedding vector; constructing a temporal high-frequency consistency constraint based on the joint feature vector; and optimizing the Gaussian ellipsoid properties based on the temporal high-frequency consistency constraint include: The joint feature vector is obtained by concatenating the high-frequency temporal embedding, spatial embedding vector, and instance embedding vector. By inputting the joint feature vector into the Gaussian deformation network, the property residuals of the Gaussian ellipsoid in the time dimension are predicted. The center position, scale, and rotation parameters of the Gaussian ellipsoid are updated recursively over time based on the residuals of the Gaussian ellipsoid's properties.
[0013] By deeply fusing high-frequency temporal embeddings, spatial embedding vectors, and instance embedding vectors through feature concatenation, we can comprehensively represent temporal dynamics, spatial geometry, and instance identity information, achieve collaborative modeling of multi-dimensional features, and improve the richness and completeness of the representation of joint feature vectors.
[0014] By using Gaussian deformation network prediction and relying on the strong representational ability of joint feature vectors, the property residuals of Gaussian ellipsoids in the time dimension are accurately predicted, achieving high-precision fitting of the pose, position and scale changes of dynamic objects in autonomous driving scenarios, and significantly improving the detail accuracy of dynamic reconstruction.
[0015] By using time-recursive updates, the center position, scale, and rotation parameters of the Gaussian ellipsoid are updated in a temporal manner based on attribute residuals, establishing a continuous evolution relationship between frames. This effectively solves the problems of flickering, misalignment, and breakage that occur in dynamic reconstruction, ensuring the temporal consistency and visual smoothness of the reconstruction results.
[0016] Furthermore, the step of constructing spatial high-frequency consistency constraints and optimizing the properties of the static Gaussian ellipsoid includes: Obtain the Gaussian ellipsoid of the static region and project the Gaussian ellipsoid of the static region onto the image plane at different viewpoints to obtain the static region rendering image; Low-frequency components are obtained by bilateral filtering of real images captured by vehicles, and image gradient information is calculated based on real images captured by vehicles. A weight map for emphasizing high-frequency spatial structures is constructed by combining the image gradient information. A structural similarity loss function is constructed based on real images of vehicles and static region rendered images. Construct a multi-view spatial high-frequency consistency loss function; Optimize the properties of static Gaussian ellipsoids based on multi-view spatial high-frequency consistency loss function and structural similarity loss function.
[0017] By projecting and rendering Gaussian spheres in static regions, independent modeling and accurate rendering of static backgrounds in autonomous driving scenarios are achieved, providing a high-quality reference benchmark for subsequent loss calculations, effectively distinguishing between dynamic and static regions, and reducing the interference of dynamic foregrounds on the reconstruction of static backgrounds.
[0018] By constructing a weighted graph based on image gradient information, and using the high-frequency structural features of the image gradient, it can adaptively emphasize important spatial regions such as lane lines, vehicle outlines, and traffic signs. At the same time, by smoothing low-frequency noise through bilateral filtering, it achieves the technical goal of noise reduction and edge preservation, laying the foundation for accurate calculation of detail loss.
[0019] A high-frequency consistency loss in multi-view space is constructed. By averaging the losses of all views and fusing edge-aware loss and structural loss, the structural consistency and high-frequency detail alignment of static scene reconstruction under different views are guaranteed. This solves the problems of texture fragmentation and view misalignment that are prone to occur in multi-view rendering, and improves the global coherence and engineering practicality of multi-view scene reconstruction.
[0020] Furthermore, the step of training the Gaussian sputtering model based on the optimized Gaussian ellipsoid and spatiotemporal high-frequency consistency constraints includes: The rendered image of the autonomous driving scene is obtained based on the optimized Gaussian ellipsoid rendering. Construct a joint objective function based on rendered images from autonomous driving scenarios; The parameters of the Gaussian ellipsoid and the model are optimized by using a joint objective function.
[0021] By rendering the optimized Gaussian ellipsoid, a high-quality image synthesis of autonomous driving scenes was achieved, providing a benchmark for reconstruction results that conform to real visual perception for subsequent joint optimization.
[0022] Secondly, the present invention provides an autonomous driving scene reconstruction system based on spatiotemporal consistency constraints, applicable to the aforementioned autonomous driving scene reconstruction method based on spatiotemporal consistency constraints, characterized in that it comprises, in sequence, the following: The input unit is used to input multi-source data of dynamic urban scenes; The Gaussian initialization and scene representation construction unit is used to obtain the initial geometric representation of the dynamic urban scene based on multi-source data of the dynamic urban scene and construct the Gaussian ellipsoid. Spatiotemporal embedding construction unit, used to construct high-frequency temporal embedding, spatial embedding vector and instance embedding vector based on the three-dimensional spatial coordinates and temporal information coordinates of the Gaussian ellipsoid; Spatiotemporal high-frequency consistency constraint unit, used to generate temporal high-frequency consistency constraints and spatial high-frequency consistency constraints; The Gaussian ellipsoid optimization element is used to optimize the properties of a Gaussian ellipsoid based on temporal high-frequency consistency constraints and spatial high-frequency consistency constraints. The projection and rasterization rendering unit is used to render the optimized Gaussian ellipsoid to obtain the rendered image of the autonomous driving scene from the corresponding viewpoint.
[0023] Furthermore, the spatiotemporal high-frequency consistency constraint unit includes a temporal high-frequency consistency constraint module and a spatial high-frequency consistency constraint module connected in parallel; The time-high frequency consistency constraint module is used to model the displacement, rotation and scale changes of the Gaussian ellipsoid in the time dimension by fusing high-frequency time embedding, spatial embedding vector and instance embedding vector to generate time-high frequency consistency constraints. The spatial high-frequency consistency constraint module is used to generate spatial high-frequency consistency constraints by constructing an edge-aware weighted loss function to supervise high-frequency structural regions.
[0024] Thirdly, the present invention provides an autonomous driving scene reconstruction device based on spatiotemporal consistency constraints, including a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the above-mentioned autonomous driving scene reconstruction method based on spatiotemporal consistency constraints.
[0025] Fourthly, the present invention provides a computer-readable storage medium containing a computer program, wherein the computer program is stored thereon, characterized in that, when the computer program is executed by one or more processors, it implements the steps of the above-described method for reconstructing autonomous driving scenarios based on spatiotemporal consistency constraints. Attached Figure Description
[0026] Figure 1 This is a flowchart of the autonomous driving scene reconstruction method based on spatiotemporal consistency constraints in an embodiment of the present invention; Figure 2 This is a schematic diagram of the Gaussian deformation network structure in an embodiment of the present invention; Figure 3 This is a comparison diagram of the temporal rendering effect on the dynamic vehicle target of autonomous driving in an embodiment of the present invention; Figure 4 This is a comparison chart of reconstruction effects and accuracy indicators in complex urban scenarios with autonomous driving, as shown in this embodiment of the invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, a clear and complete description will be provided below in conjunction with the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the protection scope of the present invention.
[0028] See appendix Figure 1 The autonomous driving scene reconstruction method based on spatiotemporal consistency constraints shown includes the following steps: S1. Acquire multi-source data of dynamic urban scenes, establish multiple Gaussian ellipsoids based on the multi-source data, and construct a Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints. The multi-source data of dynamic urban scenes includes images and 3D point cloud data of the dynamic urban scene. Specifically, a continuous image sequence of the dynamic urban scene is acquired through a multi-view camera system, and corresponding 3D point cloud information and camera pose are acquired through LiDAR. When constructing the Gaussian ellipsoids, the scene is initialized in 3D space based on the acquired multi-source data of the dynamic urban scene. A set of Gaussian ellipsoids is initialized in the point cloud or scene space, and the attributes of each Gaussian ellipsoid are defined. The attributes of the Gaussian ellipsoid include its spatial position, scale, rotation, color, and transparency, to uniformly represent the static background and dynamic targets of the Gaussian ellipsoid. The specific details of acquiring dynamic urban multi-source data and initializing Gaussian ellipsoids are existing technologies and will not be elaborated here.
[0029] S2. Construct four-dimensional spatiotemporal coordinates by combining the three-dimensional spatial coordinates of each Gaussian ellipsoid with the time dimension. Based on these four-dimensional spatiotemporal coordinates, obtain spatial embedding vectors and high-frequency temporal embeddings, and construct instance embedding vectors. When constructing the spatial embedding vectors and high-frequency temporal embeddings, multi-resolution four-dimensional hash grid encoding is used to jointly embed and encode the spatial and temporal dimensions of the four-dimensional spatiotemporal coordinates. The resolution growth rate of the time dimension is lower than that of the spatial dimension to adapt to the sparse time sampling characteristics in dynamic scenes, thereby obtaining spatial embedding vectors for subsequent Gaussian deformation values. The time signal obtained by normalizing the four-dimensional spatiotemporal coordinates is fitted using a periodic function to model the periodic turn signal, obtaining the corresponding high-frequency temporal embedding. To distinguish sudden signals from different dynamic vehicles, such as brake lights, a set of Gaussian deformations is initialized as instance embedding vectors to distinguish vehicles with on and off brake lights from noise. The construction of the spatial embedding vectors includes the following steps: A1. Construct a four-dimensional spatiotemporal coordinate system, representing the spatial and temporal position of each Gaussian point in the scene containing the Gaussian ellipsoid as a normalized four-dimensional vector: In the formula, It is a four-dimensional spacetime coordinate vector. For a Gaussian point in three-dimensional space Coordinate position on the axis For a Gaussian point in three-dimensional space Coordinate position on the axis For a Gaussian point in three-dimensional space Coordinate position on the axis This refers to the time dimension or frame sequence information corresponding to the Gaussian points. This is a matrix transpose operation.
[0030] The three-dimensional spatial coordinates and time information in the four-dimensional spatiotemporal coordinates are linearly normalized to the interval [0,1].
[0031] A2. Calculate the four-dimensional resolution vector of the four-dimensional spatiotemporal coordinates at each resolution level. Specifically, the four-dimensional resolution vector is calculated using hash grid coordinates. The hash grid coordinate calculation is used to map the normalized four-dimensional spatiotemporal coordinates to a multi-resolution four-dimensional hash grid to support subsequent feature queries and implicit representation learning. The expression for the four-dimensional resolution vector is as follows: In the formula, For the first The four-dimensional resolution vector corresponding to each resolution level For the first At each resolution level, the four-dimensional hash grid is in spatial horizontal... The resolution of the axis, For the first At each resolution level, the four-dimensional hash grid is vertical in space. The resolution of the axis, For the first At each resolution level, the four-dimensional hash grid has a spatial depth The resolution of the axis, For the first At each resolution level, the four-dimensional hash grid in time The resolution of the axis.
[0032] A3. Calculate the fractional offset of each resolution coordinate in the four-dimensional resolution vector within the grid points of the hash grid: In the formula, For the first At each resolution level, the fractional offset of the four-dimensional scaling coordinates within the corresponding grid point of the hash grid. For the first At each resolution level, the normalized four-dimensional spatiotemporal coordinates are scaled to obtain the four-dimensional scaled coordinates. For the first At each resolution level, the hash grid integer grid coordinates corresponding to the four-dimensional scaling coordinates. This is a round-down operation. This refers to the element-wise multiplication operation of a matrix.
[0033] A4. Spatiotemporal features at each resolution level are generated based on the fractional offsets of the coordinates at each resolution in the four-dimensional resolution vector. Specifically, quadlinear interpolation weighting is performed on 16 adjacent grid points within the four-dimensional hypercube in four-dimensional space based on the calculated fractional offsets to generate the spatiotemporal features at each level. The calculation formula is as follows: In the formula, For the first The spatiotemporal characteristics of the layer For a combination of neighborhood indices that take 0 or 1 in four dimensions, For fractional offset, For fractional offset The calculated quadlinear interpolation weights, For the hash table The eigenvalues of the corresponding grid points in the layer.
[0034] The specific steps involved in performing quadlinear interpolation include: Obtain the fractional offset of each coordinate in the four-dimensional resolution vector at each resolution level: In the formula, For the first At each resolution level, the fractional offset is spatially horizontal. Components of the axial dimension, For the first At each resolution level, the fractional offset is perpendicular in space. Components of the axial dimension, For the first At each resolution level, the fractional offset in spatial depth Components of the axial dimension, For the first At each resolution level, the fraction offset over time Components of the axial dimension, For the first At each resolution level, the fractional offset is at the th resolution level. Components in a dimension.
[0035] For any neighborhood corner point: In the formula, This is the offset index vector of the corner points of the neighborhood of the four-dimensional hash grid. This is the offset index vector of the corner points of the neighborhood of the four-dimensional hash grid. The offset index vector is perpendicular in space Components of the axial dimension, For the offset index vector in spatial depth Components of the axial dimension, For the offset index vector in time Components of the axial dimension.
[0036] The four-linear interpolation weights are defined as the product of the one-dimensional linear interpolation weights in each dimension: In the formula, For the first The four-linear interpolation weights corresponding to each adjacent grid point For the four-dimensional grid offset index at the th Dimensional components, For the first At each resolution level, the fractional offset In the Components in a dimension.
[0037] A5. Concatenate the spatiotemporal features from all resolution levels to obtain the spatial embedding vector of the Gaussian ellipsoid: In the formula, For spatial embedding vectors, For concatenation functions, For the first Spatiotemporal features at different resolution levels Let be a real vector space.
[0038] When constructing the high-frequency temporal embedding, one or more time frequency components are generated based on the four-dimensional spatiotemporal coordinates. These time frequency components are angular frequencies of a pre-defined time frequency base distributed at exponential intervals, generated based on the angular frequency formula. Their core function is to construct a high-frequency feature space that can cover the full dynamic range of autonomous driving scenarios. All time frequency components are accumulated to obtain the final high-frequency temporal embedding, which is used to describe the variation characteristics of Gaussian points at multiple time frequencies. The expression for the high-frequency temporal embedding is: In the formula, for High-frequency temporal embedding of moments For the first A sine function can learn parameters. For the first The angular frequency of each time frequency base is determined using an exponentially growing multi-scale frequency setting method. For the time step in autonomous driving scenarios For the first Learnable parameters of a cosine function; where and As learnable parameters, they are updated along with the deformation network and Gaussian property optimization parameters via gradient descent, thereby adaptively selecting the time-frequency components that are more important to the current dynamic scene. This represents the total number of high-frequency time components.
[0039] When constructing the instance embedding vector, the instance features of each dynamic vehicle are initialized with a corresponding Gaussian distribution: In the formula, Initialize the Gaussian distribution for the features of dynamic vehicle instances. For probability distribution symbols, Let be the probability density function of a normal distribution. Let be the variance of the Gaussian distribution; where the instance embedding is a learnable variable that is optimized during backpropagation.
[0040] S3. Obtain the joint feature vector based on the spatial embedding vector, high-frequency temporal embedding, and instance embedding vector; construct temporal high-frequency consistency constraints based on the joint feature vector, and optimize the Gaussian ellipsoid properties based on the temporal high-frequency consistency constraints; specifically including the following steps: S301. Concatenate the high-frequency temporal embedding, spatial embedding vector, and instance embedding vector to obtain the joint feature vector: In the formula, For joint feature vectors, For high-frequency time embedding, For spatial embedding vectors, Embed vectors for instances, This is the acquisition timestamp corresponding to the Gaussian ellipsoid. is the four-dimensional spacetime coordinate of the Gaussian ellipsoid.
[0041] S302. Input the joint feature vector into the Gaussian deformation network to predict the property residuals of the Gaussian ellipsoid in the time dimension: In the formula, Let Gaussian ellipsoid be the set of global deformation residuals in its current state. For Gaussian deformation network prediction operations, The displacement residual is at the center of the Gaussian ellipsoid. The residuals are at the Gaussian ellipsoid scale. This represents the residual term from the rotation of the Gaussian ellipsoid.
[0042] The Gaussian deformation network is a multilayer perceptron network, as shown in the attached figure. Figure 2As shown, it includes an input layer, a linear layer, an activation layer, and an output layer connected in sequence. The input layer is used to input the joint feature vector, and the linear layer is used to perform linear transformation and dimensionality mapping on the input feature vector to extract high-dimensional abstract features. The specific calculation steps are: performing matrix multiplication and bias addition operations on the input feature vector, and the calculation formula is: In the formula, The feature vector output by the linear layer. For the first The input features of the layer For the first The weight matrix of each linear layer For the first Bias vectors for a linear layer.
[0043] The activation layer is used to introduce nonlinear transformation capabilities into the network, enabling the network to fit the complex nonlinear spatiotemporal transformation relationships of dynamic scenes. Specifically, it performs nonlinear activation operations on the feature vectors output by the linear layer, using the ReLU activation function to transform the feature vectors output by the linear layer into feature vectors output by the activation layer.
[0044] The output layer is used to output the predicted displacement residuals, scale residuals, and rotation residuals.
[0045] S303. The center position, scale, and rotation parameters of the Gaussian ellipsoid are recursively updated over time based on the residuals of the Gaussian ellipsoid's properties; the expression for updating the center position of the Gaussian ellipsoid is: In the formula, For the Gaussian ellipsoid at time step Center position parameters at time For the Gaussian ellipsoid at time step Center position parameters at time The displacement residuals predicted by the deformation network are used to update the spatial position of the Gaussian ellipsoid; The expression for updating the scale parameter of the Gaussian ellipsoid is: In the formula, For the Gaussian ellipsoid at time step The scale parameter of time, For the Gaussian ellipsoid at time step The scale parameter of time, The scale residuals, predicted by the deformation network, are used to characterize the change in target size over time.
[0046] The expression for updating the rotation parameters of the Gaussian ellipsoid is: In the formula, For the Gaussian ellipsoid at time step Rotation parameters at time, For the Gaussian ellipsoid at time step Rotation parameters at time, The rotational residuals predicted by the deformation network are used to update the attitude changes of the Gaussian ellipsoid.
[0047] The high-frequency consistency constraint module suppresses the accumulation of frame-by-frame noise, enhances the continuity and stability of dynamic targets between adjacent time frames, and avoids problems such as light flickering distortion or motion trajectory jitter.
[0048] S4. Construct spatial high-frequency consistency constraints and optimize Gaussian ellipsoid properties; wherein in step S4, when optimizing the Gaussian ellipsoid, static Gaussian ellipsoids such as lane lines and building edges are optimized; specifically including the following steps: S401. Obtain the Gaussian ellipsoid of the static region and project the Gaussian ellipsoid of the static region onto the image plane at different viewpoints. When obtaining the Gaussian ellipsoid of the static region, the dynamic vehicle region and the static region in the collected dynamic urban scene multi-source data can be distinguished by the 3D bounding box of datasets such as nuScenes. The Gaussian ellipsoid of the distinguished static region is projected onto the image plane at different viewpoints using a camera projection model to generate the rendering results of the corresponding viewpoints. The specific content of distinguishing and projecting the static Gaussian ellipsoid is existing technology and will not be described in detail here.
[0049] S402. Perform bilateral filtering on the real image captured by the vehicle to obtain low-frequency components, and calculate image gradient information based on the real image captured by the vehicle. The method for obtaining image gradient information is as follows: First, convert the RGB real image captured by the vehicle into a grayscale image; second, use the Sobel gradient operator to perform convolution operations with the grayscale image by using horizontal and vertical convolution kernels respectively to obtain horizontal and vertical gradient components; finally, calculate the gradient magnitude of each pixel and normalize the gradient magnitude to the [0,1] interval to obtain the final image gradient information.
[0050] By incorporating image gradient information to highlight edge regions, a weight map is constructed to emphasize spatial high-frequency structures. The formula for calculating the spatial high-frequency structure weight map is as follows: In the formula, This is a weighted graph of high-frequency spatial structures. This is a normalization function used to linearly map the weight matrix values to the [0,1] interval, unifying the range of weight values and ensuring the stability of subsequent weighted loss calculations. Real images of vehicles, This is a bilateral filtering operation. For fixed hyperparameters, it is preferred to use them in this embodiment. , This refers to the gradient information calculated by applying a fixed Sobel operator to a real image.
[0051] S403. For high-frequency structural regions such as lane lines, curb boundaries, and building textures, an edge-aware weighted loss function is constructed, and multi-view consistency supervision is introduced to impose stronger constraints on high-frequency structural regions, thereby improving the reconstruction accuracy and consistency of geometric structures under sparse view conditions. The edge-aware weighted loss function is as follows: In the formula, The weighted loss function is based on edge awareness. For a norm, This is the weighting amplification factor. This refers to the element-wise multiplication operation of a matrix. Render the image for the static area.
[0052] When performing multi-view consistency supervision, a spatial high-frequency consistency loss function is established, which includes weighted photometric loss and structural similarity loss. The structural similarity loss function is as follows: In the formula, For structural similarity loss, It is a standard structural similarity index used to comprehensively compare brightness, contrast, and structural information.
[0053] In step S403, a structural loss function for a single image is also established to process image scenes captured by a single camera. The structural loss function for a single image is as follows: In the formula, The structural loss function for a single image. The scoring coefficient is used to adjust the weight distribution between pixel-level edge-weighted loss and structural similarity loss. The weighted loss function is based on edge awareness. This is the structural similarity loss function.
[0054] S404. Based on the structural loss function of a single image, a multi-view weighted consistency loss function is constructed. For scenes with multiple camera inputs, a shared high-frequency weight map is used to jointly constrain the pixel errors between the rendering results of each viewpoint and the real image by calculating the spatial high-frequency consistency loss. The multi-view spatial high-frequency consistency loss function is as follows: In the formula, For multi-view spatial high-frequency consistency loss, This is to collect the total number of camera viewpoints. A shared spatial high-frequency structure weight map from multiple perspectives. Render the image for the static area corresponding to the viewpoint. These are real-world images of vehicles taken from the corresponding perspective.
[0055] The aforementioned multi-view spatial high-frequency consistency loss function and single-image structural loss function are both used for attribute optimization of static Gaussian ellipsoids. The specific optimization steps are as follows: Forward propagation rendering: Input the attribute parameters of the current static Gaussian ellipsoid into the Gaussian sputtering rendering pipeline to render static region rendering images under single view / multi-view conditions respectively; Loss calculation: Based on the rendered image and the corresponding real-world captured image, calculate the corresponding single-frame structural loss / multi-view consistency loss to obtain the total loss value under the current parameters; Backpropagation gradient calculation: With minimizing the loss value as the optimization objective, the gradient of the loss function with respect to the attribute parameters such as the center position, scale, rotation, color, and opacity of the static Gaussian ellipsoid is calculated through the backpropagation algorithm. Parameter Iterative Update: The Adam gradient descent optimizer iteratively updates the attribute parameters of the static Gaussian ellipsoid based on the calculated gradient until the loss value converges to the preset threshold, thus completing the optimization of the static Gaussian ellipsoid and achieving high-fidelity reconstruction of the static scene.
[0056] S5. Train a Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints using the optimized Gaussian ellipsoid. This includes the following steps: S501. The dynamic Gaussian ellipsoid of the vehicle portion and the static Gaussian ellipsoid of the background portion, optimized in steps S3 and S4, are blended, and differentiable rasterization rendering is performed. The contributions of each Gaussian ellipsoid on the image plane are accumulated, resulting in accumulated color and transparency contributions from each ellipsoid in the image, thus obtaining the rendered image of the autonomous driving scene from the corresponding viewpoint. The gradient transfer relationship from rendering error to parameters is expressed as a function of the Gaussian ellipsoid parameters through the differentiable rasterization rendering process. In the formula, Rendering images for autonomous driving scenarios, For differentiable rasterization rendering functions, For the first Spatial position parameters of a Gaussian ellipsoid For the first The scale parameters of a Gaussian ellipsoid For the first The rotation vector parameters of a Gaussian ellipsoid For the first Opacity parameters of a Gaussian ellipsoid For the first Color or appearance-related parameters of a Gaussian ellipsoid For the first An instance embedding vector associated with a Gaussian ellipsoid and its corresponding vehicle instance. This represents the total number of Gaussian ellipsoids in the autonomous driving scenario.
[0057] S502. Construct a joint objective function based on the rendered image of the autonomous driving scene; specifically, compare the rendered image of the autonomous driving scene with the real image, and construct a joint loss function including depth consistency loss, opacity constraint loss, foreground constraint loss, and regularization loss; where the depth consistency loss is used to constrain the consistency between the depth result obtained by rendering and the real depth, and the depth consistency loss function is: In the formula, For deep consistency loss, A collection of image pixels. The pixel position in the image. To render pixels obtained by projecting from a Gaussian ellipsoid using differentiable rendering The predicted depth value at that location, To the pixel position of the rendered image The true depth value at the corresponding pixel position of the aligned real image or the pseudo-true depth value calculated from multi-view geometric relationships.
[0058] Opacity constraint loss is used to limit the range and distribution of the opacity parameter of the Gaussian ellipsoid. If the target opacity value is not explicitly set, this loss can also suppress the unconstrained growth of the opacity parameter, thereby improving rendering stability. The opacity constraint loss function is: In the formula, For the loss due to opacity constraints, The number of Gaussian ellipsoids. For the first The opacity parameters corresponding to each Gaussian ellipsoid, among which , The target opacity value is determined based on prior rules or rendering consistency constraints. It is the square of the L2 norm.
[0059] Foreground constraint loss is used to enhance the distinguishing ability between dynamic target regions and background regions. The foreground constraint loss function is: In the formula, Loss due to forward constraint A collection of image pixels. The pixels are obtained by weighted accumulation of the opacity of the Gaussian ellipsoid. Predicted forward value at the location, The reference foreground value is generated by the segmentation model. for Norm.
[0060] Regularization loss is used to suppress unconstrained variations in model parameters and improve the stability of the training process. The regularization loss function is: In the formula, For regularization loss, The number of Gaussian ellipsoids. For the first The scale parameters of a Gaussian ellipsoid For the first The rotation parameters of a Gaussian ellipsoid For the first Vehicle instance embedding vectors associated with a Gaussian ellipsoid.
[0061] The joint objective function is obtained by weighting and combining the above-mentioned constraint loss and the loss function in step S4: In the formula, For the common goal, For structural similarity loss loss of basic reconstruction The weighting balancing coefficients are used to adjust the proportion of single-frame image structural similarity constraints and basic pixel reconstruction constraints in the global loss. The weights for the deep consistency loss, The weights for the opacity constraint loss are... The weights for the high-frequency consistency loss in the multi-view space. The weights for the prospect constraint loss.
[0062] S503. Optimize the parameters of the Gaussian ellipsoid and the model by backpropagating the joint objective function. When updating the parameters of the Gaussian ellipsoid, the joint objective function is used as the optimization objective to update the learnable parameters of the Gaussian ellipsoid. Its gradient with respect to each learnable parameter can be expressed as: In the formula, For the partial derivative of the total loss of the joint objective function, For the partial derivative of the set of learnable parameters of the model, For the first Partial derivatives of rendered images of autonomous driving scenes from various perspectives The set of learnable parameters for the model. For the first Spatial position parameters of a Gaussian ellipsoid For the first The scale parameters of a Gaussian ellipsoid For the first The rotation vector parameters of a Gaussian ellipsoid For the first Opacity parameters of a Gaussian ellipsoid For the first Color or appearance-related parameters of a Gaussian ellipsoid For the first The instance embedding vector associated with each Gaussian ellipsoid and its corresponding vehicle instance is given.
[0063] Based on the gradient results of the learnable parameters, the parameters of the Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints are updated in each training iteration: In the formula, To determine the values of the model parameters for the next iteration after the update. To assign values to the model parameters in the current iteration, This is the learning rate.
[0064] The training process includes density adaptive control, with Gaussian attribute optimization and density adaptive control performed alternately. Density adaptive control comprises two processes: densification and pruning. Densification adaptively increases the Gaussian representation density in high-frequency changing regions, while pruning removes less contributing or redundant Gaussian ellipsoids to improve overall modeling efficiency and reconstruction quality. The specific process of density adaptive control is existing technology and will not be elaborated upon here.
[0065] S6. Render the multi-source data of the dynamic urban scene to be processed using the Gaussian sputtering model based on spatio-temporal high-frequency consistency constraints after training. The rendering results of this solution are as shown in the appendix Figure 3 and the appendix Figure 4 as shown, the appendix Figure 3 is a comparison chart of the temporal rendering effects of this solution and the mainstream dynamic scene rendering solution on the dynamic vehicle targets in autonomous driving, the appendix Figure 4 is a comparison chart of the reconstruction effects and accuracy metrics of this solution and the mainstream solution in complex urban scenes of autonomous driving.
[0066] The descriptions of other mainstream rendering methods compared in this solution are as follows: GT (Ground Truth) Chinese name: True Value Benchmark Function Introduction: Real image data collected for autonomous driving scenarios, used as the accuracy evaluation benchmark for all reconstruction and rendering solutions, to measure the degree of fit between the rendering results of each solution and the visual effects and geometric structures of the real scene.
[0067] DeformableGS (Deformable 3D Gaussian Splatting) Chinese name: Deformable 3D Gaussian Sputtering Function Introduction: A mainstream solution for dynamic scene reconstruction based on 3D Gaussian sputtering. By introducing a Gaussian deformation model, it realizes the temporal modeling of dynamic targets and supports the new view synthesis of non-rigid dynamic scenes. It is a basic comparison solution in the field of dynamic 3DGS.
[0068] PVG (Pixel-level Volumetric Gaussian) Chinese name: Pixel-level Voxel Gaussian Function Introduction: A dynamic scene reconstruction solution that combines voxel implicit representation and explicit Gaussian sputtering. By constraining the spatial range of the Gaussian distribution through a voxel grid, it realizes the multi-view rendering of dynamic scenes.
[0069] StreetGS Chinese name: Street Scene 3D Gaussian Sputtering Function Introduction: A 3D Gaussian sputtering solution optimized for autonomous driving urban scenes, which is adapted and optimized for the large-scale and long-distance characteristics of street scene static scenes. It is a mainstream comparison solution for autonomous driving static scene reconstruction.
[0070] OmniRe Chinese name: Omnidirectional Dynamic Scene Reconstruction Solution Function Introduction: A dynamic scene reconstruction solution based on the fusion of neural radiance fields and Gaussian sputtering, which is optimized for the spatio-temporal consistency of multi-view dynamic scenes and supports the new view synthesis of complex dynamic scenes.
[0071] Compared with the above rendering methods, the core advantages of this solution are as follows: The temporal reconstruction accuracy of dynamic targets is significantly leading, as shown in the appendix Figure 3The temporal rendering comparison results show that DeformableGS, PVG, StreetGS, OmniRe, and other comparative solutions all exhibit varying degrees of motion blur, misalignment, and loss of details such as headlights when rendering dynamic vehicle targets, resulting in poor consistency of rendering results at consecutive time points. In contrast, our solution achieves accurate temporal modeling of dynamic vehicle targets through high-frequency temporal embedding, instance embedding, and Gaussian deformation networks. The rendering results at consecutive time points are free of motion blur and misalignment, and the detail reproduction of vehicle headlights, body contours, and other details is completely consistent with the ground truth benchmark (GT). Our solution significantly outperforms the comparative solutions in reconstructing dynamic targets, which is of core concern to autonomous driving.
[0072] The high-frequency structure reproduction of urban scenes is more accurate, as shown in the attached image. Figure 4 The comparison results of complex urban scenes show that the comparison scheme has problems such as blurred edges and color distortion in rendering core high-frequency structures of autonomous driving, such as traffic lights and lane lines, and insufficient ability to restore details of small targets such as sudden traffic lights. In contrast, this scheme strengthens the reconstruction constraints of high-frequency structures in urban scenes by using edge perception weighted loss and multi-view spatial high-frequency consistency constraints. The rendered traffic lights, lane lines and building edges are clear and sharp, and the detail restoration is highly matched with the ground truth.
[0073] The overall reconstruction accuracy and rendering quality are superior (see attached image). Figure 4 As shown in the accuracy comparison, the PSNR (Peak Signal-to-Noise Ratio) of this solution reaches 34.75dB, significantly higher than all other comparison solutions including DeformableGS (27.56dB), PVG (29.45dB), StreetGS (30.11dB), and OmniRe (31.44dB). The improved PSNR demonstrates that the rendering results of this solution have smaller pixel-level errors compared to the real scene, and both overall reconstruction accuracy and visual quality have been significantly improved.
[0074] In the comparison of solutions with stronger joint modeling capabilities for static and dynamic scenes, StreetGS is only suitable for static street scene scenes and lacks the ability to reconstruct dynamic targets; DeformableGS, PVG and other solutions only focus on dynamic target modeling and lack the ability to restore the high-frequency structure of static backgrounds; while this solution, through joint encoding of spatial and temporal embedding and de-path optimization of static and dynamic Gaussian ellipsoids, simultaneously achieves high-precision reconstruction of static urban scenes and temporal consistency modeling of dynamic targets, and is more suitable for the full-element reconstruction needs of dynamic urban scenes for autonomous driving.
[0075] This invention also aims to provide an autonomous driving scene reconstruction system based on spatiotemporal consistency constraints, applicable to the aforementioned autonomous driving scene reconstruction method based on spatiotemporal consistency constraints, comprising: The input module is used to input multi-source data of dynamic urban scenes; The Gaussian initialization and scene representation construction module is used to obtain the initial geometric representation of the dynamic urban scene based on multi-source data of the dynamic urban scene and construct the Gaussian ellipsoid. The spatiotemporal embedding construction module is used to construct high-frequency temporal embedding, spatial embedding vectors, and instance embedding vectors based on the three-dimensional spatial coordinates and temporal information coordinates of the Gaussian ellipsoid. The temporal high-frequency consistency constraint module models the displacement, rotation, and scale changes of a Gaussian ellipsoid in the temporal dimension by fusing high-frequency temporal embeddings, spatial embedding vectors, and instance embedding vectors. This suppresses frame-by-frame noise and enhances the temporal consistency of dynamic targets between adjacent frames. The module jointly encodes the position coordinates and temporal information of the Gaussian ellipsoid in three-dimensional space using multi-resolution four-dimensional hash embedding, enabling the model to simultaneously express low-frequency overall trends and high-frequency detailed changes, thereby enhancing the modeling ability for high-frequency temporal behavior in dynamic scenes.
[0076] The Spatial High-Frequency Consistency Constraint Module is used to apply stronger supervision to high-frequency structural regions in the rendering results by constructing an edge-aware weighted loss function, and further introduces multi-view consistency constraints to improve the alignment accuracy and detail fidelity of spatial geometry under different viewpoints.
[0077] The Gaussian ellipsoid optimization module is used to optimize the properties of the Gaussian ellipsoid based on the modeling of the temporal high-frequency consistency constraint module or the weighted loss function constructed by the spatial high-frequency consistency constraint module.
[0078] The projection and rasterization rendering module is used to project the optimized Gaussian ellipsoids onto the image plane according to the intrinsic and extrinsic parameters of the target view camera, and render them through a differentiable rasterization process, so that each Gaussian ellipsoid contributes to the cumulative color and transparency in the image, thus obtaining the rendered image of the autonomous driving scene under the corresponding view.
[0079] The present invention also aims to provide an autonomous driving scene reconstruction device based on spatiotemporal consistency constraints, including a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the above-described autonomous driving scene reconstruction method based on spatiotemporal consistency constraints.
[0080] The present invention also aims to provide a computer-readable storage medium containing a computer program, wherein the computer program is stored thereon, characterized in that, when the computer program is executed by one or more processors, it implements the steps of the above-described method for reconstructing autonomous driving scenarios based on spatiotemporal consistency constraints.
[0081] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the scope of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for reconstructing autonomous driving scenarios based on spatiotemporal consistency constraints, characterized in that, Includes the following steps: Acquire multi-source data of dynamic urban scenes, establish multiple Gaussian ellipsoids based on the multi-source data of dynamic urban scenes, and construct a Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints; The three-dimensional spatial coordinates of each Gaussian ellipsoid are combined with the time dimension to form a four-dimensional spatiotemporal coordinate. Based on the four-dimensional spatiotemporal coordinate, a spatial embedding vector and a high-frequency temporal embedding are constructed, and an instance embedding vector is constructed based on multi-source data of dynamic urban scenes. A joint feature vector is obtained based on the spatial embedding vector, high-frequency temporal embedding, and instance embedding vector; a temporal high-frequency consistency constraint is constructed based on the joint feature vector, and the properties of the Gaussian ellipsoid are optimized based on the temporal high-frequency consistency constraint. Construct spatial high-frequency consistency constraints and optimize the properties of static Gaussian ellipsoids; A Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints is trained using the optimized Gaussian ellipsoid. The trained Gaussian sputtering model based on spatiotemporal high-frequency consistency constraints is used to render multi-source data of dynamic urban scenes.
2. The autonomous driving scene reconstruction method according to claim 1, characterized in that, The step of constructing a spatial embedding vector based on four-dimensional spatiotemporal coordinates includes: Construct four-dimensional spatiotemporal coordinates based on the spatial and temporal positions of each Gaussian point in the scene containing the Gaussian ellipsoid; Calculate the four-dimensional resolution vector of the four-dimensional spatiotemporal coordinates at each resolution level; Calculate the fractional offset of each resolution coordinate in the four-dimensional resolution vector within a grid point in the hash grid; Spatiotemporal features of each resolution level are generated based on the fractional offset of each resolution coordinate in the four-dimensional resolution vector. The spatial embedding vector of the Gaussian ellipsoid is obtained by concatenating the spatiotemporal features at all resolution levels.
3. The autonomous driving scene reconstruction method according to claim 1, characterized in that, When constructing a high-frequency time embedding based on four-dimensional spatiotemporal coordinates, one or more time frequency components are generated based on the four-dimensional spatiotemporal coordinates, and all time frequency components are accumulated to obtain the high-frequency time embedding.
4. The autonomous driving scene reconstruction method according to claim 1, characterized in that, The steps of obtaining a joint feature vector based on the spatial embedding vector, high-frequency temporal embedding, and instance embedding vector; constructing a temporal high-frequency consistency constraint based on the joint feature vector; and optimizing the properties of the Gaussian ellipsoid based on the temporal high-frequency consistency constraint include: The joint feature vector is obtained by concatenating the high-frequency temporal embedding, spatial embedding vector, and instance embedding vector. By inputting the joint feature vector into the Gaussian deformation network, the property residuals of the Gaussian ellipsoid in the time dimension are predicted. The center position, scale, and rotation parameters of the Gaussian ellipsoid are updated recursively over time based on the residuals of the Gaussian ellipsoid's properties.
5. The autonomous driving scene reconstruction method according to claim 1, characterized in that, The steps of constructing spatial high-frequency consistency constraints and optimizing the properties of the static Gaussian ellipsoid include: Obtain the Gaussian ellipsoid of the static region and project the Gaussian ellipsoid of the static region onto the image plane at different viewpoints to obtain the static region rendering image; Low-frequency components are obtained by bilateral filtering of real images of vehicles, and image gradient information is calculated based on real images of vehicles. A weight map for emphasizing high-frequency spatial structures is constructed by combining the image gradient information. A structural similarity loss function is constructed based on real images of vehicles and static region rendered images. Construct a multi-view spatial high-frequency consistency loss function; Optimize the properties of static Gaussian ellipsoids based on multi-view spatial high-frequency consistency loss function and structural similarity loss function.
6. The autonomous driving scene reconstruction method according to claim 1, characterized in that, The steps of training the Gaussian sputtering model based on the optimized Gaussian ellipsoid and spatiotemporal high-frequency consistency constraints include: The rendered image of the autonomous driving scene is obtained based on the optimized Gaussian ellipsoid rendering. Construct a joint objective function based on rendered images from autonomous driving scenarios; The parameters of the Gaussian ellipsoid and the model are optimized by using a joint objective function.
7. An autonomous driving scene reconstruction system based on spatiotemporal consistency constraints, applicable to the autonomous driving scene reconstruction method based on spatiotemporal consistency constraints as described in any one of claims 1-6, characterized in that, include: The input unit is used to input multi-source data of dynamic urban scenes; The Gaussian initialization and scene representation construction unit is used to obtain the initial geometric representation of the dynamic urban scene based on multi-source data of the dynamic urban scene and construct the Gaussian ellipsoid. Spatiotemporal embedding construction unit, used to construct high-frequency temporal embedding, spatial embedding vector and instance embedding vector based on the three-dimensional spatial coordinates and temporal information coordinates of the Gaussian ellipsoid; Spatiotemporal high-frequency consistency constraint unit, used to generate temporal high-frequency consistency constraints and spatial high-frequency consistency constraints; The Gaussian ellipsoid optimization element is used to optimize the properties of a Gaussian ellipsoid based on temporal high-frequency consistency constraints and spatial high-frequency consistency constraints. The projection and rasterization rendering unit is used to render the optimized Gaussian ellipsoid to obtain the rendered image of the autonomous driving scene from the corresponding viewpoint.
8. The autonomous driving scene reconstruction system according to claim 7, characterized in that, The spatiotemporal high-frequency consistency constraint unit includes a time high-frequency consistency constraint module and a spatial high-frequency consistency constraint module connected in parallel. The time-high frequency consistency constraint module is used to model the displacement, rotation and scale changes of the Gaussian ellipsoid in the time dimension by fusing high-frequency time embedding, spatial embedding vector and instance embedding vector to generate time-high frequency consistency constraints. The spatial high-frequency consistency constraint module is used to generate spatial high-frequency consistency constraints by constructing an edge-aware weighted loss function to supervise high-frequency structural regions.
9. An autonomous driving scene reconstruction device based on spatiotemporal consistency constraints, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the autonomous driving scene reconstruction method based on spatiotemporal consistency constraints as described in any one of claims 1-6.
10. A computer-readable storage medium containing a computer program, wherein the computer program is stored thereon, characterized in that, When the computer program is executed by one or more processors, it implements the steps of the autonomous driving scene reconstruction method based on spatiotemporal consistency constraints as described in any one of claims 1-6.