Dynamic street view reconstruction method and device under two-dimensional semantic prior, equipment and medium

By using a dynamic street scene reconstruction method based on two-dimensional semantic priors, and leveraging multi-view images and sparse LiDAR depth maps, combined with semantic segmentation and feedforward distillation techniques, we optimize street representations, solve the problems of computational complexity and annotation dependency in dynamic street scene reconstruction, and achieve high-precision instance-level decomposition and spatiotemporally consistent reconstruction.

CN122199739APending Publication Date: 2026-06-12ZHEJIANG YOULU ROBOT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG YOULU ROBOT TECH CO LTD
Filing Date
2026-05-12
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

Existing simulation methods have high computational complexity in reconstructing dynamic street scenes, making it difficult to meet real-time requirements. Furthermore, they rely on costly 3D bounding box annotations, which leads to confusion between the motion of the vehicle and the motion of foreground objects, resulting in poor decomposition quality and failing to meet the needs of accurate interactive simulation.

Method used

A dynamic street scene reconstruction method based on two-dimensional semantic prior is adopted. By collecting multi-view images and sparse LiDAR depth maps, a semantic segmentation model is used to identify potential moving objects, initialize interactive street representations, and a pre-trained feedforward 3D Gaussian generator network is used to assist in feedforward distillation. The street representations are optimized by combining a temporal contrastive loss function to generate a four-dimensional instance segmentation mask.

Benefits of technology

It achieves high-quality dynamic scene decoupling and reconstruction, improves the decoupling quality and editing flexibility of dynamic scenes, solves Gaussian drift and artifact problems, ensures high-fidelity rendering under interactive editing, and reduces the dependence on high-cost 3D bounding box annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122199739A_ABST
    Figure CN122199739A_ABST
Patent Text Reader

Abstract

The application relates to the field of simulation technology and provides a dynamic street view reconstruction method and device under two-dimensional semantic prior, equipment and medium, the interactive street representation includes a static background and a dynamic foreground represented by a free space-time Gaussian element with an explicit speed parameter, effectively improving the decoupling quality and editing flexibility of the dynamic scene; with the aid of a feedforward three-dimensional Gaussian generation network, the interactive street representation is fed forward and distilled based on a new view angle based on multi-view images and sparse laser radar depth maps, the Gaussian drift and artifact problems of the dynamic scene under a new track view angle are solved, and high-fidelity rendering under interactive editing is ensured; the current street representation is optimized by using a time sequence comparison loss function and a two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask, the dependence on high-cost three-dimensional bounding box labeling is overcome, and high-precision instance-level decomposition and space-time consistency reconstruction can be realized by using only two-dimensional labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of simulation technology, and in particular to a method, apparatus, device and medium for dynamic street scene reconstruction under two-dimensional semantic prior. Background Technology

[0002] In closed-loop simulations such as autonomous driving, constructing an interactive street environment is crucial.

[0003] However, among current simulation methods, the neural radiation field method has high computational complexity, making it difficult to meet real-time requirements. While 3D Gaussian sputtering performs excellently in static scenes, it also faces challenges when handling dynamic street scenes. Furthermore, existing dynamic scene reconstruction methods typically rely heavily on costly 3D bounding box annotations to achieve object decoupling. Self-supervised methods often struggle to handle complex vehicle motion and crowded dynamic objects, leading to confusion between vehicle motion and foreground object motion, resulting in poor decomposition quality and failing to meet the needs of accurate interactive simulation.

[0004] Therefore, how to achieve high-quality dynamic scene decoupling and reconstruction relying solely on easily obtainable two-dimensional annotations is an urgent problem to be solved. Summary of the Invention

[0005] In view of the above, it is necessary to provide a method, device, equipment and medium for dynamic street scene reconstruction under two-dimensional semantic prior, in order to solve the problem that high-quality dynamic scene decoupling and reconstruction cannot be achieved based on two-dimensional annotation.

[0006] A dynamic street scene reconstruction method based on two-dimensional semantic prior, the method comprising: In response to a dynamic street scene reconstruction command triggered based on a target street scene, multi-view images and sparse LiDAR depth maps of the target street scene are acquired. A semantic segmentation model is used to identify potential moving objects in the multi-view images to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image. Initialize an interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-space-time Gaussian units with explicit velocity parameters; Using a pre-trained feedforward 3D Gaussian generator network as an aid, the interactive street representation is subjected to feedforward distillation based on a new perspective based on the multi-view image and the sparse lidar depth map to obtain the optimized current street representation. A learnable feature vector is initialized for each free spatiotemporal Gaussian unit in the current street representation, and the current street representation is optimized using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

[0007] A dynamic street scene reconstruction device based on two-dimensional semantic prior, the dynamic street scene reconstruction device based on two-dimensional semantic prior includes: The acquisition unit is used to acquire multi-view images and sparse lidar depth maps of the target street scene in response to a dynamic street scene reconstruction command triggered based on the target street scene. The generation unit is used to identify potential moving objects in the multi-view images using a semantic segmentation model, so as to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image. An initialization unit is used to initialize an interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-space-time Gaussian units with explicit velocity parameters. The distillation unit is used to perform feedforward distillation of the interactive street representation based on the new perspective, using a pre-trained feedforward 3D Gaussian generator network as an aid, based on the multi-view image and the sparse lidar depth map, to obtain the optimized current street representation. An optimization unit is used to initialize a learnable feature vector for each free spatiotemporal Gaussian unit in the current street representation, and optimize the current street representation using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

[0008] A computer device, the computer device comprising: A memory for storing at least one instruction; and a processor for executing the instructions stored in the memory to implement the dynamic street scene reconstruction method under two-dimensional semantic prior.

[0009] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the dynamic street scene reconstruction method under two-dimensional semantic prior.

[0010] As can be seen from the above technical solutions, the interactive street representation of the present invention includes a static background represented by standard three-dimensional Gaussian primitives and a dynamic foreground represented by free spatiotemporal Gaussian primitives with explicit velocity parameters, which effectively improves the decoupling quality and editing flexibility of dynamic scenes. With the aid of a pre-trained feedforward three-dimensional Gaussian generator network, the interactive street representation is distilled based on a new perspective using multi-view images and sparse LiDAR depth maps, which solves the Gaussian drift and artifact problems of dynamic scenes under the new trajectory perspective, and ensures high-fidelity rendering under interactive editing. The current street representation is optimized by using the constructed temporal contrast loss function and two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask, which overcomes the dependence on high-cost three-dimensional bounding box annotation and achieves high-precision instance-level decomposition and spatiotemporal consistency reconstruction using only two-dimensional segmentation mask. Attached Figure Description

[0011] Figure 1 This is a flowchart of a preferred embodiment of the dynamic street scene reconstruction method under two-dimensional semantic prior of the present invention; Figure 2 This is a functional block diagram of a preferred embodiment of the dynamic street scene reconstruction device under two-dimensional semantic prior of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the dynamic street scene reconstruction method under two-dimensional semantic prior according to the present invention. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0013] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the dynamic street scene reconstruction method under two-dimensional semantic prior according to the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.

[0014] The dynamic street scene reconstruction method based on two-dimensional semantic prior is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0015] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.

[0016] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0017] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0018] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0019] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0020] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).

[0021] S10, in response to a dynamic street scene reconstruction command triggered based on the target street scene, acquire multi-view images and sparse lidar depth maps of the target street scene.

[0022] In this embodiment, the target street scene can be a street scene under an autonomous driving scenario.

[0023] In this embodiment, the dynamic street view reconstruction command can be triggered by relevant personnel according to actual reconstruction needs.

[0024] In this embodiment, the multi-view image can be multi-view video image data collected by a high-definition camera or the like, and is an RGB (Red Green Blue) image.

[0025] In this embodiment, the sparse lidar depth map can be obtained by projecting and converting the lidar point cloud collected by the lidar.

[0026] S11, using a semantic segmentation model to identify potential moving objects in the multi-view images, in order to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image.

[0027] In this embodiment, the semantic segmentation model may include models such as Fully Convolutional Network (FCN), U-Net, and SegNet, for object detection and segmentation of the multi-view images.

[0028] The above models can all adopt their own general structures, and take the multi-view image as input and the dynamic foreground mask and the two-dimensional instance segmentation mask as output.

[0029] The above model can be trained using images labeled with potential moving objects in a conventional manner.

[0030] In this embodiment, the semantic segmentation model is directly applied to the two-dimensional color multi-view image to identify potential moving objects, thereby generating a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image.

[0031] The dynamic foreground mask is used to distinguish between dynamic and static areas.

[0032] The two-dimensional instance segmentation mask is used to distinguish different individuals.

[0033] S12, Initialize the interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-space-time Gaussian units with explicit velocity parameters.

[0034] In this embodiment, the Gaussian parameters of the static background may include position, scale, rotation quaternion, opacity, and spherical harmonic coefficient.

[0035] In this embodiment, the Gaussian parameters of the dynamic foreground may include position, time center, four-dimensional scale, quaternion of encoding direction, opacity, spherical harmonic coefficients, and explicit velocity vector.

[0036] In the dynamic foreground, the spatial position of each free-space-time Gaussian element at any time is equal to the initial position of each free-space-time Gaussian element plus the product of the explicit velocity vector and the time offset.

[0037] The explicit velocity parameters allow each free-spacetime Gaussian element to move freely in spacetime to fit rigid and non-rigid body motions.

[0038] S13, using a pre-trained feedforward 3D Gaussian generator network as an aid, the interactive street representation is subjected to feedforward distillation based on a new perspective based on the multi-view image and the sparse lidar depth map to obtain the optimized current street representation.

[0039] In this embodiment, the feedforward 3D Gaussian generation network can be composed of PDAv2 (Prompting DepthAnything for 4K Resolution Accurate Metric Depth Estimation, version 2) and CNN (Convolutional Neural Network), and can be fully pre-trained using a high-fidelity synthetic dataset for holistic indoor scene understanding (Hypersim: A Photorealistic Synthetic Dataset for Holistic Indoor Scene Understanding).

[0040] In this embodiment, the optimized current street representation, obtained by using a pre-trained feedforward 3D Gaussian generator network as an aid and performing feedforward distillation based on the multi-view images and the sparse lidar depth map to obtain the current street representation, includes: The sparse lidar depth map is converted into a dense depth map using a depth completion network. The multi-view image and the dense depth map are input into the encoder-decoder network for regression prediction to obtain pixel-aligned three-dimensional Gaussian parameters; Based on the three-dimensional Gaussian parameters, the new viewpoint is rendered using the feedforward three-dimensional Gaussian generator network to obtain the new viewpoint image and the cumulative opacity mask. The interactive street representation is rendered to the new perspective to obtain the current reconstructed image; Calculate the reconstruction loss between the current reconstructed image and the new viewpoint image, and apply the cumulative opacity mask as a spatial weight matrix to the reconstruction loss to obtain the distillation loss that masks the gradient backpropagation of the invisible region. Based on the distillation loss, the interactive street representation is optimized through backpropagation to obtain the current street representation.

[0041] The depth completion network and the encoder-decoder network are general network structures. Furthermore, the depth completion network can be trained using training samples corresponding to the sparse LiDAR depth map, and the encoder-decoder network can be trained using training samples corresponding to the dense depth map.

[0042] The new perspective can be a configured virtual camera trajectory perspective, such as the perspective after a vehicle changes lanes.

[0043] The pixel alignment refers to the direct regression of a complete set of three-dimensional Gaussian parameters (including position offset, scale, rotation, color, opacity, etc.) for each pixel position in the feature map output by the encoder-decoder network, and the rapid construction of a local three-dimensional Gaussian field based on a single frame input.

[0044] In this embodiment, the step of rendering the new viewpoint image and the cumulative opacity mask based on the feedforward 3D Gaussian generation network, based on the 3D Gaussian parameters, includes: The three-dimensional Gaussian primitives corresponding to the three-dimensional Gaussian parameters are projected from the world coordinate system onto the two-dimensional image plane of the new perspective, and forward α-mix rendering is performed along the ray in depth order based on the three-dimensional Gaussian parameters to obtain the new perspective image obtained by color rendering, and the cumulative opacity mask obtained by rendering through the opacity channel.

[0045] For example, when the new viewpoint is a new camera pose (such as a lateral offset pose) within a given virtual camera trajectory, the predicted 3D Gaussian primitives are projected from the world coordinate system onto the 2D image plane of the new camera. Further, based on the predicted 3D Gaussian primitives' position, color, and opacity, forward alpha-blending volume rendering is performed along the ray in depth order. Specifically, the color is rendered to obtain the new viewpoint image, and the opacity channel is rendered to obtain the cumulative opacity mask.

[0046] In this embodiment, the step of calculating the reconstruction loss between the current reconstructed image and the new viewpoint image, and applying the accumulated opacity mask as a spatial weight matrix to the reconstruction loss to obtain the distillation loss that masks the gradient backpropagation of the invisible region includes: The current reconstructed image and the new viewpoint image are compared pixel by pixel to calculate the L1 loss as the reconstruction loss; The distillation loss is obtained by multiplying the cumulative opacity mask by the reconstruction loss.

[0047] The invisible region refers to those regions in the new viewpoint image rendered by the feedforward 3D Gaussian generator network that are occluded or outside the field of view in the original input single frame viewpoint, causing the feedforward 3D Gaussian generator network to be unable to predict effective geometric information (these regions have extremely low values ​​in the cumulative opacity mask).

[0048] Because the feedforward 3D Gaussian generator network (GFGN) is based on single-frame prediction, it cannot see what is behind the occlusion. Therefore, if the entire new perspective image is used to supervise the interactive street representation, the regions where the GFGN makes wild guesses or are blank will introduce incorrect gradients, causing the main representation to forcibly fit into these blank regions and produce artifacts. Thus, this embodiment uses the cumulative opacity mask as a spatial weight matrix applied to the reconstruction loss, constraining the cumulative opacity of the foreground in invisible regions to 0 (indicating that the GFGN considers this region invisible or without effective geometry), thereby shielding gradient backpropagation from invisible regions and eliminating their interference.

[0049] This embodiment uses the new perspective image as a supervision signal (pseudo-ground value) to train the interactive street representation, and uses the cumulative opacity mask to cover invisible areas to avoid artifacts, thereby enabling the distillation of single-frame geometric priors into the full scene representation, significantly improving the rendering quality of the new perspective.

[0050] S14, initialize learnable feature vectors for each free spatiotemporal Gaussian unit in the current street representation, and optimize the current street representation using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

[0051] In this embodiment, each feature vector obtained during initialization is random and unformed, and will be continuously optimized in subsequent processing.

[0052] In this embodiment, before optimizing each feature vector using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask, the method further includes: The reconstruction loss is constructed by weighting the absolute error between the rendered image and the real image with the structural similarity loss. The foreground mask loss is constructed based on the binary cross-entropy loss between the two-dimensional contour of the rendered image and the dynamic foreground mask. The reconstruction loss, the foreground mask loss, and the distillation loss are weighted and calculated based on a weighting mechanism to obtain the temporal comparison loss function.

[0053] The two-dimensional contour is the opacity projection.

[0054] The reconstruction loss ensures the rendering quality of the basic image.

[0055] Specifically, the foreground mask loss can constrain the free-space Gaussian elements to accurately cover the foreground object region.

[0056] In the weighted calculation, the weight coefficients used can be the optimal values ​​selected through a large number of experiments.

[0057] In this embodiment, the optimization of the current street representation using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including the four-dimensional instance segmentation mask includes: Each feature vector is rendered to the camera view corresponding to the multi-view image to obtain a feature map; The feature map is sampled according to the two-dimensional instance segmentation mask to obtain pixel feature points; In the first optimization stage, the geometric and appearance parameters of each Gaussian unit in the current street representation are optimized using the temporal contrast loss function to perform binary decoupling between the static background and dynamic foreground in the current street representation. In the second optimization stage, the geometric and appearance parameters of each Gaussian unit in the current street representation are frozen. The temporal contrast loss function is used to narrow the feature distance between pixel feature points belonging to the same instance and widen the feature distance between pixel feature points belonging to different instances. Based on the explicit velocity of each free spatiotemporal Gaussian unit, the two-dimensional instance constraints of a single frame are propagated into the spatiotemporally consistent four-dimensional instance segmentation mask. After optimization, a reconstructed street representation including the four-dimensional instance segmentation mask is obtained.

[0058] When sampling the feature map based on the two-dimensional instance segmentation mask, corresponding pixel feature points can be sampled in the feature map according to the spatial location provided by the two-dimensional instance segmentation mask. Specifically, the two-dimensional instance segmentation mask provides the exact pixel coordinate range of each independent object (such as a specific car) on the two-dimensional plane of the image. Using these coordinates as indices, the corresponding pixel feature points are extracted at the same position in the newly rendered feature map. After sampling, it is known which pixel feature points belong to the same object (by narrowing their feature distance) and which pixel feature points belong to different objects (by widening their feature distance), thus completing the loss calculation for contrastive learning.

[0059] The geometric parameters and appearance parameters are the underlying parameters of the interactive street representation. Specifically, the geometric parameters may include the position, scale, and rotation of Gaussian elements, and the appearance parameters may include the color, spherical harmonic coefficient, and opacity of Gaussian elements.

[0060] In the second optimization stage, the calculation of the geometric and appearance parameters can be stopped (e.g., the gradient calculation control flag of the corresponding tensor is set to False, indicating that gradient tracking is turned off, the parameters remain fixed, and they do not participate in training), and they are removed from the optimizer. Only the feature vector of each free spatiotemporal Gaussian unit is retained in the optimizer. In this way, during backpropagation, only the value of the feature vector is updated, thereby freezing the geometric and appearance parameters of each Gaussian unit in the current street representation.

[0061] The temporal contrastive loss function is used to achieve instance-level multi-objective decomposition. This function can bring pixel features belonging to the same instance closer together and widen the distance between pixel features belonging to different instances.

[0062] Among them, the short-range spatiotemporal correspondence established by the explicit velocity of each free spatiotemporal Gaussian unit can be used to propagate the two-dimensional instance constraint of a single frame into a spatiotemporally consistent four-dimensional segmentation by sampling different frames on the time series and iteratively optimizing the local feature constraints.

[0063] The short-range spatiotemporal correspondence refers to the association and mapping relationship of the physical positions of the same 3D object (i.e., the same Gaussian primitive) between several adjacent or nearby video frames. Because it is the same object, although the positions of the pixels rendered on the 2D image of the object change in adjacent frames, they should correspond to the same instance features. Specifically, since the spatial position of each free-space-time Gaussian primitive at any time is equal to the initial position of each free-space-time Gaussian primitive plus the product of the explicit velocity vector and the time offset, the specific pixel position corresponding to a certain Gaussian primitive on the feature map from frame t to frame t+1 can be accurately calculated in the two frames. Using this position as a bridge (i.e., the short-range spatiotemporal correspondence), the temporal contrastive loss function can be used to force the features of the same primitive to remain consistent in different frames, thereby stringing together the independent 2D masks of a single frame into a temporally continuous and consistent 4D instance segmentation. For example, when propagating the two-dimensional instance constraints of a single frame into a spatiotemporally consistent four-dimensional instance segmentation mask based on the explicit velocity of each free-space-time Gaussian unit, the t-th frame can be sampled on the time axis at each iteration, and the spatial position of the corresponding free-space-time Gaussian unit at time t can be calculated. Then, its corresponding feature vector is rendered into the image space of that frame to obtain the feature map of the current frame. Further, in the feature map of the current frame, pixel sampling is performed directly based on the two-dimensional instance segmentation mask of the current frame, and the temporal contrast loss is calculated. Since the feature vectors of Gaussian units are globally shared, and their physical motion is bound by explicit velocity, when the same Gaussian unit moves to different frames over time, it will be repeatedly constrained by the single-frame two-dimensional masks of these different frames. Because it is the same entity, this single-frame training based on temporal sampling objectively propagates the independent single-frame two-dimensional constraints naturally into a spatiotemporally consistent four-dimensional instance segmentation mask through the Gaussian unit itself.

[0064] By introducing a contrastive learning mechanism, it is possible to force similar features within the same instance and mutually exclusive features between different instances based on forward-rendered feature maps and calculation of temporal contrastive loss functions. This iteratively optimizes the feature vector, thereby elevating the two-dimensional mask to a spatiotemporally consistent four-dimensional instance segmentation, achieving instance-level decoupling and segmentation across frames.

[0065] In the above embodiments, the task of the first optimization stage is purely 3D reconstruction, aiming to build a solid geometric base and render the image realistically. If complex contrast loss is introduced at this time to stretch the features, it can easily interfere with the convergence of the Gaussian primitive spatial positions, producing a large number of artifacts. Therefore, it is necessary to first fix the spatial physical properties of the Gaussian primitives in the first optimization stage, and then focus on affixing feature labels for instance segmentation to the dynamic objects on this base in the second optimization stage. In this way, a win-win situation of maximizing reconstruction fidelity and segmentation accuracy can be achieved.

[0066] This embodiment combines interactive street representation with feedforward network distillation. It does not rely on costly 3D bounding box annotation. It can achieve high-fidelity dynamic scene reconstruction and instance-level decoupling using only 2D semantic priors. It can support trajectory editing and interactive simulation (such as high-fidelity rendering under interactive editing such as lane changing, acceleration and deceleration). It can be applied to fields such as autonomous driving closed-loop simulation and multi-view dynamic scene reconstruction.

[0067] As can be seen from the above technical solutions, the interactive street representation of the present invention includes a static background represented by standard three-dimensional Gaussian primitives and a dynamic foreground represented by free spatiotemporal Gaussian primitives with explicit velocity parameters, which effectively improves the decoupling quality and editing flexibility of dynamic scenes. With the aid of a pre-trained feedforward three-dimensional Gaussian generator network, the interactive street representation is distilled based on a new perspective using multi-view images and sparse LiDAR depth maps, which solves the Gaussian drift and artifact problems of dynamic scenes under the new trajectory perspective, and ensures high-fidelity rendering under interactive editing. The current street representation is optimized by using the constructed temporal contrast loss function and two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask, which overcomes the dependence on high-cost three-dimensional bounding box annotation and achieves high-precision instance-level decomposition and spatiotemporal consistency reconstruction using only two-dimensional segmentation mask.

[0068] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the dynamic street scene reconstruction device under two-dimensional semantic prior according to the present invention. The dynamic street scene reconstruction device 11 under two-dimensional semantic prior includes a data acquisition unit 110, a generation unit 111, an initialization unit 112, a distillation unit 113, and an optimization unit 114. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0069] The acquisition unit 110 is used to acquire multi-view images and sparse lidar depth maps of the target street scene in response to a dynamic street scene reconstruction command triggered based on the target street scene. The generation unit 111 is used to identify potential moving objects in the multi-view image using a semantic segmentation model, so as to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame image. The initialization unit 112 is used to initialize the interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-spacetime Gaussian units with explicit velocity parameters. The distillation unit 113 is used to perform feedforward distillation of the interactive street representation based on a new perspective, using a pre-trained feedforward 3D Gaussian generator network as an aid, based on the multi-view image and the sparse lidar depth map, to obtain the optimized current street representation. The optimization unit 114 is used to initialize a learnable feature vector for each free spatiotemporal Gaussian unit in the current street representation, and optimize the current street representation using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

[0070] As can be seen from the above technical solutions, the interactive street representation of the present invention includes a static background represented by standard three-dimensional Gaussian primitives and a dynamic foreground represented by free spatiotemporal Gaussian primitives with explicit velocity parameters, which effectively improves the decoupling quality and editing flexibility of dynamic scenes. With the aid of a pre-trained feedforward three-dimensional Gaussian generator network, the interactive street representation is distilled based on a new perspective using multi-view images and sparse LiDAR depth maps, which solves the Gaussian drift and artifact problems of dynamic scenes under the new trajectory perspective, and ensures high-fidelity rendering under interactive editing. The current street representation is optimized by using the constructed temporal contrast loss function and two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask, which overcomes the dependence on high-cost three-dimensional bounding box annotation and achieves high-precision instance-level decomposition and spatiotemporal consistency reconstruction using only two-dimensional segmentation mask.

[0071] like Figure 3 The diagram shown is a schematic representation of the computer device used in a preferred embodiment of the dynamic street scene reconstruction method based on two-dimensional semantic priors according to the present invention.

[0072] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a dynamic street scene reconstruction program under two-dimensional semantic prior.

[0073] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.

[0074] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.

[0075] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of a dynamic street scene reconstruction program under two-dimensional semantic prior, but also to temporarily store data that has been output or will be output.

[0076] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing a dynamic street scene reconstruction program based on two-dimensional semantic priors) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.

[0077] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps in the above-described embodiments of the dynamic street scene reconstruction method under two-dimensional semantic priors, for example... Figure 1 The steps are shown.

[0078] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a data acquisition unit 110, a generation unit 111, an initialization unit 112, a distillation unit 113, and an optimization unit 114.

[0079] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the dynamic street scene reconstruction method under two-dimensional semantic prior as described in the various embodiments of this invention.

[0080] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0081] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.

[0082] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0083] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0084] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.

[0085] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0086] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the computer device 1 and other computer devices.

[0087] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.

[0088] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0089] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0090] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a dynamic street scene reconstruction method under two-dimensional semantic prior, and the processor 13 can execute the multiple instructions to achieve: In response to a dynamic street scene reconstruction command triggered based on a target street scene, multi-view images and sparse LiDAR depth maps of the target street scene are acquired. A semantic segmentation model is used to identify potential moving objects in the multi-view images to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image. Initialize an interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-space-time Gaussian units with explicit velocity parameters; Using a pre-trained feedforward 3D Gaussian generator network as an aid, the interactive street representation is subjected to feedforward distillation based on a new perspective based on the multi-view image and the sparse lidar depth map to obtain the optimized current street representation. A learnable feature vector is initialized for each free spatiotemporal Gaussian unit in the current street representation, and the current street representation is optimized using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

[0091] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0092] It should be noted that all the data involved in this case was legally obtained.

[0093] If any AI models, software tools, or components not belonging to this company appear in the embodiments of this invention, they are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this invention has been obtained by an entity authorized (with the knowledge and consent) or fully authorized by all parties through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.

[0094] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0095] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0096] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0097] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0098] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0099] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0100] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A dynamic street scene reconstruction method based on two-dimensional semantic prior, characterized in that, The dynamic street scene reconstruction method based on two-dimensional semantic prior includes: In response to a dynamic street scene reconstruction command triggered based on a target street scene, multi-view images and sparse lidar depth maps of the target street scene are acquired. A semantic segmentation model is used to identify potential moving objects in the multi-view images to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image. Initialize an interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-space-time Gaussian units with explicit velocity parameters; Using a pre-trained feedforward 3D Gaussian generator network as an aid, the interactive street representation is subjected to feedforward distillation based on a new perspective based on the multi-view image and the sparse lidar depth map to obtain the optimized current street representation. A learnable feature vector is initialized for each free spatiotemporal Gaussian unit in the current street representation, and the current street representation is optimized using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

2. The dynamic street scene reconstruction method under two-dimensional semantic prior as described in claim 1, characterized in that, For the dynamic foreground, the spatial position of each free-spacetime Gaussian element at any given time is equal to the initial position of each free-spacetime Gaussian element plus the product of the explicit velocity vector and the time offset.

3. The dynamic street scene reconstruction method under two-dimensional semantic prior as described in claim 1, characterized in that, The method, aided by a pre-trained feedforward 3D Gaussian generator network, performs feedforward distillation on the interactive street representation based on the multi-view images and the sparse LiDAR depth map, resulting in an optimized current street representation including: The sparse lidar depth map is converted into a dense depth map using a depth completion network. The multi-view image and the dense depth map are input into the encoder-decoder network for regression prediction to obtain pixel-aligned three-dimensional Gaussian parameters; Based on the three-dimensional Gaussian parameters, the new viewpoint is rendered using the feedforward three-dimensional Gaussian generator network to obtain the new viewpoint image and the cumulative opacity mask. The interactive street representation is rendered to the new perspective to obtain the current reconstructed image; Calculate the reconstruction loss between the current reconstructed image and the new viewpoint image, and apply the cumulative opacity mask as a spatial weight matrix to the reconstruction loss to obtain the distillation loss that masks the gradient backpropagation of the invisible region. Based on the distillation loss, the interactive street representation is optimized through backpropagation to obtain the current street representation.

4. The dynamic street scene reconstruction method under two-dimensional semantic prior as described in claim 3, characterized in that, Based on the three-dimensional Gaussian parameters, the process of rendering the new viewpoint using the feedforward three-dimensional Gaussian generation network to obtain the new viewpoint image and cumulative opacity mask includes: The three-dimensional Gaussian primitives corresponding to the three-dimensional Gaussian parameters are projected from the world coordinate system onto the two-dimensional image plane of the new perspective, and forward α-mix rendering is performed along the ray in depth order based on the three-dimensional Gaussian parameters to obtain the new perspective image obtained by color rendering, and the cumulative opacity mask obtained by rendering through the opacity channel.

5. The dynamic street scene reconstruction method under two-dimensional semantic prior as described in claim 3, characterized in that, The process of calculating the reconstruction loss between the current reconstructed image and the new viewpoint image, and applying the accumulated opacity mask as a spatial weight matrix to the reconstruction loss to obtain the distillation loss that masks the gradient backpropagation of invisible regions, includes: The current reconstructed image and the new viewpoint image are compared pixel by pixel to calculate the L1 loss as the reconstruction loss; The distillation loss is obtained by multiplying the cumulative opacity mask by the reconstruction loss.

6. The dynamic street scene reconstruction method under two-dimensional semantic prior as described in claim 5, characterized in that, Before optimizing each feature vector using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask, the method further includes: The reconstruction loss is constructed by weighting the absolute error between the rendered image and the real image with the structural similarity loss. The foreground mask loss is constructed based on the binary cross-entropy loss between the two-dimensional contour of the rendered image and the dynamic foreground mask. The reconstruction loss, the foreground mask loss, and the distillation loss are weighted and calculated based on a weighting mechanism to obtain the temporal comparison loss function.

7. The dynamic street scene reconstruction method under two-dimensional semantic prior as described in claim 1, characterized in that, The optimization of the current street representation using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask yields a reconstructed street representation including a four-dimensional instance segmentation mask, comprising: Each feature vector is rendered to the camera view corresponding to the multi-view image to obtain a feature map; The feature map is sampled according to the two-dimensional instance segmentation mask to obtain pixel feature points; In the first optimization stage, the geometric and appearance parameters of each Gaussian unit in the current street representation are optimized using the temporal contrast loss function to perform binary decoupling between the static background and dynamic foreground in the current street representation. In the second optimization stage, the geometric and appearance parameters of each Gaussian unit in the current street representation are frozen. The temporal contrast loss function is used to narrow the feature distance between pixel feature points belonging to the same instance and widen the feature distance between pixel feature points belonging to different instances. Based on the explicit velocity of each free spatiotemporal Gaussian unit, the two-dimensional instance constraints of a single frame are propagated into the spatiotemporally consistent four-dimensional instance segmentation mask. After optimization, a reconstructed street representation including the four-dimensional instance segmentation mask is obtained.

8. A dynamic street scene reconstruction device based on two-dimensional semantic prior, characterized in that, The dynamic street scene reconstruction device based on two-dimensional semantic prior includes: The acquisition unit is used to acquire multi-view images and sparse lidar depth maps of the target street scene in response to a dynamic street scene reconstruction command triggered based on the target street scene. The generation unit is used to identify potential moving objects in the multi-view images using a semantic segmentation model, so as to generate a dynamic foreground mask and a two-dimensional instance segmentation mask for each frame of the image. An initialization unit is used to initialize an interactive street representation; wherein the interactive street representation includes a static background represented by standard three-dimensional Gaussian units and a dynamic foreground represented by free-space-time Gaussian units with explicit velocity parameters. The distillation unit is used to perform feedforward distillation of the interactive street representation based on the new perspective, using a pre-trained feedforward 3D Gaussian generator network as an aid, based on the multi-view image and the sparse lidar depth map, to obtain the optimized current street representation. An optimization unit is used to initialize a learnable feature vector for each free spatiotemporal Gaussian unit in the current street representation, and optimize the current street representation using the constructed temporal contrastive loss function and the two-dimensional instance segmentation mask to obtain a reconstructed street representation including a four-dimensional instance segmentation mask.

9. A computer device, characterized in that, The computer device includes: A memory for storing at least one instruction; and a processor for executing the instructions stored in the memory to implement the dynamic street scene reconstruction method under two-dimensional semantic prior as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the dynamic street scene reconstruction method under two-dimensional semantic prior as described in any one of claims 1 to 7.