Scene reconstruction model training method and device, equipment, medium and product
By training the scene reconstruction model and combining video repair models, the problem of poor rendering effect in the reconstruction of autonomous driving scenes is solved, and the rendering quality and model robustness are improved.
Patent Information
- Application Number
- CN202510335831.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-20
AI Technical Summary
In the reconstruction of autonomous driving scenes, the lack of multi-view images and dynamic complexity leads to poor rendering effects, which are prone to shadowing and blurring.
By acquiring the image data set acquired by the vehicle under the original trajectory and the synchronized point cloud data set, the first reconstruction model is obtained, and the rendered image is repaired using the pre-trained video repair model, and the second reconstruction model is finally trained to generate the rendered image under the new trajectory.
It improves the image rendering quality in autonomous driving scenarios, reduces the shadow and blurring in the rendering effect, and enhances the robustness of the model.
Smart Images

Figure CN120182500A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical fields of artificial intelligence, scene reconstruction, and autonomous driving technology, and in particular to a method, device, equipment, medium, and product for training a scene reconstruction model. Background Art
[0002] In the task of autonomous driving scene reconstruction, generally, the reconstruction of a driving scene is completed by inputting a section of driving video. Due to the perspective sparsity of autonomous driving data and the dynamic complexity of driving scenes, compared with traditional 3D (three-dimensional) scene reconstruction, the 3D reconstruction of autonomous driving scenes faces greater challenges.
[0003] In related technologies, there are no rich multi-perspective images in the autonomous driving data set. Its test set is generally obtained by extracting frames from the data of the original trajectory, while the training set uses the remaining images after frame extraction. This makes related technologies mainly focus on the scene reconstruction quality under the original trajectory. However, when using the reconstructed autonomous driving scene to render new trajectories such as gradual lane changes and lane translations, the rendering effect is poor, and phenomena such as ghosting and blurring are likely to occur. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure provides a method, device, equipment, medium, and product for training a scene reconstruction model to improve the image rendering quality in autonomous driving scenes by using the trained scene reconstruction model.
[0005] In the first aspect of the present disclosure, a method for training a scene reconstruction model is provided, including:
[0006] Obtain an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized in time with the image data set, where the first point cloud data set is collected by a lidar;
[0007] Use the image data set and the first point cloud data set to train a first reconstruction model;
[0008] Use the first reconstruction model to render an image under a new trajectory to obtain a rendered image under the new trajectory;
[0009] Use a pre-trained video repair model to perform repair processing on the rendered image under the new trajectory to obtain a repaired image under the new trajectory;
[0010] Use the repaired image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory to train a second reconstruction model, where the second reconstruction model is used to generate a rendered image under the new trajectory.
[0011] In some alternative embodiments, training the second reconstruction model by using the repaired image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory includes:
[0012] Perform position encoding on the Gaussian parameters and time steps of the first reconstruction model respectively to obtain Gaussian parameter position encoding and time step position encoding, where the time step is used to indicate the frame number of the captured images under the original trajectory in the image dataset;
[0013] Using the first multi-layer perceptron, based on the Gaussian parameter position encoding and time step position encoding, obtain a first feature;
[0014] Using a second multi-layer perceptron, based on the first pose of the vehicle under the original trajectory and the second pose of the vehicle under the new trajectory, obtain a second feature;
[0015] Based on the first feature, the second feature, and the Gaussian parameters of the first reconstruction model, obtain a first initial reconstruction model;
[0016] Use the repaired image under the new trajectory to train the first initial reconstruction model to obtain the second reconstruction model.
[0017] In some alternative embodiments, training the first reconstruction model by using the image dataset and the first point cloud dataset includes:
[0018] Fuse multiple frames of point clouds in the first point cloud dataset to obtain fused point cloud data;
[0019] Respectively segment the road surface point cloud data, non-road surface static background point cloud data, and dynamic object point cloud data from the fused point cloud data;
[0020] Based on the road surface point cloud data, perform initialization processing on the parameters of the three-dimensional Gaussian splash model to obtain an initial road surface model;
[0021] Based on the non-road surface static background point cloud data, perform initialization processing on the parameters of the three-dimensional Gaussian splash model to obtain an initial non-road surface background model;
[0022] Based on the dynamic object point cloud data, perform initialization processing on the parameters of the three-dimensional Gaussian splash model to obtain an initial dynamic object model;
[0023] Perform combination processing on the initial road surface model, the initial non-road surface background model, and the initial dynamic object model to obtain a second initial reconstruction model;
[0024] Using the acquisition images under the original trajectories of each frame in the image dataset, train the second initial reconstruction model to obtain the first reconstruction model.
[0025] In some alternative embodiments, after obtaining the second initial reconstruction model, it further includes:
[0026] Using the acquisition images under the original trajectories of each frame in the image dataset, insufficiently train the second initial reconstruction model to obtain a first reconstruction model with insufficient training;
[0027] Using the first reconstruction model with insufficient training and the acquisition images under the original trajectories of each frame in the image dataset, train the initial diffusion model to obtain the video restoration model.
[0028] In some alternative embodiments, the step of using the first reconstruction model with insufficient training and the images under the original trajectories of each frame in the image dataset to train the initial diffusion model to obtain the video restoration model includes:
[0029] Using the first reconstruction model with insufficient training to render the images under the original trajectories of each frame to obtain the corresponding rendered images under the original trajectories of each frame;
[0030] Using the rendered images under the original trajectories and the acquisition images under the original trajectories to construct a video restoration dataset;
[0031] Using the video restoration dataset to train the initial diffusion model to obtain the video restoration model.
[0032] In some alternative embodiments, the road surface point cloud data includes a plurality of sub-point cloud data;
[0033] The step of initializing the parameters of the three-dimensional Gaussian splash model based on the road surface point cloud data to obtain an initial road surface model includes:
[0034] Obtain the central positions of each sub-point cloud data to obtain the central positions of each Gaussian sphere;
[0035] Determine the central positions of each Gaussian sphere and the number of Gaussian spheres as three-dimensional structure prior information;
[0036] Integrate the three-dimensional structure prior information into the three-dimensional Gaussian splash model to obtain the initial road surface model, and the three-dimensional structure prior information remains fixed during the operation of training the second initial reconstruction model.
[0037] In some alternative embodiments, the initial dynamic object model is a model based on the local coordinate system of the dynamic object;
[0038] Performing combined processing on the initial road surface model, the initial non-road background model, and the initial dynamic object model to obtain a second initial reconstruction model, including:
[0039] Based on the pose information of each dynamic object in the world coordinate system, obtaining a rotation matrix and a translation vector for transforming each dynamic object from the local coordinate system to the world coordinate system;
[0040] Based on the rotation matrix and the translation vector, performing transformation processing on the initial dynamic object model to obtain a transformed dynamic object model, where the transformed dynamic object model is a model based on the world coordinate system;
[0041] Performing combined processing on the initial road surface model, the initial non-road background model, and the transformed dynamic object model to obtain the second initial reconstruction model.
[0042] In some optional embodiments, before rendering an image under the new trajectory using the first reconstruction model to obtain a rendered image under the new trajectory, the method further includes:
[0043] Sampling the new trajectory based on a preset sampling period, and the offset of the new trajectory sampled in each sampling period increases sequentially relative to the original trajectory.
[0044] A second aspect of the present disclosure provides a training device for a scene reconstruction model, including:
[0045] A first acquisition module, configured to acquire an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized with the image data set in time, where the first point cloud data set is collected by a lidar;
[0046] A first training module, configured to train a first reconstruction model using the image data set and the first point cloud data set;
[0047] A first rendering module, configured to render an image under the new trajectory using the first reconstruction model to obtain a rendered image under the new trajectory;
[0048] A repair module, configured to perform repair processing on the rendered image under the new trajectory using a pre-trained video repair model to obtain a repaired image under the new trajectory;
[0049] A second training module, configured to train a second reconstruction model using the repaired image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory, where the second reconstruction model is used to generate a rendered image under the new trajectory.
[0050] In some alternative embodiments, the second training module includes:
[0051] An encoding sub-module for performing position encoding on the Gaussian parameters and time steps of the first reconstruction model respectively to obtain Gaussian parameter position encoding and time step position encoding, where the time step is used to indicate the frame number of the captured image under the original trajectory in the image dataset;
[0052] A first processing sub-module for using the first multi-layer perceptron to obtain a first feature based on the Gaussian parameter position encoding and time step position encoding;
[0053] A second processing sub-module for using a second multi-layer perceptron to obtain a second feature based on the first pose of the vehicle under the original trajectory and the second pose of the vehicle under the new trajectory;
[0054] A parameter update sub-module for obtaining a first initial reconstruction model based on the first feature, the second feature, and the Gaussian parameters of the first reconstruction model;
[0055] A first training sub-module for training the first initial reconstruction model using the repaired image under the new trajectory to obtain the second reconstruction model.
[0056] In some alternative embodiments, the first training module includes:
[0057] A point cloud fusion sub-module for fusing multiple frames of point clouds in the first point cloud dataset to obtain fused point cloud data;
[0058] A point cloud segmentation sub-module for respectively segmenting road surface point cloud data, non-road surface static background point cloud data, and dynamic object point cloud data from the fused point cloud data;
[0059] A first initialization sub-module for initializing the parameters of the three-dimensional Gaussian splash model based on the road surface point cloud data to obtain an initial road surface model;
[0060] A second initialization sub-module for initializing the parameters of the three-dimensional Gaussian splash model based on the non-road surface static background point cloud data to obtain an initial non-road surface background model;
[0061] A third initialization sub-module for initializing the parameters of the three-dimensional Gaussian splash model based on the dynamic object point cloud data to obtain an initial dynamic object model;
[0062] A combination sub-module for performing combination processing on the initial road surface model, the initial non-road surface background model, and the initial dynamic object model to obtain a second initial reconstruction model;
[0063] The second training sub-module is used to train the second initial reconstruction model by using the captured images under the original trajectories of each frame in the image dataset, so as to obtain the first reconstruction model.
[0064] In some alternative embodiments, the device further includes:
[0065] The third training module is used to insufficiently train the second initial reconstruction model by using the captured images under the original trajectories of each frame in the image dataset, so as to obtain an insufficiently trained first reconstruction model;
[0066] The fourth training module is used to train the initial diffusion model by using the insufficiently trained first reconstruction model and the captured images under the original trajectories of each frame in the image dataset, so as to obtain the video repair model.
[0067] In some alternative embodiments, the fourth training module includes:
[0068] The first rendering sub-module is used to render the images under the original trajectories of each frame by using the insufficiently trained first reconstruction model, so as to obtain the corresponding rendered images under the original trajectories of each frame;
[0069] The dataset construction sub-module is used to construct a video repair dataset by using the rendered images under the original trajectories and the captured images under the original trajectories;
[0070] The third training sub-module is used to train the initial diffusion model by using the video repair dataset, so as to obtain the video repair model.
[0071] In some alternative embodiments, the road surface point cloud data includes a plurality of sub-point cloud data;
[0072] The first initialization sub-module includes:
[0073] The first acquisition unit is used to acquire the central positions of the sub-point cloud data, so as to obtain the central positions of the Gaussian spheres;
[0074] The prior information determination unit is used to determine the central positions of the Gaussian spheres and the number of Gaussian spheres as three-dimensional structure prior information;
[0075] The integration unit is used to integrate the three-dimensional structure prior information into the three-dimensional Gaussian splash model, so as to obtain the initial road surface model, and the three-dimensional structure prior information is fixed during the operation of training the second initial reconstruction model.
[0076] In some alternative embodiments, the initial dynamic object model is a model based on the local coordinate system of the dynamic object;
[0077] The combined sub-module includes:
[0078] A second acquisition unit, configured to acquire a rotation matrix and a translation vector for transforming each dynamic object from a local coordinate system to a world coordinate system based on the pose information of each dynamic object in the world coordinate system;
[0079] A conversion unit, configured to perform a conversion process on the initial dynamic object model based on the rotation matrix and the translation vector to obtain a converted dynamic object model, where the converted dynamic object model is a model based on the world coordinate system;
[0080] A combination unit, configured to perform a combination process on the initial road surface model, the initial non-road surface background model, and the converted dynamic object model to obtain the second initial reconstruction model.
[0081] In some alternative embodiments, the device further includes:
[0082] A sampling module, configured to sample new trajectories based on a preset sampling period, and the offset of the new trajectories sampled in each sampling period with respect to the original trajectory increases sequentially.
[0083] In a third aspect of the present disclosure, there is provided a computer-readable storage medium storing computer program instructions, which when executed, implement the training method of the scene reconstruction model proposed in the first aspect embodiment of the present disclosure.
[0084] In a fourth aspect of the present disclosure, there is provided an electronic device, where the electronic device includes:
[0085] A memory, configured to store a computer program product;
[0086] A processor, configured to execute the computer program product stored in the memory, and when the computer program product is executed, implement the training method of the scene reconstruction model proposed in the first aspect embodiment of the present disclosure.
[0087] In a fifth aspect of the present disclosure, there is provided a computer program product including computer program instructions, which when executed by a processor, implement the training method of the scene reconstruction model proposed in the first aspect embodiment above.
[0088] Based on the embodiments of the present disclosure, for the task of autonomous driving scene reconstruction, an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized with the image data set in time can be obtained; then, using the image data set and the first point cloud data set, a first reconstruction model is trained; and the first reconstruction model is used to render images under the new trajectory to obtain rendered images under the new trajectory; then, the rendered images under the new trajectory are repaired using a pre-trained video repair model to obtain repaired images under the new trajectory; finally, using the repaired images under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory, a second reconstruction model is trained, and the second reconstruction model is used to generate rendered images under the new trajectory. Thus, by using the repaired images under the new trajectory repaired by the video repair model as training samples to further train the scene reconstruction model, it helps to bridge the domain gap between the repaired images under the new trajectory and the collected images under the original trajectory, and improve the rendering quality of the reconstruction model and the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] The drawings forming a part of the specification depict embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.
[0090] Referring to the drawings, the present disclosure can be more clearly understood from the following detailed description, wherein:
[0091] Figure 1 FIG. is a scene reconstruction effect diagram implemented using an existing method and the scene reconstruction model of the present disclosure;
[0092] Figure 2 FIG. is a flowchart of an embodiment of the training method of the scene reconstruction model of the present disclosure;
[0093] Figure 3 FIG. is a schematic diagram of the training process of the training method of the scene reconstruction model provided by an exemplary embodiment of the present disclosure;
[0094] Figure 4 FIG. is a schematic diagram of the process of step 205 provided by an exemplary embodiment of the present disclosure;
[0095] Figure 5 FIG. is a schematic diagram of the process of step 202 provided by an exemplary embodiment of the present disclosure;
[0096] Figure 6 FIG. is a schematic diagram of the process of the training process of the video reconstruction model provided by an exemplary embodiment of the present disclosure;
[0097] Figure 7 FIG. is a schematic diagram of the structure of the training device of the scene reconstruction model provided by an exemplary embodiment of the present disclosure;
[0098] Figure 8 It is a schematic structural diagram of a training device for a scene reconstruction model provided by another exemplary embodiment of the present disclosure;
[0099] Figure 9 It is a schematic structural diagram of a training device for a scene reconstruction model provided by another exemplary embodiment of the present disclosure;
[0100] Figure 10 It is a structural diagram of an electronic device provided by an exemplary embodiment of the present disclosure. Detailed implementation manners
[0101] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.
[0102] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present disclosure.
[0103] Overview of the present disclosure
[0104] In the process of implementing the present disclosure, the inventors found through research that in the field of autonomous driving scenarios, the data collection for autonomous driving generally collects data for a driving process of a vehicle, and the trajectory of this driving process is called the original trajectory. If the vehicle only travels in one lane during this driving process, the collected data may only be data of one lane, resulting in only data of the original trajectory in the autonomous driving dataset and lacking rich multi-view images, leading to the view sparsity of autonomous driving data. If you want to render a video in a new view, for example, when rendering new trajectories such as gradual lane changes and lane translations, the video generated under the new trajectory will exhibit obvious ghosting phenomena, such as trailing, fragmentation, blurring, etc.
[0105] Such as Figure 1 In the left figure in [reference], it is the rendered image under the original trajectory (the image indicated by label 11) and the rendered image under the new trajectory of lane offset (lane shift 3 meters) (the image indicated by label 12) generated using Street Gaussians (a method for modeling dynamic urban street scenes). According to the rendering effect indicated by the square frame in label 12, it can be seen that the image quality of the rendered image under the new trajectory generated using Street Gaussians is poor.
[0106] Such as Figure 1The middle figure shows the rendered image under the original trajectory generated by ReconDreamer (a scene reconstruction model constructed through online restoration technology) (the image indicated by label 13) and the rendered image under the new trajectory with a lane shift of 3 meters (the image indicated by label 14). It can be seen that the rendered image under the new trajectory generated by the online restoration technology and the reconstruction model is of relatively high quality as a whole. However, according to the rendering effect indicated by the square frame in label 14, it can be seen that there are problems of quality degradation in some local positions of the rendered image under the new trajectory generated by ReconDreamer.
[0107] Based on the technical problems existing in the above-mentioned prior art, the embodiment of the present application provides a training method for a scene reconstruction model. First, an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized with the image data set in time are obtained; then, using the image data set and the first point cloud data set, a first reconstruction model is trained; and the first reconstruction model is used to render the image under the new trajectory to obtain the rendered image under the new trajectory; then, the rendered image under the new trajectory is repaired using a pre-trained video repair model to obtain the repaired image under the new trajectory; finally, using the repaired image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory, a second reconstruction model is trained, and the second reconstruction model is used to generate the rendered image under the new trajectory. Thus, by using the repaired image under the new trajectory repaired by the video repair model as a training sample to further train the scene reconstruction model, it helps to bridge the domain gap between the repaired image under the new trajectory and the collected image under the original trajectory, and improve the rendering quality of the reconstruction model and the robustness of the model.
[0108] See Figure 1 The right figure shows the rendered image under the original trajectory generated by the scene reconstruction model generated by the technical solution of the present disclosure (the image indicated by label 15) and the rendered image under the new trajectory with a lane shift of 3 meters (the image indicated by label 16). It can be seen that the quality of the rendered image under the original trajectory and the rendered image under the new trajectory generated by the scene reconstruction model trained based on the technical solution of the present disclosure is very high.
[0109] In addition, the rendering effect of the rendered image under the original trajectory can be measured by a deep learning-based image similarity metric (Learned Perceptual Image Patch Similarity, abbreviated as LPIPS). This metric is used to measure the similarity between the rendered image under the original trajectory and the acquired image under the original trajectory. The smaller this value, the more similar the two images (the rendered image under the original trajectory and the acquired image under the original trajectory), and the better the rendering quality. For the rendering effect of the rendered image under the new trajectory, it can be measured by the FID (Fréchet Inception Distance) metric. This metric is used to measure the similarity of the image features between the rendered image under the new trajectory and the acquired image under the original trajectory. The smaller this value, the more similar the two images (the rendered image under the new trajectory and the acquired image under the original trajectory), and the better the rendering quality. Based on the Figure 1 As can be seen from the rightmost metric shown in Figure 1 , the rendering effect (the metrics indicated by markers 17 and 18) of the scene reconstruction model trained by the method provided by the technical solution of the present disclosure is optimal.
[0110] Exemplary method
[0111] Figure 2 FIG. 9 is a flowchart of a method for training a scene reconstruction model provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to an electronic device, such as Figure 2 shown, and includes the following steps:
[0112] Step 201, obtain an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized with the image data set in time, where the first point cloud data set is collected by a lidar.
[0113] In this embodiment, the image data set collected by the vehicle under the original trajectory can be multiple frames of images collected by an image collection device such as a camera on the vehicle for the external environment during the vehicle driving process. It can be obtained by extracting frames from the video data collected during a driving process, or each frame image in the video data can be used as the image in the above image data set. The multiple frames of images in the image data set are driving data in the real driving environment, also called the GT (ground truth, real collected data) video under the original trajectory. The scene reconstruction model training tool, which is the execution subject of the training method of the first reconstruction model of the present disclosure, first obtains the GT video of the original trajectory and uses this video to train the first reconstruction model in subsequent steps.
[0114] Among them, the original trajectory is the collected trajectory. For example, it is obtained by positioning with the Global Positioning System (GPS) on the vehicle.
[0115] In this embodiment, the first point cloud dataset that is temporally synchronized with the image dataset refers to the three-dimensional point cloud collected by the lidar on the vehicle. The three-dimensional point cloud can be a set of three-dimensional points with color information, which is stored in a general three-dimensional point cloud format.
[0116] Among them, after obtaining the three-dimensional point cloud data, preprocessing operations such as point cloud thinning, homogenization, and denoising can be performed on the point cloud collected by the lidar. Through the preprocessing operations, the data volume of the three-dimensional point cloud data can be reduced, and the task processing efficiency can be improved.
[0117] In specific implementation, time synchronization is performed between the camera and the lidar, and calibration between the two is also performed to obtain the geometric mapping between the two.
[0118] Step 202: Use the image dataset and the first point cloud dataset to train and obtain the first reconstruction model.
[0119] In this embodiment, the training process of the first reconstruction model can be divided into the process of initializing the model parameters through the first point cloud dataset and the process of training the model through the image dataset to adjust the model parameters.
[0120] In specific implementation, the driving scene expression can be first decoupled into an initial road surface model, an initial non-road surface background model, and an initial dynamic object model; then the initial road surface model, the initial non-road surface background model, and the initial dynamic object model are combined into a second initial reconstruction model; and then the second initial reconstruction model is trained using the collected images under the original trajectories in the image dataset, and the first reconstruction model can be obtained.
[0121] Among them, the specific process of training the first reconstruction model can be specifically referred to Figure 5 the embodiments shown, which will not be elaborated here in detail.
[0122] Step 203: Use the first reconstruction model to render the image under the new trajectory to obtain the rendered image under the new trajectory.
[0123] In the embodiments of the present disclosure, a new trajectory can be first sampled. For example, based on the original trajectory, switch to a perspective where the lane is shifted 1.5 meters to the left or 2 meters to the right, and sample a new trajectory from this perspective.
[0124] In some embodiments, before using the first reconstruction model to render the image under the new trajectory, a new trajectory can also be sampled based on a preset sampling period, and the offset of the new trajectory sampled in each sampling period relative to the original trajectory increases sequentially.
[0125] In specific implementation, a progressive strategy can be adopted for sampling new trajectories, that is, during the process of model training, the offset of the new trajectory relative to the original trajectory gradually increases. For example, first based on the original trajectory, switch to the perspective of shifting the lane 1.5 meters to the left, sample a new trajectory from this perspective, and then render the image under the new trajectory to obtain the rendered image under the new trajectory. Then perform step 204 to repair the rendered image under the new trajectory using the video repair model to obtain the repaired image under the new trajectory. Then, based on a preset sampling period, switch from the original trajectory to the perspective of shifting the lane 3.0 meters to the left, and sample a new trajectory with a larger offset from this perspective. In the above sampling strategy, the offset of the new trajectory sampled in each sampling period increases successively relative to the original trajectory. First, train under the trajectory with a smaller offset, and then gradually switch to train under the trajectory with a larger offset, which can enable the model to gradually increase the learning difficulty and achieve better training results.
[0126] It can be understood that since the first reconstruction model is trained using the captured images under the original trajectories in the image dataset and is not sufficiently trained for the new trajectories, the quality of the rendered images under the new trajectories generated by the first reconstruction model is poor, and problems such as ghosting and blurring may exist. Therefore, the rendered images under the new trajectories can continue to be repaired through step 204.
[0127] Step 204: Use the pre-trained video repair model to perform repair processing on the rendered image under the new trajectory to obtain the repaired image under the new trajectory.
[0128] Among them, the pre-trained video repair model is an image used to perform repair processing on the rendered image under the new trajectory, so that the repaired image under the new trajectory can have a better rendering effect as a whole.
[0129] Specifically, in the dataset of autonomous driving, there is usually only the image dataset of the original trajectory, lacking the video repair dataset for training the video repair model. Therefore, a video repair dataset can be constructed by using the under-trained reconstruction model to render the images under the original trajectory. Specifically, use the under-trained scene reconstruction model to render images under the original trajectory, construct a video repair dataset with the captured images of each frame of the original trajectory, and use this dataset to train a video repair model.
[0130] Among them, the specific implementation method of training the video repair model can refer to Figure 6 the embodiments shown, which will not be elaborated here in detail.
[0131] Step 205: Use the repaired image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory to train a second reconstruction model, which is used to generate a rendered image under the new trajectory.
[0132] Among them, the Gaussian parameters of the first reconstruction model are used to define a three-dimensional Gaussian ellipsoid to represent the scene. The Gaussian parameters specifically include position, covariance, color, and transparency. The first pose of the vehicle under the original trajectory is used to indicate the position information and orientation information collected by the positioning module of the vehicle. The new trajectory of the vehicle is used to indicate a new trajectory sampled from the perspective of shifting the lane 1.5 meters to the left or 2 meters to the right based on the original trajectory. The second pose of the vehicle under the new trajectory is used to indicate the position information and orientation information of the vehicle in the new trajectory.
[0133] Among them, the specific implementation of training the second reconstruction model can be referred to Figure 4 the embodiments shown, which will not be elaborated here.
[0134] Through the above steps 201 - 205, for the autonomous driving scene reconstruction task, an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized with the image data set in time can be obtained; then, use the image data set and the first point cloud data set to train a first reconstruction model; and use the first reconstruction model to render an image under the new trajectory to obtain a rendered image under the new trajectory; then use a pre-trained video repair model to repair the rendered image under the new trajectory to obtain a repaired image under the new trajectory; finally, use the repaired image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory to train a second reconstruction model, which is used to generate a rendered image under the new trajectory. Thus, by using the repaired image under the new trajectory repaired by the video repair model as a training sample to further train the scene reconstruction model, it helps to bridge the domain gap between the repaired image under the new trajectory and the collected image under the original trajectory, and improve the rendering quality of the reconstruction model and the robustness of the model.
[0135] In some optional implementation manners, the training processes of the first reconstruction model and the second reconstruction model can be referred to Figure 3 .
[0136] Figure 3In it, first, the Dynamic Driving Scene Representation is decoupled into a Ground Model, a Background Model, and an Object Model, and then through Composition processing, a second initial reconstruction model (Gaussians) is obtained. By inputting the internal and external camera parameters when acquiring the captured images of each original trajectory into the second initial reconstruction model (Gaussians), the rendered images under the original trajectories can be rendered. The rendered images under multiple frames of original trajectories are stitched together to obtain the Original Trajectory Rendered Video. After that, the loss function of the scene reconstruction model is used to constrain the model. Specifically, according to the output result of the above second initial reconstruction model and the captured images under the original trajectories, the loss value of the second initial reconstruction model (Gaussians) is calculated, and the model parameters of the second initial reconstruction model are adjusted according to the loss value to obtain the first reconstruction model.
[0137] The second reconstruction model (Novel Trajectory Deformable Network) consists of a Pose Feature Module and a Temporal Field Module. The input to the Temporal Field Module is the Gaussian parameter position encoding corresponding to the Gaussian parameters of the first reconstruction model and the temporal step position encoding corresponding to the temporal step (which is used to indicate the frame number of the captured image under the original trajectory in the image dataset). The Temporal Field Module implicitly mixes the Gaussian parameter position encoding and the temporal step position encoding to obtain the first feature. The input to the Pose Feature Module is the Delta Pose determined based on the first pose of the vehicle under the original trajectory and the second pose of the vehicle under the new trajectory. After the Pose Feature Module performs feature extraction and transformation on this position change amount, the second feature is obtained. After the first feature and the second feature are element-wise added through an output linear layer, the Gaussian parameter change amount can be obtained. By adding the Gaussian parameters of the first reconstruction model and the Gaussian parameter change amount element-wise again, the first initial reconstruction model (Updated Gaussians) can be obtained. By inputting the internal and external camera parameters corresponding to each new trajectory into the first initial reconstruction model (Updated Gaussians), it is possible to render the rendered images under the new trajectory using the first initial reconstruction model. The rendered images under multiple frames of the original trajectory are stitched together to obtain the rendered video (Novel Trajectory Rendered Video) under the new trajectory. Then, the loss function of the scene reconstruction model is used to constrain the model. Specifically, according to the output result of the first initial reconstruction model and the restored image under the new trajectory (the image extracted from the Novel Trajectory Restored Video), the loss value of the first initial reconstruction model (Updated Gaussians) is calculated, and the model parameters of the first initial reconstruction model are adjusted according to the loss value to obtain the second reconstruction model.
[0138] Among them, the restored image under the new trajectory (the image extracted from the Novel Trajectory Restored Video) is an image generated by a video restoration model (Drive Restorer) based on the degraded video (Novel Trajectory Degraded Video) under the new trajectory. Each frame image in the Novel Trajectory Degraded Video is the rendered image under the new trajectory rendered using the first reconstruction model.
[0139] Based on Figure 3As can be seen from the schematic diagram in , the first reconstruction model is obtained by first decoupling the dynamic driving scene representation into a road surface model, a non-road surface background model, and a dynamic object model and initializing the model parameters, and then training. By separately modeling the road surface model, the rendering effect of the road surface from different perspectives can be significantly improved, its spatial consistency can be enhanced, and further the rendering effect and quality of the image under the original trajectory can be enhanced. For the training of the second reconstruction model, the Gaussian parameters are updated on the basis of the first reconstruction model, so that the trained second reconstruction model can bridge the domain gap between the generated data and the original data and improve the rendering effect and quality of the image under the new trajectory.
[0140] Figure 4 is a schematic flowchart of step 205 provided by an exemplary embodiment of the present disclosure. On the basis of the embodiment shown in , the specific implementation manner of training the second reconstruction model in step 205 includes the following steps: Figure 2 On the basis of the embodiment shown in , the specific implementation manner of training the second reconstruction model in step 205 includes the following steps:
[0141] Step 251, perform position encoding on the Gaussian parameters and time steps of the first reconstruction model respectively to obtain Gaussian parameter position encoding and time step position encoding. The time step is used to indicate the frame number of the acquired image under the original trajectory in the image dataset.
[0142] Among them, after performing position encoding on the Gaussian parameters (g) of the first reconstruction model, Gaussian parameter position encoding can be obtained; after performing position encoding on the time step, time step position encoding can be obtained.
[0143] Since when rendering an image using the reconstruction model, the rendering operation is performed for each frame of the image. In order to be able to generate a video based on the rendered image, when rendering each frame of the image, position encoding can be performed on the frame numbers of each frame of the image. Thus, after generating each frame of the rendered image, the frame numbers of different rendered images can be distinguished according to the time step position encoding parameters.
[0144] Step 252, use the first multi-layer perceptron to obtain a first feature based on the Gaussian parameter position encoding and the time step position encoding.
[0145] Among them, the first multi-layer perceptron corresponds to the Figure 3 time field module in , and is used to implicitly mix the Gaussian parameter position encoding and the time step position encoding and output a first feature. Specifically, the mixing method can be various processes such as forward propagation, loss calculation, backpropagation, and weight update for the Gaussian parameter position encoding and the time step position encoding.
[0146] Step 253, use the second multi-layer perceptron to obtain a second feature based on the first pose of the vehicle under the original trajectory and the second pose of the vehicle under the new trajectory.
[0147] In the embodiments of the present disclosure, the pose change amount can be determined based on the first pose and the second pose, and specifically, it can be calculated by Equation (1).
[0148]
[0149] In Equation (1), L is a hyperparameter for normalizing the changed pose, is the position and pose at each time step under the new trajectory, is the position and pose at each time step under the original trajectory.
[0150] Among them, the second multi-layer perceptron corresponds to Figure 3 the pose feature module in, and is used to perform various processes such as forward propagation, loss calculation, backpropagation, and weight update on Δp t to obtain the second feature.
[0151] Step 254: Based on the first feature, the second feature, and the Gaussian parameters of the first reconstruction model, obtain the first initial reconstruction model.
[0152] In this embodiment, the above first feature and second feature can be added through an output linear layer to obtain the change amount of the Gaussian parameters. Specifically, refer to Equation (2).
[0153]
[0154] In Equation (1), is the output of the pose feature module (the second feature), F θ (PE(g), PE(t)) is the output of the time field module (the first feature), F out is the output linear layer.
[0155] After obtaining the change amount of the above Gaussian parameters through Equation (2), the updated Gaussian parameter g′ = g + Δg can be obtained based on the change amount Δg of the Gaussian parameters and the Gaussian parameter (g) of the first reconstruction model.
[0156] Step 255: Use the repaired image under the new trajectory to train the first initial reconstruction model to obtain the second reconstruction model.
[0157] In some embodiments, after obtaining the above updated Gaussian parameters, the internal and external camera parameters corresponding to each new trajectory can be input into the first initial reconstruction model, so as to render the rendering image under the new trajectory by using the first initial reconstruction model, and then use the loss function of the scene reconstruction model to constrain the model. Specifically, according to the output result of the above first initial reconstruction model and the repaired image under the new trajectory, calculate the loss value of the first initial reconstruction model, and adjust the model parameters of the first initial reconstruction model according to the loss value to obtain the second reconstruction model.
[0158] Through the above steps 251 - 255, the technical solution of the present disclosure discloses an implementation manner of training the second reconstruction model. By updating the Gaussian parameters of the second reconstruction model, it helps to bridge the domain gap between the rendered images under the new trajectory generated by the second reconstruction model and the acquired images under the original trajectory, and improve the rendering quality of the images under the new trajectory.
[0159] Figure 5 is a schematic flowchart of step 202 provided by an exemplary embodiment of the present disclosure; on the basis of the Figure 2 illustrated embodiment, the specific implementation manner of training the first reconstruction model in step 202 includes the following steps:
[0160] Step 221, fuse multiple frames of point cloud in the first point cloud dataset to obtain the fused point cloud data.
[0161] In the embodiments of the present disclosure, the point cloud data of different frames can be first subjected to point cloud registration processing to align them in the world coordinate system, and then the points corresponding to the same object in different frames are determined based on the spatial position relationship, the motion model of the object, etc., thereby establishing the association relationship of the point cloud data of different frames; then, according to the association relationship of the point cloud data of different frames, the point cloud data of different frames are fused.
[0162] Step 222, respectively segment the road surface point cloud data, the non - road surface static background point cloud data, and the dynamic object point cloud data from the fused point cloud data.
[0163] In the embodiments of the present disclosure, the road surface point cloud data, the non - road surface static background point cloud data, and the dynamic object point cloud data can be segmented from the fused point cloud data through a clustering algorithm.
[0164] Specifically, the points in the point cloud data can be clustered according to the positions and features of each point. Specifically, the features of each point in the point space or local space can be extracted or transformed to obtain various attributes, such as normal vector, density, distance, elevation, intensity, etc., thereby realizing the segmentation of the point cloud data with different attributes.
[0165] Further, after segmenting the road surface point cloud data, the non - road surface static background point cloud data, and the dynamic object point cloud data, the initial road surface model, the initial non - road surface background model, and the initial dynamic object model can be generated through steps 223, 224, and 225.
[0166] Step 223, initialize the parameters of the three - dimensional Gaussian splash model based on the road surface point cloud data to obtain the initial road surface model.
[0167] Among them, the implementation of step 223 specifically includes the following steps:
[0168] Step 1: Obtain the central positions of each sub - point cloud data to get the central positions of each Gaussian sphere.
[0169] Step 2: Determine the central positions of each Gaussian sphere and the number of Gaussian spheres as the prior information of the three - dimensional structure.
[0170] Step 3: Integrate the prior information of the three - dimensional structure into the three - dimensional Gaussian splash model to obtain the initial road surface model. The prior information of the three - dimensional structure remains fixed during the operation of training the second initial reconstruction model.
[0171] Among them, the initial road surface model is represented by a set of Gaussian distributions in the world coordinate system. Each Gaussian distribution consists of Gaussian parameters (central position x, opacity γ, covariance matrix ∑, and RGB color parameters), where the RGB color parameters are controlled by spherical harmonic functions.
[0172] In the embodiments of the present disclosure, to ensure stability, each covariance matrix is decomposed into ∑ = RSS T R T , where the scaling matrix S and the rotation matrix R are learnable parameters, represented by a scaling factor and a quaternion respectively.
[0173] Specifically, the central positions of each sub - point cloud data can be determined as the central positions of each Gaussian sphere, and the central positions of each Gaussian sphere and the number of Gaussian spheres are determined as the prior information of the three - dimensional structure, and this prior information of the three - dimensional structure is integrated into the three - dimensional Gaussian splash model to obtain the initial road surface model.
[0174] It should be noted that since the road surface point cloud data in the point cloud data obtained by fusing multiple frames of point cloud data is rich, the present disclosure's technical solution directly fixes the number and position parameters (central position) of the Gaussian distributions in the road surface model. During the subsequent process of training the second initial reconstruction model with the acquired images under the original trajectory to obtain the first reconstruction model, this prior information of the three - dimensional structure also remains unchanged, and no modification is made to the number and central positions of the Gaussian spheres, thereby significantly reducing the search space of the second initial reconstruction model and enhancing its generalization ability.
[0175] Step 224: Initialize the parameters of the three - dimensional Gaussian splash model based on the non - road - surface static background point cloud data to obtain the initial non - road - surface background model.
[0176] In some embodiments, in the world coordinate system, the non - road - surface static background point cloud data can be used to initialize the parameters of the three - dimensional Gaussian splash model, that is, to initialize parameters such as the number and central positions of the Gaussian spheres to obtain the initial non - road - surface background model.
[0177] It should be noted that in the subsequent process of training the second initial reconstruction model with the collected images under the original trajectory to obtain the first reconstruction model, the Gaussian parameters of the initial non-road background model can be optimized.
[0178] Step 225: Initialize the parameters of the three-dimensional Gaussian splash model based on the dynamic object point cloud data to obtain an initial dynamic object model.
[0179] Among them, the number of dynamic object models is the same as the number of dynamic objects segmented from the point cloud data. For each dynamic object, a dynamic object model can be defined separately in the local coordinate system of each dynamic object.
[0180] In this embodiment, the point cloud data of each dynamic object can be used to initialize the parameters of the three-dimensional Gaussian splash model, that is, to initialize parameters such as the number and center position of the Gaussian spheres, so as to obtain the initial dynamic object model in the local coordinate system of each dynamic object.
[0181] Step 226: Combine the initial road model, the initial non-road background model, and the initial dynamic object model to obtain a second initial reconstruction model.
[0182] Among them, the implementation of step 226 specifically includes the following steps:
[0183] Step 1: Based on the pose information of each dynamic object in the world coordinate system, obtain the rotation matrix and translation vector for transforming each dynamic object from the local coordinate system to the world coordinate system;
[0184] Step 2: Based on the rotation matrix and translation vector, perform a transformation process on the initial dynamic object model to obtain a transformed dynamic object model, and the transformed dynamic object model is a model based on the world coordinate system;
[0185] Step 3: Combine the initial road model, the initial non-road background model, and the transformed dynamic object model to obtain a second initial reconstruction model.
[0186] Among them, since each dynamic object model is a model established based on the local coordinate system, before the combination process, it is necessary to first perform a transformation process on each initial dynamic model and transform it to the world coordinate system.
[0187] Specifically, since the positions and orientations of each dynamic object in the world coordinate system are marked in the data set, the rotation matrix and translation vector for transforming the dynamic object to the world coordinate system can be determined according to the local coordinate system of each dynamic object and the positions and orientations of each dynamic object in the world coordinate system, and then the Gaussian distribution position x for performing a transformation process on the initial dynamic model and transforming it to the world coordinate system can be obtained. w and the rotation quaternion qw . Specifically, reference can be made to Equation (3) and Equation (4).
[0188] x w = R t x o + T t Equation (3)
[0189] q w = ROT(R t , q o ) Equation (4)
[0190] In the above Equation (3), x o represents the position of the Gaussian distribution in the local coordinate system, T t represents the translation vector, and R t represents the rotation matrix.
[0191] In the above Equation (4), R t represents the rotation matrix, q o represents the rotation quaternion of the Gaussian distribution in the local object coordinate system, and ROT(·) represents the operation of rotating the quaternion by the rotation matrix.
[0192] After converting the above dynamic object model to the world coordinate system, the above converted dynamic object model, the initial road surface model, and the initial non-road surface background model can be combined to obtain the second initial reconstruction model.
[0193] Step 227: Use the acquisition images under each frame of the original trajectory in the image dataset to train the second initial reconstruction model to obtain the first reconstruction model.
[0194] In this embodiment, by inputting the internal and external camera parameters when acquiring the images of each original trajectory into the second initial reconstruction model, the rendered images under the original trajectory can be rendered. The rendered images under multiple frames of the original trajectory can be stitched together to obtain the rendered video under the original trajectory, and then the loss function of the scene reconstruction model is used to constrain the model. Specifically, according to the output result of the above second initial reconstruction model and the acquisition images under the original trajectory, the loss value of the second initial reconstruction model is calculated, and the model parameters of the second initial reconstruction model are adjusted according to the loss value to obtain the first reconstruction model.
[0195] Through the above-mentioned Step 221 - Step 227, the technical solution of the present disclosure discloses a specific implementation manner of obtaining a first reconstruction model by decoupling the driving scene representation into a road surface model, a non-road surface background model, and a dynamic object model, and then performing combined training. Since the technical solution of the present disclosure realizes the separate modeling of the road surface model, it helps to improve the rendering results of the road surface from different perspectives and enhance the spatial consistency of the road surface. In addition, during the training process, by fixing the number of Gaussian distributions and the central positions of the road surface model, the search space of the reconstruction model can be significantly reduced, and the generalization ability of the model can be enhanced.
[0196] Figure 6 is a schematic flowchart of the training process of a video reconstruction model provided by an exemplary embodiment of the present disclosure; as Figure 6 shown, the steps of training to obtain a video restoration model include Step 601 and Step 602. Each step will be described below.
[0197] Step 601, using the acquisition images under the original trajectories of each frame in the image dataset to insufficiently train the second initial reconstruction model to obtain a first reconstruction model with insufficient training.
[0198] In some embodiments, insufficient training is used to indicate that the number of iterations is not enough, the loss value between the acquisition images under the original trajectories and the rendered images under the original trajectories is still relatively large, and the model has not yet fully converged.
[0199] Step 602, using the first reconstruction model with insufficient training and the acquisition images under the original trajectories of each frame in the image dataset to train the initial diffusion model to obtain a video restoration model.
[0200] In some embodiments, the first reconstruction model with insufficient training can be used to render the images under each frame of the original trajectory to obtain the corresponding rendered images under each frame of the original trajectory; then, using the rendered images under the original trajectory and the acquisition images under the original trajectory, a video restoration dataset is constructed; finally, using the video restoration dataset to train the initial diffusion model to obtain a video restoration model.
[0201] Specifically, during the process of using the acquisition images under the original trajectory to train the second initial reconstruction model to obtain the first reconstruction model, when the model has not been fully trained, the video under the original trajectory can be directly rendered to obtain a degraded video of the original trajectory, that is, a rendered image with poor effect. This image and the acquisition image under the original trajectory form a video restoration pair to complete the construction of the video restoration dataset. Among them, the acquisition image under the original trajectory can be used as the supervision data for model training during the training process to supervise and guide the learning process of the video restoration model.
[0202] In the embodiments of the present disclosure, a method for constructing a dataset for training a video restoration model is provided by rendering the captured images under the original trajectory using a scene reconstruction model with insufficient training and constructing a video restoration dataset based on the rendered images under the original trajectory. By using this method to construct the dataset, the problem of insufficient training data in the existing autonomous driving dataset can be solved, so that the video restoration model can be fully trained using this dataset, improving the restoration effect of the model. At the same time, the trained video restoration model can assist the scene reconstruction model in image rendering, thereby improving the rendering quality.
[0203] Exemplary device
[0204] Figure 7 FIG. is a schematic structural diagram of a training device for a scene reconstruction model provided by an exemplary embodiment of the present disclosure. As Figure 7 shown, the device may include:
[0205] A first acquisition module 71, configured to acquire an image dataset captured by the vehicle under the original trajectory and a first point cloud dataset synchronized with the image dataset in time, where the first point cloud dataset is acquired by a lidar;
[0206] A first training module 72, configured to train a first reconstruction model using the image dataset and the first point cloud dataset;
[0207] A first rendering module 73, configured to render the images under the new trajectory using the first reconstruction model to obtain the rendered images under the new trajectory;
[0208] A restoration module 74, configured to perform a restoration process on the rendered images under the new trajectory using a pre-trained video restoration model to obtain the restored images under the new trajectory;
[0209] A second training module 75, configured to train a second reconstruction model using the restored images under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory, where the second reconstruction model is used to generate the rendered images under the new trajectory.
[0210] Figure 8 FIG. is a schematic structural diagram of a training device for a scene reconstruction model provided by another exemplary embodiment of the present disclosure. As Figure 8 shown, on the basis of the embodiment shown in Figure 7 In some embodiments, the second training module 75 includes:
[0211] An encoding sub-module 751, configured to perform position encoding on the Gaussian parameters and time steps of the first reconstruction model respectively to obtain Gaussian parameter position encoding and time step position encoding, where the time step is used to indicate the frame number of the captured images under the original trajectory in the image dataset;
[0212] The first processing sub-module 752 is configured to use a first multi-layer perceptron to obtain a first feature based on Gaussian parameter position encoding and time step position encoding;
[0213] The second processing sub-module 753 is configured to use a second multi-layer perceptron to obtain a second feature based on the first pose of the vehicle under the original trajectory and the second pose of the vehicle under the new trajectory;
[0214] The parameter update sub-module 754 is configured to obtain a first initial reconstruction model based on the first feature, the second feature, and the Gaussian parameters of the first reconstruction model;
[0215] The first training sub-module 755 is configured to use the repaired images under the new trajectory to train the first initial reconstruction model to obtain a second reconstruction model.
[0216] In some embodiments, the apparatus further includes:
[0217] The third training module 76 is configured to use the acquired images of each frame under the original trajectory in the image dataset to insufficiently train the second initial reconstruction model to obtain a first reconstruction model with insufficient training;
[0218] The fourth training module 77 is configured to use the first reconstruction model with insufficient training and the acquired images of each frame under the original trajectory in the image dataset to train the initial diffusion model to obtain a video repair model.
[0219] In some embodiments, the fourth training module 77 includes:
[0220] The first rendering sub-module 771 is configured to use the first reconstruction model with insufficient training to render the images of each frame under the original trajectory to obtain the corresponding rendered images of each frame under the original trajectory;
[0221] The dataset construction sub-module 772 is configured to use the rendered images under the original trajectory and the acquired images under the original trajectory to construct a video repair dataset;
[0222] The third training sub-module 773 is configured to use the video repair dataset to train the initial diffusion model to obtain a video repair model.
[0223] Figure 9 It is a schematic structural diagram of a training apparatus for a scene reconstruction model provided by another exemplary embodiment of the present disclosure. As Figure 9 shown, on the basis of the embodiments shown in Figure 7 and / or Figure 8 In some embodiments, the first training module 72 includes:
[0224] The point cloud fusion sub-module 721 is used to fuse multiple frames of point clouds in the first point cloud dataset to obtain the fused point cloud data;
[0225] The point cloud segmentation sub-module 722 is used to respectively segment the road surface point cloud data, the non-road surface static background point cloud data, and the dynamic object point cloud data from the fused point cloud data;
[0226] The first initialization sub-module 723 is used to initialize the parameters of the three-dimensional Gaussian splash model based on the road surface point cloud data to obtain the initial road surface model;
[0227] The second initialization sub-module 724 is used to initialize the parameters of the three-dimensional Gaussian splash model based on the non-road surface static background point cloud data to obtain the initial non-road surface background model;
[0228] The third initialization sub-module 725 is used to initialize the parameters of the three-dimensional Gaussian splash model based on the dynamic object point cloud data to obtain the initial dynamic object model;
[0229] The combination sub-module 726 is used to perform a combination process on the initial road surface model, the initial non-road surface background model, and the initial dynamic object model to obtain the second initial reconstruction model;
[0230] The second training sub-module 727 is used to train the second initial reconstruction model by using the acquisition images under the original trajectories in the image dataset to obtain the first reconstruction model.
[0231] In some embodiments, the road surface point cloud data includes multiple sub-point cloud data;
[0232] The first initialization sub-module 723 includes:
[0233] The first acquisition unit 7231 is used to acquire the central positions of the sub-point cloud data to obtain the central positions of the Gaussian spheres;
[0234] The prior information determination unit 7232 is used to determine the central positions of the Gaussian spheres and the number of Gaussian spheres as the three-dimensional structure prior information;
[0235] The integration unit 7233 is used to integrate the three-dimensional structure prior information into the three-dimensional Gaussian splash model to obtain the initial road surface model. The three-dimensional structure prior information is fixed and unchanged during the operation of training the second initial reconstruction model.
[0236] In some embodiments, the initial dynamic object model is a model based on the local coordinate system of the dynamic object;
[0237] The combination sub-module 726 includes:
[0238] A second acquisition unit 7261, configured to obtain a rotation matrix and a translation vector for transforming each dynamic object from a local coordinate system to a world coordinate system based on the pose information of each dynamic object in the world coordinate system;
[0239] A conversion unit 7262, configured to perform a conversion process on the initial dynamic object model based on the rotation matrix and the translation vector to obtain a converted dynamic object model, where the converted dynamic object model is a model based on the world coordinate system;
[0240] A combination unit 7263, configured to perform a combination process on the initial road surface model, the initial non-road surface background model, and the converted dynamic object model to obtain a second initial reconstruction model.
[0241] In some embodiments, the apparatus further includes:
[0242] A sampling module 78, configured to sample new trajectories based on a preset sampling period, and the offset of the new trajectories sampled in each sampling period increases sequentially with respect to the original trajectory.
[0243] Exemplary electronic device
[0244] Figure 10 A structural diagram of an electronic device provided in an embodiment of the present disclosure, including at least one processor 101 and a memory 102.
[0245] The processor 101 may be a central processing unit (CPU) or other form of processing unit having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0246] The memory 102 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage media, and the processor 101 may run one or more computer program instructions to implement the training method of the scene reconstruction model in various embodiments of the present disclosure above and / or other desired functions.
[0247] In one example, the electronic device 10 may further include: an input device 103 and an output device 104, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0248] The input device 103 may further include, for example, a keyboard, a mouse, etc.
[0249] The output device 104 can output various information to the outside, which may include, for example, a display, a speaker, a printer, a communication network, and remote output devices connected thereto, etc.
[0250] Of course, for simplicity, Figure 10 only some of the components of the electronic device 10 related to the present disclosure are shown in the figure, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.
[0251] Exemplary computer program product and computer-readable storage medium
[0252] In addition to the above methods and devices, embodiments of the present disclosure may also provide a computer program product, including computer program instructions, which when run by a processor cause the processor to execute the steps in the training method of the scenario reconstruction model of various embodiments of the present disclosure described in the above "Exemplary Method" section.
[0253] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0254] Furthermore, embodiments of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored, which when run by a processor cause the processor to execute the steps in the training method of the scenario reconstruction model of various embodiments of the present disclosure described in the above "Exemplary Method" section.
[0255] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes systems, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductors, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0256] The basic principles of the present disclosure have been described in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for the purposes of illustration and facilitating understanding, rather than limitations. The above details do not limit the present disclosure to necessarily implement using the above specific details.
[0257] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.
Claims
1. A method for training a scene reconstruction model, comprising: Acquire an image dataset collected by the vehicle under the original trajectory and a first point cloud dataset synchronized with the image dataset in time, wherein the first point cloud dataset is collected by a laser radar; Using the image dataset and the first point cloud dataset, training a first reconstruction model; Rendering the image under the new trajectory using the first reconstruction model to obtain a rendered image under the new trajectory; Using a pre-trained video restoration model to restore the rendered image under the new trajectory, to obtain a restored image under the new trajectory; A second reconstruction model is trained using the restored image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory. The second reconstruction model is used to generate a rendered image under the new trajectory.
2. The method according to claim 1, characterized in that The training of obtaining a second reconstruction model by using the restored image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory comprises: Position encoding is performed on the Gaussian parameters and the time step of the first reconstruction model respectively to obtain Gaussian parameter position encoding and time step position encoding, wherein the time step is used to indicate the frame number of the collected image under the original trajectory in the image data set; Using the first multi-layer perceptron, based on the Gaussian parameter position encoding and the time step position encoding, a first feature is obtained; Using a second multi-layer perceptron, based on the first posture of the vehicle under the original trajectory and the second posture of the vehicle under the new trajectory, a second feature is obtained; Obtaining a first initial reconstruction model based on the first feature, the second feature, and the Gaussian parameters of the first reconstruction model; The first initial reconstruction model is trained using the restored image under the new trajectory to obtain the second reconstruction model.
3. The method according to any one of claims 1-2, characterized in that: The step of training a first reconstruction model using the image dataset and the first point cloud dataset comprises: Fusing multiple frames of point cloud in the first point cloud data set to obtain fused point cloud data; Separating the road surface point cloud data, non-road surface static background point cloud data and dynamic object point cloud data from the fused point cloud data; Initializing the parameters of the three-dimensional Gaussian splash model based on the road surface point cloud data to obtain an initial road surface model; Initializing the parameters of the three-dimensional Gaussian splash model based on the non-road static background point cloud data to obtain an initial non-road background model; Initializing the parameters of the three-dimensional Gaussian splash model based on the dynamic object point cloud data to obtain an initial dynamic object model; Combining the initial road surface model, the initial non-road surface background model and the initial dynamic object model to obtain a second initial reconstruction model; The second initial reconstruction model is trained using the collected images under the original trajectory of each frame in the image data set to obtain the first reconstruction model.
4. The method according to claim 3, characterized in that After obtaining the second initial reconstruction model, the method further includes: Using the collected images under the original trajectory of each frame in the image data set, the second initial reconstruction model is insufficiently trained to obtain an insufficiently trained first reconstruction model; The initial diffusion model is trained using the insufficiently trained first reconstruction model and the collected images under the original trajectory of each frame in the image data set to obtain the video restoration model.
5. The method according to claim 4, characterized in that The step of training the initial diffusion model using the first reconstruction model that is not sufficiently trained and the images of each frame under the original trajectory in the image data set to obtain the video restoration model includes: Rendering the image of each frame under the original trajectory using the first reconstruction model that is not fully trained, to obtain the corresponding rendered image of each frame under the original trajectory; Constructing a video restoration dataset using the rendered image under the original trajectory and the captured image under the original trajectory; The video restoration data set is used to train the initial diffusion model to obtain the video restoration model.
6. The method according to claim 3, characterized in that: The road surface point cloud data includes a plurality of sub-point cloud data; The initializing the parameters of the three-dimensional Gaussian splash model based on the road surface point cloud data to obtain an initial road surface model includes: Get the center position of each sub-point cloud data, and get the center position of each Gaussian sphere; The center position of each Gaussian sphere and the number of Gaussian spheres are determined as prior information of the three-dimensional structure; The three-dimensional structural prior information is integrated into the three-dimensional Gaussian splash model to obtain the initial road surface model, and the three-dimensional structural prior information is fixed during the operation of training the second initial reconstruction model.
7. The method according to claim 3, characterized in that The initial dynamic object model is a model based on the local coordinate system of the dynamic object; The combining and processing the initial road surface model, the initial non-road surface background model and the initial dynamic object model to obtain a second initial reconstruction model includes: Based on the position information of each dynamic object in the world coordinate system, a rotation matrix and a translation vector for transforming each dynamic object from the local coordinate system to the world coordinate system are obtained; Based on the rotation matrix and the translation vector, the initial dynamic object model is transformed to obtain a transformed dynamic object model, wherein the transformed dynamic object model is a model based on a world coordinate system; The initial road surface model, the initial non-road surface background model and the converted dynamic object model are combined to obtain the second initial reconstruction model.
8. The method according to any one of claims 1 to 7, characterized in that: Before rendering the image under the new trajectory using the first reconstruction model to obtain the rendered image under the new trajectory, the method further includes: The new trajectory is sampled based on a preset sampling period, and the offset of the new trajectory sampled in each sampling period relative to the original trajectory increases sequentially.
9. A training device for a scene reconstruction model, comprising: A first acquisition module is used to acquire an image data set collected by the vehicle under the original trajectory and a first point cloud data set synchronized with the image data set in time, wherein the first point cloud data set is acquired by a laser radar; A first training module, used for training a first reconstruction model using the image dataset and the first point cloud dataset; A first rendering module, used for rendering an image under a new trajectory using the first reconstruction model to obtain a rendered image under the new trajectory; A restoration module, used to perform restoration processing on the rendered image under the new trajectory using a pre-trained video restoration model to obtain a restored image under the new trajectory; The second training module is used to train a second reconstruction model using the restored image under the new trajectory, the Gaussian parameters of the first reconstruction model, the first pose of the vehicle under the original trajectory, and the second pose of the vehicle under the new trajectory, and the second reconstruction model is used to generate a rendered image under the new trajectory.
10. An electronic device, characterized in that: include: A memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, it implements the method described in any one of claims 1 to 8.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 8 is implemented.
12. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Indoor dynamic environment map construction method and system based on image restoration and completion
CN118822906A
Methods and Apparatus for Video Completion
US20130128121A1
Cited By
Video generation method and device, electronic equipment, storage medium and program product
CN120434373A
Video generation methods, apparatuses, electronic devices, storage media, and software products
CN120434373B
Scene reconstruction method and device, electronic equipment, storage medium and program product
CN120635331A
Scene reconstruction method and device, electronic equipment, storage medium and program product
CN120635331B
Object-oriented periodic dynamic motion 4D Gaussian splash reconstruction method
CN121982187A