Scene reconstruction method and device, electronic equipment, storage medium and program product

By obtaining the initial trajectory image in the autonomous driving scene reconstruction, performing modeling and trajectory reconstruction, and using motion model verification and dynamic object separation technology, a new trajectory scene video that conforms to the laws of physics is generated, which solves the data sparsity problem and improves the performance and rendering quality of the scene reconstruction model.

CN120635331AActive Publication Date: 2025-09-12北京极佳视界科技有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510939326.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-12
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

In autonomous driving scene reconstruction, due to data sparsity, existing methods can only generate high-quality output within the area covered by the recorded camera trajectory, cannot reconstruct rich multi-view scenes, and the training results are poor.

Method used

By acquiring scene images under the initial trajectory, modeling and trajectory reconstruction are performed, and the reconstruction results are verified using motion models, new trajectory scene videos that conform to physical laws are generated. Long-tail scene data is generated through dynamic object separation and predefined interaction rules, and the scene is enriched using dynamic adversarial agents and trajectory data enhancement technology.

Benefits of technology

It achieves scene expansion and generates new trajectory videos that conform to actual conditions, which is suitable for autonomous driving algorithm training and improves the performance and rendering quality of the scene reconstruction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635331A_ABST
    Figure CN120635331A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a scene reconstruction method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: obtaining at least one frame of first scene image of a driving scene under an initial track; modeling based on the at least one frame of first scene image to obtain an initial scene video; performing track reconstruction based on the initial scene video to obtain at least one second scene video under a new track; and verifying the at least one second scene video through a motion model to obtain at least one target scene video passing verification. According to the embodiment, various new tracks which are not collected in an actual scene can be obtained through track reconstruction, scene expansion is achieved, and motion model verification is conducted on the reconstructed second scene video correspondingly, so that the motion of the second scene video corresponding to the obtained new tracks conforms to the physical law and better conforms to the actual situation, and the scene expansion efficiency is improved. And the method is more suitable for training an automatic driving algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to scene reconstruction technology, and in particular to a scene reconstruction method, device, electronic device, storage medium, and program product. Background Art

[0002] In the task of reconstructing autonomous driving scenes, the scene is typically reconstructed using a given driving video input. Compared to traditional 3D reconstruction, 3D reconstruction in autonomous driving scenarios faces greater challenges. This is primarily due to the sparse perspective of autonomous driving data. Compared to traditional 3D reconstruction, autonomous driving datasets typically lack rich multi-perspective images. Their test sets are typically generated by extracting frames from the original trajectory data, while the training set uses the remaining images after the extraction. This forces existing methods to focus primarily on reconstruction quality based on the original trajectory. Consequently, end-to-end autonomous driving algorithm training can only be performed based on the original trajectory, which results in poor training results due to the limited number of training samples. Summary of the Invention

[0003] In order to solve the above technical problems, the present disclosure is proposed. Embodiments of the present disclosure provide a scene reconstruction method, apparatus, electronic device, storage medium, and program product.

[0004] According to one aspect of an embodiment of the present disclosure, a scene reconstruction method is provided, comprising:

[0005] Acquire at least one frame of a first scene image of the driving scene under the initial trajectory;

[0006] Performing modeling based on the at least one frame of the first scene image to obtain an initial scene video;

[0007] Reconstructing a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory;

[0008] The at least one second scene video is verified using a motion model to obtain at least one target scene video that passes the verification.

[0009] Optionally, the performing modeling based on the at least one frame of the first scene image to obtain an initial scene video includes:

[0010] Performing dynamic separation processing on the at least one frame of the first scene image to obtain a first static image and at least one first dynamic image group; each first dynamic image group includes at least one frame of dynamic image, and each first dynamic image group corresponds to a dynamic object;

[0011] Modeling each group of the first dynamic images to obtain at least one dynamic trajectory;

[0012] The at least one dynamic trajectory is converted into a background coordinate system corresponding to the first static image to obtain the initial scene video.

[0013] Optionally, the dynamic objects include the vehicle and at least one other vehicle, and the modeling of each group of the first dynamic image groups to obtain at least one dynamic trajectory includes:

[0014] Classifying the at least one first dynamic image group based on the dynamic object, and determining the first dynamic image group corresponding to the own vehicle and the first dynamic image group corresponding to related vehicles; the related vehicles being the other vehicles closest to the own vehicle;

[0015] The at least one dynamic trajectory corresponding to the ego vehicle is determined based on predefined interaction rules and predefined behaviors corresponding to the related vehicles.

[0016] Optionally, acquiring at least one frame of a first scene image of the driving scene under the initial trajectory includes:

[0017] Collecting at least one frame of a third scene image of the driving scene under the initial trajectory;

[0018] The at least one frame of the first scene image is obtained based on the at least one frame of the third scene image.

[0019] Optionally, obtaining the at least one frame of the first scene image based on the at least one frame of the third scene image includes:

[0020] performing interpolation processing on the initial trajectory corresponding to the at least one frame of the third scene image to obtain the at least one frame of the first scene image; and / or,

[0021] A lateral translation operation is performed on each of the at least one position information corresponding to the at least one frame of the third scene image in the initial trajectory to obtain the at least one frame of the first scene image.

[0022] Optionally, reconstructing a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory includes:

[0023] The driving scene reconstruction model is used to reconstruct the trajectory of the initial scene video to obtain the at least one second scene video under a new trajectory.

[0024] Optionally, verifying the at least one second scene video using a motion model to obtain at least one target scene video that passes the verification includes:

[0025] For each second scene video, estimating a next frame estimated scene image based on a current frame scene image in the second scene video by using a motion model;

[0026] Matching a next frame of scene image in the second scene video with the next frame of estimated scene image to determine a verification result;

[0027] The at least one target scene video is determined according to the verification result.

[0028] Optionally, estimating a next frame estimated scene image from a current frame scene image in the second scene video using a motion model includes:

[0029] Determine the first position information corresponding to the vehicle based on the current frame scene image;

[0030] Perform pose estimation on the first pose information based on the motion model to obtain second estimated pose information corresponding to the next frame estimated scene image.

[0031] Optionally, matching the next frame scene image in the second scene video with the next frame estimated scene image to determine a verification result includes:

[0032] Obtaining second pose information corresponding to a next frame of scene image in the second scene video;

[0033] The second pose information is matched with the second estimated pose information to determine the verification result.

[0034] According to another aspect of an embodiment of the present disclosure, a scene reconstruction device is provided, comprising:

[0035] An image acquisition module, configured to acquire at least one frame of a first scene image of a driving scene under an initial trajectory;

[0036] An initial modeling module, configured to perform modeling based on the at least one frame of the first scene image to obtain an initial scene video;

[0037] A trajectory reconstruction module, configured to reconstruct a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory;

[0038] The video verification module is used to verify the at least one second scene video using a motion model to obtain at least one target scene video that passes the verification.

[0039] Optionally, the initial modeling module is specifically used to perform dynamic separation processing on the at least one frame of the first scene image to obtain a first static image and at least one group of first dynamic image groups; each group of the first dynamic image groups includes at least one frame of dynamic image, and each group of the first dynamic image groups corresponds to a dynamic object; each group of the first dynamic image groups is modeled separately to obtain at least one dynamic trajectory; the at least one dynamic trajectory is converted to the background coordinate system corresponding to the first static image to obtain the initial scene video.

[0040] Optionally, the dynamic object includes the ego vehicle and at least one other vehicle. When the initial modeling module models each group of the first dynamic image groups separately to obtain at least one dynamic trajectory, it is used to classify the at least one group of first dynamic image groups based on the dynamic object, determine the first dynamic image group corresponding to the ego vehicle and the first dynamic image group corresponding to the related vehicle; the related vehicle is the other vehicle that is closest to the ego vehicle; based on predefined interaction rules and predefined behaviors corresponding to the related vehicles, determine the at least one dynamic trajectory corresponding to the ego vehicle.

[0041] Optionally, the image acquisition module is specifically configured to capture at least one frame of a third scene image of the driving scene under the initial trajectory; and obtain the at least one frame of the first scene image based on the at least one frame of the third scene image.

[0042] Optionally, when obtaining the at least one frame of the first scene image based on the at least one frame of the third scene image, the image acquisition module is used to perform interpolation processing on the initial trajectory corresponding to the at least one frame of the third scene image to obtain the at least one frame of the first scene image; and / or perform a lateral translation operation on at least one position information corresponding to the at least one frame of the third scene image in the initial trajectory to obtain the at least one frame of the first scene image.

[0043] Optionally, the trajectory reconstruction module is specifically configured to perform trajectory reconstruction on the initial scene video using a driving scene reconstruction model to obtain the at least one second scene video under a new trajectory.

[0044] Optionally, the video verification module is specifically used to estimate the next frame estimated scene image of the current frame scene image in the second scene video for each second scene video through a motion model; match the next frame scene image in the second scene video with the next frame estimated scene image to determine a verification result; and determine the at least one target scene video based on the verification result.

[0045] Optionally, when the video verification module estimates the next frame estimated scene image of the current frame scene image in the second scene video through a motion model, it is used to determine the first pose information corresponding to the vehicle based on the current frame scene image; perform pose estimation on the first pose information based on the motion model to obtain the second estimated pose information corresponding to the next frame estimated scene image.

[0046] Optionally, when the video verification module matches the next frame scene image in the second scene video with the next frame estimated scene image to determine the verification result, it is used to obtain second posture information corresponding to the next frame scene image in the second scene video; and match the second posture information with the second estimated posture information to determine the verification result.

[0047] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0048] a memory for storing a computer program product;

[0049] The processor is configured to execute the computer program product stored in the memory, and when the computer program product is executed, the scene reconstruction method described in any one of the above embodiments is implemented.

[0050] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, characterized in that when the computer program instructions are executed by a processor, the scene reconstruction method described in any of the above embodiments is implemented.

[0051] According to another aspect of the embodiments of the present disclosure, a computer program product is provided, including computer program instructions, characterized in that when the computer program instructions are executed by a processor, the scene reconstruction method described in any of the above embodiments is implemented.

[0052] Based on the scene reconstruction method, device, electronic device, storage medium, and program product provided in the above-mentioned embodiments of the present disclosure, at least one frame of the first scene image of the driving scene under the initial trajectory is obtained; modeling is performed based on the at least one frame of the first scene image to obtain an initial scene video; trajectory reconstruction is performed based on the initial scene video to obtain at least one second scene video under the new trajectory; the at least one second scene video is verified using a motion model to obtain at least one target scene video that passes the verification. This embodiment can obtain various new trajectories that are not collected in the actual scene through trajectory reconstruction, thereby achieving scene expansion, and the motion model verification is performed on the corresponding reconstructed second scene video, so that the motion of the second scene video corresponding to the new trajectory conforms to physical laws, is more in line with actual conditions, and is more suitable for training autonomous driving algorithms.

[0053] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0055] The present disclosure can be more clearly understood from the following detailed description with reference to the accompanying drawings, in which:

[0056] Figure 1 is a flowchart of a scene reconstruction method provided by an exemplary embodiment of the present disclosure;

[0057] Figure 2 This disclosure Figure 1 A schematic flow chart of step 104 in the embodiment shown;

[0058] Figure 3 is a flowchart of an embodiment of a method for training a scene reconstruction model disclosed herein;

[0059] Figure 4 The overall process of training a scene reconstruction model using a training tool for the scene reconstruction model is shown;

[0060] Figure 5 This disclosure Figure 1 A schematic flow chart of step 108 in the embodiment shown;

[0061] Figure 6 is a structural diagram of a scene reconstruction device provided by an exemplary embodiment of the present disclosure;

[0062] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated. DETAILED DESCRIPTION

[0063] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0064] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0065] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.

[0066] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.

[0067] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0068] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship. The data referred to in this disclosure can include unstructured data such as text, images, and videos, as well as structured data.

[0069] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.

[0070] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0071] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0072] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0073] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0074] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.

[0075] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.

[0076] Application Overview

[0077] In the process of implementing the present disclosure, the inventors discovered that existing scene reconstruction methods have at least the following problems: Although existing scene reconstruction technologies can establish realistic digital twins, these methods are limited by their training data sets and can usually only generate high-quality sensor outputs within the area covered by the recorded camera trajectory.

[0078] Exemplary Methods

[0079] Figure 1 This is a flow chart of a scene reconstruction method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, such as Figure 1 As shown, the following steps are included:

[0080] Step 102: Acquire at least one frame of a first scene image of a driving scene under an initial trajectory.

[0081] The first scene image under the initial trajectory can be obtained based on data collected while the vehicle travels a certain distance. Based on real-world driving data, the image collected by the vehicle is also referred to as the groundtruth (GT) video under the original trajectory. Alternatively, the first scene image can be a directly captured image or an image obtained by subjecting the captured image to at least one type of processing (e.g., data augmentation).

[0082] Step 104 : Modeling is performed based on at least one frame of the first scene image to obtain an initial scene video.

[0083] In one embodiment, image-to-scene modeling can be implemented using any modeling method in the prior art.

[0084] Step 106 : Reconstruct the trajectory based on the initial scene video to obtain at least one second scene video under the new trajectory.

[0085] Optionally, the trajectory of the initial scene video is reconstructed using existing trajectory reconstruction methods to obtain a second scene video under the new trajectory; for example, a low-quality video is first generated from the unknown new trajectory, and then a driving scene reconstruction model (Drive Restorer) is applied to process these outputs to obtain a high-quality restored video; wherein the driving scene reconstruction model can be trained to generate a high-quality second scene video under the new trajectory.

[0086] Step 108: Verify at least one second scene video using a motion model to obtain at least one target scene video that passes the verification.

[0087] In the fields of autonomous driving and computer vision, a motion model is a mathematical model that describes the motion patterns of objects over time. It plays a crucial role in target tracking, path prediction, and decision planning. Object motion is an objective phenomenon, while a motion model is an abstract description of this phenomenon. That is, the motion of real objects must conform to the motion model. In this embodiment, to verify whether the trajectory-reconstructed second-scene video conforms to the motion patterns of real objects, the second-scene video is verified using the motion model. This ensures that the verified target scene video conforms to the object's motion patterns, making the data more reliable. Using the target scene video to train autonomous driving algorithms can achieve better training results.

[0088] The scene reconstruction method provided by the above-mentioned embodiment of the present disclosure obtains at least one frame of the first scene image of the driving scene under the initial trajectory; performs modeling based on the at least one frame of the first scene image to obtain an initial scene video; performs trajectory reconstruction based on the initial scene video to obtain at least one second scene video under the new trajectory; and verifies the at least one second scene video through a motion model to obtain at least one target scene video that passes the verification. This embodiment can obtain various new trajectories that are not collected in the actual scene through trajectory reconstruction, thereby achieving scene expansion. In addition, the motion model is verified for the reconstructed second scene video, so that the motion of the second scene video corresponding to the new trajectory conforms to the laws of physics, is more consistent with the actual situation, and is more suitable for training the autonomous driving algorithm.

[0089] like Figure 2 As shown in the above Figure 1 Based on the embodiment shown, step 104 may include the following steps:

[0090] Step 1041 : Perform dynamic separation processing on at least one frame of the first scene image to obtain a first static image and at least one group of first dynamic image groups.

[0091] Each first dynamic image group includes at least one dynamic image frame, and each first dynamic image group corresponds to a dynamic object. The dynamic object can be the vehicle capturing the image, other vehicles, people, animals, and other moving objects traveling around the vehicle.

[0092] Step 1042 : Model each first dynamic image group separately to obtain at least one dynamic trajectory.

[0093] Optionally, at least one piece of position information corresponding to the dynamic object is determined based on the first dynamic image group; and based on the at least one piece of position information, a dynamic trajectory corresponding to the dynamic object is obtained.

[0094] Step 1043 : Convert the at least one dynamic trajectory into a background coordinate system corresponding to the first static image to obtain an initial scene video.

[0095] This embodiment models the background and dynamic objects separately; models the static background and dynamic objects separately (for example, identifying dynamic objects in the image, distinguishing dynamic objects from the static background, modeling based on the static background, extracting dynamic areas based on motion residuals or pixel differences, and modeling the dynamic areas in consecutive frames to obtain dynamic object trajectories), and clarifies the spatial relationship of objects in the scene. The dynamic object real-time conversion technology is then used to convert the local model of the dynamic object into the background coordinate system in real time, thereby realizing dynamic updating and rendering of dynamic objects (such as vehicles) in the scene. The image rendering quality under the new perspective is improved through perspective synthesis and scene editing, and the visual artifacts that appear during the rendering process are corrected in combination with the driving scene reconstruction model, especially during dynamic operations such as vehicle lane changing.

[0096] Due to insufficient data, existing scene reconstruction methods are unable to reconstruct some long-tail scenarios, such as the sudden braking of the preceding vehicle, and thus cannot provide rich scenarios for reinforcement learning algorithms. This embodiment addresses this issue by proposing the following solution. In some optional embodiments, the dynamic objects include the ego vehicle and at least one other vehicle; step 1042 may include:

[0097] Classification processing is performed on at least one first dynamic image group based on the dynamic object to determine a first dynamic image group corresponding to the own vehicle and a first dynamic image group corresponding to related vehicles.

[0098] Among them, the relevant vehicles are other vehicles that are closest to the vehicle.

[0099] Optionally, classification processing can be performed on the dynamic images in the first dynamic image group based on a deep learning method, and the dynamic images can be matched according to dynamic objects to obtain the first dynamic image group corresponding to the own vehicle and the first dynamic image group corresponding to the related vehicles.

[0100] At least one dynamic trajectory corresponding to the ego vehicle is determined based on predefined interaction rules and predefined behaviors corresponding to the related vehicles.

[0101] In this embodiment, direct interaction rules between the ego vehicle and related vehicles are predefined. When a related vehicle exhibits predefined behaviors, the corresponding trajectory information of the ego vehicle is processed accordingly by searching the predefined interaction rules, resulting in at least one modified dynamic trajectory. This embodiment proposes a dynamic adversarial agent to control the trajectories of surrounding vehicles to generate some long-tail data (such as sudden stops). The dynamic adversarial agent process includes two parts: target vehicle identification and interaction trajectory generation. The dynamic adversarial agent analyzes the motion trajectory to determine the other vehicles closest to the ego vehicle. Then, based on the predefined behaviors of related vehicles specified in the predefined interaction rules (such as braking or overtaking), the trajectory information of related vehicles and the trajectory information of the ego vehicle are combined to generate a new motion trajectory for the ego vehicle. The generation of the new motion trajectory can be based on rules (for example, setting the brakes to stop the vehicle) or on a large model (for example, outputting corresponding trajectory information based on input text). The dynamic adversarial agent (DAA) proposed in this embodiment generates long-tail traffic scenarios by autonomously adjusting the trajectories of surrounding vehicles (for example, controlling the related vehicle to suddenly brake or cut in), thereby addressing the lack of coverage of long-tail scenarios in existing datasets and enabling reinforcement learning to learn in more complex situations.

[0102] In some optional embodiments, based on the above embodiment, step 102 may include:

[0103] At least one frame of a third scene image of the driving scene under the initial trajectory is collected; and at least one frame of the first scene image is obtained based on the at least one frame of the third scene image.

[0104] In this embodiment, the third scene image can be directly used as the first scene image, or, to enrich the diversity of the first scene image, at least one expansion process can be performed on the third scene image. In this embodiment, the third scene image is data collected through real-world driving. This data collection process is very complex, wastes a lot of manpower and resources, and is limited by real-world driving scenarios, resulting in a relatively simple first scene image. To address this issue, an initial trajectory corresponding to at least one frame of the third scene image can optionally be interpolated to obtain at least one frame of the first scene image. Interpolation of the initial trajectory determined from the position information corresponding to the at least one frame of the third scene image initially collected can obtain more position information, which is suitable for enriching scarce driving behavior scenarios.

[0105] And / or, performing a lateral translation operation on at least one position information corresponding to at least one frame of the third scene image in the initial trajectory to obtain at least one frame of the first scene image.

[0106] By performing a horizontal translation on at least one position information in the initial trajectory, a new trajectory can be generated based on the initial trajectory, thereby achieving the effect of expanding the data coverage.

[0107] Initial trajectory recovery can be achieved using existing technologies, for example, by comprehensively utilizing visual features, IMU data (provided by the vehicle's own IMU device), and GPS information (provided by the vehicle's own positioning device) through four core steps: feature matching, motion estimation, multi-sensor fusion, and trajectory optimization. The Cousin Trajectory Generator (CTG) proposed in this embodiment is a trajectory data enhancement method. To address the problem of an over-concentration and lack of diversity in existing training datasets for straight-line driving scenarios, CTG generates richer expert trajectory data through interpolation and lateral offsets, resulting in better imitation learning results.

[0108] The data collection for autonomous driving is generally data collected for a section of the vehicle's driving process, and the trajectory of the driving process is called the original trajectory. If the vehicle only drives in one lane during the driving process, the collected data may only be data for one lane, so that the autonomous driving data set usually only contains data from the original trajectory, lacking rich multi-perspective images, resulting in the perspective sparsity of autonomous driving data. If you want to render a video from a new perspective, for example, when rendering a new trajectory such as gradual lane change, lane translation, etc., the video generated under the new trajectory will have obvious ghosting phenomena, such as smearing, fragmentation, blurring, etc. The rendering effect of the video under the new trajectory generated by the existing method is poor. Therefore, this embodiment proposes to use a driving scene reconstruction model to reconstruct the trajectory of the initial scene video to obtain at least one second scene video under the new trajectory.

[0109] The scene reconstruction model is trained, and the training process may include: first, using an insufficiently trained scene reconstruction model and the collected images under the original trajectory to train a video restoration model, and then using the trained video restoration model to assist the reconstruction model to obtain the restored images under the new trajectory. The restored images under the new trajectory and the collected images under the original trajectory can be used to construct a reconstruction dataset, and the reconstruction dataset is used to train the scene reconstruction model. The above reconstruction dataset includes the restored images under the new trajectory, which solves the problem of perspective sparsity of autonomous driving data. Therefore, the use of the trained scene reconstruction model can improve the image rendering quality under the new trajectory, and then improve the performance of the scene reconstruction model, thereby solving the technical problem of poor rendering effect of the video generated under the new trajectory mentioned in the prior art. Therefore, based on the trained scene reconstruction model, a video under the new trajectory with better rendering effect can be generated.

[0110] In some optional embodiments, Figure 3FIG. 1 is a flow chart of an embodiment of a training method for a scene reconstruction model disclosed herein. Figure 3 As shown, the method may specifically include:

[0111] Step 310: Acquire a captured image of the original trajectory in the driving scene.

[0112] The captured images from the original trajectory can be data collected while the vehicle travels a certain distance, representing real-world driving data, also known as ground truth (GT) video from the original trajectory. The scene reconstruction model training tool, which executes the scene reconstruction model training method disclosed herein, first acquires GT video from the original trajectory and subsequently uses this video to train the video inpainting model and the scene reconstruction model.

[0113] Step 320 : training a video restoration model using the insufficiently trained scene reconstruction model and the captured images under the original trajectory.

[0114] First, the GT video of the original trajectory is used to train the scene reconstruction model. During the training process, the original trajectory is sampled, and the scene reconstruction model is used to render an image from the original perspective. Rendering can include the process of extracting an image from a scene given a perspective. When the scene reconstruction model is not fully trained, that is, when the model has not fully converged, the result data output by the model is usually an image with poor rendering effect and ghosting. The above image with poor rendering effect and the GT video of the original trajectory are used as a video pair, and the video pair is used to train the video restoration model. In the subsequent steps, the trained video restoration model is used to assist the scene reconstruction model, so that the scene reconstruction model can well render videos from some new perspectives to improve the performance of the scene reconstruction model.

[0115] Step 330 : Utilize the scene reconstruction model to render the image under the new trajectory to obtain a rendered image under the new trajectory.

[0116] In this step, a new trajectory can be sampled first. For example, based on the original trajectory, a perspective that shifts the lane 1.5 meters to the left can be used to sample a new trajectory. The scene reconstruction model is then used to render the image for the new trajectory. Because the scene reconstruction model has not been fully trained for the new trajectory, the rendering quality of the image generated by the model for the new trajectory is poor, requiring restoration using a video restoration model.

[0117] Step 340 : Use the trained video restoration model to restore the rendered image under the new trajectory to obtain a restored image under the new trajectory.

[0118] In step 320, the poorly rendered image from the original perspective, generated by the inadequately trained scene reconstruction model, is used as a video pair with the ground-truth video from the original trajectory to train a video inpainting model. After the video inpainting model is trained, the model parameters are locked and used to inpaint the rendered image from the new trajectory generated in step 330, resulting in a restored image from the new trajectory.

[0119] Step 350 : Training a scene reconstruction model using the restored image under the new trajectory and the captured image under the original trajectory.

[0120] In this step, the scene reconstruction model is trained using the restored image under the new trajectory obtained in step 340 and the captured image under the original trajectory obtained in step 310.

[0121] Figure 4 The overall process of training the scene reconstruction model using the training tool of the scene reconstruction model is shown. Figure 4 As shown in the figure, “Dynamic Scene Reconstruction” represents the scene reconstruction model. At the beginning of training, the dataset only contains GT videos of the original trajectory, see Figure 4 "OriginalTrajectory GT Video". First, use the GT video of the original trajectory to train the reconstruction model for a period of time. After the scene reconstruction model can obtain a good processing effect on the GT video of the original trajectory, a new trajectory can be sampled, for example, the original trajectory is offset by 1.5m to the left, and the scene reconstruction model is instructed to render the image under the new trajectory. Among them, a progressive trajectory sampling strategy can be adopted, and the lane offset and the offset angle of the new trajectory relative to the original trajectory are gradually increased during multiple sampling processes to obtain a progressive trajectory sample (ProgressiveTrajectory Sample), that is, a series of new trajectories with different offsets. Then use the scene reconstruction model to render (Render) the new trajectory to obtain the rendered image under the new trajectory, that is Figure 4 Since the scene reconstruction model is only trained on the original trajectory, the quality of the rendered images generated by the new trajectory at this stage may be poor. Figure 4"Online Restoration" refers to the process of online restoration by the video restoration model. A projection module is pre-set in the scene reconstruction model training tool. This module is used to re-project other vehicles and lane lines other than the vehicle into the rendered image according to the new trajectory. After projection, the "3D Box Sequence" and "HDMap Sequence" are obtained. The rendered image is then restored using the trained and locked parameter video restoration model DriveRestorer to obtain the restored image under the new trajectory. Specifically, Figure 3 In the code, "Enc" represents encoding and "Video Decoder" represents decoding. The "Noisy Images", control condition c, and the encoded "3D Box Sequence" and "HDMapSequence" are input into the video restoration model DriveRestorer. After being processed by the model, the processed results are decoded (Video Decoder) to obtain the restored image under the new trajectory, that is, Figure 3 Then add the restored image to the reconstructed dataset and update the reconstructed dataset, see Figure 4 Click "Dataset Update" in the , so that the updated dataset includes the restored images under the new trajectory. The above completes an iterative process of scene reconstruction model training.

[0122] In the next iteration, trajectories are resampled from the updated dataset, ensuring that images from the new trajectories account for a certain proportion. The above process is repeated, using the scene reconstruction model to render the images from the new trajectories. The trained video inpainting model is then used to inpaint the rendered images from the new trajectories. The inpainted images are then added to the reconstructed dataset, updating it. Because images from the new trajectories account for a certain proportion of the samples from the reconstructed dataset, these multiple iterations improve the quality of the reconstruction model's rendering of images from the new trajectories.

[0123] In summary, based on the disclosed embodiments, a video restoration model is first trained using an inadequately trained scene reconstruction model and the captured images from the original trajectory. The trained video restoration model is then used to assist the reconstruction model, obtaining restored images from the new trajectory. A reconstruction dataset is constructed using the restored images from the new trajectory and the captured images from the original trajectory, and the reconstruction dataset is used to train the scene reconstruction model. The above reconstruction dataset includes restored images from the new trajectory, and the images from the new trajectory are sampled so that they account for a certain proportion. Therefore, the trained scene reconstruction model can be used to improve the image rendering quality from the new trajectory, thereby improving the performance of the scene reconstruction model.

[0124] like Figure 5 As shown in the above Figure 1 Based on the embodiment shown, step 108 may include the following steps:

[0125] Step 1081 : for each second scene video, estimate a next frame estimated scene image based on a current frame scene image in the second scene video using a motion model.

[0126] The first pose information corresponding to the vehicle is determined based on the current frame scene image; the pose of the first pose information is estimated based on the motion model to obtain the second estimated pose information corresponding to the next frame estimated scene image.

[0127] In the digital twin environment corresponding to the real world, the motion trajectories of the ego vehicle and other vehicles in the scene need to conform to the motion model. Specifically, in the world coordinate system, the complete pose of the vehicle at time t is expressed as in, represents the rotation matrix describing the vehicle's orientation, It is the three-dimensional position information of the vehicle center in the world coordinate system.

[0128] At each time step, the vehicle's position changes according to the linear velocity v t and steering angle δ t The vehicle position is updated according to the defined kinematic bicycle model. Specifically, the vehicle position update formula is shown in formula (1):

[0129]

[0130] in, Represents the three-dimensional position information of the vehicle center in the world coordinate system at the next time step; is the rotation matrix The forward direction vector extracted from ; Δt represents the duration of a time step.

[0131] The vehicle orientation is updated by rotating around the vertical axis, as shown in formula (2):

[0132]

[0133] in, Rot represents the rotation matrix of the vehicle's orientation at the next time step; y (Δθ t ) represents the rotation matrix; Δθ t represents the incremental rotation angle, which can be calculated based on the sports bicycle module. For example, the incremental rotation angle is calculated based on the following formula (3):

[0134]

[0135] Where, L represents the vehicle wheelbase (calibrated value); v t represents the forward speed of the vehicle; δ t Represents the steering angle input. The corresponding rotation matrix Rot y (Δθ t The specific expression of ) is shown in the following formula (4):

[0136]

[0137] The linear speed, steering angle and other information involved in the above calculation process can be obtained based on vehicle sensors installed on the vehicle.

[0138] Step 1082: Match the next frame of scene image in the second scene video with the next frame of estimated scene image to determine a verification result.

[0139] Optionally, second pose information corresponding to the next frame of scene image in the second scene video is obtained; the second pose information is matched with the second estimated pose information to determine a verification result.

[0140] In this embodiment, in order to make the second scene video generated under the new trajectory conform to the motion law of the object, the pose information of the next time step (next frame) is inferred according to the motion model to obtain the second estimated pose information, and the second pose information corresponding to the next frame scene image in the second scene video (which can be determined based on the IMU device and positioning device carried by the vehicle) is matched with the second estimated pose information. If the difference between the second pose information and the second estimated pose information is less than the preset threshold, it means that the two match. At this time, it is determined that the second pose information conforms to the kinematic law, that is, the verification result is passed. If the difference between the second pose information and the second estimated pose information is greater than or equal to the preset threshold, it means that the two do not match, that is, there are trajectory points in the second scene video that do not conform to the kinematic law, and the verification result is failed. The trajectory point that fails the verification in the second scene video can be replaced with the second estimated trajectory point, or the motion trajectory corresponding to the trajectory point can be deleted.

[0141] Step 1083: Determine at least one target scene video according to the verification result.

[0142] This embodiment verifies the kinematic laws of the trajectory based on the motion model, avoiding the problem of unusable generated trajectories. The vehicle kinematic model updates the position information of the vehicle and other vehicles in real time, ensuring that the dynamic interaction between vehicles in the virtual environment is realistic and reliable, achieving closed-loop simulation.

[0143] Developed based on the ReconDreamer module (which constructs a realistic 4D world from a single-view input video), this module integrates world model knowledge to accurately represent dynamic scenes, ensuring high-quality sensor data even outside the recorded camera trajectory. It also utilizes a kinematic model to model all vehicles, enabling real-time closed-loop simulation and reinforcement learning training.

[0144] The target scene video obtained by the scene reconstruction method provided in this embodiment can be used to train an end-to-end autonomous driving algorithm. The autonomous driving algorithm trained using the target scene video provided in this embodiment is more realistic, so the trained autonomous driving algorithm is more consistent with the real scene. In addition, since the training data effectively covers complex and rare long-tail scenes, the trained autonomous driving algorithm can cope with various driving scenarios.

[0145] Any of the scene reconstruction methods provided in the embodiments of the present disclosure can be executed by any appropriate device with data processing capabilities, including but not limited to: a terminal device and a server. Alternatively, any of the scene reconstruction methods provided in the embodiments of the present disclosure can be executed by a processor, such as a processor that executes any of the scene reconstruction methods mentioned in the embodiments of the present disclosure by invoking corresponding instructions stored in a memory. This will not be further described below.

[0146] Exemplary devices

[0147] Figure 6 FIG. 1 is a schematic diagram of the structure of a scene reconstruction device provided by an exemplary embodiment of the present disclosure. Figure 6 As shown, the device provided in this embodiment includes:

[0148] The image acquisition module 61 is configured to acquire at least one frame of a first scene image of the driving scene under the initial trajectory.

[0149] The initial modeling module 62 is configured to perform modeling based on at least one frame of the first scene image to obtain an initial scene video.

[0150] The trajectory reconstruction module 63 is configured to reconstruct a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory.

[0151] The video verification module 64 is configured to verify at least one second scene video using a motion model to obtain at least one target scene video that passes the verification.

[0152] The scene reconstruction device provided by the above-mentioned embodiment of the present disclosure obtains at least one frame of the first scene image of the driving scene under the initial trajectory; performs modeling based on the at least one frame of the first scene image to obtain an initial scene video; performs trajectory reconstruction based on the initial scene video to obtain at least one second scene video under the new trajectory; and verifies the at least one second scene video through a motion model to obtain at least one target scene video that passes the verification. This embodiment can obtain various new trajectories that are not collected in the actual scene through trajectory reconstruction, thereby achieving scene expansion. In addition, the motion model is verified for the reconstructed second scene video, so that the motion of the second scene video corresponding to the new trajectory conforms to the laws of physics, is more consistent with the actual situation, and is more suitable for training the autonomous driving algorithm.

[0153] In some optional embodiments, the image acquisition module 61 is specifically configured to capture at least one frame of a third scene image of the driving scene under the initial trajectory; and obtain at least one frame of the first scene image based on the at least one frame of the third scene image.

[0154] Optionally, when obtaining at least one frame of the first scene image based on at least one frame of the third scene image, the image acquisition module 61 is used to perform interpolation processing on the initial trajectory corresponding to the at least one frame of the third scene image to obtain at least one frame of the first scene image; and / or, perform a lateral translation operation on at least one position information corresponding to the at least one frame of the third scene image in the initial trajectory to obtain at least one frame of the first scene image.

[0155] In some optional embodiments, the initial modeling module 62 is specifically used to perform dynamic separation processing on at least one frame of the first scene image to obtain a first static image and at least one group of first dynamic image groups; each group of first dynamic image groups includes at least one frame of dynamic image, and each group of first dynamic image groups corresponds to a dynamic object; each group of first dynamic image groups is modeled separately to obtain at least one dynamic trajectory; and the at least one dynamic trajectory is converted to the background coordinate system corresponding to the first static image to obtain the initial scene video.

[0156] Optionally, the dynamic objects include the ego vehicle and at least one other vehicle. When the initial modeling module 62 models each group of first dynamic image groups separately to obtain at least one dynamic trajectory, it is used to classify the at least one group of first dynamic image groups based on the dynamic objects, determine the first dynamic image group corresponding to the ego vehicle and the first dynamic image group corresponding to the related vehicles; the related vehicles are other vehicles that are closest to the ego vehicle; based on predefined interaction rules and predefined behaviors corresponding to the related vehicles, determine at least one dynamic trajectory corresponding to the ego vehicle.

[0157] In some optional embodiments, the trajectory reconstruction module 63 is specifically configured to reconstruct the trajectory of the initial scene video using the driving scene reconstruction model to obtain at least one second scene video under a new trajectory.

[0158] In some optional embodiments, the video verification module 64 is specifically used to estimate the next frame estimated scene image for the current frame scene image in the second scene video through a motion model for each second scene video; match the next frame scene image in the second scene video with the next frame estimated scene image to determine a verification result; and determine at least one target scene video based on the verification result.

[0159] Optionally, when the video verification module 64 estimates the next frame estimated scene image of the current frame scene image in the second scene video through the motion model, it is used to determine the first pose information corresponding to the vehicle based on the current frame scene image; perform pose estimation on the first pose information based on the motion model to obtain the second estimated pose information corresponding to the next frame estimated scene image.

[0160] Optionally, when the video verification module 64 matches the next frame scene image in the second scene video with the next frame estimated scene image to determine the verification result, it is used to obtain the second posture information corresponding to the next frame scene image in the second scene video; and match the second posture information with the second estimated posture information to determine the verification result.

[0161] Exemplary electronic devices

[0162] Below, reference Figure 7The electronic device according to the embodiment of the present disclosure is described. The electronic device may be either or both of the first device and the second device, or a standalone device independent of them, and the standalone device may communicate with the first device and the second device to receive collected input signals from them.

[0163] Figure 7 A block diagram of an electronic device according to an embodiment of the present disclosure is illustrated.

[0164] like Figure 7 As shown, the electronic device includes one or more processors and memory.

[0165] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.

[0166] The memory may store one or more computer program products, and the memory may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program products may be stored on the computer-readable storage medium, and the processor may execute the computer program products to implement the scene reconstruction method of the various embodiments of the present disclosure described above and / or other desired functions.

[0167] In one example, the electronic device may further include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0168] In addition, the input device may also include, for example, a keyboard, a mouse, and the like.

[0169] The output device can output various information to the outside, including determined distance information, direction information, etc. The output device can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.

[0170] Of course, to simplify, Figure 7 Only some of the components related to the present disclosure in the electronic device are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.

[0171] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to execute the steps of the scene reconstruction method according to various embodiments of the present disclosure described in the above part of this specification.

[0172] The computer program product may be written in any combination of one or more programming languages ​​to implement the operations of the disclosed embodiments, including object-oriented programming languages ​​such as Java, C++, and conventional procedural programming languages ​​such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0173] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the scene reconstruction method according to various embodiments of the present disclosure described in the above part of this specification.

[0174] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0175] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.

[0176] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.

[0177] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0178] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.

[0179] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0180] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0181] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A scene reconstruction method, characterized in that: include: Acquire at least one frame of a first scene image of the driving scene under the initial trajectory; Perform modeling based on the at least one frame of the first scene image to obtain an initial scene video; Reconstructing a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory; The at least one second scene video is verified using a motion model to obtain at least one target scene video that passes the verification.

2. The method according to claim 1, characterized in that The performing modeling based on the at least one frame of the first scene image to obtain an initial scene video includes: Performing dynamic separation processing on the at least one frame of the first scene image to obtain a first static image and at least one first dynamic image group; each first dynamic image group includes at least one frame of dynamic image, and each first dynamic image group corresponds to a dynamic object; Modeling each group of the first dynamic images to obtain at least one dynamic trajectory; The at least one dynamic trajectory is converted into a background coordinate system corresponding to the first static image to obtain the initial scene video.

3. The method according to claim 2, characterized in that The dynamic objects include the vehicle and at least one other vehicle, and modeling each group of the first dynamic image groups to obtain at least one dynamic trajectory includes: Classifying the at least one first dynamic image group based on the dynamic object, and determining the first dynamic image group corresponding to the own vehicle and the first dynamic image group corresponding to related vehicles; the related vehicles being the other vehicles closest to the own vehicle; The at least one dynamic trajectory corresponding to the ego vehicle is determined based on predefined interaction rules and predefined behaviors corresponding to the related vehicles.

4. The method according to any one of claims 1 to 3, characterized in that: The acquiring of at least one frame of a first scene image of the driving scene under the initial trajectory includes: Collecting at least one frame of a third scene image of the driving scene under the initial trajectory; The at least one frame of the first scene image is obtained based on the at least one frame of the third scene image.

5. The method according to claim 4, characterized in that The obtaining of the at least one frame of the first scene image based on the at least one frame of the third scene image includes: performing interpolation processing on the initial trajectory corresponding to the at least one frame of the third scene image to obtain the at least one frame of the first scene image; and / or, A lateral translation operation is performed on each of the at least one position information corresponding to the at least one frame of the third scene image in the initial trajectory to obtain the at least one frame of the first scene image.

6. The method according to any one of claims 1 to 5, characterized in that: The step of reconstructing a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory includes: The driving scene reconstruction model is used to reconstruct the trajectory of the initial scene video to obtain the at least one second scene video under a new trajectory.

7. The method according to any one of claims 1 to 6, characterized in that: Verifying the at least one second scene video using a motion model to obtain at least one target scene video that passes the verification includes: For each second scene video, estimating a next frame estimated scene image based on a current frame scene image in the second scene video by using a motion model; Matching a next frame of scene image in the second scene video with the next frame of estimated scene image to determine a verification result; The at least one target scene video is determined according to the verification result.

8. The method according to claim 7, characterized in that The estimating a next frame estimated scene image from a current frame scene image in the second scene video using a motion model includes: Determine the first position information corresponding to the vehicle based on the current frame scene image; Perform pose estimation on the first pose information based on the motion model to obtain second estimated pose information corresponding to the next frame estimated scene image.

9. The method according to claim 8, characterized in that The matching the next frame scene image in the second scene video with the next frame estimated scene image to determine a verification result includes: Obtaining second pose information corresponding to a next frame of scene image in the second scene video; The second pose information is matched with the second estimated pose information to determine the verification result.

10. A scene reconstruction device, characterized in that: include: An image acquisition module, configured to acquire at least one frame of a first scene image of a driving scene under an initial trajectory; An initial modeling module, configured to perform modeling based on the at least one frame of the first scene image to obtain an initial scene video; A trajectory reconstruction module, configured to reconstruct a trajectory based on the initial scene video to obtain at least one second scene video under a new trajectory; The video verification module is used to verify the at least one second scene video using a motion model to obtain at least one target scene video that passes the verification.

11. An electronic device, characterized in that: include: a memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, the scene reconstruction method described in any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the scene reconstruction method described in any one of claims 1 to 9 is implemented.

13. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the scene reconstruction method described in any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Track generation method and device, equipment and storage medium

    CN115601393A

  • Simulation scene construction method, computer equipment, storage medium and program product

    CN115690220A

  • Track planning method, device and equipment and computer readable storage medium

    CN118816927A

  • Driving scene data enhancement method based on track editing and image translation

    CN119068459A

  • A vehicle continuous positioning method, device, electronic equipment and storage medium

    CN119779318A