An environmental video generation method and related apparatus
By using the novel controllable video diffusion model StreetCrafter, environmental videos from any perspective are generated using pose modification and point cloud data, solving the problem of uncontrollable perspective in existing technologies and achieving high-quality dynamic street scene synthesis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CO WHEELS TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot generate environmental videos of vehicles from perspectives other than the actual driving viewpoint, resulting in uncontrollable viewpoints in the generated environmental videos. Furthermore, existing video diffusion models lack fine-grained controllability, which limits the application of autonomous driving simulators.
A novel controllable video diffusion model, StreetCrafter, is used to generate a second-view environmental video by acquiring a first environmental video and pose trajectory, modifying the pose, and combining point cloud data and reference images.
It enables the generation of environmental videos from any perspective, improves the controllability of the viewpoint in autonomous driving simulators and the accuracy of video generation, reduces artifacts, and supports high-quality synthesis of dynamic street scenes.
Smart Images

Figure CN122120409A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of autonomous driving technology, and in particular to an environmental video generation method and related apparatus. Background Technology
[0002] The vehicle is equipped with multiple sensors; during vehicle operation, these sensors collect information from the vehicle's perspective. Environmental video from this perspective can be synthesized based on the information gathered from these sensors. The process of obtaining this video is the process of dynamic street scene modeling. Dynamic street scene modeling is crucial for developing autonomous driving simulators. Summary of the Invention
[0003] In related technologies, environmental videos from the vehicle's driving perspective can be obtained, but environmental videos from other perspectives cannot be obtained. For example, if a vehicle is driving in the left lane, environmental videos from the left lane's driving perspective can be obtained, but environmental videos from the right lane's driving perspective cannot be obtained.
[0004] In view of the above problems, this application provides an environmental video generation method and related apparatus to achieve the purpose of obtaining environmental videos of vehicles from any viewpoint. The specific solution is as follows:
[0005] The first aspect of this application provides a method for generating environmental videos, including:
[0006] Acquire a first environmental video and a first pose trajectory; the first pose trajectory includes the first pose of the image acquisition device that acquired the first environmental video from a first viewpoint.
[0007] The first pose contained in the first pose trajectory is changed to the second pose of the image acquisition device that acquires the first environmental video from the second viewpoint, so as to obtain the second pose trajectory.
[0008] Based on the second pose trajectory and the reference image, a second environmental video from the second perspective is obtained; the reference image is the first frame video image in the first environmental video.
[0009] In one possible implementation, the first environmental video comprises multiple frames of video images, and the step of obtaining the second environmental video from the second viewpoint based on the second pose trajectory and the reference image includes:
[0010] Acquire multiple recording frames corresponding to different acquisition times, wherein the recording frames include first point cloud data and the video image marked with the second pose;
[0011] Multiple conditional images are obtained based on multiple recorded frames;
[0012] The second environmental video is obtained based on the reference image and multiple condition images.
[0013] In one possible implementation, the step of obtaining the second environmental video based on the reference image and the plurality of condition images includes:
[0014] The reference image and multiple conditional images are input into a pre-constructed novel controllable video diffusion model to obtain the second environmental video through the novel controllable video diffusion model;
[0015] The novel controllable video diffusion model is trained by taking the sample reference image and the sample condition image as input and the sample environment video from the sample perspective as the training target.
[0016] In one possible implementation, the step of obtaining multiple conditional images based on multiple recorded frames includes:
[0017] For each recorded frame, the first point cloud data in the recorded frame is projected onto the video image in the recorded frame to obtain the projection points corresponding to each point in the first point cloud data.
[0018] For each recorded frame, the pixel value of the projection point corresponding to each point in the first point cloud data in the recorded frame is determined as its own pixel value, so as to obtain colored second point cloud data;
[0019] Based on multiple second-point cloud data, the object trajectory is obtained;
[0020] The object point cloud and background point cloud are obtained from multiple second point cloud data using the object trajectory;
[0021] Multiple second point cloud data sets whose collection time falls within a preset time window are aggregated into a unified point cloud data set;
[0022] Transform multiple unified point cloud data sets into the world coordinate system;
[0023] For each of the second point cloud data, a point rasterization operation is performed in the world coordinate system based on the set radius corresponding to the second point cloud data and in the second pose corresponding to the second point cloud data to obtain the conditional image corresponding to the second point cloud data.
[0024] In one possible implementation, the novel controllable video diffusion model includes a first encoder, a second encoder, a decoder, a Video Denosing U-Net module, and a noise scheduling module; the step of inputting the reference image and multiple conditional images into the pre-constructed novel controllable video diffusion model to obtain the second environmental video through the novel controllable video diffusion model includes:
[0025] The latent spatial features of the reference image are obtained through the first encoder;
[0026] The second encoder is used to obtain the latent spatial features of the multiple conditional images;
[0027] The latent spatial features of the reference image with added preset noise and the latent spatial features of the conditional image are input to the Video Denosing U-Net module through the noise scheduling module;
[0028] The second environmental video, after noise removal and encoding, is obtained through the Video Denosing U-Net module;
[0029] The decoder is used to decode the encoded second environment video to obtain the decoded second environment video.
[0030] A second aspect of this application provides an environmental video generation apparatus, comprising:
[0031] The first acquisition module is used to acquire a first environmental video and a first pose trajectory; the first pose trajectory includes the first pose of the image acquisition device that acquires the first environmental video from a first viewpoint.
[0032] The second acquisition module is used to change the first pose contained in the first pose trajectory to the second pose of the image acquisition device that acquires the first environmental video from the second viewpoint, so as to obtain the second pose trajectory.
[0033] The third acquisition module is used to obtain a second environmental video from the second perspective based on the second pose trajectory and the reference image; the reference image is the first frame video image in the first environmental video.
[0034] In one possible implementation, the first environmental video comprises multiple frames of video images, and the third acquisition module includes:
[0035] The first acquisition unit is used to acquire multiple recording frames corresponding to different acquisition times. The recording frames include first point cloud data and the video image marked with the second pose.
[0036] The second acquisition unit is used to obtain multiple conditional images based on multiple recorded frames;
[0037] The third acquisition unit is used to obtain the second environmental video based on the reference image and multiple condition images.
[0038] A third aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the environmental video generation method described in the first aspect or any implementation thereof.
[0039] A fourth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0040] The memory is used to store computer programs;
[0041] The processor is used to execute the computer program so that the electronic device can implement the environmental video generation method of the first aspect or any implementation thereof.
[0042] The fifth aspect of this application provides a computer storage medium carrying one or more computer programs that, when executed by an electronic device, enable the electronic device to perform the environmental video generation method described in the first aspect or any implementation thereof.
[0043] By employing the above technical solution, this application provides an environmental video generation method, which acquires a first environmental video and a first pose trajectory. If a second environmental video from a second perspective is to be acquired, the first pose contained in the first pose trajectory needs to be changed to the second pose of the image acquisition device that acquired the first environmental video from the second perspective, so as to obtain the second pose trajectory. Based on the second pose trajectory and a reference image, the second environmental video from the second perspective is obtained. This achieves the goal of obtaining an environmental video from any perspective based on the first environmental video from the first perspective. Attached Figure Description
[0044] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0045] Figure 1 A schematic diagram of a system architecture is provided for this application;
[0046] Figure 2 A schematic diagram of an optional hardware structure for an in-vehicle terminal 100 provided in this application;
[0047] Figure 3 This application provides a schematic diagram of the structure of a server 200;
[0048] Figure 4 A flowchart illustrating an environmental video generation method provided in this application embodiment;
[0049] Figure 5 A schematic diagram of video images from a first perspective and a video image from a second perspective provided for embodiments of this application;
[0050] Figure 6 A schematic diagram illustrating the process of acquiring conditional images provided in an embodiment of this application;
[0051] Figure 7 A schematic diagram illustrating the training and usage process of the novel controllable video diffusion model provided in the embodiments of this application;
[0052] Figure 8 A schematic diagram illustrating the comparison of video images obtained from a new perspective using different algorithms, as provided in this application;
[0053] Figure 9 A schematic diagram illustrating the results obtained using the Waymo Open dataset, provided for an embodiment of this application;
[0054] Figure 10 This is a schematic diagram illustrating the results obtained using the PandaSet dataset in an embodiment of this application.
[0055] Figure 11 A schematic diagram of the visual ablation results of the algorithms of StreetCrafter and related technologies provided in the embodiments of this application;
[0056] Figure 12 A schematic diagram of the visual ablation results of the StreetCrafter and 3DGS combined algorithm and related technologies provided in the embodiments of this application;
[0057] Figure 13 A schematic diagram illustrating various editing operations for a movable object provided in this application;
[0058] Figure 14 This is a schematic diagram of the structure of an environmental video generation device provided in an embodiment of this application;
[0059] Figure 15 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation
[0060] The embodiments of this application are described below with reference to the accompanying drawings. The terminology used in the implementation section of this application is for explaining specific embodiments only and is not intended to limit the scope of this application.
[0061] The embodiments of this application will now be described with reference to the accompanying drawings. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.
[0062] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0063] Dynamic street scene modeling is a crucial step in developing autonomous driving simulators. It involves multiple aspects, including multi-sensor data fusion, real-time rendering, scene understanding, and dynamic element modeling. During dynamic street scene modeling, generating high-quality environmental videos is essential for closed-loop evaluation of autonomous driving systems. These videos can simulate real-world driving environments, allowing for the testing and validation of autonomous driving algorithms in a virtual environment, while also creating extreme situation data at low cost, such as rare traffic accidents or complex traffic congestion.
[0064] Understandably, vehicles are equipped with multiple sensors, such as image acquisition devices (e.g., cameras), point cloud acquisition devices (e.g., lidar, millimeter-wave radar), and IMUs (Inertial Measurement Units). The video images acquired by the image acquisition devices are RGB (Red, Green, Blue) images. The point cloud acquisition devices can acquire point cloud data; how to use RGB images and point cloud data from the same viewpoint to achieve real-time and high-quality environmental video synthesis is a key challenge in this field. How to use RGB images and point cloud data from the same viewpoint to generate coherent and realistic environmental videos is also a key challenge in this field.
[0065] It is understandable that dynamic street scene modeling includes both static and dynamic scene modeling. Dynamic scene modeling includes multiple dynamic objects, such as pedestrians, vehicles, and animals, which change over time. Static scene modeling includes multiple static objects, such as buildings, roads, sidewalks, trees, and traffic signs, which do not change over time.
[0066] Among related technologies, significant success has been achieved in novel perspective compositing for static scene modeling, providing valuable insights for dynamic street modeling. For example, 3DGS (3D Gaussian Splatting) technology achieves high-quality rendering results using sparse point cloud input by defining a set of anisotropic Gaussian kernels in the 3D world and performing adaptive density control. Furthermore, 3DGS-based techniques extend 3DGS to dynamic street scenes by modeling dynamic objects through scene graphs. This approach allows for the modeling of dynamic objects and can handle movement and changes within the scene.
[0067] While the relevant technologies can achieve high-quality, real-time video synthesis, significant artifacts appear in viewpoints far from the training trajectory. For example, a vehicle is traveling in the right lane, and the acquired video is from the perspective of the right lane. If a video from the perspective of the left lane is obtained based on the video from the perspective of the right lane, artifacts may appear in that video. Exemplary artifacts include, but are not limited to: aliasing artifacts, splatting expansion artifacts, high-frequency artifacts, blur and distortion, ringing artifacts, and overshoot artifacts.
[0068] In related technologies, video diffusion models can generate realistic environmental videos from new perspectives based on a small number of input images from other viewpoints, primarily thanks to training on large-scale video datasets. While video diffusion models have achieved success in environmental video generation, they typically rely on text prompts as control signals. As high-level instructions, text prompts lack fine-grained controllability, limiting their use in applications such as autonomous driving simulation. For example, text prompts may not be able to precisely control specific movements or subtle changes in detail within the video, which is a limitation for autonomous driving systems that require accurate simulation of specific driving scenarios.
[0069] Based on this, embodiments of this application provide an environmental video generation method, which involves a novel controllable video diffusion model, StreetCrafter. This novel controllable video diffusion model can precisely control the synthesis of videos from new perspectives in street scenes, providing important technical support for the development of autonomous driving simulators.
[0070] See Figure 1 , Figure 1 A schematic diagram of a system architecture is shown. The system may include an in-vehicle terminal 100, a server 200, and multiple sensors 300. The server 200 may include one or more servers (…). Figure 1 (The example includes a server), and the server 200 can provide the method provided in the embodiments of this application for one or more vehicle terminals.
[0071] The vehicle terminal 100 and multiple sensors 300 are deployed in the same vehicle.
[0072] For example, the multiple sensors 300 include, but are not limited to: one or more point cloud acquisition devices (e.g., lidar, millimeter-wave radar), one or more image acquisition devices, and IMU (Inertial Measurement Unit).
[0073] For example, the vehicle is also equipped with a Global Positioning System (GPS).
[0074] In some optional implementations, the vehicle terminal 100 can send information collected by multiple sensors to the server 200, and the server 200 can send the processing results back to the vehicle terminal 100.
[0075] In some alternative implementations, the vehicle terminal 100 can obtain processing results based on information collected by multiple sensors without the need for a server. This application does not limit the implementation of this method.
[0076] The following description Figure 1 The product form of the vehicle-mounted terminal 100;
[0077] The vehicle terminal 100 in this application embodiment can be a mobile phone, tablet computer, wearable device, vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc., and this application embodiment does not impose any restrictions on it.
[0078] Figure 2 A schematic diagram of an optional hardware structure for the vehicle-mounted terminal 100 is shown.
[0079] refer to Figure 2 As shown, the vehicle terminal 100 may include a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a headphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190, and other components. Those skilled in the art will understand that... Figure 2 This is merely an example of an in-vehicle terminal and does not constitute a limitation on in-vehicle terminals. It may include more or fewer components than shown in the illustration, or combine certain components, or use different components.
[0080] The input unit 130 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the vehicle terminal. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can collect touch operations performed by the user on or near it (such as operations performed by the user using fingers, knuckles, styluses, or any suitable object on or near the touch screen), and drive the corresponding connection devices according to a pre-set program. The touch screen can detect the user's touch actions, convert the touch actions into touch signals and send them to the processor 170, and can receive and execute commands sent by the processor 170; the touch signal includes at least touch point coordinate information. The touch screen 131 can provide an input interface and an output interface between the vehicle terminal 100 and the user. In addition, various types of touch screens, such as resistive, capacitive, infrared, and surface acoustic wave, can be used to implement the touch screen. In addition to the touch screen 131, the input unit 130 may also include other input devices. Specifically, other input devices 132 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0081] Among them, the input device 132 can receive input data, etc.
[0082] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the vehicle terminal 100, interactive interfaces, file display, and / or playback of any multimedia file. In this embodiment, the display unit 140 can be used to display interfaces, processing results, etc.
[0083] The memory 120 can be used to store instructions and data. The memory 120 may primarily include an instruction storage area and a data storage area. The data storage area can store various types of data, such as multimedia files and text. The instruction storage area can store software units such as operating systems, applications, and instructions required for at least one function, or subsets or extended sets thereof. It may also include non-volatile random access memory. It provides the processor 170 with hardware, software, and data resources for managing the computing device, supporting control software and applications. It is also used for storing multimedia files, as well as storing running programs and applications.
[0084] The processor 170 is the control center of the vehicle terminal 100. It connects various parts of the vehicle terminal 100 via various interfaces and lines. By running or executing instructions stored in the memory 120 and calling data stored in the memory 120, it performs various functions and processes data of the vehicle terminal 100, thereby providing overall control of the vehicle terminal device. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 170. In some embodiments, the processor and memory can be implemented on a single chip; in some embodiments, they can also be implemented separately on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to the corresponding components of the computing processing device, read and process data in the software, especially read and process data and programs in the memory 120, so that the various functional modules therein perform corresponding functions, thereby controlling the corresponding components to act according to the instructions.
[0085] The memory 120 can be used to store software code related to the environmental video generation method, and the processor 170 can execute the steps of the environmental video generation method, and can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to achieve the corresponding functions.
[0086] The radio frequency unit 110 (optional) can be used for receiving and transmitting signals during information transmission or calls. For example, it can receive downlink information from the base station and process it for the processor 170; additionally, it can transmit uplink data to the base station. Typically, the RF circuit includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the radio frequency unit 110 can also communicate wirelessly with network devices and other devices. This wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0087] In this embodiment of the application, the radio frequency unit 110 can send data to the server 200 and receive the processing results sent by the server 200.
[0088] It should be understood that the radio frequency unit 110 is optional and can be replaced with other communication interfaces, such as a network port.
[0089] The vehicle terminal 100 also includes a power supply 190 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 170 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0090] The vehicle terminal 100 also includes an external interface 180, which can be a standard Micro USB interface or a multi-pin connector, which can be used to connect the vehicle terminal 100 to other devices for communication, or to connect a charger to charge the vehicle terminal 100.
[0091] Although not shown, the vehicle terminal 100 may also include a flashlight, a wireless fidelity (WiFi) module, a Bluetooth module, and sensors with various functions, which will not be described in detail here. Some or all of the methods described below can be applied to, for example... Figure 2 The vehicle-mounted terminal 100 shown.
[0092] The following description Figure 1 The product form of the mid-range server 200;
[0093] Figure 3 A structural diagram of a server 200 is provided, as follows: Figure 3 As shown, server 200 includes bus 201, processor 202, communication interface 203, and memory 204. Processor 202, memory 204, and communication interface 203 communicate with each other via bus 201.
[0094] Bus 201 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 3 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0095] The processor 202 can be any one or more of the following processors: central processing unit (CPU), graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0096] Memory 204 may include volatile memory, such as random access memory (RAM). Memory 204 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0097] The memory 204 can be used to store software code related to the environmental video generation method, and the processor 202 can execute the steps of the chip's environmental video generation method, and can also schedule other units to achieve corresponding functions.
[0098] It should be understood that the aforementioned vehicle terminal 100 and server 200 can be centralized or distributed devices. The processors (e.g., processor 170 and processor 202) in the aforementioned vehicle terminal 100 and server 200 can be hardware circuits (such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), general-purpose processors, digital signal processors (DSPs), microprocessors or microcontrollers, etc.) or combinations of these hardware circuits. For example, the processor can be a hardware system with instruction execution capabilities, such as a CPU or DSP, or a hardware system without instruction execution capabilities, such as an ASIC or FPGA, or a combination of the aforementioned hardware systems without instruction execution capabilities and hardware systems with instruction execution capabilities.
[0099] In related technologies, environmental videos from the actual driving perspective of a vehicle can be obtained based on reference images and pose trajectories; however, environmental videos from other perspectives cannot be generated, meaning the perspective of the generated environmental videos is uncontrollable. Based on this, embodiments of this application provide an environmental video generation method. See details... Figure 4 .
[0100] Reference Figure 4 , Figure 4 This is a flowchart illustrating an environmental video generation method provided in an embodiment of this application, as shown below. Figure 4 As shown in the figure, an environmental video generation method provided in this application embodiment may include steps S401 to S403, which are described in detail below.
[0101] Step S401: Acquire the first environmental video and the first pose trajectory.
[0102] The first pose trajectory includes the first pose of the image acquisition device that acquires the first environmental video from the first viewpoint.
[0103] For example, the first environmental video includes multiple frames of video images.
[0104] The first pose trajectory includes the first pose of the image acquisition device that acquires the video image from the first viewpoint.
[0105] For example, if the first environment video includes 25 frames of video images, then the first pose trajectory includes 25 first poses.
[0106] For example, the first environmental video can be captured during the actual driving process of the vehicle. In this case, the first perspective is the driving perspective of the vehicle during its actual driving process. It can be understood that the vehicle is equipped with image acquisition devices, and the video images captured by the image acquisition devices are sorted from earliest to latest according to the acquisition time to obtain the first environmental video. It can also be understood that each video image captured by the image acquisition devices corresponds to a pose.
[0107] For example, pose can be fully represented by a 4x4 transformation matrix (also known as a homogeneous transformation matrix).
[0108] Camera pose refers to the position and orientation of an image acquisition device, such as a camera, in three-dimensional space. Specifically, it describes the rotation and translation relationship of the image acquisition device's coordinate system relative to the world coordinate system (or other reference coordinate system).
[0109] For example, the first environmental video can be obtained by the vehicle based on the environmental video corresponding to the driving perspective. In this case, the first perspective is different from the actual driving perspective.
[0110] The following examples illustrate the first-person and second-person perspectives.
[0111] like Figure 5 The image shown is a schematic diagram of a video image from a first perspective and a video image from a second perspective provided in an embodiment of this application.
[0112] Figure 5 The video image shown on the left is from the perspective of the vehicle traveling in the left lane. Figure 5 The video image shown on the right is the perspective of a vehicle traveling in the middle lane. This is an example. The first and second perspectives refer to the perspectives of a vehicle traveling in different lanes.
[0113] Step S402: Change the first pose contained in the first pose trajectory to the second pose of the image acquisition device that acquires the first environmental video from the second viewpoint, so as to obtain the second pose trajectory.
[0114] It is understandable that each frame of the first environmental video corresponds to a first pose; therefore, each pose needs to be modified to obtain the second pose.
[0115] For example, the first poses in the first pose trajectory are sorted from earliest to latest according to the acquisition time of the video images. The second poses in the second pose trajectory are also sorted from earliest to latest according to the acquisition time of the video images. The acquisition time corresponding to the second pose is the acquisition time of the first pose corresponding to that second pose.
[0116] Step S403: Based on the second pose trajectory and the reference image, obtain the second environmental video from the second perspective.
[0117] The reference image is the first frame of the video in the first environmental video.
[0118] For example, after obtaining the first frame video image from the second perspective by following the second pose trajectory which is ranked first in the second pose trajectory, the second frame video image from the second perspective can be obtained based on the first frame video image from the second perspective and the second pose which is ranked second in the second pose trajectory. The same principle applies to subsequent frames. This will not be elaborated further here.
[0119] This application provides an environmental video generation method, which acquires a first environmental video and a first pose trajectory. If a second environmental video from a second perspective is to be acquired, the first pose contained in the first pose trajectory needs to be changed to the second pose of the image acquisition device that acquired the first environmental video from the second perspective to obtain the second pose trajectory. Based on the second pose trajectory and a reference image, the second environmental video from the second perspective is obtained. This achieves the goal of obtaining an environmental video from any perspective based on the first environmental video from the first perspective.
[0120] It is understood that there are multiple ways to implement step S403. The embodiments of this application provide, but are not limited to, the following method, which includes the following steps A1 to A3.
[0121] Step A1: Acquire multiple recording frames corresponding to different acquisition times. The recording frames include first point cloud data and the video image marked with the second pose.
[0122] It is understandable that the first point cloud data and video image belonging to the same recording frame were acquired at the same time.
[0123] For example, the first point cloud data is collected by a point cloud acquisition device deployed on the vehicle.
[0124] Step A2: Obtain multiple conditional images based on the multiple recorded frames.
[0125] It is understood that there are multiple ways to implement step A2. The embodiments of this application provide, but are not limited to, the following method, which includes steps A21 to A27.
[0126] Step A21: For each recorded frame, project the first point cloud data in the recorded frame onto the video image in the recorded frame to obtain the projection points corresponding to each point in the first point cloud data.
[0127] For example, step A21 can be implemented by steps A211 to A212.
[0128] Step A211: Transform the first point cloud data from the lidar coordinate system to the image acquisition device coordinate system.
[0129] For example, this can be achieved using an extrinsic parameter matrix, i.e., a homogeneous transformation matrix.
[0130] Step A212: Project the first point cloud data in the coordinate system of the image acquisition device onto the two-dimensional video image plane.
[0131] For example, this can be achieved using a pinhole camera model.
[0132] For example, lens distortion in an image acquisition device can cause deviations in the projected points, so distortion correction is typically applied. Common distortions include radial and tangential distortion, which can be corrected using pre-calculated distortion coefficients.
[0133] Each projection point can be represented by a corresponding pixel coordinate on the video image plane. These pixel coordinates indicate the projection position of the projection point in the video image. The pixel value of the projection position of the projection point in the video image is the pixel value of the projection point.
[0134] Step A22: For each recorded frame, determine the pixel value of the projection point corresponding to each point in the first point cloud data of the recorded frame to obtain the second point cloud data in color.
[0135] For example, the points corresponding to the projected points can be colored by querying pixel values. For instance, if the pixel value of the projected point of point A in the first point cloud data is AA, then the pixel value of point A is AA. This allows the first point cloud data to be combined with the color information in the video image to obtain the second point cloud data.
[0136] Step A23: Obtain the object trajectory based on multiple second point cloud data.
[0137] Step A24: Obtain the object point cloud and background point cloud from multiple second point cloud data using the object trajectory.
[0138] For example, object point clouds (e.g., dynamic objects) can be separated from background point clouds using object trajectories. (Object point cloud) and background point clouds Defined in the canonical bounding box coordinate system of each dynamic instance o.
[0139] Where N is the total number of object point clouds and also the total number of background point clouds. This refers to the point cloud of the i-th object; It refers to the i-th background point cloud.
[0140] Step A25: Aggregate multiple second point cloud data sets whose acquisition time falls within a preset time window into a unified point cloud data set.
[0141] It is understandable that the collection time for each second point cloud data is the same as the collection time for the first point cloud data corresponding to the second point cloud data.
[0142] It is understandable that the coordinate system corresponding to the second point cloud data is the coordinate system of the image acquisition device. The pose of each first point cloud data is the pose of its corresponding second point cloud data.
[0143] Step A26: Transform the multiple unified point cloud data into the world coordinate system.
[0144] Assuming the second pose trajectory is used This indicates that K represents the total number of second poses, and K is also the total number of video images contained in the first environment video. Each second point cloud data t i Each corresponds to a second pose Ci, and the second point cloud data is aggregated within a preset time window l to form unified point cloud data P in the world coordinate system.
[0145] For example, step A25 enables the use of object trajectories in object point clouds. Transform to the world coordinate system.
[0146] Step A27: For each second point cloud data, perform point rasterization operation in the world coordinate system based on the set radius corresponding to the second point cloud data and in the second pose corresponding to the second point cloud data to obtain the conditional image corresponding to the second point cloud data.
[0147] For example, if each second point cloud data is treated as a pixel on the screen of an image acquisition device, there will be missing and occluded areas. Based on this, a set radius is assigned to each second point cloud data in the NDC (Normalized Device Coordinates) space, and a second pose C corresponding to that second point cloud data is defined. i Instead of directly projecting the uniform point cloud data P onto the video image plane, a point rasterization operation is performed, resulting in a conditional image that provides rich environmental information.
[0148] The conditional image not only contains the spatial location information of the second point cloud data, but also the color information. Furthermore, through point rasterization processing, the impact caused by occlusion and missing data can be reduced.
[0149] For example, conditional images can be used Characterization.
[0150] By using a point cloud acquisition device as a coarse scene geometry, a pixel-level connection is established between the second pose trajectory and the first environment video. Compared to conditional signals (such as pose embedding), conditional images can provide stronger guidance because the novel controllable video diffusion model only needs to recover a clean image from noisy input conditions, without having to learn the complex process of converting image acquisition device parameters into video frames.
[0151] To help those skilled in the art better understand the processes of steps A21 to A27 provided in the embodiments of this application, the following description is provided in conjunction with schematic diagrams, such as... Figure 6 The diagram shown is a schematic representation of the process of obtaining a conditional image provided in an embodiment of this application.
[0152] Figure 6 LiDAR and Object tracklets 61 represent object trajectories; Calibrated Cameras 62 represent multi-frame video images in the first environmental video; Colored LiDAR point cloud 63 represents unified point cloud data transformed to the world coordinate system.
[0153] The colorized LiDAR point cloud 63 contains list 631, which represents the novel trajectory, for example, the second environment video; list 632 contains the input trajectory, for example, the first environment video.
[0154] Figure 6 In the diagram, 64 refers to the conditional image corresponding to the first pose trajectory; 65 refers to the conditional image corresponding to the second pose trajectory.
[0155] The process of obtaining the conditional image corresponding to the first pose trajectory is the same as the process of obtaining the radar conditions corresponding to the second pose trajectory. Simply replace "second pose" with "first pose" in the method for obtaining the radar conditions corresponding to the second pose trajectory. This will not be elaborated further here.
[0156] Step A3: Based on the reference image and multiple condition images, obtain the second environmental video.
[0157] For example, a reference image can be used with I ref express.
[0158] Among related technologies, Video Diffusion Models (VDM) were the first to apply a diffusion model with a spatiotemporal decomposition U-Net to video generation tasks. Imagen Video proposed a cascaded video diffusion model to achieve higher resolution. Learning VDM in the latent space enables high-resolution video generation with relatively low computational cost. Learning specific elements and policies from street view video data enables the generation of realistic street view videos. To better support downstream applications such as reconstruction, some related technologies have proposed video generation methods controlled by the image acquisition device: such as training-free methods that achieve controllable video generation by controlling the denoising process of the video diffusion model. Although these methods do not require training or fine-tuning of the video diffusion model, their performance is limited by the ambiguity in the generation process. Some related technologies use additional inputs to fine-tune the video diffusion model. They either use an external matrix as a conditional input or convert the image acquisition device parameters into a ray coordinate graph to achieve image acquisition device control. Other related technologies use geometrically based models to construct explicit representations, such as point clouds, as guidance for the video diffusion model. However, these methods primarily focus on static scenes, while the method in this application utilizes more accurate point cloud acquisition device priors to provide guidance for dynamic street scenes.
[0159] NeRF and 3DGS have become leading methods for simulating autonomous driving scenarios in street scene representation. Block-NeRF uses a block-based modeling approach to represent large-scale static street scenes. Considering that street scenes typically include moving elements such as vehicles and pedestrians, some related techniques use time as an additional input to simulate dynamic street scenes. Other techniques decompose the scene into moving objects and static background, reconstruct them separately, and then combine the moving and static regions by tracking the moving objects at each time step. Some related techniques attempt to leverage point cloud acquisition device input to enhance the model's ability to capture scene geometry and generalize to new viewpoints. However, these primarily handle static scenes by adding supervision to the input trajectory, while the novel controlled video diffusion model in this application combines point cloud acquisition device input with generative priors to guide the generation of new trajectories for dynamic urban scenes.
[0160] Reconstruction with diffusion priors uses SDS to upscale 2D generation to 3D, achieving text-based 3D generation. Related techniques use video diffusion models to generate multi-view predictions based on single-view image input, then employ multi-view reconstruction methods to obtain the reconstructed 3D model. Some related techniques use diffusion priors to enhance sparse view reconstruction, while others utilize multi-view diffusion to improve multi-view... Figure 1 Consistency. For scenarios of a certain scale, video generation models can generate a large number of new views more efficiently than image generation models and have good multi-view consistency. Figure 1Similar to this application, some related techniques use DUSt3R to reconstruct point clouds from sparse viewpoints, leveraging geometric structures as conditions to guide video generation models in producing new viewpoint images, achieving static scene reconstruction with minimal viewpoint input. Related techniques enhance street scene reconstruction quality by integrating diffusion priors to generate additional views, similar to the method in this application. However, their video diffusion models lack precise control over the image acquisition device, limiting their reconstruction accuracy and visual quality.
[0161] Based on this, this application provides a novel controllable video diffusion model, which will be described below.
[0162] It is understood that there are multiple ways to implement step A3. The embodiments of this application provide, but are not limited to, the following method: inputting the reference image and multiple condition images into a pre-constructed novel controllable video diffusion model, and obtaining the second environmental video through the novel controllable video diffusion model.
[0163] The novel controllable video diffusion model is trained by taking a sample reference image and a sample condition image as input, and using a sample environment video from the sample's perspective as the training target.
[0164] It is understandable that the method for obtaining the sample condition image is the same as the method for obtaining the condition image, and will not be repeated here. The sample reference image is the first frame of the sample environment video. The sample environment video is the environmental video captured during the vehicle's movement (corresponding to the sample viewpoint).
[0165] In related technologies, reference image I ref and the second pose trajectory The input is fed into a video diffusion model, which yields an environmental video with a new perspective containing K frames of video images. However, the video diffusion model relies on the pose of the image acquisition device as a control signal, which is insufficient for street scenes with complex backgrounds and multiple dynamic objects. To address this issue, this application proposes a novel controllable video diffusion model, StreetCrafter, which utilizes point cloud acquisition device input to provide precise control over viewpoint changes during the diffusion denoising process.
[0166] For example, the value of K can be determined based on the actual situation, such as K=25.
[0167] The video diffusion model will be introduced below.
[0168] Video diffusion models have emerged as a cutting-edge approach for video generation in recent years. These models learn the underlying data distribution through forward and backward processes. During forward diffusion, Gaussian noise ϵ∼N(0, 1) is gradually added to the initial latent x0∼p(x) to obtain the noisy latent x. tNoise potential x t The calculation formula is as follows:
[0169] Where t∈{1, ...,T} refers to the diffusion time step. These are noise scheduling parameters. During backdiffusion, the video diffusion model learns to use a trained network. Iterative denoising of latent data. The video diffusion model is built on Vista, a driving world model finely tuned from the Stable Video Diffusion (SVD) model, following a continuous time-step formula. Given the input image (a video image labeled with the first pose) c 11 ,network Optimization is performed using the following loss function:
[0170] .
[0171] For example, the video diffusion model described above can diffuse a reference image I ref The conditional image corresponding to the first pose trajectory is used as input, and the first environment video corresponding to the first pose trajectory is used as the training result to obtain the result.
[0172] The following describes in detail the process of "inputting the reference image and multiple conditional images into a pre-constructed novel controllable video diffusion model, and obtaining the second environmental video through the novel controllable video diffusion model". This process includes steps B1 to B5.
[0173] Step B1: Obtain the latent spatial features of the reference image through the first encoder.
[0174] Step B2: Obtain the latent spatial features of the conditional image through the second encoder.
[0175] Step B3: The latent spatial features of the reference image with added preset noise and the latent spatial features of the conditional image are input to the Video Denosing U-Net module through the noise scheduling module.
[0176] Step B4: Obtain the second environmental video after noise removal and encoding through the Video Denosing U-Net module.
[0177] Step B5: Decode the encoded second environment video using the decoder to obtain the decoded second environment video.
[0178] To help those skilled in the art better understand the novel controllable video diffusion model provided in the embodiments of this application, the training and usage processes of the novel controllable video diffusion model are described below.
[0179] like Figure 7 The diagram shown illustrates the training and usage processes of the novel controllable video diffusion model provided in this application embodiment.
[0180] Figure 7 The left side is Figure 6 On the right side, Figure 7 Figure (b) is a schematic diagram of the process of training the new controllable video diffusion model, and Figure (c) is the process of using the new controllable video diffusion model after training.
[0181] like Figure 7 As shown, the novel controllable video diffusion model includes a first encoder 71, a second encoder 72, a decoder 73, a Video Denosing U-Net module 74, and a noise scheduling module 75.
[0182] For example, the first encoder and the second encoder can be a Pre-trained VAE (Pre-trained Variational Autoencoder).
[0183] For example, decoder 73 can be a pre-trained VAE decoder (Pre-trained Variational Auto decoder).
[0184] For example, the first encoder 71, the second encoder 72, and the decoder 73 do not require training, while the Video Denosing U-Net module 74 and the noise scheduling module 75 do require training.
[0185] Combination Figure 7 As shown in Figure (b), the conditional image 64 corresponding to the sample pose trajectory is input into the second encoder 72, and the latent spatial features (Latents1) of the conditional image are obtained through the second encoder 72. Among them, Z i C The latent spatial features (Latents1) of the i-th frame conditional image are represented; the sample environment video is input to the first encoder 71, and the latent spatial features (Latents2) of the sample environment video are obtained through the first encoder 71. Among them, Z i Let Latents2 represent the latent spatial features of the i-th video image.
[0186] For example, during actual driving, the vehicle collects sample environment videos from the sample perspective, and the sample environment videos correspond to sample pose trajectories. Figure 7 (b) The first pose trajectory As the sample pose trajectory, the conditional image of the first pose trajectory. The conditional image used as the sample pose trajectory, with the first environmental video As a sample environment video.
[0187] pass Figure 7 As shown in Figure (b), the noise scheduling module 75 adds noise to the latent spatial features (Latents2) of the sample environment video, and inputs the latent spatial features (Latents2) of the noisy sample environment video and the latent spatial features (Latents1) of the conditional image into the Video Denosing U-Net module 74; that is, by inputting... Network (Video Denosing U-Net module 74 includes) Add trainable zero-convolutional layers to the network. and will The signal is injected into the first layer of the Video Denosing U-Net module 74 architecture, and element-wise addition is performed. The noise-removed sample environment video is output through the Video Denosing U-Net module 74.
[0188] "By directing" Add trainable zero-convolutional layers to the network and will The result of injecting into the first layer of the VideoDenosing U-Net module's 74-layer architecture and performing element-wise addition. The calculation formula is as follows:
[0189] ;in, Indicates a zero convolutional layer. It comes from The potential noise x at diffusion time step t t .
[0190] The Video Denosing U-Net module 74 is optimized by minimizing the following loss objective. The formula for the loss objective is as follows:
[0191] , where C ref and C p These refer to the CLIP embedding of the reference image and the conditional image, respectively.
[0192] The following is combined with Figure 7 Figure (c) illustrates the process of using the novel controlled video diffusion model.
[0193] Figure 7 The novel controllable video diffusion model shown in Figure (c) has been trained.
[0194] pass Figure 7 As can be seen in Figure (b), the second pose trajectory Corresponding conditional image The input is fed into the second encoder to obtain the latent spatial features (Latents3) of the conditional image. The reference image I... ref The input is fed into the first encoder to obtain the latent spatial features (Latents4) of the reference image. Latents4 is not... Figure 7 As shown in the image. Among them, This represents the i-th second pose; This represents the conditional image corresponding to the i-th second pose.
[0195] For example, reference image I ref It can be the first frame of the first environment video. Or, it can be... The closest input image acquisition device is used as the reference image. The preset noise added by the noise scheduling module is the noise potential. From sampling noise Initially, the Video Denosing U-Net module is used iteratively, based on the conditional image. and reference image I ref Potential for noise The video is then denoised and converted into a clean second-view environment video. Decoder 73 is then used to decode it into a second-view environment video (i.e., Novel Views). .in, Let i be the i-th frame of the video in the second environment video.
[0196] Understandably, the Video Denosing U-Net module in the novel controllable video diffusion model has a relatively slow denoising speed. Assuming K=25, the novel controllable video diffusion model may take more than one minute to obtain the second environmental video from the second perspective. If more than 100 frames of second environmental video need to be generated, the continuity of these 100+ frames of video images may not be guaranteed. Based on this, this application also provides the following method to incorporate the generation prior of the controllable video diffusion model into a more consistent 3DGS (3D Gaussian Splatting) representation to achieve real-time rendering, that is, to distill the novel controllable video diffusion model StreetCrafter into a 3D representation to achieve real-time rendering. In this application, the novel controllable video diffusion model can be effectively distilled into a dynamic 3D scene representation, achieving state-of-the-art performance in street view synthesis.
[0197] The following is an explanation of 3DGS.
[0198] 3DGS uses a set of anisotropic Gaussians defined in the 3D world to represent a scene. Each Gaussian... Assigned with: Opacity spherical harmonic coefficient Position vector µ∈ Rotation Quaternions and scale factor The formula for the Gaussian kernel distribution is:
[0199] ;in, S is the scaling matrix determined by s, and R is the rotation matrix determined by q. Given the external W and internal K of an image acquisition device such as a camera, the 2D covariance matrix Σ in screen space... * The calculation is as follows: The color C of each pixel is rendered by alpha composition of the view-dependent color c in depth order, using the following formula: .
[0200] To simulate dynamic street scenes using 3DGS, relevant techniques were combined, and different Gaussian parameter sets were used to model background objects (e.g., background point clouds) and each foreground moving object (e.g., object point clouds). Object Gaussian Defined in the canonical coordinate system determined by the object trajectory. SE(3) pose given timestamp t. object Gauss It can be mapped to the world coordinate system for global rendering, as shown in the following formula:
[0201] ; where µ v R v, , They represent Gaussian respectively Position and rotation in local and world coordinate systems. The far region of the scene is modeled using a high-resolution cubemap, which is related to the equations... Rendering colors Used in combination, with rendering opacity .
[0202] To better align video images generated by a novel controllable video diffusion model with 3DGS scene representations, inspired by related technologies, potential generated samples can be rendered from noise instead of pure noise. Embodiments of this application find that this helps maintain the overall scene structure and accelerates the training process due to fewer denoising steps. Specifically, the second pose trajectory corresponds to a new perspective (i.e., a second perspective). Rendering images These images are encoded and perturbed into noise potential for a given diffusion time step t. t originates from the noise scale s. Samples are then generated from the Video Denosing U-Net module by running DDIM (Denoising Diffusion Implicit Models) sampling. The samples are further decoded into a second environmental video representing the second pose trajectory corresponding to the new viewpoint. To facilitate supervision, a stepwise optimization strategy is adopted based on the following generation process to gradually reduce the noise scale s. This helps the model rely more on diffusion priors to remove artifacts in the early training phases and gradually focus on refining details as training progresses.
[0203] The training set is constructed by combining an input view (e.g., the input to a novel controlled video diffusion model) with environmental videos from a new perspective generated by the model. In each training iteration, an image acquisition device, such as a camera (CC), is randomly sampled, with p representing a proportion of the new perspective image acquisition device set. Gaussian scene representation. Optimization is performed using the following loss function:
[0204] ;in , and These represent L1, SSIM, and LPIPS losses, respectively. The L1, SSIM, and LPIPS losses are selected based on whether the image acquisition device (e.g., camera CC) provides environmental video from a new perspective. input or L novel As a loss function, the external context introduces an additional LPIPS loss compared to the original 3DGS loss function, because it emphasizes high-level semantics.
[0205] Similarity, rather than photometric consistency, is used. This application also adds an additional loss for all input views. This includes depth loss from point cloud acquisition equipment. Sky mask loss and moving object regularization loss To further enhance the scene geometry.
[0206] This application provides an environmental video generation method involving a novel controllable video diffusion model, StreetCrafter. StreetCrafter can precisely control the synthesis of environmental videos from new perspectives in street scenes. Point cloud rendering from point cloud acquisition devices provides accurate geometric information, which, although incomplete and noisy, can serve as an accurate pose representation. To utilize this representation, this application uses point cloud rendering as a condition for StreetCrafter. Specifically, this application aggregates color radar point cloud data from adjacent frames to form a global point cloud (e.g., unified point cloud data) in the world coordinate system, and then renders it as an RGB image based on a given pose input, serving as a pixel-level pose condition in image space. Thanks to the pixel-level pose condition, even if the sample environmental video is a single-lane video, high-quality environmental videos spanning multiple lanes can be synthesized during testing by changing the conditions based on a new perspective (i.e., a second perspective) input. Furthermore, the proposed pixel-level pose condition can be used to implement scene editing operations without scene-by-scene optimization, simply by manipulating radar point cloud data, as shown in Figure 8.
[0207] However, the novel controlled video diffusion model StreetCrafter (hereinafter referred to as StreetCrafter) faces the challenge of high rendering latency, encountering a performance bottleneck when rendering large images (576×1024 pixels), with a frame rate of only 0.2fps. This means that the rendering speed is very slow and cannot meet the requirements of real-time rendering. To address this issue, this application distills the novel controlled video diffusion model StreetCrafter into a dynamic 3DGS representation. 3DGS is a technique capable of handling 3D scene representation and rendering, particularly suitable for handling scenes with large viewpoint changes. By distilling into a dynamic 3DGS representation, StreetCrafter can perform real-time, high-quality video image synthesis under large viewpoint changes, meaning that it can quickly render scenes from different perspectives while maintaining video image quality. To enhance the scene representation, StreetCrafter is used to generate a series of video images along a new trajectory. These video images supplement the input data of the original training trajectory as additional supervision information. The distillation process combines the advantages of 3D scene representation and the novel controlled video diffusion model StreetCrafter. 3D representation provides depth and structure information about a scene, while the novel, controllable video diffusion model, StreetCrafter, captures temporal changes and dynamic effects. Through this distillation and combination, StreetCrafter achieves state-of-the-art rendering performance, significantly improving rendering speed while maintaining video image quality, thus meeting the demands of real-time rendering.
[0208] In summary, by distilling StreetCrafter into a dynamic 3DGS representation and utilizing additional video image sequences as supervision, the rendering efficiency and quality of the system can be significantly improved, enabling real-time high-quality view composition under large viewpoint changes. This is particularly important for applications requiring real-time rendering, such as virtual reality, augmented reality, and games.
[0209] Figure 8 The diagram shows a comparison of the second environmental video with a new perspective obtained by StreetCrafter, the result obtained by the algorithm combining StreetCrafter and 3DGS (called Ours-G), and the second environmental video with a new perspective obtained by the algorithm of related technologies (called Street Gaussians).
[0210] For example, algorithms for related technologies, StreetCrafter, and algorithms combining StreetCrafter with 3DGS can be tested using the Waymo Open dataset and the PandaSet dataset to obtain... Figure 8 The results are shown.
[0211] Figure 8 Figure (a) shows the second environment video (Novel trajectory) with a new perspective (such as a second perspective) output by StreetCrafter after the recorded frames are input to StreetCrafter.
[0212] Figure 8 The left side of Figure (b) shows a second environmental video from a new perspective obtained by related technologies; Figure 8 The Ours-G representation on the right side of Figure (b) represents a new perspective in the second environmental video; from Figure 8 As can be seen from Figure (b) in this application, the second environmental video obtained by the method provided in this embodiment is of higher quality.
[0213] pass Figure 8 In Figure (c), some target objects in the second environment video generated by the embodiments of this application can be removed, and some target objects can be added. For example, relative to Figure 8 The diagram on the right side of Figure (b) Figure 8 In the left-hand image of diagram (c), the vehicle enclosed in a box is removed; this is target object removal. Figure 8 The diagram on the right side of Figure (b) Figure 8 In the middle (c) diagram, the two vehicles on the right have been replaced; this is a target object replacement.
[0214] Figure 8 The experimental results shown demonstrate that the StreetCrafter algorithm and the StreetCrafter-3DGS combined algorithm provided in this application outperform related algorithms in terms of image quality, particularly in view extrapolation. The StreetCrafter algorithm and the StreetCrafter-3DGS combined algorithm also support various scene editing operations without requiring scene-by-scene optimization, such as target object removal and replacement. In summary, the novel controllable video diffusion model StreetCrafter provided in this application offers precise image acquisition device control for the synthesis of new perspectives in street scenes. The StreetCrafter in this application can effectively distill dynamic 3D scene representations, achieving state-of-the-art performance in street view synthesis.
[0215] The implementation details of the novel controllable video diffusion model are explained below.
[0216] This application initializes StreetCrafter from Vista's pre-trained checkpoints. First, all parameters of the Video Denosing U-Net module are trained at a resolution of 320×576, with a batch size of 16 and a learning rate of 5e. −5The iterations were performed 30,000 times. Then, the time layer was fixed, and the spatial layers of the Video Denosing U-Net module were fine-tuned at a resolution of 576×1024, with a batch size of 8 and a learning rate of 1e. −5 The process is iterated 3000 times. During training, the reference and conditional images are randomly discarded with a 15% probability. Training StreetCrafter on eight NVIDIA A800 GPUs (Graphics Processing Units) using the Adam optimizer takes two days. During inference, the sampling step is set to 50, the classifier free-guided (CFG) ratio is set to 2.5, and videos of length n = 25 are generated at a resolution of 576×1024. For new trajectories longer than n (such as the second pose trajectory), the video images of length n are iteratively sampled with an overlap frame length of 5 to construct a full-length environment video. The preset time window size l is set to cover the second point cloud data within ±1 s of each other.
[0217] In the process of distilling StreetCrafter to 3DGS, this application follows the Street Gaussians setup and trains the model 30,000 times. New trajectories (such as the second pose trajectory) are constructed by laterally moving the input image acquisition device 3 meters in the direction of the vehicle's front end. This application samples the new viewpoint image acquisition device with a scale p set to 0.4. Coefficients λ1, λ2... ssim , λ lpips and λ novel The values were set to 0.2, 0.8, 0.5, and 0.1 respectively. Training on an A800 GPU took approximately 1.5 hours.
[0218] To help those skilled in the art better understand the performance of the StreetCrafter and StreetCrafter combined with 3DGS algorithms provided in the embodiments of this application, experiments are conducted below.
[0219] This application refers to the results of StreetCrafter as Ours-V; the results of the combined StreetCrafter and 3DGS algorithm are referred to as the StreetCrafter distillation framework, and the results of the StreetCrafter distillation framework are referred to as Ours-G.
[0220] In this application, the image input to StreetCrafter and the StreetCrafter distillation frame is referred to as the input image. During the evaluation experiment, the input image was cropped and resized to 576×1024 to match the output of StreetCrafter.
[0221] This application conducts experiments on the Waymo Open and PandaSet datasets, using their 10Hz front-view image acquisition devices and synchronous point cloud acquisition devices. Fifteen sequences (approximately 100 frames) from the Waymo Open validation set and five sequences (80 frames) from the PandaSet are selected to test the synthesis results of environmental videos from new perspectives (such as second-person view). This application uniformly samples half of the images from each sequence as test frames and uses the remaining images for training. The input image resolutions for the Waymo Open and PandaSet datasets are set to 1066×1600 and 900×1600, respectively. The Waymo Open training set and the remaining PandaSet sequences are used to train StreetCrafter, resulting in approximately 35,000 training samples in total.
[0222] To help those skilled in the art understand the advantages of the method provided in this application compared to algorithms in related technologies, this application also provides environmental videos with novel perspectives obtained by algorithms from various related technologies. Algorithms in related technologies include 3DGS, Street Gaussians, EmerNeRF, UniSim, and NeuRAD; among them, Street Gaussians uses separate Gaussian models to model the background and each moving object. EmerNeRF hierarchically divides the scene into static and dynamic fields, each modeled using a hash grid. UniSim and NeuRAD utilize neural feature grids to simulate dynamic driving scenes and use CNN renderers to enhance view extrapolation capabilities. This application enhances 3DGS by integrating depth supervision from point cloud acquisition devices and point cloud data initialization. For other methods, this application can evaluate the results based on their official implementations.
[0223] The comparison results of Ours-V, Ours-G, 3DGS, Street Gaussians, EmerNeRF, UniSim, and NeuRAD are explained below.
[0224] Table 1
[0225]
[0226] Table 2
[0227]
[0228] Table 1 shows the results obtained using the Waymo Open dataset. The rendered image resolution is 1066×1600.
[0229] Table 2 shows the results obtained using the Pandaset dataset. The rendered image resolution is 900×1600.
[0230] Tables 1 and 2 present the comparative results of the proposed method and related algorithms in terms of rendering quality and rendering speed. This application evaluates rendering quality under both view interpolation and extrapolation settings. It uses PSNR (Peak Signal-to-Noise Ratio) and LPIPS (Learned Perceptual Image Patch Similarity) as evaluation metrics for view interpolation and reports FID (Fréchet Inception Distance) under lane offset settings for view extrapolation, since a true image is unavailable. The proposed method achieves state-of-the-art performance in extrapolation scenarios while maintaining the comparability of interpolation results.
[0231] like Figure 9 The image shown is a schematic diagram illustrating the results obtained using the Waymo Open dataset as provided in an embodiment of this application.
[0232] Figure 9 The images shown are video images from the new perspectives obtained using the Waymo Open dataset, corresponding to the various algorithms. Each new perspective was achieved by moving the image acquisition device 3 meters laterally to the left or right.
[0233] like Figure 10 The diagram shown is a schematic representation of the results obtained using the PandaSet dataset according to an embodiment of this application.
[0234] Figure 9 The images shown are video images from the new perspectives obtained using the PandaSet dataset for each algorithm. The new perspective was achieved by moving the image acquisition device 3 meters laterally to the left or right.
[0235] Figure 9 and Figure 10 Qualitative differences were observed. The algorithms of related techniques tend to produce blurry results and artifacts in challenging and under-observed areas, such as lanes and moving vehicles, while Ours-G and Ours-V both generate high-fidelity new perspective images thanks to conditional images, diffusion priors, and distillation processes. Furthermore, Ours-G achieves similar rendering speeds to Street Gaussians, but as the FID metric indicates, Ours-G demonstrates superior generalization ability for large views.
[0236] Figure 9 and Figure 10 The input view includes the input image.
[0237] The following section will conduct an ablation study. An ablation study is a method of systematically removing, disabling, or replacing certain parts or functions of a model to assess the impact of these parts on overall performance. This study allows for a better understanding of which components in the model are important and which may not be necessary.
[0238] This application first analyzes the design choices of StreetCrafter. Several variants were analyzed using two types of settings on the validation set of the Waymo Open dataset, as shown in Table 3. For the random set, 40 video clips were randomly selected from 10 sequences obtained from the Waymo Open dataset. Further, 10 video clips from all scenes containing complex behaviors (such as turning and lane changing) were specifically selected to form the hard set. All videos were 25 frames long, with the first frame selected as the reference image. All variants were trained under the same settings of StreetCrafter. Finally, several optimization strategies were ablated during the distillation process (the process of distilling StreetCrafter to 3DGS).
[0239] Table 3
[0240]
[0241] The study on the design choice of StreetCrafter uses the average value of all sampled video images as the metric. Table 3 shows that the integrated model in this application achieves optimal performance. Ours can be Ours-V.
[0242] This application's StreetCrafter uses pose conditions. This application compares StreetCrafter with two pose-conditional variants, as shown in rows 1 and 2 of Table 3. This application first replaces the conditional image with the Plücker embedding of the image acquisition device's rays as the pose representation. Although ray embedding serves as a pixel-level pose condition, it fails to establish a relationship between the image acquisition device's pose and the scene geometry. (As...) Figure 11 As shown, the model that replaces the conditional image as a pose representation with the Plücker embedding of the image acquisition device ray lacks precise control and becomes blurry as the viewpoint deviates from the reference image. This application then treats the image acquisition device parameters as a vector and injects it into the temporal attention layer of the Denoising-VideoDenosing U-Net module. To better simulate object motion, this application replaces the conditional image with the projected object trajectory, similar to worldview models in related art. Figure 11As shown, the experimental results indicate that, due to the complexity of driving scenarios, the model that replaces the conditional image with the Plücker embedding of the image acquisition device ray as the pose representation lacks controllability.
[0243] like Figure 11 The diagram shown is a schematic representation of the visual ablation results of the algorithms of StreetCrafter and related technologies provided in the embodiments of this application.
[0244] This application compares StreetCrafter with two variants conditioned on point cloud acquisition devices, as shown in rows 3 and 4 of Table 3. This application first converts the aggregated point cloud data into single-frame LiDAR input. For example... Figure 11 As shown, the generated environmental video lacks fidelity in representing scene texture details, such as icons on the road, because single-frame LiDAR is very sparse even when using radius rendering. This application then proposes another variant that directly projects aggregated point cloud data without assigning a fixed radius to each point in the screen space of the video image. The results in Table 3 show that the StreetCrafter method mentioned in this application performs better because the projected point cloud struggles to handle occlusion relationships.
[0245] like Figure 12 The diagram shown illustrates the visual ablation results of the StreetCrafter and 3DGS combined algorithm and related technologies provided in this embodiment of the application. For example, the new perspective was obtained by moving the image acquisition device 3 meters to the left.
[0246] Table 4
[0247]
[0248] This application presents ablation studies on two sequences from the Waymo Open dataset, as shown in Table 4 and Figure 12 This application analyzes the StreetCrafter distillation framework. This application will use λ lpips Set to 0, and Change to L1 loss. Figure 12 The results show that LPIPS loss helps recover details from a new perspective. This application uses λ novel Set to 1.0 to remove the new weights. All metrics showed a significant drop, and numerous artifacts appeared on the moving object, such as... Figure 12 As shown, this highlights the importance of independently handling input and new views. This application sets the noise scale s to 1 throughout the training process so that the model always starts from pure noise. Artifacts appear in areas with insufficient LiDAR conditions, such as road signs on traffic lights.
[0249] The StreetCrafter distillation frame provided in this application embodiment maintains the 3D structure in areas where LiDAR conditions are absent, while refining the texture details of nearby areas.
[0250] The following explains scene editing.
[0251] StreetCrafter supports various editing operations on moving objects, i.e., dynamic objects. This application can achieve object translation (e.g., adjusting the object's bounding box properties during multi-frame point cloud aggregation). Figure 13 (a) moving the vehicle from the right lane to the middle lane), object replacement (such as...) Figure 13 (b) The vehicle's color and model were changed) and objects were removed (such as...). Figure 13 (c) Removes vehicles from the middle lane, providing StreetCrafter with modified LiDAR conditions. Unlike previous reconstruction methods that modeled each object separately, StreetCrafter can perform editing operations without scene-by-scene optimization.
[0252] Figure 13 The image on the right is the original image before editing. Figure 13 The image on the left is the edited version.
[0253] Figure 13 The dataset used is the Waymo Open dataset.
[0254] The StreetCrafter model provided in this application is a controllable video diffusion model for environmental video synthesis. It utilizes a sparse but geometrically accurate point cloud acquisition device to provide pixel-level pose conditions, enabling precise control of the image acquisition device and allowing StreetCrafter to generate consistent video frames that match the input of the image acquisition device. By further distilling StreetCrafter into a 3DGS model, this application enables environmental video synthesis in challenging scenarios, such as lane departures. Furthermore, scene editing is achieved by providing StreetCrafter with modified LiDAR conditions. Detailed ablation and comparisons on several datasets demonstrate the effectiveness of the proposed method. This work also has some known limitations. First, the cost of collecting and processing large amounts of data is high due to the large amount of point cloud data and object trajectories required to train StreetCrafter. Second, the inference speed of StreetCrafter is still far from real-time, attributed to the architecture of the Video Denosing U-Net module. Future work could consider using more advanced models for real-time inference.
[0255] This application aims to address the problem of synthesizing realistic views from vehicle sensor data. Recent advances in neural scene representation have achieved significant success in rendering high-quality autonomous driving scenes, but performance degrades significantly when the viewpoint deviates from the training trajectory. To mitigate this issue, this application introduces StreetCrafter, a novel controllable video diffusion model that utilizes point cloud rendering from a point cloud acquisition device as pixel-level pose conditions. It leverages generative priors for new viewpoint synthesis while maintaining precise control of the image acquisition device. Furthermore, the use of pixel-level pose conditions enables precise pixel-level editing of the target scene. StreetCrafter's generative priors can be effectively integrated into dynamic scene representations, achieving real-time rendering. Experiments on the Waymo Open dataset and the Pandaset dataset demonstrate that the model in this application can flexibly control viewpoint changes and expand the view synthesis area to meet rendering requirements, outperforming algorithms in related technologies.
[0256] The above describes an environmental video generation method provided by the embodiments of this application. The following will describe the apparatus for performing the above environmental video generation method.
[0257] Please see Figure 14 , Figure 14 This is a schematic diagram of the structure of an environmental video generation device provided in an embodiment of this application. Figure 14 As shown, the environmental video generation device includes:
[0258] The first acquisition module 1401 is used to acquire a first environmental video and a first pose trajectory; the first pose trajectory includes the first pose of the image acquisition device that acquires the first environmental video from a first viewpoint.
[0259] The second acquisition module 1402 is used to change the first pose contained in the first pose trajectory to the second pose of the image acquisition device that acquires the first environmental video from the second viewpoint, so as to obtain the second pose trajectory.
[0260] The third acquisition module 1403 is used to obtain a second environmental video from the second perspective based on the second pose trajectory and the reference image; the reference image is the first frame video image in the first environmental video.
[0261] In one optional implementation, the third acquisition module includes:
[0262] The first acquisition unit is used to acquire multiple recording frames corresponding to different acquisition times. The recording frames include first point cloud data and the video image marked with the second pose.
[0263] The second acquisition unit is used to obtain multiple conditional images based on multiple recorded frames;
[0264] The third acquisition unit is used to obtain the second environmental video based on the reference image and multiple condition images.
[0265] In one optional implementation, the third acquisition unit includes:
[0266] The first acquisition subunit is used to input the reference image and multiple condition images into a pre-constructed novel controllable video diffusion model, and obtain the second environmental video through the novel controllable video diffusion model;
[0267] The novel controllable video diffusion model is trained by taking the sample reference image and the sample condition image as input and the sample environment video from the sample perspective as the training target.
[0268] In one optional implementation, the second acquisition unit includes:
[0269] The second acquisition subunit is used to project the first point cloud data in the recording frame onto the video image in the recording frame for each recording frame, so as to obtain the projection points corresponding to each point in the first point cloud data.
[0270] A determining subunit is used to determine, for each recorded frame, the pixel value of the projection point corresponding to each point in the first point cloud data in the recorded frame is its own pixel value, so as to obtain colored second point cloud data;
[0271] The third acquisition subunit is used to acquire the object trajectory based on multiple second point cloud data;
[0272] The fourth acquisition subunit is used to obtain object point cloud and background point cloud from multiple second point cloud data through the object trajectory;
[0273] The aggregation subunit is used to aggregate multiple second point cloud data that were collected within a preset time window into a unified point cloud data;
[0274] The coordinate transformation subunit is used to transform multiple unified point cloud data into the world coordinate system;
[0275] The fifth acquisition subunit is used to perform point rasterization operation in the world coordinate system based on the set radius corresponding to the second point cloud data and in the second pose corresponding to the second point cloud data for each second point cloud data, so as to obtain the conditional image corresponding to the second point cloud data.
[0276] In one optional implementation, the novel controllable video diffusion model includes a first encoder, a second encoder, a decoder, a Video Denosing U-Net module, and a noise scheduling module; the first acquisition subunit includes:
[0277] The first acquisition submodule is used to acquire the latent spatial features of the reference image through the first encoder;
[0278] The second acquisition submodule is used to acquire the latent spatial features of the multiple conditional images through the second encoder;
[0279] The input submodule is used to input the latent spatial features of the reference image with added preset noise and the latent spatial features of the conditional image into the Video Denosing U-Net module through the noise scheduling module;
[0280] The third acquisition submodule is used to acquire the second environmental video after noise removal and encoding through the Video Denosing U-Net module;
[0281] The fourth acquisition submodule is used to decode the encoded second environment video through the decoder to obtain the decoded second environment video.
[0282] This application also provides an electronic device in its embodiments. (See reference...) Figure 15 The diagram illustrates a structural schematic suitable for implementing the electronic device in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 15 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0283] like Figure 15 As shown, the electronic device may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 1501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1502 or a program loaded from a storage device 1508 into a random access memory (RAM) 1503. When the electronic device is powered on, the RAM 1503 also stores various programs and data required for the operation of the electronic device. The processing unit 1501, ROM 1502, and RAM 1503 are interconnected via a bus 1504. An input / output (I / O) interface 1505 is also connected to the bus 1504.
[0284] Typically, the following devices can be connected to I / O interface 1505: input devices 1506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 1507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1508 including, for example, memory cards, hard drives, etc.; and communication devices 1509. Communication device 1509 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 15 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0285] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the environmental video generation methods provided in this application.
[0286] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the environmental video generation methods provided in this application.
[0287] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0288] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0289] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0290] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
Claims
1. A method for generating environmental videos, characterized in that, include: Acquire a first environmental video and a first pose trajectory; the first pose trajectory includes the first pose of the image acquisition device that acquired the first environmental video from a first viewpoint. The first pose contained in the first pose trajectory is changed to the second pose of the image acquisition device that acquires the first environmental video from the second viewpoint, so as to obtain the second pose trajectory. Based on the second pose trajectory and the reference image, a second environmental video from the second perspective is obtained; The reference image is the first frame of the video in the first environmental video.
2. The environmental video generation method according to claim 1, characterized in that, The first environmental video includes multiple video images, and the step of obtaining the second environmental video from the second viewpoint based on the second pose trajectory and the reference image includes: Acquire multiple recording frames corresponding to different acquisition times, wherein the recording frames include first point cloud data and the video image marked with the second pose; Multiple conditional images are obtained based on multiple recorded frames; The second environmental video is obtained based on the reference image and multiple condition images.
3. The environmental video generation method according to claim 2, characterized in that, The step of obtaining the second environmental video based on the reference image and multiple condition images includes: The reference image and multiple conditional images are input into a pre-constructed novel controllable video diffusion model to obtain the second environmental video through the novel controllable video diffusion model; The novel controllable video diffusion model is trained by taking the sample reference image and the sample condition image as input and the sample environment video from the sample perspective as the training target.
4. The environmental video generation method according to any one of claims 2 or 3, characterized in that, The step of obtaining multiple conditional images based on multiple recorded frames includes: For each recorded frame, the first point cloud data in the recorded frame is projected onto the video image in the recorded frame to obtain the projection points corresponding to each point in the first point cloud data. For each recorded frame, the pixel value of the projection point corresponding to each point in the first point cloud data in the recorded frame is determined as its own pixel value, so as to obtain colored second point cloud data; Based on multiple second-point cloud data, the object trajectory is obtained; The object point cloud and background point cloud are obtained from multiple second point cloud data using the object trajectory; Multiple second point cloud data sets whose collection time falls within a preset time window are aggregated into a unified point cloud data set; Transform multiple unified point cloud data sets into the world coordinate system; For each of the second point cloud data, a point rasterization operation is performed in the world coordinate system based on the set radius corresponding to the second point cloud data and in the second pose corresponding to the second point cloud data to obtain the conditional image corresponding to the second point cloud data.
5. The environmental video generation method according to claim 3, characterized in that, The novel controllable video diffusion model includes a first encoder, a second encoder, a decoder, a Video Denosing U-Net module, and a noise scheduling module; the step of inputting the reference image and multiple conditional images into the pre-constructed novel controllable video diffusion model to obtain the second environmental video through the novel controllable video diffusion model includes: The latent spatial features of the reference image are obtained through the first encoder; The second encoder is used to obtain the latent spatial features of the multiple conditional images; The latent spatial features of the reference image with added preset noise and the latent spatial features of the conditional image are input to the Video Denosing U-Net module through the noise scheduling module; The second environmental video, after noise removal and encoding, is obtained through the Video Denosing U-Net module; The decoder is used to decode the encoded second environment video to obtain the decoded second environment video.
6. An environmental video generation device, characterized in that, include: The first acquisition module is used to acquire the first environmental video and the first pose trajectory; The first pose trajectory includes the first pose of the image acquisition device that acquires the first environmental video from the first viewpoint; The second acquisition module is used to change the first pose contained in the first pose trajectory to the second pose of the image acquisition device that acquires the first environmental video from the second viewpoint, so as to obtain the second pose trajectory. The third acquisition module is used to obtain a second environmental video from the second perspective based on the second pose trajectory and the reference image; The reference image is the first frame of the video in the first environmental video.
7. The environmental video generation apparatus according to claim 6, characterized in that, The first environmental video includes multiple frames of video images, and the third acquisition module includes: The first acquisition unit is used to acquire multiple recording frames corresponding to different acquisition times. The recording frames include first point cloud data and the video image marked with the second pose. The second acquisition unit is used to obtain multiple conditional images based on multiple recorded frames; The third acquisition unit is used to obtain the second environmental video based on the reference image and multiple condition images.
8. A computer program product, characterized in that, It includes computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the environmental video generation method as described in any one of claims 1 to 5.
9. An electronic device, characterized in that, It includes at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program to enable the electronic device to implement the environmental video generation method as described in any one of claims 1 to 5.
10. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the environmental video generation method as described in any one of claims 1 to 5.