Dynamic video generation method and device, storage medium and electronic equipment
By denoising and adjusting the lighting of the target point cloud set collected by the lidar, dynamically controllable videos are generated, which solves the consistency problem of multimodal data fusion in complex urban scenes and achieves high-quality video generation and three-dimensional scene modeling.
Patent Information
- Application Number
- CN202510769461.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Existing technologies have difficulty achieving efficient fusion of multimodal data and instance-level spatiotemporal consistency expression in complex urban dynamic scenarios, which affects the depth of understanding and responsiveness of autonomous driving systems.
By denoising the target point cloud set collected by the lidar, using the pre-trained denoising network and time smoothing algorithm to eliminate noise, adjusting the light intensity and direction, and combining the vehicle's movement trajectory to generate dynamically controllable video.
The stability of image quality and the natural transition of lighting effects have been improved. The generated video is easy for users to observe and intelligently identify, and supports high-precision three-dimensional scene modeling and interactive operations.
Smart Images

Figure CN120634897A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of images, and in particular to a method, device, storage medium and electronic equipment for generating dynamic videos. Background Art
[0002] With the deep integration of artificial intelligence and digital twin technology, the accuracy and real-time performance of environmental perception and scene modeling have become core elements for improving the reliability of autonomous driving systems. Especially in dynamic urban scenarios with dense traffic, variable lighting, and frequent perspective switching, the efficient fusion of multimodal data and instance-level spatiotemporal consistency remain key bottlenecks hindering the implementation of autonomous driving technology. As the core vehicle connecting environmental perception and decision-making planning, the quality of vector maps directly impacts the depth of understanding and responsiveness of autonomous driving systems to complex scenarios. Existing technology systems still face significant shortcomings in meeting these challenges.
[0003] Shortcomings include how to construct dynamic video images for user observation and intelligent recognition. Summary of the Invention
[0004] The object of the present invention is to provide a method, device, storage medium and electronic device for generating dynamic video to improve the above-mentioned problem.
[0005] In order to achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for generating a dynamic video, the method comprising:
[0007] Denoising the conditional image to obtain a denoised image, comprising: using a pre-trained denoising network to remove noise points in the conditional image, and processing an output image of the denoising network using a temporal smoothing algorithm to eliminate image fluctuations and instability caused by noise to obtain a denoised image;
[0008] The conditional image is an image generated by rasterizing the target point cloud set collected by the laser radar deployed on the vehicle.
[0009] Adjusting the light intensity and light direction in the denoised image based on the vehicle's movement trajectory and the light intensity and light direction in the historical denoised image to obtain a light-adjusted image, wherein the historical denoised image is a plurality of frames of denoised images preceding the currently processed denoised image;
[0010] The illumination-adjusted image and its corresponding driving trajectory are input into a joint optimization framework of time series to generate a dynamic and controllable video.
[0011] In a second aspect, an embodiment of the present invention provides a dynamic video generation device, the device comprising:
[0012] a first processing unit, configured to perform denoising on the conditional image to obtain a denoised image, comprising: using a pre-trained denoising network to remove noise points in the conditional image, and using a temporal smoothing algorithm to process an output image of the denoising network to eliminate image fluctuations and instability caused by noise, thereby obtaining a denoised image;
[0013] The conditional image is an image generated by rasterizing the target point cloud set collected by the laser radar deployed on the vehicle.
[0014] The first processing unit is further configured to adjust the light intensity and light direction in the denoised image according to the vehicle's movement trajectory and the light intensity and light direction in the historical denoised image to obtain a light-adjusted image, wherein the historical denoised image is a plurality of frames of denoised images preceding the currently processed denoised image;
[0015] The second processing unit is configured to input the illumination adjustment image and its corresponding driving trajectory into a joint optimization framework of a time series to generate a dynamic controllable video.
[0016] In a third aspect, an embodiment of the present invention provides a storage medium having a computer program stored thereon, which implements the above method when executed by a processor.
[0017] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising: a processor and a memory, wherein the memory is used to store one or more programs; when the one or more programs are executed by the processor, the above method is implemented.
[0018] Compared with the prior art, the embodiments of the present invention provide a dynamic video generation method, device, storage medium and electronic device, which denoise the conditional image, use a pre-trained denoising network to delete noise points in the conditional image, and use a time series smoothing algorithm to process the output image of the denoising network to eliminate image fluctuations and instability caused by noise to obtain a denoised image; the conditional image is an image generated by point rasterization processing of a target point cloud set collected by a laser radar deployed on a vehicle; the light intensity and light direction in the denoised image are adjusted according to the vehicle's movement trajectory and the light intensity and light direction in the historical denoised image to obtain a light-adjusted image, where the historical denoised image is a multi-frame denoised image before the currently processed denoised image; the light-adjusted image and its corresponding driving trajectory are input into a joint optimization framework of a time series to generate a dynamic, controllable video for user observation and intelligent recognition.
[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A schematic structural diagram of an electronic device provided by an embodiment of the present invention.
[0022] Figure 2 This is a flow chart of a method for generating dynamic video according to an embodiment of the present invention.
[0023] Figure 3 The second flowchart of the dynamic video generation method provided by the embodiment of the present invention.
[0024] Figure 4 A schematic diagram of units of a dynamic video generation device provided by an embodiment of the present invention.
[0025] In the figure: 10 - processor; 11 - memory; 12 - bus; 13 - communication interface; 501 - first processing unit; 502 - second processing unit. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.
[0027] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but rather merely represents selected embodiments of the present invention. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort shall fall within the scope of protection of the present invention.
[0028] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0029] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0030] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0031] An embodiment of the present invention provides an electronic device, which may be a vehicle computer device or a server device connected to the vehicle computer device for communication. Figure 1 , a schematic diagram of the structure of an electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected via the bus 12. The processor 10 is used to execute executable modules stored in the memory 11, such as computer programs.
[0032] The processor 10 can be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the dynamic video generation method can be completed by the hardware integrated logic circuit in the processor 10 or the instructions in the form of software. The above-mentioned processor 10 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0033] The memory 11 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0034] The bus 12 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Figure 1 Only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus 12 or one type of bus 12.
[0035] The memory 11 is used to store programs, such as a program corresponding to the dynamic video generation device. The dynamic video generation device includes at least one software functional module that can be stored in the memory 11 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device. Upon receiving an execution instruction, the processor 10 executes the program to implement the dynamic video generation method.
[0036] Possibly, the electronic device provided by the embodiment of the present invention further includes a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.
[0037] It should be understood that Figure 1 The structure shown is only a schematic diagram of a portion of the electronic device. The electronic device may also include Figure 1More or fewer components than shown, or with Figure 1 Different configurations shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0038] The embodiment of the present invention provides a method for generating dynamic video, which can be applied to, but is not limited to, Figure 1 For detailed procedures, please refer to the electronic equipment shown in Figure 2 The dynamic video generation method includes: S11, S12 and S13, which are specifically described as follows.
[0039] S11, performing denoising processing on the conditional image to obtain a denoised image, including: using a pre-trained denoising network to delete noise points in the conditional image, and using a temporal smoothing algorithm to process the output image of the denoising network to eliminate image fluctuations and instability caused by noise to obtain a denoised image.
[0040] Among them, the conditional image is an image generated by point rasterization processing of the target point cloud set collected by the laser radar deployed on the vehicle.
[0041] This denoising process is particularly well-suited for dynamic urban scenes, where visual artifacts caused by camera motion or object movement can be effectively avoided, effectively avoiding artifacts and blur during the generation process and improving image quality stability. The denoising network utilizes a deep learning framework that combines a convolutional neural network (CNN) with residual connections. Its core task is to backpropagate the noise from the LiDAR-conditioned image, gradually optimizing image detail. Through multiple iterations of learning, the network fine-tunes the residual features of different image contents and continuously optimizes the depth and texture effects generated for each frame of video, ensuring high-quality visual results.
[0042] S12 , adjusting the illumination intensity and illumination direction in the denoised image according to the vehicle's movement trajectory and the illumination intensity and illumination direction in the historical denoised image to obtain an illumination-adjusted image.
[0043] The historical denoised images are multiple frames of denoised images before the currently processed denoised image.
[0044] Video generation not only involves geometric and structural consistency but also requires optimizing lighting effects. In complex urban scenes, lighting conditions often fluctuate significantly over time, especially considering changes in weather, time of day, and light source position. To overcome this issue, an optimization mechanism based on the lighting control module is introduced. This adjusts the lighting intensity and direction in the denoised image, ensuring that the generated video has natural lighting transitions and avoids lighting distortion or visual disharmony.
[0045] S13 , inputting the illumination-adjusted image and its corresponding driving trajectory (the posture trajectory of the vehicle when acquiring the target point cloud set) into a joint optimization framework of the time series to generate a dynamic controllable video.
[0046] The driving trajectory can be understood as the posture trajectory of the vehicle when acquiring the target point cloud set, including vehicle position information and vehicle orientation information.
[0047] In the dynamic video generation method provided in the embodiment of the present invention, a dynamic video image is constructed to facilitate user observation and intelligent recognition.
[0048] In an optional implementation, the training process of the time series joint optimization framework includes: S21 and S22, which are specifically described as follows.
[0049] S21, inputting the conditional image in the training phase and its corresponding driving trajectory into the joint optimization framework of the time series to obtain the controllable video in the training phase.
[0050] S22, constructing a first type of loss function according to the disassembled images of the controllable video in the training phase and the real environment images in the training phase to optimize the joint optimization framework of the time series.
[0051] Optionally, the first type of loss function includes a reconstruction loss function, an SSIM loss function, and a LPIPS loss function;
[0052] The reconstruction loss function is used to measure the pixel-level difference between the disassembled images of the controllable video during training and the real environment images during training.
[0053] The SSIM loss function is used to determine the structural consistency between the disassembled images of the controllable video in the training phase and the real environment images in the training phase.
[0054] The LPIPS loss function is used to measure the high-level semantic consistency between the disassembled images of the controllable video in the training phase and the real environment images in the training phase, further improving the realism of the visual quality.
[0055] Optionally, during the training phase or the actual operation phase, the process of generating a controllable video using the time series joint optimization framework includes: S31 and S32, which are specifically described as follows.
[0056] S31 , performing interpolation processing according to the driving trajectory to obtain a trajectory connecting line between the current input image and the next input image.
[0057] Ensure smooth transitions, ensure that the generated image is consistent with the physical trajectory, and reduce visual inconsistencies caused by irregular trajectory changes.
[0058] S32 , based on the current input image, the view angle is adjusted using the trajectory connection line as a geometric constraint to obtain a controllable video corresponding to the trajectory connection line.
[0059] In an alternative embodiment, please refer to Figure 3 After the illumination adjustment image and its corresponding driving trajectory are input into the joint optimization framework of the time series to generate a dynamic controllable video, the dynamic video generation method further includes: S14, S15, S16 and S17, which are specifically described as follows.
[0060] S14, performing three-dimensional scene modeling on sequence images in the controllable video based on Gaussian distribution;
[0061] S15 , determining a target foreground object (in motion) in the three-dimensional scene modeling data according to the depth information and optical flow estimation information of the sequence images in the controllable video.
[0062] The optical flow estimation information is defined as estimating the motion information of objects in the image by calculating the pixel displacement between adjacent frames, and the target foreground object is an identified object in motion.
[0063] S16, performing spatiotemporal modeling on each dynamic target region to obtain a corresponding spatiotemporal Gaussian cluster as a Gaussian foreground representation of the target foreground object.
[0064] The dynamic target area is a Gaussian ellipsoid centered on the target foreground object.
[0065] S17, using the controllable video as a supervisory signal, combined with Gaussian foreground representation and Gaussian background representation, to reconstruct the three-dimensional scene image to obtain a three-dimensional reconstructed video.
[0066] The Gaussian background representation is the Gaussian representation of the identified object in three-dimensional scene modeling except the Gaussian foreground representation.
[0067] It can display large-scale, dynamically changing 3D urban scenes in real time and supports interactive operations from different perspectives. Users can switch freely between different perspectives, and the system will display the complete scene structure and dynamic changes while maintaining high-precision image quality.
[0068] Please continue to refer to Figure 3 In an optional implementation, after obtaining the three-dimensional reconstructed video, the dynamic video generation method further includes: S18, as follows.
[0069] S18, rendering is performed based on the Gaussian weighted average of each pixel in the 3D reconstructed video to ensure the connection and consistency of the video in the time dimension.
[0070] In an optional implementation, after the training stage, the dynamic video generation method further includes: S19, as follows.
[0071] S19, obtaining evaluation optimization indicators of the controllable video in the training phase.
[0072] The evaluation optimization refers to the PSNR indicator and the FID indicator, which are used to optimize the joint optimization framework of the time series. The specific formula is as follows.
[0073]
[0074] Here, R is the maximum possible pixel value of the preset image (255 for 8-bit images), and MSE (mean squared error) is the mean squared difference between the decompressed image of the controllable video and the real environment image. By optimizing the PSNR metric, the system ensures that the generated image meets the expected standards of detail and clarity, especially for high-quality reproduction of complex scenes.
[0075] PSNR is used to evaluate the clarity of each frame in the generated scene, and the generation model is optimized by comparing the PSNR value of each frame to reduce noise in the image and ensure high-quality transitions between frames.
[0076] The FID metric is an important indicator for measuring the visual difference between the decomposition of a controllable video and the real-world image. It is widely used in the performance evaluation of generative adversarial networks (GANs) and diffusion models. It evaluates the authenticity of the generated results by calculating the difference between the feature distribution of the decomposition of the controllable video and the feature distribution of the real-world image. Specifically, FID calculates the distance between the mean and covariance matrix of the feature vectors of the generated image extracted by the pre-trained Inception network and the real image.
[0077] The formula is as follows:
[0078] FID=||μ r -μ g || 2 +Tr(∑r+∑g-2(∑r∑g) 1 / 2 )
[0079] Among them, μ r is the feature mean of the disassembled image of the controllable video, μ g is the feature mean of the real environment image, ∑r is the covariance matrix of the feature distribution of the disassembled image of the controllable video, ∑g is the covariance matrix of the feature distribution of the real environment image, and Tr is the trace of the sum and square root product of the covariance matrices, which more comprehensively evaluates the degree of difference between the two distributions. A lower FID value indicates that the feature distribution of the generated image is closer to that of the real image, that is, the generated result is more visually realistic and diverse.
[0080] Regarding the method of generating the conditional image, the embodiment of the present invention also provides an optional implementation method, please refer to the following.
[0081] The target point cloud set is colored according to the environment image corresponding to the target moment.
[0082] Among them, the target point cloud set is the point cloud data set collected by the laser radar deployed on the vehicle at the target time. The coloring processing refers to adding the corresponding color value in the environmental image to each point in the target point cloud set. The environmental image is the two-dimensional visual image obtained by the vehicle's image acquisition system at the target time, and the environmental image includes the color value of each pixel.
[0083] The target point cloud set after coloring is subjected to foreground and background separation processing to determine the dynamic foreground area and the static background area.
[0084] To highlight the saliency of target objects (moving objects) in dynamic foreground regions within the point cloud conditional image, a foreground-background separation mechanism based on geometric segmentation and semantic guidance (annotation) is introduced. This separation is performed on the colorized target point cloud set. This separation process not only ensures the overall structural coherence of the conditional image but also enhances the expressiveness of local regions, enabling the raw vector map generation model to maintain structural consistency for key target objects while focusing on perspective changes.
[0085] Combine the historical point cloud collection within the time window to complete the incomplete area.
[0086] Among them, the time window is a window of preset length between target moments, and the incomplete area is a dynamic foreground area with occlusion and incompleteness in the target point cloud set.
[0087] The completed target point cloud set is projected into the vehicle's image coordinate system using the perspective projection mapping relationship to obtain a first projection image.
[0088] The raster area in the first projection image is rendered to obtain a point rasterized conditional image.
[0089] Among them, the point rasterized conditional image is used as a training reference image for the vector map generation model to improve the training effect of the model and improve the quality of the vector map it generates.
[0090] Optionally, the target point cloud set is colored according to the environmental image corresponding to the target moment, including: using a perspective projection mapping relationship to project the target point cloud set to the image coordinate system of the vehicle to obtain a second projection image; wherein the perspective projection mapping relationship is a mapping relationship between the lidar coordinate system and the image coordinate system of the vehicle; obtaining the second projection image by projection to complete the mapping from the three-dimensional space structure to the two-dimensional image structure; using the color value of the i-th pixel in the environmental image as the color value of the i-th pixel in the second projection image; wherein 1≤i≤I, I is the total number of pixels in the environmental image; determining the corresponding point of the i-th pixel in the second projection image in the target point cloud set according to the perspective projection mapping relationship; adding the color value of the i-th pixel in the second projection image to its corresponding point in the target point cloud set.
[0091] Optionally, the environmental image is an image that has completed dynamic object annotation (such as cars, motorcycles, pedestrians, and bicycles) and static object annotation (such as roads, buildings, and trees). On this basis, the target point cloud set after coloring is subjected to foreground and background separation processing to determine the dynamic foreground area and the static background area therein, including: using spatial density, depth gradient, and normal vector as clustering reference factors (low-order geometric features) to preliminarily cluster the target point cloud set after coloring to obtain preliminary clustering areas; wherein the preliminary clustering areas are areas where the spatial density is greater than a density threshold, the fluctuation value of the depth gradient is greater than a gradient threshold, and the change value of the normal vector is greater than a vector threshold; when the proportion of dynamic points in the preliminary clustering areas is greater than or equal to a preset ratio, it is determined to be a dynamic foreground area; wherein the coordinates of the points in the preliminary clustering areas after being projected onto the image coordinate system of the vehicle are projection coordinates, and when the label of the pixel point corresponding to the projection coordinate in the environmental image is a dynamic object label, the point corresponding to the projection coordinate is a dynamic point; when the proportion of dynamic points in the preliminary clustering areas is less than a preset ratio, it is determined to be a static background area.
[0092] Since a single-frame lidar point cloud set may have problems such as sparse sampling, severe occlusion, and texture loss, it is difficult to achieve continuous and complete geometric expression in the image space. To this end, an embodiment of the present invention introduces a cross-frame point cloud aggregation mechanism to complete the target point cloud set and uniformly construct a dense point cloud that integrates multi-time observation information. Optionally, the incomplete area is completed in combination with the historical point cloud set within the time window, including: combining the vehicle motion trajectory, comparing the foreground dynamic area and background static area of the historical point cloud frame within the time window with the foreground dynamic area and background static area of the target point cloud set to determine the incomplete area; combining the posture change reference information and the vehicle motion trajectory to obtain the supplementary content of the incomplete area from the historical point cloud frame within the time window; adding the supplementary content to the incomplete area in the target point cloud set to complete the completion of the incomplete area.
[0093] Optionally, the raster region in the first projection image is rendered to obtain a point rasterized conditional image, including: constructing a target view cone model based on the corresponding position information of the laser radar at the target time; projecting the raster points into the vehicle's image coordinate system to obtain raster pixels, where the raster points are points in the target point cloud set that fall within the target view cone model; constructing a two-dimensional Gaussian kernel at the location of the raster pixel point, and determining the scale range of the two-dimensional Gaussian kernel based on the depth of the raster point; using the area within the scale range corresponding to the two-dimensional Gaussian kernel as the raster region; using the RGB, depth, point density, and normal vector corresponding to the raster region as the rendering target, and performing weighted fusion on a raster region basis to complete raster region rendering, thereby generating a conditional image that integrates geometric, texture, and shape information. By rendering the raster region in the first projection image, the final output point rasterized conditional image significantly outperforms traditional point cloud visualization results in terms of resolution, spatial consistency, and information density. It can be directly input into a vector map generation model as a high-quality prior condition, playing a key role in controlling view angle changes and lighting optimization.
[0094] See also Figure 4 , Figure 4 An embodiment of the present invention provides a dynamic video generation device. Optionally, the dynamic video generation device is applied to the electronic device described above.
[0095] A dynamic video generating device includes: a first processing unit 501 and a second processing unit 502.
[0096] The first processing unit 501 is configured to perform denoising on the conditional image to obtain a denoised image, including: using a pre-trained denoising network to remove noise points in the conditional image, and using a temporal smoothing algorithm to process the output image of the denoising network to eliminate image fluctuations and instability caused by noise, thereby obtaining a denoised image.
[0097] The conditional image is an image generated by rasterizing the target point cloud set collected by the laser radar deployed on the vehicle.
[0098] The first processing unit 501 is further configured to adjust the light intensity and light direction in the denoised image according to the vehicle's movement trajectory and the light intensity and light direction in the historical denoised images to obtain a light-adjusted image, wherein the historical denoised images are denoised images of multiple frames preceding the currently processed denoised image.
[0099] The second processing unit 502 is configured to input the illumination adjustment image and its corresponding driving trajectory into a joint optimization framework of a time series to generate a dynamic controllable video.
[0100] Optionally, the training process of the joint optimization framework of time series includes: inputting the conditional image of the training phase and its corresponding driving trajectory into the joint optimization framework of time series to obtain a controllable video of the training phase; constructing a first-class loss function based on the disassembled image of the controllable video of the training phase and the real environment image of the training phase to optimize the joint optimization framework of time series.
[0101] The second processing unit 502 may execute the above-mentioned S13 , and the first processing unit 501 may execute other steps in the above-mentioned method embodiment.
[0102] It should be noted that the dynamic video generation device provided in this embodiment can execute the method flow shown in the above method flow embodiment to achieve the corresponding technical effects. For the sake of brevity, for parts not mentioned in this embodiment, please refer to the corresponding content in the above embodiment.
[0103] An embodiment of the present invention further provides a storage medium storing computer instructions and programs that, when read and executed, execute the dynamic video generation method of the above embodiment. The storage medium may include a memory, a flash memory, a register, or a combination thereof.
[0104] The following provides an electronic device, which may be a vehicle computer device or a server device connected to the vehicle computer device for communication. Figure 1 As shown, the above-mentioned method for generating dynamic video can be implemented. Specifically, the electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs. When the one or more programs are executed by the processor 10, the method for generating dynamic video in the above-mentioned embodiment is executed.
[0105] In summary, the embodiments of the present invention provide a dynamic video generation method, device, storage medium and electronic device, which denoise a conditional image to obtain a denoised image, including: using a pre-trained denoising network to delete noise points in the conditional image, using a time series smoothing algorithm to process the output image of the denoising network, eliminating image fluctuations and instability caused by noise, so as to obtain a denoised image; wherein the conditional image is an image generated after point rasterization processing of a target point cloud set collected by a laser radar deployed on a vehicle; according to the vehicle's movement trajectory and the light intensity and light direction in the historical denoised image, the light intensity and light direction in the denoised image are adjusted to obtain a light-adjusted image, wherein the historical denoised image is a multi-frame denoised image before the currently processed denoised image; the light-adjusted image and its corresponding driving trajectory are input into a joint optimization framework of a time series to generate a dynamic controllable video for user observation and intelligent recognition.
[0106] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
[0107] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.
Claims
1. A method for generating dynamic video, characterized in that: The method comprises: Denoising the conditional image to obtain a denoised image, comprising: using a pre-trained denoising network to remove noise points in the conditional image, and processing an output image of the denoising network using a temporal smoothing algorithm to eliminate image fluctuations and instability caused by noise to obtain a denoised image; The conditional image is an image generated by rasterizing the target point cloud set collected by the laser radar deployed on the vehicle. Adjusting the light intensity and light direction in the denoised image based on the vehicle's movement trajectory and the light intensity and light direction in the historical denoised image to obtain a light-adjusted image, wherein the historical denoised image is a plurality of frames of denoised images preceding the currently processed denoised image; The illumination-adjusted image and its corresponding driving trajectory are input into a joint optimization framework of time series to generate a dynamic and controllable video.
2. The dynamic video generation method according to claim 1, wherein: The training process of the joint optimization framework of time series includes: The conditional images and their corresponding driving trajectories in the training phase are input into the joint optimization framework of time series to obtain the controllable video in the training phase; A first type of loss function is constructed according to the disassembled images of the controllable video in the training phase and the real environment images in the training phase to optimize the joint optimization framework of the time series.
3. The dynamic video generation method according to claim 2, wherein: The first type of loss function includes reconstruction loss function, SSIM loss function and LPIPS loss function; The reconstruction loss function is used to measure the pixel-level difference between the disassembled image of the controllable video in the training phase and the real environment image in the training phase; The SSIM loss function is used to determine the structural consistency between the disassembled images of the controllable video in the training phase and the real environment images in the training phase; The LPIPS loss function is used to measure the high-level semantic consistency between the disassembled images of the controllable video in the training phase and the real environment images in the training phase.
4. The dynamic video generation method according to claim 2, wherein: The process of generating a controllable video using the joint optimization framework of the time series includes: Perform interpolation processing based on the driving trajectory to obtain a trajectory connection line between the current input image and the next input image; Based on the current input image, the viewing angle is adjusted using the trajectory connection line as a geometric constraint to obtain a controllable video corresponding to the trajectory connection line.
5. The dynamic video generation method according to claim 1, wherein: After inputting the illumination-adjusted image and its corresponding driving trajectory into a time series joint optimization framework to generate a dynamic controllable video, the method further includes: Performing three-dimensional scene modeling on the sequence images in the controllable video based on Gaussian distribution; Determining a target foreground object in three-dimensional scene modeling data based on depth information and optical flow estimation information of a sequence of images in the controllable video; Performing spatiotemporal modeling on each dynamic target region to obtain a corresponding spatiotemporal Gaussian cluster as a Gaussian foreground representation of the target foreground object; wherein the dynamic target region is a Gaussian ellipsoid centered on the target foreground object; The controllable video is used as a supervisory signal, and a Gaussian foreground representation and a Gaussian background representation are combined to perform three-dimensional scene image reconstruction to obtain a three-dimensional reconstructed video.
6. The method for generating dynamic video according to claim 5, wherein: After obtaining the 3D reconstructed video, the method further includes: Rendering is performed based on the Gaussian weighted average of each pixel in the 3D reconstructed video to ensure the connection and consistency of the video in the time dimension.
7. A dynamic video generation device, characterized in that: The device comprises: a first processing unit, configured to perform denoising on the conditional image to obtain a denoised image, comprising: using a pre-trained denoising network to remove noise points in the conditional image, and using a temporal smoothing algorithm to process an output image of the denoising network to eliminate image fluctuations and instability caused by noise, thereby obtaining a denoised image; The conditional image is an image generated by rasterizing the target point cloud set collected by the laser radar deployed on the vehicle. The first processing unit is further configured to adjust the light intensity and light direction in the denoised image according to the vehicle's movement trajectory and the light intensity and light direction in the historical denoised image to obtain a light-adjusted image, wherein the historical denoised image is a plurality of frames of denoised images preceding the currently processed denoised image; The second processing unit is configured to input the illumination adjustment image and its corresponding driving trajectory into a joint optimization framework of a time series to generate a dynamic controllable video.
8. The dynamic video generation device according to claim 7, wherein: The training process of the joint optimization framework of the time series includes: inputting the conditional image and the corresponding driving trajectory of the training phase into the joint optimization framework of the time series to obtain a controllable video of the training phase; constructing a first-class loss function based on the disassembled image of the controllable video of the training phase and the real environment image of the training phase to optimize the joint optimization framework of the time series.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. An electronic device, characterized in that: include: a processor and a memory, the memory being configured to store one or more programs; When the one or more programs are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
SLAM and target tracking method
CN111060924A
Point cloud data denoising method and device, storage medium and electronic equipment
CN112435193A
Low-illumination video enhancement method and system and storage medium
CN115619674A
Point cloud data noise filtering method and device, storage medium and electronic equipment
CN117934856A
Optimization method and device for automatic driving simulation test
CN119918169A