Video generation method and device, electronic equipment, storage medium and chip

By acquiring and stitching depth maps, line drawings, and contour map videos, and combining them with target weather information to generate target videos, the problem of video generation in existing technologies not matching autonomous driving scenarios is solved, and the physical rationality and accuracy of the videos are improved.

CN120751074APending Publication Date: 2025-10-03XIAOMI TECH (WUHAN) CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510947498.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, the video generated based on images does not meet the requirements of autonomous driving scenarios for time-series video, and there is a contradiction between the intensity of weather effects and the physical rationality of objects and their movements in the video.

Method used

By obtaining the depth map video, line drawing video and contour map video of the video to be processed, and combining them with the target weather information, they are spliced ​​according to the video dimension channel to generate the target video, which contains the video features of the target weather information.

Benefits of technology

The physical rationality of video generation is improved, the contradiction between the intensity of weather effects and object movement is reduced, and the generated video is more in line with the needs of autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751074A_ABST
    Figure CN120751074A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment, a storage medium and a chip, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a depth map video, a line map video and a contour map video of a to-be-processed video, and obtaining the target weather information needed by the to-be-processed video; splicing the depth map video, the line draft map video and the contour map video according to a video dimension channel to obtain a spliced video; generating a target video according to the spliced video and the target weather information; the target video comprises video features of the target weather information. Processing is carried out based on the video, the demand of the automatic driving scene for the time sequence video is met, the depth map video provides the scene space structure, the line draft map video defines the object edge, the contour map video defines the main body shape, the physical rationality can be improved through joint constraint of the three, the target weather information is combined, and the weather effect intensity is adjusted. And the contradiction between the intensity of the weather effect of the video and the physical rationality of the object and the motion thereof in the video is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a video generation method, device, electronic device, storage medium, and chip. Background Art

[0002] With the continuous development of artificial intelligence (AI) technology, more and more vehicles with smart cockpits are equipped with autonomous driving systems, which require videos of different weather conditions for training.

[0003] Among the related technologies of video generation, video generation is usually based on images. The videos generated based on images do not meet the requirements of autonomous driving scenarios for time-series videos, and the intensity of the weather effects of the videos generated based on images is inconsistent with the physical rationality of the objects and their movements in the video. Summary of the Invention

[0004] The present disclosure provides a video generation method, device, electronic device, storage medium and chip to solve the problems in the related art.

[0005] A first embodiment of the present disclosure provides a video generation method, the method comprising:

[0006] Obtaining a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtaining target weather information required for the video to be processed;

[0007] splicing the depth map video, the line drawing video, and the outline map video according to the video dimension channel to obtain a spliced ​​video;

[0008] A target video is generated according to the spliced ​​video and the target weather information; the target video includes video features of the target weather information.

[0009] In some embodiments, the stitching of the depth map video, the line drawing video, and the outline map video according to the video dimension channel to obtain the stitched video includes:

[0010] Obtaining a first mask for the line drawing video and a second mask for the outline drawing video; the first mask is used to extract traffic elements in the line drawing video, and the second mask is used to clear sky elements in the outline drawing video;

[0011] According to the first mask and the second mask, the depth map video, the line drawing video, and the outline map video of the same video dimension channel are spliced ​​to obtain the spliced ​​video.

[0012] In some embodiments, stitching the depth map video, the line drawing video, and the outline map video of the same video dimension channel according to the first mask and the second mask to obtain the stitched video includes:

[0013] Performing a product calculation on the first mask and the line drawing video to obtain a line drawing video after the product calculation;

[0014] Performing a product calculation on the second mask and the contour map video to obtain a contour map video after the product calculation;

[0015] The depth map video, the line drawing video after the product calculation, and the contour map video after the product calculation of the same video dimension channel are spliced ​​together to obtain the spliced ​​video.

[0016] In some embodiments, generating a target video based on the spliced ​​video and the target weather information includes:

[0017] The spliced ​​video, the target weather information and the video to be processed are input into a trained video generation model to generate a video to obtain the target video.

[0018] In some embodiments, the video generation model includes a codec and a video generation network, and the step of inputting the spliced ​​video, the target weather information, and the to-be-processed video into the trained video generation model to generate the target video includes:

[0019] Using the codec, respectively determining a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed;

[0020] Using the video generation network, generate a video latent vector for the first latent vector, the second latent vector, and the third latent vector to obtain a video latent vector;

[0021] The codec is used to perform latent vector decoding processing on the video latent vector to obtain the target video.

[0022] In some embodiments, the determining, using the codec, a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed includes:

[0023] Inputting the spliced ​​video into the codec for latent vector encoding processing to obtain the first latent vector;

[0024] Inputting the target weather information into the codec for latent vector encoding processing to obtain the second latent vector;

[0025] The video to be processed is input into the codec for latent vector encoding processing to obtain the third latent vector.

[0026] In some embodiments, generating a video latent vector for the first latent vector, the second latent vector, and the third latent vector using the video generation network to obtain the video latent vector includes:

[0027] performing noise processing on the third latent vector to obtain a noisy third latent vector;

[0028] Inputting the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector, thereby obtaining a video latent vector to be processed;

[0029] Calling a preset denoising algorithm to perform denoising on the latent vector of the video to be processed to obtain the latent vector of the video.

[0030] In some embodiments, performing noise processing on the third latent vector to obtain the noisy third latent vector includes:

[0031] Calling a preset noising algorithm to perform noise processing on the third latent vector to obtain the noisy third latent vector; and / or,

[0032] The noisy third latent vector of the third latent vector is generated by using a preset normal distribution algorithm.

[0033] In some embodiments, inputting the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector to obtain a to-be-processed video latent vector includes:

[0034] In accordance with the direction in which the weather latent vector in the latent vector of the video to be processed tends toward the second latent vector, the latent vector generated by the video generation network is iteratively processed until the number of iterations reaches a preset number, thereby obtaining the latent vector of the video to be processed.

[0035] In some embodiments, the method further comprises:

[0036] Obtaining training stitched videos, training target weather information, and training videos to be processed;

[0037] The initial video generation model is trained based on the training spliced ​​video, the training target weather information and the training video to be processed to obtain the trained video generation model.

[0038] A second aspect of the present disclosure provides a video generation device, the device comprising:

[0039] An acquisition unit, configured to acquire a depth map video, a line drawing video, and a contour map video of a video to be processed, and to acquire target weather information required for the video to be processed;

[0040] a splicing unit, configured to splice the depth map video, the line drawing video, and the outline map video according to video dimension channels to obtain a spliced ​​video;

[0041] A generating unit is configured to generate a target video based on the spliced ​​video and the target weather information; the target video includes video features of the target weather information.

[0042] In some embodiments, the splicing unit includes:

[0043] an acquisition module, configured to acquire a first mask for the line drawing video and a second mask for the outline drawing video; the first mask is used to extract traffic elements in the line drawing video, and the second mask is used to clear sky elements in the outline drawing video;

[0044] A stitching module is used to stitch the depth map video, the line drawing video and the contour map video of the same video dimension channel according to the first mask and the second mask to obtain the stitched video.

[0045] In some embodiments, the splicing module is further configured to:

[0046] Performing a product calculation on the first mask and the line drawing video to obtain a line drawing video after the product calculation;

[0047] Performing a product calculation on the second mask and the contour map video to obtain a contour map video after the product calculation;

[0048] The depth map video, the line drawing video after the product calculation, and the contour map video after the product calculation of the same video dimension channel are spliced ​​together to obtain the spliced ​​video.

[0049] In some embodiments, the generating unit includes:

[0050] A generation module is used to input the spliced ​​video, the target weather information and the video to be processed into a trained video generation model to generate a video, so as to obtain the target video.

[0051] In some embodiments, the video generation model includes a codec and a video generation network, and the generation module is further configured to:

[0052] Using the codec, respectively determining a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed;

[0053] Using the video generation network, generate a video latent vector for the first latent vector, the second latent vector, and the third latent vector to obtain a video latent vector;

[0054] The codec is used to perform latent vector decoding processing on the video latent vector to obtain the target video.

[0055] In some embodiments, the generating module is further configured to:

[0056] Inputting the spliced ​​video into the codec for latent vector encoding processing to obtain the first latent vector;

[0057] Inputting the target weather information into the codec for latent vector encoding processing to obtain the second latent vector;

[0058] The video to be processed is input into the codec for latent vector encoding processing to obtain the third latent vector.

[0059] In some embodiments, the generating module is further configured to:

[0060] performing noise processing on the third latent vector to obtain a noisy third latent vector;

[0061] Inputting the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector, thereby obtaining a video latent vector to be processed;

[0062] Calling a preset denoising algorithm to perform denoising on the latent vector of the video to be processed to obtain the latent vector of the video.

[0063] In some embodiments, the generating module is further configured to:

[0064] Calling a preset noising algorithm to perform noise processing on the third latent vector to obtain the noisy third latent vector; and / or,

[0065] The noisy third latent vector of the third latent vector is generated by using a preset normal distribution algorithm.

[0066] In some embodiments, the generating module is further configured to:

[0067] In accordance with the direction in which the weather latent vector in the latent vector of the video to be processed tends toward the second latent vector, the latent vector generated by the video generation network is iteratively processed until the number of iterations reaches a preset number, thereby obtaining the latent vector of the video to be processed.

[0068] In some embodiments, the apparatus further comprises:

[0069] The acquisition unit is further used to acquire a spliced ​​video for training, target weather information for training, and a video to be processed for training;

[0070] The training unit is used to train the initial video generation model based on the training spliced ​​video, the training target weather information and the training video to be processed to obtain the trained video generation model.

[0071] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect embodiment of the present disclosure.

[0072] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the first aspect embodiment of the present disclosure.

[0073] The fifth aspect embodiment of the present disclosure proposes a chip, which includes one or more interfaces and one or more processors; the interface is used to receive signals from the memory of the electronic device and send signals to the processor, the signals including computer instructions stored in the memory, and when the processor executes the computer instructions, the electronic device executes the method described in the first aspect embodiment of the present disclosure.

[0074] In summary, according to the video generation method proposed in the present disclosure, the method includes obtaining a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtaining target weather information required for the video to be processed; splicing the depth map video, the line drawing video, and the contour map video according to the video dimension channel to obtain a spliced ​​video; generating a target video based on the spliced ​​video and the target weather information; the target video contains the video features of the target weather information. The solution disclosed in the present disclosure is based on video for processing, which meets the requirements of autonomous driving scenarios for time-series videos, and the depth map video provides the spatial structure of the scene, the line drawing video clarifies the edges of objects, and the contour map video defines the main shape. The three jointly constrain can improve physical rationality, and combined with the target weather information, the intensity of the weather effect is adjusted, reducing the contradiction between the intensity of the weather effect of the video and the physical rationality of the objects and their movements in the video.

[0075] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0077] Figure 1 A flowchart of a video generation method provided in an embodiment of the present disclosure;

[0078] Figure 2 A flowchart of a display diagram of a frame of an image in a video to be processed, a depth map video, a line drawing video, and a contour map video provided by an embodiment of the present disclosure;

[0079] Figure 3 A flowchart of another video generation method provided by an embodiment of the present disclosure;

[0080] Figure 4 A schematic diagram of a line drawing frame in a line drawing video provided by an embodiment of the present disclosure;

[0081] Figure 5 A schematic diagram of a frame of a contour image in a contour image video provided by an embodiment of the present disclosure;

[0082] Figure 6 A display diagram of a spliced ​​image of any frame of a spliced ​​video provided by an embodiment of the present disclosure;

[0083] Figure 7 A flowchart of another video generation method provided by an embodiment of the present disclosure;

[0084] Figure 8 A flowchart of another video generation method provided by an embodiment of the present disclosure;

[0085] Figure 9 A flowchart of another video generation method provided by an embodiment of the present disclosure;

[0086] Figure 10 A flowchart of a training process of a video generation model provided in an embodiment of the present disclosure;

[0087] Figure 11 A schematic structural diagram of a video generation device provided in an embodiment of the present disclosure;

[0088] Figure 12 A schematic structural diagram of another video generation device provided by an embodiment of the present disclosure;

[0089] Figure 13 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure;

[0090] Figure 14 A schematic diagram of the structure of a chip provided in an embodiment of the present disclosure;

[0091] Figure 15 A schematic diagram of the structure of another chip provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0092] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0093] With the continuous development of artificial intelligence (AI) technology, more and more vehicles with smart cockpits are equipped with autonomous driving systems, which require videos of different weather conditions for training.

[0094] Among the related technologies of video generation, video generation is usually based on images. The videos generated based on images do not meet the requirements of autonomous driving scenarios for time-series videos, and the intensity of the weather effects of the videos generated based on images is inconsistent with the physical rationality of the objects and their movements in the video.

[0095] Therefore, in order to solve the problems existing in the related art, the present disclosure proposes a video generation method, which obtains a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtains the target weather information required for the video to be processed; splices the depth map video, the line drawing video, and the contour map video according to the video dimension channel to obtain a spliced ​​video; generates a target video based on the spliced ​​video and the target weather information; the target video contains the video features of the target weather information. The solution disclosed in the present disclosure is based on video for processing, which meets the requirements of autonomous driving scenarios for time-series videos, and the depth map video provides the spatial structure of the scene, the line drawing video clarifies the edges of objects, and the contour map video defines the main shape. The three jointly constrain can improve physical rationality, and combined with the target weather information, the intensity of the weather effect is adjusted, which reduces the contradiction between the intensity of the weather effect of the video and the physical rationality of the objects and their movements in the video.

[0096] The embodiments of the present disclosure are not exhaustive and are merely illustrative of some embodiments, and are not intended to be a specific limitation on the scope of protection of the present disclosure. In the absence of contradiction, each step in a certain embodiment can be implemented as an independent embodiment, and the steps can be arbitrarily combined. For example, a solution after removing some steps in a certain embodiment can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment can be arbitrarily exchanged. In addition, the optional implementation methods in a certain embodiment can be arbitrarily combined; in addition, the embodiments can be arbitrarily combined. For example, some or all steps of different embodiments can be arbitrarily combined, and a certain embodiment can be arbitrarily combined with the optional implementation methods of other embodiments.

[0097] In each embodiment of the present disclosure, unless otherwise specified or provided for by logic, the terms and / or descriptions between the embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their inherent logical relationships.

[0098] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure.

[0099] In the embodiments of the present disclosure, unless otherwise specified, elements expressed in the singular, such as "a", "an", "the", "above", "said", "the", "the", etc., may mean "one and only one", or "one or more", "at least one", etc. For example, when using articles such as "a", "an", "the" in English in translation, the noun following the article may be understood as a singular expression or a plural expression.

[0100] In some embodiments, terms such as "in response to...", "in response to determining...", "in the case of...", "at the time of...", "when...", "if...", "if...", etc. can be used interchangeably.

[0101] In some embodiments, terms such as "greater than", "greater than or equal to", "not less than", "more than", "more than or equal to", "not less than", "higher than", "higher than or equal to", "not less than", and "above" can be replaced with each other, and terms such as "less than", "less than or equal to", "not greater than", "less than", "less than or equal to", "not more than", "lower than", "lower than or equal to", "not higher than", and "below" can be replaced with each other.

[0102] The prefixes such as "first" and "second" in the embodiments of the present disclosure are only used to distinguish different description objects and do not constitute any restrictions on the position, order, priority, quantity or content of the description objects. For the statement of the description objects, please refer to the description in the context of the claims or embodiments, and no unnecessary restrictions should be constituted due to the use of prefixes.

[0103] In the embodiments of the present disclosure, “plurality” refers to two or more.

[0104] In the embodiments of the present disclosure, terms such as “import”, “input”, and “read in” can be used interchangeably.

[0105] In some embodiments, devices, etc. can be interpreted as physical or virtual, and their names are not limited to the names recorded in the embodiments. Terms such as "device", "equipment", "device", "circuit", "network element", "node", "function", "unit", "section", "system", "network", "chip", "chip system", "entity", and "subject" can be used interchangeably.

[0106] In some embodiments, the terms "terminal", "terminal device", "user equipment (UE)", "user terminal", "mobile station (MS)", "mobile terminal (MT)", subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, etc. can be used interchangeably.

[0107] Figure 1 This is a flow chart of a video generation method provided by an embodiment of the present disclosure. This method can be applied to application scenarios such as smart terminals, for example, by a terminal with integrated video generation function or a video generator in the terminal, or by other devices suitable for video generation, and this disclosure is not limited thereto. Figure 1 As shown, the video generation method includes steps 101-103.

[0108] Step 101: Obtain a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtain target weather information required for the video to be processed.

[0109] The video to be processed refers to the original video that needs to be enhanced with weather effects. It is usually collected from the on-board camera in the autonomous driving scene and contains natural scenes (roads, sky, buildings) and traffic elements (vehicles, pedestrians, signs). The depth map is an image that represents the distance between the objects in the scene and the observation point. The depth map video consists of a series of depth map frames. Each depth map frame corresponds to the corresponding frame of the video to be processed, reflecting the distance information of each pixel in the video relative to the camera. It is usually presented in the form of a grayscale image. The higher the grayscale value, the farther the pixel is from the camera. The line drawing is an image that only contains the edges and contour lines of the object. The line drawing video is obtained by processing each frame of the video to be processed with an edge detection algorithm. It highlights the outline and structure of the objects in the scene, removes detailed information such as texture and color, and can clearly show the shape and boundaries of the object. The contour map is extracted from the depth map video. The contour map reflects the outline information of the objects in the scene. By processing the depth map, the contour lines of the objects can be obtained, thereby constructing the contour map video. Unlike the line drawing, the contour map focuses more on dividing the boundaries of the object based on the depth information, while the line drawing focuses on detecting edges based on the brightness changes of the image.

[0110] Target weather information refers to the weather conditions that users expect to appear in the generated video, such as sunny, rainy, snowy, foggy, etc. Weather information can be represented by a set of predefined weather parameters. Each weather parameter set contains characteristics related to a specific weather condition, such as light intensity, raindrop density, snowflake size, fog concentration, etc.

[0111] A depth estimation algorithm, such as one based on a convolutional neural network, is used to process each frame of the video to estimate a depth map. A deep learning-based edge detection model is used to detect edges in the video to generate a line drawing video. The depth map video is processed by first smoothing the depth map to reduce noise. The gradient of the depth map is then calculated to identify areas with significant depth variations, i.e., the contour locations of the object. A gradient operator is used to calculate the gradient magnitude of the depth map. Pixels with gradient magnitudes greater than a certain threshold are marked as contour points, thus generating a contour map video.

[0112] In order to facilitate a better understanding of the processed video, depth map video, line drawing video and contour map video, such as Figure 2 As shown, Figure 2 This is a display of a frame of image from a video to be processed, a depth map video, a line drawing video, and a contour map video provided by the embodiments of the present disclosure. Traffic elements such as vehicles, pedestrians, traffic signs, and lane characteristics are very important for self-driving scenarios. To ensure the consistency of the generated content with the traffic elements of the input video, multiple control conditions are extracted from a frame of the video to be processed. The depth map, contour map, and line drawing image are extracted from the original image frame as control conditions.

[0113] Through a user interface, such as a graphical interface or command line interface, the user can select or input the desired weather type. For example, a drop-down menu may be provided, including options such as sunny, rainy, snowy, and foggy. After the user selects a setting, the system retrieves the corresponding target weather information based on a predefined set of weather parameters. Alternatively, based on user needs or system preset conditions, the system can query and extract the corresponding target weather information from a database containing various weather characteristic parameters. For example, to train an autonomous driving model, parameters under specific weather conditions can be selected from a database containing various weather characteristic parameters as target weather information.

[0114] By acquiring depth map videos, line drawing videos, and outline map videos, we can provide richer and more accurate scene information for video generation. The depth map video provides the depth relationship of the scene, the line drawing video highlights the outline and structure of objects, and the outline video further emphasizes the boundaries of objects. The combination of depth map videos, line drawing videos, and outline map videos helps generate videos that are more realistic and visually realistic, improving the accuracy and quality of video generation.

[0115] Step 102: splice the depth map video, the line drawing video, and the outline map video according to the video dimension channel to obtain a spliced ​​video.

[0116] Video dimension channels refer to the type of information represented by each dimension of different video layers or channels during multidimensional video processing. For example, the three color channels (Red, Green, and Blue, RGB) of a video, the time dimension, the depth dimension, etc. In this embodiment, video dimension channels specifically refer to the information levels of the depth map video, line drawing video, and outline map video. The information levels are the different channel data in the video splicing process.

[0117] Video stitching combines independent video data from multiple sources according to specific rules to form a new, continuous video file. In this embodiment, a composite video containing multi-dimensional information is formed by stitching together depth map videos, line drawing videos, and outline videos. The stitching process can be performed using time synchronization, spatial alignment, or dimensional channel merging.

[0118] To stitch together the depth map video, line drawing video, and outline image video, the timelines of the three videos must first be aligned. The timestamps of each frame must be consistent to ensure temporal alignment of the data across the three video channels. Once aligned, computer vision algorithms (such as those based on optical flow or other frame-to-frame matching techniques) are used to synchronize the spatial positions of the three videos to ensure visual consistency. After aligning each dimension, a custom stitching algorithm is used to spatially combine the depth map, line drawing, and outline image information from each frame. A common stitching method is to stack or overlay each image information channel in a specific channel order, for example: the depth map video as the first channel, the line drawing video as the second channel, and the outline video as the third channel. During this process, techniques such as transparency adjustment and color mapping can be used to ensure that each video dimension is merged without information loss and that the data from each dimension is fully displayed. After stitching, the stitched video channels are combined to generate a composite video file. This video file not only contains information from each dimension but also dynamically displays data such as the object's shape, depth, and outline.

[0119] Stitching depth map videos, line drawing videos, and silhouette image videos can integrate scene information from different modalities into a single video. The depth map provides depth relationships in the scene, the line drawing highlights the edges and outlines of objects, and the silhouette image further emphasizes object boundaries. This fusion results in a stitched video with richer and more comprehensive scene features, which is beneficial for subsequent video processing tasks such as object detection, recognition, and tracking, improving the accuracy and robustness of these tasks.

[0120] Step 103: Generate a target video based on the spliced ​​video and the target weather information; the target video includes video features of the target weather information.

[0121] The target video is a composite video with specific weather characteristics generated from the stitched video and target weather information. The target video is generated by combining the stitched video with the target weather information, so that the visual features in the video match the target weather conditions, creating a dynamic scene with realistic weather effects. The video features of the target weather information refer to the changes associated with specific weather conditions that are presented after the visual effects of the stitched video are adjusted based on the target weather information. For example, on a rainy day, the video may show rain effects, cloud changes, and air humidity; on a sunny day, the video may adjust the lighting, shadows, and color temperature. The video features of the target weather information are the specific manifestations presented through these adjustments.

[0122] Using pre-trained neural network-based diffusion models (Diffusion Models with Transformers, DiT) or other advanced video generation models, the spliced ​​video and target weather information are used as input conditions for the generation model, and the target video is directly generated through the generator part of the model. During the training phase, a large amount of real video data with different weather annotations is used to train the model, so that the model learns how to generate corresponding video content based on the input scene features and weather conditions. During the generation process, the generator automatically adjusts the pixel values ​​of the generated image to make it conform to the visual characteristics of the target weather and match the scene information in the spliced ​​video. In addition, by introducing an adversarial training mechanism, the discriminator can distinguish between the generated video and the real video, thereby continuously improving the generation quality of the generator, making the generated target video closer to the real scene in terms of details and overall effect.

[0123] By combining multimodal information such as depth, line art, and outlines from the stitched video with target weather information, it is possible to accurately generate videos that match specific weather characteristics. Whether it's the reflection of water on a rainy day, the accumulation of snow on a snowy day, or the blurred vision in fog, these effects are matched to the objects and their motion in the scene, making the generated videos more visually realistic and believable. This provides autonomous driving systems with training data that is closer to actual driving environments, improving their reliability and safety in diverse weather conditions.

[0124] In summary, according to the video generation method proposed in the present disclosure, the method includes obtaining a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtaining target weather information required for the video to be processed; splicing the depth map video, the line drawing video, and the contour map video according to the video dimension channel to obtain a spliced ​​video; generating a target video based on the spliced ​​video and the target weather information; the target video contains the video features of the target weather information. The solution disclosed in the present disclosure is based on video for processing, which meets the requirements of autonomous driving scenarios for time-series videos, and the depth map video provides the spatial structure of the scene, the line drawing video clarifies the edges of objects, and the contour map video defines the main shape. The three jointly constrain can improve physical rationality, and combined with the target weather information, the intensity of the weather effect is adjusted, reducing the contradiction between the intensity of the weather effect of the video and the physical rationality of the objects and their movements in the video.

[0125] As a refinement of step 102, when executing the stitching of the depth map video, the line drawing video and the contour map video according to the video dimension channel to obtain the stitched video, it can be implemented in but not limited to the following ways, including: obtaining a first mask of the line drawing video, and obtaining a second mask of the contour map video; the first mask is used to extract traffic elements in the line drawing video, and the second mask is used to clear the sky elements in the contour map video; according to the first mask and the second mask, the depth map video, the line drawing video and the contour map video of the same video dimension channel are stitched to obtain the stitched video.

[0126] The first mask is a binary mask used to extract traffic elements from the line art video. Areas with a value of 1 represent traffic elements (vehicles, pedestrians, and road signs), and areas with a value of 0 represent the background. The second mask is a binary mask used to eliminate sky elements from the outline video. Areas with a value of 0 represent the sky, and areas with a value of 1 represent other areas (ground, buildings, etc.). The mask is multiplied pixel by pixel with the corresponding video to selectively retain or suppress specific areas in the video.

[0127] The splicing process can be achieved by the following formula:

[0128] C a =concat{C d ,C l ⊙M obj ,C s ⊙(1-M sky )}

[0129] Among them, C a is any frame image in the spliced ​​video, concat is splicing, C dis a frame of depth map in the depth map video corresponding to any frame image in the spliced ​​video, C l is a line drawing in the line drawing video corresponding to any frame image in the spliced ​​video, C s is a frame of contour image in the contour image video corresponding to any frame image in the spliced ​​video, M obj is the first mask, (1-M sky ) is the second mask, M sky Mask for the sky.

[0130] When processing line-art video, a first mask is extracted using features based on traffic elements (such as vehicles and pedestrians). The areas containing traffic elements are identified using deep learning or traditional image processing methods (such as color or shape detection). When processing silhouette video, a second mask is obtained by identifying and detecting the sky area. The sky is typically a bright spot in a scene, and a mask can be created through color analysis or feature-based region recognition. Traffic elements are extracted using the first mask. In specific implementations, the mask can be applied to retain elements such as vehicles and roads while removing irrelevant portions. A second mask is used to clear the sky elements in the silhouette image. This mask eliminates the outline information of the sky, retaining only the outlines of other elements such as the ground and buildings. The processed depth image video, line-art video, and silhouette image video are then stitched together. The stitching process combines information from each video into a new composite video based on the matching of the video dimensional channels. Typically, the dimensional information complements each other visually. For example, the depth image provides three-dimensional spatial information, the line-art image emphasizes object outlines, and the silhouette image provides shape features. The resulting stitched video not only captures multi-dimensional information like depth, lines, and contours, but also removes unnecessary background elements, such as the sky, ensuring that each element of the video is more accurate and clear. The resulting stitched video can then be processed or displayed.

[0131] By extracting and blanking out specific elements, traffic elements and scene outlines in the video can be presented more accurately. After removing irrelevant background, the visibility and clarity of the target elements are enhanced.

[0132] Figure 3 The flow chart of a video generation method proposed in the present disclosure is further shown. Figure 1 The embodiment shown further explains the above embodiment. Figure 3 The following steps may be included:

[0133] Step 201 : performing a product calculation on the first mask and the line drawing video to obtain a line drawing video after the product calculation.

[0134] A product calculation involves element-by-element multiplication of two matrices of the same size (such as images or masks). For each position in the matrix, the values ​​at the corresponding position in the two matrices are multiplied together to produce a new matrix. In this embodiment, the product calculation is used to perform a pixel-by-pixel multiplication of the first mask with each frame of the line art video to extract the traffic elements and leave the non-traffic elements blank.

[0135] Each frame of the line drawing video is processed using a semantic segmentation model. The model is trained to recognize and label traffic elements (such as vehicle outlines, road markings, traffic signs, etc.) in the line drawing. The semantic segmentation model outputs a binary mask of the same size as the input line drawing, where the pixel value of the traffic element area is 1 and the pixel value of the non-traffic element area is 0. The binary masks obtained for each frame are combined into a sequence to form a first mask video, and its frame rate is consistent with the line drawing video. A pixel-by-pixel product is calculated for each frame of the line drawing video and the corresponding first mask frame. Specifically, for each pixel in the line drawing, if the value of the corresponding position in the first mask is 1 (that is, the pixel belongs to a traffic element), the value of the pixel in the line drawing is retained; if the value of the corresponding position in the first mask is 0 (that is, the pixel does not belong to a traffic element), the value of the pixel in the line drawing is set to 0 (that is, blank). For example, if a pixel in the line drawing has a value of 255 (indicating an edge) and the corresponding position in the first mask has a value of 1, the pixel value remains 255 after the product calculation. If the corresponding position in the first mask has a value of 0, the pixel value becomes 0 after the product calculation. The line drawing frames after the product calculation are arranged in chronological order to form a new video sequence, namely the line drawing video after the product calculation. This video only retains the traffic elements in the original line drawing. Non-traffic element areas are left blank, and the background becomes pure black (pixel value 0) or another specified background color.

[0136] In order to better understand the line drawing video after product calculation, Figure 4 As shown, Figure 4 This is a schematic diagram of a line drawing frame in a line drawing video provided by an embodiment of the present disclosure. The 2D detection frame and image segmentation mask of the object in the original image frame are marked by manual or automatic annotation. These masks are used to selectively remove non-essential areas in the control image. For the line drawing, with the help of the 2D frame of object detection, only the object areas that are more important to the self-driving scene are retained, such as pedestrians, vehicles, and traffic element areas of traffic lights.

[0137] The line drawing video after product calculation reduces the amount of data to be processed because non-traffic element areas are cleared. This can significantly reduce computing resource consumption and increase processing speed during subsequent video generation and processing, especially when processing high-resolution or long-duration videos.

[0138] Step 202: Perform product calculation on the second mask and the contour image video to obtain a contour image video after product calculation.

[0139] In this embodiment, the product calculation is used to multiply the second mask by each frame of the contour map video pixel by pixel, thereby clearing the sky element area and retaining the non-sky element area.

[0140] Each frame of the contour map video is analyzed to determine the area of ​​sky elements. Since the sky is typically represented by a relatively flat area with low grayscale values ​​in the contour map, an appropriate threshold can be set based on these characteristics. For example, pixels with grayscale values ​​below a certain threshold (such as 50) are marked as 0 (sky elements), and other areas are marked as 1 (non-sky elements). In this way, a second mask corresponding to each frame of the contour map video is generated. Each frame of the contour map video is processed using a pre-trained semantic segmentation model. The model outputs a binary mask with the same size as the input contour map, where the pixel values ​​of the sky element area are 0 and the pixel values ​​of the non-sky element area are 1. The binary masks obtained for each frame are combined into a sequence to form a second mask video with the same frame rate as the contour map video. A pixel-by-pixel product is calculated for each frame of the contour map video and the corresponding second mask frame. Specifically, for each pixel in the contour map, if the value of the corresponding position in the second mask is 1 (i.e., the pixel belongs to a non-sky element), the value of the pixel in the contour map is retained; if the value of the corresponding position in the second mask is 0 (i.e., the pixel belongs to a sky element), the value of the pixel in the contour map is set to 0 (i.e., blank). For example, if the value of a pixel in the contour map is 255 (indicating an object outline) and the value of the corresponding position in the second mask is 1, the pixel value remains 255 after the product calculation; if the value of the corresponding position in the second mask is 0, the pixel value becomes 0 after the product calculation. The contour map frames after the product calculation are arranged in chronological order to form a new video sequence, namely the contour map video after the product calculation. This video retains the non-sky element portion of the original contour map, the sky element area is blanked, and the background becomes pure black (pixel value 0) or another specified background color.

[0141] To better understand the contour video after product calculation, Figure 5 As shown, Figure 5 This is a schematic diagram of a frame of a contour image in a contour image video provided by an embodiment of the present disclosure. For the contour image, a sky mask is used to remove all elements in the sky area.

[0142] The contour image video after product calculation is more concise and clear, highlighting the outline information of the ground and objects. This makes the generated target video more visually clear, making it easier for human observers or subsequent image analysis algorithms to understand the scene content. Especially in complex scenes or in adverse weather conditions, key information such as road edges and vehicle outlines can be captured more quickly.

[0143] Step 203 : splicing the depth map video, the line drawing video after the product calculation, and the contour map video after the product calculation of the same video dimension channel to obtain the spliced ​​video.

[0144] The line drawing video after product calculation is the video obtained by processing the original line drawing video. Specifically, it extracts the traffic element part by performing pixel-by-pixel product calculation on the line drawing video and the first mask, while the non-traffic element area is cleared (pixel value is 0). The contour map video after product calculation is the video obtained by processing the original contour map video. By performing pixel-by-pixel product calculation on the contour map video and the second mask, the sky element area is cleared (pixel value is 0), retaining the contour information of the ground and objects.

[0145] First, ensure that the depth map video, the post-product line drawing video, and the post-product contour video have the same frame rate and duration. If the frame rates are different, interpolate or decimate the videos to bring them to the same frame rate. Also, ensure that all videos have consistent resolution in the spatial dimensions (height and width). For example, if the depth map video has a frame rate of 30 fps, while the post-product line drawing and contour video have frame rates of 25 fps, an interpolation algorithm can be used to interpolate the frames of the latter two videos to bring their frame rates up to 30 fps, synchronizing them with the depth map video. The depth map video, post-product line drawing video, and post-product contour video are each read into memory and represented as a multidimensional array. Concatenate the three videos along the temporal dimension. Specifically, stack each frame of the three videos along the channel dimension. For example, for frame t along the temporal dimension, stack the depth map frame, post-product line drawing frame, and post-product contour frame sequentially to form a stitched frame containing multiple channels. All stitched frames are combined to form the final stitched video.

[0146] In order to better understand the spliced ​​video, Figure 6 As shown, Figure 6This is a display of a stitched image from any frame of a stitched video provided in an embodiment of the present disclosure. A stitched video is a composite video obtained by stitching together a depth map video, a line drawing video after product calculation, and a contour map video after product calculation, along the video dimension channels. It contains information from multiple modalities, including depth information, traffic element outlines, and ground and object outlines, providing a rich data foundation for subsequent video processing and analysis.

[0147] By stitching together the depth map video, the line drawing video after product calculation, and the contour map video after product calculation, we can integrate scene information from different modalities into a single video. The depth map provides the depth relationship of the scene, the line drawing highlights the outline and structure of traffic elements, and the contour map further emphasizes the boundaries between the ground and objects. This fusion enables the stitched video to contain richer and more comprehensive scene features, providing a more solid data foundation for subsequent video processing tasks such as adding weather effects and object detection and recognition.

[0148] As a refinement of step 103, when generating the target video based on the spliced ​​video and the target weather information, it can be implemented in but not limited to the following ways, including: inputting the spliced ​​video, the target weather information and the video to be processed into a trained video generation model to generate the video and obtain the target video.

[0149] The model obtained by training the initial model with a large amount of video data with weather annotations has learned the mapping relationship between different weather characteristics and video content, and can generate target videos that meet the requirements based on the input spliced ​​video, target weather information and the video to be processed.

[0150] Standardize the stitched video, target weather information, and videos to be processed to have the same frame rate and resolution to ensure data consistency and compatibility. For example, adjust the frame rate of all videos to 30 fps and the resolution to 1280 × 720 pixels.

[0151] Encode the target weather information and convert it into a numerical value or vector form that the model can recognize. For example, using one-hot encoding, a sunny day is represented as [1,0,0,0], a rainy day is represented as [0,1,0,0], a snowy day is represented as [0,0,1,0], and a foggy day is represented as [0,0,0,1].

[0152] The preprocessed spliced ​​video, the target weather information, and the video to be processed are fed into a trained video generation model as input. The video generation model first extracts features from these input data, extracting features such as depth, line art, and outlines from the spliced ​​video, weather features from the target weather information, and content features from the video to be processed. The model then fuses these features using an internal network structure, such as a codec-decoder architecture or a generative adversarial network (GAN). For example, in an adversarial network-based model, the generator fuses features from the spliced ​​video, the target weather information, and the video to be processed to generate a preliminary target video. The discriminator then evaluates the generated video to determine whether it conforms to the target weather characteristics and the plausibility of the video content. The generator continuously adjusts and optimizes the generated video based on the discriminator's feedback. After multiple iterations and optimizations within the video generation model, the target video is ultimately output. This target video retains basic information from the video to be processed, such as objects and their motion trajectories, while incorporating weather features described by the target weather information. For example, in a rainy target video, raindrops fall from the sky, puddles form on the ground, and vehicles and pedestrians get wet.

[0153] This method provides a rich and diverse source of weather scene video data for autonomous driving systems. In practical applications, collecting real-world video data in various weather conditions often faces numerous challenges, such as high cost, high risk, and data imbalance. However, this video generation method can generate a large number of videos under different weather conditions from limited raw video data, effectively expanding the training dataset for autonomous driving systems. This helps improve their performance and reliability in complex and changing real-world driving environments, accelerating the development and application of autonomous driving technology.

[0154] Figure 7 The flow chart of a video generation method proposed in the present disclosure is further shown. The video generation model includes a codec, a video generation network, and a video generation network based on the video generation model. Figure 7 The embodiment shown further explains the above embodiment. Figure 7 The following steps may be included:

[0155] Step 301: Using the codec, determine a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed.

[0156] A codec is a neural network structure that converts input data (such as a spliced ​​video, target weather information, and the video to be processed) into a low-dimensional latent vector representation. The codec learns features and patterns in the data, extracting key information for use in subsequent processing and generation. For example, the codec can convert each frame or sequence of video into a fixed-length vector containing the video's characteristic information at that moment, such as color, shape, and motion. A latent vector is the low-dimensional vector representation obtained after the codec encodes the input data. The latent vector can be viewed as a compressed representation of the input data within the model, capturing the data's key features and inherent patterns. For example, the first latent vector for the spliced ​​video contains the fused feature information of the depth map, line drawing, and outline map in the spliced ​​video; the second latent vector for the target weather information contains features such as the target weather type and intensity; and the third latent vector for the video to be processed contains content features of the original video, such as the appearance and motion trajectory of objects. For example, the latent vector for a sunny day might contain information related to sunny weather, such as sunlight intensity and sky color.

[0157] Normalize the spliced ​​video, target weather information, and the processed video to the same range (e.g., [0, 1] or [-1, 1]) to improve model training stability and convergence speed. For example, divide the value of each pixel in the spliced ​​video by 255 to normalize it to the range [0, 1]. Standardize the numerical parameters in the target weather information to have zero mean and unit variance. Encode the target weather information to a format suitable for the input codec. One-hot encoding or embedded encoding can be used. For example, one-hot encoding maps different weather types to different binary vectors, such as [1, 0, 0, 0] for sunny days, [0, 1, 0, 0] for rainy days, [0, 1, 0, 0] for snowy days, and [0, 0, 0, 1] for foggy days. If the target weather information also contains continuous numerical parameters (such as rainfall amount), it can be concatenated with the one-hot encoded vector to form a complete vector representation. The spliced ​​video, target weather information, and the video to be processed are divided into multiple small segments or frame sequences to facilitate model processing and training. For example, the video is divided into sequence segments of 16 or 32 frames at a fixed time interval (such as 30 frames per second), and each segment is input into the codec as an independent input sample for encoding. The preprocessed spliced ​​video is input into the codec, and the codec extracts and encodes features of each frame of the spliced ​​video and its temporal relationship through convolutional layers and recurrent layers (if any). The convolutional layer can extract local features in each frame of the image, such as the boundaries of objects in the depth map, the shape of lines in the line drawing, and the outline of objects in the contour map; the recurrent layer can capture the motion and change information between video frames, such as the movement trajectory of objects and the evolution of weather phenomena. After multiple layers of convolution and recurrence, the codec outputs the first latent vector of the spliced ​​video. This latent vector contains comprehensive feature information such as depth, line drawings, outlines, and time series. The target weather information is input into the codec, which encodes the vector representation of the target weather information based on its structure. If the target weather information is a vector obtained through one-hot encoding, the weather type features are extracted. Finally, the codec outputs the second latent vector of the target weather information, which represents the characteristic information of the target weather, such as weather type and intensity.

[0158] The video to be processed is fed into the codec, which similarly uses convolutional layers and recurrent layers (if any) to extract features and encode them. The convolutional layers extract features such as the appearance, color, and texture of objects in the video frames, while the recurrent layers capture the object's motion trajectory and the temporal dynamics of the video. The codec outputs a third latent vector for the video to be processed. This latent vector contains features of the original content of the video, such as the object's shape, position, and motion state.

[0159] By using codecs to determine the latent vectors for the spliced ​​video, target weather information, and the video to be processed, different types of input data can be converted into a unified low-dimensional vector representation, effectively extracting key features and information from the data. This latent vector representation not only reduces the dimensionality and complexity of the data, but also preserves the data's key characteristics and inherent patterns, providing a concise and informative representation for subsequent video generation. Compared to directly processing the raw data, the latent vector can more efficiently capture the essential characteristics of the data, improving the model's computational efficiency and processing speed, while also contributing to improved quality and accuracy of the generated video.

[0160] Step 302: Use the video generation network to generate a video latent vector for the first latent vector, the second latent vector, and the third latent vector to obtain a video latent vector.

[0161] A video generation network (VGN) is a deep learning model specifically designed to generate high-quality video content from latent vectors. It typically consists of a generator that generates a corresponding sequence of video frames from a latent vector input. VGNs utilize various network structures, such as convolutional neural networks and recurrent neural networks, to process the latent vectors and generate videos that meet the desired quality. The first latent vector is a low-dimensional vector generated by the codec from the spliced ​​video. It contains information such as the structure, depth, object outlines, and scene layout of the spliced ​​video. This first latent vector represents the global characteristics of the spliced ​​video and serves as one of the inputs to the VGN. The second latent vector is a low-dimensional vector generated by the codec from the target weather information. It contains abstract features such as weather type and climate conditions. This second latent vector represents the characteristics of the target weather and is intended to control the presentation of weather effects in the generated video. The third latent vector is a low-dimensional vector generated by the codec from the target video. It contains spatial information, temporal information, and dynamic changes in the video. This third latent vector represents the core characteristics of the target video and serves as one of the basic inputs for generating the target video. Video latent vector generation refers to the use of a video generation network to generate a new low-dimensional representation (i.e., a video latent vector) based on the input latent vectors (such as the first latent vector, the second latent vector, and the third latent vector) through various processing modules in the model (such as convolutional and fully connected layers). This video latent vector contains the required features of the target video and is used to guide the subsequent video generation process.

[0162] The three latent vectors are directly concatenated into a single long vector as the network input. This concatenation and fusion preserves the original information of each latent vector, but may result in higher dimensionality and increased computational complexity. In implementation, the concatenated vector can be fed into a fully connected or convolutional layer for further processing. When the three latent vectors have the same dimensionality, element-wise addition can be performed directly. This approach assumes that the features of different latent vectors at corresponding positions complement each other, and the dimensionality of the fused vector is the same as that of a single latent vector, which helps maintain information compactness. An activation function is often added after additive fusion to introduce nonlinearity. An attention mechanism dynamically adjusts the weights of the contributions of different latent vectors to the generated output. The attention mechanism can automatically learn the importance of each latent vector at different positions or time steps based on the target weather information or content characteristics of the video being processed, thereby achieving adaptive feature fusion. For example, a multi-head attention mechanism can be used to model the interaction between the three latent vectors to generate a weighted fused latent vector for the video.

[0163] By fusing and processing these latent vectors through the video generation network, a high-quality latent vector for the target video can be generated. This latent vector comprehensively considers the multi-dimensional characteristics of the spliced ​​video, the target weather information, and the content characteristics of the processed video, making the generated video more realistic and reasonable in terms of object appearance, motion trajectory, and weather effects. For example, when generating rainy scenes, the video latent vector can guide the decoder to generate videos with raindrops, puddles, and slippery road effects, while maintaining the consistency and naturalness of object motion.

[0164] Step 303: Use the codec to perform latent vector decoding processing on the video latent vector to obtain the target video.

[0165] In addition to the encoding function, the codec also has the encoding function. The latent vector decoding process is to convert the latent vector back into high-dimensional data, such as a video frame sequence, through the codec (usually a decoding network).

[0166] The video latent vector is input into the codec. The video latent vector can be a fixed-length vector or a sequence (for time-series video data). The codec maps the video latent vector into a high-dimensional feature space through its network structure (such as deconvolutional layers and fully connected layers). In the deconvolutional network, the video latent vector is first mapped to an initial feature map through a fully connected layer. This is then gradually upsampled through multiple layers of deconvolution operations, increasing the size and number of channels of the feature map. Based on the feature map, the codec generates each frame of the video. For each frame, the codec uses an activation function to convert the feature map into image data within a pixel value range. For example, the Tanh function can clamp the output values ​​to the interval [-1, 1], corresponding to the normalized range of pixel values. For video data, the codec needs to handle the temporal dependencies between frames. Recurrent neural networks or their variants can be used to model time series. At each time step, the codec generates the data for the current frame based on the generated results of the previous frame and the video latent vector, thereby maintaining temporal consistency and coherence of the video.

[0167] The codec's decoding process ensures that weather effects are consistent with the video content. The video latent vector combines the target weather information with the video content features, allowing the decoder to appropriately apply the weather effects to every object and scene in the video when generating the video. For example, in a foggy video, the outlines of objects will be blurred by the fog. The decoder uses the feature information in the latent vector to ensure that this blurring effect is consistent with the object outline.

[0168] Figure 8 The flow chart of a video generation method proposed in the present disclosure is further shown. Figure 8 The embodiment shown further explains the above embodiment. Figure 8 The following steps may be included:

[0169] Step 401: Input the spliced ​​video into the codec for latent vector encoding to obtain the first latent vector.

[0170] First, the spliced ​​video is used as input data and passed into the encoding module of the codec. The video data is first processed through multiple convolutional layers to extract the spatial features in the video frames. Then, the information is further compressed through network structures such as fully connected layers and converted into latent vector form. The encoder converts the spatial and temporal features in the spliced ​​video (such as motion and changes between video frames) into a low-dimensional latent vector. This latent vector not only contains the key content of the video, but also can effectively represent the overall structure of the spliced ​​video. The first latent vector output by the encoder after processing contains the important features of the spliced ​​video and can be restored to the video content in the subsequent decoding process. This latent vector is an efficient representation of the spliced ​​video and can support applications such as video generation, video retrieval, and video analysis.

[0171] By converting the spliced ​​video into latent vectors, the storage and transmission requirements of the video can be effectively reduced. The dimensions of the latent vectors are usually much smaller than the original video, so they can save storage space and reduce bandwidth consumption during network transmission.

[0172] Step 402: Input the target weather information into the codec for latent vector encoding to obtain the second latent vector.

[0173] The target weather information is fed into the encoder / decoder's encoding module as input data. The data first passes through a multi-layered neural network structure, gradually extracting spatial and temporal features from the weather information, such as temperature trends and humidity fluctuations. The encoder uses a series of nonlinear transformations to convert the raw weather data into a low-dimensional representation, known as the second latent vector. This latent vector not only contains key weather information but also abstracts weather trends and patterns. After processing, the encoder outputs a second latent vector that retains the key features of the target weather information while being compressed into a low-dimensional vector. This latent vector can then be converted into a specific weather forecast, weather analysis results, or used in other weather applications during the subsequent decoding process.

[0174] By encoding the target weather information into a latent vector, we can effectively capture the complex characteristics and time series patterns in weather data. This second latent vector, as a low-dimensional representation, can be used for further weather forecasting, trend analysis, and anomaly detection, improving the accuracy and reliability of weather forecasts.

[0175] Step 403: Input the video to be processed into the codec for latent vector encoding to obtain the third latent vector.

[0176] Multiple frames of the video to be processed are fed frame by frame through the codec's input module. During this input process, the features of each pixel or local area of ​​the video frame are extracted through multiple layers of a neural network to construct a deep representation of the video content. The encoder converts the video frame data into a low-dimensional representation through a series of nonlinear transformations. This representation is the third latent vector. The third latent vector incorporates the spatiotemporal features of the video, such as scene changes, action patterns, and correlations between audio and video. Through compression processing, the latent vector can effectively represent key information in the video and support subsequent tasks such as video analysis, behavior recognition, and content generation.

[0177] By converting the video to be processed into latent vectors, the storage and transmission requirements of video data can be significantly reduced. The low-dimensional representation of latent vectors effectively reduces the complexity of computation and communication, making it particularly suitable for applications that require real-time processing of large amounts of video data, such as intelligent surveillance and video streaming.

[0178] Figure 9 The flow chart of a video generation method proposed in the present disclosure is further shown. Figure 9 The embodiment shown further explains the above embodiment. Figure 9 The following steps may be included:

[0179] Step 501: Perform noise processing on the third latent vector to obtain a noisy third latent vector.

[0180] Noising involves introducing a certain amount of noise into the latent vector. The goal is to improve the model's robustness and enhance its resistance to uncertainty and interference by perturbing the original data. Noising is typically achieved by adding random noise or other methods. The noisy third latent vector is the new latent vector obtained by adding noise to the original third latent vector. The noisy third latent vector often contains key information from the original latent vector.

[0181] Determine the noise processing parameters, such as noise intensity, distribution type, and addition method, based on actual needs and application scenarios. The noise intensity determines the degree of difference between the noisy third latent vector and the original third latent vector; the distribution type determines the statistical properties of the noise; and the addition method determines how the noise is fused with the third latent vector. Generate the corresponding noise data based on the selected noise algorithm and the determined noise parameters. For example, if the Gaussian noise algorithm is selected, a random number generation library can be used to generate Gaussian random numbers with a specified mean and variance as noise data. The generated noise data is fused with the third latent vector to obtain the noisy third latent vector. Common addition methods include element-wise addition and weighted fusion. Element-wise addition directly adds the noise data to the corresponding elements of the third latent vector. Weighted fusion linearly combines the noise data with the third latent vector based on a predefined weight coefficient. The weight coefficient determines the relative contribution of the noise and original latent vectors to the noisy latent vector.

[0182] Noising can improve the generalization capabilities of video generation models, enabling them to better adapt to diverse data distributions and scene variations. During training, the noisy third latent vector provides the model with more variation and uncertainty, forcing it to learn more general and adaptable feature representations. This helps the model more accurately generate plausible videos when faced with new, unseen video data, without overly relying on the specific features and patterns of the training data. For example, a model trained with noise can more flexibly adjust its generation strategy when generating videos under different weather conditions, producing videos that are more consistent with real-world scenarios.

[0183] Step 502: Input the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector, thereby obtaining a video latent vector to be processed.

[0184] The input layer of the video generation network receives the first latent vector, the second latent vector, and the third latent vector after noise addition. The middle layer of the network extracts, transforms, and integrates the input latent vectors through the adopted network structure (such as fully connected layers, convolutional layers, or recurrent layers). In this process, the network learns the feature associations and fusion patterns between different latent vectors, extracting information useful for generating the latent vector of the video to be processed. After multiple layers of feature extraction and integration, the video generation network maps the input latent vector to a new feature space to generate the latent vector of the video to be processed. This latent vector contains the depth, line drawing, and contour features of the spliced ​​video, the weather features of the target weather information, and the video content features of the video to be processed after noise addition. It integrates all key information to guide the subsequent video generation process.

[0185] By combining the first latent vector, the second latent vector, and the noised third latent vector, the video generation network can obtain a more complete and accurate latent vector for the video being processed. This multi-dimensional feature representation enhances information transfer during the video generation process, making the generated video more realistic, higher quality, and richer in detail.

[0186] Step 503: Call a preset denoising algorithm to perform denoising on the latent vector of the video to be processed to obtain the latent vector of the video.

[0187] A preset denoising algorithm is an algorithm pre-designed and configured in the system to remove noise components from video latent vectors. The preset denoising algorithm can effectively identify and filter out noise information in the latent vector, improving the purity and accuracy of the latent vector. Common denoising algorithms include convolutional neural network-based denoising networks, autoencoder denoising, non-local mean filtering, wavelet transform, and other methods. The processed video latent vector refers to the latent vector output by the previous video generation network. The processed video latent vector contains noise components. The noise comes from the third latent vector after adding noise to the input latent vector and the uncertainty of the generation network itself. The video latent vector is a purer and more accurate latent vector obtained after processing by the preset denoising algorithm. It represents the key information and features of the video and is suitable for subsequent video generation, reconstruction, or analysis tasks.

[0188] Determine the parameters of the denoising algorithm, such as the filter window size, wavelet basis function, and threshold, based on actual needs and data characteristics. For example, when using Gaussian filtering, it is necessary to determine the filter window size and the standard deviation of the Gaussian kernel. When using wavelet transform denoising, it is necessary to select an appropriate wavelet basis function and number of decomposition layers, as well as determine the thresholding method and threshold value for the wavelet coefficients. Denoise the latent vector of the video to be processed based on the selected denoising algorithm and the determined parameters. For example, if an autoencoder denoising algorithm is selected, the latent vector of the video to be processed is input into a trained autoencoder. The encoder portion of the autoencoder maps the noisy latent vector to a low-dimensional latent space, and the decoder portion restores the representation of the low-dimensional latent space to the denoised video latent vector. Evaluate the denoised video latent vector to check whether the denoising effect meets the requirements. The denoising effect can be evaluated by calculating error metrics before and after denoising (such as mean square error, signal-to-noise ratio, etc.), visualizing the latent vectors before and after denoising (such as plotting the distribution of latent vectors or heat maps), or evaluating the performance of downstream tasks (such as using the denoised latent vectors for video generation and observing the quality of the generated video). If the denoising effect is not good, you can adjust the parameters of the denoising algorithm or choose another denoising algorithm for processing.

[0189] Removing noise from the latent vector of the processed video can improve the quality of the generated video. The denoised video latent vector is purer and contains more accurate and clear video feature information, guiding the decoder to generate a higher-quality target video. For example, when generating a video of a rainy scene, the denoised video latent vector can make the generated raindrop distribution more natural and the water accumulation more realistic, reducing blur, distortion, and anomalies caused by noise, thereby improving the realism and visual quality of the generated video.

[0190] As a refinement of step 501, when performing the noise processing on the third latent vector to obtain the noisy third latent vector, it can be implemented in, but not limited to, the following manner, including: calling a preset noising algorithm to perform noise processing on the third latent vector to obtain the noisy third latent vector; and / or using a preset normal distribution algorithm to generate the noisy third latent vector of the third latent vector.

[0191] A preset noise addition algorithm is a pre-set algorithm used to introduce noise into the data. Common preset noise addition algorithms include Gaussian noise addition, uniform noise addition, and salt and pepper noise addition. These algorithms simulate interference factors in real-world scenarios by adding random noise of a specific distribution to the data, thereby enhancing the robustness and generalization ability of the model. The normal distribution algorithm is an algorithm that generates random numbers or random vectors based on the normal distribution (also known as the Gaussian distribution). The normal distribution is a continuous probability distribution characterized by data being concentrated around the mean and symmetrically distributed. In noise addition processing, the normal distribution algorithm can generate noise data based on the set mean and standard deviation, and add it to the original data to simulate random phenomena and uncertainty in nature.

[0192] The system selects and calls the preset noise adding algorithm to perform noise adding processing on the third latent vector according to the demand.

[0193] For example, when applying Gaussian noise, the system can add a noise value that conforms to a normal distribution to each element of the third latent vector according to the set noise variance, obtaining the noisy third latent vector. This process can be achieved through the following steps: a. Select the noise addition algorithm type (such as Gaussian noise); b. Add a noise value generated by the noise model to each element of the third latent vector; c. Output the noisy third latent vector for subsequent model processing. If normally distributed noise is used, the system generates noise using a preset normal distribution algorithm. The specific steps are as follows: a. Generate noise that conforms to a standard normal distribution for each element of the third latent vector, with a mean of 0 and a standard deviation adjusted according to the set value; b. Add the generated noise to each element of the third latent vector to obtain the noisy third latent vector; c. Output the processed noisy latent vector for further training or testing.

[0194] Noising can artificially introduce noise perturbations, preventing the model from over-relying on training data during training, thereby reducing the risk of overfitting. The noisy latent vector provides the model with diverse training data, helping it generalize better.

[0195] As a refinement of step 502, when inputting the first latent vector, the second latent vector and the noisy third latent vector into the video generation network to generate a video latent vector to obtain a latent vector for the video to be processed, it can be implemented in but not limited to the following manner, including: iteratively processing the latent vector generated by the video generation network in a direction that makes the weather latent vector in the latent vector for the video to be processed tend toward the second latent vector until the number of iterations reaches a preset number to obtain the latent vector for the video to be processed.

[0196] The weather latent vector is a sub-component of the latent vector of the video to be processed, specifically a latent vector representing weather characteristics. The weather latent vector contains the characteristics of the target weather information, such as weather type and intensity, and is used to guide the generation of weather effects in the video. Iterative processing is a process of gradually optimizing and improving data or model output. In video generation, by performing multiple iterative processes on the latent vector generated by the video generation network, the weather latent vector in the latent vector is gradually adjusted and optimized, tending towards the direction of the second latent vector, thereby improving the quality and accuracy of the generated video. The preset number of iterations refers to the upper limit of the number of iterations pre-set during the iterative processing process. When the number of iterations reaches the preset number, the iterative processing process stops and the final latent vector of the video to be processed is output.

[0197] The latent vector of the video to be processed is iteratively processed. In each iteration, the latent vector is adjusted according to the target direction so that it gradually converges to the direction of the second latent vector. The specific iterative steps are as follows: obtain the current latent vector of the video to be processed and the weather latent vector; calculate the distance or difference between the weather latent vector and the second latent vector; adjust the weather latent vector in the latent vector of the video to be processed according to the difference so that it gradually converges to the second latent vector; repeat the above process until the number of iterations reaches the preset number; if the weather latent vector is close enough to the second latent vector when the preset number is reached, the iteration is terminated.

[0198] By iteratively aligning the weather latent vector with the second latent vector, the generated video's weather authenticity and consistency can be improved. The resulting weather effects will better align with the target weather characteristics and remain consistent with the target weather information, avoiding unrealistic or inconsistent weather effects. For example, when generating a video of a rainy scene, the weather latent vector will align with the second latent vector representing rainy weather, resulting in realistic raindrops, puddles, and other effects in the generated video, enhancing the video's realism.

[0199] In practical applications, the video generation model needs to be trained before it can be used. The training method of the video generation model can be implemented by, but is not limited to, the following methods, including: obtaining a spliced ​​video for training, target weather information for training, and a video to be processed for training; training the initial video generation model based on the spliced ​​video for training, the target weather information for training, and the video to be processed for training to obtain the trained video generation model.

[0200] Training stitched videos refer to the stitched video data used for model training. A stitched video is a video created by stitching together the depth map, line drawing, and contour map videos of the target video according to the video dimension channels. The depth map video reflects the depth information of objects in the video scene, the line drawing video depicts the object's outline and shape structure, and the contour map video displays the object's edge contour information. The stitched training video provides the model with rich multi-dimensional feature information for learning video features such as depth, shape, and contour. Training target weather information refers to the target weather information data used for model training. This data describes the specific weather conditions that are expected to appear in the generated video. It can be represented by a specific numerical value, vector, or encoded form and includes weather type (such as sunny, rainy, snowy, foggy, etc.), intensity (such as rain weight, fog density, etc.), and other relevant parameters (such as wind speed and visibility). The target weather information for training provides the model with supervision signals for weather characteristics, allowing it to learn video features under different weather conditions. Training target videos refer to the original video footage used for model training. This can be a video of a vehicle driving under normal weather conditions, for example, and includes various objects and their motion in the road scene, such as vehicles, pedestrians, and the relative motion of objects. The training video provides the model with raw video content and dynamic information, serving as the basis for generating the target video. The initial video generation model refers to an untrained video generation model. It typically consists of multiple neural network components, such as codecs and video generation networks, and is used to convert the input spliced ​​video, target weather information, and the processed video into the target video. The parameters of the initial video generation model are randomly initialized and need to be optimized through the training process.

[0201] Define an appropriate loss function based on the characteristics and requirements of the video generation task. Common loss functions include mean squared error (MSE), structural similarity loss, perceptual loss, and adversarial loss. MSE measures the pixel-wise difference between the generated video and the ground-truth target video; structural similarity loss assesses the structural similarity between the generated video and the ground-truth video; perceptual loss, based on features extracted by a pre-trained deep convolutional network, measures the difference in high-level semantic features between the generated and ground-truth videos; and adversarial loss, by introducing the adversarial training mechanism of a generative adversarial network, brings the distribution of generated videos closer to that of ground-truth videos, improving their realism and diversity. For example, MSE and perceptual loss can be combined to comprehensively consider pixel-level and feature-level differences, improving the quality of generated videos. Select an appropriate optimization algorithm to minimize the loss function. Common optimization algorithms include stochastic gradient descent (SGD), which is widely used due to its adaptability and fast convergence. Optimization algorithms calculate the gradient of the loss function with respect to model parameters and update them to minimize the difference between the generated and ground-truth videos. The training set data is fed into the model in batches. A forward propagation is performed to generate the video, calculate the loss function, and then a backpropagation is performed to calculate the gradients and update the model parameters. This process is repeated until the preset number of training iterations is reached or the model's performance on the validation set stops improving. During training, the model's performance can be regularly evaluated on the validation set. Based on the results, the model's hyperparameters, such as the learning rate and batch size, can be adjusted to optimize the model's training results. For example, a learning rate of 0.001 and a batch size of 32 can be set. Multiple iterations are performed after each pass through the training set until the model's performance metrics on the validation set (such as mean squared error and structural similarity) reach satisfactory levels.

[0202] By training on a large number of stitched training videos, target weather information, and pre-processed training videos, the trained video generation model accurately captures the complex relationships between the multi-dimensional features of videos and weather characteristics, improving the performance and accuracy of generated videos. The generated videos are more realistic and natural in terms of object appearance, motion trajectories, and weather effects, maintaining a high degree of consistency with the target weather information and video content, and can better meet the video quality requirements of practical applications such as autonomous driving systems.

[0203] In order to better understand the training process of the video generation model, Figure 10 As shown, Figure 10This is a flowchart of the training process of a video generation model provided by an embodiment of the present disclosure. Leveraging the powerful capabilities of diffusion models in the field of video generation, the video generation model uses DiT. First, the required multiple control conditions and masks for different regions are extracted from the original video. Following the description in 1 above, control conditions are adaptively selected based on different regions and fused. Finally, the prompt text containing weather condition prompts, the fused control conditions, and the noise-processed video are input into the DiT network to predict the output video.

[0204] Specifically, the video generation model consists of three components: a 3D Variational Autoencoder (3DVAE), a 3D Diffusion Probabilistic Model (DDP), and a text encoding model. During training, only the weights of the 3D DiT module are adjusted, while the weights of the other modules are fixed. The training process is as follows: 1) The fused control map is fed into the 3D VAE in the form of a video to obtain its latent space vector c_latent; 2) The sentence describing the weather in the video is fed into the text encoder model, encoding a latent vector prompt_embeds that differs from the video latent vector latent only in channel size, serving as the language control condition vector; 3) The original video is fed into the 3D VAE to obtain the latent vector latent, which is then denoised using a denoising diffusion probabilistic model (DDPM) to obtain noise_latent; 4) The control map latent vector c_latent, the weather text latent vector prompt_embeds, and noise_latent are fed into the DDP module to predict the noise pred_noise on noise_latent. 5) Use pred_noise and noise as the mean square error loss mse_loss, backpropagate the gradient, and use gradient descent to update the weights of 3D DiT.

[0205] In the inference phase, the whole process is similar to Figure 5The process is similar, except that the initial input model's noise_latent is not obtained by adding noise to the video's latent vector latent, but is instead generated directly from a standard normal distribution. The DiT model predicts the noise and gradually removes it using Denoising Diffusion Implicit Models (DDIM), ultimately generating a video with weather effects consistent with the text description of the weather. The specific process is as follows: 1) Generate initial noise: We use a standard normal distribution to generate the initial noise_latent; 2) Prepare control conditions: In this step, we encode the textual information describing the target weather state into a latent vector called prompt_embeds. Furthermore, consistent with the above, we extract and adaptively fuse the control map from the original video and feed it into the 3D VAE module to obtain the control map latent vector c_latent; 3) Input to the DiT model: The generated noise_latent, prompt_embeds, and c_latent are all input into the 3D DiT model. The task of the model is to gradually improve the image quality by predicting noise; 4) Noise prediction and gradual denoising: The DiT model first analyzes the characteristics of noise_latent under given control conditions and predicts the corresponding noise pred_noise. Then, the model uses the DDIM (Denoising Diffusion Implicit Models) algorithm to gradually denoise; 5) Generate the final output: In each round of iteration, the denoised result gradually approaches the goal of meeting the control conditions. When the iteration reaches the set number of rounds, the DiT model outputs the final video latent vector and sends it to the 3D VAE decoder to obtain the final video frame. This video frame shows the weather effect that matches the input text description, successfully transforming the scene under normal weather conditions into the desired severe weather conditions.

[0206] Autonomous driving technology relies heavily on data-driven deep neural networks. However, data in the autonomous driving field often exhibits a long-tail distribution, which means that critical driving data in adverse conditions is relatively scarce, and collecting this data presents a significant challenge. Our proposed solution can convert videos in normal weather conditions into videos in adverse weather conditions, thereby generating a large amount of long-tail data that can be used for autonomous driving model training and evaluation. This approach effectively compensates for the lack of real-world data for certain special scenarios.

[0207] By leveraging this simulation data, we can significantly improve the autonomous driving system's ability to identify complex and dangerous situations. For example, the system can better cope with driving challenges in extreme weather conditions such as heavy rain, fog, and snow. This not only helps improve the accuracy and reliability of the model but also has a significant impact on improving the overall robustness of the system, laying the foundation for safe and efficient autonomous driving.

[0208] Corresponding to the above-mentioned video generation method, the present invention also provides a video generation device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment and will not be repeated in this invention.

[0209] Figure 11 This is a schematic structural diagram of a video generation device 600 provided in an embodiment of the present disclosure, the video generation device comprising:

[0210] An acquisition unit 61 is configured to acquire a depth map video, a line drawing video, and a contour map video of a video to be processed, and to acquire target weather information required for the video to be processed;

[0211] a stitching unit 62 for stitching the depth map video, the line drawing video, and the outline map video according to video dimension channels to obtain a stitched video;

[0212] The generating unit 63 is configured to generate a target video according to the spliced ​​video and the target weather information; the target video includes video features of the target weather information.

[0213] In summary, according to the video generation device proposed in the present disclosure, the device includes obtaining a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtaining target weather information required for the video to be processed; splicing the depth map video, the line drawing video, and the contour map video according to the video dimension channel to obtain a spliced ​​video; generating a target video based on the spliced ​​video and the target weather information; the target video contains the video features of the target weather information. The solution disclosed in the present disclosure is based on video for processing, which meets the requirements of autonomous driving scenarios for time-series videos, and the depth map video provides the spatial structure of the scene, the line drawing video clarifies the edges of objects, and the contour map video defines the main shape. The three jointly constrain can improve physical rationality, and combined with the target weather information, the intensity of the weather effect is adjusted, reducing the contradiction between the intensity of the weather effect of the video and the physical rationality of the objects and their movements in the video.

[0214] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the splicing unit 62 includes:

[0215] An acquisition module 621 is configured to acquire a first mask for the line drawing video and a second mask for the outline drawing video; the first mask is used to extract traffic elements in the line drawing video, and the second mask is used to clear sky elements in the outline drawing video;

[0216] The stitching module 622 is configured to stitch the depth map video, the line drawing video, and the contour map video of the same video dimension channel together according to the first mask and the second mask to obtain the stitched video.

[0217] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the splicing module 622 is further used for:

[0218] Performing a product calculation on the first mask and the line drawing video to obtain a line drawing video after the product calculation;

[0219] Performing a product calculation on the second mask and the contour map video to obtain a contour map video after the product calculation;

[0220] The depth map video, the line drawing video after the product calculation, and the contour map video after the product calculation of the same video dimension channel are spliced ​​together to obtain the spliced ​​video.

[0221] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the generating unit 63 includes:

[0222] The generation module 631 is used to input the spliced ​​video, the target weather information and the video to be processed into a trained video generation model to generate a video, so as to obtain the target video.

[0223] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the video generation model includes a codec and a video generation network, and the generation module 631 is further used to:

[0224] Using the codec, respectively determining a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed;

[0225] Using the video generation network, performing video latent vector generation on the first latent vector, the second latent vector, and the third latent vector to obtain a video latent vector;

[0226] The codec is used to perform latent vector decoding processing on the video latent vector to obtain the target video.

[0227] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the generating module 631 is further used for:

[0228] Inputting the spliced ​​video into the codec for latent vector encoding processing to obtain the first latent vector;

[0229] Inputting the target weather information into the codec for latent vector encoding processing to obtain the second latent vector;

[0230] The video to be processed is input into the codec for latent vector encoding processing to obtain the third latent vector.

[0231] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the generating module 631 is further used for:

[0232] performing noise processing on the third latent vector to obtain a noisy third latent vector;

[0233] Inputting the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector, thereby obtaining a video latent vector to be processed;

[0234] Calling a preset denoising algorithm to perform denoising on the latent vector of the video to be processed to obtain the latent vector of the video.

[0235] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the generating module 631 is further used for:

[0236] Calling a preset noising algorithm to perform noise processing on the third latent vector to obtain the noisy third latent vector; and / or,

[0237] The noisy third latent vector of the third latent vector is generated by using a preset normal distribution algorithm.

[0238] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the generating module 631 is further used for:

[0239] In accordance with the direction in which the weather latent vector in the latent vector of the video to be processed tends toward the second latent vector, the latent vector generated by the video generation network is iteratively processed until the number of iterations reaches a preset number, thereby obtaining the latent vector of the video to be processed.

[0240] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 12 As shown, the device also includes:

[0241] The acquisition unit 61 is further used to acquire the spliced ​​video for training, the target weather information for training, and the video to be processed for training;

[0242] The training unit 64 is configured to train the initial video generation model based on the training spliced ​​video, the training target weather information, and the training video to be processed, to obtain the trained video generation model.

[0243] Since the device provided in the embodiment of the present disclosure corresponds to the methods provided in the above embodiments, the implementation of the method is also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.

[0244] In the embodiments provided above, the methods and devices provided in the embodiments of the present application are introduced. In order to implement the various functions of the methods provided in the embodiments of the present application, the electronic device may include a hardware structure and a software module, and implement the aforementioned functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. One of the aforementioned functions may be executed in the form of a hardware structure, a software module, or a hardware structure plus a software module.

[0245] Figure 13 FIG2 is a block diagram of an electronic device 700 for implementing the above-described video generation method according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0246] Reference Figure 13 , electronic device 700 may include one or more of the following components: a processing component 702 , a memory 704 , a power component 706 , a multimedia component 708 , an audio component 710 , an input / output (I / O) interface 712 , a sensor component 714 , and a communication component 716 .

[0247] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 702 may include one or more modules to facilitate interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate interaction between the multimedia component 708 and the processing component 702.

[0248] The memory 704 is configured to store various types of data to support operations on the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0249] The power supply component 706 provides power to the various components of the electronic device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 700.

[0250] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0251] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 also includes a speaker for outputting audio signals.

[0252] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0253] The sensor assembly 714 includes one or more sensors for providing various aspects of status assessment for the electronic device 700. For example, the sensor assembly 714 can detect the open / closed state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor assembly 714 can also detect changes in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and temperature changes of the electronic device 700. The sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 714 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0254] The communication component 716 is configured to facilitate wired or wireless communication between the electronic device 700 and other devices. The electronic device 700 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0255] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0256] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by the processor 720 of the electronic device 700 to perform the above method for video generation. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0257] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the above embodiments of the present disclosure.

[0258] In order to implement the above embodiments, the present disclosure further proposes a chip, including: the chip includes a processing circuit, and the processing circuit is configured to execute the method provided in the above embodiments.

[0259] Figure 14 This is a schematic diagram of the structure of a chip proposed in an embodiment of the present disclosure. Figure 14 The structure of the chip 800 is shown, but is not limited thereto.

[0260] The chip 800 includes a processing circuit 801 and an interface circuit 802 . The interface circuit 802 is used to read instructions and send the instructions to the processing circuit 801 so that the processing circuit 801 executes the above method.

[0261] Alternatively, as Figure 15 As shown, Figure 15 The chip 800 may further include a memory 803 for storing instructions, and the interface circuit 802 may be used to read the instructions stored in the memory 803 .

[0262] Optionally, the interface circuit 802 is connected to the memory 803. The interface circuit 802 can be used to receive signals from the memory 803 or other devices, and can be used to send signals to the memory 803 or other devices. For example, the interface circuit 802 can read instructions stored in the memory 803 and send the instructions to the processing circuit 801.

[0263] Optionally, the number of memories 803 can be one or more, and the number of interface circuits 802 can also be one or more.

[0264] In some embodiments, the interface circuit 802 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 801 performs the other steps.

[0265] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.

[0266] Optionally, all or part of the memory 803 may also be located outside the chip 800 .

[0267] Those skilled in the art will also appreciate that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present application.

[0268] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0269] Throughout this specification, references to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" indicate that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0270] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0271] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having one or more wires (control method), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0272] It should be understood that various parts of the embodiments of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0273] Those skilled in the art will understand that all or part of the steps of the method for implementing the above-mentioned embodiment can be completed by instructing the relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0274] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in either hardware or software functional modules. If the integrated modules are implemented as software functional modules and sold or used as standalone products, they may also be stored in a computer-readable storage medium. The aforementioned storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0275] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A video generation method, characterized in that: The method comprises: Obtaining a depth map video, a line drawing video, and a contour map video of a video to be processed, and obtaining target weather information required for the video to be processed; splicing the depth map video, the line drawing video, and the outline map video according to the video dimension channel to obtain a spliced ​​video; A target video is generated according to the spliced ​​video and the target weather information; the target video includes video features of the target weather information.

2. The method according to claim 1, characterized in that The step of splicing the depth map video, the line drawing video, and the outline map video according to the video dimension channel to obtain a spliced ​​video includes: Obtaining a first mask for the line drawing video and a second mask for the outline drawing video; the first mask is used to extract traffic elements in the line drawing video, and the second mask is used to clear sky elements in the outline drawing video; According to the first mask and the second mask, the depth map video, the line drawing video, and the outline map video of the same video dimension channel are spliced ​​to obtain the spliced ​​video.

3. The method according to claim 2, characterized in that The step of stitching the depth map video, the line drawing video, and the outline map video of the same video dimension channel according to the first mask and the second mask to obtain the stitched video includes: Performing a product calculation on the first mask and the line drawing video to obtain a line drawing video after the product calculation; Performing a product calculation on the second mask and the contour map video to obtain a contour map video after the product calculation; The depth map video, the line drawing video after the product calculation, and the contour map video after the product calculation of the same video dimension channel are spliced ​​together to obtain the spliced ​​video.

4. The method according to claim 1, wherein Generating a target video according to the spliced ​​video and the target weather information includes: The spliced ​​video, the target weather information and the video to be processed are input into a trained video generation model to generate a video to obtain the target video.

5. The method according to claim 4, characterized in that The video generation model includes a codec and a video generation network. The spliced ​​video, the target weather information, and the to-be-processed video are input into the trained video generation model to generate the target video. Using the codec, respectively determining a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed; Using the video generation network, performing video latent vector generation on the first latent vector, the second latent vector, and the third latent vector to obtain a video latent vector; The codec is used to perform latent vector decoding processing on the video latent vector to obtain the target video.

6. The method according to claim 5, characterized in that The step of using the codec to respectively determine a first latent vector of the spliced ​​video, a second latent vector of the target weather information, and a third latent vector of the video to be processed includes: Inputting the spliced ​​video into the codec for latent vector encoding processing to obtain the first latent vector; Inputting the target weather information into the codec for latent vector encoding processing to obtain the second latent vector; The video to be processed is input into the codec for latent vector encoding processing to obtain the third latent vector.

7. The method according to claim 5, characterized in that The step of generating a video latent vector from the first latent vector, the second latent vector, and the third latent vector using the video generation network to obtain the video latent vector includes: performing noise processing on the third latent vector to obtain a noisy third latent vector; Inputting the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector, thereby obtaining a video latent vector to be processed; Calling a preset denoising algorithm to perform denoising on the latent vector of the video to be processed to obtain the latent vector of the video.

8. The method according to claim 7, characterized in that The performing noise processing on the third latent vector to obtain the noisy third latent vector includes: Calling a preset noising algorithm to perform noise processing on the third latent vector to obtain the noisy third latent vector; and / or, The noisy third latent vector of the third latent vector is generated by using a preset normal distribution algorithm.

9. The method according to claim 7, characterized in that Inputting the first latent vector, the second latent vector, and the noise-added third latent vector into the video generation network to generate a video latent vector to obtain the to-be-processed video latent vector includes: In accordance with the direction in which the weather latent vector in the latent vector of the video to be processed tends toward the second latent vector, the latent vector generated by the video generation network is iteratively processed until the number of iterations reaches a preset number, thereby obtaining the latent vector of the video to be processed.

10. The method according to claim 4, characterized in that The method further comprises: Obtaining training stitched videos, training target weather information, and training videos to be processed; The initial video generation model is trained based on the training spliced ​​video, the training target weather information and the training video to be processed to obtain the trained video generation model.

11. A video generating device, characterized in that: The device comprises: An acquisition unit, configured to acquire a depth map video, a line drawing video, and a contour map video of a video to be processed, and to acquire target weather information required for the video to be processed; a splicing unit, configured to splice the depth map video, the line drawing video, and the outline map video according to video dimension channels to obtain a spliced ​​video; A generating unit is configured to generate a target video based on the spliced ​​video and the target weather information; the target video includes video features of the target weather information.

12. The device according to claim 11, characterized in that The splicing unit includes: an acquisition module, configured to acquire a first mask for the line drawing video and a second mask for the outline drawing video; the first mask is used to extract traffic elements in the line drawing video, and the second mask is used to clear sky elements in the outline drawing video; A stitching module is used to stitch the depth map video, the line drawing video and the contour map video of the same video dimension channel according to the first mask and the second mask to obtain the stitched video.

13. The device according to claim 12, characterized in that The splicing module is also used for: Performing a product calculation on the first mask and the line drawing video to obtain a line drawing video after the product calculation; Performing a product calculation on the second mask and the contour map video to obtain a contour map video after the product calculation; The depth map video, the line drawing video after the product calculation, and the contour map video after the product calculation of the same video dimension channel are spliced ​​together to obtain the spliced ​​video.

14. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

15. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

16. A chip, characterized in that: The chip includes a processing circuit and an interface circuit; wherein the interface circuit is used to read instructions, and the interface circuit sends the instructions to the processing circuit, so that the processing circuit executes the method according to any one of claims 1 to 10.