DiT condition guidance-based multi-label weather image data expansion method
By adopting DiT conditional guidance method and LDM technology in the generation of multi-label weather images, combined with multi-label condition control, data scarcity and labeling complexity problems are solved, and high-quality multi-label weather images are generated, which significantly improves the performance of the multi-label weather classification model.
Patent Information
- Application Number
- CN202510212406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing multi-label weather classification technology faces data scarcity and labeling complexity. Traditional data augmentation technology cannot effectively generate diverse multi-label weather images, and the existing diffusion model has shortcomings in multi-label control.
A conditional guidance method based on Diffusion Transformers (DiT) is adopted, combining the latent diffusion model (LDM) and multi-label condition control mechanism, high-quality multi-label weather images are generated, and the generalization ability and robustness of the model are improved through a hybrid data-driven learning strategy.
The diversity and quality of multi-label weather images are significantly improved, the scale and diversity of the data set are enhanced, and the performance and accuracy of the multi-label weather classification model are improved.
Smart Images

Figure CN120219873A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and further relates to data augmentation technology. Specifically, it is a multi-label weather image data augmentation method guided by Diffusion Transformers (DiT for short), which can be used for multi-label weather classification tasks in deep learning. Background Technique
[0002] With the rapid development of deep learning technology, multi-label weather classification shows important application value in fields such as meteorological forecasting, intelligent transportation, and autonomous driving. Different from traditional single-label classification, multi-label weather classification can assign multiple weather labels (such as sunny, rainy, foggy, etc.) to each image, thus more comprehensively describing complex weather scenarios. However, the current multi-label weather classification faces the following key problems: First, training a deep learning model requires a large-scale high-quality dataset, while the existing multi-label weather datasets are seriously insufficient in scale. Taking the Multi-label Weather Dataset as an example, this dataset only contains about 10,000 labeled images, covering five weather types (cloudy, sunny, rainy, snowy, foggy). Such a limited amount of data is difficult to support the model to learn the diversity and complex features of weather scenarios. Second, the annotation process of weather images is complex and costly. Due to the variability and non-exclusivity of weather phenomena (for example, rainy and foggy weather may occur simultaneously in the same scene), annotation requires professional personnel to make detailed judgments on the intensity of each label, which is not only time-consuming and laborious but also prone to errors.
[0003] To solve the problem of data scarcity, researchers have proposed a variety of data augmentation techniques. Traditional methods such as rotation, flipping, cropping, scaling, and brightness adjustment can, to a certain extent, expand the dataset, but the generated samples have limited changes and cannot create new weather patterns, especially in multi-label scenarios, it is difficult to maintain the complex correlation between labels. In addition, generative adversarial networks (GANs) have been widely used in data generation. For example, Weather GAN generates images under specific weather conditions through image style transfer. However, GAN training is unstable, prone to mode collapse, the generated samples have insufficient diversity, and it has limited ability in multi-label control, making it difficult to precisely adjust the combination and intensity of multiple weather phenomena. Variational autoencoders (VAEs) are stable in training, but the generated image quality is low and difficult to meet the requirements of high-resolution weather images.
[0004] In recent years, diffusion models have received attention for their ability to generate high-quality images. Diffusion models based on the U-Net architecture generate realistic images through a step-by-step noise addition and denoising process and have made some progress in weather image enhancement. For example, Zhu et al. used diffusion models to restore images under adverse weather conditions, and Cheng et al. used diffusion models to perceive the data distribution under different weather conditions to expand the dataset. However, the inductive bias of the U-Net architecture limits its performance, and existing diffusion models mostly focus on single-label generation and lack effective support for multi-label weather scenarios. Diffusion Transformers (DiT) proposed by Peebles and Xie replaced the U-Net by introducing the Transformer architecture and significantly improved the generation quality using its long-range dependence modeling ability, providing new possibilities for multi-label weather image generation.
[0005] Existing patented technologies have also tried to solve similar problems. For example, the "Hybrid Degraded Image Restoration Method Based on Joint Conditional Diffusion Model" (CN202410827420.0) of Beijing Institute of Technology focuses on image restoration rather than data augmentation, which is different from the objective and application scenario of the present invention. The "Two-Stage Rain and Fog Image Restoration Method Based on Conditional Diffusion Model" (CN202410835030.8) of China University of Petroleum uses a diffusion model but only focuses on the restoration of specific weather and lacks multi-label generality. These solutions have not fully explored the potential of multi-label weather image data augmentation. Summary of the Invention
[0006] The objective of the present invention is to propose a multi-label weather image data augmentation method based on DiT conditional guidance in view of the above deficiencies of the existing technologies. The method aims to generate high-quality multi-label weather images through the Latent Diffusion Models (LDM) and DiT technologies, combined with a multi-label conditional control mechanism, expand the dataset, and improve the performance of the multi-label weather classification model. The present invention first efficiently generates diverse weather images in the latent space through LDM to alleviate the problem of data scarcity; then uses the multi-label conditional control mechanism of DiT to ensure that the generated images highly match multiple weather labels; finally, through a hybrid data-driven learning strategy, combines the generated pseudo-data with the original data to train the classification model, improving the generalization ability and robustness of the model.
[0007] To achieve the above objective, the technical solution of the present invention is as follows:
[0008] A multi-label weather image data augmentation method based on DiT conditional guidance, characterized by comprising the following steps:
[0009] 1) Obtain multi-label weather image data from the public dataset, and use half of it with the original image order shuffled as the training set;
[0010] 2) Construct a diffusion model DiTc based on DiT conditional guidance. Inside this model, through a conditional fusion mechanism, embed the multi-label condition y label and the conditional embedding y t of the diffusion time step t to generate comprehensive conditional information for guiding the model output;
[0011] 3) Train the model DiTc:
[0012] (3.1) Initialize the encoder and decoder parameters of the variational autoencoder VAE, initialize the parameters of the diffusion model DiTc, set the maximum time step in the training phase as T, and initialize the current training time step t = 0;
[0013] (3.2) Use the training set as the model input data, map it to the low-dimensional latent space through the VAE encoder to generate the initial latent representation z0 of the input image, and apply Gaussian noise to z0 to generate the noisy latent representation z t ;
[0014] (3.3) Use the model DiTc to predict the noise ∈ θ :
[0015] ∈ θ = DiTc(z t , t, y label );
[0016] (3.4) Use the hybrid loss function L to optimize the model DiTc, and calculate the gradient through backpropagation Update the parameters of the model DiTc according to this gradient; loop and iterate the update process until the model converges to obtain the optimized diffusion model based on DiT conditional guidance;
[0017] 4) Use the optimized diffusion model to augment the multi-label weather image data, that is, use the multi-label condition as the input of the optimized diffusion model to generate the multi-label synthetic weather image image';
[0018] 5) Mix the image image' generated in step 4) with the training set images in step 1) to obtain an augmented dataset of multi-label weather image data based on DiT conditional guidance for optimizing the training of the multi-label classification model.
[0019] Compared with the prior art, the present invention has the following advantages:
[0020] First, since the present invention uses LDM to generate images in the latent space, compared with the traditional pixel space diffusion model, it significantly reduces the computational cost while maintaining high-quality generation effects, and the generated synthetic images have higher diversity.
[0021] Second, through the multi-label conditional control mechanism of DiT, the present invention uses an independent Embedding layer and adaptive layer normalization adaLN to achieve precise control of various combinations of weather phenomena, ensuring the consistency between the generated images and the labels, and improving the data augmentation effect in multi-label scenarios.
[0022] Third, the present invention adopts a hybrid data-driven learning strategy, combining synthetic data and original data for training, which not only alleviates the problem of data scarcity but also fully utilizes their complementarity, thereby significantly improving the accuracy and robustness of the classification model. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is the overall implementation flowchart of the method of the present invention;
[0024] Figure 2 is the training flowchart of the diffusion model DiTc based on DiT conditional guidance in the present invention;
[0025] Figure 3 is the flowchart for inference using the optimized diffusion model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further described below with reference to the accompanying drawings.
[0027] Example 1. Referring to Figure 1 , a multi-label weather image data augmentation method based on DiT conditional guidance proposed by the present invention includes the following steps:
[0028] Step 1) Obtain multi-label weather image data from a public dataset, and shuffle the original image order of half of them as the training set; specifically, this training set is obtained by acquiring N weather images containing multiple weather categories, and each image contains multiple weather labels. A dataset is obtained by randomly shuffling half of the N images for training a deep learning model; in this embodiment, the weather categories for obtaining weather images include at least cloudy, sunny, rainy, foggy, and snowy, and a public dataset containing other weather types such as lightning and typhoons can also be used for acquisition.
[0029] Step 2) Construct a diffusion model DiTc based on DiT conditional guidance, and embed multi-label conditions y label and the conditional embedding y tIntegrate to generate comprehensive conditional information for guiding the model output. This information is specifically generated according to the following steps:
[0030] (2.1) Embed the multi-label condition y label and the embedding of the diffusion time step t are both used as separate inputs;
[0031] (2.2) Generate the multi-label condition embedding through an independent Embedding layer Generate the embedding of the time step t through the embedding layer
[0032] (2.3) Inside the model DiTc, merge y label and y t through the adaptive layer normalization adaLN to fuse and obtain the comprehensive conditional information Condition merged :
[0033] Condition merged = adaLN(z t , y label + y t ).
[0034] Step 3) Train the model DiTc:
[0035] (3.1) Initialize the parameters of the encoder and decoder of the variational autoencoder VAE, and initialize the parameters of the diffusion model DiTc, which include the conditional network parameters required for the self-attention mechanism, the feed-forward neural network, and the adaptive layer normalization adaLN, etc.; set the maximum time step in the training stage to T, with a value range of 1000 - 2000. In this embodiment, it is preferably set that T = 1500; initialize the current training time step t = 0;
[0036] (3.2) Use the training set as the model input data, map it to the low-dimensional latent space through the VAE encoder to generate the initial latent representation z0 of the input image, and apply Gaussian noise to z0 to generate the noise-added latent representation z t , specifically as follows:
[0037] z0 = Encoder(x),
[0038]
[0039] where x represents the multi-label weather image data in the training set, and α t is a preset noise scheduling parameter used to control the degree of noise addition; represents the random noise of the standard normal distribution.
[0040] (3.3) Use the model DiTc to predict the noise ∈ during the noise addition process θ :
[0041] ∈ θ = DiTc(z t , t, y label );
[0042] (3.4) Use the hybrid loss function L to optimize the model DiTc, and calculate the gradient through backpropagation Update the parameters of the model DiTc according to this gradient; loop and iterate the update process until the model converges to obtain the optimized diffusion model guided by the DiT condition.
[0043] In this embodiment, the hybrid loss function L used in this step has the following expression:
[0044] L = λ simple ||∈ θ - ∈|| 2 + λ vlb L vlb ;
[0045] Among them, λ simple and λ vlb are the weight coefficients for controlling the simple mean square error and the variational lower bound loss respectively, and L vlb represents the variational lower bound loss, which is used to measure the distribution difference between the latent space representation and the target.
[0046] The above update of the parameters of the model DiTc according to the gradient is implemented as follows:
[0047]
[0048] Among them, θ′ represents the set of updated parameters of the model DiTc, θ represents the current set of parameters of the model DiTc; η is the learning rate.
[0049] Step 4) Use the optimized diffusion model to augment the multi-label weather image data, that is, use the multi-label condition as the input of the optimized diffusion model to generate the multi-label synthetic weather image image′, and the implementation steps are as follows:
[0050] (4.1) Initialize the parameters of the encoder and decoder of the VAE, and initialize the parameters of the optimized diffusion model. These model parameters include the conditional network parameters required for the self-attention mechanism, the feed-forward neural network, and the adaptive layer normalization adaLN, etc.; set the maximum number of time steps in the generation stage to T′, and the value range is 500 - 1000. In this embodiment, it is preferably set that T′ = 800; initialize the current generation time step t′ = 0; initialize the latent representation and the multi-label condition;
[0051] (4.2) Gradually add Gaussian noise to the latent representations of the training set images in the latent space to generate a noisy sequence z1, z2, …, z T′ ;
[0052] (4.3) Design independent Embedding layers for each weather category to generate multi-label conditions y′ label ; Generate the embedding y of time step t′ through the embedding layer t′ = Embed(t′), Generate comprehensive conditional information through the adaptive layer normalization adaLN operation inside the optimized diffusion model;
[0053] (4.4) The optimized diffusion model predicts the noise ∈′ label and time step t′ for (z θ , t, y t ), and fuses the conditional information through a merging operation, gradually denoising to the latent representation z′0 at the initial time step, and the expression is as follows: label )
[0054] z t′-1 = z t′ - ∈ θ (z t′ , t′, y label ),
[0055] where z t′-1 is the denoised latent representation at the (t′ - 1)th time step;
[0056] (4.5) Decode the image through the VAE decoder, that is, convert z′0 into a synthetic weather image image′;
[0057] image′ = Decoder(z′0).
[0058] Step 5) Mix the image image′ generated in Step 4) with the training set images in Step 1) to obtain an augmented dataset of multi-label weather image data based on DiT conditional guidance for the optimized training of the multi-label classification model.
[0059] Example 2. The overall implementation steps of the image data augmentation method proposed in this example are the same as those in Example 1. Now, with reference to the attached Figures 1 - 3 give specific examples to further describe the implementation process of the present invention in detail:
[0060] Step 1: Obtain a publicly available multi-label weather image dataset, where the number of images is N and the number of categories is C; select N / 2 images from it, shuffle their original order, and construct a training set. In this embodiment, the Multi-labelWeather Dataset is used, which contains 10,000 images covering 5 weather categories (cloudy, sunny, rainy, snowy, foggy). Select 5,000 images from it and randomly shuffle them to obtain a training set.
[0061] Step 2: Construct a diffusion model DiTc based on DiT conditional guidance and train it:
[0062] Refer to Figure 2 , and perform model training according to the following steps:
[0063] (2a) Randomly initialize model parameters: Initialize the parameters of the encoder and decoder of the variational autoencoder (VAE) to ensure the stability and high-quality generation of the latent space representation. Initialize the DiT model parameters θ, including the parameters of the Transformer blocks (such as self-attention mechanism, feed-forward neural network, and adaptive layer normalization adaLN).
[0064] (2b) Add Gaussian noise: Map the multi-label weather image x in the input training set to the low-dimensional latent space through the VAE encoder to generate an initial latent representation:
[0065] z0 = Encoder(x)
[0066] Subsequently, apply Gaussian noise to z0 to generate the noisy latent representation at the t-th moment:
[0067]
[0068] where, α t is a preset noise scheduling parameter that controls the degree of noise addition, and t is the diffusion time step (from 0 to the maximum number of steps T).
[0069] (2c) Model prediction of noise addition: Use the DiT model to predict the noise during the noise addition process, and the formula is:
[0070] ∈ θ = DiT(z t , t, Condition)
[0071] where, Condition contains multi-label conditional embedding and time step conditional embedding:
[0072] Multi-label conditional embedding: Design an independent Embedding layer for each weather category (such as sunny, rainy, foggy, etc.) to generate a feature vector representing the feature representation of multi-labels.
[0073] Time-step conditional embedding: Design an Embedding layer for each diffusion time step t to generate time feature vectors to guide the model to understand different stages in the diffusion process. The combined conditional vector is generated by adding y {label} and y t to ensure that the DiT model can combine multi-label and time information for conditional guidance.
[0074] (2d) Calculate loss and gradient: Use a hybrid loss function to optimize the DiT model. The loss function includes simple mean squared error and variational lower bound loss, and the formula is:
[0075] L = λ simple ||∈ θ - ∈|| 2 + λ vlb L vlb
[0076] where λ simple and λ vlb are weight coefficients that respectively control the contributions of simple mean squared error and variational lower bound loss; calculate the gradient through backpropagation
[0077] (2e) Training completed: Update the DiT model parameters according to the calculated gradient:
[0078]
[0079] where η is the learning rate. Iterate the training process until the model converges to obtain the optimized DiT model.
[0080] Through the above training process, the DiT model can effectively handle the interaction between multi-label conditional information and image features, learn the mapping from noise to high-quality multi-label weather images, and provide support for subsequent data augmentation.
[0081] Step 3: Data augmentation process
[0082] Refer to Figure 3 and perform data augmentation training according to the following steps:
[0083] (3a) Parameter initialization: Initialize the parameters of the VAE encoder and decoder, the parameters of the DiT model (including Transformer blocks, self-attention mechanisms, adaLN layers, etc.), and the number of diffusion steps T = 1000.
[0084] (3b) Forward diffusion: Apply Gaussian noise to z0 to generate a noisy sequence z t , and the formula is:
[0085]
[0086] Among them, α t is a preset noise scheduling parameter.
[0087] (3c) Conditional embedding: Design an Embedding layer for each weather category to generate feature vectors; add the label vector to the Embedding at time step t to obtain the Condition.
[0088] (3d) Reverse denoising: The DiT model predicts the noise ∈ θ (z t , t, Condition), and gradually denoises to z′0.
[0089] (3e) Image decoding: Convert z′0 into a synthetic image through the VAE decoder.
[0090] Step 4: Hybrid data training
[0091] Refer to Figure 3 , mix the generated 5,000 synthetic images with the original 5,000 training sets, and input them into a multi-label classification model (such as ResNet) for training to optimize the classification performance.
[0092] The effects of the present invention will be further described below in conjunction with simulation experiments.
[0093] 1. Simulation conditions:
[0094] The simulation hardware environment is as follows: The GPU is NVIDIA RTX 4090, the CPU is AMD EPYC 9654 96-Core Processor, the operating system is Ubuntu 20.04, and the software environment is Python 3.9 and PyTorch 2.0.
[0095] 2. Simulation content:
[0096] Use 5,000 training sets of the Multi-label Weather Dataset, and respectively generate new multi-label weather image data by using the diffusion model method based on U-Net and the method of the present invention (the multi-label weather image data augmentation method based on DiT conditional guidance). After mixing the generated image data with the original training sets, input them into a pre-trained ResNet-50 model for multi-label weather classification training, and compare the classification performances of the two methods.
[0097] To evaluate the quality and diversity of the generated images, the Fréchet Inception Distance (FID) was used as the main metric to compare the distribution similarity between the generated images and the real images (real samples in the Multi-label Weather Dataset). Additionally, FIDs (Fréchet Inception Distance Standard Deviation) were calculated to evaluate the stability of the generated image distribution. Meanwhile, the mean average precision (mAP) and training time of the ResNet-50 model on the multi-label weather classification task were recorded as supplementary metrics for classification performance.
[0098] 3. Simulation Results:
[0099] Table 1 shows the Fréchet Inception Distance (FID) and FIDs values of different methods:
[0100] Table 1
[0101] Method FID FIDs Unet 4.1832 14.31 DiT (the method of the present invention) 2.6872 13.02
[0102] As can be seen from Table 1, the method of the present invention based on DiT is significantly superior to the diffusion model method based on U-Net in terms of the FID metric. The FID value decreased from 4.1832 to 2.6872, indicating that the generated synthetic images are closer to the distribution of real multi-label weather images, with higher image quality and diversity. The FIDs value (standard deviation) decreased from 14.31 to 13.02, showing that the image distribution generated by the method of the present invention is more stable, with smaller fluctuations and a more reliable generation process.
[0103] Further analyzing the classification performance of ResNet-50, Table 2 shows the mean average precision (mAP) of different methods on the multi-label weather classification task:
[0104] Table 2
[0105] Method mAP (%) Unet 78.5 DiT (the method of the present invention) 82.3
[0106] As can be seen from Table 2, the method of the present invention increased the mAP of ResNet-50 to 82.3%, which is 3.8% higher than that of the diffusion model method based on U-Net, indicating that the synthetic data generated by the present invention can better improve the classification performance of the model.
[0107] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0108] The above simulation analysis proves the correctness and effectiveness of the method proposed by the present invention.
[0109] The parts not detailed in the present invention belong to the common general knowledge of those skilled in the art.
[0110] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Obviously, for professionals in the field, after understanding the content and principle of the present invention, various modifications and changes in form and details may be made without departing from the principle and structure of the present invention. However, these modifications and changes based on the idea of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A multi-label weather image data expansion method based on DiT condition guidance, characterized in that: The steps include: 1) Obtain multi-label weather image data from public datasets, and shuffle the original image order of half of them as training sets; 2) Construct a diffusion model DiTc based on DiT conditional guidance, in which multi-label conditions are embedded into y through conditional fusion mechanism. label and the conditional embedding y at diffusion time step t t Perform integration to generate comprehensive condition information for guiding model output; 3) Training model DiTc: (3.1) Initialize the encoder and decoder parameters of the variational autoencoder VAE, initialize the parameters of the diffusion model DiTc, set the maximum time step of the training phase to T, and initialize the current training time step t = 0; (3.2) The training set is used as the model input data, mapped to a low-dimensional latent space through the VAE encoder, and the initial latent representation z0 of the input image is generated. Gaussian noise is applied to z0 to generate the noisy latent representation z at the tth time step. t ; (3.3) Using the model DiTc to predict the noise ∈ in the noise adding process θ : ∈ θ =DiTc(z t ,t,y label ); (3.4) Use the mixed loss function L to optimize the model DiTc and calculate the gradient through back propagation The DiTc parameters of the model are updated according to the gradient; the updating process is iterated repeatedly until the model converges, and an optimized diffusion model guided by the DiT condition is obtained; 4) Use the optimized diffusion model to expand the multi-label weather image data, that is, use the multi-label conditions as the input of the optimized diffusion model to generate a multi-label synthetic weather image image′; 5) The image image′ generated in step 4) is mixed with the training set image in step 1) to obtain an expanded dataset of multi-label weather image data guided by DiT conditions for optimization training of the multi-label classification model.
2. The method according to claim 1, characterized in that: The training set in step (1) is obtained by obtaining N weather images containing multiple weather categories, each image containing multiple weather labels; half of the N images are randomly shuffled to obtain a data set for training the deep learning model; the weather categories include at least cloudy, sunny, rainy, foggy and snowy.
3. The method according to claim 1, characterized in that: In step 2), the multi-label conditions are embedded into y through the conditional fusion mechanism label and the conditional embedding y at diffusion time step t t Integrate and generate comprehensive condition information, which is implemented as follows: (2.1) Embed multi-label conditions into y label and the embedding of diffusion time step t are taken as separate inputs; (2.2) Generate multi-label conditional embedding through independent Embedding layer Generate an embedding at time step t through an embedding layer (2.3) Inside the model DiTc, y is normalized by adaptive layer adaLN. label and t The two are combined to obtain comprehensive condition information Condition merged : Condition merged =adaLN(z t ,y label +y t )。 4. The method according to claim 1, characterized in that: The diffusion model DiTc parameters initialized in step (3.1) and the optimized diffusion model parameters initialized in step (4.1) both include the conditional network parameters required for the self-attention mechanism, feedforward neural network, and adaptive layer normalization adaLN.
5. The method according to claim 1, characterized in that: In step (3.2), the initial latent representation z0 of the input image is generated, and Gaussian noise is applied to z0 to generate the noisy latent representation z at the tth time step t , as follows: z0=Encoder(x), Among them, x represents the multi-label weather image data in the training set, α t It is a preset noise scheduling parameter used to control the degree of noise addition; Represents random noise from a standard normal distribution.
6. The method according to claim 1, characterized in that: The mixed loss function L described in step (3.4) is expressed as follows: L=λ simple ||∈ θ -∈|| 2 +λ vlb L vlb ; Among them, λ simple and λ vlb are the weight coefficients controlling the simple mean square error and variational lower bound loss, L vlb represents the variational lower bound loss, which is used to measure the distribution difference between the latent space representation and the target.
7. The method according to claim 6, characterized in that: In step (3.4), according to the gradient Update the DiTc parameters of the model as follows: Among them, θ′ represents the updated parameter set of model DiTc, θ represents the current parameter set of model DiTc; η is the learning rate.
8. The method according to claim 1, characterized in that: In step (4), a multi-label synthetic weather image image′ is generated, and the implementation steps are as follows: (4.1) Initialize the encoder and decoder parameters of VAE, initialize the optimized diffusion model parameters, set the maximum time step of the generation phase to T′, initialize the current generation time step t′=0; initialize the latent representation and multi-label conditions; (4.2) In the latent space, Gaussian noise is gradually added to the latent representation of the training set images to generate a noisy sequence z1,z2,…,z T′ ; (4.3) Design an independent Embedding layer for each weather category to generate multi-label conditions y′ label ; Generate the embedding y at time step t′ through the embedding layer t′ =Embed(t′), The comprehensive condition information is generated by adaptive layer normalization adaLN operation inside the optimized diffusion model; (4.4) The optimized diffusion model is based on the multi-label condition y′ label and prediction noise ∈ ′ at time step t′ θ (z t ,t,y label ), the conditional information is fused through the merging operation and gradually denoised to the potential representation z′0 at the initial time step; (4.5) Decode the image through the VAE decoder, that is, convert z′0 into a synthetic weather image image′; image′=Decoder(z′0).
9. The method according to claim 8, characterized in that: The stepwise denoising to the potential representation z′0 of the initial time step described in step (4.4) is implemented according to the following formula: z t′-1 =z t′ -∈ θ (z t′ ,t′,y label ), Among them, z t′-1 is the potential representation after denoising at the t′-1th time step.
10. The method according to claim 1, characterized in that: The maximum time step of the training phase in step (3.1) is T, which ranges from 1000 to 2000; the maximum time step of the generation phase in step (4.1) is T′, which ranges from 500 to 1000.
Citation Information
Patent Citations
Hybrid degraded image restoration method based on joint conditional diffusion model
CN118864291A
Two-stage rain and fog image restoration method based on conditional diffusion model
CN118864319A
Cited By
Typhoon rapid enhancement prediction method based on time-space sequence and multi-modal feature fusion
CN120633957A