Self-adaptive snow and rain removal method based on feature-driven GAN
Through the adaptive snow rain removal method of feature-driven GAN, the problem of insufficient distinction between snow/rain coverage areas in the prior art is solved, and efficient image recovery effect is achieved, which is suitable for snow removal and rain removal tasks in complex scenes.
Patent Information
- Application Number
- CN202510285634.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-08
AI Technical Summary
The existing generator architecture and loss functions fail to effectively distinguish the snow/rain coverage areas during the removal and rain removal process, resulting in limited removal effects, and traditional methods have blurring problems when dealing with complex degraded images.
Adaptive snow rain removal method based on feature-driven GAN is adopted, and precise removal of snowflakes and raindrops are achieved through feature extraction network architecture, spatially adaptable hierarchical constraint loss system and multi-stage collaborative optimization dual-path reasoning framework, combined with industrial-grade trusted evaluation system.
It significantly improves the recovery performance of rain and snow degraded images, and can generate clear rain and snow-free images in complex scenes, which are highly adaptable and meet professional film and television color grading standards.
Smart Images

Figure CN120278913A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image snow and rain removal, and specifically relates to an adaptive snow and rain removal method based on feature-driven GAN. Background Art
[0002] In recent years, significant progress has been made in single-image snow and rain removal methods based on generative adversarial networks (GANs). These methods improve the restoration performance by introducing unique network architectures, descriptors, and loss functions (such as Cai et al., 2021; Ding et al., 2021; Zhang et al., 2019). A typical GAN framework consists of a generator network and a discriminator network. The main task of the discriminator is to distinguish whether the image generated by the generator is real or fake.
[0003] However, the optimization of the generator not only depends on the feedback of the discriminator but is also constrained by various loss functions, which provide important guidance for the accuracy of the images it generates. For example, multi-scale pixel-level loss, structural similarity index (SSIM) loss, perceptual optimization loss, and multi-scale perceptual loss are usually integrated into the objective function of the GAN framework to improve the quality of the generated images.
[0004] Nevertheless, these loss functions have a significant limitation: they are indiscriminate about image regions, that is, the regions not covered by snow / rain are treated equally as the regions covered by snow / rain during the loss minimization process. This indiscriminateness causes the generator to fail to effectively focus on the regions mainly occupied by snow / rain particles, thus limiting the further improvement of snow and rain removal effects. At the same time, existing generator architectures, such as U-Net and ResNet, also have deficiencies: U-Net is limited by the depth of feature extraction and computing resources and is difficult to adapt to the diversity of snowflakes and raindrops; although ResNet enhances feature extraction through residual blocks, it has a blurring problem when generating images due to the lack of skip connections. Therefore, we provide an adaptive snow and rain removal method based on feature-driven GAN. Summary of the Invention
[0005] Aiming at the above-mentioned shortcomings of the prior art, the first object of the present invention is to provide an adaptive snow and rain removal method based on feature-driven GAN to solve the problems in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] An adaptive snow and rain removal method based on feature-driven GAN, comprising the following steps:
[0008] S1. Construct a feature extraction network architecture for complex degradation features;
[0009] S2, Spatially Adaptive Hierarchical Constraint Loss System;
[0010] S3, Multi-Stage Cooperative Optimization Dual-Path Inference Framework;
[0011] S4, Construction of Industrial-Level Trusted Evaluation System.
[0012] The present invention is further configured as follows: In the construction of the feature extraction network architecture for complex degradation features in step S1:
[0013] S1.1, Embed multi-branch Inception modules in the U-Net encoder stage to construct a feature extraction network with a deep pyramid structure. Each Inception module consists of 4 parallel branches:
[0014] (1) Use a 1×1 convolutional kernel for basic feature projection;
[0015] (2) Use a 3×3 depthwise separable convolution to extract local spatial features;
[0016] (3) Use a 5×5 dilated convolution to expand the receptive field;
[0017] (4) Channel adaptive attention layer;
[0018] The multi-branch outputs are fused through a gated weighting mechanism after normalization. The number of feature channels is controlled within 256 dimensions to reduce video memory occupancy. This design enables the network to achieve long-range correlation modeling of full-size snowflakes (≥128px) with 10.3M parameters;
[0019] S1.2, Introduce a deformable convolution-guided feature alignment module at the skip connection to address the dynamic snowflake offset problem in combination with an optical flow correction algorithm. Before each cross-layer connection, perform scale-aware resampling on the encoder features, and use a joint operation of bilinear interpolation fusion and 3×3 deformable convolution to ensure that the feature alignment error is ≤0.8px at 4K resolution input. At the same time, introduce a channel attention (CBAM) mechanism with second-order momentum constraint to dynamically adjust the feature weights at each level to improve the utilization rate of effective information;
[0020] S1.3, Adopt a 3-level cascade structure in the decoding stage, with each level containing 2 residual LSTM modules with deep supervision:
[0021] The first level processes the 256×256 feature map and uses ConvGRU to model temporal features;
[0022] The second level applies a spatial transformation network (STN) for local affine transformation correction;
[0023] The third level introduces a spectrum repair module guided by a physical degradation model;
[0024] Thresholded gradient clipping is set between each layer to control the range of backpropagation gradients within the interval [1e-4, 1e-2], ensuring the stability of model convergence;
[0025] S1.4. For RAW file input, design an end-to-end adaptive preprocessing layer based on the Bayer pattern, dynamically generate demosaicing parameters using a conditional generative adversarial mechanism, use phase correlation loss to constrain color consistency, jointly train the preprocessing module with the main network, support an input dynamic range of 12bit / 16bit professional image formats, and control the processing delay within 23ms / frame (NVIDIA A6000).
[0026] The present invention is further configured as follows: in the step S2, the spatially adaptive hierarchical constraint loss system:
[0027] S2.1. Construct a prior knowledge base of snowflake morphology based on deep pre-training, generate a three-dimensional snow depth map under the guidance of the YOLOv7-Snow model output, and derive a dynamic mask by combining the atmospheric scattering physical model: M = αSigmoid(GT-LQ difference map) + β semantic segmentation result, where α = 0.7 and β = 0.3 are empirical coefficients, calculate the temporal continuity constraint through the optical flow pyramid, and locate the difference region of consecutive frames to achieve sub-pixel mask accuracy;
[0028] S2.2. Construct pyramid perception loss at the 3rd, 5th, and 7th layers from the bottom of VGG19, with the weight ratio of each layer being 3:2:1. Introduce the middle layer of SWIN-Transformer as an auxiliary discriminator, extract spatial frequency features to form a 64-dimensional description vector, calculate the difference in feature distribution through the Wasserstein distance, and fuse spectral normalization constraints to enhance the consistency of local region gradient directions;
[0029] S2.3. Propose an optical flow consistency loss term:
[0030] L f low = ‖Φ(G(X t )) - Φ(G(X t+1 ))‖1
[0031] where Φ is a pre-trained RAFT optical flow network. For high-speed moving snowflakes (speed > 10px / frame), introduce an acceleration constraint term to strengthen the continuity of the motion trajectory and avoid jitter in the generated sequence;
[0032] S2.4. Construct a 4-channel input (RGB + near-infrared), establish a material reflection feature library in the spectral range of 350 - 1100nm, introduce a spectral weighting module in each residual block, dynamically fuse multi-spectral features through a learnable parameter matrix, and the spectral fidelity term constrains the ΔE in the CIE-Lab space between the generated image and GT to be < 2.5, meeting the professional film and television color grading standard.
[0033] The present invention is further configured such that in the step S3, the two-path inference framework with multi-stage collaborative optimization:
[0034] S3.1. The backbone network adopts a dual-encoder architecture. The SnowPath encoder focuses on extracting high-frequency information in the wavelet domain and is equipped with Haar wavelet initialization convolutional kernels. The RainPath encoder strengthens the spatial attention mechanism and is initialized with an anisotropic Gabor filter bank. The features of the two paths are fused through an adaptive gating network in the decoding stage. The input of the gating network includes meteorological sensor data (metadata such as precipitation type and intensity).
[0035] S3.2. A dynamic hyperparameter optimization space is established, and periodic parameter tuning is performed through a DWave quantum annealing machine. Key parameters such as the learning rate and decay coefficient are encoded into 512-bit quantum states, and the optimal combination is automatically searched on the ImageDesnow-2025 benchmark dataset. The optimization period is set to restart every 50 epochs.
[0036] S3.3. Develop a physical-level snowflake motion simulation system based on the NS equation, control the number of particles at the order of 1e6, establish a database of 7 types of snow crystals (plate-like, columnar, etc.), generate multi-angle scattering effects in combination with real atmospheric refractive index parameters, support dynamic adjustment of precipitation rate (0 - 50 mm / h), wind speed (0 - 20 m / s), and ambient temperature (-30°C to +5°C), and generate 10240 hours of equivalent training samples.
[0037] S3.4. Customize dedicated operators for the NVIDIA Ampere architecture, adopt a mixed-precision calculation process deeply optimized by cuDNN 8.5, design a model parallel strategy based on Triton, split the generator into 3 segments and deploy them on a 4-GPU cluster, and the overall inference speed reaches 128 FPS @ 4K resolution, with the video memory occupancy stable within 18 GB.
[0038] The present invention is further configured such that in the step S4, the construction of the industrial-level trusted evaluation system:
[0039] S4.1. Define the comprehensive index of Snow-Removal-Score (SRS):
[0040] SRS = 0.4PSNR + 0.3SSIM + 0.2LPIPS + 0.1HumanScore
[0041] Construct a test set containing 10,000 professional images, covering extreme scenarios (such as blizzard weather, low-light environment, etc.), introduce a frequency band evaluation mechanism, and calculate the recovery ability indicators at high frequency (30 lp / mm), medium frequency (10 - 30 lp / mm), and low frequency (<10 lp / mm) respectively;
[0042] S4.2. Develop an Xilinx Versal ACAP evaluation kit, integrate 4-channel 4K@60fps video input, deploy a dedicated preprocessing IP core to implement pipeline operations such as Bayer array decoding and white balance correction, and form a hardware closed-loop verification system with the main model, with a latency jitter <0.5 ms;
[0043] S4.3. Design an adversarial attack test set, including 100 adversarial samples based on the FGSM and PGD methods, establish a degradation parameter perturbation matrix, and conduct Monte Carlo sampling tests on key parameters such as snowfall density (±30%) and snowflake scale (±50%) to ensure that the model performance fluctuation is within ±3 dB PSNR;
[0044] S4.4. Integrate lidar point cloud assisted evaluation, obtain the real geometric information of the scene through a ToF depth camera, and construct three-dimensional reconstruction quality indicators: the point cloud coincidence degree ≥92%, the edge sharpness loss ≤15%. At the same time, use an infrared thermal imager to verify the physical consistency, and require the temperature distribution variance ≤0.3℃.
[0045] The present invention is further configured as: in the step S2, the spatially adaptive hierarchical constraint loss system, the rain removal network includes two generators: the first generator Gd is used to remove raindrops and rain lines from the input rain image Xr and generate an intermediate image The second generator Gr further corrects the color deviation or artifacts in the image generated by Gd to generate the final rain-removed image The two generators share the same discriminator D, and generate more realistic images through collaborative optimization, so as to jointly "fool" the discriminator. The objective function LGd of the generator Gd includes LG which is the same as the formula, and a newly introduced loss function used to guide the generator to generate better rain-removed images, considering the unique shape and appearance characteristics of raindrops and rain lines, designed based on the gradient magnitude maps of the rain-free image C and the generated image which is calculated as follows:
[0046]
[0047] Similarly, the objective function LGr of the generator Gr consists of the basic loss LG and the loss LSSIM based on the structural similarity index (SSIM), and its definition is as follows:
[0048]
[0049] Among them, C and respectively represent the real rain-free image and the enhanced image generated by the generator Gr. μ and σ respectively represent the mean and variance of the image, and c1 and c2 are constants. Therefore, the objective function of the generator Gr is defined as:
[0050] LGr = LG + LSSIM
[0051] Real snow image generation: Let the video tensor A ∈ Rm×n×k, where m represents the number of rows, n represents the number of columns, and k represents the number of frames. According to the selected dimension, three slicing modes can be extracted: horizontal slice Ai:: ∈ Rk×n, side slice A:j: ∈ Rk×m, and frontal slice A::k ∈ Rm×n. Considering the spatio-temporal characteristics of A:j and A::k, low-rank and sparse decomposition of one of them can achieve the separation of the background layer and the snow layer. Specifically, the SVD of the spatio-temporal slice Ai:: is expressed as:
[0052] Ai:: = μ1σ1V1T +... + μdσdVdT
[0053] Temporal filtering and snow removal: By applying appropriate temporal filtering to the foreground slice, the snow represented by short-term motion patterns can be removed. Specifically, the filtering operation performs temporal filtering on the left singular vectors representing motion information:
[0054] μl' = f(μl) for l = 2, 3,..., q
[0055]
[0056] AGT = B + Fdesnowed + N.
[0057] Beneficial effects
[0058] Adopting the technical solution provided by the present invention, compared with the known public technology, it has the following beneficial effects:
[0059] The rain / snow removal framework based on generative adversarial networks proposed by the present invention innovates in terms of the generator architecture, loss function design, and collaborative training of the generator and discriminator. By introducing a dual-generator structure, an attention mechanism, and a spatial guidance loss, the restoration performance of rain / snow degraded images is significantly improved, providing an effective solution for practical applications. Through these innovations, the model demonstrates excellent image restoration effects in complex rain / snow scenes. The framework includes two generators, which respectively process the degraded parts in rain and snow images. The first generator focuses on removing raindrops and rain lines to generate an intermediate image, and the second generator further corrects the color deviation or artifacts in the image to finally generate a clear rain-free image. The two generators share the same discriminator and jointly optimize to generate higher-quality images. In addition, a spatial guidance loss is introduced into the framework. Through this mechanism, the generator is forced to pay more attention to the areas covered by raindrops or snowflakes in the image. This method effectively avoids the limitation of traditional loss functions that treat all areas equally, improving the restoration performance of the model in the rain / snow removal task. The generator and discriminator jointly learn through adversarial training. The discriminator helps the generator optimize the image generation process through feedback, while the generator generates results as close to real images as possible by optimizing the loss function. Through such training, the model can gradually improve the rain / snow removal effect and show strong adaptability in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 It is a schematic flow chart of the method steps of the adaptive rain / snow removal method based on feature-driven GAN of the present invention;
[0061] Figure 2 It is a schematic diagram of the SGAN model structure of the adaptive rain / snow removal method based on feature-driven GAN of the present invention;
[0062] Figure 3 It is a schematic diagram of the collaborative framework structure of the adaptive rain / snow removal method based on feature-driven GAN of the present invention;
[0063] Figure 4 It is a schematic diagram of the comparison of rain removal results of the adaptive rain / snow removal method based on feature-driven GAN of the present invention;
[0064] Figure 5 It is a schematic diagram of the comparison of snow removal results of the adaptive rain / snow removal method based on feature-driven GAN of the present invention;
[0065] Figure 6 It is a schematic diagram of the comparison of defogging effects of different methods of the adaptive rain / snow removal method based on feature-driven GAN of the present invention;
[0066] Figure 7Schematic diagram for comparing the snow removal effects of different methods of the adaptive snow and rain removal method based on feature-driven GAN of the present invention. Detailed implementation manners
[0067] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] The present invention will be further described below in conjunction with embodiments.
[0069] Embodiment 1
[0070] As Figure 1-7 shown, the present invention provides a technical solution: an adaptive snow and rain removal method based on feature-driven GAN, including the following steps:
[0071] S1. Construction of a feature extraction network architecture for complex degradation features;
[0072] S2. A spatially adaptive hierarchical constraint loss system;
[0073] S3. A two-path inference framework for multi-stage collaborative optimization;
[0074] S4. Construction of an industrial-level reliable evaluation system;
[0075] In the step S1, the construction of the feature extraction network architecture for complex degradation features:
[0076] S1.1. Embed multi-branch Inception modules in the U-Net encoder stage to construct a feature extraction network with a deep pyramid structure. Each Inception module consists of 4 parallel branches:
[0077] (1) 1×1 convolutional kernel for basic feature projection;
[0078] (2) 3×3 depthwise separable convolution to extract local spatial features;
[0079] (3) 5×5 dilated convolution to expand the receptive field;
[0080] (4) Channel adaptive attention layer;
[0081] The multi-branch outputs are fused through a gated weighting mechanism after normalization. The number of feature channels is controlled within 256 dimensions to reduce video memory occupancy. This design enables the network to achieve long-range correlation modeling of full-size snowflakes (≥128px) with 10.3M parameters.
[0082] S1.2. Introduce a deformable convolution-guided feature alignment module at the jump connection, combine with an optical flow correction algorithm to handle the dynamic snowflake offset problem. Before each cross-layer connection, perform scale-aware resampling on the encoder features, and adopt the joint operation of bilinear interpolation fusion and 3×3 deformable convolution to ensure that the feature alignment error is ≤0.8px under 4K resolution input. At the same time, introduce a channel attention (CBAM) mechanism with second-order momentum constraint to dynamically adjust the feature weights at each level to improve the utilization rate of effective information;
[0083] S1.3. Adopt a three-level cascade structure in the decoding stage, and each level contains 2 residual LSTM modules with deep supervision:
[0084] The first level processes the 256×256 feature map and uses ConvGRU to model the temporal features;
[0085] The second level applies a spatial transformation network (STN) for local affine transformation correction;
[0086] The third level introduces a spectrum repair module guided by a physical degradation model;
[0087] Set thresholded gradient clipping between each level to control the range of the backpropagation gradient in the interval [1e-4, 1e-2] to ensure the convergence stability of the model;
[0088] S1.4. For RAW file input, design an end-to-end adaptive preprocessing layer based on the Bayer pattern, adopt a conditional generative adversarial mechanism to dynamically generate demosaicing parameters, use the phase correlation loss to constrain the color consistency, and jointly train the preprocessing module and the main network to support an input dynamic range of 12bit / 16bit professional image formats, with the processing delay controlled within 23ms / frame (NVIDIA A6000);
[0089] In the step S2, the spatially adaptive hierarchical constraint loss system:
[0090] S2.1. Construct a prior knowledge base of snowflake morphology based on deep pre-training, generate a three-dimensional snow depth map under the guidance of the YOLOv7-Snow model output, and derive a dynamic mask by combining the atmospheric scattering physical model: M = αSigmoid(GT - LQ difference map) + β semantic segmentation result, where α = 0.7 and β = 0.3 are empirical coefficients. Calculate the temporal continuity constraint through the optical flow pyramid to locate the continuous frame difference region and achieve sub-pixel mask accuracy;
[0091] S2.2. Construct pyramid perception loss at the 3rd, 5th, and 7th layers from the end of VGG19. The weight ratio of each layer is 3:2:1. Introduce the middle layer of SWIN-Transformer as an auxiliary discriminator, extract spatial frequency features to form a 64-dimensional description vector, calculate the difference in feature distribution through the Wasserstein distance, and fuse spectral normalization constraints to enhance the consistency of gradient directions in local regions;
[0092] S2.3. Propose an optical flow consistency loss term:
[0093] L f low=‖Φ(G(X t ))-Φ(G(X t+1 ))‖1
[0094] where Φ is a pre-trained RAFT optical flow network. For high-speed moving snowflakes (speed > 10px / frame), introduce an acceleration constraint term to strengthen the continuity of the motion trajectory and avoid jitter in the generated sequence;
[0095] S2.4. Construct a 4-channel input (RGB + near-infrared), establish a material reflection feature library in the spectral range of 350 - 1100nm, introduce a spectral weighting module in each residual block, dynamically fuse multi-spectral features through a learnable parameter matrix, and the spectral fidelity term constrains that the ΔE in the CIE-Lab space between the generated image and the GT is < 2.5, meeting the professional film and television color grading standard;
[0096] In the step S3, the multi-stage collaborative optimization dual-path inference framework:
[0097] S3.1. The backbone network adopts a dual-encoder architecture. The SnowPath encoder focuses on extracting high-frequency information in the wavelet domain and is equipped with Haar wavelet initialization convolution kernels. The RainPath encoder strengthens the spatial attention mechanism and is initialized with an anisotropic Gabor filter bank. The dual-path features are fused through an adaptive gating network in the decoding stage. The input of the gating network includes meteorological sensor data (metadata such as precipitation type and intensity);
[0098] S3.2. Establish a dynamic hyperparameter optimization space, perform periodic parameter tuning through a DWave quantum annealing machine, encode key parameters such as the learning rate and decay coefficient into 512-bit quantum states, automatically search for the optimal combination on the ImageDesnow-2025 benchmark dataset, and set the optimization period to restart every 50 epochs;
[0099] S3.3. Develop a physical-level snowflake motion simulation system based on the NS equation, control the number of particles at the order of 1e6, establish a database of 7 types of snow crystal (plate-like, columnar, etc.), generate multi-angle scattering effects by combining real atmospheric refractive index parameters, support dynamic adjustment of precipitation rate (0 - 50 mm / h), wind speed (0 - 20 m / s), and environmental temperature (-30°C to +5°C), and generate 10240 hours of equivalent training samples;
[0100] S3.4. Customize a dedicated operator for the NVIDIA Ampere architecture, adopt a mixed-precision calculation process with deep optimization by cuDNN 8.5, design a model parallel strategy based on Triton, split the generator into 3 segments and deploy them on a 4-GPU cluster, with the overall inference speed reaching 128 FPS @ 4K resolution and the video memory occupancy stable within 18 GB;
[0101] In the step S4, constructing an industrial-level trusted evaluation system:
[0102] S4.1. Define the comprehensive index of Snow-Removal-Score (SRS):
[0103] SRS = 0.4PSNR + 0.3SSIM + 0.2LPIPS + 0.1HumanScore
[0104] Construct a test set containing 10,000 professional images, covering extreme scenarios (blizzard weather, low-light environment, etc.), introduce a frequency band evaluation mechanism, and calculate the restoration ability indicators at high frequency (30 lp / mm), medium frequency (10 - 30 lp / mm), and low frequency (<10 lp / mm) respectively;
[0105] S4.2. Develop an Xilinx Versal ACAP evaluation kit, integrate 4-channel 4K@60fps video input, deploy a dedicated preprocessing IP core to implement pipeline operations such as Bayer array decoding and white balance correction, and form a hardware closed-loop verification system with the main model, with the delay jitter <0.5 ms;
[0106] S4.3. Design an adversarial attack test set, including 100 types of adversarial samples based on the FGSM and PGD methods, establish a degradation parameter perturbation matrix, and conduct Monte Carlo sampling tests on key parameters such as snowfall density (±30%) and snowflake scale (±50%) to ensure that the model performance fluctuates within ±3 dB PSNR;
[0107] S4.4. Integrate lidar point cloud assisted evaluation, obtain the real geometric information of the scene through a ToF depth camera, construct three-dimensional reconstruction quality indicators: the point cloud coincidence degree ≥92%, the edge sharpness loss ≤15%, and at the same time use an infrared thermal imager to verify the physical consistency, requiring the temperature distribution variance ≤0.3°C;
[0108] In the above-mentioned step S2, in the spatially adaptive hierarchical constraint loss system, the rain removal network includes two generators: the first generator Gd is used to remove raindrops and rain lines from the input rainy image Xr and generate an intermediate image The second generator Gr further corrects the color deviation or artifacts in the image generated by Gd to generate the final rain-removed image The two generators share the same discriminator D. By collaborating and optimizing, they generate more realistic images, thus jointly "deceiving" the discriminator. The objective function LGd of generator Gd includes LG which is the same as the formula, and a newly introduced loss function used to guide the generator to generate better rain-removed images. Considering the unique shapes and appearance features of raindrops and rain lines Based on the gradient magnitude map design of the rain-free image C and the generated image It is calculated as follows:
[0109]
[0110] Similarly, the objective function LGr of generator Gr consists of the basic loss LG and the loss LSSIM based on the structural similarity index (SSIM), and its definition is as follows:
[0111]
[0112] Among them, C and represent the real rain-free image and the enhanced image generated by generator Gr respectively, μ and σ represent the mean and variance of the image respectively, and c1 and c2 are constants. Therefore, the objective function of generator Gr is defined as:
[0113] LGr = LG + LSSIM
[0114] Generation of real snow images Let the video tensor A ∈ Rm×n×k, where m represents the number of rows, n represents the number of columns, and k represents the number of frames. According to the selected dimensions, three slicing modes can be extracted: horizontal slice Ai:: ∈ Rk×n, side slice A:j: ∈ Rk×m, and frontal slice A::k ∈ Rm×n. Considering the spatio-temporal characteristics of A:j and A::k, separating the background layer and the snow layer can be achieved by performing low-rank and sparse decomposition on one of them. Specifically, the SVD of the spatio-temporal slice Ai:: is expressed as:
[0115] Ai:: = μ1σ1V1T +... + μdσdVdT
[0116] Temporal filtering and snow removal By applying appropriate temporal filtering to the foreground slice, the snow represented by short-term motion patterns can be removed. Specifically, the filtering operation performs temporal filtering on the left singular vectors representing motion information:
[0117] μl' = f(μl) for l = 2, 3, ..., q
[0118]
[0119] AGT = B + Fdesnowed + N。
[0120] In this embodiment, by introducing innovative network architectures, descriptors, and loss functions, these methods exhibit superior performance in processing complex rain and snow degraded images. However, a significant limitation of existing methods is that they fail to effectively focus on the image regions occluded by snowflakes or raindrops, which restricts the further improvement of the removal effect. The loss mechanism significantly enhances the recovery performance by forcing the generator to pay more attention to the regions where snow / rain particles are located. Combining the improved U-Net generator architecture and the GAN architecture with dual-generator collaboration, the model in this paper can not only effectively remove raindrops and snowflakes but also perform excellently on multiple datasets. Especially when dealing with complex degradation scenarios, our method demonstrates the ability to outperform existing rain and snow removal techniques. Through extensive experimental verification, the PSNR and SSIM metrics of the model in this paper on multiple standard datasets (such as Rain200H, Rain200L, DID-Data, etc.) exceed those of existing state-of-the-art methods, proving the effectiveness and applicability of the proposed framework in practical applications. We propose a brand-new method to solve the problem of simultaneous rain and snow removal through innovative network design, loss function, and co-training method, and provide new ideas for future image restoration research.
[0121] The present invention constructs a feature extraction network through an Inception-Unet hybrid architecture. In the encoding stage, a multi-branch structure is adopted to process rain and snow particles of different scales: 1×1 convolution compresses the feature dimension, 3×3 depthwise separable convolution captures local textures, 5×5 dilated convolution expands the receptive field to adapt to the long-range correlation characteristics of snowflakes, and the channel adaptive attention layer dynamically adjusts the feature weights. This network combines the advantages of the skip connections of U-Net and the multi-scale feature extraction ability of the Inception module. When processing 4K resolution inputs, the deformable convolution alignment module in the S1.2 stage can control the feature offset error within 0.8px, and the important channel information is retained with the cooperation of the CBAM mechanism. In the preprocessing link, for professional RAW format inputs, the demosaicing parameters are dynamically optimized through conditional GAN, and a real-time processing ability of 23ms / frame is achieved on an NVIDIA A6000 GPU.
[0122] During the execution process, the system adopts a hierarchical constraint loss system to strengthen physical rationality. Based on the three-dimensional snow depth map constructed in S2.1, combined with the dynamic mask calculated by the optical flow pyramid
[0123] M = 0.7Sigmoid(ΔGT) + 0.3SegMap
[0124] Precisely locate the snow - covered area. For raindrop removal, a collaborative mechanism of two generators (S3.1) is adopted to form an iterative optimization loop: in the first - stage generator Gd, the gradient magnitude constraint is used to optimize the preservation of raindrop edges, and in the second - stage generator Gr, the SSIM metric is introduced
[0125]
[0126] to correct color deviation. For dynamic snowflakes, combined with the physical simulation system based on the NS equation in S3.3 stage, 10240 - hour equivalent training data containing 1e6 snow crystal particles and seven crystal morphologies (plate - like, columnar, etc.) are generated, and the spatio - temporal slicing separation of the video tensor A is achieved through SVD decomposition: Ai::=ΣμlσlvlT. The time - domain filtering of the left - singular vector μl is used to remove the dynamic snowflake components and retain the low - rank background information.
[0127] In the industrial deployment stage, the system ensures reliability through a quantitative evaluation system. Multidimensional evaluation indicators are adopted
[0128] SRS = 0.4PSNR + 0.3SSIM + 0.2LPIPS + 0.1HumanScore
[0129] Verify the performance on a test set containing 10,000 professional images (covering extreme scenarios such as blizzards and low - light conditions). At the hardware level, a dedicated evaluation kit is developed based on Xilinx Versal ACAP, integrating 4 - channel 4K@60fps input and deploying the Bayer array decoding IP core to achieve an end - to - end delay jitter < 0.5ms. In the adversarial test session, the robustness of the model under perturbations such as snowfall density ±30% and snowflake size ±50% is verified through Monte Carlo sampling to ensure that the PSNR fluctuation does not exceed 3dB. The final output is verified for physical consistency through 3D lidar point cloud (coincidence degree ≥92%) and infrared thermal imaging (temperature variance ≤0.3℃), meeting the requirements of industrial - level applications.
[0130] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An adaptive snow and rain removal method based on feature-driven GAN, characterized in that It includes the following steps: S1. Construction of a feature extraction network architecture for complex degradation features; S2. A spatially adaptive hierarchical constraint loss system; S3. A dual-path inference framework for multi-stage collaborative optimization; S4. Construction of an industrial-level trustworthy evaluation system.
2. The adaptive snow and rain removal method based on feature-driven GAN according to claim 1, characterized in that: In the step S1, the construction of the feature extraction network architecture for complex degradation features: S1.
1. Embed multi-branch Inception modules in the U-Net encoder stage to construct a feature extraction network with a deep pyramid structure. Each Inception module consists of 4 parallel branches: (1) 1×1 convolutional kernel for basic feature projection; (2) 3×3 depthwise separable convolution to extract local spatial features; (3) 5×5 dilated convolution to expand the receptive field; (4) Channel adaptive attention layer; The multi-branch outputs are fused through a gated weighting mechanism after normalization. The number of feature channels is controlled within 256 dimensions to reduce video memory occupancy. This design enables the network to achieve long-range correlation modeling of full-size snowflakes (≥128px) with 10.3M parameters; S1.
2. Introduce a deformable convolution-guided feature alignment module at the skip connection to address the dynamic snowflake offset problem in combination with an optical flow correction algorithm. Before each cross-layer connection, perform scale-aware resampling on the encoder features, and adopt a joint operation of bilinear interpolation fusion and 3×3 deformable convolution to ensure that the feature alignment error is ≤0.8px under a 4K resolution input. At the same time, introduce a channel attention (CBAM) mechanism with second-order momentum constraint to dynamically adjust the feature weights at each level to improve the utilization rate of effective information; S1.
3. Adopt a 3-level cascaded structure in the decoding stage, and each level contains 2 residual LSTM modules with deep supervision: The first level processes the 256×256 feature map and uses ConvGRU to model temporal features; The second level applies a spatial transformation network (STN) for local affine transformation correction; The third level introduces a spectrum repair module guided by a physical degradation model; Set thresholded gradient clipping between each level to control the range of the backpropagation gradient in the interval [1e-4, 1e-2] to ensure the convergence stability of the model; S1.
4. For RAW file input, design an end-to-end adaptive preprocessing layer based on the Bayer pattern, dynamically generate demosaicing parameters using a conditional generative adversarial mechanism, use a phase correlation loss to constrain color consistency, and jointly train the preprocessing module and the main network to support an input dynamic range of 12bit / 16bit professional image formats, with the processing delay controlled within 23ms / frame (NVIDIA A6000).
3. The adaptive snow and rain removal method based on feature-driven GAN according to claim 1, wherein: In the step S2, the spatially adaptive hierarchical constraint loss system: S2.
1. Construct a snowflake morphology prior knowledge base based on deep pre-training, generate a three-dimensional snow depth map under the guidance of the YOLOv7-Snow model output, and deduce a dynamic mask by combining the atmospheric scattering physical model: M = αSigmoid(GT-LQ difference map) + β semantic segmentation result, where α = 0.7 and β = 0.3 are empirical coefficients. Calculate the temporal continuity constraint through the optical flow pyramid, locate the continuous frame difference region, and achieve sub-pixel mask accuracy; S2.
2. Construct pyramid perceptual losses at the 3rd, 5th, and 7th layers from the end of VGG19, with the weight ratio of each layer being 3:2:
1. Introduce the middle layer of SWIN-Transformer as an auxiliary discriminator, extract spatial frequency features to form a 64-dimensional description vector, calculate the feature distribution difference through the Wasserstein distance, and fuse the spectral normalization constraint to enhance the consistency of the local region gradient direction; S2.
3. Propose an optical flow consistency loss term: L f low = ‖Φ(G(X t )) - Φ(G(X t+1 ))‖1 where Φ is the pre-trained RAFT optical flow network. For high-speed moving snowflakes (speed > 10px / frame), introduce an acceleration constraint term to strengthen the continuity of the motion trajectory and avoid generating sequence jitter; S2.
4. Construct a 4-channel input (RGB + near-infrared), establish a material reflection feature library in the spectral range of 350 - 1100nm, introduce a spectral weighting module in each residual block, dynamically fuse multi-spectral features through a learnable parameter matrix, and the spectral fidelity term constrains that the ΔE in the CIE-Lab space between the generated image and GT is < 2.5, meeting the professional film and television color grading standard.
4. The adaptive snow and rain removal method based on feature-driven GAN according to claim 1, wherein: In the step S3, the multi-stage collaborative optimization dual-path inference framework: S3.
1. The backbone network adopts a dual-encoder architecture. The SnowPath encoder focuses on extracting high-frequency information in the wavelet domain and is equipped with Haar wavelet initialization convolution kernels. The RainPath encoder strengthens the spatial attention mechanism and is initialized with an anisotropic Gabor filter bank. The dual-path features are fused through an adaptive gating network in the decoding stage, and the input of the gating network includes meteorological sensor data (precipitation type, intensity, etc. metadata); S3.
2. Establish a dynamic hyperparameter optimization space, perform periodic parameter tuning through a DWave quantum annealing machine, encode key parameters such as the learning rate and decay coefficient into 512-bit quantum states, and automatically search for the optimal combination on the ImageDesnow-2025 benchmark dataset. The optimization period is set to restart every 50 epochs; S3.
3. Develop a physical-level snowflake motion simulation system based on the NS equation, control the number of particles at the order of 1e6, establish a database of 7 types of snow crystals (plate-like, columnar, etc.), generate multi-angle scattering effects by combining real atmospheric refractive index parameters, support dynamically adjusting the precipitation rate (0 - 50mm / h), wind speed (0 - 20m / s), and ambient temperature (-30℃~+5℃), and generate 10240 hours of equivalent training samples; S3.
4. Customize dedicated operators for the NVIDIA Ampere architecture, adopt a mixed-precision calculation process deeply optimized by cuDNN 8.5, design a model parallel strategy based on Triton, split the generator into 3 segments and deploy it on a 4-GPU cluster. The overall inference speed reaches 128 FPS@4K resolution, and the video memory occupancy is stable within 18GB.
5. The adaptive snow and rain removal method based on feature-driven GAN according to claim 1, wherein: In the above step S4, in the construction of the industrial-level trusted evaluation system: S4.
1. Define the comprehensive index of Snow-Removal-Score (SRS): SRS = 0.4PSNR + 0.3SSIM + 0.2LPIPS + 0.1HumanScore Construct a test set containing 10,000 professional images, covering extreme scenarios (such as blizzard weather, low-light environment, etc.), introduce a frequency-band evaluation mechanism, and calculate the recovery ability indicators at high frequency (30 lp / mm), medium frequency (10 - 30 lp / mm), and low frequency (<10 lp / mm) respectively; S4.
2. Develop the Xilinx Versal ACAP evaluation kit, integrate 4-channel 4K@60fps video input, deploy dedicated preprocessing IP cores to implement pipeline operations such as Bayer array decoding and white balance correction, and form a hardware closed-loop verification system with the main model, with a delay jitter <0.5ms; S4.
3. Design an adversarial attack test set, including 100 adversarial samples based on the FGSM and PGD methods, establish a degradation parameter perturbation matrix, and conduct Monte Carlo sampling tests on key parameters such as snowfall density (±30%) and snowflake scale (±50%) to ensure that the model performance fluctuates within ±3dB PSNR; S4.
4. Integrate lidar point cloud assisted evaluation, obtain the real geometric information of the scene through a ToF depth camera, construct three-dimensional reconstruction quality indicators: the point cloud coincidence degree ≥92%, the edge sharpness loss ≤15%, and at the same time use an infrared thermal imager to verify the physical consistency, with the requirement that the temperature distribution variance ≤0.3°C.
6. The adaptive snow and rain removal method based on feature-driven GAN according to claim 3, wherein, In the above step S2, in the spatially adaptive hierarchical constraint loss system, the deraining network includes two generators: the first generator Gd is used to remove raindrops and rain streaks from the input rainy image Xr to generate an intermediate image The second generator Gr further corrects the color deviations or artifacts in the image generated by Gd to generate the final derained image The two generators share the same discriminator D, and through collaborative optimization, they generate more realistic images, thus jointly "fooling" the discriminator. The objective function LGd of generator Gd includes LG which is the same as the formula, and a newly introduced loss function which is used to guide the generator to generate better derained images. Considering the unique shapes and appearance features of raindrops and rain streaks Based on the gradient magnitude map design of the rain-free image C and the generated image it is calculated as follows: Similarly, the objective function LGr of the generator Gr consists of the basic loss LG and the loss LSSIM based on the structural similarity index (SSIM), and its definition is as follows: Among them, C and respectively represent the real rain-free image and the enhanced image generated by the generator Gr. μ and σ respectively represent the mean and variance of the image, and c1 and c2 are constants. Therefore, the objective function of the generator Gr is defined as: LGr = LG + LSSIM For real snow image generation, assume the video tensor A ∈ Rm×n×k, where m represents the number of rows, n represents the number of columns, and k represents the number of frames. According to the selected dimension, three slicing modes can be extracted: horizontal slice Ai:: ∈ Rk×n, side slice A:j: ∈ Rk×m, and front slice A::k ∈ Rm×n. Considering the spatio-temporal characteristics of A:j and A::k, low-rank and sparse decomposition of one of them can achieve the separation of the background layer and the snowflake layer. Specifically, the SVD of the spatio-temporal slice Ai:: is expressed as: Ai:: = μ1σ1V1T +... + μdσdVdT Time filtering and snow removal. By applying appropriate time filtering to the foreground slice, the snowflakes represented by short-term motion patterns can be removed. Specifically, the filtering operation performs time filtering on the left singular vectors representing motion information: μl' = f(μl) for l = 2, 3,..., q AGT = B + Fdesnowed + N。
Citation Information
Cited By
Mixed shooting image correction method and system based on deep learning
CN120765486A
A deep learning-based mixed image correction method and system
CN120765486B
Novel intelligent oil and gas reservoir numerical simulation method based on physical mechanism driving
CN121835402A