Semantic guidance-based two-stage raindrop removal method
Through semantic guidance of the two-stage raindrop removal method, combined with semantic guidance and background recovery modules, the problems of raindrop residues and background blur in multi-scene raindrop removal are solved, achieving more efficient raindrop removal and background recovery effects.
Patent Information
- Application Number
- CN202510431959.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art has problems of raindrop residues, artifacts and background blur in multi-scene raindrop removal, especially in domain offsets between day and night scenes and densely distributed raindrop removal effects.
The semantic guidance dual-stage raindrop removal method is adopted, and the semantic guidance module is introduced to fine repairs combined with image features, and a background recovery module is added to the network output. The multi-frame median fusion and semi-supervised fine-tuning strategies of the same scene are used to enhance the performance of the model on unlabeled data.
Raindrop artifacts and residues are significantly reduced, raindrop removal effects are improved in multiple scenarios, especially between day and night scenes, and clear background details are restored.
Smart Images

Figure CN120278900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to the technical field of a semantic-guided two-stage rain drop removal method. Background Art
[0002] Image de-raining is an important research direction in the field of low-level vision, aiming to solve problems such as image blurring and detail loss caused by weather interferences such as rain drops and rain streaks. In a rainy environment, the dynamics of rain lines (such as density and direction differences) and the optical interaction between rain drops and the background (such as refraction and reflection) will significantly reduce the image quality and affect the reliability of high-level vision tasks such as autonomous driving and video surveillance.
[0003] To solve the above problems, Garg et al. [1] were the first to systematically analyze the visual impact of rainfall on imaging systems. They developed a physical model that combines rain drop dynamics with a lighting model, and on this basis, proposed an effective video rain drop detection and removal algorithm, which showed excellent performance in complex dynamic scenes. Subsequently, Kang et al. [2] proposed a method for simultaneously detecting and removing rain drops from a single image. By combining bilateral filtering with sparse coding, they regarded rain drop removal as an image decomposition problem under morphological component analysis, and achieved effective rain drop removal while retaining image details.
[0004] Qian et al. [3] were the first to apply a generative adversarial network (GAN) to rain drop removal. By introducing an attention mechanism into the generator and the image processor, they achieved precise restoration of the areas occluded by rain drops, thus significantly improving the image quality. Chai et al. [4] proposed a rain drop removal method called MARR-GAN, which integrates the attention mechanism with a recursive residual structure and uses memristor-based hardware acceleration to implement a software-hardware co-design solution tailored for edge devices. Although the existing methods have made progress, the current GAN-based rain drop removal models still face limitations, such as difficulty in capturing fine-grained texture details in severely degraded areas and a tendency to produce over-smoothed results in the background areas.
[0005] Transformer [5] has gradually emerged in the field of computer vision due to its powerful ability to capture global dependencies and has begun to be applied to image inpainting tasks, including raindrop removal. Zamir et al. [6] proposed an efficient Transformer-based model called Restormer, which enhanced the multi-head self-attention mechanism and feed-forward network modules to effectively capture long-range pixel interactions in high-resolution images, thus achieving state-of-the-art performance in various inpainting tasks. Fu et al. [7] introduced an Uncertainty-aware Sparse Transformer Network (USTN), which combines multi-scale sparse feature extraction, dynamic top-k attention mechanism, and uncertainty-guided dual decoder architecture, enabling precise removal of raindrops and rain streaks.
[0006] Frequency domain transformation is a mathematical technique that can map a signal or image from the time domain (or spatial domain) to the frequency domain. In the task of image raindrop removal, raindrops usually exhibit characteristics such as strong locality, semi-transparency, and blurred boundaries, which can severely damage the structural information of the image. Frequency domain analysis can effectively capture the unique spectral anomalies associated with the raindrop regions, providing key clues for accurate restoration of occluded areas. Yang et al. [8] proposed an image de-raining method that combines wavelet transform with a recursive structure, effectively addressing the challenges brought by heavy rain streaks and accumulated rain or fog. This method significantly improves the de-raining performance in real-world scenarios while preserving image details. Huang et al. [9] proposed a single-image de-raining method based on a selective wavelet attention mechanism, which jointly utilizes spatial and frequency domain information to guide the separation of rain streaks from the background. Their method improves the accuracy and visual quality of de-raining while maintaining background details. Yang et al.
[10] further developed a convolutional neural network that combines wavelet transform and channel attention mechanism. By replacing traditional upsampling and downsampling operations with wavelet decomposition and fusing multi-frequency features, this method achieves high-quality single-image de-raining and outperforms existing methods on both synthetic and real-world datasets. Gao et al.
[11] proposed FADformer, an effective Transformer-based de-raining framework that incorporates frequency domain features. By introducing frequency domain convolution and contrastive learning, this model balances global modeling ability and computational efficiency. It effectively utilizes the rain streak information from negative samples, thus significantly enhancing the de-raining performance of single images.
[0007] However, the above methods mainly focus on single-scene daytime raindrop images and perform poorly in multi-scene raindrop removal. Common problems include raindrop residue, raindrop artifacts, and background blur. Raindrop Clarity is a large multi-scene raindrop removal dataset that opens up a new research direction for this task. The dataset contains images from four different scenes. This dataset introduces three major challenges: Domain shift between daytime and nighttime scenes: Previous raindrop removal datasets only contain daytime scenes. In contrast, this dataset contains a large number of nighttime images. Due to the inherent domain gap between daytime and nighttime scenes, the model needs to be highly robust to achieve effective raindrop removal under both conditions. Blurred background in raindrop-centered images: When the camera focuses on raindrops, the background naturally becomes blurred. Therefore, even if the raindrops are successfully removed, the model may still not be able to restore a clear and clean background. Densely distributed and slender raindrops: Some images contain extremely dense raindrop occlusions where the raindrops are slender and overlap with each other. This makes it difficult for the model to completely eliminate all raindrop residues and artifacts, resulting in visual distortions still present in the output image.
[0008] References:
[0009] [1] Garg K, Nayar S K. Detection and removal of rain from videos[C] / / Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004. IEEE, 2004, 1: I-I.
[0010] [2] Kang L W, Lin C W, Fu Y H. Automatic single-image-based rain streaks removal via image decomposition[J]. IEEE transactions on image processing, 2011, 21(4): 1742 - 1755.
[0011] [3]Qian R,Tan R T,Yang W,et al.Attentive generative adversarial network for raindrop removal from a single image[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition.2018:2482-2491.
[0012] [4]Chai Q,Liu Y.MARR-GAN:Memristive Attention Recurrent Residual Generative Adversarial Network for Raindrop Removal[J].Micromachines,2024,15(2):217.
[0013] [5]Vaswani A,Shazeer N,Parmar N,et al.Attention is all you need[J].Advances in neural information processing systems,2017,30.
[0014] [6]Zamir S W,Arora A,Khan S,et al.Restormer:Efficient transformer for high-resolution image restoration[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition.2022:5728-5739.
[0015] [7]Fu B,Jiang Y,Wang D,et al.Uncertainty-aware sparse transformer network for single image deraindrop[J].IEEE Transactions on Instrumentation and Measurement,2024.
[0016] [8]Yang W,Liu J,Yang S,et al.Scale-free single image deraining via visibility-enhanced recurrent wavelet learning[J].IEEE Transactions on Image Processing,2019,28(6):2948-2961.
[0017] [9]Huang H,Yu A,Chai Z,et al.Selective wavelet attention learning for single image deraining[J].International Journal of Computer Vision,2021,129(4):1282-1300.
[0018]
[10] Yang H H,Yang C H H,Wang Y C F.Wavelet channel attention module with a fusion network for single image deraining[C] / / 2020 IEEE International Conference on Image Processing(ICIP).IEEE,2020:883-887.
[0019]
[11] Gao N,Jiang X,Zhang X,et al.Efficient Frequency-Domain Image Deraining with Contrastive Regularization[C] / / European Conference on Computer Vision.Cham:Springer Nature Switzerland,2024:240-257. Summary of the Invention
[0020] To solve the above technical problems, the present invention proposes a semantic-guided two-stage raindrop removal method, which introduces a semantic guidance module that combines semantic information with image features to guide the model for more refined restoration. A background restoration module is added at the network output end, which performs secondary restoration to repair blurred background regions. In different frames of the same scene, by applying median fusion to multiple frames from the same scene, this strategy significantly reduces raindrop artifacts and residues. The fused image is regarded as pseudo ground truth to implicitly expand the training dataset and fine-tune the model in a semi-supervised manner, thus greatly improving its performance on unlabeled data.
[0021] The present invention proposes a semantic-guided two-stage raindrop removal method, including preparing the training set and test set of the RaindropClarity dataset, and further including the following steps:
[0022] Step 1: Use the images in the training set to train the STRRNet network to generate a training model;
[0023] Step 2: Save the training model to a local folder, use the images in the test set to test the effect of the training model, and if satisfied, save the training model as a satisfactory training model;
[0024] Step 3: Save the satisfactory training model to a local folder and use the satisfactory training model to test rainy images.
[0025] Preferably, the STRRNet network includes a Transformer encoder module, a semantic guidance module, a Transformer decoder module, and a background restoration module.
[0026] In any of the above solutions, preferably, the encoder module is used to extract the feature map of the blurred image and lift the feature map to a high-dimensional space to process the feature information.
[0027] In any of the above solutions, preferably, the encoder and decoder are composed of multiple Transformer blocks, and each block contains a multi-head transposed attention module and a gated feed-forward network.
[0028] In any of the above solutions, preferably, the semantic guidance module combines semantic information with image features to achieve more refined restoration.
[0029] In any of the above solutions, preferably, the background restoration module aims to perform secondary restoration on the output image to reconstruct the blurred background.
[0030] Preferably, in any of the above solutions, the training strategy is to train a pre-trained model on the original training data set, perform median fusion on multiple frames from the same scene to obtain a median fusion image, regard the median fusion image as the pseudo ground truth, and use it in the semi-supervised fine-tuning stage to enhance the raindrop removal ability of the model on unlabeled images.
[0031] Preferably, in any of the above solutions, the semantic guidance module has two stages: in the first stage, the cross-entropy is used to train the text embedder to align semantics with image features; in the second stage, the text embedder is frozen and the query matrix Q is inferred, and the self-attention is combined with the image features K / V to partition the scene feature space.
[0032] Preferably, in any of the above solutions, the decoder module is used to restore the clear image.
[0033] Preferably, in any of the above solutions, the decoder module is further used to restore the detail information of the rain-free image and the low-dimensional rain-free image according to the high-dimensional feature map.
[0034] Preferably, in any of the above solutions, the output image sequentially passes through a 3×3 convolution, 3 enhanced residual pixel-level attention blocks and a 3×3 convolution, and is added to the output of the encoder to obtain the finally restored image.
[0035] The present invention proposes a semantic-guided two-stage raindrop removal method, which can effectively retain detailed information and restore high-quality rain-free images.
[0036] STRRNet refers to the Semantics-guided Two-stage Raindrop Removal Network, that is, the semantic-guided two-stage raindrop removal network. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of a preferred embodiment of the semantic-guided two-stage raindrop removal method according to the present invention.
[0038] Figure 2 It is a schematic diagram of the overall network structure of an embodiment of the STRRNet network of the semantic-guided two-stage raindrop removal method according to the present invention.
[0039] Figure 3 It is a schematic diagram of the structure of the semantic guidance module of the STRRNet network of the semantic-guided two-stage raindrop removal method according to the present invention.
[0040] Figure 4 It is a schematic diagram of the training strategy of an embodiment of the STRRNet network of the semantic-guided two-stage raindrop removal method according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0041] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0042] Embodiment 1
[0043] As Figure 1 shown, perform step 100 to prepare the training set and test set of the Raindrop Clarity dataset.
[0044] Perform step 110 to train the STRRNet network using the images in the training set to generate a training model.
[0045] The STRRNet network includes a Transformer encoder module, a semantic guidance module, a Transformer decoder module, and a background restoration module.
[0046] Both the Transformer encoder and the Transformer decoder are composed of multiple Transformer blocks, and each block contains a multi-head transposed attention MDTA module and a gated feed-forward network GDFN. The multi-head transposed attention MDTA applies self-attention in the channel dimension. In the gated feed-forward network GDFN, depthwise separable convolutions are also used to encode information from spatially adjacent pixel positions, which helps to learn the local image structure for effective inpainting.
[0047] The semantic guidance module is divided into two stages: in the first stage, the text embedder is trained with cross-entropy to align semantics with image features; in the second stage, the text embedder is frozen and the query matrix Q is inferred, and self-attention is calculated by combining the image features K / V to partition the scene feature space.
[0048] The background restoration module performs secondary restoration on the output image, that is, the output image passes through a 3×3 convolution, 3 enhanced residual pixel-level attention blocks, and a 3×3 convolution in sequence, and the final restored image is obtained by adding the output image.
[0049] Perform step 120 to save the training model to a local folder, and use the images in the test set to test the effect of the training model. If satisfied, save the training model as a satisfactory training model.
[0050] Perform step 130 to save the satisfactory training model to a local folder, and use the satisfactory training model to test rainy images.
[0051] Embodiment 2
[0052] The present invention is a semantics-guided two-stage raindrop removal network, a trainable end-to-end rain removal network for multi-scene raindrop removal, named "Semantics-guided Two-stage Raindrop Removal Network" (STRRNet).
[0053] The content of STRRNet includes: a Transformer encoder module, a semantics guidance module, a Transformer decoder module, and a background restoration module.
[0054] The training of STRRNet requires a labeled pair dataset of "rainy - rainless" images to train a training model. By bringing the network hyperparameters in the trained model into the network model, the rain removal function for rainy images is realized, achieving the rain removal effect. It effectively preserves detailed information and restores high-quality rain-removed images.
[0055] Usage steps of STRRNet (implemented using the Python programming language):
[0056] 1. Prepare the training set and test set of the Raindrop Clarity dataset, and set the image format to ".jpg" or ".png";
[0057] 2. Use the training file "train.py" to start the training of the network, and adjust parameters such as the batch size and lr of the network according to requirements.
[0058] 3. Save the trained model to a local folder, use the "vali.py" file and the test set to test the effect of the trained model of the network. If not satisfied, the "checkpoint" technology can be used to continue training to obtain a satisfactory model;
[0059] 4. Save the satisfactory trained model to a local folder, and use "test.py" to test rainy images.
[0060] Example 3
[0061] Multi-scene raindrop removal includes four scenarios: daytime background focused images, daytime raindrop focused images, nighttime background focused images, and nighttime raindrop focused images. This task presents three main challenges: how to achieve effective raindrop removal under daytime and nighttime conditions, how to restore the blurred background while removing raindrops, and how to handle densely distributed raindrops. To address these challenges, a semantics-guided two-stage raindrop removal network, called STRRNet, is proposed.
[0062] As Figure 2As shown, the overall network architecture is an architecture improved based on Restormer. However, different from these architectures, both the Transformer encoder and the Transformer decoder are composed of multiple Transformer blocks, each block contains a multi-head transposed attention module and a gated feed-forward network, a semantic guidance module is embedded at the connection between the encoder and decoder modules, and a background restoration module is used after the decoder. The STRRNet network includes a Transformer encoder module, a semantic guidance module, a Transformer decoder module, and a background restoration module. The Transformer encoder-decoder improved based on Restormer serves as the backbone network, responsible for global and local feature modeling. The semantic guidance module is used to integrate scene semantic information and enhance the robustness of the model to different scenes (day / night, background / raindrop focus). The background restoration module is used to perform secondary restoration on the image after preliminary rain removal to solve the problem of background blur caused by raindrop focus. The training strategy uses temporal redundancy information to eliminate rain residue and expand training data.
[0063] The working of each module is as follows:
[0064] 1. The Transformer encoder module (as Figure 2 shown)
[0065] The Transformer encoder-decoder is one of the core components of STRRNet, used to extract and reconstruct image features. The encoder part is composed of multiple Transformer Blocks, and each Block contains two key modules: multi-Dconv head transposed attention (MDTA) and Gated-DConv feed-forward network (GDFN).
[0066] The MDTA module can capture local and global information of image features through multi-dimensional sparse feature extraction and transposed attention mechanism, thereby enhancing the model's perception ability of complex scenes. The GDFN module further enhances the feature expression ability through a gated mechanism and dilated convolution, while reducing information loss.
[0067] The decoder part is also composed of multiple Transformer Blocks, and its role is to gradually reconstruct the features extracted by the encoder into a clear image. Through the self-attention mechanism and feature fusion, the decoder can effectively restore background details and remove raindrop artifacts.
[0068] The overall design of the Transformer encoder-decoder aims to address complex issues in multi-scene raindrop removal, such as background blurring and raindrop residue. By combining a semantic guidance module and a background restoration module, this architecture can achieve refined restoration for different scenes while enhancing the model's robustness in day and night scenarios.
[0069] 2. Semantic Guidance Module (as Figure 3 shown)
[0070] The semantic guidance module combines semantic information with image features to achieve more refined restoration. As Figure 3 shown, the semantic guidance module consists of two stages. In the first stage, a text embedder is trained using cross-entropy loss to associate specific semantic descriptions with corresponding image features. Specifically, images from four different scenes are first input into a lightweight image feature extraction network to obtain image embeddings. We first downsample the input images and then use ResNet-18 to extract image features. These features are then passed through two linear layers and then through a pooling layer to obtain the final image embeddings. At the same time, the corresponding scene descriptors are processed by the text embedder to generate text embeddings. The parameters of the text embedder are optimized by minimizing the cross-entropy loss between the image embeddings and the text embeddings. In the second stage, the text embedder is frozen and used to infer the scene description query matrix Q for each image. Matrices K and V represent the key and value features extracted from the image. By calculating the self-attention between the semantic information and the image features, the image feature space can be effectively partitioned according to different scene categories, enabling the decoder to perform more refined and context-aware restoration.
[0071] 3. Background Restoration Module (as Figure 2 shown)
[0072] In multi-scene raindrop removal, some images contain focused raindrops, resulting in background blurring. Even if the model successfully removes the raindrops, it may still not be able to restore a clear background. To address this issue, a background restoration module is used in the network to perform a secondary restoration on the output image to reconstruct the blurred background. The background restoration module consists of three ERPAB blocks and two 3×3 convolutional layers. The Enhanced Residual Pixel-level Attention Block (ERPAB) was originally proposed in, and it consists of a Parallel Dilated Fusion Block (PDFB) and a Pixel-level Attention Block (PAB). The PDFB contains three parallel dilated convolutional layers for feature extraction and a convolutional layer for feature aggregation.
[0073] 4. Training Strategy (as Figure 4 shown)
[0074] First, a pre-trained model is trained on the original training dataset. Then, this model is used to perform inference on all frames within the same scene in the test set. The inferred images may still contain residual raindrops and artifacts. Since the background remains consistent across different time frames while the positions of raindrops change over time, this motivates us to perform median fusion on multiple frames from the same scene to obtain a median-fused image. Due to the dynamic nature of raindrop artifacts, their positions are inconsistent across different frames, which usually prevents them from appearing in the median, while the stable background information is retained. Then, we treat the median-fused image as pseudo-ground truth and use it in the semi-supervised fine-tuning stage to enhance the model's raindrop removal ability on unlabeled images.
[0075] To better understand the present invention, the above has been described in detail in conjunction with specific embodiments of the present invention, but it is not a limitation of the present invention. Any simple modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention. Each embodiment in this specification focuses on the differences from other embodiments, and the same or similar parts among the various embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
Claims
1. A semantic-guided two-stage raindrop removal method, including preparing the training set and test set of the Raindrop Clarity dataset, characterized in that, It also includes the following steps: Step 1: Use the images in the training set to train the STRRNet network to generate a training model; Step 2: Save the training model to a local folder, use the images in the test set to test the effect of the training model, and if satisfied, save the training model as a satisfactory training model; Step 3: Save the satisfactory training model to a local folder and use the satisfactory training model to test rainy images.
2. The semantic-guided two-stage raindrop removal method according to claim 1, wherein The STRRNet network includes a Transformer encoder module, a semantic guidance module, a Transformer decoder module, and a background restoration module.
3. The semantic-guided two-stage raindrop removal method according to claim 2, wherein, The encoder and decoder are composed of multiple Transformer blocks, and each block contains a multi-head transposed attention module and a gated feed-forward network.
4. The semantic guidance-based two-stage raindrop removal method according to claim 3, wherein The semantic guidance module combines semantic information with image features to achieve more refined restoration.
5. The semantic guidance-based two-stage raindrop removal method according to claim 4, wherein The background restoration module aims to perform secondary restoration on the output image to reconstruct the blurred background.
6. The semantic-guided two-stage raindrop removal method according to claim 5, wherein The training strategy is to train a pre-trained model on the original training dataset, perform median fusion on multiple frames from the same scene to obtain a median fusion image, regard the median fusion image as pseudo-ground truth, and use it in the semi-supervised fine-tuning stage to enhance the model's rain drop removal ability on unlabeled images.
7. The semantic-guided two-stage raindrop removal method according to claim 6, wherein, The semantic guidance module has two stages: in the first stage, the text embedder is trained with cross-entropy to align semantics with image features; In the second stage, the text embedder is frozen and the query matrix Q is inferred, and the self-attention is combined with the image features K / V to partition the scene feature space.
8. The semantic-guided two-stage raindrop removal method according to claim 7, wherein The output image sequentially passes through a 3×3 convolution, 3 enhanced residual pixel-level attention blocks, and a 3×3 convolution, and the output of the encoder is added to obtain the finally restored image.