A method for removing rain from traffic images based on reference images

By constructing a deep learning model based on reference images and combining feature alignment matching strategies, the problem of insufficient rain removal methods in complex scenarios is solved, and the effect of efficiently removing rain line noise and retaining image details is achieved.

CN119600299BActive Publication Date: 2025-07-22TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411710821.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-27
Publication Date
2025-07-22
Estimated Expiration
2044-11-27

AI Technical Summary

Technical Problem

Existing traffic image rain removal methods require accurate rain line models or a large number of computing resources, and it is difficult to effectively remove rain line noise in complex traffic scenarios while maintaining the details and edge features of the image.

Method used

Using a reference image-based rain removal method, a deep learning model including an encoder, a multi-scale feature matching module and a decoder is constructed, combined with a feature alignment matching strategy, features are extracted using the PyTorch framework and the pre-trained VGG16 model, and the reference image with the highest correlation coefficient is selected for rain removal.

Benefits of technology

It significantly improves the reconstruction quality of traffic images, can better preserve image details and edge features, improves the accuracy and efficiency of rain removal effects, and is suitable for rain removal tasks of multiple types of traffic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600299B_ABST
    Figure CN119600299B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and particularly relates to a method for removing rain from traffic images based on reference images, comprising the following steps: selecting and establishing a data set; preprocessing the data set; constructing a rain removal model, where the backbone network extracts rich features by stacking multiple encoder modules, multi-scale feature matching modules, and decoder modules; training the rain removal model using the training set in the target data set, and obtaining a trained rain removal model after verifying that the trained rain removal model is qualified through the test set; inputting the traffic rain image to be processed into the trained rain removal model for background image reconstruction, and the rain removal model outputs the reconstructed traffic road image. The present invention constructs a new network model for natural image rain removal research. This network model focuses on feature extraction, aiming to accurately capture the latent features in the rain image and its corresponding reference image, ensuring that more long-distance dependence information can be captured to enhance the restoration accuracy of traffic background images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for removing rain from traffic images based on a reference image. Background Art

[0002] With the rapid development of the transportation industry, intelligent transportation systems have gradually become an important means to improve traffic efficiency and ensure traffic safety. However, under adverse weather conditions such as rainy days, the quality of traffic images will be severely affected, and rain line noise will interfere with the clarity of the images, thereby affecting the performance of intelligent transportation systems. Therefore, how to effectively remove rain line noise from traffic images and improve the clarity of the images has become an urgent problem to be solved in the transportation field. Clear traffic images not only help intelligent transportation systems to more accurately identify traffic elements such as vehicles and pedestrians, improve the efficiency of traffic management, but also provide more powerful guarantees for traffic safety. However, due to the complexity and diversity of traffic scenes, as well as the randomness and uncertainty of rain line noise, removing rain from traffic images faces many challenges. How to maintain the details and edge features of the images while ensuring the rain removal effect has become the key to traffic image rain removal technology.

[0003] Currently, image rain removal methods mainly include methods based on physical models, image processing knowledge, sparse coding dictionary learning and classifiers, and deep convolutional neural networks. These methods can, to a certain extent, remove rain line noise from traffic images and improve the clarity of the images. However, in the face of complex traffic scenes and diverse rain line noise, existing methods still have some limitations. For example, methods based on physical models require accurate rain line models, while methods based on deep learning require a large amount of training data and computing resources. In view of the limitations of existing methods, it is particularly important to propose a new image rain removal method. Summary of the Invention

[0004] In view of the technical problems that the methods based on physical models in the above existing methods require accurate rain line models, while the methods based on deep learning require a large amount of training data and computing resources, the present invention provides a method for removing rain from traffic images based on a reference image.

[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0006] A method for removing rain from traffic images based on a reference image, comprising the following steps:

[0007] S1. Select and establish a data set;

[0008] S2. Preprocess the data set;

[0009] S3. Build a rain removal model. The backbone network extracts rich features by stacking multiple encoder modules, multi-scale feature matching modules, and decoder modules. These features can capture the spatially varying rainfall distribution. Train the rain removal model using the training set in the target dataset, and after verifying that the trained rain removal model is qualified through the test set, obtain the trained rain removal model.

[0010] S4. Input the traffic rain map to be processed into the trained rain removal model for background map reconstruction, and the rain removal model outputs the reconstructed traffic road image.

[0011] The method of selecting and establishing the dataset in S1 is: select and collect diverse traffic images to construct the initial dataset of the rain removal model, and divide the initial dataset into a training set and a test set. The initial dataset includes multiple rain scene images of traffic roads.

[0012] The method of preprocessing the dataset in S2 is: using the PyTorch framework, load the pre-trained VGG16 model and intercept the first 29 layers of the model as the feature extractor. These layers contain most of the convolutional operations of the model and can effectively extract the high-level features of the image. For each traffic road rain map and the traffic road image under normal weather, after preprocessing, input them into the VGG16 feature extractor to obtain the corresponding feature vectors. The preprocessing steps include resizing the image, central cropping, converting the image to a tensor, and normalizing. These feature vectors reflect the high-level visual information of the image. Use the Pearson correlation coefficient to measure the similarity between the features of the rain map and the features of the normal image. For each input rain map, traverse all normal images in the training set except the ground truth, calculate the correlation coefficient between its feature vector and the feature vector of the input rain map, and select the target image with the highest correlation coefficient as the reference image.

[0013] The training set of the target dataset in S3 includes multiple traffic road rain pattern images and the traffic road images corresponding to each traffic road rain pattern image under a clean background. The test set of the target dataset is the same as the test set of the initial dataset, both including multiple traffic road rain pattern images.

[0014] The method of building the rain removal model in S3 is:

[0015] S3.1. Image Encoding: The encoder module is composed of multiple Transformer blocks connected in series. Each Transformer module integrates a multi-head self-attention mechanism and a feed-forward neural network to effectively capture the global dependencies in the image. In the initial stage, the image undergoes preliminary feature extraction through L1 Transformer blocks, followed by downsampling to effectively reduce the data dimension while extracting key features. These downsampled feature maps are further subjected to deep feature extraction and more refined downsampling through L2 Transformer blocks, which significantly enhances the model's ability to understand the image content.

[0016] S3.2. Image Decoding and Reconstruction: The decoder module performs advanced feature extraction through L2 Transformer blocks, followed by an upsampling process and then enters L1 Transformer blocks for further refinement. During the downsampling stage of the entire model, a feature alignment and matching module is innovatively added between the encoder module and the decoder module. Through residual connections, the output of the encoder module and the output of the decoder module are concatenated to fuse local features and aligned features. The recursive numbers L1, L2, and L3 of the Transformer blocks that achieve the best rain removal performance are set to 2, 3, and 3 respectively.

[0017] S3.3. Feature Alignment and Matching Module: A two-stage matching strategy is adopted, combining a method based on a reference image with a feature alignment and accelerated matching strategy, aiming to overcome the problem of image detail loss caused by traditional rain removal algorithms that only consider single-image rain removal. The feature alignment and matching module can locate and extract feature blocks that are crucial for rain removal, thereby optimizing the rain removal effect. The rain image features are unfolded into multiple non-overlapping blocks, and the most similar reference block is found for each block in the rain image. To improve accuracy, local block matching is further performed for each pair of rain blocks and reference blocks. Finally, useful reference features are extracted based on the obtained corresponding information.

[0018] The method of unfolding the rain image features into multiple non-overlapping blocks and finding the most similar reference block for each block in the rain image in S3.3 is as follows:

[0019] S3.3.1. In the first stage, first unfold the rain image features into N non-overlapping blocks, each block consisting of multiple w×h local blocks. For each rain block, find its most relevant reference block. First, take out the central local block of the rain block and calculate the cosine similarity with each local block in the reference block. According to the similarity score, find the local block in the reference image features that is most similar to the central local block of the rain block. Crop out a block of size w×h from the surrounding area as the most relevant reference block. After the first-stage matching, each rain block obtains the most relevant reference block.

[0020] S3.3.2. In the second stage, perform local block matching on each rainy block and the reference block to obtain their index maps and similarity maps; calculate the similarity scores between each local block in the rainy block and each local block in the reference block; extract relevant blocks according to the index map, multiply these blocks with the corresponding similarity maps to obtain weighted features; finally, merge these N features to obtain the reference feature.

[0021] In S3, it also includes using Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM) to evaluate the quality of the reconstruction result; the larger the PSNR, the better the image quality; SSIM is a number between 0 and 1, and the larger it is, the smaller the gap between the output image and the distortion-free image, that is, the better the image quality.

[0022] The beneficial effects of the present invention compared with the prior art are as follows:

[0023] The present invention constructs a new network model for natural image deraining research. This network model focuses on feature extraction, aiming to accurately capture the potential features in the rainy image and its corresponding reference image, ensuring that more long-distance dependence information can be captured to enhance the restoration accuracy of traffic background images. In addition, the present invention also integrates a feature alignment and matching strategy, which effectively incorporates the visual information of multiple traffic road images under normal weather conditions as reference images into the traffic rain pattern image, further refining the potential features of the image and achieving efficient style transfer. This strategy enables the traffic rain image to fully absorb the valuable guiding information of the rain-free reference image, and then reconstructs a derained image with richer details and better effects. Through the processing of this deraining model, the quality of the reconstructed image is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.

[0025] The structures, ratios, sizes, etc. depicted in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have technical essence. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.

[0026] Figure 1Flowchart of the traffic image de-raining method based on a reference image provided by the present invention.

[0027] Figure 2 Schematic diagram of the composition structure of the de-raining model in the present invention.

[0028] Figure 3 Schematic diagram of the composition structure of the encoder in the present invention.

[0029] Figure 4 Schematic diagram of the composition structure of the decoder in the present invention.

[0030] Figure 5 Schematic diagram of the composition structure of the Transformer block in the present invention.

[0031] Figure 6 Schematic diagram of the execution process of the reference image selection strategy in the present invention. Detailed implementation manners

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. These descriptions are only for further explaining the features and advantages of the present invention, rather than limiting the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0033] The following will further describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0034] A traffic image de-raining method based on a reference image provided by an embodiment of the present invention includes the following steps:

[0035] Step 1: Selection and establishment of a data set: Select and collect diverse traffic images to construct an initial data set for the de-raining model, and divide the initial data set into a training set and a test set. The initial data set includes multiple rainy scene images of traffic roads.

[0036] Step 2: Data set preprocessing:

[0037] Using the PyTorch framework, the pre-trained VGG16 model is loaded, and the first 29 layers of the model are intercepted as feature extractors. These layers contain most of the convolution operations of the model and can effectively extract high-level features of the image. For each traffic road rain image and traffic road image in normal weather, after preprocessing, they are input into the VGG16 feature extractor to obtain the corresponding feature vector. The preprocessing steps include resizing the image, center cropping, converting the image to a tensor, and normalizing it. These feature vectors reflect the high-level visual information of the image. The Pearson correlation coefficient is used to measure the similarity between the features of the rain image and the features of the normal image. For each input rain image, all normal images in the training set except the ground truth are traversed, and the correlation coefficient between their feature vectors and the input rain image feature vector is calculated, and the target image with the highest correlation coefficient is selected as the reference image.

[0038] This strategy ensures that the selected reference image is highly consistent with the image block to be processed at the feature level, providing valuable reference information for subsequent rain removal processing. This strategy shows significant advantages in terms of efficiency and accuracy. By integrating the pre-trained deep network for feature extraction and the Pearson correlation coefficient for accurate similarity calculation, we can quickly and accurately lock the reference image that is most similar to the image block to be processed. This innovative measure significantly improves the efficiency of rain removal and is expected to achieve a significant improvement in the quality of the image after rain removal, especially in restoring image details that are severely obscured by raindrops. In addition, this strategy shows wide versatility and can be flexibly applied to rain removal tasks for various types of traffic images, showing strong practical value.

[0039] The training set of the target dataset includes multiple traffic road rain streak images and traffic road images under clean backgrounds corresponding to each traffic road rain streak image. The test set of the target dataset is the same as the test set of the initial dataset, both of which include multiple traffic road rain streak images.

[0040] Step 3: Model construction: Build a rain removal model, such as Figure 2 As shown in the figure, the backbone network extracts rich features by stacking multiple encoder modules, multi-scale feature matching modules and decoder modules, which can capture the spatially varying rainfall distribution. The deraining model is trained using the training set in the target dataset, and the trained deraining model is verified to be qualified by the test set to obtain the trained deraining model.

[0041] Step 3.1: Image encoding: Figure 3As shown, the encoder module is composed of multiple Transformer blocks connected in series. Inside each Transformer module, a multi-head self-attention mechanism and a feed-forward neural network are integrated to effectively capture the global dependencies in the image. In the initial stage, the image undergoes preliminary feature extraction through L1 Transformer blocks, followed by downsampling to effectively reduce the data dimension while extracting key features. These downsampled feature maps are further subjected to deep feature extraction and more refined downsampling through L2 Transformer blocks, which significantly enhances the model's ability to understand the image content.

[0042] Step 3.2: Image decoding and reconstruction: As Figure 4 , Figure 5 shown, the decoder module performs advanced feature extraction through L2 Transformer blocks, followed by an upsampling process and then enters L1 Transformer blocks for further refinement. In the downsampling stage of the entire model, a feature alignment and matching module is innovatively introduced between the encoder module and the decoder module. Through residual connections, the output of the encoder module is concatenated with the output of the decoder module to fuse local features and aligned features. Experimental results have shown that the optimal number of recursive Transformer blocks L1, L2, and L3 for achieving the best rain removal performance are set to 2, 3, and 3 respectively.

[0043] Step 3.3: Feature alignment and matching module: The present invention adopts a two-stage matching strategy. As Figure 6 shown, it innovatively combines the method based on the reference image with a feature alignment and accelerated matching strategy, aiming to overcome the problem of image detail loss caused by traditional rain removal algorithms that only consider single-image rain removal. The feature alignment and matching module can locate and extract the feature blocks that are crucial for rain removal, thereby optimizing the rain removal effect. Specifically, the rain image features are unfolded into multiple non-overlapping blocks, and the most similar reference block is found for each block in the rain image. This significantly reduces the matching calculation cost compared to traditional methods. To improve the accuracy, local block matching is further performed for each pair of rain blocks and reference blocks. Finally, useful reference features are extracted based on the obtained correspondence information.

[0044] Specifically, the rain image features are unfolded into multiple non-overlapping blocks, and the most similar reference block is found for each block in the rain image. This significantly reduces the matching calculation cost compared to traditional methods. To improve the accuracy, local block matching is further performed for each pair of rain blocks and reference blocks. Finally, relevant reference features are extracted based on the matching information.

[0045] Step 3.3.1: In the first stage, first expand the rain image features into N non-overlapping blocks, and each block consists of multiple w×h local blocks. For each rain block, find its most relevant reference block. First, take out the central local block of the rain block and calculate the cosine similarity with each local block in the reference block. According to the similarity score, find the local block in the reference image features that is most similar to the central local block of the rain block. Crop a block of size w×h from the surroundings as the most relevant reference block. After the first-stage matching, each rain block obtains the most relevant reference block.

[0046] Step 3.3.2: In the second stage, perform local block matching on each rain block and the reference block to obtain their index map and similarity map. Calculate the similarity scores between each local block in the rain block and each local block in the reference block. Extract the relevant blocks according to the index map, multiply these blocks with the corresponding similarity map to obtain the weighted features. Finally, merge these N features to obtain the reference feature.

[0047] After training the rain removal model, the embodiments of the present invention also use two metrics, namely peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM), to evaluate the quality of the reconstruction results. The larger the PSNR, the better the image quality. SSIM is a number between 0 and 1, and the larger it is, the smaller the gap between the output image and the distortion-free image, that is, the better the image quality.

[0048] Step 4: Input the traffic rain image to be processed into the trained rain removal model for background image reconstruction, and the rain removal model outputs the reconstructed traffic road image.

[0049] In summary, the present invention provides an innovative rain image reconstruction method. By combining the reference image and the feature alignment and matching strategy, this method significantly improves the processing effect of traffic images under rainy conditions. The present invention uses technical means of deep learning and feature engineering to construct a rain removal model focusing on feature extraction and efficient matching. By introducing the reference image, this model can accurately capture the potential relationship between the rain image and the clean background image, effectively overcoming the deficiency of traditional single-image rain removal algorithms in retaining image details. At the same time, the introduction of the feature alignment and matching module further improves the rain removal accuracy and efficiency of the model, realizing the accurate transfer and efficient utilization of visual information. This method can not only remove rain line noise, but also enhance the detail features of the image during the reconstruction process, generating clearer and more realistic rain-removed images, providing a solid technical support for the stable operation of intelligent transportation systems.

[0050] The above only elaborates on the preferred embodiments of the present invention in detail. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the purpose of the present invention, and all such changes should be included within the protection scope of the present invention.

Claims

1. A method for removing rain from traffic images based on a reference image, characterized in that It includes the following steps: S1. Select and establish a dataset; S2. Preprocess the dataset; S3. Construct a rain removal model. The backbone network extracts rich features by stacking an encoder module, a multi-scale feature matching module, and a decoder module. These features can capture the spatially varying rainfall distribution. Use the training set in the target dataset to train the rain removal model, and after verifying that the trained rain removal model is qualified through the test set, obtain the trained rain removal model; The method for constructing the rain removal model in S3 is as follows: S3.

1. Image encoding: The encoder module is composed of multiple Transformer blocks connected in series. Each Transformer module integrates a multi-head self-attention mechanism and a feed-forward neural network inside to effectively capture the global dependencies in the image; In the initial stage, the image undergoes preliminary feature extraction through L1 Transformer blocks, and then downsampling is performed to effectively reduce the data dimension while extracting key features. These downsampled feature maps are further subjected to deep feature extraction and more refined downsampling through L2 Transformer blocks; S3.

2. Image decoding and reconstruction: The decoder module performs advanced feature extraction through L2 Transformer blocks, then undergoes an upsampling process, and then enters L1 Transformer blocks for further refinement processing. In the downsampling stage of the entire model, a feature alignment and matching module is innovatively added between the encoder module and the decoder module; Through residual connection, the output of the encoder module and the output of the decoder module are concatenated to fuse local features and aligned features. The recursive numbers L1, L2, and L3 of the Transformer blocks that achieve the best rain removal performance are set to 2, 3, and 3 respectively; S3.

3. Feature alignment and matching module: Adopt a two-stage matching strategy, combining a method based on a reference image with a feature alignment and acceleration matching strategy, aiming to overcome the problem of image detail loss caused by traditional rain removal algorithms due to only considering single-image rain removal. The feature alignment and matching module can locate and extract feature blocks that are crucial for rain removal, thereby optimizing the rain removal effect. Expand the rain image features into multiple non-overlapping blocks, and find the most similar reference blocks for each block in the rain image. To improve the accuracy, further perform local block matching in each pair of rain blocks and reference blocks. Finally, extract useful reference features based on the obtained corresponding information; S4. Input the traffic rain image to be processed into the trained rain removal model for background image reconstruction, and the rain removal model outputs the reconstructed traffic road image.

2. The method for removing rain from traffic images based on a reference image according to claim 1, wherein The method for selecting and establishing the dataset in S1 is as follows: Select and collect diverse traffic images to construct the initial dataset for the rain removal model, and divide the initial dataset into a training set and a test set. The initial dataset includes multiple rain scene images of traffic road surfaces.

3. A traffic image de-raining method based on a reference image according to claim 1, wherein, The method for preprocessing the dataset in S2 is as follows: Using the PyTorch framework, load the pre-trained VGG16 model and intercept the first 29 layers of the model as the feature extractor; these layers contain most of the convolutional operations of the model and can effectively extract the high-level features of the image; for each traffic road rain image and traffic road image under normal weather, after preprocessing, input them into the VGG16 feature extractor to obtain the corresponding feature vectors; the preprocessing steps include resizing the image, central cropping, converting the image to a tensor, and performing normalization processing. These feature vectors reflect the high-level visual information of the image; the Pearson correlation coefficient is used to measure the similarity between the rain image features and the normal image features; for each input rain image, traverse all normal images in the training set except the ground truth, calculate the correlation coefficient between its feature vector and the feature vector of the input rain image, and select the target image with the highest correlation coefficient as the reference image.

4. The traffic image de-raining method based on a reference image according to claim 1, wherein, The training set of the target dataset in S3 includes multiple traffic road rain pattern images and the traffic road images corresponding to each traffic road rain pattern image under a clean background; the test set of the target dataset is the same as the test set of the initial dataset and both include multiple traffic road rain pattern images.

5. A traffic image de-raining method based on a reference image according to claim 1, characterized in that, The method in S3.3 for expanding the rain image features into multiple non-overlapping blocks and finding the most similar reference block for each block in the rain image is as follows: S3.3.

1. In the first stage, first expand the rain image features into N non-overlapping blocks, and each block consists of multiple w×h local blocks; for each rain block, find its most relevant reference block; first take out the central local block of the rain block and calculate the cosine similarity with each local block in the reference block; according to the similarity score, find the local block in the reference image features that is most similar to the central local block of the rain block. Crop out a block of size w×h from the surrounding area as the most relevant reference block. After the first stage of matching, each rain block obtains the most relevant reference block. S3.3.

2. In the second stage, perform local block matching on each rain block and the reference block to obtain their index map and similarity map; calculate the similarity score between each local block in the rain block and each local block in the reference block; extract the relevant blocks according to the index map, multiply these blocks by the corresponding similarity map to obtain the weighted features. Finally, merge these N weighted features to obtain the reference features.

6. A traffic image de-raining method based on a reference image according to claim 1, characterized in that S4 also includes using the peak signal-to-noise ratio PSNR and the structural similarity index SSIM to evaluate the quality of the reconstruction result; the larger the PSNR, the better the image quality; SSIM is a number between 0 and 1, and the larger it is, the smaller the gap between the output image and the distortion-free image, that is, the better the image quality.

Citation Information

Patent Citations

  • Real-time video rain removal method based on attention deformation convolution automatic search

    CN112734672A

  • Rain removal method based on high-resolution single image rain removal network

    CN116664445A