An Image De-raining Method and Device Based on Attention Mechanism Encoder-Decoder

By constructing a gradual image degradation database and optimizing network model, combined with the attention mechanism encoding and decoding method, the generalization and computing efficiency of image rain removal methods in the existing technology in real-time systems is solved, and efficient image rain removal effect is achieved.

CN116452460BActive Publication Date: 2025-07-25VOYAH AUTOMOBILE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310457825.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-25
Publication Date
2025-07-25
Estimated Expiration
2043-04-25

AI Technical Summary

Technical Problem

The existing CNN-based image rain removal technology is insufficient in real-time systems, and cannot effectively deal with image degradation caused by different degrees of rainy weather, and has low computing efficiency.

Method used

The image rain removal method based on the attention mechanism is coded. By constructing a stepwise image degradation database, combining the correspondence between the rain removal images and the original images of different degrees, a network model is constructed, and the weight is activated through the Gate MLP layer and the penalty item sparse network to optimize the model to improve computing efficiency.

Benefits of technology

On the premise of ensuring image rain removal effect, effectively reduce the model size, improve computing efficiency, and enhance the generalization and real-time processing capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116452460B_ABST
    Figure CN116452460B_ABST
Patent Text Reader

Abstract

The present application relates to an image de-raining method and device based on an attention mechanism encoder-decoder, and relates to the technical field of image processing. The method includes the following steps: constructing a degradation database based on different real rain-free images and corresponding real light-rain images and real heavy-rain images; constructing a network model based on the degradation database; performing de-raining operations on the real light-rain images or real heavy-rain images based on the network model to obtain corresponding model de-rained images; optimizing the network model based on the model de-rained images and corresponding real rain-free images; and performing de-raining operations on the target image based on the optimized network model. The present application establishes a database based on the progressive image degradation effect, and then constructs a corresponding network model. Combining the correspondence between de-rained images of different degrees and the original images, image de-raining operations are performed. On the premise of ensuring the image de-raining effect, the model size is effectively reduced and the calculation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and in particular to an image de-raining method and device based on an attention mechanism encoder-decoder. Background Art

[0002] Digital images are ubiquitous in daily life and production. At present, research on digital images focuses on image enhancement for image degradation and information loss caused by weather reasons such as rain, fog, and snow.

[0003] Image degradation caused by weather causes visual interference and has a great impact on perception functions such as detection, segmentation, and depth evaluation. In recent years, technologies for de-raining, de-fogging, and de-snowing based on CNNs have developed rapidly and have obvious effects. However, research shows that most methods focus on the current single task or are fine-tuned separately for each task. Although the effects are ideal, they are not general solutions and the generalization is not reflected. Therefore, it is unlikely to be applied in real-time systems.

[0004] Therefore, to meet the digital image processing requirements involving weather factors, an image de-raining technology based on an attention mechanism encoder-decoder is provided. Summary of the Invention

[0005] The present application provides an image de-raining method and device based on an attention mechanism encoder-decoder. A database is established based on the gradual image degradation effect, and then a corresponding network model is constructed. Combining the correspondence between de-rained images of different degrees and the original images, image de-raining operations are performed. On the premise of ensuring the image de-raining effect, the model size is effectively reduced and the calculation efficiency is improved.

[0006] To achieve the above object, the present application provides the following solutions.

[0007] In a first aspect, the present application provides an image de-raining method based on an attention mechanism encoder-decoder, and the method includes the following steps:

[0008] Construct a degradation database based on different real rain-free images and corresponding real light-rain images and real heavy-rain images;

[0009] Construct a network model based on the degradation database;

[0010] Perform de-raining operations on the real light-rain image or the real heavy-rain image based on the network model to obtain a corresponding model de-rained image;

[0011] Optimize the network model based on the model de-rained image and the corresponding real rain-free image;

[0012] Perform de-raining operations on the target image based on the network model that has been optimized.

[0013] Further, in constructing the degradation database based on different real rain - free images and corresponding real light - rain images and real heavy - rain images, the following steps are included;

[0014] Based on the real rain - free images, perform light - rain spraying simulation to obtain the corresponding real light - rain images;

[0015] Based on the real rain - free images, perform heavy - rain spraying simulation to obtain the corresponding real heavy - rain images;

[0016] Associate different real rain - free images with the corresponding real light - rain images and real heavy - rain images, and construct the degradation database.

[0017] Further, in constructing the network model based on the degradation database, the following steps are included:

[0018] Take the real rain - free images as the original - state images, take the real light - rain images as the shallow - degradation images, and take the real heavy - rain images as the deep - degradation images;

[0019] Based on the changes reflected by the original - state images, the corresponding shallow - degradation images, and the corresponding deep - degradation images, construct the network model.

[0020] Further, in constructing the network model based on the changes reflected by the original - state images, the corresponding shallow - degradation images, and the corresponding deep - degradation images, the following steps are included:

[0021] Based on the original - state images, the corresponding shallow - degradation images, and the corresponding deep - degradation images, reverse - infer and simulate the picture changes of the rain - removal operation;

[0022] Based on the picture changes of the rain - removal operation obtained by simulation, construct the network model.

[0023] Further, in optimizing the network model based on the model rain - removed images and the corresponding real rain - free images, the following steps are included:

[0024] Based on the model rain - removed images and the corresponding real rain - free images, obtain the corresponding peak signal - to - noise ratio and the structural similarity between the model rain - removed images and the real rain - free images;

[0025] Based on the peak signal - to - noise ratio and the structural similarity, determine whether the regional operation of the network model meets the preset requirements. If not, optimize the network model.

[0026] Second aspect, the present application provides an image de-raining device based on attention mechanism encoding and decoding, and the device includes:

[0027] A database construction module, which is used to construct a degradation database based on different real rain-free images and corresponding real light rain images and real heavy rain images;

[0028] A model construction module, which is used to construct a network model based on the degradation database;

[0029] A de-raining model module, which is used to perform de-raining operations on the real light rain image or the real heavy rain image based on the network model to obtain a corresponding model de-rained image;

[0030] A model optimization module, which is used to optimize the network model based on the model de-rained image and the corresponding real rain-free image;

[0031] A de-raining execution module, which is used to perform de-raining operations on a target image based on the network model that has completed optimization.

[0032] Further, the database construction module is used to perform light rain spraying simulation based on the real rain-free image to obtain the corresponding real light rain image;

[0033] The database construction module is used to perform heavy rain spraying simulation based on the real rain-free image to obtain the corresponding real heavy rain image;

[0034] The database construction module is used to associate different real rain-free images and the corresponding real light rain images and real heavy rain images, and construct the degradation database.

[0035] Further, the model construction module is used to use the real rain-free image as the original state image, the real light rain image as the shallow degradation image, and the real heavy rain image as the deep degradation image;

[0036] The model construction module is used to construct the network model based on the change situations reflected by the original state image, the corresponding shallow degradation image, and the corresponding deep degradation image.

[0037] Further, the model construction module is used to reverse infer the picture change situation of the simulated de-raining operation based on the original state image, the corresponding shallow degradation image, and the corresponding deep degradation image;

[0038] The model construction module is used to construct the network model based on the picture change situation of the simulated de-raining operation.

[0039] Further, the model optimization module is used to obtain the corresponding peak signal-to-noise ratio and the structural similarity between the model de-rained image and the real rain-free image based on the model de-rained image and the corresponding real rain-free image;

[0040] The model optimization module is used to determine whether the regional operation of the network model meets the preset requirements based on the peak signal-to-noise ratio and the structural similarity, and if not, optimize the network model.

[0041] The beneficial effects brought by the technical solution provided in this application include:

[0042] This application establishes a database based on the gradual image degradation effect, and then constructs a corresponding network model. Combining the correspondence between de-rained images of different degrees and the original images, image de-raining operations are performed. On the premise of ensuring the image de-raining effect, the model size is effectively reduced and the calculation efficiency is improved. Description of the Drawings

[0043] Term Explanation:

[0044] MLP: Multilayer Perceptron, Multilayer Perceptron;

[0045] CNN: Convolutional Neural Networks, Convolutional Neural Networks;

[0046] PSNR: Peak Signal to Noise Ratio, Peak Signal to Noise Ratio

[0047] SSIM: structural similarity index, Structural Similarity Index

[0048] MSA: Multi-head Self-attention Layer, Multi-head Self-attention Layer.

[0049] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0050] Figure 1 It is a flowchart of the steps of the image de-raining method based on attention mechanism encoding and decoding provided in the embodiments of this application;

[0051] Figure 2 It is a principle structure block diagram of the image de-raining method based on attention mechanism encoding and decoding provided in the embodiments of this application;

[0052] Figure 3 This is the principle flowchart of the image de-raining method based on attention mechanism encoding and decoding provided in the embodiments of the present application;

[0053] Figure 4 This is the structural block diagram of the image de-raining device based on attention mechanism encoding and decoding provided in the embodiments of the present application. Detailed implementation manners

[0054] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0055] The following further elaborates on the embodiments of the present application with reference to the accompanying drawings.

[0056] The embodiments of the present application provide an image de-raining method and device based on attention mechanism encoding and decoding. A database is established based on the gradual image degradation effect, and then a corresponding network model is constructed. Combining the correspondence between de-rained images of different degrees and the original images, image de-raining operations are performed. On the premise of ensuring the image de-raining effect, the model size is effectively reduced and the calculation efficiency is improved.

[0057] To achieve the above technical effects, the overall idea of the present application is as follows:

[0058] An image de-raining method based on attention mechanism encoding and decoding, the method comprising the following steps:

[0059] S1. Construct a degradation database based on different real rain-free images and corresponding real light-rain images and real heavy-rain images;

[0060] S2. Construct a network model based on the degradation database;

[0061] S3. Perform de-raining operations on the real light-rain image or real heavy-rain image based on the network model to obtain the corresponding model de-rained image;

[0062] S4. Optimize the network model based on the model de-rained image and the corresponding real rain-free image;

[0063] S5. Perform de-raining operations on the target image based on the optimized network model.

[0064] The following further elaborates on the embodiments of the present application with reference to the accompanying drawings.

[0065] See Figures 1 to 3 As shown, the embodiment of the present application provides an image de-raining method based on an attention mechanism encoder-decoder, and the method includes the following steps:

[0066] S1. Based on different real rain-free images and corresponding real light rain images and real heavy rain images, construct a degradation database;

[0067] S2. Based on the degradation database, construct a network model;

[0068] S3. Perform de-raining operations on the real light rain image or the real heavy rain image based on the network model to obtain the corresponding model de-rained image;

[0069] S4. Optimize the network model based on the model de-rained image and the corresponding real rain-free image;

[0070] S5. Perform de-raining operations on the target image based on the optimized network model.

[0071] In the embodiment of the present application, a database is established based on the gradual image degradation effect, and then a corresponding network model is constructed. Combining the correspondence between different degrees of de-rained images and the original images, image de-raining operations are performed. On the premise of ensuring the image de-raining effect, the model size is effectively reduced and the calculation efficiency is improved.

[0072] It should be noted that the framework based on transformer is increasingly used in the field of machine vision. The encoder-decoder model of the multi-head self-attention mechanism of the present invention, through a self-built dataset strengthened by a degradation mode, after training and introducing the GateMLP sparse coefficient, the performance and effect of the model have been improved;

[0073] In the mode of gradually increasing the image degradation effect, the degradation degree includes three images from shallow to deep, namely the true value image, the light rain droplet image, and the heavy rain droplet image. The degradation transition mode makes the application range of the dataset wider and the generalization ability of the final model stronger;

[0074] The Gate MLP layer adds a Gate and a penalty term sparse network activation weight. The number of parameters of the model is reduced by 5%, the calculation speed of the model is improved, and it has better real-time processing performance, making this function better deployed and implemented.

[0075] Further, in the construction of the degradation database based on different real rain-free images and corresponding real light rain images and real heavy rain images, the following steps are included;

[0076] Based on the real rain-free image, perform light rain spraying simulation to obtain the corresponding real light rain image;

[0077] Based on the real rainless image, a heavy rain spraying simulation is carried out to obtain the corresponding real heavy rain image;

[0078] Associate different real rainless images, the corresponding real light rain images and the corresponding real heavy rain images, and construct the degradation database.

[0079] Further, in constructing the network model based on the degradation database, the following steps are included:

[0080] Take the real rainless image as the original state image, the real light rain image as the shallow degradation image, and the real heavy rain image as the deep degradation image;

[0081] Construct the network model based on the change situation reflected by the original state image, the corresponding shallow degradation image and the corresponding deep degradation image.

[0082] Further, in constructing the network model based on the change situation reflected by the original state image, the corresponding shallow degradation image and the corresponding deep degradation image, the following steps are included:

[0083] Based on the original state image, the corresponding shallow degradation image and the corresponding deep degradation image, reverse-infer the change situation of the image in the rain removal operation simulation;

[0084] Construct the network model based on the change situation of the image in the rain removal operation obtained by simulation.

[0085] Further, in optimizing the network model based on the model rain removal image and the corresponding real rainless image, the following steps are included:

[0086] Based on the model rain removal image and the corresponding real rainless image, obtain the corresponding peak signal-to-noise ratio and the structural similarity between the model rain removal image and the real rainless image;

[0087] Judge whether the regional operation of the network model meets the preset requirements based on the peak signal-to-noise ratio and the structural similarity, and if not, optimize the network model.

[0088] It should be noted that in the embodiments of the present application, traditional digital image processing techniques need to be used:

[0089] Since in current digital image processing, based on the CNN model framework, the parameters are huge, and vision transformers are increasingly used in visual detection, segmentation, and enhancement tasks;

[0090] The technical solution of the embodiment of the present application is based on the Transformer framework and improves the Trans Weather, an enhancement algorithm for image degradation caused by weather such as rain, fog, snow, and haze.

[0091] An individual model composed of an encoding-decoding module is used to complete the enhancement of images degraded by all types of weather.

[0092] Generally, the evaluation metrics for image enhancement accuracy are PSNR (Peak Signal to Noise Ratio) and SSIM (structural similarity index).

[0093] PSNR is an objective criterion for evaluating images, and the calculation formula is:

[0094]

[0095] SSIM is an index for measuring the similarity between two images. Among the two images used, one is an uncompressed and distortion-free image, and the other is a distorted image. The value ranges from 0 to 1, where 1 means completely identical and 0 means completely different.

[0096] Suppose two given M*N images are X and Y. The mean, standard deviation of X, and the covariance between X and Y are represented by ux, σx, and σxy respectively. The comparison functions for luminance, contrast, and structure are defined as:

[0097]

[0098]

[0099]

[0100] SSIM(X,Y) = [l(X,Y)] α [c(X,Y)] β [s(X,Y)] γ (5)

[0101] The Transformer structure proposed in the technical solution of the embodiment of the present application extracts features by serially and parallely connecting Transformer blocks with different patch sizes, proposes a method for a sparse perception layer, improves the raindrop removal effect, enhances PSNR and SSIM, and reduces the Flops and the number of network parameters of the model, thereby improving the processing speed of a single image.

[0102] The technical solution of the embodiment of this application proposes an improvement solution from the dataset to the network model. Currently, most datasets are synthesized based on rain lines and the background, or true value and degraded image pairs are constructed by manually spraying raindrops. These datasets are lacking in terms of authenticity and practicality. Based on the undistorted fisheye vehicle-mounted camera images, the present invention proposes a mode of gradually increasing the image degradation effect. The degradation degree includes three images from shallow to deep, namely the true value image, the light raindrop image, and the heavy raindrop image. As follows Figure 1 As shown, the dataset mode of the degradation transition mode is novel, has a wider application range, and the generalization ability of the final model is stronger.

[0103] In the technical solution of the embodiment of this application, the network structure of attention encoding and decoding is improved on the all-in-one mode. By resizing and Embedding the input degraded image and true value image pair, and inputting different-sized Transformer Blocks to calculate self-attention. In order to reduce the model size and improve the calculation efficiency, a Gate and penalty term sparse network activation weight are added to the MLP layer, and the Gate MLP structure is proposed.

[0104] The operation process of the embodiment of this application can be as follows:

[0105] Construct a dataset, and collect 360-degree panoramic images of similar scenes with information relevance as the network training dataset;

[0106] Build a network model, including a convolutional neural network and a Transformer network;

[0107] Extract the features of the input image of the network training dataset through the convolutional neural network to obtain a feature map;

[0108] Pay attention to the feature elements in the feature map through the Transformer network and associate them with the most relevant feature elements;

[0109] The feature elements output by the Transformer network extract features through the discriminator and perform element-wise multiplication to generate an attention map, which is the rain-free image generated by the Transformer network. Train by inputting the rain-free image generated by the Transformer network and the real rain-free image into the discriminator, and use maximum likelihood estimation to describe the gap between the two rain-free images and converge;

[0110] After the training is completed, remove the discriminator, evaluate the effect of image restoration with peak signal-to-noise ratio and structural similarity, and obtain the 360-degree panoramic image after rain removal.

[0111] Based on the technical solution of the embodiment of this application, the specific implementation process is as follows:

[0112] First, create a dataset. The technical solution of the embodiment of the present application adopts a progressive method to obtain degraded images, and the specific situation is as follows:

[0113] First, obtain the ground truth image, which is clear and non-degraded. Then, artificially spray a small amount of water droplets on the ground truth image to form a first-stage slightly degraded image. Finally, spray large water droplets on the basis of the slightly degraded image to form a severely degraded image; among them,

[0114] Correspond the ground truth, slightly degraded, and severely degraded images one by one, and combine them into input image pairs. The dataset is an in-vehicle fisheye image that has not undergone distortion correction, which better contains the details and deformations of the original image;

[0115] The progressive-mode degradation dataset can help the model better learn various raindrop features. The learning effect of migrating to large raindrop features on the small rain dataset is more obvious, the application range is wider, and finally the generalization ability of the model is stronger.

[0116] Furthermore, create a corresponding algorithm model, and the specific situation is as follows:

[0117] Input the image pair composed of the ground truth picture and the degraded picture. First, perform patch embedding to convert the picture into a vector;

[0118] Use a preset vector input encoding module, which passes through 3 encoder blocks and 3 intra encoder blocks. Each block is composed of a multi-head self-attention layer (MSA), Depth-wise, and an improved GateMLP;

[0119] During the forward propagation of the MSA, the self-attention feature is calculated as shown in formula 1 below

[0120] E(I i ) = Gate MLP(GELU(DWC(Gate MLP((MSA(I i ) + I i )))) + (MSA(I i ))

[0121] + I i )(6); where,

[0122] E() represents the encoder block, I i represents the input image vector, GELU is the Gaussian linear error unit, and the formula used by the Intra encoder block is the same as formula (6).

[0123] The Gate MLP layer adds a gate and a penalty term to sparsify the activation weights of the network. The gate mainly adds activation units to the perception mechanism, which can sparsify the coefficients with low weights.

[0124]

[0125] Output = gate * x (8)

[0126] In formula (7), μ is the added random noise, and logα represents the initialized weight. In the inference stage, the parameters have converged, so no noise is introduced.

[0127] Formula (8) indicates that the gate acts on the MLP layer. In the Gate MLP layer, penalty terms L0 and L2 are applied to further sparsify the MLP layer coefficients.

[0128] Simply put, the L0 penalty term is the number of non-zero components in the weight vector, and the L2 penalty term is the maximum singular value of the weight matrix. When the L0 and L2 penalty terms are obtained, the sparsity ratio of the weights can be calculated, and some weight coefficients can be compressed through the sparsity ratio.

[0129] Among them, the preset decoding module is composed of embedding to transform the feature vector, followed by MSA feature decoding and GateMLP feature sparsification.

[0130] The feature vector passes through the upsampling layer and the connection layer and is finally transformed into an image encoding.

[0131] The output encoded image is transformed into an RGB image, and the PSNR and SSIM between the ground truth image and the processed output image are calculated through formulas (1) to (5) to judge the performance of the algorithm processing.

[0132] In summary, the innovations of the embodiments of this application at least include:

[0133] First, a dataset is made, that is, a progressively degenerating dataset is constructed, which helps to strengthen feature learning.

[0134] Second, the MLP layer is improved, and Gate MLP is proposed. By introducing a gate and a penalty term, the weight matrix is sparsified, making the weights of the model smaller and the overall computational amount smaller, which helps to improve the inference speed of the model.

[0135] See Figure 4 As shown, based on the same inventive concept as the method embodiment, the embodiments of this application provide an image de-raining device based on an attention mechanism for encoding and decoding. The device includes:

[0136] A database construction module, which is used to construct a degradation database based on different real rainless images and corresponding real light rain images and real heavy rain images;

[0137] A model construction module, which is used to construct a network model based on the degradation database;

[0138] A rain removal model module, which is used to perform rain removal operations on the real light rain image or the real heavy rain image based on the network model to obtain a corresponding model rain removal image;

[0139] A model optimization module, which is used to optimize the network model based on the model rain removal image and the corresponding real rainless image;

[0140] A rain removal execution module, which is used to perform rain removal operations on the target image based on the optimized network model.

[0141] In the embodiments of the present application, a database is established based on the gradual image degradation effect, and then a corresponding network model is constructed. Combining the correspondence between different degrees of rain removal images and the original images, image rain removal operations are performed. On the premise of ensuring the rain removal effect of the image, the model size is effectively reduced and the calculation efficiency is improved.

[0142] It should be noted that frameworks based on transformers are increasingly used in the field of machine vision. The encoder-decoder model of the multi-head self-attention mechanism of the present invention, through a self-built dataset strengthened by a degradation mode, after training and introducing the GateMLP sparse coefficient, the performance and effect of the model have been improved;

[0143] In the mode of gradually increasing the image degradation effect, the degradation degree includes three images from shallow to deep, namely the ground truth image, the light rain droplet image, and the heavy rain droplet image. The degradation transition mode makes the application range of the dataset wider, and finally the generalization ability of the model is stronger;

[0144] The Gate MLP layer adds a Gate and a penalty term to sparse the network activation weights. The number of parameters of the model is reduced by 5%, improving the calculation speed of the model, having better real-time processing performance, and enabling better deployment and implementation of this function.

[0145] Further, the database construction module is used to perform light rain spraying simulation based on the real rainless image to obtain the corresponding real light rain image;

[0146] The database construction module is used to perform heavy rain spraying simulation based on the real rainless image to obtain the corresponding real heavy rain image;

[0147] The database construction module is used to associate different real rainless images, corresponding real light rain images, and real heavy rain images, and construct the degradation database.

[0148] Further, the model construction module is used to take the real rainless image as the original state image, the real light rain image as the shallow degradation image, and the real heavy rain image as the deep degradation image;

[0149] The model construction module is used to construct the network model based on the changes reflected by the original state image, the corresponding shallow degradation image, and the corresponding deep degradation image.

[0150] Further, the model construction module is used to reverse and simulate the image change situation of the rain removal operation based on the original state image, the corresponding shallow degradation image, and the corresponding deep degradation image;

[0151] The model construction module is used to construct the network model based on the simulated image change situation of the rain removal operation.

[0152] Further, the model optimization module is used to obtain the corresponding peak signal-to-noise ratio and the structural similarity between the model rain removal image and the real rainless image based on the model rain removal image and the corresponding real rainless image;

[0153] The model optimization module is used to judge whether the regional operation of the network model meets the preset requirements based on the peak signal-to-noise ratio and the structural similarity, and if not, optimize the network model.

[0154] It should be noted that for the image rain removal device based on attention mechanism encoding and decoding provided in the embodiments of the present application, its corresponding technical problems, technical means, and technical effects are similar in principle to those of the image rain removal method based on attention mechanism encoding and decoding.

[0155] It should be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0156] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. An image de-raining method based on attention mechanism encoding and decoding, characterized in that, The method includes the following steps: Construct a degradation database based on different real rain-free images and corresponding real light-rain images and real heavy-rain images; Construct a network model based on the degradation database; Perform rain removal on the real light-rain image or the real heavy-rain image based on the network model to obtain a corresponding model rain-removed image; Optimize the network model based on the model rain-removed image and the corresponding real rain-free image; Perform rain removal on the target image based on the optimized network model; The constructing a degradation database based on different real rain-free images and corresponding real light-rain images and real heavy-rain images, and constructing a network model based on the degradation database specifically includes: First, obtain real rain-free images, then perform light-rain spraying simulation on the real rain-free images to form real light-rain images, forming a first-stage shallow degradation image, and finally perform heavy-rain spraying simulation on the basis of the shallow degradation image to form a deep degradation image; Correspond the real rain-free images, shallow degradation images, and deep degradation images one by one, combine them into input image pairs, and construct the network model based on the changes reflected by the real rain-free images, the corresponding shallow degradation images, and the corresponding deep degradation images; The constructing the network model specifically includes the following steps: Input the input image pair composed of real rain-free images, shallow degradation images, and deep degradation images. First, perform patch embedding to convert the pictures into vectors; Use a preset vector input encoding module, through 3 encoder blocks and 3 intra encoder blocks. Each block is composed of a multi-head self-attention layer (MSA), Depth-wise, and an improved Gate MLP. The Gate MLP layer adds a Gate and a penalty term sparse network activation weight. The Gate adds an activation unit to the perception mechanism to sparse out the coefficients with low weights.

2. The image rain removal method based on attention mechanism encoding and decoding according to claim 1, wherein Based on the real rain-free image, the corresponding shallow degradation image, and the corresponding deep degradation image, reverse-infer the picture change situation of the simulated rain removal operation; Construct the network model based on the picture change situation of the simulated rain removal operation.

3. The image de-raining method based on attention mechanism encoding and decoding according to claim 1, characterized in that, In the optimizing the network model based on the model rain-removed image and the corresponding real rain-free image, the following steps are included: Based on the model rain-removed image and the corresponding real rain-free image, obtain the corresponding peak signal-to-noise ratio and the structural similarity between the model rain-removed image and the real rain-free image; Judge whether the rain removal operation of the network model meets the preset requirements based on the peak signal-to-noise ratio and the structural similarity. If not, optimize the network model.

4. An image de-raining device based on attention mechanism encoding and decoding using the image de-raining method based on attention mechanism encoding and decoding according to any one of claims 1-3, characterized in that, The device includes: A database construction module, which is used to construct a degradation database based on different real rain-free images and corresponding real light-rain images and real heavy-rain images; A model construction module, which is used to construct a network model based on the degradation database; A rain removal model module, which is used to perform rain removal operations on the real light rain image or the real heavy rain image based on the network model to obtain the corresponding model rain-removed image; A model optimization module, which is used to optimize the network model based on the model rain-removed image and the corresponding real rain-free image; A rain removal execution module, which is used to perform rain removal operations on the target image based on the optimized network model.

5. The image rain removal device based on attention mechanism encoding and decoding according to claim 4, characterized in that: The database construction module is used to perform light rain spraying simulation based on the real rain-free image to obtain the corresponding real light rain image; The database construction module is used to perform heavy rain spraying simulation based on the real rain-free image to obtain the corresponding real heavy rain image; The database construction module is used to associate different real rain-free images and the corresponding real light rain images and real heavy rain images, and construct the degradation database.

6. The image rain removal device based on attention mechanism encoding and decoding according to claim 4, characterized in that: The model construction module is used to use the real rain-free image as the original state image, the real light rain image as the shallow degradation image, and the real heavy rain image as the deep degradation image; The model construction module is used to construct the network model based on the change situations reflected by the original state image, the corresponding shallow degradation image, and the corresponding deep degradation image.

7. The image rain removal device based on attention mechanism encoding and decoding according to claim 6, characterized in that: The model construction module is used to reverse-infer and simulate the picture change situation of the rain removal operation based on the original state image, the corresponding shallow degradation image, and the corresponding deep degradation image; The model construction module is used to construct the network model based on the simulated picture change situation of the rain removal operation.

8. The image rain removal device based on attention mechanism encoding and decoding according to claim 4, characterized in that: The model optimization module is used to obtain the corresponding peak signal-to-noise ratio and the structural similarity between the model rain-removed image and the real rain-free image based on the model rain-removed image and the corresponding real rain-free image; The model optimization module is used to judge whether the rain removal operation of the network model meets the preset requirements based on the peak signal-to-noise ratio and the structural similarity, and if not, optimize the network model.

Citation Information

Patent Citations

  • Image rain removal network, and training method and device of image rain removal network

    CN115984143A