Rain and fog image restoration method based on frequency attention and uncertainty guidance network

Through the method of guiding the network with frequency attention and uncertainty, the problem of missing details and degradation of accuracy during image restoration in various bad weather is solved, and a better image recovery effect is achieved, especially in rainy and foggy weather preserving the edges and texture details of the image.

CN120298256APending Publication Date: 2025-07-11CHANGAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510403890.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the process of image restoration in various bad weather, there are problems such as missing details and degradation of recovery accuracy. Especially when image restoration in rainy and fog weather, the existing methods fail to effectively consider the degradation of rainy and fog superposition, resulting in the recovery image losing details and introducing artifacts.

Method used

Using a method based on frequency attention and uncertainty guidance network, the Laplace pyramid structure jump connection unit and confidence feature feedback unit are designed to improve image restoration effect through feature encoding, feature refinement enhancement and confidence feature feedback decoding processing.

Benefits of technology

It effectively improves the details retention and accuracy of image restoration, alleviates the uncertainty of image restoration in rainy and foggy weather, and improves the quality of image restoration, especially in terms of vision blurring and color degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298256A_ABST
    Figure CN120298256A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a rain and fog image restoration method based on a frequency attention and uncertainty guidance network, which comprises the following steps: acquiring a rain and fog image to be restored; and inputting a to-be-restored rain and fog image into the network model based on frequency attention and uncertainty guidance, and obtaining a restored rain and fog image after feature coding, feature refinement enhancement and self-confidence feature feedback decoding processing in sequence. After the architecture of the frequency attention and uncertainty guide network model is coded, residual information among different scale features is gradually superposed from a low scale to a high scale, in addition, in order to deeply mine features, from the angle of a frequency domain, feature refinement enhancement is carried out, and self-confidence features are fed back in real time in the feature decoding process to guide feature optimization. Therefore, the image restoration effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a rain and fog image restoration method based on frequency attention and uncertainty guidance network. Background Technique

[0002] With the development of information technology, digital image processing technology and deep learning in recent years, there have been more and more new algorithms and theories in the field of image restoration, and they are widely applied to various fields related to image processing. For example, in the field of public security, traffic monitoring images can be used to investigate and obtain evidence for accidents; in the field of remote sensing, remote sensing images can be used to identify and monitor various geographical and environmental targets, etc. Therefore, it is crucial to obtain clear images. However, in real life, the quality of the acquired images is not high due to various reasons. Especially in harsh weather conditions, such as haze days, rainy days, etc.

[0003] Deep learning-based methods have good performance in both single-image rain removal and single-image fog removal tasks. However, most methods only design network structures for single-image rain removal or single-image fog removal. Some methods design a unified processing framework for image restoration tasks under various harsh weather conditions, but train weights separately for each restoration task without considering the situation of multiple harsh weather conditions superimposed and degraded, such as rain and fog weather; secondly, there will be uncertainty problems in the image restoration process. If the uncertainty problems are not considered, the restored image may lose details and introduce artifacts to a certain extent; at the same time, due to the large differences in the feature levels of rain and fog, it poses a challenge to extract the key features of the background of rain and fog images; finally, the current deep learning networks mainly mine spatial features, but the degraded image and the non-degraded image often have large differences in the frequency domain. Designing a model from the perspective of the frequency domain is a research direction worthy of study.

[0004] Therefore, how to effectively represent the detailed information of images and design corresponding image restoration algorithms has become a challenging problem at present. Summary of the Invention

[0005] The purpose of the present invention is to provide a rain and fog image restoration method based on frequency attention and uncertainty guidance network to solve the problems of missing detailed information and reduced restoration accuracy in image restoration under multiple harsh weather conditions in the prior art.

[0006] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present application discloses a rain and fog image restoration method based on frequency attention and uncertainty guidance network, including: Obtain the rain and fog image to be restored; Input the rainy and foggy image to be restored into the frequency attention and uncertainty-guided network model. After passing through feature encoding, feature refinement and enhancement, and confidence feature feedback decoding processes in sequence, the restored rainy and foggy image is obtained.

[0007] Preferably, after obtaining the rainy and foggy image to be restored, preprocess the rainy and foggy image, which specifically includes: Crop the rainy and foggy image to be restored to obtain a cropped image dataset; Divide the cropped image dataset into a training set and a test set.

[0008] Preferably, the frequency attention and uncertainty-guided network model includes: An encoding unit, which is used to extract the initial features and shallow features of the input sample image respectively, concatenate the initial features and shallow features with the same scale and number of channels to obtain features of different scales, and transmit the features of different scales to the decoding unit respectively; the features of different scales are convolved to obtain encoded features, and the encoded features are transmitted to the decoding unit and the feature refinement module respectively; A feature refinement unit, which is used to receive the encoded features, refine and enhance the encoded features to obtain detail-enhanced features, and send the detail-enhanced features to the decoding unit; A Laplacian pyramid structure skip connection unit, which is used to connect the encoding unit and the decoding unit, calculate the residual information between features of different scales, and concatenate the residual information to the decoding unit; A confidence feature feedback unit, which is used to obtain optimized features by feedback confidence optimization processing of the fused features; and transmit the optimized features to the decoding unit; A decoding unit, which is used to receive the encoded features, detail-enhanced features and optimized features, concatenate the received encoded features and detail-enhanced features to obtain fused features, stack the fused features and optimized features to obtain feedback fused features, superimpose the feedback fused features and the residual information to obtain output features, and decode the output features to obtain the restored rainy and foggy image.

[0009] Preferably, the refinement of the encoded features in the feature refinement unit specifically includes: Obtain a confidence map and an uncertainty map by performing uncertainty estimation on the encoded features; extract local features and global features from the encoded features through uncertainty local-global feature extraction; The local features are guided by the uncertainty map to obtain uncertainty local features, the overall features refined from the uncertainty local features are obtained according to the uncertainty local features, the difference between the overall features and the encoded features is weighted based on the confidence map to obtain a weighted difference, and an attention-enhanced feature is obtained based on the weighted difference; The attention-enhanced feature is processed through frequency component decomposition and fusion, and a multi-branch compact selection frequency module MCSF to obtain the detail-enhanced feature.

[0010] Preferably, the confidence map and the uncertainty map are obtained through uncertainty estimation of the encoded features, which specifically includes: The encoded features are successively processed through channel shuffling and channel grouping to obtain a number of feature layer groups; The number of feature layer groups respectively pass through a convolutional subnet with shared weights to obtain a number of convolutional output results; Calculate the mean and standard deviation of the number of convolutional output results, and the standard deviation is the estimation of image uncertainty; Solve the confidence map and the uncertainty map according to the estimation of image uncertainty:

[0011]

[0012] Among them, is the confidence map, is the uncertainty map, represents a fixed constant, 1 represents a tensor of corresponding size with all values being 1, and Tanh represents the nn.Tanh() activation function; is the estimation of image uncertainty.

[0013] Preferably, in the confidence feature feedback unit, the fused feature is processed through feedback confidence optimization to obtain an optimized feature, which specifically includes: The fused feature is processed through uncertainty to obtain a confidence map; the confidence map is processed through dimensionality reduction to obtain a query vector for querying the confidence feature in the fused feature; and the fused feature is processed through mapping dimensionality reduction to obtain a key vector and a value vector; Determine the confidence feedback feature based on the query vector, the key vector, and the value vector; Stack the confidence feedback features and cascade them with the fused feature to obtain an optimized feature.

[0014] Preferably, the Laplacian pyramid structure skip connection unit calculates the residual information between features of different scales and superimposes the residual information on the decoding unit, which specifically includes: Calculate the residual information between features of different scales:

[0015] Superimpose the calculated residual information and the feedback fused feature of the corresponding encoding unit to obtain the output feature of the encoding unit B :

[0016] In the formula, represents the residual information between features of different scales; represents the scale of , and the number of channels is Features; The scale is , the number of channels is C features; ConvTranspose represents transposed convolution.

[0017] In the second aspect, the present application also discloses a rain and fog image restoration system based on frequency attention and uncertainty guidance network, which is characterized by comprising: An acquisition unit, used for acquiring the rain and fog image to be restored; The processing unit is used to input the rain and fog image to be restored into the frequency attention and uncertainty guided network model, and obtain the restored rain and fog image after feature encoding, feature refinement and enhancement, and confident feature feedback decoding processing in sequence.

[0018] In a third aspect, the present application also discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned methods for restoring rain and fog image based on frequency attention and uncertainty guided network when executing the computer program.

[0019] In a fourth aspect, the present application also discloses a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the rain and fog image restoration method based on frequency attention and uncertainty guided network as described in any of the above items.

[0020] Compared with the prior art, the present invention has the following beneficial effects: The rain and fog image restoration method based on frequency attention and uncertainty guided network proposed in this application adopts a frequency attention and uncertainty guided network model, and obtains the restored rain and fog image after feature encoding, feature refinement enhancement and confident feature feedback decoding processing in sequence; the architecture of the frequency attention and uncertainty guided network model is based on the U-net architecture, including encoding units and decoding units, and a feature refinement unit UGFR is added between the encoding and decoding units. In the jump connection part of the network, a Laplacian pyramid structure jump connection unit is designed to gradually superimpose the residual information between features of different scales from low scale to high scale, and merge it with the corresponding features of the decoding part. In addition, in order to deeply mine features, a feature refinement unit UGFR is proposed from the perspective of frequency domain. Finally, the feature decoding unit of the network is further improved, and a confident feature feedback unit CFF is proposed to instantly feedback confident features to guide feature optimization during feature decoding. Thereby improving the restoration effect of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic flowchart of the method of the present invention; Figure 2 It is an architecture diagram of the frequency attention and uncertainty guidance network model of the present invention; Figure 3 It is a schematic diagram of the Laplacian pyramid structure skip connection of the present invention; Figure 4 It is an architecture diagram of the feature refinement unit of the present invention; Figure 5 It is an architecture diagram of the uncertainty estimation module of the present invention; Figure 6 It is an architecture diagram of the ULG module of the present invention; Figure 7 It is an architecture diagram of the frequency component decomposition and fusion module of the present invention; Figure 8 It is an architecture diagram of the confident feature feedback module of the present invention; Figure 9 It is a schematic diagram of the filter setting of the present invention. Detailed implementation manners

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the accompanying drawings here can be arranged and designed in various different configurations.

[0024] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0025] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0026] In the description of the embodiments of the present invention, it should be noted that if terms such as "upper", "lower", "horizontal", "inner", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the invention product is usually placed during use. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, terms such as "first", "second", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.

[0027] In addition, if the term "horizontal" appears, it does not mean that the component is required to be absolutely horizontal, but it can be slightly inclined. For example, "horizontal" only means that its direction is more horizontal relative to "vertical", and it does not mean that the structure must be completely horizontal, but it can be slightly inclined.

[0028] In the description of the embodiments of the present invention, it should also be noted that unless otherwise clearly defined and limited, if terms such as "set", "installed", "connected", "coupled" are used, they should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0029] The following further describes the present invention in detail with reference to the drawings: See Figure 1 , in an embodiment of the present invention, a rain and fog image restoration method based on a frequency attention and uncertainty guidance network is provided, which effectively solves the problems of missing detail information and reduced restoration accuracy during image restoration in various bad weather conditions. The specific steps are as follows: S1: Obtain the rain and fog image to be restored; S2: Input the rain and fog image to be restored into the frequency attention and uncertainty guidance network model. After passing through feature encoding, feature refinement enhancement, and confidence feature feedback decoding processing in sequence, the restored rain and fog image is obtained.

[0030] The architecture of the frequency attention and uncertainty-guided network model proposed in this application is based on the U-net architecture, including an encoding part and a decoding part, and a feature refinement module UGFR is added between the encoding and decoding parts. In the jump connection part of the network, a Laplacian pyramid structure jump connection is designed to gradually superimpose the residual information between features of different scales from low scale to high scale, and merge it with the corresponding features of the decoding part. In addition, in order to deeply mine features, a feature refinement module UGFR is proposed from the perspective of the frequency domain. Finally, the feature decoding part of the network is further improved, and a confident feature feedback module CFF is proposed to instantly feedback confident features during the feature decoding process to guide feature optimization. Thereby improving the restoration effect of the image.

[0031] In some embodiments, after obtaining the rain and fog image to be restored, preprocessing the rain and fog image specifically includes: Crop the rain and fog image to be restored to obtain a cropped image dataset; The cropped image dataset is divided into training and testing sets.

[0032] In some embodiments, the rain and fog image is preprocessed, specifically including: S101: Inputting a degraded and clear image pair to be restored , N is the number of samples. This example uses but is not limited to 128×128 randomly cropped blocks to train the network.

[0033] S102: Randomly shuffle the dataset and divide it into a training set and a test set.

[0034] S103: End the preprocessing of the input image data and obtain the corresponding training set and test set.

[0035] In some embodiments, the frequency-attention and uncertainty-guided network model includes: The frequency attention and uncertainty guided network model includes: The encoding unit is used to extract the initial features and shallow features of the input sample image respectively, cascade the initial features and shallow features with the same scale and number of channels to obtain features of different scales, and transmit the features of different scales to the decoding unit respectively; convolute the features of different scales to obtain encoding features, and transmit the encoding features to the decoding unit and feature refinement module respectively; A feature refinement unit is used to receive the encoded features, refine and enhance the encoded features to obtain detail enhancement features, and send the detail enhancement features to the decoding unit; The Laplacian pyramid structure jump connection unit is used to connect the encoding unit and the decoding unit, calculate the residual information between features of different scales, and cascade the residual information to the decoding unit; A confidence feature feedback unit, configured to optimize the fused features through feedback confidence processing to obtain optimized features; and transmit the optimized features to the decoding unit; A decoding unit, configured to receive the encoded features, the detail enhancement features, and the optimized features, cascade the received encoded features and the detail enhancement features to obtain fused features, stack the fused features and the optimized features to obtain feedback fused features, superimpose the feedback fused features and the residual information to obtain output features, and decode the output features to obtain the restored rainy and foggy image.

[0036] In some embodiments, the Laplacian pyramid structure skip connection unit has the following specific processing flow: S301: Calculate the residual between features of different scales in the encoding part R , as shown in Equation (1): (1) where represents the feature with scale and number of channels , represents the feature with scale and number of channels C , and ConvTranspose represents transposed convolution, which makes the two terms for calculating the residual match in both scale and number of channels.

[0037] S302: Superimpose the residual obtained in S301 and the corresponding encoding part features to obtain the enhanced features between features of different scales in the encoding part B , as shown in Equation (2): (2) As Figure 3 shown, the features stacked by the Laplacian pyramid structure skip connection in FD-Up_1 and FD-Up_2 can be represented by Equations (3) and (4) respectively: (3) (4) where FD represents the transposed convolution ConvTranspose.

[0038] S303: Then stack the enhanced features in the encoding part and the features in the decoding part and fuse them with 3×3 Conv. For example, fuse the texture detail enhanced encoding part feature B 1 obtained from Equation (3) with the output feature of the first upsampling layer FD-Up_1 in the decoding part; Equation (4) uses transposed convolution to upsample the low-scale residual R 1, halve its number of channels, and make it match and add with the higher-scale feature residual R 2 to obtain the texture detail enhanced encoding part featureB It is fused with the output features of the second upsampling layer FD-Up_2 in the decoding part.

[0039] In some embodiments, the encoded features are refined and enhanced within the feature refinement unit, specifically including: The encoded features are passed through uncertainty estimation to obtain a confidence map and an uncertainty map; the encoded features are passed through uncertainty local-global feature extraction to obtain local features and global features; The local features are guided by the uncertainty map to obtain uncertainty local features, and the overall features refined from the uncertainty local features are obtained according to the uncertainty local features. The difference between the overall features and the encoded features is weighted based on the confidence map to obtain a weighted difference, and an attention-enhanced feature is obtained based on the weighted difference; The attention-enhanced feature undergoes frequency component decomposition and fusion processing, as well as a multi-branch compact selection frequency module MCSF to obtain a detail-enhanced feature.

[0040] In some embodiments, a feature refinement module UGFR is constructed to enrich the detailed information of the image, and the refined features are sent to the decoding part; it includes: S401: Construct an Uncertainty Estimation (UE) module in the feature refinement module UGFR. The overall architecture of the UGFR module is as Figure 4 shown.

[0041] S402: Obtaining the confidence map and the uncertainty map.

[0042] The uncertainty output by the UE module is simply processed to obtain a confidence map and an uncertainty map to guide feature learning. The uncertainty map is used to guide the expression of uncertain features in the local features in the local information extraction branch of the MHSA module in ULG, and the confidence map is used to enhance the confidence features in the overall feature transmission of the UGFR module. The calculation of the confidence map and the uncertainty map is as shown in (5) and (6): (5) (6) Among them, is the confidence map, is the uncertainty map, represents a fixed constant, which is set to 5 in the experiment, 1 represents a tensor of the corresponding size with all values being 1, and Tanh is the nn.Tanh() activation function in the experiment.

[0043] S403: Construct the Uncertainty Local Global feature extraction module (ULG) in the feature refinement module UGFR.

[0044] S404: Confidence feature enhancement operation Enhance the expression of confidence features during the overall transmission of features. The specific operation is shown in Equation (7): (7) where is the input feature, is the output feature of ULG. For the difference between the output feature of ULG and the original input feature , it is weighted by the confidence map . The weighting operation will assign higher weights to the features with enhanced attention in the confidence region. Finally, the weighted difference is superimposed back onto the features with enhanced attention, thereby enhancing the expression of the features with enhanced attention in the confidence region within the overall features with enhanced attention.

[0045] Step 405: Construct the Frequency component decomposition and fusion (FCDF) module in the feature refinement module UGFR.

[0046] Step 406: Complete the construction of the feature refinement module UGFR and output the feature map after feature refinement. This example includes, but is not limited to, the stacking of 3 UGFR modules.

[0047] In some embodiments, Step 401 includes the following steps: Step 501: The present invention mainly considers arbitrary uncertainties, that is, the pixels of the collected images show a certain distribution due to factors such as noise. The present invention regards it as the uncertainty caused by the degradation of the image pixels due to rain and fog, and proposes an Uncertainty Estimation (UE) module as Figure 5 shown.

[0048] The UE module first performs a channel shuffle operation on each feature layer of the feature map in the channel dimension. Figure 5 An example of channel shuffle is to reorganize the input number of channels in groups of 4.

[0049] Step 502: Next, perform a channel grouping operation (Channel Split) to divide the feature layer into several groups. It should be noted that the feature layers in each group are non - adjacent before channel shuffling and channel grouping, which can ensure the diversity and integrity of each group of features as much as possible.

[0050] Step 503: In the present invention, each group of features passes through a group of convolution sub - networks with shared weights, which includes Conv3×3 + BN layer + ReLU activation function + Conv3×3 + BN layer + Sigmoid activation function. The convolution kernels corresponding to each group of features for convolution operation in the UE are weight - shared. Since the main function of the UE module is to estimate uncertainty, each group of features passes through a convolution sub - network with the same structure and weights. After that, the mean and standard deviation are calculated for the convolution output results of these grouped features. The obtained mean is the prediction of the low - scale restored image, and the obtained standard deviation is used as an estimate of the image uncertainty.

[0051] In some embodiments, step 403 is specifically operated as follows: As Figure 6 shown, in the MHSA module of ULG, a Talking Head fully - connected layer is used to interact with multi - head information. In practice, a 1×1Conv layer is used to replace the fully - connected layer; a learnable attention bias Attention Bias is added; in the V - branch, DW convolution is used to extract local features, and the local features are injected into the attention features. ULG uses the uncertainty map to guide the expression of uncertainty features in local features, and the uncertainty map is passed into the local information injection part of the ULG module. The specific operation is shown in Equation (8): (8) Where represents Figure 5 the features of the V - vector branch in , is a 3×3 DW convolution, is the extracted local information.

[0052] The difference between the obtained local features and the original features is weighted. Higher weight values are assigned to pixel positions with higher uncertainty. Finally, the weighted difference is superimposed back on the local features to obtain the final local features, and the final local features are injected into the attention features. This operation is a feature refinement operation, which will enhance the expression of local features in the uncertain region in the attention features output by the MHSA module.

[0053] In some embodiments, step 405 includes the following steps: Step 601: construct a frequency component decomposition module FCD of the FCDF module, wherein the architecture of the FCDF module is as follows: Figure 7 shown.

[0054] In the FCD module, the present invention performs input feature X The dynamic low-pass filter is generated using the following equation, ,in, , GAP, W and BN represent dynamic low-pass filter, global average pooling, convolution layer and BN layer respectively. X Subtract the acquired low-frequency components The initial high-frequency characteristic components are obtained, and the operation is shown in formula (9): (9) in, represents the high-frequency characteristic component, represents the low-frequency characteristic component, c represents the index of the channel, h, w Represents the spatial coordinate value, . The present invention introduces a 3×3 Laplacian operator as a high-frequency feature extraction filter to supplement the high-frequency components. The edges (contours), noise, and details of the image are the parts where the image changes dramatically, and usually correspond to the high-frequency components in the frequency components. As a classic image edge extraction operator, the Laplacian operator can easily obtain the edge information of the image with a small amount of calculation. The present invention uses the Laplacian operator to perform a convolution operation on the input features to extract the high-frequency components as a supplement to the initial high-frequency components. The specific operation is shown in formula (10): (10) in, represents the Laplacian operator with a positive central coefficient, c represents the index of the channel, h, w Represents the spatial coordinate value, , is the weighting parameter.

[0055] To improve the feature representation capability of the Laplacian operator in the 45-degree direction, the filter is set as follows: Figure 9 shown.

[0056] Step 602: Construct the frequency attention module FCA of the FCDF module The present invention uses a multi-head attention mechanism to re-integrate high- and low-frequency components and extract key features. FCA uses a 1×1 convolutional layer to map the frequency components, and the 1×1 convolutional layer is used to replace the linear mapping layer of the self-attention module. Low-frequency components The query vector is mapped through a 1×1 convolutional layerQ (query), and use a dimensionality transformation operation to convert it into a multi-head form, , M represents the number of heads, with a value of 8, res represents the dimensionality adjustment parameter, with a value of 8, and the high-frequency components are mapped to obtain key vectors through a 1×1 convolutional layer K (key), after dimensionality transformation and transposition, , for the low-frequency components are mapped to obtain value vectors through another set of 1×1 convolutions V (value), after dimensionality transformation, , next, obtain the attention features based on frequency components according to Equation (11): (11) where, F represents the attention features based on frequency components, ab represents the learnable attention shift. Then, based on the value vectors V use DW convolution to extract local features and stack them to F , and finally use a 1×1 convolutional layer to fuse the multi-head information.

[0057] Step 603: Construct the MCSF module of the FCDF module The specific principle of the Multi-branch Compact Selective Frequency (MCSF) module can be found in Reference [2]. This module splits the input features into two groups based on the channel dimension, and the number of channels changes from C to . One group of features passes through the global processing branch, and one group passes through the window processing branch. The processing of the window branch is to divide the input features into 4 windows in the spatial dimension, with the dimension of each window being ( ). Next, apply global average pooling GAP to each window to obtain the low-frequency components, subtract the low-frequency components from the windowed input features to obtain the high-frequency components, give a learnable channel dimension weight to the low-frequency components and the high-frequency components respectively for weighted fusion, and finally reshape the dimension of the weighted fusion features to the original dimension. The global processing branch is similar to the window processing branch, except that there is no windowing processing.

[0058] In some embodiments, the specific process of obtaining the confidence map and the uncertainty map from the encoded features through uncertainty processing includes: The encoded features are successively processed through channel shuffling and channel grouping to obtain several groups of feature layers; Several groups of feature layers respectively pass through convolutional subnets with shared weights to obtain several convolutional output results; Calculate the mean and standard deviation of several convolution output results, where the standard deviation is an estimate of the image uncertainty; Solve the confidence map and uncertainty map according to the estimate of the image uncertainty:

[0059]

[0060] where, is the confidence map, is the uncertainty map, represents a fixed constant, 1 represents a tensor of corresponding size with all values being 1, and Tanh represents the nn.Tanh() activation function; is the estimate of the image uncertainty.

[0061] In some embodiments, in the confidence feature feedback unit, the fused feature is processed through feedback confidence optimization to obtain an optimized feature, which specifically includes: The fused feature is processed through uncertainty processing to obtain a confidence map; and the fused feature is processed through mapping and dimensionality reduction to obtain a key vector and a value vector; The confidence map is processed through dimensionality reduction to obtain a query vector for querying the confidence feature in the fused feature; Determine the confidence feedback feature based on the query vector, key vector, and value vector; Stack the confidence feedback features and concatenate them with the fused feature to obtain the optimized feature.

[0062] In some embodiments, step 106 includes the following steps: Step 701: The confidence feature feedback module CFF is as Figure 8 shown.

[0063] In the CFF module, the output of the UE module is expanded into a one-dimensional vector in the channel dimension, changing from a three-dimensional vector to a two-dimensional vector, and this is used as the query vector (query) to query the confidence feature in the input feature . C 1 is the number of channels of the multi-scale predicted restoration map, which is 3. The CFF module uses a 1×1 Conv layer to replace the linear mapping layer in the self-attention module of ViT. At the same time, the 1×1 Conv layer also plays a role in dimensionality reduction, mapping and reducing the dimensionality to the number of channels C 2 ( C 2 = 64). Next, the mapped feature is also expanded into a one-dimensional vector in the channel dimension, and finally two two-dimensional vectors are generated: the key vector and the value vector . The operations of obtaining the confidence map and confidence feature using the attention mechanism are as shown in Equation (12): (12) Among them, = HW . Q and The matrix multiplication is to query Q , K the similarity degree of the vector in the channel dimension. The CFF module stacks the confidence features and cascades them with the features of the decoding part, which plays a role in feedback and guidance for the network learning process.

[0064] In some embodiments, the Laplacian pyramid structure skip connection unit calculates the residual information between features of different scales and superimposes the residual information on the decoding unit, specifically including: Calculating the residual information between features of different scales:

[0065] Superimposing the calculated residual information and the feedback fusion features of the corresponding encoding unit to obtain the output features of the encoding unit B :

[0066] In the formula, represents the residual information between features of different scales; represents the feature with scale and number of channels ; represents the feature with scale and number of channels C ; ConvTransp ose represents transposed convolution.

[0067] The specific algorithm steps are as follows: Step 101: Start image restoration based on frequency attention and uncertainty guidance network; Step 102: Preprocess the input image data to obtain a training set and a test set; Step 103: Send the input sample image into the encoding part for feature extraction (FE), shallow feature extraction (SFE), and cascade it with the features of the encoding part with the same scale and number of channels; Step 104: Construct a Laplacian pyramid structure skip connection, and gradually stack the residual information between features of different scales from low scale to high scale for fusion with the corresponding features of the decoding part; Step 105: Construct a feature refinement module UGFR to enrich the detail information of the image and send the refined features into the decoding part; Step 106: Construct a confidence feature feedback module CFF to instantaneously feedback confidence features during the feature decoding process to guide feature optimization.

[0068] Step 107: Construct a network structure based on frequency attention and uncertainty guidance. Input the training data for network training; Step 108: Input the test data into the trained network model to end the image restoration of the network based on frequency attention and uncertainty guidance.

[0069] In some embodiments, step 401 includes the following steps: Step 501: The present invention mainly considers arbitrary uncertainty, that is, the acquired image pixels show a certain distribution due to factors such as noise. The present invention regards it as the uncertainty caused by rain and fog degradation of the image pixels, and proposes an uncertainty estimation (UE) module as Figure 5 shown.

[0070] The UE module first performs a channel shuffle operation on each feature layer of the feature map in the channel dimension. Figure 5 An example of channel shuffle is to reorganize the input number of channels with a group number of 4.

[0071] Step 502: Next, perform a channel split operation to divide the feature layer into several groups. It should be noted that each group of feature layers is not adjacent before channel shuffle and channel split, so as to ensure the diversity and integrity of each group of features as much as possible. Step 503: The present invention passes each group of features through a group of convolution subnets with shared weights, which includes Conv3×3 + BN layer + ReLU activation function + Conv3×3 + BN layer + Sigmoid activation function. Each group of convolution kernels corresponding to the convolution operation of each group of features in the UE is weight-shared. Since the main function of the UE module is to estimate uncertainty, each group of features is passed through a convolution subnet with the same structure and weights. Thereafter, the mean and standard deviation are calculated for the convolution output results of these grouped features. The obtained mean is the prediction of the low-scale restored image, and the obtained standard deviation is used as the estimation of image uncertainty.

[0072] In some embodiments, the specific algorithm steps are as follows: Step 101: Start the image restoration of the network based on frequency attention and uncertainty guidance; Step 102: Preprocess the input image data and obtain the training set and test set; Step 103: Feed the input sample image into the encoding part for feature extraction (FE), shallow feature extraction (SFE), and concatenate it with the features of the encoding part with the same scale and number of channels; Step 104: Construct a Laplacian pyramid structure skip connection, and gradually stack the residual information between different scale features from low scale to high scale for fusion with the corresponding features in the decoding part; Step 105: Construct a feature refinement module UGFR to enrich the detailed information of the image and feed the refined features into the decoding part; Step 106: Construct a confidence feature feedback module CFF to instantaneously feedback confidence features during the feature decoding process to guide feature optimization.

[0073] Step 107: Construct a network structure based on frequency attention and uncertainty guidance. Input the training data for network training; Step 108: Input the test data into the trained network model to end the image restoration based on the network of frequency attention and uncertainty guidance.

[0074] The said Step 107 includes the following steps: Step 801: Construct a network structure based on frequency attention and uncertainty guidance. This network model includes operations such as feature extraction (FE), shallow feature extraction (SFE), Laplacian filter, frequency component decomposition module FCD, frequency attention module FCA, etc. The network structure can be expressed as [FE, FE-Down_1, SFE_1, , FE-Down_2, SFE_2, , UE, ULG, FCD FCA,MCSF, FD1, FD2, FD3, Mix, FD-Up_1, , FD-Up_2, , FR, . Step 802: Input the training set constructed in Step 102 into the network of frequency attention and uncertainty guidance.

[0075] Step 803: In the encoder part, first extract shallow features in sequence and concatenate them with the corresponding features , , , , , , , and then calculate the intermediate feature information through formulas (5) - (11) , , ; Step 804: In the decoder part, restore the image Successively pass through , , , , , to obtain; Step 806: Train the model, and the entire network model performs parameter learning through error backpropagation.

[0076] Further preferably, step 108 includes the following steps: Step 901: Use the frequency attention and uncertainty-guided network model trained in step 107 to perform image restoration on the test set constructed in step 102.

[0077] The total loss used during network training is expressed by the following formula (13), which includes the MAE loss and , the uncertainty penalty loss , the uncertainty constraint loss , the contrast loss , the structural similarity loss , and the frequency loss function .

[0078] (13) where represents the total loss function for training the network, , , , respectively represent the weighted hyperparameters for , , , and , and are set to 0.02, 0.1, 0.4, and 0.05 respectively during training.

[0079] The calculation of f ( I ) and the clear image is represented by the MAE between: (14) where represents the calculation of MAE. At the same time, the present invention also uses the clear image Downsample to multiple scales and calculate the MAE between the clear image and the predicted maps output by multiple UE modules in the network. Among them, the scales of the predicted maps output by UE_1, UE_2, and UE_3 in the feature refinement part are 0.25, and the scales of the predicted maps output by UE_4 and UE_5 in the decoding part are 0.5 and 1 respectively.

[0080] The calculation is as follows: (15) where, represents downsampling the clear image to the scale of 0.25, represents downsampling the clear image to the scale of 0.5, , , , , respectively represent the average predicted maps output by UE_1, UE_2, UE_3, UE_4, and UE_5. , , are weighted hyperparameters, which are set to 0.1, 0.15, and 0.2 respectively during training.

[0081] Uncertainty penalty loss The calculation formula is as shown in (16): (16) where, taking as the weight to control the MAE loss of the predicted restored image . Pixels considered to be in the uncertainty region will be given a larger weight for penalty to drive the network to learn. Taking as the weight to control the MAE loss of the output of the UE module. Pixels considered to be in the uncertainty region will be given a larger weight for penalty.

[0082] Uncertainty constraint loss The calculation formula is as shown in (17): (17) where, D represents the total number of all pixels, k represents the pixel label, represents downsampling the clear image to the scale of 0.25, represents UE_ i ( i = 1, 2, 3)the average predicted restored map output, represents UE_i ( i = 1, 2, 3) Output image uncertainty.

[0083] The present invention extracts the features of the clear image, the degraded image, and the restored image at multiple scales by means of a Vgg16 feature extractor. A total of three scales are based on: 0.5, 0.25, 0.125 to calculate the contrast loss . The calculation formula is as shown in (18): (18) Among them, represents the clear image, represents the restored image output by the network, represents the rain and fog image, and Vgg represents the Vgg16 feature extractor. In a specific experiment, , and are respectively input into the Vgg16 feature extractor, and their features at the scales of 0.5, 0.25, and 0.125 are respectively extracted for calculation.

[0084] Structural similarity (SSIM) loss The calculation method is as shown in (19): (19) Among them, J represents the clear reference image, I represents the input rain and fog image, is the restored image output by the network.

[0085] Frequency loss function The calculation formula is as follows: (20) The calculation method of is: first perform a two-dimensional fast Fourier transform (fft) on the restored image of the network to obtain a complex vector , and then splice the real part vector and the imaginary part vector of the complex vector on the last dimension to obtain the frequency vector of the output image. Similarly,

[0086] Through the present invention, the limitations existing in the single-image rain and fog removal tasks based on deep learning are alleviated, and good restoration effects are achieved on the long-distance blur and color degradation caused by rain and fog effects in the publicly available rain and fog datasets Outdoor-Rain, RainCityscapes, and DQA datasets, while retaining the edge and texture details of the image.

[0087] The present application also discloses a rain and fog image restoration system based on frequency attention and uncertainty guidance network, including: An acquisition unit for acquiring the rain and fog image to be restored; A processing unit for inputting the rain and fog image to be restored into the frequency attention and uncertainty guidance network model, and after sequentially passing through feature encoding, feature refinement enhancement, and confidence feature feedback decoding processing, obtaining the restored rain and fog image.

[0088] The present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the rain and fog image restoration method based on the frequency attention and uncertainty guidance network described in any one of the above are implemented.

[0089] The present application also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the rain and fog image restoration method based on the frequency attention and uncertainty guidance network described in any one of the above are implemented.

[0090] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0091] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0092] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions in the processFigure 1 one process or multiple processes and / or boxes Figure 1 the functions specified in one box or multiple boxes.

[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or boxes Figure 1 one box or multiple boxes.

[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A rain and fog image restoration method based on frequency attention and uncertainty-guided network, characterized in that Including: Obtain a rain and fog image to be restored; Input the rain and fog image to be restored into a frequency attention and uncertainty-guided network model. After passing through feature encoding, feature refinement enhancement, and confidence feature feedback decoding processes in sequence, a restored rain and fog image is obtained.

2. The rain and fog image restoration method based on frequency attention and uncertainty guidance network according to claim 1, wherein, After obtaining the rain and fog image to be restored, preprocess the rain and fog image, specifically including: Crop the rain and fog image to be restored to obtain a cropped image dataset; Divide the cropped image dataset into a training set and a test set.

3. A rain and fog image restoration method based on a frequency attention and uncertainty guidance network according to claim 1, characterized in that The frequency attention and uncertainty-guided network model includes: An encoding unit for respectively extracting the initial features and shallow features of the input sample image, concatenating the initial features and shallow features with the same scale and number of channels to obtain different-scale features, and respectively transmitting the different-scale features to the decoding unit; performing convolutional processing on the different-scale features to obtain encoded features, and respectively transmitting the encoded features to the decoding unit and the feature refinement module; A feature refinement unit for receiving the encoded features, refining and enhancing the encoded features to obtain detail-enhanced features, and sending the detail-enhanced features to the decoding unit; A Laplacian pyramid structure skip connection unit for connecting the encoding unit and the decoding unit, calculating the residual information between different-scale features, and cascading the residual information to the decoding unit; A confidence feature feedback unit for obtaining optimized features by performing feedback confidence optimization processing on the fused features; and transmitting the optimized features to the decoding unit; A decoding unit for receiving the encoded features, detail-enhanced features, and optimized features, concatenating the received encoded features and detail-enhanced features to obtain fused features, stacking the fused features and the optimized features to obtain feedback fused features, superimposing the feedback fused features and the residual information to obtain output features, and decoding the output features to obtain a restored rain and fog image.

4. A rain and fog image restoration method based on a frequency attention and uncertainty guidance network according to claim 3, wherein In the feature refinement unit, the refinement and enhancement of the encoded features specifically include: Pass the encoded features through uncertainty estimation to obtain a confidence map and an uncertainty map; pass the encoded features through uncertainty local-global feature extraction to obtain local features and global features; The local features are guided by the uncertainty map to obtain uncertainty local features, the overall features refined from the uncertainty local features are obtained according to the uncertainty local features, the difference between the overall features and the encoded features is weighted based on the confidence map to obtain a weighted difference, and attention-enhanced features are obtained based on the weighted difference; The attention-enhanced features are processed through frequency component decomposition and fusion, and a multi-branch compact selection frequency module MCSF to obtain detail-enhanced features.

5. A method for restoring rainy and foggy images based on a frequency attention and uncertainty guidance network according to claim 4, characterized in that, Specifically, the encoded features passing through uncertainty estimation to obtain a confidence map and an uncertainty map include: The encoded features are sequentially processed through channel shuffling and channel grouping to obtain a number of feature layer groups; The number of feature layer groups respectively pass through convolutional subnets with shared weights to obtain a number of convolutional output results; Calculate the mean and standard deviation of the number of convolutional output results, and the standard deviation is an estimate of the image uncertainty; Solve the confidence map and the uncertainty map according to the estimate of the image uncertainty: Among them, is the confidence map, is the uncertainty map, represents a fixed constant, 1 represents the corresponding size tensor with all values being 1, and Tanh represents the nn.Tanh() activation function; is the estimate of image uncertainty.

6. The rain and fog image restoration method based on frequency attention and uncertainty guidance network according to claim 3, wherein, In the confidence feature feedback unit, obtaining optimized features by performing feedback confidence optimization processing on the fused features specifically includes: The fused features are processed through uncertainty processing to obtain a confidence map; the confidence map is processed through dimensionality reduction to obtain a query vector for querying the confidence features in the fused features; and the fused features are processed through mapping dimensionality reduction to obtain key vectors and value vectors; Determine the confidence feedback features based on the query vectors, key vectors, and value vectors; Stack the confidence feedback features and concatenate them with the fused features to obtain optimized features.

7. A rain and fog image restoration method based on a frequency attention and uncertainty guidance network according to claim 3, characterized in that The Laplacian pyramid structure skip connection unit calculates the residual information between features of different scales and superimposes the residual information on the decoding unit, which specifically includes: Calculate the residual information between features of different scales: Superimpose the calculated residual information and the feedback fusion feature of the corresponding coding unit to obtain the output feature of the coding unit B : In the formula, represents the residual information between different scale features; represents the scale of , and the number of channels is features; represents the scale of , and the number of channels is C features; ConvTranspose represents transposed convolution.

8. A rain and fog image restoration system based on frequency attention and uncertainty guidance network, characterized in that, Including: An acquisition unit for acquiring the rainy and foggy image to be restored; A processing unit for inputting the rainy and foggy image to be restored into a frequency attention and uncertainty-guided network model, and obtaining the restored rainy and foggy image after feature encoding, feature refinement enhancement, and confidence feature feedback decoding processing in sequence.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the rainy and foggy image restoration method based on a frequency attention and uncertainty-guided network according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the steps of the rainy and foggy image restoration method based on a frequency attention and uncertainty-guided network according to any one of claims 1-7 are implemented.