Deep learning defogging method based on feature grid affine transformation
Through a deep learning method based on feature grid affine transformation, high-resolution images acquired by driverless cars in foggy traffic environments are defogged, which solves the problems of slow fogging speed and incomplete effects in the prior art, and achieves fast and effective image defogging.
Patent Information
- Application Number
- CN202510320509.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art has slow fog removal speed for high-resolution images in foggy traffic environments in driverless cars, and the fog removal effect is not thorough, making it difficult to meet the demand for real-time fog removal.
Demist treatment of fogging images is achieved by using a deep learning defogging method based on feature grid affine transformation by fitting the image feature grid model, full resolution feature map and RGB three-color channel feature map through multi-convolution neural network branches.
It realizes rapid fog removal for high-resolution images, achieves good fog removal effect, while maintaining the naturalness and harmony of the image, meeting the real-time fog removal needs of driverless cars in foggy traffic environments.
Smart Images

Figure CN120163723A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more particularly to a deep learning defogging method based on feature grid affine transformation. Background Art
[0002] The recognition accuracy of the target by a driverless vehicle in a traffic scene on a foggy day is relatively low. In order to improve the recognition accuracy of the driverless vehicle in a foggy traffic environment and enhance the driving safety of the driverless vehicle in a foggy traffic environment, it is necessary to perform defogging processing on the images obtained by the visual sensor in a haze weather.
[0003] The image defogging technology based on the physical model has poor defogging effect and slow defogging speed, and it is difficult to meet the needs of real-time defogging tasks. The current defogging methods based on deep learning technology have problems such as incomplete defogging and color distortion. When facing the defogging task of high-resolution images, the processing speed is slow, and it is difficult to meet the needs of real-time defogging of driverless vehicles. Summary of the Invention
[0004] The object of the present invention is to provide a deep learning defogging method based on feature grid affine transformation, which solves the problems disclosed in the background art. The method includes: obtaining a foggy image through the visual sensor of the driverless vehicle, and using the foggy image to fit an image feature grid model, a full-resolution feature map, and an RGB three-color channel feature map through three convolutional neural network branches, and using the image feature grid model, the full-resolution feature map, and the RGB three-color channel feature map to process the foggy image into a fog-free image through a fourth convolutional neural network branch.
[0005] To achieve the above object, the present invention provides a deep learning defogging method based on feature grid affine transformation, including the following steps:
[0006] Step S1: Input the original haze image into the first convolutional neural network to reduce the resolution of the original haze image and construct a feature grid;
[0007] Step S2: Input the original haze image into the second convolutional neural network to extract a full-resolution feature map;
[0008] Step S3: The third convolutional neural network fits an RGB three-channel feature map from the RGB three color channels of the original haze image;
[0009] Step S4: The fourth convolutional neural network processes the feature grid, the full-resolution feature map, and the RGB three-channel feature map to obtain a defogged image.
[0010] Preferably, the resolution of the original haze image is greater than 1080p.
[0011] Preferably, in step S1, the original haze image is input into the first convolutional neural network to reduce the resolution of the original haze image and construct a feature grid. The specific operations are as follows:
[0012] Step S11: Reduce the resolution of the original haze image to 256×256 with 3 channels, extract features through the UNet network and perform feature fusion to generate the feature map x 10 , and the processing process of the UNet network is as follows:
[0013] x1 = ReLU(BatchNorm(Conv(x, W1, b1)));
[0014] x2 = Down(x1) = MaxPool(ReLU(BatchNorm(Conv(x1, W2, b2))));
[0015] x3 = Down(x2), x4 = Down(x3), x5 = Down(x4);
[0016]
[0017] x7 = UP(x6, x3), x8 = UP(x7, x2), x9 = UP(x8, x1);
[0018] x 10 = Sigmoid(Conv(x9, W out , b out ));
[0019] Among them, x represents the haze image with a resolution of 256×256; x1, x2, x3, x4, x5, x6, x7, x8, x9, x 10 all represent feature maps; Conv(.) represents the convolution operation, W1 represents the convolution kernel, b1 represents the bias term, ReLU(.) represents the ReLU activation function, W out represents the convolution kernel, b out represents the bias term, W6 represents the convolution kernel, b6 represents the bias term; BatchNorm represents the batch normalization layer; represents the skip connection; ConvTransposed(.) represents the transposed convolution operation; Sigmoid represents the activation function that limits the output value within the range of [0, 1]; MaxPool(.) represents the max pooling; UP(.) represents the upsampling; Down(.) represents the downsampling; DoubleConv(.) represents;
[0020] Step S12: The feature map x extracted from the UNet network 10A feature grid is generated through a downsampling operation and an operation to change the shape, obtaining the feature grid. The specific process is expressed as:
[0021] y = f downsample (x 10 , (64, 256));
[0022] Grid = reshape(y, (12, 16, 16, 16));
[0023] Among them, f downsample represents the downsampling operation, which downsamples the feature map x 10 to the feature tensor y of (64, 256); Grid represents the generated feature grid, whose shape is (12, 16, 16, 16), and reshape(.) represents the operation to change the shape.
[0024] Preferably, in step S2, the haze image is input into the second convolutional neural network to extract the full-resolution feature map. The specific operation is as follows:
[0025] Step S21: Reduce the resolution of the haze image to 320×320, with the number of channels being 3. Extract features through the UNet network and perform feature fusion to generate the feature map x1'0;
[0026] Step S22: Process the feature map x 10 ' to obtain the full-resolution feature map. The processing process is expressed as follows:
[0027] x u = Interpolate(x1'0, H, W);
[0028]
[0029] c = PReLU(x u ');
[0030] Among them, x u is the feature map obtained by upsampling x1'0 to the same size as the original haze image. Interpolate(·) represents the interpolation function, which is the method to implement upsampling; H and W respectively represent the length and width of the original haze image; W conv represents the 1×1 convolutional kernel; b conv represents the bias term; represents the convolution operation; x u ' is to perform a linear transformation on the channels of each pixel of x u without changing the spatial dimensions, and the number of channels is 3; PReLU(·) represents the activation operation using the PReLU activation function; c represents the full-resolution feature map; Conv2D(.) represents the two-dimensional convolution operation.
[0031] Preferably, in step S3, the third convolutional neural network fits RGB three-channel feature maps from the RGB three color channels of the original haze image, and the specific operations are as follows:
[0032] n1 = Conv(input, W c1 , b c1 );
[0033] n1' = ReLU(BatchNorm(n1));
[0034] n2 = Conv(n1', W c2 , b c2 );
[0035] n2' = ReLU(BatchNorm(n2));
[0036] n3 = Conv(n2', W c3 , b c3 );
[0037] n3' = Tanh(n3);
[0038] Among them, input is the input single color channel; W c1 , W c2 , W c3 are the convolution kernels in the corresponding convolution operations respectively; b c1 , b c2 , b c3 are the bias terms in the corresponding convolution operations respectively; n1, n1', n2, n2', n3, n3' are the feature maps output after the above formula operations respectively; BatchNorm(·) represents performing batch normalization operation; Tanh(·) represents performing activation operation using the Tanh activation function;
[0039] The feature maps obtained after performing the above operations on the three color channels of the original haze image are denoted as x R ', x G ', x B '.
[0040] Preferably, in step S4, the fourth convolutional neural network processes the feature grid, the full-resolution feature map, and the RGB three-channel feature map to obtain the dehazed image, and the specific operations are as follows:
[0041] Step S41: Map the feature grid Grid to the obtained x R ', x G ', x B ' three-color channel feature maps, x R ', x G ', xB The height and width of the feature map are H and W respectively. Let m be the same as x R ′, x G ′, x B ′ feature maps with the same size and the same number of channels. The specific operation of the mapping process is expressed as:
[0042] hg(i, j) = i, i = 0, 1,..., H - 1;
[0043] wg(i, j) = j, j = 0, 1,..., W - 1;
[0044]
[0045] hg″ = 2·hg′ - 1, wg″ = 2·wg′ - 1;
[0046] guidemap = [wg″, hg″, m];
[0047] m′ = Coeff(guidemap, Grid);
[0048] Among them, hg and wg represent feature maps with the same size as m. The elements of hg are the width of m, and the elements of wg represent the height of m; hg(i, j) represents the element at coordinates (i, j) in hg; wg(i, j) represents the element at coordinates (i, j) in wg; i and j respectively represent the abscissa and ordinate of each element in the feature map; hg′ and wg′ represent the feature maps obtained after normalizing the values of the elements in the hg and wg feature maps to the range [0, 1]; hg″ and wg″ represent the feature maps obtained after mapping the range of the elements in the hg′ and wg′ feature maps to the range [-1, 1]; guidemap represents a tensor obtained by concatenating hg″, wg″ and m; Coeff(·) represents a trilinear interpolation operation; m′ represents the feature map tensor obtained by mapping Grid to m;
[0049] Step S42: After operating on x R ′, x G ′, x B ′ respectively according to the operation process of step S41, x R ″, x G ″, x B ″ are obtained. Then, an affine transformation operation is performed on x R ″, x G ″, x B ″ and the full-resolution feature map c obtained in step S21. Let q be the same as x R ″, x G ″, x BFeature maps with the same size and the same number of channels. The specific operation in the affine transformation process is expressed as:
[0050]
[0051] q' = Concat(α, β, γ);
[0052] Among them, represents a feature map tensor with a height and width of H and W respectively and 12 channels; q z represents the feature map of the z-th channel of q; Concat(·) represents the concatenation operation; q' represents the feature map tensor obtained by concatenating the three feature maps α, β, and γ.
[0053] Step S43: Process x R ″, x G ″, x B ″ respectively according to the operation process of step S42 to obtain x R ″′, x G ″′, x B ″′. Then, perform convolution and feature fusion operations on x R ″′, x G ″′, x B ″′ to obtain the final output fog-free image. The specific operation process is expressed as:
[0054] P = Concat(x R ″′, x G ″′, x B ″′);
[0055] P1 = PReLU(Conv(P, W o1 , b o1 ));
[0056] P2 = PReLU(Conv(P1, W o2 , b o2 ));
[0057] P3 = PReLU(Conv(P2, W o3 , b o3 ));
[0058] P4 = PReLU(Conv(P3, W o4 , b o4 ));
[0059] Output = PReLU(P4 ⊙ x - P4 + 1);
[0060] Among them, W o1 , W o2 , Wo3 , W o4 and b o1 , b o2 , b o3 , b o4 are respectively the convolution kernel and the bias term in the convolution operation; ⊙ represents element-wise multiplication; Output represents the dehazed image of the output; P represents the feature map after splicing x R ″, x G ″, x B ″; P1, P2, P3, and P4 respectively represent the feature maps obtained after the corresponding operations;
[0061] Step S44: The training dataset used for training is several pairs of hazy and haze-free pictures. During the training process, the L1 loss function is used to calculate the overall difference between the network output and the target image, and the parameters are transmitted and updated layer by layer through backpropagation;
[0062] Among them, the L1 loss function is as follows:
[0063]
[0064] Among them, k represents each pixel value of the image output by the network, that is, the output image of the hazy picture input into the network after being processed by the network; l represents each pixel value of the target image, that is, the clear haze-free image input during training; k r represents each pixel value of the network output; y i represents each pixel value of the target image; t represents the total number of pixels in the image; L1Loss(x, y) is the loss value of the loss function.
[0065] Therefore, the present invention adopts the above-mentioned deep learning dehazing method based on feature grid affine transformation, which can maintain the naturalness and harmony of the overall image while achieving a good dehazing effect, and at the same time has a very high image processing speed to meet the requirements of real-time image dehazing. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a system block diagram of a deep learning dehazing method based on feature grid affine transformation of the present invention;
[0067] Figure 2 is a network structure diagram of a deep learning dehazing method based on feature grid affine transformation of the present invention;
[0068] Figure 3 is a UNet network structure diagram of a deep learning dehazing method based on feature grid affine transformation of the present invention;
[0069] Figure 4Schematic diagram of the affine transformation of a deep learning haze removal method based on feature grid affine transformation according to the present invention;
[0070] Figure 5 Haze removal effect diagram of a deep learning haze removal method based on feature grid affine transformation according to the present invention; wherein, Figure 5 (a) in is the image before haze removal; Figure 5 (b) in is the image after haze removal. Detailed implementation manners
[0071] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0072] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.
[0073] Embodiment 1
[0074] As Figure 1 、 Figure 2 The present invention provides a method for removing haze from a hazy image. By using deep learning and neural networks, and through a multi-branch convolutional neural network, haze removal of a hazy image in a traffic environment is achieved. This method can quickly remove haze from high-resolution pictures, can achieve a good haze removal effect on high-resolution pictures, and at the same time meet the requirements of real-time haze removal for driverless cars, improving the recognition accuracy of driverless vehicles in a haze traffic environment. The specific implementation method is as follows:
[0075] Step S1: Input the original hazy image (resolution: 1080p.) into the first convolutional neural network to reduce the resolution of the original hazy image and construct a feature grid. The specific operations are as follows:
[0076] Step S11: Reduce the resolution of the original hazy image to 256×256, with 3 channels. Extract features through the UNet network and perform feature fusion to generate the feature map x 10 。As Figure 3 shown, the processing process of the UNet network is as follows:
[0077] x1 = DoubleConv(x) = ReLU(BatchNorm(Conv(x, W1, b1)));
[0078] x2 = Down(x1) = MaxPool(ReLU(BatchNorm(Conv(x1, W2, b2))));
[0079] x3 = Down(x2), x4 = Down(x3), x5 = Down(x4);
[0080]
[0081] x7 = UP(x6, x3), x8 = UP(x7, x2), x9 = UP(x8, x1);
[0082] x 10 = Sigmoid(Conv(x9, W out , b out ));
[0083] Among them, x represents the hazy image with a resolution of 256×256; x1, x2, x3, x4, x5, x6, x7, x8, x9, x 10 all represent feature maps; Conv(.) represents the convolution operation, W1 represents the convolution kernel, b1 represents the bias term, ReLU(.) represents the ReLU activation function, W out represents the convolution kernel, b out represents the bias term, W6 represents the convolution kernel, b6 represents the bias term; BatchNorm represents the batch normalization layer; represents the skip connection; ConvTransposed(.) represents the transposed convolution operation; Sigmoid represents the activation function that limits the output value within the range of [0, 1]; MaxPool(.) represents the max pooling. In Figure 3 "Downsampling 1" represents the DoubleConv(·) operation in the formula; "Downsampling 2" represents the Down(·) operation in the formula; "Upsampling" represents the UP(·) operation in the formula;
[0084] Step S12: Generate a feature grid from the feature map x 10 extracted from the UNet network through a downsampling operation and a shape-changing operation. The specific process is expressed as:
[0085] y = f downsample (x 10 , (64, 256));
[0086] Grid = reshape(y, (12, 16, 16, 16));
[0087] Among them, f downsample represents the downsampling operation that downsamples the feature map x 10 to the feature tensor y of (64, 256); Grid represents the generated feature grid with a shape of (12, 16, 16, 16), and reshape(.) represents the shape-changing operation.
[0088] Step S2: Input the original hazy image into the second convolutional neural network to extract the full-resolution feature map. The specific operations are as follows:
[0089] Step S21: Reduce the resolution of the haze image to 320×320 with 3 channels, extract features through the UNet network and generate a feature map x1'0 through feature fusion;
[0090] Step S22: Process the feature map x 10 ' to obtain a full-resolution feature map. The processing process is as follows:
[0091] x u = Interpolate(x1'0, H, W);
[0092]
[0093] c = PReLU(x u ');
[0094] Among them, x u is the feature map obtained by upsampling x1'0 to the same size as the original haze image. Interpolate(·) represents the interpolation function; H and W respectively represent the length and width of the original haze image; W conv represents a 1×1 convolution kernel; b conv represents the bias term; represents the convolution operation; x u ' is to perform a linear transformation on the channels of each pixel of x u without changing the spatial dimensions, with 3 channels; PReLU(·) represents the activation operation using the PReLU activation function; c represents the full-resolution feature map; Conv2D(.) represents the two-dimensional convolution operation.
[0095] Step S3: The third convolutional neural network fits out the RGB three-channel feature map from the RGB three color channels of the original haze image. The specific operations are as follows:
[0096] n1 = Conv(input, W c1 , b c1 );
[0097] n1' = ReLU(BatchNorm(n1));
[0098] n2 = Conv(n1', W c2 , b c2 );
[0099] n2' = ReLU(BatchNorm(n2));
[0100] n3 = Conv(n2', W c3 , b c3 );
[0101] n3′ = Tanh(n3);
[0102] where input is a single input color channel; W c1 , W c2 , W c3 are the convolution kernels in the corresponding convolution operations respectively; b c1 , b c2 , b c3 are the bias terms in the corresponding convolution operations; n1, n1′, n2, n2′, n3, n3′ are the feature maps output after the operations of the above formulas respectively; BatchNorm(·) represents performing batch normalization operation; Tanh(·) represents performing activation operation using the Tanh activation function;
[0103] The feature maps obtained after performing the above operations on the three color channels of the original hazy image are denoted as x R ′, x G ′, x B ′.
[0104] Step S4. The fourth convolutional neural network processes the feature grid, full-resolution feature map, and RGB three-channel feature map to obtain a dehazed image. The specific operations are as follows:
[0105] Step S41. Map the feature grid Grid to the obtained x R ′, x G ′, x B ′ three-color channel feature maps. The height and width of the x R ′, x G ′, x B ′ feature maps are H and W respectively. Let m be a feature map with the same size and number of channels as x R ′, x G ′, x B ′. The specific operations of the mapping process are expressed as:
[0106] hg(i,j) = i, i = 0,1,...,H - 1;
[0107] wg(i,j) = j, j = 0,1,...,W - 1;
[0108]
[0109] hg″ = 2·hg′ - 1, wg″ = 2·wg′ - 1;
[0110] guidemap = [wg″, hg″, m];
[0111] m′ = Coeff(guidemap, Grid);
[0112] Among them, hg and wg represent feature maps with the same size as m, where the elements of hg are the width of m, and the elements of wg represent the height of m; hg(i,j) represents the element at coordinates (i,j) in hg; wg(i,j) represents the element at coordinates (i,j) in wg; i and j respectively represent the abscissa and ordinate of each element in the feature map; hg′ and wg′ represent the feature maps obtained after normalizing the values of the elements in the hg and wg feature maps to the range [0,1]; hg″ and wg″ represent the feature maps obtained after mapping the range of the elements in the hg′ and wg′ feature maps to the range [-1,1]; guidemap represents a tensor obtained by concatenating hg″, wg″ and m; Coeff(·) represents a trilinear interpolation operation; m′ represents the feature map tensor obtained by mapping Grid to m;
[0113] Step S42: Operate on x R ′, x G ′, x B ′ respectively according to the operation process of step S41 to obtain x R ″, x G ″, x B ″, and then use x R ″, x G ″, x B ″ and the full-resolution feature map c obtained in step S21 to perform an affine transformation operation. Let q be a feature map with the same size and number of channels as x R ″, x G ″, x B ″. As shown in Figure 4 Figure, the specific operation of the affine transformation process is expressed as:
[0114]
[0115] q′ = Concat(α,β,γ);
[0116] Among them, represents a feature map tensor with a height and width of H and W respectively and 12 channels; q z represents the feature map of the z-th channel of q; Concat(·) represents a concatenation operation; q′ represents the feature map tensor obtained by concatenating the three feature maps of α, β, and γ,
[0117] Step S43: Process x R ″, x G ″, x B ″ respectively according to the operation process of step S42 to obtain x R ″′, x G ″′, x B ″′, and use xR ″′, x G ″′, x B ″′ undergoes convolution and feature fusion operations to obtain the final output fog-free image. The specific operation process is expressed as:
[0118] P = Concat(x R ″′, x G ″′, x B ″′);
[0119] P1 = PReLU(Conv(P, W o1 , b o1 ));
[0120] P2 = PReLU(Conv(P1, W o2 , b o2 ));
[0121] P3 = PReLU(Conv(P2, W o3 , b o3 ));
[0122] P4 = PReLU(Conv(P3, W o4 , b o4 ));
[0123] Output = PReLU(P4 ⊙ x - P4 + 1);
[0124] Among them, W o1 , W o2 , W o3 , W o4 and b o1 , b o2 , b o3 , b o4 are the convolution kernel and bias term in the convolution operation respectively; ⊙ represents element-wise multiplication; Output represents the de-fogged image of the output; P represents the feature map after splicing x R ″, x G ″, x B ″; P1, P2, P3, P4 represent the feature maps obtained after the corresponding operations respectively;
[0125] Step S44: The training dataset used for training is several pairs of foggy and fog-free pictures. During the training process, the L1 loss function is used to calculate the overall difference between the network output and the target image, and the parameters are passed and updated layer by layer through backpropagation.
[0126] After building the neural network model according to the above steps, the training of the model weights is carried out next. The training dataset used for training this deep learning dehazing method based on feature grid affine transformation is several pairs of hazy and haze-free images. During the forward propagation of training, the input data passes through each layer in the neural network, and is calculated layer by layer from the input layer to the output layer until the predicted output of the model is obtained. Then the loss is calculated and the backpropagation is carried out. The backpropagation starts from the output layer and calculates the gradient of the loss function with respect to each parameter (weight and bias) layer by layer backward. The chain rule is used to calculate the gradient of each layer's parameters. Through the output and gradient information of each layer, it is passed and updated layer by layer. The loss function used in this method is the L1 loss function (mean absolute error), and the specific operation process is as follows. The L1 loss function is as follows:
[0127]
[0128] Among them, k represents each pixel value of the image output by the network, that is, the output image of the hazy image passed into the network after being processed by the network; l represents each pixel value of the target image, that is, the clear haze-free image passed in during training; k r represents each pixel value output by the network; y i represents each pixel value of the target image; t represents the total number of pixels in the image; L1Loss(x, y) is the loss value of the loss function.
[0129] As Figure 5 shown, the hazy image input to the dehazing network is first a 1080p 30FPS road traffic image taken by a drone, and then a hazy image obtained by performing a synthetic fog operation on the haze-free image using the synthetic fog method based on the atmospheric scattering model theory ( Figure 5 (a) in). Then, the hazy image is input into the dehazing network frame by frame for dehazing, Figure 5 (b) in is the comparison picture before and after dehazing.
[0130] Therefore, the present invention adopts the above-mentioned deep learning dehazing method based on feature grid affine transformation, which can maintain the naturalness and harmony of the overall image while achieving a good dehazing effect, and at the same time has a very high image processing speed to meet the requirements of real-time image dehazing.
[0131] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A deep learning defogging method based on feature grid affine transformation, characterized in that: The following steps are involved: Step S1, inputting the original haze image into the first convolutional neural network, reducing the resolution of the original haze image and constructing a feature grid; Step S2: inputting the original haze image into the second convolutional neural network to extract a full-resolution feature map; Step S3, the third convolutional neural network fits the RGB three-channel feature map from the RGB three color channels of the original haze image; Step S4: The fourth convolutional neural network processes the feature grid, the full-resolution feature map, and the RGB three-channel feature map to obtain a dehazed image.
2. According to claim 1, a deep learning defogging method based on feature grid affine transformation is characterized in that: The resolution of the original haze image is greater than 1080p.
3. The deep learning defogging method based on feature grid affine transformation according to claim 1 is characterized in that: In step S1, the original haze image is input into the first convolutional neural network, the resolution of the original haze image is reduced and a feature grid is constructed. The specific operations are as follows: Step S11: Reduce the resolution of the original haze image to 256×256, the number of channels is 3, extract features through the UNet network, and fuse the features to generate a feature map x 10 , the processing of the UNet network is as follows: x1=ReLU(BatchNorm(Conv(x,W1,b1))); x2=Down(x1)=MaxPool(ReLU(BatchNorm(Conv(x1,W2,b2)))); x3=Down(x2), x4=Down(x3), x5=Down(x4); x7=UP(x6,x3), x8=UP(x7,x2), x9=UP(x8,x1); x 10 =Sigmoid(Conv(x9,W out ,b out )); Among them, x represents the haze image with a resolution of 256×256; x1, x2, x3, x4, x5, x6, x7, x8, x9, x 10 All represent feature maps; Conv(.) represents the convolution operation, W1 represents the convolution kernel, b1 represents the bias term, ReLU(.) represents the ReLU activation function, W out represents the convolution kernel, b out represents the bias term, W6 represents the convolution kernel, b6 represents the bias term; BatchNorm represents the batch normalization layer; Indicates skip connection; ConvTransposed(.) indicates transposed convolution operation; Sigmoid indicates activation function, which limits the output value to the range [0,1]; MaxPool(.) indicates maximum pooling; UP(.) indicates upsampling; Down(.) indicates downsampling; Step S12: extract the feature map x from the UNet network 10 After a downsampling operation and a shape-changing operation, a feature grid is generated to obtain a feature grid. The specific process is expressed as follows: y=f downsample (x 10 ,(64,256)); Grid=reshape(y,(12,16,16,16)); Among them, f downsample Represents the downsampling operation, which converts the feature map x 10 The feature tensor y is downsampled to (64, 256); Grid represents the generated feature grid, whose shape is (12, 16, 16, 16), and reshape(.) represents the operation of changing the shape.
4. The deep learning defogging method based on feature grid affine transformation according to claim 3 is characterized in that: In step S2, the haze image is input into the second convolutional neural network to extract the full-resolution feature map. The specific operations are as follows: Step S21, reduce the resolution of the haze image to 320×320, the number of channels is 3, extract features through the UNet network and fuse the features to generate a feature map x1'0; Step S22: transform the feature map x 10 'Processing is performed to obtain a full-resolution feature map. The processing process is as follows: x u =Interpolate(x1'0,H,W); c=PRELU(x u ′)? Among them, x u is to upsample x1'0 to a feature map of the same size as the original haze image. Interpolate(·) represents the interpolation function; H and W represent the length and width of the original haze image, respectively; W conv represents a 1×1 convolution kernel; b conv represents the bias term; represents the convolution operation; x u ′ is x u The channel of each pixel is linearly transformed without changing the spatial size. The number of channels is 3. PReLU(·) indicates activation using the PReLU activation function. c indicates the full-resolution feature map. Conv2D(.) indicates a two-dimensional convolution operation.
5. The deep learning defogging method based on feature grid affine transformation according to claim 4 is characterized in that: In step S3, the third convolutional neural network fits the RGB three-channel feature map from the RGB three color channels of the original haze image. The specific operation is as follows: n1=Conv(input,W c1 ,b c1 ); n1′=ReLU(BatchNorm(n1)); n2=Conv(n1′,W c2 ,b c2 ); n2′=ReLU(BatchNorm(n2)); n3=Conv(n2′,W c3 ,b c3 ); n3′=Tanh(n3); Among them, input is a single color channel of the input; W c1 , W c2 , W c3 are the convolution kernels in the corresponding convolution operations; b c1 、b c2 、b c3 is the bias term in the corresponding convolution operation; n1, n1′, n2, n2′, n3, n3′ are the feature maps output after the above formula operation; BatchNorm(·) indicates batch normalization operation; Tanh(·) indicates activation operation using Tanh activation function; The feature map obtained by performing the above operations on the three color channels of the original haze image is represented as x R ′、x G ′、x B ′.
6. The deep learning defogging method based on feature grid affine transformation according to claim 5 is characterized in that: In step S4, the fourth convolutional neural network processes the feature grid, the full-resolution feature map, and the RGB three-channel feature map to obtain a dehazed image. The specific operations are as follows: Step S41: Map the feature grid Grid to the obtained x R ′、x G ′、x B ′On the three color channel feature map, x R ′、x G ′、x B ′The height and width of the feature map are H and W respectively, let m be and x R ′、x G ′、x B ′For feature maps with the same size and the same number of channels, the specific operation of the mapping process is expressed as: hg(i,j)=i, i=0,1,...,H-1; wg(i,j)=j, j=0,1,...,W-1; hg″=2·hg′-1, wg″=2·wg′-1; guidemap = [wg″,hg″,m]; m′=Coeff(guidemap,Grid); Among them, hg and wg represent feature maps of the same size as m, where the elements of hg are the width of m and the elements of wg represent the height of m; hg(i,j) represents the element with coordinates (i,j) in hg; wg(i,j) represents the element with coordinates (i,j) in wg; i and j represent the horizontal and vertical coordinates of each element in the feature map, respectively; hg′ and wg′ represent the feature maps obtained by normalizing the values of the elements in the hg and wg feature maps to the range [0,1]; hg″ and wg″ represent the feature maps obtained by mapping the range of the elements in the hg′ and wg′ feature maps to the range [-1,1]; guidemap represents a tensor obtained by concatenating hg″, wg″ and m; Coeff(·) represents a trilinear interpolation operation; m′ represents the feature map tensor obtained by mapping Grid to m; Step S42: According to the operation process of step S41, x R ′、x G ′、x B ′After separate operations, we get x R ″、x G ″、x B ″, and then use x R ″、x G ″、x B ″ and the full-resolution feature map c obtained in step S21 perform an affine transformation operation, assuming that q is the same as x R ″、x G ″、x B For feature maps with the same size and the same number of channels, the specific operation of the affine transformation process is expressed as: q′=Concat(α,β,γ); in, Represented as a feature map tensor with a height and width of H and W and a channel number of 12; q z represents the feature map of the zth channel of q; Concat(·) represents the concatenation operation; q′ represents the feature map tensor obtained by concatenating the three feature maps α, β, and γ. Step S43: x R ″、x G ″、x B ″ are processed according to the operation process of step S42 to obtain x R ″′、x G ″′、x B ″′, x R ″′、x G ″′、x B After convolution and feature fusion operations, the final output fog-free image is obtained. The specific operation process is expressed as follows: P=Concat(x R ″′,x G ″′,x B ″′); P1=PReLU(Conv(P,W o1 ,b o1 )); P2=PReLU(Conv(P1,W o2 ,b o2 )); P3=PReLU(Conv(P2,W o3 ,b o3 )); P4=PReLU(Conv(P3,W o4 ,b o4 )); Output=PReLU(P4⊙x-P4+1); Among them, W o1 , W o2 , W o3 , W o4 and b o1 , b o2 , b o3 , b o4 They are the convolution kernel and bias term in the convolution operation respectively; ⊙ represents element-by-element multiplication; Output represents the output dehazed image; P represents x R ″、x G ″、x B ″Feature map after splicing; P1; P2; P3; P4 respectively represent the feature maps obtained after the corresponding operations; Step S44: The training data set used for training is a number of pairs of foggy and non-foggy pictures. During the training process, the L1 loss function is used to calculate the overall difference between the network output and the target image, and the parameters are transferred and updated layer by layer through back propagation; Among them, the L1 loss function is as follows: Where k represents the pixel value of the output image of the network, i.e., the foggy image input to the network after being processed by the network; l represents the pixel value of the target image, i.e., the clear fog-free image input during training; k r Represents the value of each pixel output by the network; y i Represents the value of each pixel of the target image; t represents the total number of pixels in the image; L1Loss(x,y) is the loss value of the loss function.
Citation Information
Patent Citations
Multi-scale connected image defogging algorithm based on UNet3+
CN113034445A
Image defogging method based on contextual information aggregation and fusion feature attention
CN115713473A
Single image defogging method and system based on pyramid efficient channel attention mechanism
CN116468625A
End-to-end rapid defogging detection combination method and device
CN118864325A
RGB-FIR multispectral image defogging method based on generative adversarial network
CN119090773A