Image defogging method and device

By building a dual-branch feature fusion network, combining global feature extraction and local feature supplementation, the problem of poor results of the existing technology in different foggy environments is solved, and a higher quality image fogging effect and computing efficiency are achieved.

CN120163731APending Publication Date: 2025-06-17JIANGXI UNIV OF SCI & TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510339882.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing image defogging method is not effective when dealing with different foggy environments, and the traditional method has defects such as over-enhancement, image edge distortion, and color misalignment.

Method used

Build a dual-branch feature fusion network combining global feature extraction and local feature supplementation, and enhance the learning and processing capabilities of complex images through multi-scale parallel modules and parallel attention feature fusion modules.

Benefits of technology

It significantly improves the image defogging performance, can understand image features more comprehensively and accurately, and the restored image details are more complete, making the defogging image closer to the original image, while improving computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163731A_ABST
    Figure CN120163731A_ABST
Patent Text Reader

Abstract

The invention discloses an image defogging method and device. The method comprises the steps that a double-branch feature fusion network combining global feature extraction and local feature supplementation is built; collecting a plurality of original images in a plurality of different application environments; processing the original image to generate a foggy image; taking the original image and the foggy image as a data set, and training the double-branch feature fusion network by using the data set to obtain an image defogging model; and carrying out defogging processing on the to-be-defogged image by using the image defogging model. By means of the scheme, foggy images in different environments can be processed, and the defogging performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image dehazing method and apparatus. Background Art

[0002] Image dehazing is crucial in computer vision and is used to solve the problem of atmospheric haze caused by particles such as water vapor, smoke, and dust. This haze reduces image contrast, blurs image details, and significantly degrades image quality, making tasks such as object detection, semantic segmentation, and autonomous driving complex. Most traditional image dehazing directions are based on prior knowledge to calculate the atmospheric diffraction model, calculate the ambient light and transmittance through the atmospheric diffraction model to solve the process of image degradation, and finally restore the image clarity through inverse calculation. However, dehazing images with prior methods will have defects such as over-enhancement, image edge distortion, and color misalignment.

[0003] In the prior art, there are mainly two categories of image dehazing methods: one is mainly based on physical prior models. Based on physical prior models, by estimating the atmospheric diffraction model more accurately, a dehazed restored image can be obtained. However, physical prior models are often designed and optimized for a certain type or several types of fog conditions, and cannot achieve ideal results in various foggy environments. When facing changing fog conditions, the model needs to be continuously adjusted or different prior strategies need to be replaced, and the application is not flexible and convenient enough. The other is based on deep learning methods. However, single deep network models have different focuses in image processing, so the disadvantages are relatively obvious and they cannot adapt to various different foggy environments. For example, the CNN network has excellent performance in extracting local features, but has insufficient grasp of global features, while the Transformer network has a multi-head attention mechanism and has excellent performance in parallel processing and extracting global features, but it is easy to ignore image details (such as edges, textures, etc.). It also cannot adapt to various different foggy environments. Summary of the Invention

[0004] The present invention provides an image dehazing method and apparatus, which can effectively process foggy images in different environments and significantly improve the dehazing performance.

[0005] To this end, the present invention provides the following technical solutions:

[0006] The present invention provides an image dehazing method, and the method includes:

[0007] Construct a dual-branch feature fusion network that combines global feature extraction and local feature supplementation;

[0008] Collect multiple original images in various different application environments;

[0009] Process the original images to generate foggy images;

[0010] Using the original image and the hazy image as a data set, training the dual-branch feature fusion network with the data set to obtain an image dehazing model;

[0011] Using the image dehazing model to perform dehazing processing on the image to be dehazed.

[0012] Optionally, the dual-branch feature fusion network includes: a global feature extraction module, a local feature supplement module, and a parallel attention feature fusion module respectively connected to the global feature extraction module and the local feature extraction module.

[0013] Optionally, the global feature extraction module includes a multi-scale parallel module, and the multi-scale parallel module expands the receptive field in a parallel manner using dilated convolutional kernels of multiple different sizes.

[0014] Optionally, the local feature supplement module uses convolutional layers and ReLU layers with dense residual connections to achieve the fusion of local features and the learning of local residuals, forming a continuous memory mechanism.

[0015] Optionally, the parallel attention feature fusion module includes two channel attention structures and two pixel attention structures;

[0016] The channel attention structure is used to extract global information and change the channel dimension of the features;

[0017] The pixel attention structure is used to extract information features related to position.

[0018] Optionally, the dual-branch feature fusion network further includes: a normalization processing layer.

[0019] Optionally, the method further includes:

[0020] Training the dual-branch feature fusion network on a general large-scale data set to obtain network initial parameters;

[0021] The training the dual-branch feature fusion network with the data set to obtain an image dehazing model includes:

[0022] Based on the network initial parameters, training the dual-branch feature fusion network with the data set to obtain an image dehazing model.

[0023] Optionally, the general large-scale data set at least includes: an indoor data set, an outdoor data set, and a comprehensive target test set.

[0024] The present invention also provides an image dehazing device, and the device includes:

[0025] A network construction module for building a dual-branch feature fusion network that combines global feature extraction and local feature supplementation;

[0026] An image collection module for collecting multiple original images in various different application environments;

[0027] An image processing module for processing the original images to generate hazy images;

[0028] A training module for using the original images and the hazy images as a data set to train the dual-branch feature fusion network with the data set to obtain an image dehazing model;

[0029] A dehazing module for dehazing the image to be dehazed using the image dehazing model.

[0030] Optionally, the device further includes:

[0031] A pre-training module for training the dual-branch feature fusion network on a general large data set to obtain network initial parameters;

[0032] The training module trains the dual-branch feature fusion network with the data set based on the network initial parameters to obtain an image dehazing model.

[0033] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps of the image dehazing method.

[0034] The image dehazing method and device provided by the present invention have the following beneficial effects:

[0035] The network structure of the image dehazing model adopts a dual-branch parallel feature fusion network, which can process the dependency relationship of global features and the supplementation of local features at the same level, thereby enhancing the learning and processing ability of complex images, enabling a more comprehensive and accurate understanding of image features, making the restored image details more complete, and making the dehazed image closer to the original image. Moreover, it can make full use of computing resources, improve computing efficiency, reduce the training and inference time of the model, and can better be applicable to scenarios that require real-time processing or processing of a large amount of image data in practical applications.

[0036] Furthermore, by adopting a parallel attention feature fusion module, different attentions process features from different angles. Channel attention can better extract and encode shared global information, and pixel attention can better extract and encode position-related information. By parallelly extracting the global and local information of the original features, it can more effectively process hazy images in different environments and can significantly improve the dehazing performance.

[0037] Furthermore, during model training, first train on a general large dataset to obtain the initial network parameters. Using these parameters as the starting point for parameters is beneficial to accelerating the training convergence speed of the image dehazing model and quickly finding a better solution for model parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 is a flowchart of the image dehazing method according to an embodiment of the present invention;

[0040] Figure 2 is a schematic structural diagram of a dual-branch feature fusion network according to an embodiment of the present invention;

[0041] Figure 3 is a specific example of a dual-branch feature fusion network according to an embodiment of the present invention;

[0042] Figure 4 is Figure 3 a schematic structural diagram of the global feature extraction module in the dual-branch feature fusion network shown;

[0043] Figure 5 is Figure 3 a schematic structural diagram of the local feature supplementation module in the dual-branch feature fusion network shown;

[0044] Figure 6 is Figure 3 a schematic structural diagram of the parallel attention feature fusion module in the dual-branch feature fusion network shown;

[0045] Figure 7 is a schematic structural diagram of one PA branch in the pixel attention according to an embodiment of the present invention;

[0046] Figure 8 is a schematic structural diagram of the channel attention according to an embodiment of the present invention;

[0047] Figure 9 is a comparison diagram of the dehazing effect of a foggy image using the image dehazing method according to an embodiment of the present invention;

[0048] Figure 10 is a schematic structural diagram of an image dehazing device according to an embodiment of the present invention;

[0049] Figure 11 is another schematic structural diagram of an image dehazing device according to an embodiment of the present invention. Specific Embodiments

[0050] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.

[0051] In order to enable those skilled in the art of this technology to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0052] With the development of deep learning and the emergence of technologies such as Convolutional Neural Network (CNN) and Transformer, the image dehazing technology has been greatly improved in feature extraction, end-to-end learning, and generalization ability. The CNN method can increase the receptive field of the image by increasing the depth and width of the network or using large convolutional kernels. After training with a large amount of data, CNN has strong robustness and adaptability to foggy images with different types, different environments, and different distributions. However, due to the existence of multiple layers of convolution, CNN requires a lot of resources and training time when training large models and a large amount of data, and there will be certain deficiencies in the real-time performance of the algorithm. The Transformer method, on the other hand, increases the parallel computing ability, greatly improves the computing efficiency, and has more obvious advantages when dealing with a large amount of data. In terms of feature extraction, Transformer introduces the Multi-Head Attention mechanism. Each "head" can learn different feature representations of the foggy image, and these different "heads" can finally be concatenated or fused together to obtain richer and more comprehensive feature information of the haze image. However, the multi-head parallel ability means that Transformer will occupy a large amount of memory space and has a large space complexity; and the parallel fusion calculation of the multi-head attention mechanism for global features may relatively weaken the local feature information. For example, the detailed information at the edges of the image may have disadvantages such as over-enhancement or blurring.

[0053] According to the atmospheric diffraction model: I(x) = J(x)t(x) + A(1 - t(x)) (where A is the global atmospheric light, t(x) is the medium transmittance map, and J(x) is the haze-free image), and I(x) is the foggy image. After obtaining A and t(x) using the network, J(x) is inversely derived: The dehazed image J(x) can be output.

[0054] To this end, an embodiment of the present invention provides an image dehazing method and device, and constructs a dual-branch feature fusion network (DBFF-Net) for global feature extraction and local feature supplementation. An image dehazing model is trained based on the dual-branch feature fusion dehazing network, and the image dehazing model is used to perform dehazing processing on the image to be dehazed, so as to obtain a high-quality dehazed image.

[0055] As Figure 1 shown, it is a flowchart of an image dehazing method provided by an embodiment of the present invention, including the following steps:

[0056] Step 101, construct a dual-branch feature fusion network that combines global feature extraction and local feature supplementation.

[0057] As Figure 2 shown, in a non-limiting embodiment, the dual-branch feature fusion network may include: a global feature extraction module, a local feature supplementation module, and a parallel attention feature fusion module respectively connected to the global feature extraction module and the local feature extraction module. Among them, the global feature extraction module and the local feature supplementation module respectively perform global feature extraction and local feature supplementation of the image, and the features extracted by the global feature extraction module and the local feature supplementation module are input in parallel to the parallel attention feature fusion module, and finally dehazing features with global features and taking into account local details of the image are output.

[0058] The global feature extraction module is mainly used to capture the overall information of the image and long-distance dependence relationships (such as the global atmospheric light A and the distribution of global color and brightness, etc.). In some embodiments, the global feature extraction module may include a multi-scale parallel module, and the multi-scale parallel module uses a parallel manner of multiple different-sized dilated convolution kernels to expand the receptive field.

[0059] The local feature supplementation module is mainly used to extract the detailed information of the image (such as the transmittance t(x), the edges and textures of objects, etc.). In some embodiments, the local feature supplementation module may use convolutional layers and RELU layers with dense residual connections to achieve local feature fusion and local residual learning, forming a continuous memory mechanism.

[0060] The parallel attention feature fusion module allows the network to process the dependence relationships of global features and local feature supplementation at the same level by parallelly processing different attention mechanisms, thereby enhancing the learning and processing capabilities for complex images. In some embodiments, the parallel attention feature fusion module may include two channel attention structures and two pixel attention structures; wherein, the channel attention structure is used to extract global information and change the channel dimension of the features; the pixel attention structure is used to extract position-related information features.

[0061] In the embodiments of the present invention, the structure of the dual-branch feature fusion network needs to have an appropriate depth to ensure sufficient extraction of features and effectively retain the features extracted at each layer.

[0062] Figure 3 FIG. shows a schematic diagram of a dual-branch feature fusion network in the embodiments of the present invention.

[0063] In this example, the dual-branch feature fusion network is composed of 5 down-samples, 5 up-samples, and a global residual layer connected through skip connections. The down-sampling gradually reduces the resolution layer by layer through multiple convolutional layers and pooling operations, reducing the size of the feature map, enabling the network to learn deeper features; the up-sampling gradually restores the resolution through deconvolution operations, making the output image size match the input image size; the skip connection helps the network to simultaneously focus on the features extracted from the shallow and deep layers, preventing information loss and improving image quality; the global residual layer receives the final dehazed feature A,t(x) and finally outputs a clear dehazed image.

[0064] Correspondingly, Figure 4 FIG. shows Figure 3 a schematic diagram of the structure of the global feature extraction module in the shown dual-branch feature fusion network.

[0065] Figure 4 In the shown example, the global feature extraction module uses parallel dilated convolutional kernels of three different sizes (large, medium, and small) to expand the receptive field to improve the ability to extract global complex features: the large and medium dilated convolutional kernels have long-range modeling and large receptive fields, and they can focus on the thick fog area and the overall features of the image (such as atmospheric light, etc.), and the small dilated convolution can focus on the small fog area and the details of the image (such as object edge textures, etc.).

[0066] At the same time, in combination with Figure 2 and Figure 4 illustrate the process of the global feature extraction module extracting features from the input image.

[0067] Let x be the original feature map. First, normalize it using batch normalization (BatchNorm), that is Batch normalization can accelerate network convergence, improve generalization ability, and prevent overfitting. Then obtain x1 through pointwise convolution (PWConv), and then obtain x2 through convolution (Conv, kernel size is 5). Next, perform three different sizes of dilated convolutions (DWDConv19, DWDConv13, DWDConv7) on x2 respectively, with a dilation rate of 3 for all of them, and concatenate (Concat) the results of these three convolutions in the channel dimension to obtain x3, and its feature dimension becomes three times that of x. The formulas for the above output data can be expressed as:

[0068]

[0069] x2 = Conv(x1);

[0070] x3 =

[0071] Concat(DWDConv19(x2), DWDConv13(x2), DWDConv7(x2));

[0072] Wherein, PWConv represents pointwise convolution; Conv represents convolution with a convolution kernel size of 5; DWDConv19 represents depthwise separable dilated convolution with a dilated convolution kernel size of 19 and a dilation rate of 3, DWDConv13 represents depthwise separable dilated convolution with a dilated convolution kernel size of 13 and a dilation rate of 3, DWDConv7 represents depthwise separable dilated convolution with a dilated convolution kernel size of 7 and a dilation rate of 3, and Concat represents concatenating features in the channel dimension. Finally, x3 passes through a multi-layer perceptron (including two pointwise convolutions with the activation function GELU) to convert its feature dimension to be the same as x and add it to x. Let the number of channels of x be C, then the number of channels of x3 is 3C. The multi-layer perceptron first inputs x3 into the first pointwise convolution, which can be understood as a fully connected layer in the channel dimension. By compressing information, the number of channels 3C is reduced to 2C and input into the GELU activation function to increase the non-linear transformation and improve the network's learning ability for complex features. The result is input into the last pointwise convolution to reduce the number of channels from 2C to C again, so as to be consistent with x; finally, it is added to x item by item. The formula of this multi-layer perceptron can be expressed as: y = x + PWConv(GELU(PWConv(x3))). The multi-layer perceptron can combine three different types of features to extract the overall features of the image more completely.

[0073] Correspondingly, Figure 5 shows Figure 3 a schematic structural diagram of a local feature supplementation module in the shown double-branch feature fusion network.

[0074] Figure 5 In the shown example, the local feature supplementation module selects a residual dense block, and uses the connection of a convolutional layer and a RELU layer with dense residual connections to achieve the fusion of local features and the learning of local residuals, thereby forming a continuous memory mechanism. The continuous memory mechanism is realized by transmitting the state of the previous local feature supplementation module to each layer of the current local feature supplementation module. It can not only read the state from the previous local feature supplementation module through the continuous memory mechanism, but also make full use of all layers therein through local dense links.

[0075] At the same time, combined with Figure 2 andFigure 5 Describe the process of the global feature extraction module extracting features from the input image.

[0076] While global feature extraction is being performed, the local feature supplementation module extracts local features from the original feature map for supplementing image details. Let F d-1 and F d be the input and output of the d-th local feature supplementation module respectively, both having G0 feature maps. The output F d,c of the c-th convolutional layer of the d-th local feature supplementation module can be expressed as:

[0077] F d,c = σ(W d,c [F d-1 , F d,1 … F d,c-1 ) ;

[0078] Among them, σ represents the RELU activation function, and W d,c is the weight of the c-th convolutional layer. Assume that F d,c is composed of G feature maps, and [F d-1 , F d,1 … F d,c-1 refers to the concatenation of the feature maps generated by the convolutional layers 1, …, c - 1 in the (d - 1)-th local feature supplementation module and the d-th local feature supplementation module, thus obtaining G0+(c - 1)×G feature maps.

[0079] The output of the previous local feature supplementation module and each layer has a direct connection to all subsequent layers, which not only preserves the feed-forward property but also extracts local dense features; then local feature fusion is applied to adaptively fuse the state of the previous local feature supplementation module and the entire convolutional layer in the current local feature supplementation module, that is, the feature maps of the (d - 1)-th local feature supplementation module are directly introduced into the d-th local feature supplementation module in a concatenated manner. After that, a 1×1 convolutional layer is introduced to adaptively control the output information for feature fusion, and the formula is as follows:

[0080]

[0081] Among them, represents the function of the 1×1 convolutional layer in the d-th local feature supplementation module, and F d,LF represents the features extracted by the d-th local feature supplementation module.

[0082] The final output of the d-th local feature supplementation module can be obtained through the following formula:

[0083] F d = F d-1 + F d,LF .

[0084] Through the densely connected convolutional layer, the local feature supplementation module can extract rich local features, which enables the network to make full use of the information of each layer, better capture the details and features in the image, and thus provide more useful information for subsequent image details.

[0085] Correspondingly, Figure 6 shows Figure 3 a schematic structural diagram of the parallel attention feature fusion module in the shown double-branch feature fusion network.

[0086] Figure 6 In the shown example, the parallel attention feature fusion module adopts two channel attentions and two pixel attentions. The channel attention structure can effectively extract global information and change the channel dimension of the features, and the pixel attention structure can effectively extract position-related information features, such as different fog distributions in the image.

[0087] The parallel attention feature fusion module mixes different types of attention mechanisms. Figure 6 In the shown example, two channel attention structures and two pixel attention structures are adopted.

[0088] Let x1 and x2 be the feature maps extracted by the global feature extraction module and the local feature supplementation module respectively, and use batch normalization (BatchNorm) to normalize them, that is

[0089] The two pixel attention structures are the same. For convenience of description, each pixel attention is called a PA branch. Figure 7 shows the schematic structural diagram of the pixel attention.

[0090] The two PA branches respectively extract position-related information features from the output information of the global feature extraction module and the local feature supplementation module, and their role is to calculate the weights of the pixel attention. The calculation formula is as follows:

[0091]

[0092] where PWConv represents pointwise convolution, Conv represents convolution with a convolution kernel size of 3, GELU represents the GELU activation function, Sigmoid represents the Sigmoid activation function, which is used to extract the global pixel gating feature. PA1 and PA2 represent the weights of the pixel attention, F S and F p represent the final outputs, which are the feature maps of the input feature maps x1 and x2 after being weighted by the corresponding pixel attentions respectively.

[0093] The two PA branches use PWConv-GELU-PWConv to fit the features.

[0094] The two channel attention structures are the same. For the convenience of description, each channel attention is called the CA branch. Figure 8 The structural schematic diagram of the channel attention is shown.

[0095] The two CA branches respectively extract the features of the entire channel from the output information of the global feature extraction module and the local feature supplement module as the global channel gating signals. The formula is as follows:

[0096]

[0097] Among them, GAP represents global average pooling, PWConv represents pointwise convolution, Conv represents convolution with a convolution kernel size of 3, GELU represents the GELU activation function, and Sigmoid represents the Sigmoid activation function. CA1 and CA2 represent the weights of the channel attention, F c and F D represent the final outputs, which are the feature maps obtained by weighting the input feature maps x1 and x2 with the corresponding channel attention respectively.

[0098] The two CA branches use global average pooling, PWConv-GELU-PWConv, and the Sigmoid function to extract the global channel gating features.

[0099] Finally, the four attention processing results are concatenated and input into a multi-layer perceptron. The multi-layer perceptron reduces the channel dimension of the concatenated features to the same dimension as the input image feature x and adds it to x to obtain the final global feature, that is:

[0100] F = Concat(F S , F P , F C , F D );

[0101] y = x + (PWConv(GELU(PWConv(F)).

[0102] The atmospheric light A is a shared global variable, while the transmittance t(x) is a local variable dependent on the position. Channel attention can better extract the shared global information and encode the atmospheric light A, and pixel attention can better extract the position-dependent information and encode t(x). This parallel attention feature fusion module can better fuse the defogging features.

[0103] Step 102, collect multiple original images in a variety of different application environments.

[0104] For example, a camera can be used to take a number of (e.g., 50) photos in multiple different application environments. When taking photos, sufficient light needs to be ensured to avoid photos with solid color blanks.

[0105] Step 103: Process the original image to generate a foggy image.

[0106] Process the obtained images to have a unified resolution, and use software for synthesizing fog to synthesize the images with fog of different concentrations to generate foggy images. For example, there are three kinds of fog concentrations: low concentration (10%-30%), medium concentration (30%-50%), and high concentration (50%-70%). Use these three fog concentrations to synthesize the original images respectively to obtain 150 fog-synthesized images, that is, foggy images.

[0107] Step 104: Use the original image and the foggy image as a data set, and use the data set to train a double-branch feature fusion network to obtain an image defogging model.

[0108] Make corresponding annotations for the synthesized foggy images and the original images to obtain an image defogging data set in which the original images and the foggy images correspond one by one. The annotations made for the original images are mainly used to make the foggy images and their original images correspond one by one, and the present invention does not limit the annotation method. For example, use the naming of the pictures as annotations. For example, an original image is named A, and the foggy image generated corresponding to A is named A1; another original image is named B, and the foggy image corresponding to B is named B1. As long as the annotations can distinguish different pictures and their corresponding synthesized foggy images.

[0109] Furthermore, the image defogging data set can also be divided into a training set and a test set according to a ratio of 7:3. The training set is used to train the defogging model, and the validation set is used to test the performance of the model in real time and determine whether to terminate the training.

[0110] Before the training starts, set the loss function, initial learning rate, maximum number of training iterations, learning rate adjustment strategy, parameter optimization strategy based on gradient descent, weight decay strategy, data augmentation strategy, etc.

[0111] Among them, the loss function can select the perceptual loss with strong robustness in image processing and capable of ensuring the overall quality of the image and the perceptual contrast loss with good performance in retaining the texture details of the image. According to the atmospheric diffraction model: I = Jt + A(1 - t), where I is the synthesized fog image, J is the original image, A is the global atmospheric light, and t is the transmittance. Let the defogged image output by the DBFF-NET model be The formula of the loss function can be expressed as:

[0112]

[0113] Among them, φ i (i = 0, 1, 2…n) represents extracting the i-th layer features from the saved pre-trained model. represents the mean squared error, D(x, y) is a distance metric function for measuring the xy feature space, ω i is the weight coefficient, and β is a hyperparameter that balances the perceptual contrast loss and the perceptual loss.

[0114] In a non-limiting embodiment, the image (including the original image and the foggy image) can be cropped into a format of 256×256 size, and the weight coefficient ω i is sequentially set to 1, and β is set to 0.1.

[0115] For example, the exponential decay rates β1 and β2 can be optimized to 0.9 and 0.999 respectively using the AdamW optimizer. AdamW is an improved Adam optimization algorithm that modifies the weight decay component of Adam so that the weight decay is no longer added to the gradient but directly updates the parameters, which can bring better training stability and generalization performance. At the same time, the initial learning rate is set to 2×10 -4 , and the cosine annealing strategy is used to gradually reduce the initial learning rate to 2×10 -6 .

[0116] During the model training process, after each training cycle on the training dataset, the changing trends of the PSNR and SSIM of the model can be detected using the validation set. If the PSNR and SSIM keep increasing, the model continues to be trained; otherwise, the training is terminated.

[0117] Since the parameters of the base model are pre-trained based on a general dataset, using these parameters as the parameter starting point is beneficial to accelerating the training convergence of the model and finding a better model parameter solution. The PSNR of the image dehazing model obtained after training can reach 37.26 on the validation set, and the SSIM can reach 0.983.

[0118] In a non - restrictive embodiment, the dual - branch feature fusion network can also be first trained on a general large - scale dataset to obtain initial network parameters. For example, but not limited to, the RESIDE dataset, which is widely recognized in image de - hazing, can be preferentially selected for pre - training. The RESIDE dataset includes an indoor dataset (ITS), an outdoor dataset (OTS), a comprehensive target test set (SOTS), etc. Among them, the indoor dataset (ITS) includes 13,990 pairs of images, the outdoor dataset (OTS) includes 313,950 pairs of images, and the comprehensive target test set (SOTS) includes 500 pairs of indoor datasets and 500 pairs of outdoor datasets. This dataset is a general dataset, which is easy to obtain and can simulate hazy conditions in various complex environments. The data volume is large enough to significantly improve the adaptability of the model in complex environments. After the DBFF - Net model is pre - trained on this dataset, the obtained basic model can adapt to the image de - hazing ability of various environments.

[0119] For example, randomly select 12,800 pairs of images from the ITS dataset for training, randomly select 500 pairs of images from the SOTS dataset for testing. Set the batch size to 32 and the number of training epochs to 400. Use PSNR (Peak Signal - to - Noise Ratio) and SSIM (Structural Similarity Index) as evaluation criteria to evaluate and adjust the model, so that the PSNR of the model on the SOTS validation set can reach above 35 dB and the SSIM can reach above 0.95. After training, save the parameters of the feature extraction structure of the pre - trained DBFF - Net model to a binary file.

[0120] Taking the initial network parameters obtained by pre - training as the parameter starting point, and then using the dataset to train the dual - branch feature fusion network to obtain an image de - hazing model is beneficial to accelerating the training convergence speed of the image de - hazing model and quickly finding a better model parameter solution.

[0121] Step 105, use the image de - hazing model to perform de - hazing processing on the image to be de - hazed.

[0122] Figure 9 A comparison diagram of the de - hazing effect of a hazy image using the image de - hazing method of the embodiment of the present invention is shown. Among them, (a) is the original image (i.e., the hazy image), and (b) is the image after de - hazing processing.

[0123] As Figure 9 can be seen, using the solution of the present invention, a hazy image can be effectively processed to obtain a high - quality de - hazed image.

[0124] The image defogging method provided by the present invention constructs a dual-branch feature fusion network that combines global feature extraction and local feature supplementation. This network can process the dependency relationship of global features and the supplementation of local features at the same level, thereby enhancing the learning and processing ability of complex images and greatly improving the computational efficiency. When training the model parameters, first train on a general large-scale dataset to obtain initial parameters, use the dual-branch feature fusion network based on the initial parameters as the basic model, and then collect corresponding images for different application environments to train the basic model to obtain an image defogging model, effectively improving the model quality and training efficiency. Using this defogging model to defog the image to be defogged can not only ensure the balance of the overall image color and brightness, but also supplement and enhance local details, effectively improving the defogging effect, increasing the computational efficiency, and ensuring the accuracy and efficiency of image defogging.

[0125] Correspondingly, an embodiment of the present invention further provides an image defogging device, as Figure 10 shown, which is a schematic structural diagram of the image defogging device.

[0126] In this example, the image defogging device 800 includes the following modules:

[0127] A network construction module 801, configured to construct a dual-branch feature fusion network that combines global feature extraction and local feature supplementation;

[0128] An image collection module 802, configured to collect multiple original images in various different application environments;

[0129] An image processing module 803, configured to process the original image to generate a foggy image;

[0130] A training module 804, configured to use the original image and the foggy image as a dataset, and use the dataset to train the dual-branch feature fusion network to obtain an image defogging model 80;

[0131] A defogging module 805, configured to use the image defogging model 80 to defog the image to be defogged.

[0132] As Figure 11 shown, which is another schematic structural diagram of the image defogging device provided by an embodiment of the present invention.

[0133] In this example, the image defogging device 800 further includes: a pre-training module 806, configured to train the dual-branch feature fusion network on a general large-scale dataset to obtain network initial parameters.

[0134] Accordingly, the training module 804 can train the dual-branch feature fusion network using the dataset based on the initial network parameters to obtain an image defogging model.

[0135] It should be noted that in specific implementation, the above training module 804 and pre-training module 806 can be the same module, that is, the above functions are implemented by the same physical entity, or implemented by different physical entities. The present invention does not limit this.

[0136] For more content about the above modules, reference can be made to the description in the method embodiments of the present invention above, and details will not be repeated here.

[0137] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0138] The present invention also provides a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program runs, it can execute Figure 1 some or all of the steps of the method shown. The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc, etc. The storage medium can also include a non-volatile memory or a non-transitory memory, etc.

[0139] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data provider to another website, computer, server, or data provider in a wired or wireless manner.

[0140] The above has introduced the embodiments of the present invention in detail. In this article, specific implementation manners are used to expound the present invention. The description of the above embodiments is only used to help understand the method and system of the present invention. They are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention. The content of this specification should not be construed as a limitation to the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image defogging method, characterized in that: The method comprises: Build a dual-branch feature fusion network that combines global feature extraction and local feature supplementation; Collect multiple original images from various application environments; Processing the original image to generate a foggy image; The original image and the foggy image are used as data sets, and the dual-branch feature fusion network is trained using the data sets to obtain an image defogging model; The image to be defogged is defogged using the image defogging model.

2. The image defogging method according to claim 1, characterized in that: The dual-branch feature fusion network includes: a global feature extraction module, a local feature supplementation module, and a parallel attention feature fusion module respectively connected to the global feature extraction module and the local feature extraction module.

3. The image defogging method according to claim 2, characterized in that: The global feature extraction module includes a multi-scale parallel module, which uses a plurality of hole convolution kernels of different sizes in parallel to expand the receptive field.

4. The image defogging method according to claim 2, characterized in that: The local feature supplementation module uses a convolutional layer and a RELU layer with dense residual connections to achieve local feature fusion and local residual learning, forming a continuous memory mechanism.

5. The image defogging method according to claim 2, characterized in that: The parallel attention feature fusion module includes two channel attention structures and two pixel attention structures; The channel attention structure is used to extract global information and change the channel dimension of the feature; The pixel attention structure is used to extract position-related informative features.

6. The image defogging method according to any one of claims 1 to 5, characterized in that: The method further comprises: Training the dual-branch feature fusion network on a general large-scale data set to obtain initial network parameters; The step of using the data set to train the dual-branch feature fusion network to obtain an image defogging model comprises: Based on the initial parameters of the network, the dual-branch feature fusion network is trained using the data set to obtain an image defogging model.

7. The image defogging method according to claim 6, characterized in that: The general large-scale data set includes at least: an indoor data set, an outdoor data set, and a comprehensive target test set.

8. An image defogging device, characterized in that: The device comprises: Network construction module, used to build a dual-branch feature fusion network that combines global feature extraction and local feature supplementation; An image collection module, used to collect multiple original images in various application environments; An image processing module, used for processing the original image to generate a foggy image; A training module, used to take the original image and the foggy image as data sets, and train the dual-branch feature fusion network using the data sets to obtain an image defogging model; The defogging module is used to perform defogging processing on the image to be defogged using the image defogging model.

9. The image defogging device according to claim 8, characterized in that: The device also includes: A pre-training module, used to train the dual-branch feature fusion network on a general large-scale data set to obtain initial network parameters; The training module trains the dual-branch feature fusion network based on the initial parameters of the network and uses the data set to obtain an image defogging model.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image defogging method according to any one of claims 1 to 7 are executed.

Citation Information

Cited By

  • Channel image defogging method and system, and computing device

    CN120807366A

  • A channel image defogging method, system, and computing device

    CN120807366B

  • Efficient image defogging method based on adaptive frequency enhancement and global-local feature aggregation

    CN120852232A

  • Cross-modal interaction conversion method and device for haze remote sensing image target detection

    CN121033643A

  • Power grid inspection data analysis method and system based on artificial intelligence

    CN121169759A