A semi-supervised image dehazing method

By building a semi-supervised image restoration network, combining multi-scale extraction and attention mechanism fusion module, the accuracy and robustness of image defogging processing in the prior art are solved, and efficient and reliable image automatic restoration effect is achieved.

CN114155165BActive Publication Date: 2025-07-01WENZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111436173.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-29
Publication Date
2025-07-01
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

In the prior art, image defogging processing has problems such as insufficient prediction accuracy and poor robustness. Especially when the real haze image lacks corresponding clear images as training data, the model generalization ability is poor, and the differences and complementarity of the global and local features and multi-scale features of the image are not fully explored.

Method used

Using a semi-supervised image defog method, a semi-supervised image recovery network is constructed, combined with supervised and unsupervised networks, a multi-scale extraction module and attention mechanism fusion module are used to perform feature separation and fusion, a high-dimensional deep learning feature is mapped using convolutional layers and pre-trained models, and a supervised and unsupervised loss function is trained to achieve automatic restoration of image features.

Benefits of technology

It achieves higher image fog removal efficiency, more reliable image quality, and can automatically restore clear foggy-free images, improving the model's generalization ability in real haze scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155165B_ABST
    Figure CN114155165B_ABST
Patent Text Reader

Abstract

The present invention provides a semi-supervised image defogging method, which includes: Step 1: Obtain an image data set, and preprocess the training set to obtain synthetic foggy images and real fog maps; Step 2: Construct an image restoration model based on deep learning, input the foggy images extracted after preprocessing into the image restoration model, perform extraction and analysis of image features, and obtain defogged image information; Step 3: Use the image restoration model to construct a semi-supervised image restoration network, and train the semi-supervised image restoration network to obtain a trained image restoration model; Step 4: Obtain an original image set composed of foggy images to be restored, and input the original image set into the trained semi-supervised image restoration network to obtain the final restored images. The present invention can achieve automatic image restoration, with higher image defogging efficiency and more reliable restored image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a semi-supervised image defogging method. Background Art

[0002] With the wide popularization and application of visual recognition technology, a large amount of image data needs to be analyzed and processed. Due to the interference of the acquisition environment, some of these image data are degraded in appearance due to weather factors. Among these degraded images, haze images constitute the main part due to the high frequency of this weather. Therefore, how to analyze and process haze images to restore clear and real fog-free images is very important for the application process of image information.

[0003] Currently, there are problems such as insufficient prediction accuracy and weak robustness in image defogging processing.

[0004] In summary, it is an urgent problem for those skilled in the art to provide a semi-supervised image defogging method that can achieve automatic image restoration, has higher image defogging efficiency, and more reliable restored image quality. Summary of the Invention

[0005] In view of the above problems and needs, the present solution proposes a semi-supervised image defogging method, which can solve the above technical problems due to the following technical solutions.

[0006] To achieve the above object, the present invention provides the following technical solution: A semi-supervised image defogging method, comprising:

[0007] Step Step1: Obtain an image data set, which includes a training set and a test set, and preprocess the training set to obtain synthetic foggy images and real fog maps;

[0008] Step Step2: Construct an image restoration model based on deep learning, input the foggy images extracted after preprocessing into the image restoration model, the image restoration model separates the features of the image content and image details in the foggy images, and fuses the two features to achieve the extraction and analysis of image features, and obtain defogged image information;

[0009] Step Step3: Use the image restoration model to construct a semi-supervised image restoration network, and train the semi-supervised image restoration network to obtain a trained image restoration model;

[0010] Step Step4: Obtain an original image set composed of foggy images to be restored, and input the original image set into the trained semi-supervised image restoration network to obtain the final restored image.

[0011] Further, a pre-trained model of a convolutional layer and an image restoration model is used to map an input image into high-dimensional deep learning features and content-feature separation weights; different-level convolution and separation operations are performed on the deep learning features; and the proportion of different components in the deep features is continuously adjusted according to the separation weights.

[0012] Furthermore, the image restoration model includes a multi-scale extraction module and an attention mechanism fusion module; the multi-scale extraction module is used to extract feature information at a fine scale and feature information at a coarse scale; the attention mechanism fusion module is used for feature fusion; the multi-scale extraction module includes sub-module a and sub-module b, both sub-module a and sub-module b have four branches, and include an average pooling and a concatenation layer, the concatenation layer concatenates the feature maps output by the corresponding four branches together, in the first layer of each branch, multiple 1×1 convolutions are used to change the dimension of the input picture, according to the network structure from bottom to top and from left to right, there are two 3×3 convolutional layers on the leftmost side of sub-module a, and for the extraction of coarse-scale features in sub-module b, two convolutional pairs of 7×1 and 1×7 are set on the leftmost branch, and a convolutional pair of 1×7 and 7×1 is set on the second left branch, the use of convolutional pairs reduces the number of parameters of the model, the size of the average pooling is 3×3, and all convolutional layers are followed by block regularization and the activation function ReLU.

[0013] Furthermore, the attention mechanism fusion module includes three feature fusion modules, each feature fusion module includes a fusion module and an attention module, all features first pass through the attention module composed of Conv+Sigmoid structure to extract attention feature maps before passing through the fusion module, each fusion module fuses the high-level feature f1 and the low-level feature f2 through a multiplication operation, so that the fused feature simultaneously captures the common properties of the high-level and low-level features, and each time the fused feature is forwarded to a convolutional layer, a batch normalization layer and a ReLU layer to obtain the final output feature f out , f out =O(F b (f1)*F R (f2), where O(.), F b and F R respectively represent the convolutional layer, the batch normalization layer and the ReLU layer, and then are processed by the next fusion module.

[0014] Further, the semi-supervised image restoration network includes a supervised network and an unsupervised network that shares weights with the supervised network. The supervised network obtains a dehazed image through an image restoration model, and trains the parameters of the supervised network through a supervised loss function and the hazy images synthesized in the training set. At the same time, the parameters of the unsupervised network with shared weights are optimized, and the unsupervised network is trained on real hazy images.

[0015] Furthermore, the supervised loss function includes a mean squared error loss function, a perceptual loss function, and an adversarial loss function; the unsupervised loss function includes a ranking loss function.

[0016] Furthermore, when the unsupervised network is being trained, first, the input hazy image is restored to a clear image J through the image restoration model, and then the dehazed image J is randomly cropped to obtain a local image. The classifier outputs a clear probability between [0,1], predicts the clarity of the dehazed image J output and the local image. Then, the loss is calculated using these two output probabilities. The ranking loss is expressed as

[0017] An image restoration device includes: an image preprocessing module, an input module, a restoration network construction module, and an output module;

[0018] The image preprocessing module is used to obtain an image data set, which includes a training set and a test set, and preprocess the training set to obtain synthesized hazy images and real fog images;

[0019] The input module is used to construct an image restoration model based on deep learning, input the hazy image extracted after preprocessing into the image restoration model. The image restoration model separates the image content and image details in the hazy image, and fuses the two features to realize the extraction and analysis of image features, and obtains the dehazed image information.

[0020] The restoration network construction module is used to construct a semi-supervised image restoration network using the image restoration model, and train the semi-supervised image restoration network to obtain a trained image restoration model;

[0021] The output module is used to obtain the original image set composed of the foggy images to be restored, and input the original image set into the trained image restoration model to obtain the final restored image.

[0022] From the above technical solutions, it can be seen that the beneficial effects of the present invention are: by establishing a semi-supervised restoration network composed of a supervised network and an unsupervised network, the present invention can realize automatic image restoration, with higher image dehazing efficiency and more reliable restored image quality.

[0023] In addition to the purposes, features and advantages described above, the preferred embodiments of implementing the present invention will be described in more detail below in conjunction with the accompanying drawings, so as to easily understand the features and advantages of the present invention. Description of the Drawings

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings required for describing the embodiments of the present invention or the prior art will be briefly introduced below. Among them, the accompanying drawings are only used to show some embodiments of the present invention, rather than limiting all embodiments of the present invention thereto.

[0025] Figure 1 It is a schematic diagram of the specific steps of a semi-supervised image dehazing method of the present invention.

[0026] Figure 2 It is a schematic diagram of the structure of the multi-scale extraction module in this embodiment.

[0027] Figure 3 It is a schematic diagram of the process of the unsupervised network branch in this embodiment.

[0028] Figure 4 It is a structural diagram of the semi-supervised image restoration network in this embodiment.

[0029] Figure 5 It is a schematic diagram of the composition structure of an image restoration device of the present invention.

[0030] Reference numerals: Image preprocessing module 1, input module 2, restoration network construction module 3, and output module 4. Detailed Description of the Invention

[0031] In order to make the purposes, technical solutions and advantages of the technical solutions of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the specific embodiments of the present invention. The same reference numerals in the drawings represent the same components. It should be noted that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0032] With the development and application of visual sensor technology, people often need to capture, process, and analyze a large number of image data with various angles and qualities, such as various surveillance images, Internet pictures, remote sensing satellite images, etc. However, there are often some images in the large amount of image data that are apparently degraded due to weather factors. Among these degraded images, due to the high frequency of haze weather, haze images constitute the main part of these degraded images. Therefore, how to analyze and process haze images to restore clear and real haze-free images has become a long-standing key problem. Deep learning algorithms can extract high-dimensional abstract features from input images and restore the clear appearance of the input images, so they have received extensive attention. However, the haze removal algorithms based on deep learning mainly use a fully supervised method to train the model, which requires a large number of hazy images and their corresponding clear images as training data. Since there are no corresponding clear images for real hazy images, they cannot be used as training data in the fully supervised method. Moreover, when foggy images are formed, the appearances of different objects are affected by haze to varying degrees, making the degradation distribution in real images have a certain local variability. This local variability also makes the domain attribute differences between real hazy images and synthetic images with a globally consistent degradation distribution.

[0033] To address the above problems, the embodiments of this application provide a semi-supervised based image haze removal method and an image restoration device, which can avoid the problem that using a large number of hazy images and their corresponding clear images as training data leads to poor generalization ability of the trained model for real hazy scenes, and the problem of not fully exploring the differences and complementarities of global and local features and multi-scale features of images. This method can be implemented in corresponding software, hardware, and a combination of software and hardware. The following provides a detailed introduction to the embodiments of this application.

[0034] As shown in the appendix Figure 1 shown, the appendix Figure 1 This application embodiment provides a semi-supervised based image haze removal method for automatically removing haze from hazy images. The method includes the following steps:

[0035] Step 1: Obtain an image dataset, where the image dataset contains a training set and a test set, and preprocess the training set to obtain synthetic hazy images and real foggy images.

[0036] The image dataset is the basis for method research and testing. A complete image dataset can reasonably and effectively reflect the performance of the method. Therefore, the image dataset is a very important part of the research task. In this embodiment, the image dataset includes two parts: (1) The RESIDE dataset, which contains synthetic and real-world blurred images. Its training set contains 13,990 synthetic blurred images; the test set consists of the Synthetic Objective Test Set (SOTS) and the Hybrid Subjective Test Set (HSTS). The RESIDE dataset introduces a new single-image dehazing benchmark, with a large-scale comprehensive training set, and two different datasets designed for objective and subjective quality assessment respectively. (2) The D-HAZY dataset, which contains ground truth reference images and blurred images of the same scene. It builds the dataset by synthesizing the haze in real images of complex scenes and contains multiple scenes.

[0037] Step Step2: Build a deep learning-based image restoration model, and input the hazy image extracted after preprocessing into the image restoration model. The image restoration model separates the image content and image details in the hazy image, and fuses the two features to achieve the extraction and analysis of image features, and obtains the dehazed image information.

[0038] In this embodiment, the image content and image details in the input foggy image are separated and feature-learned, and then the two different types of features are processed, and the two features are continuously fused during this process to utilize the complementary relationship between them, so as to achieve the effective extraction of image features.

[0039] According to the atmospheric scattering model, the foggy image I(z) = J(z)t(z) + A[1 - t(z)], where I and J represent the hazy image and the clear image respectively, and A and t represent the atmospheric light intensity and the transmittance. From the above formula, it can be found that during imaging, some appearances in the clear image will be affected by the atmospheric scattering rate, and the multiplication of 1 - t and A often causes the weakening effect of the local appearance of the clear image. That is, during imaging, the main area of the real appearance will be affected by atmospheric scattering and become dull, and its details may be lost due to the weakening effect. Then the restoration process of the clear picture can be obtained as where, That is, J = F1(I, t) + F2(A, T) = J1 + J2, where J1 and J2 respectively represent the deep features related to content and details. That is, the restoration process of the clear picture J consists of two parts: the first part is to restore the general content from I, and the second part is to restore the details of the image from A and T.

[0040] Specifically, in order to extract two types of images, a fixed convolutional layer and a pre-trained model of an image restoration model are used to map the input image into high-dimensional deep learning features and content-feature separation weights. Then, different levels of convolution and separation operations are performed on the deep learning features, which mainly include: i) using different convolutional layers to operate on the deep learning features to obtain features with different precisions; ii) extracting information about different components in the features according to the above-mentioned features with different precisions and separation weights; and then continuously adjusting the proportion of different components in the deep features according to the separation weights.

[0041] The image restoration model includes a multi-scale extraction module and an attention mechanism fusion module. The multi-scale extraction module is used to extract feature information at a fine scale and feature information at a coarse scale, that is, the above-mentioned different convolutional layers are used to operate on the deep learning features to obtain features with different precisions. The attention mechanism fusion module is used for feature fusion. As shown in the appendix Figure 2 shown, the appendix Figure 2 is a schematic structural diagram of the multi-scale extraction module. The multi-scale extraction module includes sub-module a and sub-module b. Both sub-module a and sub-module b have four branches, and include an average pooling layer and a concatenation layer. The concatenation layer concatenates the feature maps output by the corresponding four branches. In the first layer of each branch, multiple 1×1 convolutions are used to change the dimension of the input image. According to the network structure from bottom to top and from left to right, there are two 3×3 convolutional layers on the leftmost side of sub-module a, while for the extraction of coarse-scale features in sub-module b, two convolutional pairs of 7×1 and 1×7 are set on the leftmost branch, and a convolutional pair of 1×7 and 7×1 is set on the second left branch. The use of convolutional pairs reduces the number of parameters of the model. The size of the average pooling layer is 3×3. All convolutional layers are followed by block regularization and the activation function ReLU. The block regularization operation can prevent overfitting and accelerate network convergence. Finally, for sub-blocks a and b, the outputs of the four branches converge at the concatenation layer, so that the network can simultaneously extract feature information of the input image at different scales in the next stage.

[0042] The attention mechanism fusion module includes three feature fusion modules. Each feature fusion module includes a fusion module and an attention module. Before all features pass through the fusion module, they first pass through the attention module composed of the Conv+Sigmoid structure to extract the attention feature map. Each fusion module fuses the high-level feature f1 and the low-level feature f2 through a multiplication operation, so that the fused feature can simultaneously capture the common properties of the high-level and low-level features. Each time the fused feature is forwarded to the convolutional layer, the batch normalization layer, and the ReLU layer to obtain the final output feature f out , f out = O(F b (f1)*FR (f2), where O(.), F b and F R respectively represent the convolutional layer, the batch normalization layer, and the ReLU layer, and then are processed by the next fusion module.

[0043] The low-level features extracted by the deep neural network, that is, Figure 2 the features extracted from the upper two layers in the (a)(b) stages in [reference], usually represent local clues of the image, such as the edges and other patterns of the image. As the receptive field increases, the high-level features extracted by the network can capture the semantics within the global range of the image. If only low-level features are used, although they mainly contain detailed information, the semantics in a more global range cannot be well restored. On the contrary, if only high-level features are used, the details will be lost. Therefore, in the feature fusion method adopted in this embodiment, stage 1 corresponds to Figure 2 the first-layer features in the (a)(b) stages in [reference], stage 2 corresponds to Figure 2 the second-layer features in the (a)(b) stages in [reference], stage 3 corresponds to Figure 2 the third-layer features in the (a)(b) stages in [reference], stage 4 corresponds to Figure 2 the bottommost-layer features in the (a)(b) stages in [reference]. There are three feature fusion modules from top to bottom. The first module fuses the high-level features (stage 4) and the low-level features (stage 3), and takes the obtained features as high-level features, which are fused with the sub-high-level features of stage 2 through the second feature fusion module. Similarly, the obtained features are processed as high-level features and fused with the low-level features of stage 1 through the third feature fusion module. For each feature fusion module, given the high-level features and the low-level features, feature fusion is achieved by performing an element-wise multiplication operation on these two types of features. The fused features will be forwarded to the convolutional layer, the batch normalization layer (BatchNormalization), and the ReLU layer, and then are processed by the next fusion module.

[0044] Step Step3: Use the image restoration model to construct a semi-supervised image restoration network, and train the semi-supervised image restoration network to obtain a trained image restoration model.

[0045] As shown in the appendix Figure 4 shown, the appendix Figure 4 is the structural diagram of the semi-supervised image restoration network. The semi-supervised image restoration network includes a supervised network and an unsupervised network that shares weights with the supervised network. The supervised network obtains a defogged image through the image restoration model, and trains the parameters of the supervised network through the supervised loss function and the synthetic foggy images in the training set. At the same time, it optimizes the parameters of the unsupervised network with weight sharing. The unsupervised network is trained on real foggy images.

[0046] In this embodiment, the supervised loss function includes the mean squared error loss function, the perceptual loss function, and the adversarial loss function; the unsupervised loss function includes the ranking loss function. And the training process is divided into three stages: unsupervised pre-training, semi-supervised training, and fine-tuning. First, the network is pre-trained unsupervised using unlabeled data to learn the general visual representation of the images, and no specific task of the images is specified at this stage; then, semi-supervised learning is performed according to the results of the pre-training stage for a specific task (image dehazing) to ensure the restoration quality of the predicted images; finally, in order to further improve the performance of the network for image dehazing, the unlabeled data is directly used for the target task, and the network is fine-tuned using the unsupervised loss function. Adding a fine-tuning stage after the semi-supervised training to use the unlabeled data in a task-specific manner can enhance the consistency of the model's predictions for the unlabeled data between different training stages.

[0047] As shown in the Figure 3 appendix, Figure 3 FIG. is a schematic flow diagram of the unsupervised network branch. When the unsupervised network is trained, first, the input hazy image is restored to a clear image J by the generator, and then the dehazed image J is randomly cropped to obtain a local image. The classifier is used to output the clear probability between [0, 1], predict the clarity of the output of the dehazed image J and the local image. Then, the loss is calculated using these two output probabilities. The ranking loss is expressed as

[0048] Step 4: Obtain the original image set composed of the foggy images to be restored, and input the original image set into the trained semi-supervised image restoration network to obtain the final restored images.

[0049] As shown in the Figure 5 appendix, in the second embodiment, the present application further includes an image restoration device, which specifically includes: an image preprocessing module 1, an input module 2, a restoration network construction module 3, and an output module 4;

[0050] The image preprocessing module 1 is used to obtain an image data set, which includes a training set and a test set, and preprocess the training set to obtain synthetic foggy images and real fog maps;

[0051] The input module 2 is used to construct an image restoration model based on deep learning, input the foggy images extracted after preprocessing into the image restoration model, and the image restoration model separates the image content and image details in the foggy images, and fuses the two features to implement the extraction and analysis of image features to obtain the dehazed image information;

[0052] The restoration network construction module 3 is used to construct a semi-supervised image restoration network by using the image restoration model, and train the semi-supervised image restoration network to obtain a trained image restoration model;

[0053] The output module 4 is used to obtain the original image set composed of the foggy images to be restored, and input the original image set into the trained image restoration model to obtain the final restored images.

[0054] The method disclosed in this application can be embodied in the form of a software product. The computer software product is stored in a memory, and computer instructions that can run on the processor are stored on the memory. When the processor runs the computer instructions, the steps of the above-mentioned semi-supervised image dehazing method can be executed.

[0055] It should be noted that the embodiments described in the present invention are only the preferred ways to implement the present invention. Any obvious changes that belong to the overall concept of the present invention should fall within the protection scope of the present invention.

Claims

1. A semi-supervised image dehazing method, characterized in that, It includes the following steps: Step 1: Obtain an image dataset, which contains a training set and a test set, and preprocess the training set to obtain synthetic foggy images and real fog images; Step 2: Construct an image restoration model based on deep learning, and input the foggy images extracted after preprocessing into the image restoration model. The image restoration model separates and extracts the features of the image content and image details in the foggy images, and fuses the two features to achieve the extraction and analysis of image features, and obtains the de-fogged image information; Step 3: Use the image restoration model to construct a semi-supervised image restoration network, and train the semi-supervised image restoration network to obtain a trained image restoration model; Step 4: Obtain an original image set composed of foggy images to be restored, and input the original image set into the trained semi-supervised image restoration network to obtain the final restored image; The image restoration model includes a multi-scale extraction module and an attention mechanism fusion module; the multi-scale extraction module is used to extract feature information at a fine scale and feature information at a coarse scale; The attention mechanism fusion module is used for feature fusion; the multi-scale extraction module includes sub-module a and sub-module b. Both sub-module a and sub-module b have four branches, and contain an average pooling and a concatenation layer. The concatenation layer concatenates the feature maps output by the corresponding four branches. In the first layer of each branch, multiple 1×1 convolutions are used to change the dimension of the input image. According to the network structure from bottom to top and from left to right, there are two 3×3 convolutional layers on the leftmost side of sub-module a, and for the extraction of coarse-scale features in sub-module b, two convolutional pairs of 7×1 and 1×7 are set on the leftmost branch, and a convolutional pair of 1×7 and 7×1 is set on the second left branch. The use of convolutional pairs reduces the number of parameters of the model. The size of the average pooling is 3×3, and all convolutional layers are followed by block regularization and the activation function ReLU; The attention mechanism fusion module includes three feature fusion modules. Each feature fusion module includes a fusion module and an attention module. Before all features pass through the fusion module, they first pass through the attention module composed of the Conv+Sigmoid structure to extract the attention feature map. Each fusion module fuses the high-level feature f1 and the low-level feature f2 through a multiplication operation, so that the fused feature can capture the common properties of the high-level and low-level features at the same time. Each time the fused feature is forwarded to the convolutional layer, the batch normalization layer, and the ReLU layer to obtain the final output feature , , where 、 and represent the convolutional layer, the batch normalization layer, and the ReLU layer respectively, and then are processed by the next fusion module.

2. The semi-supervised image defogging method according to claim 1, wherein The feature separation of the image content and image details in the foggy image includes: using a convolutional layer and a pre-trained model of the image restoration model to map the input image into high-dimensional deep learning features and content-feature separation weights; performing convolutional and separation operations at different levels on the deep learning features; continuously adjusting the proportion of different components in the deep features according to the separation weights.

3. The semi-supervised image defogging method according to claim 1, characterized in that The semi-supervised image restoration network includes a supervised network and an unsupervised network that shares weights with the supervised network. The supervised network obtains de-fogged images through the image restoration model, and trains the parameters of the supervised network through the supervised loss function and the synthetic foggy images in the training set. At the same time, it optimizes the parameters of the unsupervised network with shared weights. The unsupervised network is trained on real foggy images.

4. The semi-supervised based image dehazing method according to claim 3, characterized in that, The supervised loss function includes a mean square error loss function, a perceptual loss function, and an adversarial loss function; the unsupervised loss function includes a ranking loss function.

5. The semi-supervised image defogging method according to claim 4, characterized in that When the unsupervised network is trained, first, the input hazy image is restored to a clear image J through the image restoration model, and then the dehazed image J is randomly cropped to obtain a local image , and the classifier outputs a clear probability between [0, 1] to predict the clarity of the dehazed image J output and the local image . Then, the loss is calculated using these two output probabilities, and the ranking loss is expressed as .

Citation Information

Patent Citations

  • Feature domain single image defogging method based on domain transformation

    CN112907478A

  • Image defogging method and system based on global feature fusion attention network

    CN113344806A