Single image defogging method and device based on style recovery and detail supplement, equipment and medium
By introducing style recovery and detail supplement modules into the image defog network, the problem of poor image defog removal effect in the prior art is solved, and more efficient image defog removal effect and performance improvement are achieved.
Patent Information
- Application Number
- CN202510226147.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The existing single-image defog removal method is poor in dealing with complex and changeable actual scenes, resulting in unsatisfactory image defog removal effect.
The image defog method based on style recovery and detail supplement is adopted. The image is processed through the pre-trained style recovery module, the detail supplement module to be trained and the cross-fusion module to be trained to restore the style and supplement details of the image respectively, reducing the difficulty of learning in the image defog network.
It effectively improves the effect and performance of image defog removal, significantly improves image quality, makes the boundaries of objects more distinct, and improves detection accuracy and classification accuracy.
Smart Images

Figure CN120147186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a single image defogging method, device, equipment and medium based on style restoration and detail supplementation. Background Art
[0002] Computer vision technology is rapidly expanding its influence in various industries and has become one of the key technologies driving the development of cutting-edge fields such as smart cities, autonomous driving, and intelligent safety systems. However, image quality largely determines the success rate of these applications. Various factors such as bad weather (smog, rain and snow), mechanical problems (camera shake), physical phenomena (thermal noise, shadows), and insufficient lighting or image data loss can significantly affect the clarity and integrity of the image, thereby limiting the effectiveness of subsequent analysis tasks. Among the many image damage problems, fog is the most common type of problem. Effectively restoring this type of damaged image can effectively improve the results of subsequent visual tasks.
[0003] As a key low-level visual task, single image dehazing plays a vital role in substantial applications such as object detection and autonomous driving. This technology aims to remove the effects of haze from a single input image through algorithms, restore the real details and clarity of the scene, and thus provide a more reliable foundation for high-level visual analysis. In terms of object detection, haze reduces image contrast and introduces noise, resulting in blurred object contours, which in turn affects the performance of the detection algorithm. Through effective dehazing processing, the image quality can be significantly improved, making the object boundaries more distinct, which not only improves the detection accuracy, but also enhances the classification accuracy. Clear images are helpful for feature extraction. For tasks that rely on texture or shape information for classification (such as pedestrian, vehicle and other target recognition), accurate image restoration can bring better classification results. Single image dehazing has important and widespread applications. For autonomous driving, the system relies on accurate environmental perception, and haze weather seriously affects the information obtained by the camera. Effective dehazing algorithms can help vehicles better understand the surrounding environment, avoid potential dangers, and ensure driving safety. In addition, image dehazing can enhance the data quality of visual sensors, making data fusion with other sensors such as LiDAR and millimeter-wave radar more effective, thereby improving the robustness and reliability of the entire system.
[0004] At present, single image defogging methods usually directly learn the overall difference between foggy images and clear images, which means that they directly learn from a very difficult task, resulting in a high learning difficulty, which in turn damages the image defogging effect to a certain extent, especially resulting in poor image defogging effect when processing complex and changeable actual scenes. Therefore, how to improve the image defogging effect is a technical problem to be solved by the present invention. Summary of the invention
[0005] Based on the above technical problems, the present invention provides a single-image defogging method, device, equipment and medium based on style restoration and detail supplementation, aiming to improve the effect and performance of image defogging.
[0006] In the first aspect of the present invention, a single-image defogging method based on style restoration and detail supplementation is provided. The method includes: Input a sample foggy image into the image defogging network to be trained. The image defogging network to be trained at least includes: a pre-trained style restoration module, a detail supplementation module to be trained, and a cross-fusion module to be trained. The detail supplementation module consists of multi-directional and multi-dimensional convolutions and is used to extract and supplement image details. The style restoration module is based on a large visual pre-trained model and is used to restore the style of the image. The cross-fusion module is used to repeatedly aggregate the output information of the detail supplementation module and the style restoration module; Process the sample foggy image through the pre-trained style restoration module to obtain a sample style-restored image; Process the sample foggy image through the detail supplementation module to be trained to obtain detail-supplemented image features; Input the sample style-restored image and the detail-supplemented image features into the cross-fusion module to be trained to obtain a sample defogged image; Based on the sample defogged image and the sample clean image corresponding to the sample foggy image, update the model parameters of the detail supplementation module to be trained and the cross-fusion module to be trained until a trained image defogging network is obtained; Input the image to be defogged into the trained image defogging network to obtain the defogged image output by the trained image defogging network.
[0007] In the second aspect of the present invention, a single-image defogging device based on style restoration and detail supplementation is provided. The device includes: An image input module, which is used to input a sample foggy image into the image defogging network to be trained. The image defogging network to be trained at least includes: a pre-trained style restoration module, a detail supplementation module to be trained, and a cross-fusion module to be trained. The detail supplementation module consists of multi-directional and multi-dimensional convolutions and is used to extract and supplement image details. The style restoration module is based on a large visual pre-trained model and is used to restore the style of the image. The cross-fusion module is used to repeatedly aggregate the output information of the detail supplementation module and the style restoration module; A style processing module, which is used to process the sample foggy image through the pre-trained style restoration module to obtain a sample style-restored image; A detail processing module, configured to process the sample hazy image through the to-be-trained detail supplementation module to obtain detail supplementation image features; A feature fusion module, configured to input the sample style restoration image and the detail supplementation image features into the to-be-trained cross-fusion module to obtain a sample haze-removed image; A model training module, configured to update the model parameters of the to-be-trained detail supplementation module and the to-be-trained cross-fusion module based on the sample haze-removed image and the sample clean image corresponding to the sample hazy image until a trained image haze-removal network is obtained; An image haze-removal module, configured to input a to-be-haze-removed image into the trained image haze-removal network to obtain a haze-removed image output by the trained image haze-removal network.
[0008] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the single-image haze-removal method based on style restoration and detail supplementation in the first aspect of the embodiments of the present invention.
[0009] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the single-image haze-removal method based on style restoration and detail supplementation in the first aspect of the embodiments of the present invention.
[0010] In the single-image haze-removal method based on style restoration and detail supplementation provided by the present invention, the sample hazy image is respectively subjected to style restoration and detail supplementation through a style restoration module and a detail supplementation module to obtain a sample style restoration image and detail supplementation image features, and then the sample style restoration image and the detail supplementation image features are combined through a cross-fusion module, and through repeated aggregation of high-level information, a clear haze-removed image is finally obtained. In this way, the present invention proposes an image haze-removal network with a style restoration module and a detail supplementation module as two sub-networks. In this image haze-removal network, the image haze-removal task is divided into two sub-tasks of "style restoration" and "detail supplementation" to respectively perform style restoration and detail supplementation on the hazy image, thereby reducing the learning difficulty of the image haze-removal network, effectively realizing image haze-removal, and improving the effect and performance of image haze-removal. Description of the Drawings
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It is a schematic diagram showing the style and content differences between a hazy image and a clear image pair in different scenarios shown in an embodiment of the present invention; Figure 2 It is a flowchart of the steps of a single-image dehazing method based on style restoration and detail supplementation shown in an embodiment of the present invention; Figure 3 It is a schematic diagram showing the implementation details of the detail supplementation module shown in an embodiment of the present invention; Figure 4 It is a schematic diagram showing the overall structure of an image dehazing network shown in an embodiment of the present invention; Figure 5 It is a comparison schematic diagram between three training methods shown in an embodiment of the present invention; Figure 6 It is a training comparison result diagram of SRM trained in three different ways on SOTS-indoor in the RESIDE dataset shown in an embodiment of the present invention; Figure 7 It is a flowchart of the training and inference of an image dehazing network shown in an embodiment of the present invention; Figure 8 It is a flowchart of the training and inference of an image dehazing network shown in another embodiment of the present invention; Figure 9 It is a visualization comparison diagram between a single-image dehazing method based on style restoration and detail supplementation and other mainstream methods shown in an embodiment of the present invention; Figure 10 It is another visualization comparison diagram between a single-image dehazing method based on style restoration and detail supplementation and other mainstream methods shown in an embodiment of the present invention; Figure 11 It is a structural block diagram of a single-image dehazing device based on style restoration and detail supplementation provided in an embodiment of the present invention. Detailed implementation manners
[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0014] In view of the technical defect that current deep learning-based image defogging methods usually directly learn the overall differences between hazy images and clear image pairs, which means learning directly from a very difficult task, thus resulting in the impairment of the image defogging effect, the present invention discovers that the degradation of hazy images usually involves two aspects: the transformation of style and the hiding of details. To illustrate this, the present invention evaluates the style and content differences between image pairs composed of hazy images and clear images under two different scenarios. The style difference can be represented by the Gram matrix, which is an index in the style transfer task. If the mean square distance of the Gram matrix between image pairs is smaller, it can be considered that the styles of the two images are more similar. The calculation method of the Gram matrix is as follows: ; where H, W, and C respectively represent the dimensions H×W×C of the image, represents the k-th element of the i-th channel, represents the element in the i-th row and j-th column of the Gram matrix. To calculate the content difference between two images, the present invention uses the function Contextual Loss ( ) to calculate. The specific calculation process is as follows, that is, the content and detail differences between can be represented by the contextual loss ( ): ; where CX(X,Y) is used to calculate the number of features in Y that are the nearest neighbors of a certain feature in X, and X and Y are the feature values obtained after passing through a certain feature extraction network Φ. The Φ of the present invention uses the Swin network. Through the above definitions of the style and content difference measurement methods, it is possible to calculate how much the image has specifically changed in terms of style and content after being affected by heavy fog.
[0015] As Figure 1 shown, Figure 1 is a schematic diagram of the style and content differences between hazy images and clear image pairs under different scenarios shown in an embodiment of the present invention. As can be seen from Figure 1 , when both images in the same image pair are clear or hazy, the Gram matrix is smaller, although the details shown in these images are completely different ( is relatively large). When a hazy image and a clear image in the same scenario are paired, the Gram matrix is larger, while is relatively small. This shows that: 1. The style difference between hazy images and clear images is very significant, and style restoration can play a key role in single-task image defogging; 2. There are also content differences between hazy images and clear images, and detail supplementation can be used as an important auxiliary means to improve the defogging performance.
[0016] Therefore, in order to at least partially solve one or more of the above problems and other potential problems, the present invention proposes an image defogging network with style restoration and detail supplementation as two sub-networks. In this image defogging network, the image defogging task is divided into two sub-tasks of "style restoration" and "detail supplementation" to perform style restoration and detail supplementation on the foggy image respectively, thereby reducing the learning difficulty of the image defogging network, effectively realizing image defogging, and improving the effect and performance of image defogging.
[0017] In the following, specific examples of this solution will be described in more detail with reference to the accompanying drawings.
[0018] Please refer to Figure 2 , Figure 2 which is a flowchart of the steps of a single-image defogging method based on style restoration and detail supplementation shown in an embodiment of the present invention. As Figure 2 shown, the single-image defogging method based on style restoration and detail supplementation provided in this embodiment at least includes the following steps: Step S11: Input the sample foggy image into the image defogging network to be trained, where the image defogging network to be trained at least includes: a pre-trained style restoration module, a detail supplementation module to be trained, and a cross-fusion module to be trained.
[0019] In this embodiment, the image defogging network to be trained at least includes: a pre-trained style restoration module, a detail supplementation module to be trained, and a cross-fusion module to be trained. Among them, the detail supplementation module is composed of multi-directional and multi-dimensional convolutions and is used to extract and supplement image details; the style restoration module is implemented based on a large visual pre-trained model and is used to restore the style of the image; the cross-fusion module is used to repeatedly aggregate the output information of the detail supplementation module and the style restoration module.
[0020] This embodiment prepares a sample data set for training the image defogging network, and this sample data set includes: sample foggy images and sample clean images corresponding to the sample foggy images; the sample clean images are the labels corresponding to the sample foggy images and are clear images after defogging corresponding to the sample foggy images; the sample foggy images are foggy image samples for training the image defogging network.
[0021] In this embodiment, during the training process of the image defogging network to be trained, first input the sample foggy image into the image defogging network to be trained.
[0022] Step S12: Process the sample foggy image through the pre-trained style restoration module to obtain a sample style restoration image.
[0023] In this embodiment, after the hazy sample image is input into the image dehazing network to be trained, the hazy sample image is respectively input into the pre-trained style restoration module and the detail supplement module to be trained, and the hazy sample image is processed in parallel by the pre-trained style restoration module and the detail supplement module to be trained. For example, when the hazy sample image enters the image dehazing network to be trained, it will be copied into two copies, one is input into the pre-trained style restoration module, and the other is input into the detail supplement module to be trained.
[0024] In this embodiment, after the hazy sample image is input into the pre-trained style restoration module, the pre-trained style restoration module can perform style restoration processing on the hazy sample image to obtain the sample style restoration image output by the style restoration module. The sample style restoration image is the image after style restoration of the hazy sample image. Among them, the pre-trained style restoration module in this embodiment is a pre-trained style restoration module. During the training process of the image dehazing network to be trained, the network parameters of the pre-trained style restoration module are fixed and will not be updated.
[0025] Step S13: Process the hazy sample image through the detail supplement module to be trained to obtain detail supplement image features.
[0026] In this embodiment, after the hazy sample image is input into the detail supplement module to be trained, the detail supplement module to be trained can process the hazy sample image to obtain the detail supplement image features output by the detail supplement module. The detail supplement image features are the image features after extracting and supplementing details in multiple directions and dimensions of the hazy sample image.
[0027] Step S14: Input the sample style restoration image and the detail supplement image features into the cross-fusion module to be trained to obtain a dehazed sample image.
[0028] In this embodiment, after obtaining the sample style restoration image and the detail supplement image features, the sample style restoration image and the detail supplement image features can be input into the cross-fusion module to be trained. The cross-fusion module fuses the sample style restoration image and the detail supplement image features to obtain the dehazed sample image output by the cross-fusion module. For example, the cross-fusion module to be trained can combine the sample style restoration image and the detail supplement image features and finally generate a clear image (i.e., the dehazed sample image) through repeated aggregation of high-level information. Among them, the dehazed sample image is the clean image obtained after the image dehazing network to be trained performs dehazing processing on the hazy sample image during the training process.
[0029] Step S15: Based on the sample haze-removed image and the sample clean image corresponding to the sample hazy image, update the model parameters of the to-be-trained detail supplementation module and the to-be-trained cross-fusion module until a trained image haze-removal network is obtained.
[0030] In this embodiment, based on the sample haze-removed image corresponding to the sample hazy image and the sample clean image corresponding to the sample hazy image, the model parameters of the to-be-trained detail supplementation module and the to-be-trained cross-fusion module can be updated until a trained image haze-removal network is obtained.
[0031] Step S16: Input the to-be-haze-removed image into the trained image haze-removal network to obtain the haze-removed image output by the trained image haze-removal network.
[0032] In this embodiment, the trained image haze-removal network is used to perform haze-removal processing on an image. During the application process of the trained image haze-removal network, the to-be-haze-removed image can be input into the trained image haze-removal model to obtain the haze-removed image output by the trained image haze-removal model. Among them, the to-be-haze-removed image is an image that needs to be haze-removed, and the haze-removed image is the image after haze-removal processing on the to-be-haze-removed image.
[0033] In this embodiment, the sample hazy image is respectively subjected to style restoration and detail supplementation through the style restoration module and the detail supplementation module to obtain the sample style-restored image and the detail-supplemented image features. Then, the sample style-restored image and the detail-supplemented image features are combined through the cross-fusion module, and through repeated aggregation of high-level information, a clear haze-removed image is finally obtained. In this way, the present invention proposes an image haze-removal network with the style restoration module and the detail supplementation module as two sub-networks. In this image haze-removal network, the image haze-removal task is divided into two sub-tasks of "style restoration" and "detail supplementation" to respectively perform style restoration and detail supplementation on the hazy image, thereby reducing the learning difficulty of the image haze-removal network, effectively realizing image haze-removal, and improving the effect and performance of image haze-removal.
[0034] Currently, in the field of single-image haze-removal tasks, simple convolutions are usually used for feature extraction, which limits the network feature expression and learning ability, and thus limits its ability to effectively express and model complex information. To solve this problem, in combination with the above embodiments, in one implementation, the present invention also provides a single-image haze-removal method based on style restoration and detail supplementation. In this method, the above step S13 can specifically include steps S21 to S23: Step S21: Input the sample hazy image into the first convolution, the second convolution, the third convolution, and the fourth convolution respectively for processing to obtain the first image feature, the second image feature, the third image feature, and the fourth image feature.
[0035] In this embodiment, the goal of the detail supplement module is to focus on extracting and supplementing details from all directions. To achieve this, the detail supplement module in this embodiment does not have any image generation function to ensure the effective extraction of eigenvalue in all directions and dimensions, facilitating the detail supplement in different directions and sizes.
[0036] In this embodiment, the detail supplement module uses multiple convolutions with different sizes and directions. The detail supplement module at least includes: the first convolution, the second convolution, the third convolution, and the fourth convolution. The first convolution, the second convolution, the third convolution, and the fourth convolution are respectively responsible for detail supplement in four dimensions: point level, horizontal level, vertical level, and plane level. Among them, the first convolution can be Pixel-Wise Convolution (P-WC), the second convolution can be Horizontal-Wise Convolution (H-WC), the third convolution can be Vertical-Wise Convolution (V-WC), and the fourth convolution can be Plane-Wise Convolution (PL-WC).
[0037] In the detail supplement module, the sample foggy image is respectively input into the first convolution, the second convolution, the third convolution, and the fourth convolution for processing, and the first image feature output by the first convolution, the second image feature output by the second convolution, the third image feature output by the third convolution, and the fourth image feature output by the fourth convolution are respectively obtained.
[0038] Step S22: Connect the sample foggy image, the first image feature, the second image feature, the third image feature, and the fourth image feature to obtain a fifth image feature.
[0039] In this embodiment, after obtaining the first image feature, the second image feature, the third image feature, and the fourth image feature, the sample foggy image, the first image feature, the second image feature, the third image feature, and the fourth image feature are connected to obtain a comprehensive eigenvalue, that is, the fifth image feature is obtained.
[0040] Step S23: Adjust the spatial resolution of the fifth image feature to obtain the detail supplement image feature.
[0041] In this embodiment, after obtaining the fifth image feature, it is necessary to adjust the spatial resolution of the fifth image feature to obtain the detailed supplementary image feature output by the detailed supplementary module. In a specific example, this embodiment adjusts the spatial resolution of the fifth image feature through the PixelUnshuffle operation. Among them, PixelUnshuffle and PixelShuffle are reverse operations. PixelUnshuffle is an effective downsampling method, which is crucial for converting a low-dimensional, large-size tensor into a high-dimensional, small-size tensor, while PixelShuffle performs the opposite conversion.
[0042] In addition, in a specific example, the implementation details of various convolutions in the detailed supplementary module are shown in Table 1, where padding is used to keep the output and output size of the convolution the same.
[0043] Table 1 Detailed parameter table of each convolution in the detailed supplementary module
[0044] In one embodiment, as Figure 3 shown, Figure 3 is a schematic diagram of the implementation details of the detailed supplementary module shown in an embodiment of the present invention. In Figure 3 , the detailed supplementary module (Detail Replenishment Module, DRM) uses a variety of convolutions with different sizes and directions. These convolutions include four types: P-WC, H-WC, V-WC, and PL-WC, and these convolutions are responsible for detailed supplementation in the four dimensions of point level, horizontal level, vertical level, and plane level respectively. In DRM, the features extracted by P-WC, H-WC, V-WC, and PL-WC from the input image (such as the sample hazy image) are concatenated with the input image to form a comprehensive eigenvalue, and the spatial resolution is adjusted through the PixelUnshuffle operation to obtain the detailed supplementary image feature corresponding to the input image.
[0045] Combining the above embodiments, in one implementation manner, the present invention also provides a single-image defogging method based on style restoration and detail supplementation. In this method, step S14 above may specifically include steps S31 to S34: Step S31: In the cross-fusion module to be trained, perform an upsampling operation and a first spatial resolution adjustment operation on the sample style restoration image to obtain a sixth image feature.
[0046] In this embodiment, in the cross-fusion module to be trained, the sample style-restored image output by the style restoration module can be successively subjected to an upsampling operation and a first spatial resolution adjustment operation to obtain a sixth image feature. The first spatial resolution adjustment operation in this embodiment can be a PixelUnshuffle operation.
[0047] Step S32: Connect the sixth image feature and the detail-complemented image feature to obtain a seventh image feature.
[0048] In this embodiment, after obtaining the sixth image feature, the sixth image feature and the detail-complemented image feature output by the detail complement module can be connected to obtain a seventh image feature.
[0049] Step S33: Process the seventh image feature through multiple dimensionality reduction modules to obtain an intermediate feature map.
[0050] In this embodiment, the cross-fusion module includes multiple dimensionality reduction modules, and each dimensionality reduction module includes: a dimensionality reduction convolution (DRC) and an activation function (such as GeLU). After obtaining the seventh image feature, the seventh image feature can be successively input into multiple dimensionality reduction modules for processing to perform successive dimensionality reduction to obtain an intermediate feature map. Among them, when the feature is input into each dimensionality reduction module, it is first processed by the dimensionality reduction convolution and then processed by the activation function.
[0051] Step S34: Perform a second spatial resolution adjustment operation on the intermediate feature map to obtain the sample defogged image.
[0052] In this embodiment, a second spatial resolution adjustment operation is performed on the obtained intermediate feature map to obtain the sample defogged image output by the cross-fusion module. Among them, the second spatial resolution adjustment operation can be a PixelShuffle operation.
[0053] Combined with the above embodiments, in one implementation, the present invention also provides a single-image defogging method based on style restoration and detail complement. In this embodiment, the style restoration module at least includes: an improved MAE encoder and an MAE decoder, and step S12 above can specifically include step S41: Step S41: Process the sample foggy image successively through the improved MAE encoder and the MAE decoder to obtain the sample style-restored image.
[0054] In this embodiment, the main task of the style restoration module is to restore the style of the image. Since style features belong to a type of high-level semantic features, using a deeper and larger network is beneficial for achieving better style feature restoration. Based on this, this embodiment uses a large vision model to extract semantic features to achieve style restoration, which can alleviate common problems such as overfitting and poor generalization performance in single-task image restoration networks. Therefore, the style restoration module in this embodiment is implemented based on the large vision pre-trained model MAE. MAE (Masked Auto Encoder) is a self-supervised pre-trained large vision model, and MAE includes: an MAE decoder and an MAE encoder. This embodiment uses MAE as the benchmark network for the style restoration module, and the style restoration module includes: an MAE decoder and an improved MAE encoder.
[0055] In this embodiment, after the sample hazy image is input into the style restoration module, the sample hazy image is processed successively through the improved MAE encoder and the MAE decoder to obtain the sample style-restored image output by the style restoration module. In a specific implementation, in order to keep the sizes of the input and output images consistent with the original MAE, this embodiment needs to reduce the resolution of the input image (such as the sample hazy image) before feeding it into the style restoration module to adapt to the MAE network.
[0056] To enable MAE to have the ability to perform image dehazing in style restoration, this embodiment realizes dehazing MAE by using an adapter before and after each MAE encoder layer in the MAE encoder. Specifically, the improved MAE encoder in this embodiment includes: a plurality of dehazing MAE encoding layer structures connected in sequence. The input sample hazy image is processed successively through a plurality of dehazing MAE encoding layer structures and then processed by the MAE decoder to finally obtain the sample style-restored image. Among them, each dehazing MAE encoding layer structure includes: a first adapter, an MAE encoder layer, and a second adapter connected in sequence.
[0057] In this embodiment, in order to enable MAE to obtain the ability of style restoration, according to the data transmission direction in the MAE encoder, a first adapter is added before each MAE encoder layer, and a second adapter is added after each MAE encoder layer. Among them, the first adapter can be a Channel-Wise Adapter (CW-A), and the second adapter can be a Patch-Wise Adapter (PW-A). The first adapter includes a linear layer and an activation layer, which are responsible for feature transformation across multiple patches; the second adapter involves a single linear transformation that modifies the features of each eigenvalue before being processed by the encoder layer.
[0058] In this embodiment, through the first adapter and the second adapter, the original self-attention mechanism can effectively identify information of different dimensions, enabling the original MAE to acquire new capabilities, namely, the ability to remove fog. The eigenvalue output by each dehazing MAE encoding layer structure in this embodiment can be expressed as: batch_size (the number of data in each batch of training) * patch_num (the number of patches each image is divided into) * patch_size (the size of each patch); specifically, the output of each dehazing MAE encoding layer structure can be obtained through the following formula: ; Among them, represents the eigenvalue fed into the t-th dehazing MAE encoding layer structure, represents the t-th layer encoder of MAE, respectively represent the linear layer weights of the Adapters before and after the t-th layer encoder of MAE, that is, the weights of the first adapter and the second adapter corresponding to the t-th layer encoder of MAE respectively; represents the bias of the linear layer of the Adapter before and after the t-th layer encoder, and GeLU is the activation function. It should be noted that this embodiment uses the original decoder of the pre-trained model MAE, rather than only using the high-level information generated by the encoder for downstream tasks. Therefore, when designing the Adapter in this embodiment, a simple linear transformation is selected to retain the original solution space of the decoder, and the Adapter weights are initialized with the identity matrix to ensure its normal operation. That is to say, in the above formula will be initialized as the identity matrix.
[0059] In this embodiment, during the training process of the style restoration module, in order to retain more pre-trained knowledge, the network parameters of both the MAE encoder layer and the MAE decoder are frozen to prevent any gradient calculation or weight update. At the same time, only the network parameters of the adapters (i.e., the first adapter and the second adapter) are updated, that is, only the adapters perform gradient calculation and weight update.
[0060] Currently, the datasets used for dehazing mainly consist of synthetic data. However, synthetic datasets may not fully cover all possible real degradation situations, resulting in unstable performance of the image dehazing network when encountering unseen situations. In addition, these datasets are often limited in scale, which restricts the training volume and may lead to overfitting, that is, the model overfits to the training data and cannot generalize well to new data. Therefore, the generalization ability of the image dehazing network trained with these datasets is easily limited. Based on this, in view of the problem of poor generalization performance of the current mainstream image dehazing network, a dehazing module (i.e., style restoration module) is designed based on the self-masked autoencoder MAE to leverage the advantages of PTMs (Pre-training models) in the image dehazing task, enabling the network to acquire rich prior knowledge, thereby improving the generalization ability of the image dehazing network and further ensuring the practical application of image dehazing in various scenarios.
[0061] In one embodiment, as Figure 4 shown, Figure 4 is the overall structural schematic diagram of the image dehazing network shown in an embodiment of the present invention. In Figure 4In , the image defogging network is mainly composed of three modules: Style Restoration Module (SRM), Detail Replenishment Module (DRM) and Cross Fusion Module (CFM). When a foggy image enters the image defogging network, it will be copied into two copies. One is sent to DRM to extract and supplement details and obtain detail supplement image features. The other is sent to SRM to achieve style restoration and obtain a style restored image. After the foggy image passes through SRM and DRM, all the output features will be input into CFM to combine the output information of SRM and DRM, and finally generate a clear and clean image through repeated aggregation of high-level information. Among them, SRM includes MAE decoder and improved MAE encoder. The improved MAE encoder includes N defogging MAE coding layers (defogging MAE coding layer structure), and each defogging MAE coding layer includes: a CW-A, a MAE encoder layer and a PW-A. CFM combines the outputs of DRM and SRM to fuse style and detail information. In CFM, the dimensionality reduction operation is achieved by adjusting the number of input channels and output channels. The goal of this is to enable the network to autonomously learn how to utilize style information and detail information. To achieve this, this embodiment increases the feature dimensions of DRM and SRM to a very high level when designing the network, and then connects them together to obtain a 672-dimensional feature map. Next, a dimensionality reduction convolution is used to continuously reduce the dimensionality of the feature map to obtain a 48-dimensional feature map. Finally, a PixelShuffle operation is performed, and this process will produce a clear image. In this way, the single image dehazing method proposed in the embodiment of the present invention has better generalization performance and stronger dehazing and restoration capabilities.
[0062] In combination with the above embodiments, in one implementation, the present invention further provides a single image defogging method based on style restoration and detail supplementation. In this embodiment, the pre-trained style restoration module is obtained by training based on the style restoration module to be trained, and the style restoration module to be trained includes: a pre-trained MAE encoder layer, a pre-trained MAE decoder, a first adapter to be trained, and a second adapter to be trained; and the training steps of the style restoration module to be trained may specifically include steps S51 to S57: Step S51: inputting a sample clean image corresponding to the sample foggy image into the style restoration module to be trained to obtain a first sample image.
[0063] In this embodiment, during the training of the image dehazing network, in order to enable MAE to effectively learn the dehazing ability through style restoration, a two-stage progressive multi-task training method is adopted to train the style restoration module and the detail supplementation module according to different sub-tasks respectively, so that they respectively have the abilities of style restoration and detail supplementation, which helps to effectively apply the pre-trained large model to the image dehazing task.
[0064] In the first-stage training, first complete the training of the style restoration module, and during the process of training the style restoration module, freeze all weights except the first adapter to be trained and the second adapter to be trained.
[0065] When training the style restoration module, first input the sample clean image corresponding to the sample hazy image into the style restoration module to be trained, and obtain the first sample image output by the style restoration module to be trained. Among them, the sample data set used to train the style restoration module is the same as the sample data set used to train the image dehazing network, that is, the sample hazy image used to train the style restoration module is the same as the sample hazy image used to train the image dehazing network.
[0066] Step S52: Based on the sample clean image and the first sample image, obtain the first loss.
[0067] In this embodiment, after obtaining the first sample image, the first loss can be calculated based on the sample clean image corresponding to the sample hazy image and the first sample image corresponding to the sample hazy image.
[0068] Step S53: Input the sample hazy image into the style restoration module to be trained, and obtain the second sample image.
[0069] In this embodiment, the sample hazy image can also be input into the style restoration module to be trained, and the second sample image output by the style restoration module to be trained is obtained.
[0070] Step S54: Based on the second sample image and the first sample image, obtain the second loss.
[0071] In this embodiment, after obtaining the second sample image, the second loss can be calculated based on the first sample image corresponding to the sample hazy image and the second sample image corresponding to the sample hazy image.
[0072] Step S55: Obtain the image reconstruction loss based on the first loss and the second loss.
[0073] In this embodiment, after obtaining the first loss and the second loss, the image reconstruction loss can be obtained based on the first loss and the second loss.
[0074] Step S56: Obtain a style loss based on the second sample image and the sample clean image.
[0075] In this embodiment, a style loss can also be calculated based on the obtained second sample image and sample clean image.
[0076] In a specific example, the style loss can be represented by the following formula:
[0077] where l represents an intermediate layer of a certain feature extraction network (such as the style restoration module in this embodiment). In one embodiment, 5 layers of Relu-1, Relu-6, Relu-11, Relu-20, and Relu-26 in VGG19 are adopted. represents the Gram matrix of the sample clean image corresponding to the l-th layer, represents the Gram matrix of the second sample image output by the l-th layer.
[0078] Step S57: Update the network parameters of the to-be-trained first adapter and the to-be-trained second adapter based on at least the image reconstruction loss and the style loss until the pre-trained style restoration module is obtained.
[0079] In this embodiment, after obtaining the image reconstruction loss and the style loss, the network parameters of the to-be-trained first adapter and the to-be-trained second adapter can be updated based on at least the image reconstruction loss and the style loss until the trained first adapter and the trained second adapter are obtained, so as to obtain the pre-trained style restoration module based on the trained first adapter, the trained second adapter, the pre-trained MAE encoder layer, and the pre-trained MAE decoder.
[0080] In an example, in order to train the style restoration module, this embodiment proposes a progressive multi-task training method. In this stage (i.e., the first-stage training), in order to obtain the image reconstruction loss, two propagations are performed. In the first propagation process, the sample clean image Clear is input into the to-be-trained style restoration module, and the task of the style restoration module is to restore the sample clean image to obtain the first sample image . The mean squared error (MSE) is used to evaluate the gap between the sample clean image Clear and the first sample image to obtain the first loss . In the second propagation process, the sample hazy image Hazy is fed into the to-be-trained style restoration module, and the second sample image is obtained after the sample hazy image is processed by the style restoration module. In this embodiment, the second sample image Instead of comparing with the sample clean image Clear, the second sample image is subjected to MSE evaluation with the first sample image , denoted as the second loss . The image reconstruction loss is composed of the first loss + the second loss . By backpropagating the image reconstruction loss to update the style restoration module, this process is the "progressive multi-task" training method. Through this progressive multi-task loss of this embodiment, the ability of the style restoration network in the image restoration task is improved, thereby enhancing the defogging ability.
[0081] To show the advantages of the progressive multi-task loss in the training process, in one embodiment, this embodiment also compares the training method using the progressive multi-task loss with two other training methods that respectively use the "single-task direct loss" and the "multi-task direct loss". Among them, the "single-task direct loss" can be described as the loss between the second sample image and the sample clean image, that is ; the "multi-task direct loss" can be described as . The specific relationship among the three can be referred to Figure 5 as shown Figure 5 , which is a comparison schematic diagram among the three training methods shown in one embodiment of the present invention.
[0082] Combining the above embodiments, in one implementation manner, the present invention also provides a single-image defogging method based on style restoration and detail supplementation. In this method, the above step S57 can specifically include step S61 and step S62: Step S61: Calculate the total loss based on the image reconstruction loss and the image reconstruction loss weight, and the style loss and the style loss weight.
[0083] In this embodiment, the image reconstruction loss corresponds to an image reconstruction loss weight, and the style loss corresponds to a style loss weight. The image reconstruction loss weight and the style loss weight can be freely set. The total loss can be calculated based on the image reconstruction loss and the image reconstruction loss weight, and the style loss and the style loss weight.
[0084] In a specific example, the total loss is: ; where is the image reconstruction loss, is the style loss, is the image reconstruction loss weight, is the style loss weight.
[0085] Step S62: Update the network parameters of the first adapter to be trained and the second adapter to be trained based on the total loss until the pre-trained style restoration module is obtained.
[0086] In this embodiment, after obtaining the total loss, the network parameters of the first adapter to be trained and the second adapter to be trained can be updated based on the total loss until the trained first adapter and the trained second adapter are obtained, thereby obtaining the pre-trained style restoration module.
[0087] Among them, in the initial stage of training the style restoration module, that is, during the first convergence of the total loss, the image reconstruction loss weight is set to 1 and the style loss weight is set to 0, so that the style restoration module has the image reconstruction ability; after the first convergence of the total loss, the image reconstruction loss weight is set to 1 and the style loss weight is set to 1000, and the accuracy of the model will increase again and converge again to enhance the style restoration ability of the style restoration module.
[0088] In addition, to illustrate the rationality and effectiveness of the training strategy proposed in this embodiment, key metrics during the training process were also measured. As Figure 6 shown, Figure 6 is a training comparison result graph of the SRM trained in three different ways on SOTS-indoor in the RESIDE dataset shown in an embodiment of the present invention. In Figure 6 , by comparing the images generated by SRM with the Clear images, the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) are calculated respectively. Compared with the other two methods ("direct single-task learning" and "direct multi-task learning"), the "progressive multi-task" training method has obvious advantages in terms of training speed and accuracy. After further training for 26 epochs, the total loss converges, and can be set to 1, and can be set to 1000, and the model further converges. The present invention designs the dehazing MAE for style restoration by adjusting the pre-trained model MAE, and at the same time adopts a two-stage "progressive multi-task" method to train the network, so that MAE can effectively learn the dehazing ability through style restoration.
[0089] Combining the above embodiments, in one implementation manner, the present invention also provides a single-image dehazing method based on style restoration and detail supplementation. In this method, the above step S15 may specifically include steps S71 to S73: Step S71: Based on the sample dehazed image and the corresponding sample clean image of the sample hazy image, calculate the first loss function to obtain the third loss, and update the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module based on the third loss.
[0090] In this embodiment, in the two-stage training method, after the to-be-trained style restoration module is trained in the first stage to obtain the pre-trained style restoration module, the pre-trained style restoration module is completely frozen, and its gradients and parameters are no longer updated, and then the second stage of training is carried out. In the second stage of training, the parameters of the pre-trained style restoration module are fixed, and the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module are updated.
[0091] In this embodiment, based on the sample dehazed image and the corresponding sample clean image of the sample hazy image, the first loss function can be calculated to obtain the third loss. Then, based on the third loss, the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module are updated. In a specific example, the first loss function is L2Loss, that is is: ; where y represents the sample clean image, represents the sample dehazed image, and i represents the i-th pixel value of the image.
[0092] Step S72: After the third loss converges, based on the sample dehazed image and the corresponding sample clean image of the sample hazy image, calculate the second loss function to obtain the fourth loss, and continue to update the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module based on the fourth loss until the fourth loss converges, and obtain the trained detail supplement module and the trained cross-fusion module.
[0093] In this embodiment, after the third loss converges, based on the sample dehazed image and the corresponding sample clean image of the sample hazy image, calculate the second loss function to obtain the fourth loss. Among them, the second loss function is a loss function incorporating SSIM. In a specific example, the second loss function is: ; where y represents the sample clean image, represents the sample dehazed image, and respectively represent the average pixel values of the sample clean image and the sample dehazed image, and represent the variances of the sample dehazed image and the sample clean image respectively, represents the covariance between the sample clean image and the sample dehazed image. and is to prevent the denominator from becoming 0. In this embodiment, take and .
[0094] After obtaining the fourth loss, the model parameters of the detail supplement module to be trained and the cross-fusion module to be trained can be continuously updated based on the fourth loss until the fourth loss converges, and the trained detail supplement module and the trained cross-fusion module are obtained.
[0095] Step S73: Based on the pre-trained style restoration module, the trained detail supplement module, and the trained cross-fusion module, obtain the trained image dehazing network.
[0096] In this embodiment, after obtaining the trained detail supplement module and the trained cross-fusion module, the trained image dehazing network can be obtained based on the pre-trained style restoration module, the trained detail supplement module, and the trained cross-fusion module.
[0097] In a specific example, in the second-stage training, the batch_size is set to 4, and the learning rate is initialized to . Continue to use cosine decay to update the learning rate, and keep the minimum learning rate .
[0098] Image dehazing plays a crucial role in security and surveillance systems, and can significantly improve the monitoring effect of these systems. By removing image blurring and color distortion caused by haze, the dehazing algorithm can restore clearer and more realistic scene details, thereby improving the quality and reliability of video surveillance. In one embodiment, an image restoration preprocessing module (such as an image dehazing preprocessing module) can be deployed in the system through the single-image dehazing method based on style restoration and detail supplement shown in any of the foregoing embodiments. Visual systems in these fields can obtain more accurate data support and thus make more informed decisions. The deployment of the system is divided into a training stage and an inference stage, as Figure 7 shown, Figure 7 is the training and inference flowchart of the image dehazing network shown in an embodiment of the present invention. The training process is as follows: Building the dataset: In the experimental part, in this embodiment, the performance of the model was evaluated on three different datasets: RESIDE, Haze4K, and NH-Haze. The RESIDE dataset is a widely used dataset for real and synthetic image dehazing. It includes five subsets: Indoor Training Set (ITS), Outdoor Training Set (OTS), Synthetic Object Test Set (SOTS), Real Task-driven Test Set (RTTS), and Hybrid Subjective Test Set (HSTS). In the experiment, this embodiment mainly used the ITS and OTS subsets for training and evaluated the model on the SOTS subset. The Haze4K dataset includes 3000 synthetic training images and 1000 synthetic test images, and this embodiment also evaluated the key metrics of the model on it. In addition, training and testing were also carried out on the real-scene dataset NH-Haze.
[0099] Reading image data: Read the corresponding image data and its ground truth data from the training dataset. The batch size of loading training image data for each GPU is 32.
[0100] Configuring training parameters: Use PyTorch 2.0 to build the model and train it on a system with 4 GPUs configured with NVIDIA 3090 or above. The specific training parameters are shown in Table 2. The specific training strategy refers to the single-image dehazing method based on style restoration and detail supplementation shown in any of the above embodiments.
[0101] The model outputs the prediction result. The restored result obtained by model inference is an RGB image identical to the input image. Calculating the loss and backpropagating: During the training process, calculate the loss between the input image and the output image. The design of the specific loss function refers to the single-image dehazing method based on style restoration and detail supplementation shown in any of the above embodiments.
[0102] Saving the training result: After training is completed, the final model is obtained and saved offline, such as in HDF5, TensorFlow SavedModel, ONNX, etc., and the model is exported through the corresponding API or tool. At the same time, the model can be verified on the validation set.
[0103] The inference process during the implementation can be divided into the following steps: Loading the model: Load the offline model using methods such as ONNX and TensorRT. Using tools such as ONNX and TensorRT can significantly improve the efficiency of model loading and inference and support cross-platform flexibility.
[0104] Data Reading: Read and decode images from the camera in real time. During the data reading process, the program will initialize the camera connection, capture image frames in real time and perform decoding processing, converting them into digital information that can be analyzed for subsequent image processing and analysis algorithms. Initialize the camera through OpenCV or a similar computer vision library. Then, use the camera's API to set parameters such as frame rate and resolution, and start capturing the image stream in real time. After each frame is captured, use the built-in decoder to convert the original image data (such as H.264 encoded) into matrix data in RGB or grayscale format for subsequent processing. This process is usually continuously executed in a loop to ensure the continuity and real-time nature of the image data.
[0105] Send the read image into the model for inference The output of the model is the dehazed image. Further send the dehazed image into the subsequent visual tasks to be completed in the next step.
[0106] Table 2 System Hardware Configuration and Software Version Information Table
[0107] In addition, in another embodiment, the single-image dehazing method based on style restoration and detail supplementation proposed in the above embodiment is often deployed and implemented as part of image processing software. In the past two years, with the rise of computational photography, removing haze has become one of the basic AI capabilities. Whether it is professional graphic design software or lightweight applications for ordinary users, they can provide users with a powerful tool to improve the quality of photos and other visual content. Especially in post-processing of photography, the dehazing function will play an irreplaceable role. The system deployment is divided into a training stage and an inference stage, as Figure 8 shown, Figure 8 is the training and inference flowchart of the image dehazing network shown in another embodiment of the present invention.
[0108] 1. Training Stage: Since the process and parameter settings in the training stage are the same as those in the previous embodiment, they will not be elaborated here. Briefly, it includes steps such as data preparation, model selection and initialization, parameter configuration, training process, verification and tuning, and saving the best model.
[0109] 2. Deployment and Inference Stage: Load the model: Load the offline model using methods such as ONNX and TensorRT. Using tools such as ONNX and TensorRT can significantly improve the efficiency of model loading and inference, and support cross-platform flexibility. When this inference module is deployed as part of image processing software, there are some specific considerations. The trained model and its inference logic can be seamlessly integrated into the existing image processing software framework to ensure good cooperation with other functional modules.
[0110] Data Loading: Real-time data: Read image frames from the camera, and after necessary preprocessing (such as scaling, normalization), send them into the model. Custom data: Users upload image or video files, and the system automatically parses and converts them into a format acceptable to the model, ensuring the same preprocessing standards as the real-time data.
[0111] Saving Results: Ensure that the inference results are presented to users in an intuitive way, display the detection results or classification labels through a graphical interface, and enhance the overall user experience.
[0112] Through the above deployment implementation steps, the inference module can be effectively integrated into the image processing software, not only maintaining consistency with the training stage, but also fully considering the requirements of different data sources and actual application scenarios. This deployment method enhances the functionality and flexibility of the software, making it more adaptable to diverse user needs.
[0113] In one embodiment, in order to verify the effectiveness of the single-image dehazing method based on style restoration and detail supplementation proposed in the embodiments of the present invention, in Figure 9 and Figure 10 , different methods are used to present the visualization results of the image dehazing algorithms for indoor and outdoor scenes. The algorithms compared with the method proposed in the embodiments of the present invention are: mainstream dehazing algorithms such as ADONet, GridDehazeNet, FFANet, C2PNet, DeHamer, etc. It can be seen from the visualization comparison of the indoor scene that there are obvious color differences in the images generated by AODNet. In contrast, C2PNet performs well in both indoor and outdoor scenes. However, the method (Ours) even exceeds C2PNet in terms of naturalness and faithfulness. On the more realistic dataset NH-Haze, the method in this paper still has obvious advantages. In the thick fog area, SRDR has more realistic colors and richer texture details, being closer to the real image, indicating that the method proposed in the present invention has achieved state-of-the-art performance on many mainstream datasets. Among them, Figure 9 is a visualization comparison graph of the single-image dehazing method based on style restoration and detail supplementation shown in one embodiment of the present invention and other mainstream methods; Figure 10 is another visualization comparison graph of the single-image dehazing method based on style restoration and detail supplementation shown in one embodiment of the present invention and other mainstream methods.
[0114] In one embodiment, Table 3 shows the performance metrics of various dehazing networks across multiple datasets, where the best results are highlighted in bold. Compared with other methods, the method proposed in the present invention (Ours) has an obvious PSNR advantage on all datasets. Specifically, on the test set of RESIDE, the SRDR network achieved PSNR values of 43.46 dB and 40.63 dB in indoor and outdoor scenes respectively, which are 0.90 dB and 3.95 dB higher than those of C2PNet respectively. On another synthetic dataset, Haze 4K, the method proposed in the present invention achieved a PSNR of 35.11 dB and an SSIM of 0.991, which is also better than other existing image dehazing methods. In addition, the performance on NH-Haze, a real-world dataset, was also evaluated. SRDR still achieved the best value in terms of the PSNR metric, but the SSIM metric was slightly lower than that of C2PNet.
[0115] Table 3 Quantitative comparison table of the present invention and other mainstream methods
[0116] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0117] Based on the same inventive concept, an embodiment of the present invention provides a single-image dehazing device based on style restoration and detail supplementation. Refer to Figure 11 , Figure 11 which is the structural block diagram of a single-image dehazing device based on style restoration and detail supplementation provided by an embodiment of the present invention. As Figure 11 shown, the single-image dehazing device based on style restoration and detail supplementation in this embodiment may include: An image input module, configured to input a sample hazy image into an image dehazing network to be trained, where the image dehazing network to be trained at least includes: a pre-trained style restoration module, a detail supplementation module to be trained, and a cross-fusion module to be trained; the detail supplementation module is composed of multi-directional and multi-dimensional convolutions, and is used to extract and supplement image details; the style restoration module is based on a large vision pre-trained model and is used to restore the style of the image; the cross-fusion module is used to repeatedly aggregate the output information of the detail supplementation module and the style restoration module; A style processing module for processing the sample hazy image through the pre-trained style restoration module to obtain a sample style-restored image; A detail processing module for processing the sample hazy image through the to-be-trained detail supplement module to obtain detail supplement image features; A feature fusion module for inputting the sample style-restored image and the detail supplement image features into the to-be-trained cross-fusion module to obtain a sample haze-removed image; A model training module for updating the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module based on the sample haze-removed image and the sample clean image corresponding to the sample hazy image until a trained image haze-removal network is obtained; An image haze-removal module for inputting the to-be-haze-removed image into the trained image haze-removal network to obtain the haze-removed image output by the trained image haze-removal network.
[0118] Optionally, the detail supplement module at least includes: a first convolution, a second convolution, a third convolution, and a fourth convolution, and the first convolution, the second convolution, the third convolution, and the fourth convolution are respectively responsible for detail supplement in four dimensions of point level, horizontal level, vertical level, and plane level; The detail supplement module includes: A convolution processing module for respectively inputting the sample hazy image into the first convolution, the second convolution, the third convolution, and the fourth convolution for processing to obtain a first image feature, a second image feature, a third image feature, and a fourth image feature; A feature connection module for connecting the sample hazy image, the first image feature, the second image feature, the third image feature, and the fourth image feature to obtain a fifth image feature; A resolution adjustment module for performing spatial resolution adjustment on the fifth image feature to obtain the detail supplement image features.
[0119] Optionally, the feature fusion module includes: A first processing module for performing an upsampling operation and a first spatial resolution adjustment operation on the sample style-restored image in the to-be-trained cross-fusion module to obtain a sixth image feature, A second processing module for connecting the sixth image feature and the detail supplement image features to obtain a seventh image feature; A third processing module for processing the seventh image feature through a plurality of dimensionality reduction modules to obtain an intermediate feature map; each of the plurality of dimensionality reduction modules includes: a dimensionality reduction convolution and an activation function; A fourth processing module for performing a second spatial resolution adjustment operation on the intermediate feature map to obtain the sample defogged image.
[0120] Optionally, the style restoration module at least includes: an improved MAE encoder and an MAE decoder; The improved MAE encoder includes: a plurality of defogging MAE encoding layer structures connected in sequence, and each defogging MAE encoding layer structure includes: a first adapter, an MAE encoder layer, and a second adapter connected in sequence; The first adapter includes a linear layer and an activation layer, and is responsible for feature transformation across multiple blocks; the second adapter involves a single linear transformation; A style restoration module, including: An encoding and decoding processing module for processing the sample foggy image through the improved MAE encoder and the MAE decoder in sequence to obtain the sample style restoration image.
[0121] Optionally, the pre-trained style restoration module is trained based on the style restoration module to be trained. The style restoration module to be trained includes: a pre-trained MAE encoder layer, a pre-trained MAE decoder, a first adapter to be trained, and a second adapter to be trained; the apparatus further includes: a style restoration training module for training the style restoration module to be trained. The style restoration training module includes: A first input module for inputting the sample clean image corresponding to the sample foggy image into the style restoration module to be trained to obtain a first sample image; A first calculation module for obtaining a first loss based on the sample clean image and the first sample image; A second input module for inputting the sample foggy image into the style restoration module to be trained to obtain a second sample image; A second calculation module for obtaining a second loss based on the second sample image and the first sample image; A third calculation module for obtaining an image reconstruction loss based on the first loss and the second loss; A fourth calculation module for obtaining a style loss based on the second sample image and the sample clean image; A first parameter update module for updating the network parameters of the first adapter to be trained and the second adapter to be trained at least based on the image reconstruction loss and the style loss until the pre-trained style restoration module is obtained.
[0122] Optionally, the first parameter update module includes: The total loss calculation module is used to calculate the total loss based on the image reconstruction loss and the image reconstruction loss weight, as well as the style loss and the style loss weight; The second parameter update module is used to update the network parameters of the to-be-trained first adapter and the to-be-trained second adapter based on the total loss until the pre-trained style restoration module is obtained; Wherein, during the first convergence of the total loss, the image reconstruction loss weight is set to 1 and the style loss weight is set to 0, so that the style restoration module has the image reconstruction ability; After the first convergence of the total loss, the image reconstruction loss weight is set to 1 and the style loss weight is set to 1000 to enhance the style restoration ability of the style restoration module.
[0123] Optionally, the model training module includes: The first update module is used to calculate a third loss by calculating a first loss function based on the sample haze-removed image and the sample clean image corresponding to the sample hazy image, and update the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module based on the third loss; The second update module is used to calculate a fourth loss by calculating a second loss function based on the sample haze-removed image and the sample clean image corresponding to the sample hazy image after the third loss converges, and continue to update the model parameters of the to-be-trained detail supplement module and the to-be-trained cross-fusion module based on the fourth loss until the fourth loss converges, so as to obtain a trained detail supplement module and a trained cross-fusion module; The network training module is used to obtain the trained image haze removal network based on the pre-trained style restoration module, the trained detail supplement module, and the trained cross-fusion module.
[0124] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the single-image haze removal method based on style restoration and detail supplement as described in any one of the above embodiments of the present invention are implemented.
[0125] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, the steps in the single-image haze removal method based on style restoration and detail supplement as described in any one of the above embodiments of the present invention are implemented.
[0126] For the apparatus embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For related parts, please refer to the corresponding descriptions in the method embodiments.
[0127] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0128] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, apparatuses, or computer program products. Therefore, the embodiments of the present invention can take the form of all-hardware embodiments, all-software embodiments, or embodiments combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0129] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0130] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0131] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0132] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0133] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the element.
[0134] The above has introduced in detail a single-image defogging method, device, equipment and medium based on style restoration and detail supplementation provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A single image dehazing method based on style restoration and detail supplementation, characterized in that: The method comprises: Inputting a sample foggy image into an image defogging network to be trained, the image defogging network to be trained at least comprises: a pre-trained style restoration module, a detail supplementation module to be trained and a cross-fusion module to be trained; the detail supplementation module is composed of multi-directional and multi-dimensional convolutions, and is used to extract and supplement image details; the style restoration module is implemented based on a large visual pre-training model, and is used to restore the style of the image; the cross-fusion module is used to repeatedly aggregate the output information of the detail supplementation module and the style restoration module; Processing the sample foggy image by means of the pre-trained style restoration module to obtain a sample style restored image; Processing the sample foggy image by the detail supplementation module to be trained to obtain detail supplementation image features; Inputting the sample style restoration image and the detail supplement image features into the cross fusion module to be trained to obtain a sample defogging image; Based on the sample defogging image and the sample clean image corresponding to the sample foggy image, the model parameters of the detail supplement module to be trained and the cross fusion module to be trained are updated until a trained image defogging network is obtained; The image to be defogged is input into the trained image defogging network to obtain a defogged image output by the trained image defogging network.
2. The single image defogging method based on style restoration and detail supplementation according to claim 1, characterized in that: The detail supplement module at least includes: a first convolution, a second convolution, a third convolution and a fourth convolution, wherein the first convolution, the second convolution, the third convolution and the fourth convolution are respectively responsible for detail supplementation in four dimensions: point level, horizontal level, vertical level and plane level; The sample foggy image is processed by the detail supplementation module to be trained to obtain detail supplementation image features, including: Inputting the sample foggy image into the first convolution, the second convolution, the third convolution and the fourth convolution respectively for processing to obtain a first image feature, a second image feature, a third image feature and a fourth image feature; Connecting the sample foggy image, the first image feature, the second image feature, the third image feature and the fourth image feature to obtain a fifth image feature; The spatial resolution of the fifth image feature is adjusted to obtain the detail supplement image feature.
3. The single image defogging method based on style restoration and detail supplementation according to claim 1, characterized in that: Inputting the sample style restoration image and the detail supplement image features into the cross fusion module to be trained to obtain a sample defogging image, including: In the cross-fusion module to be trained, an upsampling operation and a first spatial resolution adjustment operation are performed on the sample style restoration image to obtain a sixth image feature, Connecting the sixth image feature and the detail supplementary image feature to obtain a seventh image feature; The seventh image feature is processed by multiple dimension reduction modules to obtain an intermediate feature map; each of the multiple dimension reduction modules includes: a dimension reduction convolution and an activation function; A second spatial resolution adjustment operation is performed on the intermediate feature map to obtain the sample defogging image.
4. The single image defogging method based on style restoration and detail supplementation according to claim 1, characterized in that: The style restoration module at least includes: an improved MAE encoder and a MAE decoder; The improved MAE encoder comprises: a plurality of defogging MAE encoding layer structures connected in sequence, each defogging MAE encoding layer structure comprising: a first adapter, a MAE encoder layer and a second adapter connected in sequence; The first adapter includes a linear layer and an activation layer, which is responsible for feature transformation across multiple blocks; the second adapter involves a single linear transformation; Processing the sample foggy image by the pre-trained style restoration module to obtain a sample style restoration image includes: The sample foggy image is processed by the improved MAE encoder and the MAE decoder in sequence to obtain the sample style restored image.
5. The single image defogging method based on style restoration and detail supplementation according to claim 4, characterized in that: The pre-trained style restoration module is obtained by training the style restoration module to be trained, and the style restoration module to be trained includes: a pre-trained MAE encoder layer, a pre-trained MAE decoder, a first adapter to be trained, and a second adapter to be trained; the training steps of the style restoration module to be trained include at least: Inputting a sample clean image corresponding to the sample foggy image into the style restoration module to be trained to obtain a first sample image; Based on the sample clean image and the first sample image, obtaining a first loss; Inputting the sample foggy image into the style restoration module to be trained to obtain a second sample image; Obtaining a second loss based on the second sample image and the first sample image; Obtaining an image reconstruction loss based on the first loss and the second loss; Obtaining a style loss based on the second sample image and the sample clean image; Based at least on the image reconstruction loss and the style loss, network parameters of the first adapter to be trained and the second adapter to be trained are updated until the pre-trained style restoration module is obtained.
6. The single image defogging method based on style restoration and detail supplementation according to claim 5, characterized in that: At least based on the image reconstruction loss and the style loss, updating the network parameters of the first adapter to be trained and the second adapter to be trained until the pre-trained style restoration module is obtained, comprising: Calculating a total loss based on the image reconstruction loss and the image reconstruction loss weight, and the style loss and the style loss weight; Based on the total loss, network parameters of the first adapter to be trained and the second adapter to be trained are updated until the pre-trained style restoration module is obtained; Wherein, during the first convergence of the total loss, the image reconstruction loss weight is set to 1, and the style loss weight is set to 0, so that the style restoration module has image reconstruction capability; After the total loss converges for the first time, the image reconstruction loss weight is set to 1 and the style loss weight is set to 1000 to enhance the style restoration capability of the style restoration module.
7. The single image defogging method based on style restoration and detail supplementation according to any one of claims 1 to 6, characterized in that: Based on the sample defogging image and the sample clean image corresponding to the sample foggy image, the model parameters of the detail supplement module to be trained and the cross fusion module to be trained are updated until a trained image defogging network is obtained, including: Based on the sample defogging image and the sample clean image corresponding to the sample foggy image, a first loss function is calculated to obtain a third loss, and based on the third loss, model parameters of the detail supplement module to be trained and the cross fusion module to be trained are updated; After the third loss converges, a second loss function is calculated based on the sample defogging image and the sample clean image corresponding to the sample foggy image to obtain a fourth loss, and model parameters of the detail supplement module to be trained and the cross fusion module to be trained are continuously updated based on the fourth loss until the fourth loss converges to obtain a trained detail supplement module and a trained cross fusion module; Based on the pre-trained style restoration module, the trained detail supplement module and the trained cross-fusion module, the trained image defogging network is obtained.
8. A single image defogging device based on style restoration and detail supplementation, characterized in that: The device comprises: An image input module is used to input a sample foggy image into an image defogging network to be trained, wherein the image defogging network to be trained comprises at least: a pre-trained style restoration module, a detail supplementation module to be trained, and a cross-fusion module to be trained; the detail supplementation module is composed of multi-directional and multi-dimensional convolutions, and is used to extract and supplement image details; the style restoration module is implemented based on a large visual pre-training model, and is used to restore the style of the image; the cross-fusion module is used to repeatedly aggregate the output information of the detail supplementation module and the style restoration module; A style processing module, used to process the sample foggy image through the pre-trained style restoration module to obtain a sample style restored image; A detail processing module, used for processing the sample foggy image through the detail supplementation module to be trained to obtain detail supplementation image features; A feature fusion module, used for inputting the features of the sample style restoration image and the detail supplement image into the cross fusion module to be trained to obtain a sample defogging image; A model training module, used for updating the model parameters of the detail supplement module to be trained and the cross fusion module to be trained based on the sample defogging image and the sample clean image corresponding to the sample foggy image, until a trained image defogging network is obtained; The image defogging module is used to input the image to be defogged into the trained image defogging network to obtain the defogged image output by the trained image defogging network.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the single image dehazing method based on style restoration and detail supplementation is implemented as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the single image defogging method based on style restoration and detail supplementation as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Single image defogging method based on detail recovery
CN116152107A
End-to-end image defogging method and device and medium
CN118037585A
Single-image defogging method based on detail restoration
WO2024178979A1