Haze image defogging model training method and device

Through the combined training method of self-supervised learning strategy and feedback optimization network, the problem of insufficient adaptability and robustness of deep learning dehazing models in hazy images is solved, and efficient dehazing effects are achieved in complex scenes.

CN120672598APending Publication Date: 2025-09-19CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510806370.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing deep learning dehazing models have poor adaptability and robustness due to their static architecture, making it difficult to cope with the uncertainty and complex scenes of hazy images. They are unable to achieve ideal dehazing effects in scenes with uneven haze distribution or rich image content.

Method used

A self-supervised learning strategy of dynamically updating the training sample set and sample label set is adopted to train the template haze image dehazing model including the feedback optimization network. The adaptive adjustment of the haze image is achieved through the combination of the feature processing network, the dehazed image prediction network and the feedback optimization network.

Benefits of technology

The robustness and adaptability of the haze image dehazing model in practical applications are enhanced, and it can better handle complex haze scenes and image content and generate high-quality dehazed images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672598A_ABST
    Figure CN120672598A_ABST
Patent Text Reader

Abstract

The invention discloses a haze image defogging model training method and device. The method comprises the following steps: acquiring a training sample set and a sample label set; a template haze image defogging model is obtained through training of the training sample set and the sample label set, and the template haze image defogging model at least comprises a feature processing network, a defogging image prediction network and a feedback optimization network; inputting a to-be-processed target haze image into the template haze image de-model to obtain a corresponding target defogged image; and respectively adding the target haze image and the target defogging image into the training sample set and the sample label set, and carrying out self-supervised training on the template haze image defogging model by using the new training sample set and the new sample label set to obtain a target haze image defogging model. According to the method and the device, the technical problem that the adaptability of the model is relatively poor as the haze image defogging model cannot be dynamically adjusted according to the content information of the haze image and the uncertainty of haze distribution in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and device for training a haze image defogging model. Background Art

[0002] In the field of computer vision, image dehazing technology is crucial for improving the quality and accuracy of image recognition in hazy weather conditions, especially in applications such as autonomous driving, drone inspections, and security monitoring. However, traditional image dehazing methods, such as atmospheric scattering models and dark channel priors, are often based on fixed assumptions and parameters, making them difficult to adapt to complex and changing hazy scenes. This results in inconsistent image quality after processing, and is particularly ineffective in scenes with uneven haze distribution or rich image content.

[0003] In recent years, with the rise of deep learning technology, image dehazing algorithms based on deep neural networks have gradually become mainstream. These algorithms automatically adjust dehazing parameters by learning image features and haze patterns, theoretically enabling better processing of various haze images. However, existing deep learning dehazing models mostly use static architectures, meaning the network structure and feature extraction methods are fixed during model design. These fixed models struggle to cope with the uncertainties of haze images, such as haze density and distribution, and the diversity of objects within the image. The static nature of these models limits their adaptability and robustness in complex scenarios. In particular, they often fail to achieve ideal dehazing results when faced with unknown haze patterns or complex image structures, hindering subsequent image recognition and analysis tasks.

[0004] Therefore, designing a haze image dehazing model that can adaptively expand to cope with the diversity of image content and the uncertainty of haze distribution is a key issue that needs to be solved urgently. Summary of the Invention

[0005] The embodiments of the present application provide a method and device for training a haze image dehazing model, so as to at least solve the technical problem that the related technology cannot dynamically adjust the haze image dehazing model according to the content information of the haze image and the uncertainty of the haze distribution, resulting in poor adaptability of the model.

[0006] According to one aspect of an embodiment of the present application, a method for training a haze image dehazing model is provided, comprising: obtaining a training sample set and a sample label set, wherein the training sample set includes a plurality of haze images as training samples, and the sample label set includes a dehazed image corresponding to each haze image as a sample label corresponding to the corresponding training sample; using the training sample set and the sample label set to train a template haze image dehazing model, wherein the template haze image dehazing model includes at least: a feature processing network, a dehazed image prediction network, and a feedback optimization network; inputting a target haze image to be processed into the template haze image dehazing model to obtain a corresponding target dehazed image; adding the target haze image and the target dehazed image to the training sample set and the sample label set, respectively, and performing self-supervised training on the template haze image dehazing model using the new training sample set and the new sample label set to obtain a target haze image dehazing model.

[0007] Optionally, the feature processing network includes at least: a first feature extraction module, a second feature extraction module and a feature enhancement module, wherein the first feature extraction module includes at least: multiple moving inverted bottleneck convolution MBConv modules connected in series, and the first feature extraction module is used to extract features from the haze image to obtain a first multi-scale fusion feature map; the second feature extraction module includes at least: three adapters, and the second feature extraction module is used to extract features from the haze image to obtain a second multi-scale fusion feature map; the feature enhancement module is used to add the first multi-scale fusion feature map output by the first feature extraction module and the second multi-scale fusion feature map output by the second feature extraction module to obtain an enhanced feature map.

[0008] Optionally, the first feature extraction module also includes a feature fusion module, wherein the MBConv module includes at least: an expansion layer, a depth-separable convolution layer, a compression and excitation unit, and a point-by-point convolution layer, and each MBConv module is used to capture the features of the haze image at different scales to obtain the corresponding first feature map; the feature fusion module includes multiple convolution layers, and the number of convolution layers is equal to the number of first feature maps. The first feature fusion module is used to perform feature fusion on the first image features after each convolution layer performs dimensional conversion on each first feature map to obtain a first multi-scale fusion feature map.

[0009] Optionally, the second feature extraction module also includes: an adapter selection unit, wherein each adapter is used to perform multiple feature extractions on the haze image to obtain corresponding multiple second feature maps; the adapter selection unit is used to generate a first selection weight corresponding to each adapter based on each second feature map, and add the multiple second feature maps output by the adapter with the largest first selection weight to obtain a second multi-scale fusion feature map.

[0010] Optionally, the feature extraction stage of the first adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual convolutional layer, a second feature extraction stage consisting of a downsampling block and at least one residual convolutional layer, and a third feature extraction stage consisting of an upsampling block and at least one residual convolutional layer; the feature extraction stage of the second adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual dense block, a second feature extraction stage consisting of a downsampling block and at least one residual dense block, and a third feature extraction stage consisting of an upsampling block and at least one residual dense block; the feature extraction stage of the third adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual block based on center difference convolution, a second feature extraction stage consisting of a downsampling block and at least one residual block based on center difference convolution, and a third feature extraction stage consisting of an upsampling block and at least one residual block based on center difference convolution.

[0011] Optionally, the dehazed image prediction network includes at least: multiple third feature extraction modules and a feature mapping layer, and the third feature extraction module is composed of multiple cascaded residual blocks, and the feature mapping layer includes a convolution layer.

[0012] Optionally, the feedback optimization network includes at least: an evaluation module, three optimization modules and an optimization selection module, wherein the evaluation module is used to evaluate the difference between the predicted defogging image output by the defogging image prediction network and the preset annotated clear image, and use the difference as the quality evaluation result of the predicted defogging image; the first optimization module includes a plurality of cascaded self-attention modules, and the first feature optimization unit is used to correct the erroneous information in the enhanced feature map when the difference is lower than a preset threshold value, so as to obtain an optimized enhanced feature map; the second optimization module includes a plurality of cascaded residual dense blocks, and the second feature optimization unit is used to correct the erroneous information in the enhanced feature map when the difference is lower than a preset threshold value, so as to obtain an optimized enhanced feature map. When the limit is reached, the local errors in the enhanced feature map are corrected to obtain the optimized enhanced feature map; the third optimization module includes multiple cascaded residual blocks based on central difference convolution, and the third feature optimization unit is used to enhance the texture details in the enhanced feature map when the difference is lower than the preset threshold value to obtain the optimized enhanced feature map; the optimization selection module includes at least: multiple convolution layers, global evaluation pooling layers, and fully connected layers, and the optimization selection module is used to generate the second selection weights of each optimization module according to the optimized enhanced feature maps output by each optimization module, and output the optimized enhanced feature map output by the feature optimization unit with the largest second selection weight.

[0013] According to another aspect of an embodiment of the present application, a haze image defogging model training device is also provided, including: an acquisition module for acquiring a training sample set and a sample label set, wherein the training sample set includes multiple haze images as training samples, and the sample label set includes a defogged image corresponding to each haze image as a sample label corresponding to the corresponding training sample; a template generation module for using the training sample set and the sample label set to train a template haze image defogging model, wherein the template haze image defogging model includes at least: a feature processing network, a defogged image prediction network, and a feedback optimization network; an inference module for inputting the target haze image to be processed into the template haze image defogging model to obtain the corresponding target defogged image; a training module for adding the target haze image and the target defogged image to the training sample set and the sample label set respectively, and using the new training sample set and the new sample label set to perform self-supervised training on the template haze image defogging model to obtain the target haze image defogging model.

[0014] According to another aspect of an embodiment of the present application, a computer program product is further provided, comprising: a computer program, wherein when the computer program is executed by a processor, the above-mentioned haze image defogging model training method is implemented.

[0015] According to another aspect of an embodiment of the present application, an electronic device is further provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned haze image defogging model training method through the computer program.

[0016] In an embodiment of the present application, a self-supervised learning strategy of dynamically updating the training sample set and the sample label set is adopted to train a template haze image dehazing model including a feedback optimization network, thereby achieving the purpose of enhancing the robustness and adaptability of the haze image dehazing model in practical applications, and solving the technical problem that the related technology cannot dynamically adjust the haze image dehazing model according to the content information of the haze image and the uncertainty of the haze distribution, resulting in poor adaptability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 This is a flow chart of an optional haze image defogging model training method according to an embodiment of the present application;

[0019] Figure 2 is a structural diagram of an optional template haze image defogging model according to an embodiment of the present application;

[0020] Figure 3 is a structural diagram of an optional haze image defogging model training device according to an embodiment of the present application;

[0021] Figure 4 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0023] It should be noted that the terms "first", "second", etc. in the specification, claims, and drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.

[0024] In order to better understand the embodiments of the present application, some nouns or terms that appear in the description of the embodiments of the present application are first translated and explained as follows:

[0025] The Residual Dense Block (RDB) is a deep learning architecture primarily used for image super-resolution tasks. It extracts rich local features through densely connected convolutional layers and allows the state of the previous RDB to be directly connected to all layers of the current RDB, forming a continuous memory mechanism. This design helps more effectively integrate local features, stabilize the network training process, and improve the quality of image reconstruction.

[0026] Central Difference Convolution (CDC): is a new type of convolution operation designed to enhance the representation of fine-grained features in convolutional neural networks. CDC combines intensity and gradient information to capture more detailed patterns, thereby improving the robustness and generalization of the model.

[0027] Bilinear interpolation method: It is a commonly used interpolation method in image processing and computer vision. It is used to estimate the value of unknown data points between known data points. Its core idea is to perform linear interpolation in two directions to obtain the value of the target point.

[0028] EffcientNet model: is a convolutional neural network (CNN) designed through neural architecture search (NAS) technology, which aims to optimize the depth, width and input resolution of the model to improve performance and efficiency. Among them, the EffcientNet series includes multiple variants from B0 to B7, each variant exhibits different performance and computing requirements under different parameter settings. Specifically, the network structure of EffcientNet consists of multiple stages (Stage), each stage contains a different number of MBConv modules, among which the MBConv module mainly consists of the following parts: a 1*1 expansion layer, a k*k Depthwise convolution layer, a SE (Squeeze and Exciatation) module, a 1*1 convolution layer (dimensionality reduction, including BN), and a Dropout layer.

[0029] Example 1

[0030] According to an embodiment of the present application, a method for training a haze image dehazing model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0031] Figure 1 FIG. 1 is a flow chart of a haze image defogging model training method according to an embodiment of the present application. Figure 1 As shown, the method includes the following steps:

[0032] Step S102: Obtain a training sample set and a sample label set, wherein the training sample set includes a plurality of haze images as training samples, and the sample label set includes a dehazed image corresponding to each haze image as a sample label corresponding to the corresponding training sample.

[0033] Step S104: Using the training sample set and the sample label set, a template haze image defogging model is trained. The template haze image defogging model includes at least: a feature processing network, a defogging image prediction network, and a feedback optimization network.

[0034] Step S106 : inputting the target haze image to be processed into the template haze image de-model to obtain a corresponding target de-hazed image.

[0035] Step S108 , adding the target haze image and the target dehazed image to the training sample set and the sample label set respectively, and performing self-supervised training on the template haze image dehazing model using the new training sample set and the new sample label set to obtain the target haze image dehazing model.

[0036] Based on the scheme defined in the above steps S102 to S108, it can be known that in an embodiment of the present application, a self-supervised learning strategy of dynamically updating the training sample set and the sample label set is adopted to train the template haze image dehazing model including the feedback optimization network, thereby enhancing the robustness and adaptability of the haze image dehazing model in practical applications.

[0037] The following describes the steps of the haze image dehazing model training method in combination with the specific implementation process.

[0038] In the technical solution provided in step S102 above, the haze images used as training samples in the training sample set can be derived from various real-world scenes, such as city streets, natural scenery, and industrial environments. The dehazed images in the sample label set are clear images obtained by dehazing these haze images used as training samples manually or through a pre-designed image generation algorithm.

[0039] In addition to the training sample acquisition schemes listed above, based on the basic concept of the present invention, those skilled in the art can also obtain training sample sets and sample label sets through other technical solutions. For example, those skilled in the art may make changes to the above implementation schemes, which should also be within the scope of protection of the present invention.

[0040] Furthermore, a template haze image dehazing model is trained using the training sample set and the sample label set, so that the model can learn the transformation rules from haze images to clear images during the training process.

[0041] In the technical solution provided in step S104 above, the specific structure of the template haze image defogging model is as follows: Figure 2 As shown below. Figure 2 The model structure shown introduces the model dehazing process.

[0042] The feature processing network in the template haze image dehazing model is used to extract, fuse, and enhance features of the haze image, mining and integrating all information that is helpful for dehazing. Therefore, the feature processing network includes at least: a first feature extraction module, a second feature extraction module, and a feature enhancement module, where:

[0043] (1) The first feature extraction module.

[0044] Optionally, the first feature extraction module is used to extract features from the haze image to obtain a first multi-scale fusion feature map, and the first feature extraction module at least includes: multiple MBConv (Mobile Inverted Bottleneck Convolution) modules connected in series and a feature fusion module.

[0045] Specifically, the multiple MBConv modules can be the first 7 layers of EfficientNet, and each MBConv module can capture the features of the input haze image at different scales to obtain the corresponding first feature map.

[0046] Among them, the structure of the MBConv module mainly includes the following parts:

[0047] Expansion Layer: This layer uses 1×1 convolution to expand the number of channels in the input haze image, forming an expanded feature map. This layer increases the network's expressive power, allowing subsequent depthwise separable convolutions to process richer information.

[0048] Depthwise Separable Convolution: Use 3×3 depthwise separable convolution to perform spatial processing on the expanded feature map.

[0049] Squeeze-and-Excitation (SE) unit: It consists of two parts, Squeeze and Excitation. The Squeeze part compresses the feature map output by the depthwise separable convolutional layer into a channel vector through global average pooling; while the Excitation part processes the channel vector through two fully connected layers to generate a weight vector, which is then multiplied channel by channel by the weight vector and the feature map output by the depthwise separable convolutional layer. This achieves dynamic adjustment of the importance of different channel information and enhances the information interaction between channels of the network.

[0050] Pointwise Convolution Layer: A 1×1 pointwise convolution is used after the SE unit to adjust the number of channels to restore the number of channels of the feature map to the size before the expansion layer, or to adjust it according to different stages of the network.

[0051] Therefore, each MBConv module in the first feature extraction module performs feature extraction and information processing through the network layer of the above-mentioned fixed architecture, and can obtain multi-level information commonly present in haze images, such as basic features such as image texture, color, and edges, and provide efficient and rich intermediate feature representations for subsequent deeper feature learning. It should be noted that the parameters of each MBConv module in the first feature extraction module (such as expansion ratio, size and number of depth-separable convolutions, setting of activation functions, etc.) can be set according to the actual application scenario. The embodiments of this application do not impose specific restrictions on this.

[0052] Furthermore, since the first feature maps output by each MBConv module carry information of the input haze image at different scales, in order to ensure that the feature maps at each scale can be effectively processed and fused, the embodiment of the present application also designs a feature fusion module in the first feature extraction module.

[0053] Optionally, the feature fusion module includes multiple convolution layers (such as the convolution kernel size can be 1*1, and the output channels are 32), and the number of convolution layers is equal to the number of first feature maps, so as to ensure that each first feature map can be processed by a dedicated convolution layer. Among them, the role of the convolution layer is to perform channel dimension conversion, that is, to convert the first feature map into a unified output dimension. On the one hand, this can reduce the computational complexity, and on the other hand, it can also maintain or enhance the key features of the image. Then, each convolution layer performs feature fusion on the first image features after the dimension conversion of each first feature map. Among them, feature fusion can be performed by adding or splicing to aggregate the first image features of different scales and after dimension adjustment to obtain a first multi-scale fusion feature map.

[0054] In addition, since the resolutions of the first feature maps output by each MBConv module are different, the first feature maps of the lower layer have a higher resolution and can therefore carry more detail information; while the first feature maps of the higher layer have a lower resolution and therefore contain more abstract and global information. Therefore, directly fusing multiple first feature maps output by multiple MBConv modules may cause information misalignment and affect the fusion effect. To this end, an embodiment of the present application proposes that before performing channel conversion, the first feature maps of different scales can be expanded to the same spatial size as the input haze image through methods such as bilinear interpolation to achieve spatial alignment.

[0055] Therefore, the first multi-scale fusion feature map extracted by the first feature extraction module is a universal and widely applicable feature. These features can be regarded as basic attributes of haze images and do not depend on specific image content or degradation type.

[0056] (2) The second feature extraction module.

[0057] Optionally, the second feature extraction module is also used to extract features from the haze image to obtain a second multi-scale fusion feature map, and the second feature extraction module at least includes: three adapters and an adapter selection unit.

[0058] Specifically, each adapter performs multiple feature extractions on the haze image to obtain corresponding multiple second feature maps, and the feature extraction of each adapter includes three stages, where:

[0059] The feature extraction stage of the first adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual convolutional layer; a second feature extraction stage consisting of a downsampling block and at least one residual convolutional layer; and a third feature extraction stage consisting of an upsampling block and at least one residual convolutional layer. The convolution kernel of the residual convolution layer is designed to be large (e.g., 7*7). This is because haze images typically have a certain degree of diffusion and non-uniformity. Using a large convolution kernel can expand the receptive field and capture a wider range of contextual information in the image, thereby better handling the degradation of these global properties and helping the model understand the overall structure and layout of the haze image.

[0060] Therefore, the second feature map extracted by the first adapter will contain richer macro information, such as the depth of the scene, lighting conditions, and the relative positions of objects, which helps the model understand the large-scale background and composition of haze images.

[0061] The feature extraction stage of the second adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual dense block (RDB); a second feature extraction stage consisting of a downsampling block and at least one residual dense block; and a third feature extraction stage consisting of an upsampling block and at least one residual dense block. The design of the residual dense block focuses on enhancing the extraction of local features and dense connection of information, aiming to preserve and enhance local details of the image.

[0062] Therefore, the second feature map extracted by the second adapter highlights local and detailed information in hazy images, such as object boundaries, textures, and color changes. This helps restore image details obscured or blurred by local haze, making the model more sensitive and accurate in local features.

[0063] The feature extraction stage of the third adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual block based on central difference convolution (CDC); a second feature extraction stage consisting of a downsampling block and multiple residual blocks based on central difference convolution; and a third feature extraction stage consisting of an upsampling block and at least one residual block based on central difference convolution. The central difference convolution can effectively capture edge information in the image, while the residual block can further enhance the extraction of this edge feature while maintaining the stability of the model.

[0064] Therefore, the second feature map extracted by the third adapter will emphasize edge and contour information, helping the model to clearly distinguish and recover object boundaries in hazy images, even when the edge details are severely obscured by haze.

[0065] In summary, the second feature maps obtained by the three adapters through the three feature extraction stages focus on the global information, local details and edge contours of the haze image respectively.

[0066] In addition, the adapter selection unit can generate a first selection weight corresponding to each adapter based on each second feature map, and add multiple second feature maps output by the adapter with the largest first selection weight to obtain a second multi-scale fusion feature map.

[0067] Specifically, the adapter selection unit can encode multiple second feature maps output by each adapter and extract key descriptors that reflect the characteristics of the feature maps, such as the degree of haze, image complexity, texture details, etc. Then, based on the extracted descriptors, the adapter selection unit calculates the first selection weight of each adapter through an internal weight generation mechanism. This mechanism can be a small neural network, such as a fully connected network or a multi-layer perceptron (MLP), which automatically learns and predicts the applicability of each adapter to the current image based on the encoded information of the feature map. Then, the adapter selection unit compares the first selection weights of each adapter and selects the adapter with the largest weight. This adapter with the largest weight is considered to be the most suitable for the processing requirements of the currently input haze image and can provide the most effective feature extraction and haze removal effects.

[0068] Therefore, the second feature extraction module is intended to adaptively adjust the feature extraction process according to the specific content and degradation condition of the input haze image (such as the distribution and concentration of haze) to obtain more targeted features.

[0069] (3) Feature enhancement module.

[0070] The feature enhancement module may add the first multi-scale fusion feature map output by the first feature extraction module and the second multi-scale fusion feature map output by the second feature extraction module to obtain an enhanced feature map.

[0071] It should be noted that the feature enhancement module can use additional feature refinement modules, such as residual blocks, dense blocks, or attention mechanisms, to perform deep learning processing on the fused features, strengthen important features, and suppress irrelevant or redundant information, thereby improving the feature expression ability and the dehazing performance of the model.

[0072] The dehazed image prediction network in the template haze image dehazing model is used to generate a predicted dehazed image based on the extracted enhanced feature map. The dehazed image prediction network includes at least: multiple third feature extraction modules and a feature mapping layer, where the third feature extraction modules are composed of multiple cascaded residual blocks, and the feature mapping layer includes a convolutional layer (the convolution kernel can be 3*3).

[0073] The feedback optimization network in the template haze image dehazing model is used to determine whether the predicted dehazed image output by the dehazed image prediction network needs further optimization and correction. Therefore, the feedback optimization network includes at least: an evaluation module, three optimization modules, and an optimization selection module, where:

[0074] (1) Evaluation module.

[0075] The evaluation module evaluates the difference between the predicted dehazed image output by the dehazed image prediction network and the preset annotated clear image, and uses the difference as the quality assessment result of the predicted dehazed image. The evaluation module can be implemented by calculating a pixel-level similarity metric between the two images, such as Peak Signal-to-Noise Ratio (PSNR) or Structural Similarity (SSIM).

[0076] (2) Three optimization modules.

[0077] The first optimization module includes multiple cascaded self-attention modules. Therefore, when the difference falls below a preset threshold, the first feature optimization unit can correct erroneous information in the enhanced feature map, generating an optimized enhanced feature map. For example, when the image has uneven brightness or large areas are obscured by haze, the first optimization module can utilize a sub-attention mechanism to enhance the model's understanding of the overall image structure.

[0078] The second optimization module includes multiple cascaded residual dense blocks. Therefore, the second feature optimization unit is used to correct local errors in the enhanced feature map when the difference falls below a preset threshold, resulting in an optimized enhanced feature map. For example, when an image is partially blurred or lacks detail, the second optimization module can use the residual dense blocks to intensively extract and correct local features, improving the clarity and detail richness of specific areas of the image.

[0079] The third optimization module includes multiple cascaded residual blocks based on center-difference convolution. Therefore, the third feature optimization unit is used to enhance texture details in the enhanced feature map when the difference falls below a preset threshold, resulting in an optimized enhanced feature map. Center-difference convolution effectively captures edge information, while residual learning helps preserve subtle textures in the image, ensuring the structural integrity of the dehazed image.

[0080] (3) Optimize the selection module.

[0081] The optimization selection module is used to generate the second selection weight of each optimization module based on the optimized enhanced feature map output by each optimization module, and output the optimized enhanced feature map output by the feature optimization unit with the largest second selection weight.

[0082] Specifically, the optimization selection module includes at least: multiple convolutional layers, a global assessment pooling layer, and a fully connected layer. The multiple convolutional layers perform preliminary processing on the optimized enhanced feature maps from the three optimization modules. The global assessment pooling layer then aggregates the information from each feature map. Finally, a fully connected layer generates the selection weights for each optimization module.

[0083] Therefore, the present embodiment of the application designs a feedback optimization network to achieve dynamic monitoring and adaptive optimization of the prediction results during the image dehazing process, ensuring the final output of high-quality haze-free images. This modular, feedback-driven design greatly improves the model's flexibility and robustness, and allows for precise control over the dehazing effect.

[0084] In the technical solution provided in the above step S106, the target haze image to be processed is input into the template haze image defogging model, and the existing defogging capability and feature extraction mechanism of the template haze image defogging model are utilized to obtain the corresponding target defogging image.

[0085] Furthermore, the target haze image and the target dehazed image are added to the training sample set and sample label set, respectively, thereby updating the training sample set and sample label set. The updated training sample set and sample label set are then used to perform self-supervised training on the template haze image dehazing model. This training process does not rely on manually labeled clear images, but instead uses the difference between the model's input and output for the same image as feedback signals to adjust and optimize the model's parameters.

[0086] During this self-supervised training process, the template haze image dehazing model continuously adjusts its parameters to more efficiently and accurately process haze images, generating high-quality dehazed images. This training approach not only enhances the model's generalization capabilities but also enables personalized adjustments based on specific haze scenarios, resulting in superior dehazing results in practical applications.

[0087] Example 2

[0088] According to an embodiment of the present application, a haze image defogging model training device for implementing the haze image defogging model training method in Example 1 is also provided. Figure 3 As shown, the haze image defogging model training device at least includes: an acquisition module 32, a template generation module 34, an inference module 36 and a training module 38, wherein:

[0089] An acquisition module 32 is configured to acquire a training sample set and a sample label set, wherein the training sample set includes a plurality of haze images as training samples, and the sample label set includes a dehazed image corresponding to each haze image as a sample label corresponding to the corresponding training sample;

[0090] The template generation module 34 is used to train a template haze image defogging model using the training sample set and the sample label set, wherein the template haze image defogging model at least includes: a feature processing network, a defogging image prediction network, and a feedback optimization network;

[0091] The inference module 36 is used to input the target haze image to be processed into the template haze image defogging model to obtain the corresponding target defogging image;

[0092] The training module 38 is used to add the target haze image and the target dehazed image to the training sample set and the sample label set respectively, and use the new training sample set and the new sample label set to perform self-supervised training on the template haze image dehazing model to obtain the target haze image dehazing model.

[0093] The following describes the functions and structures of each component in the template haze image dehazing model.

[0094] The feature processing network in the template haze image dehazing model is used to extract, fuse, and enhance features of the haze image, mining and integrating all information that is helpful for dehazing. Therefore, the feature processing network includes at least: a first feature extraction module, a second feature extraction module, and a feature enhancement module, where:

[0095] (1) The first feature extraction module.

[0096] Optionally, the first feature extraction module is used to extract features from the haze image to obtain a first multi-scale fusion feature map, and the first feature extraction module at least includes: multiple MBConv (Mobile Inverted Bottleneck Convolution) modules connected in series and a feature fusion module.

[0097] Specifically, the multiple MBConv modules can be the first 7 layers of EfficientNet, and each MBConv module can capture the features of the input haze image at different scales to obtain the corresponding first feature map.

[0098] Among them, the structure of the MBConv module mainly includes the following parts:

[0099] Expansion Layer: This layer uses 1×1 convolution to expand the number of channels in the input haze image, forming an expanded feature map. This layer increases the network's expressive power, allowing subsequent depthwise separable convolutions to process richer information.

[0100] Depthwise separable convolution layer: Use 3×3 depthwise separable convolution to perform spatial processing on the expanded feature map.

[0101] Compression and excitation unit: It consists of two parts: Squeeze and Excitation. The Squeeze part compresses the feature map output by the depthwise separable convolutional layer into a channel vector through global average pooling; while the Excitation part processes the channel vector through two fully connected layers to generate a weight vector, and then multiplies this weight vector with the feature map output by the depthwise separable convolutional layer channel by channel to achieve dynamic adjustment of the importance of different channel information and enhance the information interaction between channels of the network.

[0102] Point-wise convolution layer: A 1×1 point-wise convolution is used after the SE unit to adjust the number of channels to restore the number of channels of the feature map to the size before the expansion layer, or to adjust it according to different stages of the network.

[0103] Therefore, each MBConv module in the first feature extraction module performs feature extraction and information processing through the network layer of the above-mentioned fixed architecture, and can obtain multi-level information commonly present in haze images, such as basic features such as image texture, color, and edges, and provide efficient and rich intermediate feature representations for subsequent deeper feature learning. It should be noted that the parameters of each MBConv module in the first feature extraction module (such as expansion ratio, size and number of depth-separable convolutions, setting of activation functions, etc.) can be set according to the actual application scenario. The embodiments of this application do not impose specific restrictions on this.

[0104] Furthermore, since the first feature maps output by each MBConv module carry information of the input haze image at different scales, in order to ensure that the feature maps at each scale can be effectively processed and fused, the embodiment of the present application also designs a feature fusion module in the first feature extraction module.

[0105] Optionally, the feature fusion module includes multiple convolution layers (such as the convolution kernel size can be 1*1, and the output channels are 32), and the number of convolution layers is equal to the number of first feature maps, so as to ensure that each first feature map can be processed by a dedicated convolution layer. Among them, the role of the convolution layer is to perform channel dimension conversion, that is, to convert the first feature map into a unified output dimension. On the one hand, this can reduce the computational complexity, and on the other hand, it can also maintain or enhance the key features of the image. Then, each convolution layer performs feature fusion on the first image features after the dimension conversion of each first feature map. Among them, feature fusion can be performed by adding or splicing to aggregate the first image features of different scales and after dimension adjustment to obtain a first multi-scale fusion feature map.

[0106] In addition, since the resolutions of the first feature maps output by each MBConv module are different, the first feature maps of the lower layer have a higher resolution and can therefore carry more detail information; while the first feature maps of the higher layer have a lower resolution and therefore contain more abstract and global information. Therefore, directly fusing multiple first feature maps output by multiple MBConv modules may cause information misalignment and affect the fusion effect. To this end, an embodiment of the present application proposes that before performing channel conversion, the first feature maps of different scales can be expanded to the same spatial size as the input haze image through methods such as bilinear interpolation to achieve spatial alignment.

[0107] Therefore, the first multi-scale fusion feature map extracted by the first feature extraction module is a universal and widely applicable feature. These features can be regarded as basic attributes of haze images and do not depend on specific image content or degradation type.

[0108] (2) The second feature extraction module.

[0109] Optionally, the second feature extraction module is also used to extract features from the haze image to obtain a second multi-scale fusion feature map, and the second feature extraction module at least includes: three adapters and an adapter selection unit.

[0110] Specifically, each adapter performs multiple feature extractions on the haze image to obtain corresponding multiple second feature maps, and the feature extraction of each adapter includes three stages, where:

[0111] The feature extraction stage of the first adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual convolutional layer; a second feature extraction stage consisting of a downsampling block and at least one residual convolutional layer; and a third feature extraction stage consisting of an upsampling block and at least one residual convolutional layer. The convolution kernel of the residual convolution layer is designed to be large (e.g., 7*7). This is because haze images typically have a certain degree of diffusion and non-uniformity. Using a large convolution kernel can expand the receptive field and capture a wider range of contextual information in the image, thereby better handling the degradation of these global properties and helping the model understand the overall structure and layout of the haze image.

[0112] Therefore, the second feature map extracted by the first adapter will contain richer macro information, such as the depth of the scene, lighting conditions, and the relative positions of objects, which helps the model understand the large-scale background and composition of haze images.

[0113] The feature extraction stage of the second adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual dense block, a second feature extraction stage consisting of a downsampling block and at least one residual dense block, and a third feature extraction stage consisting of an upsampling block and at least one residual dense block. The design of the residual dense block focuses on enhancing the extraction of local features and the dense connection of information, aiming to preserve and enhance local details of the image.

[0114] Therefore, the second feature map extracted by the second adapter highlights local and detailed information in hazy images, such as object boundaries, textures, and color changes. This helps restore image details obscured or blurred by local haze, making the model more sensitive and accurate in local features.

[0115] The feature extraction stage of the third adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual block based on center-difference convolution; a second feature extraction stage consisting of a downsampling block and multiple residual blocks based on center-difference convolution; and a third feature extraction stage consisting of an upsampling block and at least one residual block based on center-difference convolution. Center-difference convolution can effectively capture edge information in images, while residual blocks can further enhance the extraction of such edge features while maintaining model stability.

[0116] Therefore, the second feature map extracted by the third adapter will emphasize edge and contour information, helping the model to clearly distinguish and recover object boundaries in hazy images, even when the edge details are severely obscured by haze.

[0117] In summary, the second feature maps obtained by the three adapters through the three feature extraction stages focus on the global information, local details and edge contours of the haze image respectively.

[0118] In addition, the adapter selection unit can generate a first selection weight corresponding to each adapter based on each second feature map, and add multiple second feature maps output by the adapter with the largest first selection weight to obtain a second multi-scale fusion feature map.

[0119] Specifically, the adapter selection unit can encode multiple second feature maps output by each adapter and extract key descriptors that reflect the characteristics of the feature maps, such as the degree of haze, image complexity, texture details, etc. Then, based on the extracted descriptors, the adapter selection unit calculates the first selection weight of each adapter through an internal weight generation mechanism. This mechanism can be a small neural network, such as a fully connected network or a multi-layer perceptron (MLP), which automatically learns and predicts the applicability of each adapter to the current image based on the encoded information of the feature map. Then, the adapter selection unit compares the first selection weights of each adapter and selects the adapter with the largest weight. This adapter with the largest weight is considered to be the most suitable for the processing requirements of the currently input haze image and can provide the most effective feature extraction and haze removal effects.

[0120] Therefore, the second feature extraction module is intended to adaptively adjust the feature extraction process according to the specific content and degradation condition of the input haze image (such as the distribution and concentration of haze) to obtain more targeted features.

[0121] (3) Feature enhancement module.

[0122] The feature enhancement module may add the first multi-scale fusion feature map output by the first feature extraction module and the second multi-scale fusion feature map output by the second feature extraction module to obtain an enhanced feature map.

[0123] It should be noted that the feature enhancement module can use additional feature refinement modules, such as residual blocks, dense blocks, or attention mechanisms, to perform deep learning processing on the fused features, strengthen important features, and suppress irrelevant or redundant information, thereby improving the feature expression ability and the dehazing performance of the model.

[0124] The dehazed image prediction network in the template haze image dehazing model is used to generate a predicted dehazed image based on the extracted enhanced feature map. The dehazed image prediction network includes at least: multiple third feature extraction modules and a feature mapping layer, where the third feature extraction modules are composed of multiple cascaded residual blocks, and the feature mapping layer includes a convolutional layer (the convolution kernel can be 3*3).

[0125] The feedback optimization network in the template haze image dehazing model is used to determine whether the predicted dehazed image output by the dehazed image prediction network needs further optimization and correction. Therefore, the feedback optimization network includes at least: an evaluation module, three optimization modules, and an optimization selection module, where:

[0126] (1) Evaluation module.

[0127] The evaluation module evaluates the difference between the predicted dehazed image output by the dehazed image prediction network and the preset annotated clear image, and uses the difference as the quality assessment result of the predicted dehazed image. The evaluation module can be implemented by calculating a pixel-level similarity metric between the two images, such as Peak Signal-to-Noise Ratio (PSNR) or Structural Similarity (SSIM).

[0128] (2) Three optimization modules.

[0129] The first optimization module includes multiple cascaded self-attention modules. Therefore, when the difference falls below a preset threshold, the first feature optimization unit can correct erroneous information in the enhanced feature map, generating an optimized enhanced feature map. For example, when the image has uneven brightness or large areas are obscured by haze, the first optimization module can utilize a sub-attention mechanism to enhance the model's understanding of the overall image structure.

[0130] The second optimization module includes multiple cascaded residual dense blocks. Therefore, the second feature optimization unit is used to correct local errors in the enhanced feature map when the difference falls below a preset threshold, resulting in an optimized enhanced feature map. For example, when an image is partially blurred or lacks detail, the second optimization module can use the residual dense blocks to intensively extract and correct local features, improving the clarity and detail richness of specific areas of the image.

[0131] The third optimization module includes multiple cascaded residual blocks based on center-difference convolution. Therefore, the third feature optimization unit is used to enhance texture details in the enhanced feature map when the difference falls below a preset threshold, resulting in an optimized enhanced feature map. Center-difference convolution effectively captures edge information, while residual learning helps preserve subtle textures in the image, ensuring the structural integrity of the dehazed image.

[0132] (3) Optimize the selection module.

[0133] The optimization selection module is used to generate the second selection weight of each optimization module based on the optimized enhanced feature map output by each optimization module, and output the optimized enhanced feature map output by the feature optimization unit with the largest second selection weight.

[0134] Specifically, the optimization selection module includes at least: multiple convolutional layers, a global assessment pooling layer, and a fully connected layer. The multiple convolutional layers perform preliminary processing on the optimized enhanced feature maps from the three optimization modules. The global assessment pooling layer then aggregates the information from each feature map. Finally, a fully connected layer generates the selection weights for each optimization module.

[0135] Therefore, the present embodiment of the application designs a feedback optimization network to achieve dynamic monitoring and adaptive optimization of the prediction results during the image dehazing process, ensuring the final output of high-quality haze-free images. This modular, feedback-driven design greatly improves the model's flexibility and robustness, and allows for precise control over the dehazing effect.

[0136] It should be noted that the modules in the haze image defogging model training device in the embodiment of the present application correspond one-to-one to the implementation steps of the haze image defogging model training method in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be elaborated here.

[0137] Example 3

[0138] According to an embodiment of the present application, a computer program product is also provided, which includes a computer program, wherein when the computer program is executed by a processor, the haze image defogging model training method in Example 1 is implemented.

[0139] According to an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the haze image defogging model training method in Example 1 by running the computer program.

[0140] According to an embodiment of the present application, a processor is further provided, which is used to run a computer program, wherein the haze image defogging model training method in Example 1 is executed when the computer program is running.

[0141] According to an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the haze image defogging model training method in Example 1 through the computer program.

[0142] Specifically, the computer program executes the following steps when it is running: obtaining a training sample set and a sample label set, wherein the training sample set includes multiple haze images as training samples, and the sample label set includes a dehazed image corresponding to each haze image as a sample label corresponding to the corresponding training sample; using the training sample set and the sample label set to train a template haze image dehazing model, wherein the template haze image dehazing model includes at least: a feature processing network, a dehazed image prediction network, and a feedback optimization network; inputting the target haze image to be processed into the template haze image dehazing model to obtain a corresponding target dehazed image; adding the target haze image and the target dehazed image to the training sample set and the sample label set respectively, and using the new training sample set and the new sample label set to perform self-supervised training on the template haze image dehazing model to obtain a target haze image dehazing model.

[0143] As an optional implementation, the electronic device may be in the form of a mobile terminal, a computer terminal or a similar computing device. Figure 4 The hardware structure block diagram of an electronic device for implementing a haze image defogging model training method is shown. Figure 4 As shown, the electronic device 40 may include one or more (402a, 402b, ..., 402n are shown in the figure) processors 402 (the processor 402 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 404 for storing data, and a transmission device 406 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 4 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 4 More or fewer components than shown, or with Figure 4 Different configurations shown.

[0144] It should be noted that the one or more processors 402 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the electronic device 40. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0145] The memory 404 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the haze image defogging model training method in the embodiment of the present application. The processor 402 executes various functional applications and data processing by running the software programs and modules stored in the memory 404, that is, implementing the vulnerability detection method of the above-mentioned application. The memory 404 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 404 may further include a memory remotely located relative to the processor 402, and these remote memories may be connected to the electronic device 40 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0146] Transmission device 406 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communications provider of electronic device 40. In one embodiment, transmission device 406 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 406 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0147] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 40 .

[0148] The serial numbers of the above embodiments are for description only and do not represent the advantages or disadvantages of the embodiments.

[0149] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0150] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0151] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.

[0152] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0153] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.

[0154] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A haze image defogging model training method, characterized in that: include: Obtaining a training sample set and a sample label set, wherein the training sample set includes a plurality of haze images as training samples, and the sample label set includes a dehazed image corresponding to each haze image as a sample label corresponding to the corresponding training sample; A template haze image defogging model is obtained by training using the training sample set and the sample label set, wherein the template haze image defogging model at least includes: a feature processing network, a defogging image prediction network, and a feedback optimization network; Inputting the target haze image to be processed into the template haze image de-model to obtain the corresponding target de-hazed image; The target haze image and the target dehazed image are added to the training sample set and the sample label set respectively, and the template haze image dehazing model is self-supervised trained using the new training sample set and the new sample label set to obtain the target haze image dehazing model.

2. The method according to claim 1, characterized in that The feature processing network at least includes: a first feature extraction module, a second feature extraction module and a feature enhancement module, wherein: The first feature extraction module at least includes: a plurality of moving inverted bottleneck convolution MBConv modules connected in series, and the first feature extraction module is used to extract features from the haze image to obtain a first multi-scale fusion feature map; The second feature extraction module includes at least three adapters, and the second feature extraction module is used to extract features from the haze image to obtain a second multi-scale fusion feature map; The feature enhancement module is used to add the first multi-scale fusion feature map output by the first feature extraction module and the second multi-scale fusion feature map output by the second feature extraction module to obtain an enhanced feature map.

3. The method according to claim 2, characterized in that The first feature extraction module also includes a feature fusion module, wherein: The MBConv module includes at least: an expansion layer, a depth-separable convolution layer, a compression and excitation unit, and a point-by-point convolution layer, and each of the MBConv modules is used to capture features of the haze image at different scales to obtain a corresponding first feature map; The feature fusion module includes multiple convolution layers, and the number of the convolution layers is equal to the number of the first feature maps. The first feature fusion module is used to perform feature fusion on the first image features after each convolution layer performs dimensional conversion on each of the first feature maps to obtain the first multi-scale fusion feature map.

4. The method according to claim 2, characterized in that The second feature extraction module further includes an adapter selection unit, wherein: Each of the adapters is configured to perform multiple feature extractions on the haze image to obtain corresponding multiple second feature maps; The adapter selection unit is used to generate a first selection weight corresponding to each of the adapters based on each of the second feature maps, and add multiple second feature maps output by the adapter with the largest first selection weight to obtain the second multi-scale fusion feature map.

5. The method according to claim 4, characterized in that The feature extraction stage of the first adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual convolutional layer, a second feature extraction stage consisting of a downsampling block and at least one residual convolutional layer, and a third feature extraction stage consisting of an upsampling block and at least one residual convolutional layer; The feature extraction stage of the second adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual dense block, a second feature extraction stage consisting of a downsampling block and at least one residual dense block, and a third feature extraction stage consisting of an upsampling block and at least one residual dense block; The feature extraction stage of the third adapter includes: a first feature extraction stage consisting of a convolutional layer and at least one residual block based on center difference convolution, a second feature extraction stage consisting of a downsampling block and at least one residual block based on center difference convolution, and a third feature extraction stage consisting of an upsampling block and at least one residual block based on center difference convolution.

6. The method according to claim 1, characterized in that The dehazed image prediction network at least includes: multiple third feature extraction modules and a feature mapping layer, and the third feature extraction module is composed of multiple cascaded residual blocks, and the feature mapping layer includes a convolution layer.

7. The method according to claim 2, characterized in that The feedback optimization network includes at least: an evaluation module, three optimization modules and an optimization selection module, wherein: The evaluation module is configured to evaluate the difference between the predicted dehazed image output by the dehazed image prediction network and a preset labeled clear image, and use the difference as a quality evaluation result of the predicted dehazed image; The first optimization module includes a plurality of cascaded self-attention modules, and the first feature optimization unit is used to correct the erroneous information in the enhanced feature map when the difference is lower than a preset threshold value to obtain the optimized enhanced feature map; The second optimization module includes a plurality of cascaded residual dense blocks, and the second feature optimization unit is used to correct local errors in the enhanced feature map when the difference is lower than a preset threshold value to obtain the optimized enhanced feature map; The third optimization module includes a plurality of cascaded residual blocks based on central difference convolution, and the third feature optimization unit is used to enhance the texture details in the enhanced feature map when the difference is lower than a preset threshold value to obtain the optimized enhanced feature map; The optimization selection module includes at least: multiple convolutional layers, global evaluation pooling layers, and fully connected layers, and the optimization selection module is used to generate the second selection weight of each optimization module based on the optimized enhanced feature map output by each optimization module, and output the optimized enhanced feature map output by the feature optimization unit with the largest second selection weight.

8. A haze image defogging model training device, characterized in that: include: an acquisition module, configured to acquire a training sample set and a sample label set, wherein the training sample set includes a plurality of haze images as training samples, and the sample label set includes a dehazed image corresponding to each haze image as a sample label corresponding to the corresponding training sample; A template generation module, configured to train a template haze image defogging model using the training sample set and the sample label set, wherein the template haze image defogging model comprises at least: a feature processing network, a defogging image prediction network, and a feedback optimization network; an inference module, configured to input the target haze image to be processed into the template haze image defogging model to obtain a corresponding target defogging image; A training module is used to add the target haze image and the target dehazed image to the training sample set and the sample label set respectively, and use the new training sample set and the new sample label set to perform self-supervised training on the template haze image dehazing model to obtain the target haze image dehazing model.

9. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the haze image defogging model training method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the haze image defogging model training method according to any one of claims 1 to 7 through the computer program.