A method, apparatus, and storage medium for evaluating image dehazing quality.
By preprocessing and feature analysis of the dehazed images to be evaluated, a training model is constructed, which solves the shortcomings of existing technologies in the quality assessment of dehazed images without reference. It realizes the collaborative perception of image content, distortion and fog density information, and improves the accuracy and consistency of image quality assessment.
Patent Information
- Application Number
- CN202511269850.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing no-reference dehazing image quality assessment methods suffer from insufficient perceptual models, limited content extraction, weak distortion recognition capabilities, lack of fog density quantification, and poor result interpretability, making it difficult to meet the requirements of high precision, reliability, and interpretability in practical applications.
By constructing a training model, the dehazing image to be evaluated is preprocessed, and feature analysis of image content, distortion, and fog density is performed to obtain the target dehazing feature map. Evaluation analysis is then conducted to assess the quality of image dehazing.
It improves the accuracy and subjective consistency of image quality assessment, has strong practical value and promotion potential, and realizes the collaborative perception of image content, distortion and fog density information.
Smart Images

Figure CN120765644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image evaluation technology, specifically to an image dehazing quality evaluation method, apparatus, and storage medium. Background Technology
[0002] Currently, image dehazing algorithms have been widely used in scenarios such as autonomous driving, remote sensing monitoring, and video security. They mainly use physical models or deep learning methods to restore degraded images, thereby improving their clarity and visibility. These technologies have achieved significant results in image contrast, detail, and color restoration.
[0003] To objectively evaluate the effectiveness of dehazing algorithms, researchers have proposed various image dehazing quality assessment methods, including full-reference, partial-reference, and no-reference methods. Since obtaining haze-free reference images is difficult in real-world scenarios, the no-reference quality assessment method is more practical because it eliminates the need for a reference image.
[0004] Existing no-reference dehazing image quality assessment methods are mostly based on image statistics, structure, or depth features. However, they suffer from shortcomings such as insufficient perceptual models, limited content extraction, weak ability to identify distortions, lack of fog density quantification, and poor interpretability of results, making it difficult to meet the requirements of high precision, reliability, and interpretability in practical applications. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide an image dehazing quality assessment method, apparatus and storage medium that address the shortcomings of the prior art.
[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: An image dehazing quality evaluation method, comprising the following steps:
[0007] Import multiple dehazed images to be evaluated, as well as the original foggy images corresponding to each of the dehazed images to be evaluated;
[0008] Each of the dehazed images to be evaluated is preprocessed to obtain a preprocessed dehazed image corresponding to each of the dehazed images to be evaluated.
[0009] A training model is constructed, and feature analysis is performed on each of the preprocessed dehazing images and the original foggy images corresponding to each of the dehazing images to be evaluated, to obtain target dehazing feature maps corresponding to each of the dehazing images to be evaluated.
[0010] The dehazing feature maps of each target are evaluated and analyzed to obtain the evaluation results of the image dehazing quality.
[0011] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: An image dehazing quality assessment device, comprising:
[0012] The import module is used to import multiple dehazing images to be evaluated and the original foggy images corresponding to each of the dehazing images to be evaluated.
[0013] The preprocessing module is used to preprocess each of the dehazing images to be evaluated to obtain a preprocessed dehazing image corresponding to each of the dehazing images to be evaluated.
[0014] The feature analysis module is used to build a training model. The training model is used to perform feature analysis on each of the preprocessed dehazing images and the original foggy images corresponding to each of the dehazing images to be evaluated, so as to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated.
[0015] The evaluation result acquisition module is used to evaluate and analyze the dehazing feature maps of each target respectively, and obtain the evaluation result of the image dehazing quality.
[0016] Based on the above-mentioned image dehazing quality assessment method, the present invention also provides an image dehazing quality assessment system.
[0017] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: an image dehazing quality assessment system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the image dehazing quality assessment method as described above.
[0018] Based on the above-described image dehazing quality assessment method, the present invention also provides a computer-readable storage medium.
[0019] Another technical solution of the present invention to solve the above-mentioned technical problems is as follows: a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the image dehazing quality assessment method as described above.
[0020] The beneficial effects of this invention are as follows: by preprocessing the image to be evaluated for dehazing, a preprocessed dehazed image is obtained; by training a model to analyze the features of the preprocessed dehazed image and the original foggy image, a target dehazed feature map is obtained; and by evaluating and analyzing the target dehazed feature map, an evaluation result of the image dehazing quality is obtained. This achieves the collaborative perception of image content, distortion, and fog density information, improves the accuracy and subjective consistency of image quality evaluation, and has strong practical value and promotion potential. Attached Figure Description
[0021] Figure 1 This is one of the flowcharts illustrating the image dehazing quality assessment method provided in this embodiment of the invention;
[0022] Figure 2 This is a second schematic flowchart of the image dehazing quality assessment method provided in an embodiment of the present invention;
[0023] Figure 3 This is a processing logic diagram of the image content-aware pre-training of the image dehazing quality assessment method provided in the embodiments of the present invention;
[0024] Figure 4 The processing logic diagram for image distortion perception pre-training of the image dehazing quality assessment method provided in the embodiments of the present invention;
[0025] Figure 5 This is a processing logic diagram of the image fog density perception pre-training of the image dehazing quality assessment method provided in the embodiments of the present invention;
[0026] Figure 6 The processing logic diagram of the content-distortion-fog density perceived feature representation self-interaction module of the image dehazing quality assessment method provided in the embodiments of the present invention is shown below.
[0027] Figure 7 This is a processing logic diagram of the dual-branch quality predictor of the image dehazing quality assessment method provided in this embodiment of the invention;
[0028] Figure 8 A block diagram of an image dehazing quality assessment device provided in an embodiment of the present invention. Detailed Implementation
[0029] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0030] Figure 1 This is a flowchart illustrating an image dehazing quality assessment method provided in an embodiment of the present invention.
[0031] like Figure 1 As shown, an image dehazing quality assessment method includes the following steps:
[0032] S1: Import multiple dehazing images to be evaluated and the original foggy images corresponding to each of the dehazing images to be evaluated;
[0033] S2: Preprocess each of the dehazing images to be evaluated to obtain preprocessed dehazing images corresponding to each of the dehazing images to be evaluated;
[0034] S3: Construct a training model, and perform feature analysis on each of the preprocessed dehazing images and the original foggy images corresponding to each of the dehazing images to be evaluated, to obtain target dehazing feature maps corresponding to each of the dehazing images to be evaluated.
[0035] S4: Evaluate and analyze each of the target dehazing feature maps to obtain the evaluation results of the image dehazing quality.
[0036] In the above embodiments, a preprocessed dehazed image is obtained by preprocessing the dehazed image to be evaluated. The target dehazed feature map is obtained by analyzing the features of the preprocessed dehazed image and the original foggy image through a trained model. The evaluation and analysis of the target dehazed feature map yields the evaluation result of the image dehazed quality. This achieves the collaborative perception of image content, distortion and fog density information, improves the accuracy and subjective consistency of image quality evaluation, and has strong practical value and promotion potential.
[0037] Optionally, as an embodiment of the present invention, the process of preprocessing each of the dehazing images to be evaluated to obtain a preprocessed dehazing image corresponding to each of the dehazing images to be evaluated includes:
[0038] Each of the dehazing images to be evaluated is resized to obtain an adjusted dehazing image corresponding to each of the dehazing images to be evaluated.
[0039] Each of the adjusted dehazed images is normalized to obtain a preprocessed dehazed image corresponding to each of the dehazed images to be evaluated.
[0040] It should be understood that the input dehazed image to be evaluated (i.e., the dehazed image to be evaluated) is subjected to standardized preprocessing, including resizing and normalization, as a uniform input to the model.
[0041] In the above embodiments, each dehazed image to be evaluated is preprocessed to obtain a preprocessed dehazed image, which realizes the collaborative perception of image content, distortion and fog density information, improves the accuracy and subjective consistency of image quality assessment, and has strong practical value and promotion potential.
[0042] Optionally, as an embodiment of the present invention, the training model includes an image content-aware network, an image distortion-aware network, and an image fog density-aware network.
[0043] The process of performing feature analysis on each of the preprocessed dehazed images and the original foggy images corresponding to each of the dehazed images to be evaluated using the trained model to obtain the target dehazed feature map corresponding to each of the dehazed images to be evaluated includes:
[0044] The image content-aware network is used to analyze the image content of each of the preprocessed dehazed images to obtain the target content-aware encoder and the original content-aware feature map corresponding to each of the dehazed images to be evaluated.
[0045] The image distortion perception network is used to analyze the image distortion perception of each of the preprocessed dehazed images to obtain the target distortion perception encoder and the original distortion perception feature set corresponding to each of the dehazed images to be evaluated.
[0046] The image fog density sensing network analyzes the fog density of each original foggy image and the preprocessed defogging image corresponding to each defogging image to be evaluated, and obtains a fog density predictor and the original fog density sensing feature map corresponding to each defogging image to be evaluated.
[0047] The target content-aware encoder, the target distortion-aware encoder, and the fog density predictor respectively perform fusion analysis on each of the original content-aware feature maps, the original distortion-aware feature sets corresponding to each of the dehazing images to be evaluated, and the original fog density-aware feature maps corresponding to each of the dehazing images to be evaluated, to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated.
[0048] In the above embodiments, the training model performs feature analysis on each preprocessed dehazed image and the original foggy image to obtain the target dehazed feature map, thereby achieving a comprehensive and scientific evaluation of the dehazed image and realizing the collaborative perception of image content, distortion and fog density information, thus improving the accuracy and subjective consistency of image quality assessment.
[0049] Optionally, as an embodiment of the present invention, the image content-aware network includes a raw content-aware encoder and a content-aware decoder.
[0050] The process of performing image content-aware analysis on each of the preprocessed dehazed images using the image content-aware network to obtain the target content-aware encoder and the original content-aware feature map corresponding to each of the dehazed images to be evaluated includes:
[0051] The original content-aware encoder is used to extract features from each of the preprocessed dehazed images to obtain the original content-aware feature map corresponding to each of the dehazed images to be evaluated.
[0052] The content-aware decoder is used to reconstruct each of the original content-aware feature maps to obtain a first reconstructed clear image corresponding to each of the dehazed images to be evaluated.
[0053] The first loss function is obtained by calculating the loss function on all the preprocessed dehazed images and all the first reconstructed clear images using the mean square error loss function algorithm.
[0054] Perceptual image patch similarity is calculated for all the preprocessed dehazed images and all the first reconstructed clear images, and the calculation result is used as the second loss function;
[0055] The third loss function is obtained by calculating the first loss function and the second loss function using the first equation. The first equation is:
[0056] ,
[0057] in, For the third loss function, For the first loss function, The first weighted hyperparameter, This is the second loss function;
[0058] The target content-aware encoder is obtained by updating the parameters of the original content-aware encoder using the third loss function.
[0059] It should be understood that the content-aware encoder-decoder network (i.e., image content-aware network) is trained with clear images (i.e., pre-processed dehazed images) to extract structural semantic features of the images and optimize reconstruction capabilities.
[0060] It should be understood that this module (i.e., the image content-aware network) aims to extract features in images that are highly relevant to the content information. Specifically, it uses a self-supervised learning mechanism to construct an encoder-decoder network structure (i.e., the image content-aware network) for perceptual modeling of the content of clear (fog-free) images (i.e., pre-processed defog images).
[0061] Specifically, the module (i.e., the image content-aware network) includes the following steps:
[0062] The input clear image (i.e., the pre-processed dehazed image) is fed into the content-aware encoder (i.e., the original content-aware encoder). The content-aware encoder (i.e., the original content-aware encoder) extracts features from the image content (i.e., the pre-processed dehazed image) to obtain content-aware features. (i.e., the original content-aware feature map);
[0063] Content-aware features The original content-aware feature map is input into the content-aware decoder (i.e., the content-aware decoder) for image reconstruction, generating a clear image (i.e., the first reconstructed clear image) that is consistent with the content structure of the input image.
[0064] Mean Squared Error (MSE) Loss Function (i.e., the first loss function) measures the pixel-level error between the input image and the reconstructed image;
[0065] Furthermore, a loss function based on perceptual similarity is introduced. (i.e., the second loss function), which is calculated based on the learned perceptual image patch similarity (LPIPS) index, enhances the model's ability to preserve content structure and texture information;
[0066] The total loss function is a weighted combination of the two, as shown in the following equation:
[0067] ,
[0068] in, is a weighting hyperparameter used to balance the sensing loss and reconstruction error.
[0069] It should be understood that the obtained content-aware encoder (i.e., the target content-aware encoder) has robust image content understanding and representation capabilities, providing structural content support for subsequent modules.
[0070] In the above embodiments, the image content-aware network analyzes the image content of each preprocessed dehazed image to obtain the target content-aware encoder and the original content-aware feature map, which has robust image content understanding and representation capabilities, providing support for subsequent structured content.
[0071] Optionally, as an embodiment of the present invention, the image distortion sensing network includes a raw distortion sensing encoder and a distortion sensing decoder.
[0072] The process of analyzing image distortion perception of each preprocessed dehazed image and the original content-aware feature map corresponding to each dehazed image to be evaluated through the image distortion perception network to obtain the target distortion perception encoder and the original distortion perception feature set corresponding to each dehazed image to be evaluated includes:
[0073] The original distortion-aware encoder extracts features from each of the preprocessed dehazed images to obtain multiple original distortion-aware feature maps corresponding to each of the dehazed images to be evaluated. The multiple original distortion-aware feature maps corresponding to each of the dehazed images to be evaluated are then combined to obtain an original distortion-aware feature set corresponding to each of the dehazed images to be evaluated.
[0074] The distortion-aware decoder performs image reconstruction on each of the original distortion-aware feature sets and the original content-aware feature maps corresponding to each of the dehazing images to be evaluated, to obtain a second reconstructed clear image corresponding to each of the dehazing images to be evaluated.
[0075] The mean squared error loss function algorithm is used to calculate the loss function for all the preprocessed dehazed images and all the second reconstructed clear images to obtain the fourth loss function.
[0076] Perceptual image patch similarity is calculated for all the preprocessed dehazed images and all the second reconstructed clear images, and the calculation result is used as the fifth loss function;
[0077] The sixth loss function is obtained by calculating the fourth and fifth loss functions using the second equation, whereby the second equation is:
[0078] ,
[0079] in, The sixth loss function, This is the fourth loss function. This is the second weight hyperparameter. This is the fifth loss function;
[0080] The original distortion-aware encoder is updated using the sixth loss function to obtain the target distortion-aware encoder.
[0081] It should be understood that, with the aid of content features, the distortion-aware encoder (i.e., the original distortion-aware encoder) is trained using the dehazed image (i.e., the pre-processed dehazed image) to learn the distortion-related features introduced by dehazing.
[0082] It should be understood that this module (i.e., the image distortion-aware network) is used to learn the potential distortion information introduced after image dehazing. Based on the pre-modeling of image content, i.e., the content-aware encoder (i.e., the original distortion-aware encoder) has been trained, it extracts distortion-related representational features. Its network structure is similar to the image content-aware pre-training module (i.e., the image content-aware network), with the main differences being in the input source and decoding strategy.
[0083] Specifically, this module (i.e., the image distortion perception network) includes:
[0084] The input image is a dehazed image (i.e., a pre-processed dehazed image), which is input to a distortion-aware encoder (i.e., the original distortion-aware encoder) for distortion-aware feature extraction.
[0085] In the process of reconstructing an image, the distortion-aware decoder (i.e., the original set of distortion-aware features) not only receives distortion-aware features from the distortion-aware decoder (i.e., the original set of distortion-aware features), but also additionally introduces content-aware features obtained from the pre-trained content-aware encoder. (i.e., the original content-aware feature map) to improve the quality of reconstruction;
[0086] The distortion-aware decoder (i.e., the distortion-aware decoder) uses the original input dehazed image (i.e., the pre-processed dehazed image) as the reconstruction target, and forces the distortion encoder to focus on feature modeling of distortion information by sharing content information;
[0087] The same hybrid loss function consisting of MSE and LPIPS is used for training, consistent with the image content-aware pre-training module.
[0088] It should be understood that, through pre-training of the module (i.e., the image distortion-aware network), the distortion-aware encoder (i.e., the original distortion-aware encoder) is able to perceive distortion features related to the restoration process, thereby enhancing the system's sensitivity to image quality degradation.
[0089] In the above embodiments, the image distortion perception network analyzes the image distortion perception of each preprocessed dehazed image and the original content-aware feature map to obtain the target distortion perception encoder and the original distortion perception feature set. This enables the perception of distortion features related to the restoration process and enhances the sensitivity to image quality degradation.
[0090] Optionally, as an embodiment of the present invention, the image fog density sensing network includes a first large kernel selective convolution, a first max pooling layer, a first average pooling layer, a second large kernel selective convolution, a 1×1 convolutional layer, and a Sigmoid activation layer.
[0091] The process of analyzing the image fog density perception of each original foggy image and the preprocessed defogging image corresponding to each defogging image to be evaluated through the image fog density perception network to obtain a fog density predictor and the original fog density perception feature map corresponding to each defogging image to be evaluated includes:
[0092] The first kernel selective convolution is used to extract features from each of the original foggy images to obtain the original fog density sensing feature map corresponding to each of the defogging images to be evaluated.
[0093] The first max pooling layer is used to perform pooling processing on each of the original fog density sensing feature maps to obtain the first pooled fog density sensing feature map corresponding to each of the dehazing images to be evaluated.
[0094] The first average pooling layer is used to pool each of the original fog density sensing feature maps to obtain a second pooled fog density sensing feature map corresponding to each of the defogging images to be evaluated.
[0095] Each first pooled fog density sensing feature map is concatenated and fused with the second pooled fog density sensing feature map corresponding to each dehazing image to be evaluated, to obtain a fused fog density sensing feature map corresponding to each dehazing image to be evaluated.
[0096] The second kernel selective convolution is used to compress the features of each of the fused fog density sensing feature maps to obtain the first compressed fog density sensing feature map corresponding to each of the dehazing images to be evaluated.
[0097] The 1×1 convolutional layer is used to compress the features of each of the first compressed fog density sensing feature maps to obtain the second compressed fog density sensing feature maps corresponding to each of the dehazing images to be evaluated.
[0098] The Sigmoid activation layer is used to smooth each of the second compressed fog density sensing feature maps to obtain image fog density maps corresponding to each of the dehazing images to be evaluated.
[0099] The structural similarity index is calculated for each of the original foggy images and the preprocessed defogging images corresponding to each of the defogging images to be evaluated, and the calculation results are used as pseudo-labels to obtain pseudo-labels corresponding to each of the defogging images to be evaluated.
[0100] The mean squared error loss function algorithm is used to calculate the loss function for all the image fog density maps and all the pseudo-labels to obtain the seventh loss function;
[0101] Pearson linear correlation coefficients are calculated for all the image fog density maps and all the pseudo-labels, and the calculation results are used as the eighth loss function;
[0102] Spearman ranking order correlation coefficients are calculated for all the image fog density maps and all the pseudo-labels, and the calculation results are used as the ninth loss function;
[0103] The tenth loss function is obtained by calculating the seventh, eighth, and ninth loss functions using the third equation. The third equation is:
[0104] ,
[0105] in, The tenth loss function, The seventh loss function, This is the third weight hyperparameter. This is the eighth loss function. This is the fourth weight hyperparameter. This is the ninth loss function;
[0106] The image fog density sensing network is updated with parameters using the tenth loss function to obtain a fog density predictor.
[0107] It should be understood that this module (i.e., the image fog density perception network) aims to achieve quantitative modeling of fog density information in images, and guides the training process of the fog density prediction network by introducing a pseudo-label supervision mechanism.
[0108] Specifically, it includes the following steps:
[0109] Using the foggy image (i.e. the original foggy image) as input, a multi-layer large kernel selective convolution (i.e. the first large kernel selective convolution) is introduced to increase the receptive field. Multi-scale context modeling is performed on the input image (i.e. the original foggy image) to achieve the purpose of fog density feature extraction and mapping.
[0110] The extracted fog density features (i.e., the original fog density sensing feature map) are processed using local statistical calculations, namely, max pooling and average pooling operations are used to process the image feature map (i.e., the original fog density sensing feature map), and the outputs of the two are cascaded and fused. Then, feature compression is performed through large kernel selective convolution (i.e., second largest kernel selective convolution) and 1×1 convolutional layers.
[0111] A sigmoid activation layer is introduced to smooth the feature map after feature compression (i.e., the second compressed fog density sensing feature map) and output the image fog density map.
[0112] The structural similarity index (SSIM) score between the input foggy image (i.e. the original foggy image) and its paired clear image (i.e. the pre-processed defogging image) is used as a training pseudo-label to represent the degree of fog density in the image.
[0113] The training objective is to minimize the difference between the SSIM pseudo-label and the global average of the fog density map output by the fog density prediction network. In addition to the MSE loss for model optimization, two differentiable list loss functions are introduced: the Pearson linear correlation coefficient (PLCC) loss and the Spearman rank-order correlation coefficient (SRCC) loss, to enhance the model's predictive linearity and monotonicity, as shown below:
[0114] ,
[0115] in, and is a weighting hyperparameter used to balance the contributions among the three losses.
[0116] It should be understood that, through the pre-training of the module (i.e., the image fog density sensing network), the obtained image fog density predictor (i.e., fog density predictor) can predict the fog density of any input image, supporting the quantitative analysis of fog residue in subsequent image quality assessment tasks.
[0117] In the above embodiments, the image fog density perception network analyzes the original foggy image and the pre-processed defogging image to obtain a fog density predictor and the original fog density perception feature map. This enables fog density prediction for any input image, supports the quantitative analysis of fog residue in subsequent image quality assessment tasks, and achieves collaborative perception of image content, distortion, and fog density information. This improves the accuracy and subjective consistency of image quality assessment and has strong practical value and promotion potential.
[0118] Optionally, as an embodiment of the present invention, the original distortion perception feature set includes a first original distortion perception feature map, a second original distortion perception feature map, a third original distortion perception feature map, and a fourth original distortion perception feature map.
[0119] The process of fusing and analyzing each of the original content-aware feature maps, the original distortion-aware feature sets corresponding to each of the dehazing images to be evaluated, and the original fog density-aware feature maps corresponding to each of the dehazing images to be evaluated through the target content-aware encoder, the target distortion-aware encoder, and the fog density predictor to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated includes:
[0120] The original content-aware feature maps are calculated using the fourth equation to obtain the original content-aware feature vectors corresponding to each of the dehazing images to be evaluated. The fourth equation is as follows:
[0121] ,
[0122] in, This is the original content-aware feature vector. This is the transformation processing function. This is the original content-aware feature map;
[0123] The fifth equation is used to calculate the first original distortion perception feature map, the second original distortion perception feature map corresponding to each of the dehazing images to be evaluated, the third original distortion perception feature map corresponding to each of the dehazing images to be evaluated, and the fourth original distortion perception feature map corresponding to each of the dehazing images to be evaluated, respectively, to obtain the distortion perception feature map to be processed corresponding to each of the dehazing images to be evaluated. The fifth equation is:
[0124] ,
[0125] in, The image shows the distorted sensory feature map to be processed. This is the first original distortion perception feature map. For upsampling processing, This is the second original distortion perception feature map. For splicing processing, This is the third original distortion perception feature map. This is the fourth original distortion perception feature map;
[0126] The sixth equation is used to calculate the target distortion perception feature vector corresponding to each of the first original distortion perception feature maps, the second original distortion perception feature maps corresponding to each of the dehazing images to be evaluated, the third original distortion perception feature maps corresponding to each of the dehazing images to be evaluated, and the fourth original distortion perception feature maps corresponding to each of the dehazing images to be evaluated, respectively, to obtain the target distortion perception feature vector corresponding to each of the dehazing images to be evaluated. The sixth equation is:
[0127] ,
[0128] in, ,
[0129] in, The feature vector for perceiving target distortion. For splicing processing, This is the first original distortion-perceived feature vector. This is the second original distortion-perceived feature vector. This is the third original distortion-perceived feature vector. This is the fourth original distortion-perceived feature vector. This refers to the first, second, third, or fourth original distortion-perceived feature vector. For spatial pyramid pooling blocks, It is either the first original distortion perception feature map, the second original distortion perception feature map, the third original distortion perception feature map, or the fourth original distortion perception feature map.
[0130] The seventh equation is used to calculate the target fog density sensing feature vector corresponding to each of the original fog density sensing feature maps, thereby obtaining the target fog density sensing feature vector corresponding to each of the dehazing images to be evaluated. The seventh equation is:
[0131] ,
[0132] in, The target fog density sensing feature vector. This is the transformation processing function. This is the original fog density sensing feature map;
[0133] The target fog density sensing feature map corresponding to each of the original fog density sensing feature maps is obtained by calculating the eighth equation:
[0134] ,
[0135] in, For target fog density sensing feature map, For splicing processing, This is the average processing function. For the maximum processing function, This is the original fog density sensing feature map;
[0136] The original content-aware feature maps, the distortion-aware feature maps corresponding to the dehazing images to be evaluated, and the target fog density-aware feature maps corresponding to the dehazing images to be evaluated are respectively stitched together to obtain the stitched-together-aware-feature maps corresponding to the dehazing images to be evaluated.
[0137] Global feature extraction is performed on each of the stitched perceptual feature maps to obtain global feature maps corresponding to each of the dehazing images to be evaluated.
[0138] Region feature extraction is performed on each of the stitched sensory feature maps to obtain the region feature maps corresponding to each of the dehazing images to be evaluated.
[0139] Local feature extraction is performed on each of the stitched sensory feature maps to obtain local feature maps corresponding to each of the dehazing images to be evaluated.
[0140] The original fused feature map corresponding to each of the stitched perceptual feature maps, the global feature map corresponding to each of the dehazed images to be evaluated, the region feature map corresponding to each of the dehazed images to be evaluated, and the local feature map corresponding to each of the dehazed images to be evaluated are calculated using the ninth formula to obtain the original fused feature map corresponding to each of the dehazed images to be evaluated. The ninth formula is as follows:
[0141] ,
[0142] in, This is the original fused feature map. It is the Sigmoid activation function. For global feature maps, For regional feature maps, For local feature maps, For element-wise addition, For element-wise multiplication, This is the spliced sensory feature map;
[0143] Max pooling is performed on each of the original fused feature maps to obtain max pooled fused feature maps corresponding to each of the dehazed images to be evaluated.
[0144] Each of the original fused feature maps is subjected to average pooling to obtain an average pooled fused feature map corresponding to each of the dehazed images to be evaluated.
[0145] The tenth equation is used to calculate the original content-aware feature vectors, the target distortion-aware feature vectors corresponding to the dehazing images to be evaluated, and the target fog density-aware feature vectors corresponding to the dehazing images to be evaluated, respectively, to obtain the original fusion feature vectors corresponding to the dehazing images to be evaluated. The tenth equation is:
[0146] ,
[0147] in, The original fused feature vector, For space expansion functions, For splicing processing, This is the original content-aware feature vector. The feature vector for perceiving target distortion. The target fog density sensing feature vector;
[0148] The max pooled fused feature maps, the average pooled fused feature maps corresponding to the dehazed images to be evaluated, and the original fused feature vectors corresponding to the dehazed images to be evaluated are concatenated to obtain the target dehazed feature maps corresponding to the dehazed images to be evaluated.
[0149] It should be understood that It includes convolutional layers, batch normalization (BN), and global average pooling (GAP), which helps map feature maps to feature vectors; It includes an average pooling layer and a convolutional layer, which helps to extract the overall average features of the feature map; It includes a max pooling layer and a convolutional layer, which helps to extract local salient features from the feature map.
[0150] It should be understood that the present invention aims to achieve robust representation of image quality by content, distortion, and fog density-aware features in complex dehazing scenarios. After three pre-trained modules (i.e., an image content-aware network, an image distortion-aware network, and an image fog density-aware network), a content-aware encoder is obtained. (i.e., target content-aware encoder), distortion-aware encoder) (i.e., target distortion sensing encoder) and fog density predictor (i.e., fog density predictor) to obtain content, distortion and fog density related features (i.e., original content-aware feature map, original distortion-aware feature set and original fog density-aware feature map).
[0151] Specifically, to enable the method to adapt to complex dehazing scenarios, three operations are defined: transformation operation, maximum operation, and average operation. Transformation operation... This includes convolutional layers, batch normalization (BN), and global average pooling (GAP), which helps map feature maps to feature vectors. Maximum operation... It includes a max-pooling layer and a convolutional layer, which helps extract salient local features from the feature map. (Averaging operation) It includes an average pooling layer and a convolutional layer, which helps to extract the overall average features of the feature map.
[0152] It should be understood that the image content-aware feature representation is composed of content-aware feature maps. (i.e., the original content-aware feature map) and content-aware feature vector (i.e., the original content-aware feature vector) is composed of the following:
[0153] ,
[0154] in, This represents the input image for the content-aware encoder.
[0155] Specifically, the image distortion perception feature representation is composed of distortion perception feature maps. (i.e., the distortion-perceived feature map to be processed) and the distortion-perceived feature vector (i.e., target distortion perception feature vector) consists of four distortion perception feature maps. (That is, the first, second, third, and fourth original distortion-perceived feature maps) are extracted from four different layers of the distortion-perceived encoder (i.e., the target distortion-perceived encoder). These feature maps (i.e., the first, second, third, and fourth original distortion-perceived feature maps) are upsampled to a consistent spatial resolution and then concatenated along the channel dimension to generate the distortion-perceived feature map. (i.e., the distortion-aware feature map to be processed). Simultaneously, each distortion-aware feature map is processed through a Spatial Pyramid Pooling (SPP) block to generate a fixed-length distortion representation. The SPP block (Spatial Pyramid Pooling Block) consists of two convolutional layers, one SPP operation, and one fully connected (FC) layer. Finally, these distortion representations are concatenated to obtain distortion-aware feature vectors. (That is, the target distortion-perceived feature vector). Its description is as follows:
[0156] ,
[0157] ,
[0158] ,
[0159] in, Indicates an upsampling operation. This represents the input image of the distortion-aware encoder. Furthermore, This indicates an SPP block.
[0160] It should be understood that the image fog density perception feature representation is derived from the fog density perception feature map. (i.e., target fog density sensing feature map) and fog density sensing feature vector (i.e., the target fog density sensing feature vector) is composed of the following:
[0161] ,
[0162] ,
[0163] in, This represents the fog density map (i.e., the original fog density perception feature map). This represents the input image for the fog density predictor.
[0164] It should be understood that, in order to further improve the accuracy of no-reference dehazed image quality assessment, this invention integrates multi-source heterogeneous perceptual features extracted by an image content-aware encoder, an image distortion-aware encoder, and an image fog density predictor, thereby simulating the overall perception process of image quality by the human visual system and achieving accurate prediction of image visual quality.
[0165] Specifically, the processing procedure is as follows:
[0166] Multi-source perceptual feature representations are fused with the input. Content-aware feature maps are then used. (i.e., original content-aware feature map), distortion-aware feature map (i.e., the distortion sensing feature map to be processed) and the fog density sensing feature map (i.e., the target fog density sensing feature map) is stitched together along the channel dimension to obtain a joint feature map. (That is, the spliced perceptual feature map), which is described as follows:
[0167] .
[0168] Joint feature map The stitched perceptual feature map is input to a multi-scale adaptive quality interaction module, which simulates the human visual system's judgment process of image quality at multiple perceptual levels: global, regional, and local. This module consists of three branches: a global branch... (Extracting overall image quality perception features), region branching (Extracting perceptual features of key regions in an image) and local branches (Extracting fine-grained texture and local structural features), the outputs of the above three branches are fused through element-wise addition to generate an intermediate fused feature, and then the intermediate fused feature is connected to the input through residual connections. (i.e., the concatenated perceptual feature maps) are multiplied together to further enhance the feature representation capability. The process is described as follows:
[0169] ,
[0170] in, This represents the Sigmoid activation function. and These represent element-wise addition and multiplication, respectively.
[0171] To reduce computational complexity and enhance the expressive power of salient regions, Perform max pooling and average pooling operations respectively to obtain two sets of feature maps. and (i.e., the fused feature map after max pooling and the fused feature map after average pooling).
[0172] In order to preserve each independent perceptual feature vector, i.e., content-aware feature vector (i.e., original content-aware feature vector), distortion-aware feature vector (i.e., target distortion perception feature vector) and fog density perception feature vector (i.e., the target fog density perception feature vector), the independent influence of these three factors on the final visual quality prediction, are concatenated along the channel dimension and expanded spatially to match the spatial size of the pooling feature map, resulting in an independent representation. (i.e., the original fused feature vector), the process is described as follows:
[0173] ,
[0174] in, This indicates a space expansion.
[0175] Will (i.e., the fused feature map after average pooling) (i.e., feature maps fused after max pooling) and The original fused feature vectors are concatenated along the channel dimension to generate the final fused feature map. (i.e., the target dehazing feature map), the process is described as follows:
[0176] ,
[0177] This feature map integrates the interactive representation of three types of information with their independent contributions, and can accurately and comprehensively reflect the overall visual quality of the image.
[0178] In the above embodiments, the target dehazing feature map is obtained by fusing and analyzing the original content-aware feature map, the original distortion-aware feature set, and the original fog density-aware feature map through the target content-aware encoder, the target distortion-aware encoder, and the fog density predictor. This can integrate the interactive expression of multiple information and their independent contributions, accurately and comprehensively reflect the overall visual quality of the image, and enhance the expressive power of the features.
[0179] Optionally, as an embodiment of the present invention, the process of evaluating and analyzing each of the target dehazing feature maps to obtain the evaluation result of image dehazing quality includes:
[0180] The eleventh equation is used to calculate the quality assessment score corresponding to each of the target dehazing feature maps, and all the quality assessment scores are used as the evaluation result of the image dehazing quality. The eleventh equation is:
[0181] ,
[0182] in, For quality assessment scores, For the target dehazing feature map, For flattening, For mapping processing.
[0183] It should be understood that, in order to achieve accurate assessment of the visual quality of dehazed images, this invention extracts spatially distributed quality-perceived information from the fused feature map (i.e., the target dehazed feature map) and combines it with a weighting mechanism to generate the final dehazed image quality prediction score (i.e., quality assessment score).
[0184] Specifically, the main structure of this invention consists of two parts: a weighted branch and a fractional branch. These branches are used to predict the influence weights of different spatial locations on the final image quality and the local quality score, respectively, thereby achieving fine-grained image quality modeling. By combining spatial sensitivity and a global weighting mechanism, it enhances the modeling ability for complex factors such as regional distortion, content differences, and fog density variations. The calculation process is as follows:
[0185] ,
[0186] in, This represents the final predicted quality score (i.e., quality assessment score) of the dehazed image. This indicates that feature maps are merged and fused along spatial dimensions. The obtained flattened feature map. and These represent the fused feature maps. Height and width. Indicates allocation to spatial location The weight, and This represents the predicted score for the corresponding location. This calculation method effectively integrates the subjective weights and objective scores of visual quality at different spatial locations, achieving an image quality evaluation mechanism that is more in line with the characteristics of the human visual system.
[0187] In the above embodiments, the image dehazing quality evaluation results are obtained by evaluating and analyzing the dehazing feature maps of each target. This effectively integrates the subjective weights and objective scores of visual quality at different spatial locations, combines spatial sensitivity characteristics and global weighting mechanisms, and improves the ability to model complex factors such as regional distortion, content differences and fog density changes. This achieves an image quality evaluation mechanism that is more in line with the characteristics of the human visual system.
[0188] Optionally, as another embodiment of the present invention, the present invention addresses the problems of inaccurate evaluation and poor robustness of current no-reference dehazing image quality evaluation models, which do not comprehensively consider the impact of image content, distortion, and fog density and their interactions on visual quality. The present invention proposes a perception-driven no-reference dehazing image quality evaluation method. The present invention achieves a comprehensive and scientific evaluation of dehazing images by accurately perceiving the three aspects of image content, distortion, and fog density through the proposed model.
[0189] Optionally, as another embodiment of the present invention, the present invention aims to solve the technical problems of insensitivity to visual perception of dehazed images, poor evaluation without reference, and lack of fog density analysis capability in existing dehazed image quality assessment. To address these problems, the present invention provides a perception-driven, reference-free dehazed image quality assessment method and a system for image content extraction, distortion perception, and fog density analysis. By introducing three sub-modules—image content perception pre-training, image distortion perception pre-training, and image fog density perception pre-training—a deep learning framework with multi-perceptual dimension feature fusion capabilities is constructed, achieving collaborative perception of image content, distortion, and fog density information, thereby improving the accuracy and subjective consistency of image quality assessment.
[0190] Alternatively, as another embodiment of the present invention, such as Figure 2As shown, the specific implementation steps of the present invention are as follows:
[0191] (1) Standardize the input dehazed image to be evaluated, including size adjustment and normalization, as a unified input for the model;
[0192] (2) Image content-aware pre-training. A content-aware encoder-decoder network is trained using clear images to extract structural semantic features of the images and optimize reconstruction capabilities;
[0193] (3) Image distortion perception pre-training. With the assistance of content features, a distortion perception encoder is trained using dehazed images to learn the distortion-related features introduced by dehazing;
[0194] (4) Image fog density perception pre-training. A fog density prediction network is constructed, fog feature maps are learned through pseudo-label supervision, and residual fog in the image is quantified;
[0195] (5) Extract content, distortion and fog density feature maps and vectors from each trained sub-module to construct a robust multidimensional perceptual representation;
[0196] (6) In the content-distortion-fog density perception feature representation self-interaction module, the three types of feature maps are fused together, and the independent influence of each of the three aspects is retained. The final fused feature map is generated by combining multi-scale branching and spatial expansion.
[0197] (7) The image quality score is obtained by predicting the local scores and their weights through a dual-branch network and calculating the weighted sum.
[0198] Alternatively, as another embodiment of the present invention, such as Figure 3 As shown, the present invention aims to extract features in images that are highly correlated with content information. Specifically, it employs a self-supervised learning mechanism to construct an encoder-decoder network structure for perceptual modeling of the content of clear (fog-free) images.
[0199] Alternatively, as another embodiment of the present invention, such as Figure 4 As shown, this invention is used to learn the potential distortion information introduced after image dehazing. Based on the pre-modeling of image content (i.e., the content-aware encoder has been trained), it extracts distortion-related representational features. Its network structure is similar to that of an image content-aware pre-training module, with the main differences lying in the input source and decoding strategy.
[0200] Alternatively, as another embodiment of the present invention, such as Figure 5 As shown, the present invention aims to achieve quantitative modeling of fog density information in images by introducing a pseudo-label supervision mechanism to guide the training process of the fog density prediction network.
[0201] Alternatively, as another embodiment of the present invention, such as Figure 6As shown, in order to further improve the accuracy of image quality assessment without reference dehazing, this invention proposes a content-distortion-fog density-aware feature representation self-interactive module, which is used to fuse multi-source heterogeneous perceptual features extracted by an image content-aware encoder, an image distortion-aware encoder, and an image fog density predictor, thereby simulating the overall perception process of human visual system on image quality and achieving accurate prediction of image visual quality.
[0202] Alternatively, as another embodiment of the present invention, such as Figure 7 As shown, to achieve accurate assessment of the visual quality of dehazed images, this invention further proposes a dual-branch quality predictor module based on spatially sensitive modeling. This module extracts spatially distributed quality-perceived information from the fused feature map and generates the final dehazed image quality prediction score using a weighted mechanism. The module's main structure consists of two parts: a weight branch and a score branch. These branches predict the influence weights of different spatial locations on the final quality and the local quality score, respectively, thus achieving fine-grained image quality modeling. This module, combining spatial sensitivity and a global weighting mechanism, enhances its ability to model complex factors such as regional distortion, content differences, and fog density variations.
[0203] Optionally, as another embodiment of the invention, the method is executed on an NVIDIA GeForce RTX 4090 equipped with PyTorch 1.11.0 and CUDA 12.6. Initially, during the pre-training phase, training is performed using a dedicated pre-training dataset, with all networks trained from scratch and all training images cropped into 256×256 image patches. RMSProp is used as an optimizer to update the network parameters, with an initial learning rate set to... In the first five iterations, a warm-up strategy is applied to gradually increase the learning rate. After the warm-up phase, a cosine annealing strategy is employed to further stabilize the training process. Optimization stops once the loss stops decreasing. Dehazed images and their score labels from the benchmark dehazed image quality assessment dataset are used to optimize the remaining parameters of the entire model, with the same configuration as used in the pre-training phase. Its loss function is the same as that of the fog density predictor.
[0204] Alternatively, as another embodiment of the present invention, the present invention has the following beneficial effects:
[0205] This invention is the first to propose using a pre-trained model to independently extract quality-perceived features for no-reference dehazing image quality assessment. It comprehensively integrates three key factors: image content preservation, distortion artifacts, and residual fog, thereby improving the accuracy of perceptual quality assessment for dehazed images. By designing a perceptual feature representation module, it achieves effective extraction and fusion of multi-scale and multi-angle features, enhancing feature representation capabilities and significantly improving assessment robustness in complex and variable dehazing scenarios. The introduction of a content-distortion-fog density feature self-interaction module and a multi-scale adaptive quality interaction mechanism optimizes the fusion and weight allocation among various factors, making quality prediction more consistent with human visual perception. Validation on a large number of publicly available datasets demonstrates that this method outperforms existing mainstream image quality assessment and dehazing quality assessment methods in terms of evaluation performance, and is highly correlated with subjective visual ratings, possessing greater practical value and potential for wider application.
[0206] Optionally, as another embodiment of the present invention, for accurately evaluating the visual quality of a dehazed image under no-reference conditions, the processing flow of the method is as follows: Figure 1 As shown, the system structure includes an image content perception module, an image distortion perception module, an image fog density perception module, a perception feature fusion module, and a dual-branch quality prediction module.
[0207] I. Image Input and Preprocessing
[0208] In this embodiment, images from a publicly available dehazing image quality assessment dataset are selected as input. The images are uniformly cropped to 256×256 and normalized to ensure the consistency of the input format and the stability of feature extraction.
[0209] II. Image Content Awareness Pre-training
[0210] A clear (haze-free) image is input into a content-aware encoder, which is trained using a self-supervised learning approach. The encoder extracts image content features. The input image is then fed into the decoder for reconstruction. The training objective is to minimize the mean square error (MSE) and perceptual error (LPIPS) between the input and reconstructed images. After training this module, the content-aware encoder parameters are frozen.
[0211] III. Image Distortion Perception Pre-training
[0212] The dehazed image is input into a distortion-aware encoder to extract distortion features. Simultaneously, content features are obtained from the trained content-aware encoder. The two inputs are combined and fed into the distortion decoder for reconstruction. Training also employs joint optimization using MSE and LPIPS loss. The parameters of the distortion-aware encoder are frozen after training.
[0213] IV. Image Fog Density Perception Pre-training
[0214] The fog image is input into the fog density predictor, and multi-scale fog density features are extracted using large kernel selective convolution. The fog density predictor is fused using max pooling and average pooling operations, and outputs a fog density map. Using a clear image as a reference, pseudo-labels are generated using the structural similarity (SSIM) metric. The training objective is to minimize the error between the predicted map and the SSIM pseudo-labels, and PLCC and SRCC losses are combined to improve the linearity and ranking consistency of the predictions. The fog density predictor is frozen after training.
[0215] V. Multi-source sensing feature extraction
[0216] The content, distortion, and fog density feature maps of the input image are extracted using a frozen content-aware encoder, a distortion-aware encoder, and a fog density predictor, respectively. , and ) and their corresponding eigenvectors ( , and Feature extraction methods include global average pooling (GAP), max pooling, and convolutional transformations.
[0217] VI. Content-Distortion-Fog Density Feature Fusion
[0218] The three types of feature maps are concatenated along the channel dimension to form a joint feature map. The input is a multi-scale adaptive quality interaction module. This module includes three branches: global, regional, and local. It simulates the multi-layered perception mechanism of the human visual system and outputs features through residual connections and fusion weights. Then, max pooling and average pooling are performed, and combined with the expanded receptive vectors to form the final fused feature map. .
[0219] VII. Image Quality Scoring
[0220] Will Flattened into a two-dimensional feature matrix, it is input into a two-branch quality prediction network. The weight branch uses sigmoid activation to output the weights at each spatial location. Fractional branch outputs local quality score The two values are weighted and summed to generate the final image quality prediction score.
[0221] VIII. Training and Optimization
[0222] This embodiment is implemented on the NVIDIA RTX 4090 platform using PyTorch 1.11.0 and the CUDA 12.6 framework, with RMSProp as the optimizer and an initial learning rate set to... The first five training rounds employed a warm-up strategy to increase the learning rate, followed by cosine annealing to decrease it. During the training phase, the content-aware encoder, distortion-aware encoder, and fog density predictor were frozen, and only the remaining parts of the model were optimized.
[0223] IX. Quality Prediction and Visualization Output
[0224] After training, this system can perform no-reference quality scoring on any dehazed image and output intermediate results such as fog density map and fused feature map to assist in quality interpretation and visualization analysis.
[0225] This embodiment fully demonstrates the adaptability and prediction accuracy of the present invention in complex no-reference dehazing image quality assessment scenarios. Experimental results show that the proposed method exhibits high consistency with subjective ratings on multiple public datasets, outperforming existing mainstream no-reference quality assessment methods.
[0226] Figure 8 This is a block diagram of an image dehazing quality assessment device provided in an embodiment of the present invention.
[0227] Alternatively, as another embodiment of the present invention, such as Figure 8 As shown, an image dehazing quality assessment device includes:
[0228] The import module is used to import multiple dehazing images to be evaluated and the original foggy images corresponding to each of the dehazing images to be evaluated.
[0229] The preprocessing module is used to preprocess each of the dehazing images to be evaluated to obtain a preprocessed dehazing image corresponding to each of the dehazing images to be evaluated.
[0230] The feature analysis module is used to build a training model. The training model is used to perform feature analysis on each of the preprocessed dehazing images and the original foggy images corresponding to each of the dehazing images to be evaluated, so as to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated.
[0231] The evaluation result acquisition module is used to evaluate and analyze the dehazing feature maps of each target respectively, and obtain the evaluation result of the image dehazing quality.
[0232] Optionally, another embodiment of the present invention provides an image dehazing quality assessment system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the image dehazing quality assessment method described above. This system can be a computer or similar system.
[0233] Optionally, another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image dehazing quality assessment method as described above.
[0234] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0235] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0236] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0237] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.
[0238] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0239] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0240] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for evaluating image dehazing quality, characterized in that, Includes the following steps: Import multiple dehazed images to be evaluated, as well as the original foggy images corresponding to each of the dehazed images to be evaluated; Each of the dehazed images to be evaluated is preprocessed to obtain a preprocessed dehazed image corresponding to each of the dehazed images to be evaluated. A training model is constructed, and feature analysis is performed on each of the preprocessed dehazing images and the original foggy images corresponding to each of the dehazing images to be evaluated, to obtain target dehazing feature maps corresponding to each of the dehazing images to be evaluated. The dehazing feature maps of each target are evaluated and analyzed to obtain the evaluation results of the image dehazing quality; The training model includes an image content-aware network, an image distortion-aware network, and an image fog density-aware network. The process of performing feature analysis on each of the preprocessed dehazed images and the original foggy images corresponding to each of the dehazed images to be evaluated using the trained model to obtain the target dehazed feature map corresponding to each of the dehazed images to be evaluated includes: The image content-aware network is used to analyze the image content of each of the preprocessed dehazed images to obtain the original content-aware feature map corresponding to each of the dehazed images to be evaluated. The image content-aware network includes an original content-aware encoder, and the parameters of the original content-aware encoder are updated by a loss function to obtain a target content-aware encoder. The image distortion perception network is used to analyze the image distortion perception of each of the preprocessed dehazed images and the original content-aware feature maps corresponding to each of the dehazed images to be evaluated, so as to obtain the original distortion perception feature set corresponding to each of the dehazed images to be evaluated. The image distortion sensing network includes an original distortion sensing encoder, and the parameters of the original distortion sensing encoder are updated by a loss function to obtain a target distortion sensing encoder. The image fog density sensing network is used to analyze the fog density of each original foggy image and the preprocessed defogging image corresponding to each defogging image to be evaluated, so as to obtain the original fog density sensing feature map corresponding to each defogging image to be evaluated. The parameters of the image fog density perception network are updated using a loss function to obtain a fog density predictor; The target content-aware encoder, the target distortion-aware encoder, and the fog density predictor respectively perform fusion analysis on each of the original content-aware feature maps, the original distortion-aware feature sets corresponding to each of the dehazing images to be evaluated, and the original fog density-aware feature maps corresponding to each of the dehazing images to be evaluated, to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated.
2. The image dehazing quality assessment method according to claim 1, characterized in that, The process of preprocessing each of the dehazed images to be evaluated to obtain a preprocessed dehazed image corresponding to each of the dehazed images to be evaluated includes: Each of the dehazing images to be evaluated is resized to obtain an adjusted dehazing image corresponding to each of the dehazing images to be evaluated. Each of the adjusted dehazed images is normalized to obtain a preprocessed dehazed image corresponding to each of the dehazed images to be evaluated.
3. The image dehazing quality assessment method according to claim 1, characterized in that, The image content-aware network includes a content-aware decoder. The process of performing image content-aware analysis on each of the preprocessed dehazed images using the image content-aware network to obtain the target content-aware encoder and the original content-aware feature map corresponding to each of the dehazed images to be evaluated includes: The original content-aware encoder is used to extract features from each of the preprocessed dehazed images to obtain the original content-aware feature map corresponding to each of the dehazed images to be evaluated. The content-aware decoder is used to reconstruct each of the original content-aware feature maps to obtain a first reconstructed clear image corresponding to each of the dehazed images to be evaluated. The first loss function is obtained by calculating the loss function on all the preprocessed dehazed images and all the first reconstructed clear images using the mean square error loss function algorithm. Perceptual image patch similarity is calculated for all the preprocessed dehazed images and all the first reconstructed clear images, and the calculation result is used as the second loss function; The third loss function is obtained by calculating the first loss function and the second loss function using the first equation. The first equation is: , in, For the third loss function, For the first loss function, The first weighted hyperparameter, This is the second loss function; The target content-aware encoder is obtained by updating the parameters of the original content-aware encoder using the third loss function.
4. The image dehazing quality assessment method according to claim 1, characterized in that, The image distortion perception network includes a distortion perception decoder. The process of analyzing image distortion perception of each preprocessed dehazed image and the original content-aware feature map corresponding to each dehazed image to be evaluated through the image distortion perception network to obtain the target distortion perception encoder and the original distortion perception feature set corresponding to each dehazed image to be evaluated includes: The original distortion-aware encoder extracts features from each of the preprocessed dehazed images to obtain multiple original distortion-aware features corresponding to each of the dehazed images to be evaluated. The multiple original distortion-aware features corresponding to each of the dehazed images to be evaluated are then combined to obtain a set of original distortion-aware features corresponding to each of the dehazed images to be evaluated. The distortion-aware decoder performs image reconstruction on each of the original distortion-aware feature sets and the original content-aware feature maps corresponding to each of the dehazing images to be evaluated, to obtain a second reconstructed clear image corresponding to each of the dehazing images to be evaluated. The mean squared error loss function algorithm is used to calculate the loss function for all the preprocessed dehazed images and all the second reconstructed clear images to obtain the fourth loss function. Perceptual image patch similarity is calculated for all the preprocessed dehazed images and all the second reconstructed clear images, and the calculation result is used as the fifth loss function; The sixth loss function is obtained by calculating the fourth and fifth loss functions using the second equation, whereby the second equation is: , in, The sixth loss function, This is the fourth loss function. This is the second weight hyperparameter. This is the fifth loss function; The original distortion-aware encoder is updated using the sixth loss function to obtain the target distortion-aware encoder.
5. The image dehazing quality assessment method according to claim 1, characterized in that, The image fog density sensing network includes a first large kernel selective convolution, a first max pooling layer, a first average pooling layer, a second large kernel selective convolution, a 1×1 convolutional layer, and a sigmoid activation layer. The process of analyzing the image fog density perception of each original foggy image and the preprocessed defogging image corresponding to each defogging image to be evaluated through the image fog density perception network to obtain a fog density predictor and the original fog density perception feature map corresponding to each defogging image to be evaluated includes: The first kernel selective convolution is used to extract features from each of the original foggy images to obtain the original fog density sensing feature map corresponding to each of the defogging images to be evaluated. The first max pooling layer is used to perform pooling processing on each of the original fog density sensing feature maps to obtain the first pooled fog density sensing feature map corresponding to each of the dehazing images to be evaluated. The first average pooling layer is used to pool each of the original fog density sensing feature maps to obtain a second pooled fog density sensing feature map corresponding to each of the defogging images to be evaluated. Each first pooled fog density sensing feature map is concatenated and fused with the second pooled fog density sensing feature map corresponding to each dehazing image to be evaluated, to obtain a fused fog density sensing feature map corresponding to each dehazing image to be evaluated. The second kernel selective convolution is used to compress the features of each of the fused fog density sensing feature maps to obtain the first compressed fog density sensing feature map corresponding to each of the dehazing images to be evaluated. The 1×1 convolutional layer is used to compress the features of each of the first compressed fog density sensing feature maps to obtain the second compressed fog density sensing feature maps corresponding to each of the dehazing images to be evaluated. The Sigmoid activation layer is used to smooth each of the second compressed fog density sensing feature maps to obtain image fog density maps corresponding to each of the dehazing images to be evaluated. The structural similarity index is calculated for each of the original foggy images and the preprocessed defogging images corresponding to each of the defogging images to be evaluated, and the calculation results are used as pseudo-labels to obtain pseudo-labels corresponding to each of the defogging images to be evaluated. The mean squared error loss function algorithm is used to calculate the loss function for all the image fog density maps and all the pseudo-labels to obtain the seventh loss function; Pearson linear correlation coefficients are calculated for all the image fog density maps and all the pseudo-labels, and the calculation results are used as the eighth loss function; Spearman ranking order correlation coefficients are calculated for all the image fog density maps and all the pseudo-labels, and the calculation results are used as the ninth loss function; The tenth loss function is obtained by calculating the seventh, eighth, and ninth loss functions using the third equation. The third equation is: , in, The tenth loss function, The seventh loss function, This is the third weight hyperparameter. This is the eighth loss function. This is the fourth weight hyperparameter. This is the ninth loss function; The image fog density sensing network is updated with parameters using the tenth loss function to obtain a fog density predictor.
6. The image dehazing quality assessment method according to claim 1, characterized in that, The original distortion perception feature set includes a first original distortion perception feature map, a second original distortion perception feature map, a third original distortion perception feature map, and a fourth original distortion perception feature map. The process of fusing and analyzing each of the original content-aware feature maps, the original distortion-aware feature sets corresponding to each of the dehazing images to be evaluated, and the original fog density-aware feature maps corresponding to each of the dehazing images to be evaluated through the target content-aware encoder, the target distortion-aware encoder, and the fog density predictor to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated includes: The original content-aware feature maps are calculated using the fourth equation to obtain the original content-aware feature vectors corresponding to each of the dehazing images to be evaluated. The fourth equation is as follows: , in, This is the original content-aware feature vector. This is the transformation processing function. This is the original content-aware feature map; The fifth equation is used to calculate the first original distortion perception feature map, the second original distortion perception feature map corresponding to each of the dehazing images to be evaluated, the third original distortion perception feature map corresponding to each of the dehazing images to be evaluated, and the fourth original distortion perception feature map corresponding to each of the dehazing images to be evaluated, respectively, to obtain the distortion perception feature map to be processed corresponding to each of the dehazing images to be evaluated. The fifth equation is: , in, The image shows the distorted sensory feature map to be processed. This is the first original distortion perception feature map. For upsampling processing, This is the second original distortion perception feature map. For splicing processing, This is the third original distortion perception feature map. This is the fourth original distortion perception feature map; The sixth equation is used to calculate the target distortion perception feature vector corresponding to each of the first original distortion perception feature maps, the second original distortion perception feature maps corresponding to each of the dehazing images to be evaluated, the third original distortion perception feature maps corresponding to each of the dehazing images to be evaluated, and the fourth original distortion perception feature maps corresponding to each of the dehazing images to be evaluated, respectively, to obtain the target distortion perception feature vector corresponding to each of the dehazing images to be evaluated. The sixth equation is: , in, , in, The feature vector for perceiving target distortion. For splicing processing, This is the first original distortion-perceived feature vector. This is the second original distortion-perceived feature vector. This is the third original distortion-perceived feature vector. This is the fourth original distortion-perceived feature vector. This refers to the first, second, third, or fourth original distortion-perceived feature vector. For spatial pyramid pooling blocks, It is either the first original distortion perception feature map, the second original distortion perception feature map, the third original distortion perception feature map, or the fourth original distortion perception feature map. The seventh equation is used to calculate the target fog density sensing feature vector corresponding to each of the original fog density sensing feature maps, thereby obtaining the target fog density sensing feature vector corresponding to each of the dehazing images to be evaluated. The seventh equation is: , in, The target fog density sensing feature vector. This is the transformation processing function. This is the original fog density sensing feature map; The target fog density sensing feature map corresponding to each of the original fog density sensing feature maps is obtained by calculating the eighth equation: , in, For target fog density sensing feature map For splicing processing, This is the average processing function. For the maximum processing function, This is the original fog density sensing feature map; The original content-aware feature maps, the distortion-aware feature maps corresponding to the dehazing images to be evaluated, and the target fog density-aware feature maps corresponding to the dehazing images to be evaluated are respectively stitched together to obtain the stitched-together-aware-feature maps corresponding to the dehazing images to be evaluated. Global feature extraction is performed on each of the stitched perceptual feature maps to obtain global feature maps corresponding to each of the dehazing images to be evaluated. Region feature extraction is performed on each of the stitched sensory feature maps to obtain the region feature maps corresponding to each of the dehazing images to be evaluated. Local feature extraction is performed on each of the stitched sensory feature maps to obtain local feature maps corresponding to each of the dehazing images to be evaluated. The original fused feature map corresponding to each of the stitched perceptual feature maps, the global feature map corresponding to each of the dehazed images to be evaluated, the region feature map corresponding to each of the dehazed images to be evaluated, and the local feature map corresponding to each of the dehazed images to be evaluated are calculated using the ninth formula to obtain the original fused feature map corresponding to each of the dehazed images to be evaluated. The ninth formula is as follows: , in, This is the original fused feature map. It is the Sigmoid activation function. For global feature maps, For regional feature maps, For local feature maps, For element-wise addition, For element-wise multiplication, This is the spliced sensory feature map; Max pooling is performed on each of the original fused feature maps to obtain max pooled fused feature maps corresponding to each of the dehazed images to be evaluated. Each of the original fused feature maps is subjected to average pooling to obtain an average pooled fused feature map corresponding to each of the dehazed images to be evaluated. The tenth equation is used to calculate the original content-aware feature vectors, the target distortion-aware feature vectors corresponding to the dehazing images to be evaluated, and the target fog density-aware feature vectors corresponding to the dehazing images to be evaluated, respectively, to obtain the original fusion feature vectors corresponding to the dehazing images to be evaluated. The tenth equation is: , in, The original fused feature vector, For space expansion functions, For splicing processing, This is the original content-aware feature vector. The feature vector for perceiving target distortion. The target fog density sensing feature vector; The max pooled fused feature maps, the average pooled fused feature maps corresponding to the dehazed images to be evaluated, and the original fused feature vectors corresponding to the dehazed images to be evaluated are concatenated to obtain the target dehazed feature maps corresponding to the dehazed images to be evaluated.
7. The image dehazing quality assessment method according to claim 1, characterized in that, The process of evaluating and analyzing each of the target dehazing feature maps to obtain the evaluation result of image dehazing quality includes: The eleventh equation is used to calculate the quality assessment score corresponding to each of the target dehazing feature maps, and all the quality assessment scores are used as the evaluation result of the image dehazing quality. The eleventh equation is: , in, For quality assessment scores, For the target dehazing feature map, For flattening, For mapping processing.
8. An image dehazing quality assessment device, characterized in that, include: The import module is used to import multiple dehazing images to be evaluated and the original foggy images corresponding to each of the dehazing images to be evaluated. The preprocessing module is used to preprocess each of the dehazing images to be evaluated to obtain a preprocessed dehazing image corresponding to each of the dehazing images to be evaluated. The feature analysis module is used to build a training model. The training model is used to perform feature analysis on each of the preprocessed dehazing images and the original foggy images corresponding to each of the dehazing images to be evaluated, so as to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated. The evaluation result acquisition module is used to evaluate and analyze each of the target dehazing feature maps to obtain the evaluation result of the image dehazing quality. The training model includes an image content-aware network, an image distortion-aware network, and an image fog density-aware network. The feature analysis module is specifically used for: The image content-aware network is used to analyze the image content of each of the preprocessed dehazed images to obtain the original content-aware feature map corresponding to each of the dehazed images to be evaluated. The image content-aware network includes an original content-aware encoder, and the parameters of the original content-aware encoder are updated by a loss function to obtain a target content-aware encoder. The image distortion perception network is used to analyze the image distortion perception of each of the preprocessed dehazed images and the original content-aware feature maps corresponding to each of the dehazed images to be evaluated, so as to obtain the original distortion perception feature set corresponding to each of the dehazed images to be evaluated. The image distortion sensing network includes an original distortion sensing encoder, and the parameters of the original distortion sensing encoder are updated by a loss function to obtain a target distortion sensing encoder. The image fog density sensing network is used to analyze the fog density of each original foggy image and the preprocessed defogging image corresponding to each defogging image to be evaluated, so as to obtain the original fog density sensing feature map corresponding to each defogging image to be evaluated. The parameters of the image fog density perception network are updated using a loss function to obtain a fog density predictor; The target content-aware encoder, the target distortion-aware encoder, and the fog density predictor respectively perform fusion analysis on each of the original content-aware feature maps, the original distortion-aware feature sets corresponding to each of the dehazing images to be evaluated, and the original fog density-aware feature maps corresponding to each of the dehazing images to be evaluated, to obtain the target dehazing feature map corresponding to each of the dehazing images to be evaluated.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the image dehazing quality assessment method as described in any one of claims 1 to 7.