Image defogging quality evaluation method and device and storage medium

By building a training model to preprocess and analyze the features of dehazed images, the shortcomings of existing technologies in quality assessment of dehazed images without reference are solved, the coordinated perception of image content, distortion and fog density is achieved, and the accuracy and consistency of image quality assessment are improved.

CN120765644AActive Publication Date: 2025-10-10WUHAN INST OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511269850.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-10-10
Estimated Expiration
2045-09-08

AI Technical Summary

Technical Problem

Existing no-reference dehazed image quality assessment methods have the disadvantages of insufficient perceptual models, limited content extraction, weak distortion recognition ability, lack of fog density quantification and poor result interpretability, making it difficult to meet the requirements of high accuracy, reliability and interpretability in practical applications.

Method used

By constructing a training model, the dehazed image to be evaluated is preprocessed, and the feature analysis of image content, distortion and fog density is performed respectively. The image content perception network, distortion perception network and fog density perception network are used to extract features, and the evaluation results of image dehazing quality are obtained through fusion analysis.

Benefits of technology

It achieves the coordinated perception of image content, distortion and fog density information, improves the accuracy and subjective consistency of image quality assessment, and has strong practical value and promotion potential.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765644A_ABST
    Figure CN120765644A_ABST
Patent Text Reader

Abstract

The invention provides an image defogging quality evaluation method and device and a storage medium, and belongs to the technical field of image evaluation, and the method comprises the steps: importing a plurality of to-be-evaluated defogged images and an original foggy day image; preprocessing each defogged image to be evaluated to obtain a preprocessed defogged image; performing feature analysis on each pre-processed defogging image and the original foggy day image through a training model to obtain a target defogging feature map; and respectively evaluating and analyzing each target defogging feature map to obtain an evaluation result of the image defogging quality. According to the method, collaborative perception of image content, distortion and fog density information is realized, the accuracy and subjective consistency of image quality evaluation are improved, and the method has very high practical value and popularization potential.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of image evaluation, and in particular to an image defogging quality evaluation method, device and storage medium. Background Art

[0002] Currently, image dehazing algorithms are widely used in scenarios such as autonomous driving, remote sensing monitoring, and video security. These algorithms primarily use physical models or deep learning methods to restore degraded images to improve clarity and visibility. These technologies have achieved remarkable results in restoring image contrast, detail, and color.

[0003] To objectively evaluate the effectiveness of dehazing algorithms, researchers have proposed a variety of image dehazing quality assessment methods, including full-reference, partial-reference, and no-reference methods. Since it is difficult to obtain a haze-free reference image in real-world scenarios, no-reference quality assessment methods are more practical because they do not require a reference image.

[0004] Existing no-reference dehazing image quality assessment methods are mostly based on image statistics, structure or depth features. However, they suffer from defects such as insufficient perception models, limited content extraction, weak distortion recognition ability, lack of fog density quantification and poor result interpretability, making it difficult to meet the requirements of high accuracy, reliability and interpretability in practical applications. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an image defogging quality assessment method, device and storage medium in response to the deficiencies of the existing technology.

[0006] The present invention solves the above technical problems with the following technical solution: a method for evaluating image defogging quality, comprising the following steps: Importing a plurality of defogging images to be evaluated and an original foggy image corresponding to each of the defogging images to be evaluated; Preprocessing each of the defogging images to be evaluated respectively to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated; Constructing a training model, and performing feature analysis on each of the preprocessed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated using the training model to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated; Each of the target defogging feature maps is evaluated and analyzed respectively to obtain an evaluation result of the image defogging quality.

[0007] Another technical solution of the present invention to solve the above technical problem is as follows: an image defogging quality assessment device, comprising: An import module, configured to import a plurality of defogging images to be evaluated and an original foggy image corresponding to each of the defogging images to be evaluated; A preprocessing module, configured to preprocess each of the defogging images to be evaluated to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated; a feature analysis module, configured to construct a training model, and perform feature analysis on each of the preprocessed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated, using the training model, to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated; The evaluation result acquisition module is used to evaluate and analyze each of the target defogging feature maps respectively to obtain an evaluation result of the image defogging quality.

[0008] Based on the above-mentioned image defogging quality assessment method, the present invention also provides an image defogging quality assessment system.

[0009] Another technical solution of the present invention to solve the above-mentioned technical problem is as follows: an image dehazing quality assessment system includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the image dehazing quality assessment method as described above is implemented.

[0010] Based on the above-mentioned image defogging quality assessment method, the present invention also provides a computer-readable storage medium.

[0011] Another technical solution of the present invention to solve the above technical problem is as follows: a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the image defogging quality assessment method as described above is implemented.

[0012] The beneficial effects of the present invention are: a preprocessed defogged image is obtained by preprocessing the defogged image to be evaluated, a target defogging feature map is obtained by feature analysis of the preprocessed defogged image and the original foggy image through a training model, and an evaluation result of the image defogging quality is obtained by evaluating and analyzing the target defogging feature map, thereby realizing the coordinated perception of image content, distortion and fog density information, improving the accuracy and subjective consistency of image quality evaluation, and having strong practical value and promotion potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a flow chart of a method for evaluating image defogging quality according to an embodiment of the present invention; Figure 2 The second flowchart of the image defogging quality assessment method provided by an embodiment of the present invention; Figure 3 A processing logic diagram of image content-aware pre-training of the image dehazing quality assessment method provided by an embodiment of the present invention; Figure 4 A processing logic diagram of image distortion perception pre-training of the image defogging quality assessment method provided by an embodiment of the present invention; Figure 5 A processing logic diagram of image fog density perception pre-training of the image defogging quality assessment method provided by an embodiment of the present invention; Figure 6 A processing logic diagram of the content-distortion-fog density perception feature representation self-interaction module of the image defogging quality assessment method provided by an embodiment of the present invention; Figure 7 A processing logic diagram of a dual-branch quality predictor of the image defogging quality assessment method provided by an embodiment of the present invention; Figure 8 This is a module block diagram of the image defogging quality assessment device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0014] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.

[0015] Figure 1 A flowchart of an image defogging quality assessment method provided by an embodiment of the present invention.

[0016] like Figure 1 As shown in FIG, a method for evaluating image defogging quality includes the following steps: S1: importing multiple defogging images to be evaluated and original foggy images corresponding to each of the defogging images to be evaluated; S2: Preprocessing each of the defogging images to be evaluated to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated; S3: constructing a training model, and performing feature analysis on each of the pre-processed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated using the training model to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated; S4: Evaluate and analyze each of the target defogging feature maps to obtain an evaluation result of the image defogging quality.

[0017] In the above embodiment, a preprocessed defogged image is obtained by preprocessing the defogged image to be evaluated, a target defogging feature map is obtained by feature analysis of the preprocessed defogged image and the original foggy image through a training model, and an evaluation result of the image defogging quality is obtained by evaluation and analysis of the target defogging feature map, thereby realizing the coordinated perception of image content, distortion and fog density information, improving the accuracy and subjective consistency of image quality evaluation, and having strong practical value and promotion potential.

[0018] Optionally, as an embodiment of the present invention, the process of preprocessing each of the defogging images to be evaluated to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated includes: Resizing each of the defogging images to be evaluated to obtain an adjusted defogging image corresponding to each of the defogging images to be evaluated; Normalization processing is performed on each of the adjusted defogging images to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated.

[0019] It should be understood that the input dehazed image to be evaluated (i.e., the dehazed image to be evaluated) is subjected to standardized preprocessing, including size adjustment, normalization, etc., and is used as a unified input of the model.

[0020] In the above embodiment, each defogging image to be evaluated is preprocessed to obtain a preprocessed defogging image, which realizes the coordinated perception of image content, distortion and fog density information, improves the accuracy and subjective consistency of image quality assessment, and has strong practical value and promotion potential.

[0021] Optionally, as an embodiment of the present invention, the training model includes an image content perception network, an image distortion perception network, and an image fog density perception network. The process of performing feature analysis on each of the pre-processed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated using the training model to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated includes: Performing image content-aware analysis on each of the preprocessed dehazed images using the image content-aware network to obtain a target content-aware encoder and an original content-aware feature map corresponding to each of the dehazed images to be evaluated; Performing image distortion perception analysis on each of the preprocessed dehazed images using the image distortion perception network to obtain a target distortion perception encoder and a set of original distortion perception features corresponding to each of the dehazed images to be evaluated; Performing image fog density perception analysis on each of the original foggy images and the pre-processed defogged images corresponding to each of the defogged images to be evaluated through the image fog density perception network to obtain a fog density predictor and an original fog density perception feature map corresponding to each of the defogged images to be evaluated; The target content perception encoder, the target distortion perception encoder, and the fog density predictor are used for fusion analysis on each original content perception feature map, an original distortion perception feature set corresponding to each to-be-evaluated defogged image, and an original fog density perception feature map corresponding to each to-be-evaluated defogged image, to obtain a target defogging feature map corresponding to each to-be-evaluated defogged image.

[0022] In the above embodiment, the model is trained to perform feature analysis on each preprocessed defogged image and original foggy image respectively to obtain a target defogging feature map, thereby achieving comprehensive and scientific evaluation of the defogged image, achieving collaborative perception of image content, distortion, and fog density information, and thereby improving the accuracy and subjective consistency of image quality evaluation.

[0023] Optionally, as an embodiment of the present application, the image content perception network comprises an original content perception encoder and a content perception decoder, The process of performing image content perception analysis on each preprocessed defogged image by the image content perception network to obtain a target content perception encoder and an original content perception feature map corresponding to each to-be-evaluated defogged image comprises: Each preprocessed defogged image is subjected to feature extraction by the original content perception encoder to obtain an original content perception feature map corresponding to each to-be-evaluated defogged image; Each original content perception feature map is subjected to image reconstruction by the content perception decoder to obtain a first reconstructed clear image corresponding to each to-be-evaluated defogged image; A mean square error loss function algorithm is used to calculate a loss function for all preprocessed defogged images and all first reconstructed clear images to obtain a first loss function; The perceptual image block similarity of all preprocessed defogged images and all first reconstructed clear images is calculated, and the calculation result is taken as a second loss function; The first loss function and the second loss function are calculated by a first formula to obtain a third loss function, the first formula being: , wherein, is the third loss function, is the first loss function, is a first weight hyperparameter, is the second loss function; The original content perception encoder is updated in parameters by the third loss function to obtain a target content perception encoder.

[0024] It should be understood that the content-aware encoder-decoder network (i.e., image content-aware network) is trained by using clear images (i.e., dehazed images after preprocessing) to extract the structural semantic features of the image and optimize the reconstruction capability.

[0025] It should be understood that this module (i.e., image content-aware network) aims to extract features in the image that are highly correlated with the content information. Specifically, it adopts a self-supervised learning mechanism to construct an encoder-decoder network structure (i.e., image content-aware network) for perceptual modeling of the content of clear (haze-free) images (i.e., dehazed images after preprocessing).

[0026] Specifically, the module (i.e., image content-aware network) includes the following steps: The input clear image (i.e., the dehazed image after preprocessing) is sent to the content-aware encoder (i.e., the original content-aware encoder), which extracts features from the image content (i.e., the dehazed image after preprocessing) to obtain content-aware features. (i.e., original content-aware feature map); Content-aware features (i.e., the original content-aware feature map) is input to a content-aware decoder (i.e., a content-aware decoder) for image reconstruction, generating a clear image that is consistent with the content structure of the input image (i.e., a first reconstructed clear image); Using the mean square error (MSE) loss function (i.e., the first loss function) measures the pixel-level error between the input image and the reconstructed image; Further introduce the loss function based on perceptual similarity (i.e., the second loss function), which is calculated based on the learned perceptual image patch similarity (LPIPS) metric, and enhances the model's ability to retain content structure and texture information; The total loss function is a weighted combination of the two, which is the following formula: , in, is a weight hyperparameter used to balance the perceptual loss and reconstruction error.

[0027] It should be understood that the obtained content-aware encoder (i.e., the target content-aware encoder) has robust image content understanding and representation capabilities, providing structural content support for subsequent modules.

[0028] In the above embodiment, the image content-aware network is used to perform image content-aware analysis on each pre-processed dehazed image to obtain a target content-aware encoder and an original content-aware feature map, which has a robust image content understanding and representation capability, and provides support for subsequent structural content.

[0029] Optionally, as an embodiment of the present invention, the image distortion perception network includes an original distortion perception encoder and a distortion perception decoder. The process of performing image distortion perception analysis on each of the preprocessed dehazed images and the original content-aware feature maps corresponding to each of the dehazed images to be evaluated through the image distortion perception network to obtain a target distortion-aware encoder and an original distortion-aware feature set corresponding to each of the dehazed images to be evaluated includes: Performing feature extraction on each of the preprocessed dehazed images using the original distortion-perceptual encoder to obtain a plurality of original distortion-perceptual feature maps corresponding to each of the dehazed images to be evaluated, and respectively aggregating the plurality of original distortion-perceptual feature maps corresponding to each of the dehazed images to be evaluated to obtain an original distortion-perceptual feature set corresponding to each of the dehazed images to be evaluated; Reconstructing each of the original distortion-perceptual feature sets and the original content-perceptual feature maps corresponding to each of the dehazed images to be evaluated by using the distortion-perceptual decoder to obtain a second reconstructed clear image corresponding to each of the dehazed images to be evaluated; Performing a loss function calculation on all of the preprocessed dehazed images and all of the second reconstructed clear images using a mean square error loss function algorithm to obtain a fourth loss function; Calculating the perceived image block similarity for all the pre-processed dehazed images and all the second reconstructed clear images, and using the calculation result as a fifth loss function; The fourth loss function and the fifth loss function are calculated by the second formula to obtain the sixth loss function, where the second formula is: , in, is the sixth loss function, is the fourth loss function, is the second weight hyperparameter, is the fifth loss function; The parameters of the original distortion-aware encoder are updated using the sixth loss function to obtain a target distortion-aware encoder.

[0030] It should be understood that, with the assistance of content features, the dehazed image (ie, the pre-processed dehazed image) is used to train the distortion-aware encoder (ie, the original distortion-aware encoder) to learn the distortion-related features introduced by dehazing.

[0031] This module (the Image Distortion-Aware Network) is designed to learn about the potential distortion introduced by image dehazing. Based on the image content modeling—the trained content-aware encoder (the original distortion-aware encoder)—it extracts distortion-related representations. Its network structure is similar to the Image Content-Aware Pretraining Module (the Image Content-Aware Network), differing primarily in the input source and decoding strategy.

[0032] Specifically, this module (i.e., image distortion perception network) includes: The input image is a dehazed image (i.e., a dehazed image after preprocessing), which is input to a distortion-aware encoder (i.e., an original distortion-aware encoder) for distortion-aware related feature extraction. In the process of reconstructing the image, the distortion-aware decoder (i.e., the distortion-aware decoder) not only accepts the distortion-aware features (i.e., the original distortion-aware feature set) from the distortion-aware decoder (i.e., the distortion-aware decoder), but also introduces the content-aware features obtained from the pre-trained content-aware encoder. (i.e., original content-aware feature maps) to improve the quality of reconstruction; The distortion-aware decoder (i.e., the distortion-aware decoder) takes the original input dehazed image (i.e., the pre-processed dehazed image) as the reconstruction target and forces the distortion encoder to focus on feature modeling of distortion information by sharing content information. The mixed loss function composed of MSE and LPIPS is also used for training, which is consistent with the image content perception pre-training module.

[0033] It should be understood that, through pre-training of the module (i.e., the image distortion-aware network), the distortion-aware encoder (i.e., the original distortion-aware encoder) is able to perceive the distortion features associated with the restoration process, thereby enhancing the system's sensitivity to image quality degradation.

[0034] In the above embodiment, the image distortion perception network is used to analyze the image distortion perception of each preprocessed dehazed image and the original content-perceived feature map, and a target distortion perception encoder and an original distortion perception feature set are obtained, which can perceive the distortion features related to the restoration process and enhance the sensitivity to image quality degradation.

[0035] Optionally, as an embodiment of the present invention, the image fog density perception network includes a first large kernel selective convolution, a first maximum pooling layer, a first average pooling layer, a second large kernel selective convolution, a 1×1 convolution layer and a Sigmoid activation layer. The process of performing image fog density perception analysis on each of the original foggy images and the pre-processed defogged images corresponding to each of the defogged images to be evaluated through the image fog density perception network to obtain a fog density predictor and an original fog density perception feature map corresponding to each of the defogged images to be evaluated includes: Performing feature extraction on each of the original foggy images through the first large kernel selective convolution to obtain an original fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing pooling processing on each of the original fog density perception feature maps through the first maximum pooling layer to obtain a first pooled fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing pooling processing on each of the original fog density perception feature maps through the first average pooling layer to obtain a second pooled fog density perception feature map corresponding to each of the defogged images to be evaluated; Cascadingly fusing each of the first pooled fog density perception feature maps with the second pooled fog density perception feature maps corresponding to each of the defogged images to be evaluated, to obtain a fused fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing feature compression on each of the fused fog density perception feature maps through selective convolution with the second large kernel, thereby obtaining a first compressed fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing feature compression on each of the first compressed fog density perception feature maps through the 1×1 convolutional layer to obtain a second compressed fog density perception feature map corresponding to each of the defogged images to be evaluated; Smoothing each of the second compressed fog density perception feature maps through the Sigmoid activation layer to obtain an image fog density map corresponding to each of the defogging images to be evaluated; Calculating a structural similarity index for each of the original foggy images and the preprocessed defogged images corresponding to each of the defogged images to be evaluated, and using the calculation results as pseudo labels, thereby obtaining pseudo labels corresponding to each of the defogged images to be evaluated; Performing loss function calculation on all the image fog density maps and all the pseudo labels using the mean square error loss function algorithm to obtain a seventh loss function; Calculating the Pearson linear correlation coefficient for all the image fog density maps and all the pseudo labels, and using the calculation result as the eighth loss function; Calculating the Spearman rank order correlation coefficient for all the image fog density maps and all the pseudo labels, and using the calculation result as the ninth loss function; The seventh loss function, the eighth loss function, and the ninth loss function are calculated by the third formula to obtain the tenth loss function, and the third formula is: , in, is the tenth loss function, is the seventh loss function, is the third weight hyperparameter, is the eighth loss function, is the fourth weight hyperparameter, is the ninth loss function; The parameters of the image fog density perception network are updated using the tenth loss function to obtain a fog density predictor.

[0036] It should be understood that this module (i.e., the image fog density perception network) aims to achieve quantitative modeling of fog density information in the image, and guide the training process of the fog density prediction network by introducing a pseudo-label supervision mechanism.

[0037] Specifically, the steps include: The foggy image (i.e., the original foggy image) is used as input. Multi-layer large-kernel selective convolution (i.e., the first large-kernel selective convolution) is introduced to increase the receptive field. Multi-scale context modeling is performed on the input image (i.e., the original foggy image) to achieve the purpose of fog density feature extraction and mapping. The extracted fog density features (i.e., the original fog density perception feature map) are processed using local statistical calculations. That is, the image feature map (i.e., the original fog density perception feature map) is processed using maximum pooling and average pooling operations. The outputs of the two are cascaded and fused. Feature compression is then performed through a large kernel selective convolution (i.e., the second large kernel selective convolution) and a 1×1 convolution layer. A Sigmoid activation layer is introduced to smooth the feature map after feature compression (i.e., the fog density perception feature map after the second compression) and output the image fog density map. The structural similarity index (SSIM) score between the input foggy image (i.e., the original foggy image) and its paired clear image (i.e., the defogged image after preprocessing) is used as a training pseudo-label to indicate the fog density of the image. The training objective is to minimize the difference between the SSIM pseudo-label and the global average of the fog density map output by the fog density prediction network. In addition to the MSE loss for model optimization, the loss function also introduces two differentiable list loss functions, namely the Pearson linear correlation coefficient (PLCC) loss and the Spearman rank order correlation coefficient (SRCC) loss, to enhance the model's prediction linearity and monotonicity, namely the following formula: , in, and is a weight hyperparameter used to balance the contribution among the three losses.

[0038] It should be understood that through the pre-training of the module (i.e., the image fog density perception network), the obtained image fog density predictor (i.e., the fog density predictor) can realize the fog density prediction of any input image, supporting the quantitative analysis of the degree of fog residue in subsequent image quality assessment tasks.

[0039] In the above embodiment, the image fog density perception network is used to perform image fog density perception analysis on the original foggy image and the pre-processed defogged image, respectively, to obtain a fog density predictor and an original fog density perception feature map, thereby realizing the fog density prediction of any input image, supporting the quantitative analysis of the degree of fog residue in subsequent image quality assessment tasks, and realizing the coordinated perception of image content, distortion and fog density information, thereby improving the accuracy and subjective consistency of image quality assessment, and having strong practical value and promotion potential.

[0040] Optionally, as an embodiment of the present invention, the original distortion perception feature set includes a first original distortion perception feature map, a second original distortion perception feature map, a third original distortion perception feature map, and a fourth original distortion perception feature map. The process of performing fusion analysis on each of the original content-aware feature maps, the original distortion-aware feature set corresponding to each of the defogged images to be evaluated, and the original fog density-aware feature map corresponding to each of the defogged images to be evaluated by the target content-aware encoder, the target distortion-aware encoder, and the fog density predictor, respectively, to obtain a target defogging feature map corresponding to each of the defogged images to be evaluated includes: The original content-aware feature maps are calculated using the fourth formula to obtain the original content-aware feature vectors corresponding to the dehazed images to be evaluated. The fourth formula is: , in, is the original content-aware feature vector, is the transformation processing function, is the original content-aware feature map; The first original distortion perception feature map, the second original distortion perception feature map corresponding to each dehazed image to be evaluated, the third original distortion perception feature map corresponding to each dehazed image to be evaluated, and the fourth original distortion perception feature map corresponding to each dehazed image to be evaluated are calculated using the fifth formula to obtain a distortion perception feature map to be processed corresponding to each dehazed image to be evaluated. The fifth formula is: , in, is the distortion perception feature map to be processed, is the first original distortion perception feature map, For upsampling processing, is the second original distortion perception feature map, For splicing processing, is the third original distortion perception feature map, is the fourth original distortion perception feature map; The target distortion perception feature vector corresponding to each dehazed image to be evaluated is obtained by respectively calculating each of the first original distortion perception feature maps, the second original distortion perception feature map corresponding to each dehazed image to be evaluated, the third original distortion perception feature map corresponding to each dehazed image to be evaluated, and the fourth original distortion perception feature map corresponding to each dehazed image to be evaluated using the sixth formula. The sixth formula is: , in, , in, is the target distortion perception feature vector, For splicing processing, is the first original distortion perception feature vector, is the second original distortion perception feature vector, is the third original distortion perception feature vector, is the fourth original distortion perception feature vector, is the first original distortion perception feature vector, the second original distortion perception feature vector, the third original distortion perception feature vector, or the fourth original distortion perception feature vector, is the spatial pyramid pooling block, is the first original distortion perception feature map, the second original distortion perception feature map, the third original distortion perception feature map, or the fourth original distortion perception feature map; The target fog density perception feature vector corresponding to each of the defogging images to be evaluated is obtained by calculating each of the original fog density perception feature maps using the seventh formula. The seventh formula is: , in, is the target fog density perception feature vector, is the transformation processing function, is the original fog density perception feature map; The target fog density perception feature map corresponding to each defogging image to be evaluated is obtained by calculating each of the original fog density perception feature maps using the eighth formula. The eighth formula is: , in, is the target fog density perception feature map, For splicing processing, is the average processing function, is the maximum processing function, is the original fog density perception feature map; respectively splicing the original content perception feature maps, the distortion perception feature maps to be processed corresponding to the defogging images to be evaluated, and the target fog density perception feature maps corresponding to the defogging images to be evaluated, to obtain spliced ​​perception feature maps corresponding to the defogging images to be evaluated; Performing global feature extraction on each of the spliced ​​perceptual feature maps to obtain a global feature map corresponding to each of the dehazed images to be evaluated; Performing regional feature extraction on each of the spliced ​​perceptual feature maps to obtain a regional feature map corresponding to each of the dehazed images to be evaluated; Performing local feature extraction on each of the spliced ​​perceptual feature maps to obtain a local feature map corresponding to each of the dehazed images to be evaluated; The original fusion feature map corresponding to each defogged image to be evaluated is obtained by respectively calculating each of the spliced ​​perceptual feature maps, the global feature map corresponding to each defogged image to be evaluated, the regional feature map corresponding to each defogged image to be evaluated, and the local feature map corresponding to each defogged image to be evaluated through the ninth formula. The ninth formula is: , in, is the original fusion feature map, is the Sigmoid activation function, is the global feature map, is the regional feature map, is the local feature map, is element-wise addition, is element-wise multiplication, is the perception feature map after splicing; Performing maximum pooling processing on each of the original fused feature maps respectively to obtain a maximum pooled fused feature map corresponding to each of the dehazed images to be evaluated; Performing average pooling processing on each of the original fused feature maps respectively to obtain an average pooled fused feature map corresponding to each of the dehazed images to be evaluated; The original fused feature vector corresponding to each defogging image to be evaluated is obtained by respectively calculating each of the original content-aware feature vectors, the target distortion-aware feature vector corresponding to each defogging image to be evaluated, and the target fog density-aware feature vector corresponding to each defogging image to be evaluated through the tenth formula. The tenth formula is: , in, is the original fusion feature vector, is the spatial extension function, For splicing processing, is the original content-aware feature vector, is the target distortion perception feature vector, is the target fog density perception feature vector; Each of the maximum pooled fused feature maps, the average pooled fused feature maps corresponding to each of the defogged images to be evaluated, and the original fused feature vectors corresponding to each of the defogged images to be evaluated are spliced ​​to obtain a target defogging feature map corresponding to each of the defogged images to be evaluated.

[0041] It should be understood that It includes convolutional layers, batch normalization (BN), and global average pooling (GAP), which helps map feature maps to feature vectors; It includes an average pooling layer and a convolution layer, which helps to extract the overall average features of the feature map; It includes a max pooling layer and a convolutional layer, which helps to extract local salient features of the feature map.

[0042] It should be understood that the present invention aims to achieve a robust representation of image quality based on content, distortion and fog density perception features in complex defogging scenes. After three pre-training modules (i.e., image content-aware network, image distortion-aware network and image fog density-aware network), a content-aware encoder is obtained. (i.e. target content-aware encoder), distortion-aware encoder (i.e., object distortion-aware encoder) and fog density predictor (i.e., fog density predictor) to obtain content-, distortion-, and fog density-related features (i.e., original content-aware feature map, original distortion-aware feature set, and original fog density-aware feature map), respectively.

[0043] Specifically, in order to adapt the method to complex dehazing scenarios, three operations are defined: transformation operation, maximum operation and average operation. : It includes convolutional layers, batch normalization (BN) and global average pooling (GAP), which helps map feature maps to feature vectors. Maximum operation : It includes a maximum pooling layer and a convolution layer, which helps to extract local salient features of the feature map. Average operation : It includes an average pooling layer and a convolutional layer, which helps to extract the overall average features of the feature map.

[0044] It should be understood that the image content-aware feature representation is composed of the content-aware feature map (i.e., original content-aware feature map) and content-aware feature vector (i.e., the original content-aware feature vector), which is described as follows: , in, Represents the input image to the content-aware encoder.

[0045] Specifically, the image distortion perception feature representation is composed of the distortion perception feature map (i.e., the distortion-perceptual feature map to be processed) and the distortion-perceptual feature vector (i.e. target distortion perception feature vector), four distortion perception feature maps (i.e., the first original distortion-perceptual feature map, the second original distortion-perceptual feature map, the third original distortion-perceptual feature map, and the fourth original distortion-perceptual feature map) are extracted from four different layers of the distortion-perceptual encoder (i.e., the target distortion-perceptual encoder). These feature maps (i.e., the first original distortion-perceptual feature map, the second original distortion-perceptual feature map, the third original distortion-perceptual feature map, and the fourth original distortion-perceptual feature map) are upsampled to a consistent spatial resolution and then concatenated along the channel dimension to generate a distortion-perceptual feature map (i.e., the distortion-aware feature map to be processed). At the same time, each distortion-aware feature map passes through the spatial pyramid pooling (SPP) block to generate a fixed-length distortion representation The SPP block (i.e., spatial pyramid pooling block) consists of two convolutional layers, an SPP operation, and a fully connected (FC) layer. Finally, these distortion representations are concatenated to obtain the distortion-aware feature vector (i.e., target distortion perception feature vector). It is described as follows: , , , in, represents the upsampling operation, represents the input image of the distortion-aware encoder, and furthermore, Indicates an SPP block.

[0046] It should be understood that the image fog density perception feature representation is composed of the fog density perception feature map (i.e., target fog density perception feature map) and fog density perception feature vector (i.e., target fog density perception feature vector), which is described as follows: , , in, represents the fog density map (i.e., the original fog density perception feature map), Represents the input image to the fog density predictor.

[0047] It should be understood that in order to further improve the accuracy of reference-free dehazed image quality assessment, the present invention simulates the overall perception process of the human visual system on image quality by fusing multi-source heterogeneous perception features extracted by the image content perception encoder, the image distortion perception encoder and the image fog density predictor, thereby achieving accurate prediction of image visual quality.

[0048] Specifically, the specific processing flow is as follows: Multi-source perceptual feature representation fusion input. (i.e. original content-aware feature map), distortion-aware feature map (i.e., distortion perception feature map to be processed) and fog density perception feature map (i.e., the target fog density perception feature map) is spliced ​​in the channel dimension to obtain the joint feature map (i.e., the concatenated perceptual feature map), which is expressed as follows: .

[0049] Joint feature map (i.e., the spliced ​​perceptual feature map) is input into the multi-scale adaptive quality interaction module, which is used to simulate the human visual system's judgment process of image quality at multiple perceptual levels: global, regional, and local. This module consists of three branches: global branch (Extracting overall image quality perception features), regional branch (extracting perceptual features of key image regions) and local branches (Extract fine-grained texture and local structural features), the outputs of the above three branches are fused by element-wise addition to generate an intermediate fusion feature, and the intermediate fusion feature is connected to the input through residual connection (i.e., the concatenated perceptual feature map) to further enhance the feature expression capability. The process is described as follows: , in, represents the Sigmoid activation function, and Represent element-wise addition and multiplication, respectively.

[0050] In order to reduce the computational complexity and enhance the expression ability of salient regions, Perform maximum pooling and average pooling operations respectively to obtain two sets of feature maps ( and ) (i.e., the fused feature map after maximum pooling and the fused feature map after average pooling).

[0051] In order to preserve each independent perceptual feature vector, that is, the content-aware feature vector (i.e. original content-aware feature vector), distortion-aware feature vector (i.e., target distortion perception feature vector) and fog density perception feature vector (i.e., target fog density perception feature vector), has an independent impact on the final visual quality prediction. The three are spliced ​​in the channel dimension and expanded along the space to be consistent with the spatial size of the pooled feature map to obtain an independent representation (i.e. the original fused feature vector), the process is described as follows: , in, Indicates spatial expansion.

[0052] Will (i.e., fusion feature map after average pooling), (i.e., fusion feature map after maximum pooling) and (i.e. the original fusion feature vector) is spliced ​​in the channel dimension to generate the final fusion feature map (i.e., target defogging feature map), the process is described as follows: , This feature map simultaneously integrates the interactive expression of the three types of information and their independent contributions, and can accurately and comprehensively reflect the comprehensive visual quality of the image.

[0053] In the above embodiment, the target dehazing feature map is obtained by performing a fusion analysis on the original content-aware feature map, the original distortion-aware feature set, and the original fog density-aware feature map through a target content-aware encoder, a target distortion-aware encoder, and a fog density predictor. This can fuse the interactive expression of multiple information and their independent contributions, accurately and comprehensively reflect the comprehensive visual quality of the image, and enhance the expressive ability of the features.

[0054] Optionally, as an embodiment of the present invention, the process of evaluating and analyzing each of the target defogging feature maps to obtain an evaluation result of the image defogging quality includes: The target defogging feature maps are calculated respectively by the eleventh formula to obtain quality assessment scores corresponding to the defogging images to be evaluated, and all the quality assessment scores are used as the evaluation results of the image defogging quality. The eleventh formula is: , in, is the quality assessment score, is the target dehazing feature map, For flattening, Processing for mapping.

[0055] It should be understood that in order to achieve accurate evaluation of the visual quality of the dehazed image, the present invention extracts spatially distributed quality perception information from the fused feature map (i.e., the target dehazed feature map) and combines it with a weighting mechanism to generate the final dehazed image quality prediction score (i.e., the quality assessment score).

[0056] Specifically, the main structure of this invention consists of two parts: a weight branch and a score branch, which are used to predict the weight of the impact of different spatial locations on the final quality and the local quality score, respectively, thereby achieving fine-grained image quality modeling. By combining spatial sensitivity and a global weighting mechanism, the modeling ability of complex factors such as regional distortion, content differences, and fog density changes is improved. The calculation process is as follows: , in, Represents the final predicted quality score (i.e., quality assessment score) of the dehazed image. Represents the fusion feature map by merging along the spatial dimension The obtained flattened feature map. and Represents the fusion feature map The height and width of the Indicates the assignment to a spatial location The weight of Represents the predicted score at the corresponding position. This calculation method effectively integrates the subjective weight and objective score of visual quality at different spatial positions, realizing an image quality evaluation mechanism that is more consistent with the characteristics of the human visual system.

[0057] In the above embodiment, the dehazing feature maps of each target are evaluated and analyzed respectively to obtain the evaluation results of the image dehazing quality, which effectively integrates the subjective weights and objective scores of visual quality at different spatial positions, combines spatial sensitivity characteristics and a global weighting mechanism, and improves the modeling capabilities of complex factors such as regional distortion, content differences, and fog density changes, thereby realizing an image quality evaluation mechanism that is more in line with the characteristics of the human visual system.

[0058] Optionally, as another embodiment of the present invention, the present invention proposes a perception-driven no-reference dehazed image quality assessment method to address the problems that the current no-reference dehazed image quality evaluation model does not comprehensively consider the impact of image content, distortion and fog density and the interaction between them on visual quality, resulting in inaccurate evaluation of dehazed images and poor robustness. The present invention realizes a comprehensive and scientific evaluation of dehazed images through the proposed model that accurately perceives the three aspects of image content, distortion and fog density.

[0059] Optionally, as another embodiment of the present invention, the present invention aims to solve the technical problems in the existing dehazed image quality assessment, such as insensitivity to the visual perception of the dehazed image, poor no-reference assessment, and lack of fog density analysis capability. In order to solve the above problems, the present invention provides a perception-driven no-reference dehazed image quality assessment method and a system for image content extraction, distortion perception, and fog density analysis. By introducing three sub-modules, namely, image content perception pre-training, image distortion perception pre-training, and image fog density perception pre-training, a deep learning framework with multi-perception dimension feature fusion capability is constructed, which realizes the coordinated perception of image content, distortion, and fog density information, thereby improving the accuracy and subjective consistency of image quality assessment.

[0060] Alternatively, as another embodiment of the present invention, Figure 2 As shown, the specific implementation steps of the present invention are as follows: (1) Perform standard preprocessing on the input dehazed image to be evaluated, including size adjustment and normalization, as the unified input of the model; (2) Image content-aware pre-training: Training a content-aware encoder-decoder network with clear images to extract the structural semantic features of the image and optimize the reconstruction capability; (3) Image distortion-aware pre-training. With the help of content features, a distortion-aware encoder is trained using dehazed images to learn the distortion-related features introduced by dehazing. (4) Image fog density perception pre-training. Construct a fog density prediction network, learn the fog feature map through pseudo-label supervision, and quantify the residual fog in the image; (5) Extracting content, distortion, and fog density feature maps and vectors from each trained submodule to construct a robust multi-dimensional perceptual representation; (6) The three types of feature maps are fused in the content-distortion-fog density perception feature representation self-interaction module, and the independent influence of the three aspects is retained. The final fused feature map is generated by combining multi-scale branches and spatial expansion; (7) The local scores and their weights are predicted separately through a dual-branch network, and the weighted sum is calculated to obtain the image quality score.

[0061] Alternatively, as another embodiment of the present invention, Figure 3As shown, the present invention aims to extract features in an image that are highly correlated with content information. Specifically, a self-supervised learning mechanism is used to construct an encoder-decoder network structure for perceptual modeling of the content of a clear (fog-free) image.

[0062] Alternatively, as another embodiment of the present invention, Figure 4 As shown, this method learns the potential distortion introduced by image dehazing. Based on the image content modeling (i.e., a trained content-aware encoder), it extracts distortion-related features. Its network structure is similar to the content-aware pre-trained image module, with the main differences being the input source and decoding strategy.

[0063] Alternatively, as another embodiment of the present invention, Figure 5 As shown, the present invention aims to achieve quantitative modeling of fog density information in images and guide the training process of the fog density prediction network by introducing a pseudo-label supervision mechanism.

[0064] Alternatively, as another embodiment of the present invention, Figure 6 As shown, in order to further improve the accuracy of reference-free dehazed image quality assessment, the present invention proposes a content-distortion-fog density perception feature representation self-interaction module, which is used to fuse multi-source heterogeneous perception features extracted by the image content perception encoder, the image distortion perception encoder and the image fog density predictor, thereby simulating the overall perception process of the human visual system on image quality and realizing accurate prediction of image visual quality.

[0065] Alternatively, as another embodiment of the present invention, Figure 7 As shown, to accurately assess the visual quality of dehazed images, the present invention further proposes a dual-branch quality predictor module based on spatially sensitive modeling. This module extracts spatially distributed quality perception information from the fused feature map and combines it with a weighting mechanism to generate a final dehazed image quality prediction score. This module's main structure consists of two parts: a weight branch and a score branch, which are used to predict the weights of the impact of different spatial locations on the final quality and the local quality score, respectively, thereby achieving fine-grained image quality modeling. This module combines spatial sensitivity with a global weighting mechanism to enhance its ability to model complex factors such as regional distortion, content differences, and changes in fog density.

[0066] Alternatively, as another embodiment of the present invention, the method of the present invention is executed on an NVIDIA GeForce RTX 4090 equipped with PyTorch 1.11.0 and CUDA 12.6. Initially, in the pre-training phase, a dedicated pre-training dataset is used for training, all networks are trained from scratch, and all training images are cropped into 256×256 image blocks. RMSProp is used as an optimizer to update the network parameters, and the initial learning rate is set to . In the first 5 iterations, a warm-up strategy is applied to gradually increase the learning rate. After the warm-up phase, a cosine annealing strategy is adopted to further stabilize the training process. Once the loss stops decreasing, the optimization stops. The dehazed images and their score labels in the benchmark dehazed image quality assessment dataset are used to optimize the remaining parameters of the entire model with the same configuration as used in the pre-training phase. Its loss function is the same as that of the fog density predictor.

[0067] Optionally, as another embodiment of the present invention, the present invention has the following beneficial effects: This invention proposes for the first time the use of a pre-trained model to independently extract quality perception features for the task of no-reference dehazing image quality assessment, comprehensively integrating the three key factors of image content preservation, distortion artifacts, and residual fog, thereby improving the accuracy of perceptual quality assessment of dehazed images. By designing a perceptual feature representation module, the effective extraction and fusion of multi-scale and multi-angle features are achieved, the feature expression capability is enhanced, and the robustness of the assessment in complex and changeable dehazing scenarios is significantly improved. The introduction of a content-distortion-fog density feature self-interaction module and a multi-scale adaptive quality interaction mechanism optimizes the fusion and weight distribution between various factors, making the quality prediction more consistent with the laws of human visual perception. Verification by a large number of public datasets shows that this method outperforms existing mainstream image quality assessment and dehazing quality assessment methods in terms of assessment performance, and is highly correlated with subjective visual scores, possessing greater practical value and promotion potential.

[0068] Optionally, as another embodiment of the present invention, for accurately evaluating the visual quality of a defogging image without a reference, the processing flow of the method is as follows: Figure 1 As shown in the figure, the system structure includes an image content perception module, an image distortion perception module, an image fog density perception module, a perception feature fusion module and a dual-branch quality prediction module.

[0069] 1. Image input and preprocessing In this embodiment, images from a public dehazing image quality assessment dataset are selected as input. The images are uniformly cropped to 256×256 and normalized to ensure consistency of the input format and stability of feature extraction.

[0070] 2. Image Content-Aware Pre-training The clear (haze-free) image is fed into a content-aware encoder, which is trained using self-supervised learning. The encoder extracts image content features. , after being input into the decoder, reconstruction is performed. The training goal is to minimize the pixel mean square error (MSE) and perceptual error (LPIPS) between the input image and the reconstructed image. After the training of this module is completed, the content-aware encoder parameters are frozen.

[0071] 3. Image Distortion Perception Pre-training Input the dehazed image into the distortion-aware encoder to extract distortion features At the same time, content features are obtained from the trained content-aware encoder The two are combined and fed into the distortion decoder for reconstruction. Training also uses a joint optimization of MSE and LPIPS losses. The parameters of the distortion-aware encoder are also frozen after training.

[0072] IV. Image Fog Density Perception Pre-training The fog image is input into the fog density predictor, and multi-scale fog density features are extracted using large kernel selective convolution. Max pooling and average pooling are used for fusion, and the fog density map is output. Using the clear image as a reference, pseudo labels are generated using the structural similarity (SSIM) metric. The training objective is to minimize the error between the predicted map and the SSIM pseudo labels. PLCC and SRCC losses are combined to improve the linearity and ranking consistency of the prediction. The fog density predictor is frozen after training.

[0073] 5. Multi-source perception feature extraction The frozen content-aware encoder, distortion-aware encoder, and fog density predictor are used to extract content, distortion, and fog density feature maps of the input image, respectively ( 、 and ) and its corresponding eigenvector ( 、 and ). Feature extraction methods include global average pooling (GAP), maximum pooling and convolution transformation.

[0074] 6. Content-Distortion-Fog Density Feature Fusion Concatenate the three types of feature maps in the channel dimension to form a joint feature map , input the multi-scale adaptive quality interaction module. This module consists of three branches: global, regional and local. It simulates the multi-level perception mechanism of the human visual system and outputs features through residual connections and fusion weights. . Then perform maximum pooling and average pooling, and combine the expanded perception vector to form the final fusion feature map .

[0075] 7. Image Quality Rating Will Flattened into a two-dimensional feature matrix and input to the dual-branch quality prediction network. The weight branch uses Sigmoid activation to output the weight of each spatial position , the score branch outputs a local quality score , and the weighted sum of the two generates the final image quality prediction score.

[0076] 8. Training and Optimization This example is implemented on the NVIDIA RTX 4090 platform using PyTorch 1.11.0 and CUDA 12.6 framework, with the optimizer being RMSProp and the initial learning rate being set to The warm-up strategy was used to increase the learning rate for the first five rounds of training, and then the cosine annealing method was used to decrease the learning rate. During the training phase, the content-aware encoder, distortion-aware encoder, and fog density predictor were frozen, and only the rest of the model was optimized.

[0077] 9. Quality Prediction and Visualization Output After training, the system can perform reference-free quality scoring on any dehazed image, and output intermediate results such as fog density maps and fused feature maps to assist in quality interpretation and visual analysis.

[0078] This example fully demonstrates the adaptability and prediction accuracy of the proposed method in complex no-reference dehazing image quality assessment scenarios. Experimental results show that the proposed method is highly consistent with subjective scores on multiple public datasets and outperforms existing mainstream no-reference quality assessment methods.

[0079] Figure 8 This is a module block diagram of an image defogging quality assessment device provided by an embodiment of the present invention.

[0080] Alternatively, as another embodiment of the present invention, Figure 8 As shown, an image defogging quality assessment device includes: An import module, configured to import a plurality of defogging images to be evaluated and an original foggy image corresponding to each of the defogging images to be evaluated; A preprocessing module, configured to preprocess each of the defogging images to be evaluated to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated; a feature analysis module, configured to construct a training model, and perform feature analysis on each of the preprocessed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated, using the training model, to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated; The evaluation result acquisition module is used to evaluate and analyze each of the target defogging feature maps respectively to obtain an evaluation result of the image defogging quality.

[0081] Alternatively, another embodiment of the present invention provides an image dehazing quality assessment system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the image dehazing quality assessment method described above is implemented. The system may be a computer or other system.

[0082] Optionally, another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image defogging quality assessment method as described above is implemented.

[0083] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0084] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0085] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.

[0086] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of the embodiments of the present invention.

[0087] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0088] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for evaluating image defogging quality, characterized in that: The steps include: Importing a plurality of defogging images to be evaluated and an original foggy image corresponding to each of the defogging images to be evaluated; Preprocessing each of the defogging images to be evaluated respectively to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated; Constructing a training model, and performing feature analysis on each of the preprocessed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated using the training model to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated; Each of the target defogging feature maps is evaluated and analyzed respectively to obtain an evaluation result of the image defogging quality.

2. The image defogging quality assessment method according to claim 1, characterized in that: The process of preprocessing each of the defogging images to be evaluated to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated comprises: Resizing each of the defogging images to be evaluated to obtain an adjusted defogging image corresponding to each of the defogging images to be evaluated; Normalization processing is performed on each of the adjusted defogging images to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated.

3. The image defogging quality assessment method according to claim 1, characterized in that: The training model includes an image content perception network, an image distortion perception network, and an image fog density perception network. The process of performing feature analysis on each of the pre-processed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated using the training model to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated includes: Performing image content-aware analysis on each of the preprocessed dehazed images using the image content-aware network to obtain a target content-aware encoder and an original content-aware feature map corresponding to each of the dehazed images to be evaluated; Performing image distortion perception analysis on each of the preprocessed dehazed images and the original content-aware feature maps corresponding to each of the dehazed images to be evaluated through the image distortion perception network to obtain a target distortion perception encoder and an original distortion perception feature set corresponding to each of the dehazed images to be evaluated; Performing image fog density perception analysis on each of the original foggy images and the pre-processed defogged images corresponding to each of the defogged images to be evaluated through the image fog density perception network to obtain a fog density predictor and an original fog density perception feature map corresponding to each of the defogged images to be evaluated; The target content-aware encoder, the target distortion-aware encoder, and the fog density predictor are used to perform fusion analysis on each of the original content-aware feature maps, the original distortion-aware feature set corresponding to each of the defogged images to be evaluated, and the original fog density-aware feature map corresponding to each of the defogged images to be evaluated, respectively, to obtain a target defogging feature map corresponding to each of the defogged images to be evaluated.

4. The image defogging quality assessment method according to claim 3, characterized in that: The image content-aware network includes an original content-aware encoder and a content-aware decoder. The process of performing image content-aware analysis on each of the pre-processed dehazed images by the image content-aware network to obtain a target content-aware encoder and an original content-aware feature map corresponding to each of the dehazed images to be evaluated includes: Performing feature extraction on each of the preprocessed dehazed images using the original content-aware encoder to obtain an original content-aware feature map corresponding to each of the dehazed images to be evaluated; Reconstructing each of the original content-aware feature maps using the content-aware decoder to obtain a first reconstructed clear image corresponding to each of the dehazed images to be evaluated; Performing loss function calculation on all the pre-processed dehazed images and all the first reconstructed clear images using a mean square error loss function algorithm to obtain a first loss function; Calculating the perceived image block similarity for all the pre-processed dehazed images and all the first reconstructed clear images, and using the calculation result as a second loss function; The first loss function and the second loss function are calculated by the first formula to obtain a third loss function, where the first formula is: , in, is the third loss function, is the first loss function, is the first weight hyperparameter, is the second loss function; Parameters of the original content-aware encoder are updated using the third loss function to obtain a target content-aware encoder.

5. The image defogging quality assessment method according to claim 3, characterized in that: The image distortion perception network includes an original distortion perception encoder and a distortion perception decoder, The process of performing image distortion perception analysis on each of the preprocessed dehazed images and the original content-aware feature maps corresponding to each of the dehazed images to be evaluated through the image distortion perception network to obtain a target distortion-aware encoder and an original distortion-aware feature set corresponding to each of the dehazed images to be evaluated includes: Performing feature extraction on each of the preprocessed dehazed images using the original distortion-perceptual encoder to obtain a plurality of original distortion-perceptual features corresponding to each of the dehazed images to be evaluated, and respectively aggregating the plurality of original distortion-perceptual features corresponding to each of the dehazed images to be evaluated to obtain an original distortion-perceptual feature set corresponding to each of the dehazed images to be evaluated; Reconstructing each of the original distortion-perceptual feature sets and the original content-perceptual feature maps corresponding to each of the dehazed images to be evaluated by using the distortion-perceptual decoder to obtain a second reconstructed clear image corresponding to each of the dehazed images to be evaluated; Performing a loss function calculation on all of the preprocessed dehazed images and all of the second reconstructed clear images using a mean square error loss function algorithm to obtain a fourth loss function; Calculating the perceived image block similarity for all the pre-processed dehazed images and all the second reconstructed clear images, and using the calculation result as a fifth loss function; The fourth loss function and the fifth loss function are calculated by the second formula to obtain the sixth loss function, where the second formula is: , in, is the sixth loss function, is the fourth loss function, is the second weight hyperparameter, is the fifth loss function; The parameters of the original distortion-aware encoder are updated using the sixth loss function to obtain a target distortion-aware encoder.

6. The image defogging quality assessment method according to claim 3, characterized in that: The image fog density perception network includes a first large kernel selective convolution, a first maximum pooling layer, a first average pooling layer, a second large kernel selective convolution, a 1×1 convolution layer and a Sigmoid activation layer. The process of performing image fog density perception analysis on each of the original foggy images and the pre-processed defogged images corresponding to each of the defogged images to be evaluated through the image fog density perception network to obtain a fog density predictor and an original fog density perception feature map corresponding to each of the defogged images to be evaluated includes: Performing feature extraction on each of the original foggy images through the first large kernel selective convolution to obtain an original fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing pooling processing on each of the original fog density perception feature maps through the first maximum pooling layer to obtain a first pooled fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing pooling processing on each of the original fog density perception feature maps through the first average pooling layer to obtain a second pooled fog density perception feature map corresponding to each of the defogged images to be evaluated; Cascadingly fusing each of the first pooled fog density perception feature maps with the second pooled fog density perception feature maps corresponding to each of the defogged images to be evaluated, to obtain a fused fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing feature compression on each of the fused fog density perception feature maps through selective convolution with the second large kernel, thereby obtaining a first compressed fog density perception feature map corresponding to each of the defogged images to be evaluated; Performing feature compression on each of the first compressed fog density perception feature maps through the 1×1 convolutional layer to obtain a second compressed fog density perception feature map corresponding to each of the defogged images to be evaluated; Smoothing each of the second compressed fog density perception feature maps through the Sigmoid activation layer to obtain an image fog density map corresponding to each of the defogging images to be evaluated; Calculating a structural similarity index for each of the original foggy images and the preprocessed defogged images corresponding to each of the defogged images to be evaluated, and using the calculation results as pseudo labels, thereby obtaining pseudo labels corresponding to each of the defogged images to be evaluated; Performing loss function calculation on all the image fog density maps and all the pseudo labels using the mean square error loss function algorithm to obtain a seventh loss function; Calculating the Pearson linear correlation coefficient for all the image fog density maps and all the pseudo labels, and using the calculation result as the eighth loss function; Calculating the Spearman rank order correlation coefficient for all the image fog density maps and all the pseudo labels, and using the calculation result as the ninth loss function; The seventh loss function, the eighth loss function, and the ninth loss function are calculated by the third formula to obtain the tenth loss function, and the third formula is: , in, is the tenth loss function, is the seventh loss function, is the third weight hyperparameter, is the eighth loss function, is the fourth weight hyperparameter, is the ninth loss function; The parameters of the image fog density perception network are updated using the tenth loss function to obtain a fog density predictor.

7. The image defogging quality assessment method according to claim 3, characterized in that: The original distortion perception feature set includes a first original distortion perception feature map, a second original distortion perception feature map, a third original distortion perception feature map, and a fourth original distortion perception feature map. The process of performing fusion analysis on each of the original content-aware feature maps, the original distortion-aware feature set corresponding to each of the defogged images to be evaluated, and the original fog density-aware feature map corresponding to each of the defogged images to be evaluated by the target content-aware encoder, the target distortion-aware encoder, and the fog density predictor, respectively, to obtain a target defogging feature map corresponding to each of the defogged images to be evaluated includes: The original content-aware feature maps are calculated using the fourth formula to obtain the original content-aware feature vectors corresponding to the dehazed images to be evaluated. The fourth formula is: , in, is the original content-aware feature vector, is the transformation processing function, is the original content-aware feature map; The first original distortion perception feature map, the second original distortion perception feature map corresponding to each dehazed image to be evaluated, the third original distortion perception feature map corresponding to each dehazed image to be evaluated, and the fourth original distortion perception feature map corresponding to each dehazed image to be evaluated are calculated using the fifth formula to obtain a distortion perception feature map to be processed corresponding to each dehazed image to be evaluated. The fifth formula is: , in, is the distortion perception feature map to be processed, is the first original distortion perception feature map, For upsampling processing, is the second original distortion perception feature map, For splicing processing, is the third original distortion perception feature map, is the fourth original distortion perception feature map; The target distortion perception feature vector corresponding to each dehazed image to be evaluated is obtained by respectively calculating each of the first original distortion perception feature maps, the second original distortion perception feature map corresponding to each dehazed image to be evaluated, the third original distortion perception feature map corresponding to each dehazed image to be evaluated, and the fourth original distortion perception feature map corresponding to each dehazed image to be evaluated using the sixth formula. The sixth formula is: , in, , in, is the target distortion perception feature vector, For splicing processing, is the first original distortion perception feature vector, is the second original distortion perception feature vector, is the third original distortion perception feature vector, is the fourth original distortion perception feature vector, is the first original distortion perception feature vector, the second original distortion perception feature vector, the third original distortion perception feature vector, or the fourth original distortion perception feature vector, is the spatial pyramid pooling block, is the first original distortion perception feature map, the second original distortion perception feature map, the third original distortion perception feature map, or the fourth original distortion perception feature map; The target fog density perception feature vector corresponding to each of the defogging images to be evaluated is obtained by calculating each of the original fog density perception feature maps using the seventh formula. The seventh formula is: , in, is the target fog density perception feature vector, is the transformation processing function, is the original fog density perception feature map; The target fog density perception feature map corresponding to each defogging image to be evaluated is obtained by calculating each of the original fog density perception feature maps using the eighth formula. The eighth formula is: , in, is the target fog density perception feature map, For splicing processing, is the average processing function, is the maximum processing function, is the original fog density perception feature map; respectively splicing the original content perception feature maps, the distortion perception feature maps to be processed corresponding to the defogging images to be evaluated, and the target fog density perception feature maps corresponding to the defogging images to be evaluated, to obtain spliced ​​perception feature maps corresponding to the defogging images to be evaluated; Performing global feature extraction on each of the spliced ​​perceptual feature maps to obtain a global feature map corresponding to each of the dehazed images to be evaluated; Performing regional feature extraction on each of the spliced ​​perceptual feature maps to obtain a regional feature map corresponding to each of the dehazed images to be evaluated; Performing local feature extraction on each of the spliced ​​perceptual feature maps to obtain a local feature map corresponding to each of the dehazed images to be evaluated; The original fusion feature map corresponding to each defogged image to be evaluated is obtained by respectively calculating each of the spliced ​​perceptual feature maps, the global feature map corresponding to each defogged image to be evaluated, the regional feature map corresponding to each defogged image to be evaluated, and the local feature map corresponding to each defogged image to be evaluated through the ninth formula. The ninth formula is: , in, is the original fusion feature map, is the Sigmoid activation function, is the global feature map, is the regional feature map, is the local feature map, is element-wise addition, is element-wise multiplication, is the perception feature map after splicing; Performing maximum pooling processing on each of the original fused feature maps respectively to obtain a maximum pooled fused feature map corresponding to each of the dehazed images to be evaluated; Performing average pooling processing on each of the original fused feature maps respectively to obtain an average pooled fused feature map corresponding to each of the dehazed images to be evaluated; The original fused feature vector corresponding to each defogging image to be evaluated is obtained by respectively calculating each of the original content-aware feature vectors, the target distortion-aware feature vector corresponding to each defogging image to be evaluated, and the target fog density-aware feature vector corresponding to each defogging image to be evaluated through the tenth formula. The tenth formula is: , in, is the original fusion feature vector, is the spatial extension function, For splicing processing, is the original content-aware feature vector, is the target distortion perception feature vector, is the target fog density perception feature vector; Each of the maximum pooled fused feature maps, the average pooled fused feature maps corresponding to each of the defogged images to be evaluated, and the original fused feature vectors corresponding to each of the defogged images to be evaluated are spliced ​​to obtain a target defogging feature map corresponding to each of the defogged images to be evaluated.

8. The image defogging quality assessment method according to claim 3, characterized in that: The process of evaluating and analyzing each of the target defogging feature maps to obtain an evaluation result of the image defogging quality includes: The target defogging feature maps are calculated respectively by the eleventh formula to obtain quality assessment scores corresponding to the defogging images to be evaluated, and all the quality assessment scores are used as the evaluation results of the image defogging quality. The eleventh formula is: , in, is the quality assessment score, is the target dehazing feature map, For flattening, Processing for mapping.

9. An image defogging quality assessment device, characterized in that: include: An import module, configured to import a plurality of defogging images to be evaluated and an original foggy image corresponding to each of the defogging images to be evaluated; A preprocessing module, configured to preprocess each of the defogging images to be evaluated to obtain a preprocessed defogging image corresponding to each of the defogging images to be evaluated; a feature analysis module, configured to construct a training model, and perform feature analysis on each of the preprocessed defogging images and the original foggy image corresponding to each of the defogging images to be evaluated, using the training model, to obtain a target defogging feature map corresponding to each of the defogging images to be evaluated; The evaluation result acquisition module is used to evaluate and analyze each of the target defogging feature maps respectively to obtain an evaluation result of the image defogging quality.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the image defogging quality assessment method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Method for judging quality effect of defogged and enhanced image

    CN101901482A

  • Quality evaluation method and device for defogged image

    CN114155198A

  • Semi-reference image defogging algorithm evaluation method

    CN115187471A

  • No-reference image quality evaluation method based on twin network and feature fusion

    CN115205196A

  • No-reference evaluation method for evaluating underwater video quality

    CN117197720A