A method and device for evaluating the quality of dehazed images

Through deep learning methods, combined with feature extraction, fusion and attention models, the limitations and targeted problems of defogging image quality evaluation are solved, and efficient and reliable defogging image quality evaluation is achieved.

CN114155198BActive Publication Date: 2025-06-27SHENZHEN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111314610.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-08
Publication Date
2025-06-27
Estimated Expiration
2041-11-08

AI Technical Summary

Technical Problem

The prior art has limitations and targeted aspects in the evaluation of defogging image quality, making it difficult to achieve real-time and accurate evaluation, especially in the absence of clear reference images.

Method used

By using deep learning method, by obtaining multiple haze-related features of the defogging sample image, feature extraction and fusion are used to extract and fusion using the feature extraction blocks in the preset defogging image initial evaluation model, deep aggregation is performed in combination with the attention model, and reverse training is performed through the sorting loss function to obtain the quality evaluation model of the defogging image.

Benefits of technology

It realizes efficient and reliable evaluation of the quality of defogging image, and can make more comprehensive use of feature information, improve model efficiency and accuracy, and enhance the robustness of the results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155198B_ABST
    Figure CN114155198B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method for evaluating the quality of a defogged image, which is characterized by including: obtaining a defogging sample image to be trained, and extracting a plurality of haze-related features from the defogging sample image; using a plurality of feature extraction blocks in a preset initial evaluation model of the defogged image to respectively extract features from the plurality of haze-related features, and performing feature fusion on the extracted features; using an attention model to perform deep aggregation on the fused features to obtain an aggregation result, and processing the aggregation result through a basic block and a fully connected layer to output an evaluation result of the defogging sample image; using a preset ranking loss function to perform backpropagation training on the evaluation result until the preset initial evaluation model of the defogged image converges, to obtain a quality evaluation model of the defogged image; obtaining a defogged image to be evaluated, and inputting the defogged image into the quality evaluation model of the defogged image to obtain a quality evaluation score of the defogged image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of computer technology, and in particular, to a method and device for evaluating the quality of defogged images. Background Art

[0002] Image Quality Assessment (IQA) refers to the analysis and research of image characteristics, such as clarity, authenticity, etc., so as to evaluate the quality of the image. It has wide practicability in the fields of image restoration, image compression, video coding and decoding, autonomous driving, etc. With the continuous emergence of defogging algorithms, how to correctly evaluate the visual quality of defogged images has become increasingly urgent and important.

[0003] Currently, there are mainly the following difficulties in researching the quality evaluation of defogged images: 1) Limitation: Subjective evaluation is the most accurate and reliable but time-consuming and laborious, and cannot be applied in real time. In objective evaluation, due to the lack of clear reference images in real life, full-reference (FR) and reduced-reference (RR) image quality evaluation methods are limited; 2) Specificity: In recent years, quality evaluation methods for general image distortion types (compression, blur, etc.) have achieved great success. However, there are differences between haze images and general distortions. With the increasing demand for image defogging, the quality evaluation indicators for defogged images are few, and there is still a lack of methods that can correctly distinguish the quality of images. Summary of the Invention

[0004] To solve the above technical problems, embodiments of the present invention provide a method for evaluating the quality of defogged images, including:

[0005] Obtain the defogged sample images to be trained, and extract multiple haze-related features from the defogged sample images;

[0006] Use multiple feature extraction blocks in the preset initial evaluation model of defogged images to respectively extract features from the multiple haze-related features, and perform feature fusion on the extracted features;

[0007] Use the attention model to perform deep aggregation on the fused features to obtain an aggregation result, and process the aggregation result through basic blocks and fully connected layers to output the evaluation result of the defogged sample images;

[0008] Use the preset sorting loss function to perform backpropagation training on the evaluation result until the preset initial evaluation model of defogged images converges to obtain a quality evaluation model of defogged images;

[0009] Obtain the dehazed image to be evaluated, and input the dehazed image into the quality evaluation model of the dehazed image to obtain the quality evaluation score of the dehazed image.

[0010] Further, the multiple feature extraction blocks include basic blocks and dilation blocks. The use of the multiple basic blocks and dilation blocks in the preset initial evaluation model of the dehazed image to respectively extract features from the multiple haze-related features and fuse the extracted features includes:

[0011] Use the basic block and the dilation block to extract each haze-related feature and respectively output different scale features;

[0012] Fuse the different scale features output by the dilation block.

[0013] Further, the basic block includes a first basic block and a second basic block, and the dilation block includes a first dilation block, a second dilation block, and a third dilation block. The use of the basic block and the dilation block to extract each haze-related feature and respectively output different scale features includes:

[0014] Use the first basic block to extract features from each haze-related feature to obtain a first haze feature, and use the second basic block to extract features from the first haze feature to obtain a second haze feature;

[0015] Use the first dilation block with a first dilation rate parameter to extract features from the second haze feature to obtain a first scale feature of each haze-related feature; use the second dilation block with a second dilation rate parameter to extract features from the first scale feature to obtain a second scale feature of each haze-related feature, and use the third dilation block with a third dilation rate parameter to extract features from the second scale feature to obtain a third scale feature of each haze-related feature;

[0016] The fusion of the different scale features output by the dilation block includes:

[0017] Feature splice each of the first scale features to obtain a first spliced feature, splice the first spliced feature with each of the second scale features to obtain a second spliced feature, and splice the second spliced feature with each of the third scale features to obtain a third spliced feature.

[0018] Further, the use of the attention model to deeply aggregate the fused features to obtain an aggregation result includes:

[0019] Use the channel attention module to aggregate the fused features to obtain a first aggregated feature;

[0020] Use the contrast attention module to aggregate the fused features to obtain a second aggregated feature;

[0021] Element-wise multiply the first aggregation feature and the second aggregation feature to obtain an aggregation result.

[0022] Further, the ranking loss function L is:

[0023] L = L rank + L1,

[0024]

[0025]

[0026] where D = -[Q pre (x i ) - Q pre (x j )][Q gt (x i ) - Q gt (x j )], N is the number of batch training samples; Q pre (x) and Q gt (x) respectively represent the predicted quality score and the true reference quality score of the input image, x i and x j are indices in the batch training images, with ranges [1, N - 1] and [2, N] respectively, and D represents the distance between the predicted and true quality differences.

[0027] An embodiment of the present invention provides a quality evaluation device for defogged images, including:

[0028] An acquisition module, configured to acquire defogging sample images to be trained, and extract a plurality of haze-related features from the defogging sample images;

[0029] A processing module, configured to respectively perform feature extraction on the plurality of haze-related features by using a plurality of feature extraction blocks in a preset initial evaluation model for defogged images, and perform feature fusion on the extracted features;

[0030] The processing module is further configured to perform deep aggregation on the fused features by using an attention model to obtain an aggregation result, and process the aggregation result through a basic block and a fully connected layer to output an evaluation result of the defogging sample image;

[0031] The processing module is further configured to perform backpropagation training on the evaluation result by using a preset ranking loss function until the preset initial evaluation model for defogged images converges, to obtain a quality evaluation model for defogged images;

[0032] An execution module, configured to obtain a defogged image to be evaluated, input the defogged image into a quality evaluation model of the defogged image, and obtain a quality evaluation score of the defogged image.

[0033] Further, the feature extraction block includes a basic block and an expansion block, and the processing module includes:

[0034] A first processing sub-module, configured to extract each haze-related feature by using the basic block and the expansion block and respectively output different scale features;

[0035] A first execution sub-module, configured to fuse the different scale features output by the expansion block.

[0036] Further, the basic block includes a first basic block and a second basic block, the expansion block includes a first expansion block, a second expansion block, and a third expansion block, and the first execution sub-module includes:

[0037] A third processing sub-module, configured to extract a first haze feature from each haze-related feature by using the first basic block, and extract a second haze feature from the first haze feature by using the second basic block;

[0038] A fourth processing sub-module, configured to extract a first scale feature of each haze-related feature from the second haze feature by using the first expansion block provided with a first expansion rate parameter; extract a second scale feature of each haze-related feature from the first scale feature by using the second expansion block provided with a second expansion rate parameter, and extract a third scale feature of each haze-related feature from the second scale feature by using the third expansion block provided with a third expansion rate parameter;

[0039] A second execution sub-module, configured to perform feature splicing on each of the first scale features to obtain a first spliced feature, splice the first spliced feature with each of the second scale features to obtain a second spliced feature, and splice the second spliced feature with each of the third scale features to obtain a third spliced feature.

[0040] Further, the processing module further includes:

[0041] A first acquisition sub-module, configured to aggregate the fused features by using a channel attention module to obtain a first aggregated feature;

[0042] A second acquisition sub-module, configured to aggregate the fused features by using a contrast attention module to obtain a second aggregated feature;

[0043] A third execution sub-module, configured to perform element-wise multiplication on the first aggregated feature and the second aggregated feature to obtain an aggregation result.

[0044] Furthermore, the sorting loss function L is as follows:

[0045] L = L rank + L1,

[0046]

[0047]

[0048] where, D = -[Q pre (x i ) - Q pre (x j )][Q gt (x i ) - Q gt (x j )], N is the number of batch training samples; Q pre (x) and Q gt (x) respectively represent the predicted quality score and the true reference quality score of the input image, x i and x j are the indices in the batch training images, with ranges [1, N - 1] and [2, N] respectively, and D represents the distance between the predicted and true quality differences.

[0049] The beneficial effects of the embodiments of the present invention are as follows: The haze - removing image quality evaluation method proposed in the embodiments of the present invention is based on deep learning, which can make the training more efficient and reliable. For haze - removing images, a variety of haze - related features (image depth, standard deviation, dark channel, edge, contrast) are selected and combined with the attention mechanism, and a new feature fusion model is proposed, enabling the network to fully learn features at different scales, thereby making more comprehensive use of feature information; considering the contrast sensitivity of the human visual system and the high correlation between image contrast and haze - removing distortion, the embodiments of the present invention use a contrast attention module to enable the network to focus on sensitive and distorted areas, thereby improving the model efficiency and accuracy; in addition, the sorting loss function proposed by the present invention can enhance the robustness of the results. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following - described drawings are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0051] Figure 1 It is a schematic flowchart of a method for evaluating the quality of a haze - removing image provided by an embodiment of the present invention;

[0052] Figure 2 Schematic diagram of a network framework applied to a method for evaluating the quality of defogged images provided by an embodiment of the present invention;

[0053] Figure 3 Schematic diagram of a process for obtaining an aggregation result by deeply aggregating features after multi-scale fusion using an attention model provided by an embodiment of the present invention;

[0054] Figure 4 Comparison chart of experimental results of different IQ methods on a real dataset (exBeDDE) provided by an embodiment of the present invention;

[0055] Figure 5 Comparison chart of experimental results of different IQ methods on a synthetic dataset (SHRQ) provided by an embodiment of the present invention;

[0056] Figure 6 Schematic diagram of the structure of a device for evaluating the quality of defogged images provided by an embodiment of the present invention. Detailed implementation manners

[0057] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0058] As Figure 1 shown, this embodiment provides a method for evaluating the quality of defogged images, including:

[0059] S1. Obtain a to-be-trained defogged sample image, and extract a plurality of haze-related features from the defogged sample image;

[0060] It should be noted that the to-be-trained defogged sample images are selected as color defogged images of the same size. Among them, the haze-related features include: depth, standard deviation, dark channel, edge, and contrast. It should be noted that the defogged sample images need to be preprocessed before training. For example, defogged images of any size can be randomly cropped so that each defogged image has the same size, for example, H*W is 320*640.

[0061] S2. Use a plurality of feature extraction blocks in a preset initial evaluation model for defogged images to respectively extract features from the plurality of haze-related features, and perform feature fusion on the extracted features;

[0062] The initial evaluation model for defogged images uses the DHIQA network model. Preferably, there are two basic blocks, and each basic block is used to extract features related to haze. Among them, the processing kernel sizes of the two basic blocks are different. It should be noted that the basic block includes: a convolutional layer (Conv), an acceleration training layer (Batch Norm), an activation function layer (Leaky RuLU), a max pooling layer (Max Pool), and a blur pooling layer (Blur Pool). Among them, the acceleration training layer is Batch Norm (fully known as Batch Normalization) batch normalization; Leaky ReLU is an activation function.

[0063] S3. Use the attention model to deeply aggregate the fused features to obtain an aggregation result, and process the aggregation result through the basic block and the fully connected layer to output the evaluation result of the defogged sample image;

[0064] In the embodiment of the present invention, there are two types of attention models. One is a multi-channel attention module, and the other is a contrast attention module. Among them, the multi-channel attention module uses the SE (Squeeze-and-Excitation) network model to learn the interdependence between different channels and automatically adjusts the feature response values of each channel; the contrast attention module enables the IQA (Image Quality Assessment) network model to focus on the areas most sensitive to the human visual system.

[0065] S4. Use the preset ranking loss function to perform backpropagation training on the evaluation result until the preset initial evaluation model for defogged images converges to obtain a quality evaluation model for defogged images;

[0066] In the embodiment of the present invention, the ranking loss function L is:

[0067] L = L rank + L1,

[0068]

[0069]

[0070] Among them, D = -[Q ptr (x i ) - Q pre (x j )][Q gt (x i ) - Q gt (x j )], N is the number of batch training samples; Q pre (x) and Q gt(x) represents the predicted quality score and the true reference quality score of the input image, respectively, where x i and x j are the indices in the batch training images, with ranges [1, N - 1] and [2, N] respectively, and D represents the distance between the predicted and true quality differences.

[0071] It should be noted that when the predicted ranking is the same as the true ranking, the loss value is zero (D ij = 0) and does not participate in backpropagation. Conversely, if the rankings are inconsistent, the value of the distance D will be passed into the network as the sorting loss value. The design of the above sorting loss function can not only feedback the inconsistency of the rankings, but also feedback the distance between the predicted and true quality gaps of the measured image pairs. By combining the above two loss functions L rank and L1, not only the absolute quality scores of the test images are considered, but also the sorting relationship between the images is considered.

[0072] S5. Obtain the defogged image to be evaluated, and input the defogged image into the quality evaluation model of the defogged image to obtain the quality evaluation score of the defogged image.

[0073] The defogged image quality evaluation method proposed in the embodiments of the present invention is based on deep learning, which can make the training more efficient and reliable. For defogged images, a variety of haze-related features (image depth, standard deviation, dark channel, edge, contrast) are selected and combined with the attention mechanism to propose a new feature fusion model, enabling the network to fully learn features at different scales, thereby making more comprehensive use of feature information; considering the contrast sensitivity of the human visual system and the high correlation between image contrast and defogging distortion, the embodiments of the present invention use a contrast attention module to enable the network to focus on sensitive and distorted areas, thereby improving the model efficiency and accuracy; in addition, the sorting loss function proposed in the present invention can enhance the robustness of the results.

[0074] An embodiment of the present invention, as Figure 2 shown, Figure 2 is to use multiple feature extraction blocks in the preset initial evaluation model of the defogged image to respectively extract features of the multiple haze-related features, and perform feature fusion on the extracted features, including:

[0075] S22. Use the basic block and the feature block to extract each haze-related feature and respectively output different scale features;

[0076] In this embodiment, multiple feature extraction blocks include a basic block and an expansion block. The basic block includes a first basic block and a second basic block. The expansion block includes a first expansion block, a second expansion block, and a third expansion block. Among them, the processing core sizes of the basic block and the expansion block are different. Each basic block is used to extract features related to multiple haze levels. Each basic block includes a convolutional layer (Conv), an accelerated training layer (Batch Norm), an activation function layer (Leaky RuLU), a max pooling layer (MaxPool), and a blur pooling layer (Blur Pool). The expansion block consists of conv and activation ReLU. By setting the dilation rate parameter of conv, multi-scale information extraction is achieved.

[0077] An embodiment of the present invention, as Figure 2 shown, after extracting five haze-related features, the first expansion block outputs features The second expansion block outputs features The third expansion block outputs features Among them, the dilation rates of the first expansion block, the second expansion block, and the third expansion block are different, so the receptive fields of the dehazed sample images are different, and feature information of different scales can be extracted respectively. The embodiment of the present invention is to solve the problem that in the multi-scale feature extraction task, the internal data structure and spatial hierarchical information are lost during the upsampling and pooling processes and small target information cannot be reconstructed. The expansion block in the embodiment of the present invention uses dilated convolutions with three different dilation rates to extract multi-scale information of features.

[0078] S22. Fuse the different-scale features output by the basic block and the expansion block.

[0079] Specifically, step S22 includes the following steps:

[0080] S221. Use the first basic block to extract the first haze feature from each haze-related feature, and use the second basic block to extract the second haze feature from the first haze feature;

[0081] S222. Use the first expansion block with the first dilation rate parameter to extract the first-scale feature of each haze-related feature from the second haze feature; use the second expansion block with the second dilation rate parameter to extract the second-scale feature of each haze-related feature from the first-scale feature, and use the third expansion block with the third dilation rate parameter to extract the third-scale feature of each haze-related feature from the second-scale feature;

[0082] S223. Feature splices are performed on each of the first-scale features to obtain first spliced features. The first spliced features are spliced with each of the second-scale features to obtain second spliced features. The second spliced features are spliced with each of the third-scale features to obtain third spliced features.

[0083] As Figure 2 shown, are the first-scale features of 5 haze-related features output by the first expansion block. Splice with x1 to obtain first spliced feature f1.

[0084] are the second-scale features of 5 haze-related features output by the second expansion block. Splice the first spliced feature f1 with to obtain second spliced feature f2.

[0085] are the third-scale features of 5 haze-related features output by the third expansion block. Splice the second spliced feature f2 with to obtain third spliced feature f3.

[0086] The embodiments of the present invention further provide a method for deeply aggregating the fused features by using an attention model to obtain an aggregation result, including:

[0087] S31. Use a channel attention module to aggregate the fused features to obtain first aggregated features;

[0088] As Figure 3 shown, in the embodiments of the present invention, the fused features are Figure 2 the third spliced feature f3 obtained after aggregation splicing in. Output this feature to a multi-channel attention feature model to capture context information and centrally learn the regions of interest, and deeply aggregate the features to obtain first aggregated features. As Figure 3 , in this embodiment, input the third spliced feature f3: C*H*W, obtain C*1*1 through average pooling operation, and after conv—ReLU—conv—sigmoid, learn the weight sizes between different channels, and multiply the obtained weights by the original input feature C*H*W to obtain the required first aggregated features. Among them, in the embodiments of the present invention, the multi-channel attention module uses an SE (Squeeze-and-Excitation) network model to learn the interdependence between different channels and automatically adjust the feature response values of each channel.

[0089] S32. Use a contrast attention module to aggregate the fused features to obtain second aggregated features;

[0090] S33. Multiply the first aggregated feature and the second aggregated feature element - by - element to obtain an aggregated result.

[0091] In the embodiments of the present invention, the contrast attention module enables the IQA network model to focus on the regions most sensitive to the human visual system. The second aggregated feature is obtained by using the contrast attention module for feature aggregation. As Figure 3 shown, multiply the first aggregated feature and the second aggregated feature element - by - element and output the aggregated result.

[0092] In the embodiments of the present invention, after obtaining the quality evaluation model for the defogged image, the defogged image to be evaluated can be evaluated, and finally the corresponding quality evaluation score is output. Experiments were carried out on the quality evaluation model for the defogged image, as Figure 4 and Figure 5 shown in Table 1 and Table 2. It can be seen that the three ratios in the two tables represent different proportions between the training set and the test set. According to the result comparison, the following conclusions can be drawn:

[0093] 1. Whether in the real or synthetic defogged image dataset, the IQA method is mainly judged by three evaluation indicators: Spearman Rank - order Correlation Coefficient (SRCC), Pearson Linear Correlation Coefficient (PLCC), and Root Mean Squared Error (RMSE). The method of the present invention is significantly better than other No - Reference (NR) IQA methods in all three indicators, indicating that DHIQA not only conforms more to human visual perception but also is a powerful defog evaluation standard.

[0094] 2. After the synthetic foggy image is processed by the defogging algorithm, there are still certain differences between its result and the image of the real scene. The DHIQA method proposed by the present invention shows superior performance on both datasets, which means that the generalization ability of this model is strong and it can be effectively applied to multiple datasets, facilitating user use.

[0095] Therefore, compared with traditional machine learning methods (such as BRISQUE, BMPRI, and DHQI), the multi - scale feature extraction in the embodiments of the present invention can learn more comprehensive image information, and the result is more accurate and reliable after being trained by the deep learning network; the contrast attention mechanism proposed by the present invention makes full use of the importance of the contrast sensitivity of the human visual system, enables the network to focus on sensitive regions, and improves the network calculation efficiency; the ranking loss function proposed by the present invention uses the ranking characteristics existing between images, which can help the network model reduce the error rate and thus improve the model performance.

[0096] An embodiment of the present invention further provides a quality evaluation device for defogged images, as Figure 6 shown, including: an acquisition module 2100, a processing module 2200, and an execution module 2300; wherein, the acquisition module is configured to acquire a defogging sample image to be trained, and extract a plurality of haze-related features from the defogging sample image; the processing module is configured to use a plurality of feature extraction blocks in a preset initial evaluation model for defogged images to respectively extract features from the plurality of haze-related features, and perform feature fusion on the extracted features; the processing module is further configured to use an attention model to perform deep aggregation on the fused features to obtain an aggregation result, and process the aggregation result through a basic block and a fully connected layer to output an evaluation result of the defogging sample image; the processing module is further configured to use a preset ranking loss function to perform backpropagation training on the evaluation result until the preset initial evaluation model for defogged images converges, to obtain a quality evaluation model for defogged images; the execution module is configured to acquire a defogged image to be evaluated, and input the defogged image into the quality evaluation model for defogged images to obtain a quality evaluation score for the defogged image.

[0097] The defogged image quality evaluation method proposed in the embodiment of the present invention is based on deep learning, which can make the training more efficient and reliable. For defogged images, a variety of haze-related features (image depth, standard deviation, dark channel, edge, contrast) are selected and combined with the attention mechanism, and a new feature fusion model is proposed, enabling the network to fully learn features at different scales, thereby making more comprehensive use of feature information; considering the contrast sensitivity of the human visual system and the high correlation between image contrast and defogging distortion, the embodiment of the present invention uses a contrast attention module to enable the network to focus on sensitive and distorted areas, thereby improving the model efficiency and accuracy; in addition, the ranking loss function proposed by the present invention can enhance the robustness of the results.

[0098] In some embodiments, the plurality of feature extraction blocks include basic blocks and dilation blocks, and the processing module includes: a first processing sub-module, configured to use the basic blocks and dilation blocks to extract each haze-related feature and respectively output different-scale features; a first execution sub-module, configured to fuse the different-scale features output by the basic blocks and the dilation blocks.

[0099] In some embodiments, the basic block includes a first basic block and a second basic block, and the dilation block includes a first dilation block, a second dilation block, and a third dilation block; the first execution sub-module includes: a third processing sub-module, configured to perform feature extraction on each haze-related feature using the first basic block to obtain a first haze feature, and perform feature extraction on the first haze feature using the second basic block to obtain a second haze feature; a fourth processing sub-module, configured to perform feature extraction on the second haze feature using the first dilation block with a first dilation rate parameter to obtain a first-scale feature of each haze-related feature; perform feature extraction on the first-scale feature using the second dilation block with a second dilation rate parameter to obtain a second-scale feature of each haze-related feature, and perform feature extraction on the second-scale feature using the third dilation block with a third dilation rate parameter to obtain a third-scale feature of each haze-related feature; a second execution sub-module, configured to perform feature concatenation on each of the first-scale features to obtain a first concatenated feature, concatenate the first concatenated feature with each of the second-scale features to obtain a second concatenated feature, and concatenate the second concatenated feature with each of the third-scale features to obtain a third concatenated feature.

[0100] In some embodiments, the processing module further includes: a first acquisition sub-module, configured to aggregate the fused features using a channel attention module to obtain a first aggregated feature; a second acquisition sub-module, configured to aggregate the fused features using a contrast attention module to obtain a second aggregated feature; a third execution sub-module, configured to perform element-wise multiplication on the first aggregated feature and the second aggregated feature to obtain an aggregation result.

[0101] In some embodiments, the ranking loss function L is: L = L rank + L1,

[0102]

[0103]

[0104] where, D = -[Q pre (x i ) - Q pre (x j )][Q gt (x i ) - Q gt (x j )], N is the number of batch training samples; Q pre (x) and Q gt (x) respectively represent the predicted quality score and the true reference quality score of the input image, x i and x jis the index in the batch training images, with ranges of [1, N - 1] and [2, N] respectively, and D represents the distance between the predicted and the true quality differences.

[0105] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.

[0106] It should be understood that although each step in the flowchart of the accompanying drawings is shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and they can be executed in other orders. Moreover, at least some of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0107] The above are only some embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for evaluating the quality of a defogged image, characterized in that, Including: Obtain a dehazing sample image to be trained, and extract multiple haze-related features from the dehazing sample image; Use multiple feature extraction blocks in a preset initial evaluation model for dehazing images to respectively extract features from the multiple haze-related features, and fuse the extracted features; Use an attention model to perform deep aggregation on the fused features to obtain an aggregation result, and process the aggregation result through a basic block and a fully connected layer to output an evaluation result of the dehazing sample image; Use a preset sorting loss function to perform backpropagation training on the evaluation result until the preset initial evaluation model for dehazing images converges, to obtain a quality evaluation model for dehazing images; Obtain a dehazing image to be evaluated, input the dehazing image into the quality evaluation model for the dehazing image, and obtain a quality evaluation score for the dehazing image.

2. The quality evaluation method according to claim 1, wherein The multiple feature extraction blocks include basic blocks and dilation blocks. Using the multiple basic blocks and dilation blocks in the preset initial evaluation model for dehazing images to respectively extract features from the multiple haze-related features, and fuse the extracted features, includes: Use the basic blocks and the dilation blocks to extract each haze-related feature and respectively output different scale features; Fuse the different scale features output by the dilation blocks.

3. The quality evaluation method according to claim 2, wherein The basic blocks include a first basic block and a second basic block, and the dilation blocks include a first dilation block, a second dilation block, and a third dilation block. Using the basic blocks and the dilation blocks to extract each haze-related feature and respectively output different scale features, includes: Use the first basic block to extract features from each haze-related feature to obtain a first haze feature, and use the second basic block to extract features from the first haze feature to obtain a second haze feature; Use the first dilation block with a first dilation rate parameter to extract features from the second haze feature to obtain a first scale feature of each haze-related feature; use the second dilation block with a second dilation rate parameter to extract features from the first scale feature to obtain a second scale feature of each haze-related feature, and use the third dilation block with a third dilation rate parameter to extract features from the second scale feature to obtain a third scale feature of each haze-related feature; The fusing the different scale features output by the dilation blocks includes: Perform feature concatenation on each of the first scale features to obtain a first concatenated feature, concatenate the first concatenated feature with each of the second scale features to obtain a second concatenated feature, and concatenate the second concatenated feature with each of the third scale features to obtain a third concatenated feature.

4. The quality evaluation method according to claim 1, wherein The using an attention model to perform deep aggregation on the fused features to obtain an aggregation result includes: Use a channel attention module to aggregate the fused features to obtain a first aggregated feature; Use a contrast attention module to aggregate the fused features to obtain a second aggregated feature; Element-wise multiply the first aggregated feature and the second aggregated feature to obtain an aggregation result.

5. The quality evaluation method according to any one of claims 1 to 4, characterized in that, The sorting loss function L is: L = L rank + L1, Among them, D = -[Q pre (x i ) - Q pre (x j )][Q gt (x i ) - Q gt (x j )], where N is the number of batch training samples; Q pre (x) and Q gt (x) respectively represent the predicted quality score and the true reference quality score of the input image, x i and x j are the indices in the batch training images, with ranges [1, N - 1] and [2, N] respectively, and D represents the distance between the predicted and true quality differences.

6. An apparatus for evaluating the quality of a defogged image, characterized in that, Including: An acquisition module, configured to obtain a dehazing sample image to be trained, and extract multiple haze-related features from the dehazing sample image; A processing module, configured to respectively perform feature extraction on the plurality of haze-related features by using a plurality of feature extraction blocks in a preset initial evaluation model for defogged images, and perform feature fusion on the extracted features; The processing module is further configured to perform depth aggregation on the fused features by using an attention model to obtain an aggregation result, and process the aggregation result through a basic block and a fully-connected layer to output an evaluation result of the defogged sample image; The processing module is further configured to perform reverse training on the evaluation result by using a preset ranking loss function until the preset initial evaluation model for defogged images converges, so as to obtain a quality evaluation model for defogged images; An execution module, configured to obtain a defogged image to be evaluated, input the defogged image into the quality evaluation model for defogged images, and obtain an evaluation score of the defogged image, i.e., the quality evaluation score of the defogged image.

7. The quality evaluation device according to claim 6, wherein The plurality of feature extraction blocks include basic blocks and dilation blocks, and the processing module includes: A first processing sub-module, configured to extract each haze-related feature by using the basic block and the feature block and respectively output different-scale features; A first execution sub-module, configured to fuse the different-scale features output by the dilation block.

8. The quality evaluation device according to claim 7, characterized in that, The basic block includes a first basic block and a second basic block, the dilation block includes a first dilation block, a second dilation block, and a third dilation block, and the first execution sub-module includes: A third processing sub-module, configured to perform feature extraction on each haze-related feature by using the first basic block to obtain a first haze feature, and perform feature extraction on the first haze feature by using the second basic block to obtain a second haze feature; A fourth processing sub-module, configured to perform feature extraction on the second haze feature by using the first dilation block provided with a first dilation rate parameter to obtain a first-scale feature of each haze-related feature; perform feature extraction on the first-scale feature by using the second dilation block provided with a second dilation rate parameter to obtain a second-scale feature of each haze-related feature, and perform feature extraction on the second-scale feature by using the third dilation block provided with a third dilation rate parameter to obtain a third-scale feature of each haze-related feature; A second execution sub-module, configured to perform feature concatenation on each of the first-scale features to obtain a first concatenated feature, concatenate the first concatenated feature with each of the second-scale features to obtain a second concatenated feature, and concatenate the second concatenated feature with each of the third-scale features to obtain a third concatenated feature.

9. The quality evaluation device according to claim 6, characterized in that The processing module further includes: A first acquisition sub-module, configured to aggregate the fused features by using a channel attention module to obtain a first aggregated feature; A second acquisition sub-module, configured to aggregate the fused features by using a contrast attention module to obtain a second aggregated feature; A third execution sub-module, configured to perform element-wise multiplication on the first aggregated feature and the second aggregated feature to obtain an aggregation result.

10. The quality evaluation device according to any one of claims 6 to 9, characterized in that The ranking loss function L is: L = L rank + L1, Among them, D = -[Q pre (x i ) - Q pre (x j )][Q gt (x i ) - Q gt (x j )], where N is the number of batch training samples; Q pre (x) and Q gt (x) represent the predicted quality score and the true reference quality score of the input image respectively, x i and x j are the indices in the batch training images, with ranges [1, N - 1] and [2, N] respectively, and D represents the distance between the predicted and true quality differences.

Citation Information

Patent Citations

  • A Multi-Scale Feature Fusion Network based on GANs for Haze Removal

    AU2020100274A4

  • Image quality evaluation method based on fusion of advanced visual perception features and depth features

    CN111429402A