An underwater image quality evaluation method guided by degradation information
By constructing an image enhancement network and a quality evaluation network, and using degraded information extraction features to evaluate underwater image quality, it solves the problem that it is difficult to effectively evaluate underwater image quality in the prior art, and achieves more accurate quality evaluation results.
Patent Information
- Application Number
- CN202210654333.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-10
AI Technical Summary
The prior art is difficult to effectively evaluate the quality of underwater images, especially in the absence of original reference information, and traditional reference quality evaluation methods cannot accurately predict the quality of underwater images.
Using the underwater image quality evaluation method guided by degradation information, two neural networks are constructed: the image enhancement network and the quality evaluation network. Image enhancement networks are used to extract degraded information in underwater images, while quality evaluation networks use these features for quality evaluation.
Automatic extraction of quality-related features is achieved, the correlation between objective evaluation results and subjective perception is improved, and the quality of underwater images can be more accurately evaluated.
Smart Images

Figure CN115170936B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image quality evaluation technology, and more particularly to an underwater image quality evaluation method guided by degradation information. Background Art
[0002] With the continuous development of digital imaging technology, the quality of captured images is getting higher and higher. The acquisition of high-quality underwater images is of great significance for ocean engineering and marine environmental protection. However, due to the complex underwater imaging environment, underwater images are inevitably affected by wavelength-dependent absorption and scattering, including forward scattering and backward scattering. Therefore, the original underwater images often have problems such as color bias, uneven illumination, reduced contrast and visibility, which seriously affect the performance of underwater application systems. Although in this case, the quality of captured images can be improved by optimizing the imaging equipment, there are limitations in the underwater environment and the investment cost. Therefore, in practical applications, underwater image enhancement has always been highly regarded by researchers and has made great progress in the past few years. However, developing a practical underwater image enhancement method is still challenging. For example, how to effectively remove color bias, effectively improve visibility, and improve efficiency are all research directions worthy of attention.
[0003] Currently, since underwater images have a wide range of applications in fields such as ocean engineering and marine environmental protection, in order to better guide underwater image enhancement methods to obtain the highest-quality underwater enhanced images, it is necessary to design an objective method that can automatically evaluate the quality of underwater enhanced images. The research on objective image quality evaluation can be roughly divided into three categories according to the need for original reference information, namely full-reference quality evaluation (FR-IQA), no-reference quality evaluation (NR-IQA), and reduced-reference quality evaluation (RR-IQA). Full-reference quality evaluation and reduced-reference quality evaluation respectively require all or part of the original reference information. For underwater images captured in the underwater environment, usually no original image is provided as the original reference information. Therefore, no-reference quality evaluation is more applicable to underwater images and has more research value.
[0004] Reference-free quality assessment aims to accurately predict the quality of an image in the absence of the original reference information provided by the original image. This topic has received attention in recent years and certain progress has been made. However, there are currently few reference-free quality assessment methods for underwater images. Most existing reference-free quality assessment methods are based on the assumption that the type of image distortion is known and are used for quality assessment tasks of images in specific scenarios, especially natural scene images, screen images, tone-mapped images, and low-light images. These methods calculate the quality of corresponding features by designing feature extraction methods for specific distortions and are used for quality assessment tasks. Most current traditional reference-free quality assessment methods are aimed at synthetic distortions or common distortions, and the research is basically carried out on some synthetic distortion benchmark databases. However, the particularity of the underwater scenario limits the performance of these traditional reference-free quality assessment methods. Due to the uneven attenuation of light during transmission in the water medium, the quality of underwater images captured by imaging devices is affected, specifically manifested as complex types of distortions in underwater images, such as color cast, detail blur, and low contrast. Therefore, it is particularly important to develop a quality assessment method for underwater images. After investigation, there are currently several large-scale underwater image databases. Taking two of these underwater image databases as examples, one contains a total of 1000 enhanced result images obtained by 10 different underwater image enhancement algorithms, and then a reference-free quality assessment method is designed according to the color and brightness characteristics of underwater images; the other contains 890 underwater images, and a reference-free quality assessment method is designed by extracting the color, contrast, and sharpness characteristics of underwater images. Compared with traditional reference-free quality assessment methods, the above two reference-free quality assessment methods have proven their effectiveness through experiments.
[0005] In addition, researchers have also successively proposed some quality assessment methods for underwater images in recent years. They basically take into account the unique distortions in the color, contrast, and sharpness of underwater images. However, these objective quality assessment methods still require artificial design of the feature extraction process, which is time-consuming and laborious. Therefore, it is very necessary to design a quality assessment method that can automatically extract quality-related features. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method for underwater image quality assessment guided by degradation information, which can automatically extract quality-related features, save time and effort, and effectively improve the correlation between objective evaluation results and subjective perception.
[0007] The technical solution adopted by the present invention to solve the above technical problems is as follows: A method for underwater image quality assessment guided by degradation information, characterized by comprising the following steps:
[0008] Step 1: Construct two neural networks. The first neural network serves as an image enhancement network, and the second neural network serves as a quality evaluation network;
[0009] The image enhancement network includes 1 first convolutional block, 4 second convolutional blocks, 4 third convolutional blocks, 4 fourth convolutional blocks, and 1 fifth convolutional block. The encoding network is composed of the first convolutional block, the 1st second convolutional block, the 2nd second convolutional block, the 3rd second convolutional block, and the 4th second convolutional block. The channel attention module is composed of the 1st third convolutional block, the 2nd third convolutional block, the 3rd third convolutional block, and the 4th third convolutional block. The decoding network is composed of the 1st fourth convolutional block, the 2nd fourth convolutional block, the 3rd fourth convolutional block, the 4th fourth convolutional block, and the fifth convolutional block. The input channel number of the first convolutional block is 3, and the output channel number is 32. The input end of the first convolutional block serves as the input end of the image enhancement network and simultaneously receives the R, G, and B channels of an RGB image with a size of H×W. Denote the feature map with a size of output from the output end of the first convolutional block as F E1 ; The input channel number of the 1st second convolutional block is 32, and the output channel number is 32. The input end of the 1st second convolutional block receives F E1 , and denote the feature map with a size of output from the output end of the 1st second convolutional block as F E2 ; The input channel number of the 2nd second convolutional block is 32, and the output channel number is 64. The input end of the 2nd second convolutional block receives F E2 , and denote the feature map with a size of output from the output end of the 2nd second convolutional block as F E3 ; The input channel number of the 3rd second convolutional block is 64, and the output channel number is 128. The input end of the 3rd second convolutional block receives F E3 , and denote the feature map with a size of output from the output end of the 3rd second convolutional block as F E4 ; The input channel number of the 4th second convolutional block is 128, and the output channel number is 256. The input end of the 4th second convolutional block receives F E4 , and denote the feature map with a size of output from the output end of the 4th second convolutional block as F E5 ; The input channel number of the 1st third convolutional block is 32, and the output channel number is 32. The input end of the 1st third convolutional block receives F E2 , and denote the feature map with a size of output from the output end of the 1st third convolutional block as F C1 ; The input channel number of the 2nd third convolutional block is 64, and the output channel number is 64. The input end of the 2nd third convolutional block receives F E3, denote the feature map with a size of output from the output end of the second third convolution block as F C2 ; The number of input channels of the third third convolution block is 128, and the number of output channels is 128. The input end of the third third convolution block receives F E4 , and denote the feature map with a size of output from the output end of the third third convolution block as F C3 ; The number of input channels of the fourth third convolution block is 256, and the number of output channels is 256. The input end of the fourth third convolution block receives F E5 , and denote the feature map with a size of output from the output end of the fourth third convolution block as F C4 ; The number of input channels of the first fourth convolution block is 256, and the number of output channels is 256. The input end of the first fourth convolution block receives F E5 , and denote the feature map with a size of output from the output end of the first fourth convolution block as F D1 ; The number of input channels of the second fourth convolution block is 512, and the number of output channels is 128. The input end of the second fourth convolution block receives the feature map F D1 obtained by performing a concatenation operation on F C4 and F with a size of DC1 , and denote the feature map with a size of output from the output end of the second fourth convolution block as F D2 , and at the same time, regard F DC1 as the first intermediate layer feature map; The number of input channels of the third fourth convolution block is 256, and the number of output channels is 64. The input end of the third fourth convolution block receives the feature map F D2 obtained by performing a concatenation operation on F C3 and F with a size of DC2 , and denote the feature map with a size of output from the output end of the third fourth convolution block as F D3 , and at the same time, regard F DC2 as the second intermediate layer feature map; The number of input channels of the fourth fourth convolution block is 128, and the number of output channels is 32. The input end of the fourth fourth convolution block receives the feature map F D3 obtained by performing a concatenation operation on F C2 and F with a size of DC3 , and denote the feature map with a size of output from the output end of the fourth fourth convolution block as F D4 , and at the same time, regard F DC3As the third intermediate layer feature map; the number of input channels of the fifth convolutional block is 64, and the number of output channels is 3. The input end of the fifth convolutional block receives the feature map F D4 and F C1 after concatenation operation, and the size of the obtained feature map is denoted as F DC4 . Denote the feature map with a size of H×W×3 output from the output end of the fifth convolutional block as F D5 . At the same time, regard F DC4 as the fourth intermediate layer feature map, regard F D5 as the image degradation information corresponding to the RGB image. Perform element-wise addition on the RGB image and its corresponding image degradation information, and regard the image obtained by element-wise addition as the enhanced result image output from the output end of the image enhancement network;
[0010] The quality evaluation network includes 1 sixth convolutional block, 4 seventh convolutional blocks, 12 eighth convolutional blocks, 4 ninth convolutional blocks, 4 global average pooling models, and 1 fully connected layer. The encoding network is composed of 1 sixth convolutional block, 4 seventh convolutional blocks, and 12 eighth convolutional blocks. The feature fusion module is composed of 4 ninth convolutional blocks. The regression network is composed of 4 global average pooling models and 1 fully connected layer. The number of input channels of the sixth convolutional block is 3, and the number of output channels is 64. The input end of the sixth convolutional block simultaneously receives the R, G, and B channels of an RGB image with a size of H×W. The RGB image received by the input end of the sixth convolutional block is the same as the RGB image received by the input end of the first convolutional block. Denote the feature map with a size of output from the output end of the sixth convolutional block as F Q1 ; the number of input channels of the first seventh convolutional block is 64, and the number of output channels is 256. The input end of the first seventh convolutional block receives F Q1 . Denote the feature map with a size of output from the output end of the first seventh convolutional block as F Q2 ; the number of input channels of the first eighth convolutional block is 256, and the number of output channels is 256. The input end of the first eighth convolutional block receives F Q2 . Denote the feature map with a size of output from the output end of the first eighth convolutional block as F Q3 ; the number of input channels of the second eighth convolutional block is 256, and the number of output channels is 256. The input end of the second eighth convolutional block receives F Q3 . Denote the feature map with a size of output from the output end of the second eighth convolutional block as F Q4 , and regard F Q4 as the fifth intermediate layer feature map; the number of input channels of the second seventh convolutional block is 256, and the number of output channels is 512. The input end of the second seventh convolutional block receives FQ4 , denote the feature map with a size of output from the output end of the second seventh convolution block as F Q5 ; the number of input channels of the third eighth convolution block is 512, and the number of output channels is 512. The input end of the third eighth convolution block receives F Q5 , denote the feature map with a size of output from the output end of the third eighth convolution block as F Q6 ; the number of input channels of the fourth eighth convolution block is 512, and the number of output channels is 512. The input end of the fourth eighth convolution block receives F Q6 , denote the feature map with a size of output from the output end of the fourth eighth convolution block as F Q7 ; the number of input channels of the fifth eighth convolution block is 512, and the number of output channels is 512. The input end of the fifth eighth convolution block receives F Q7 , denote the feature map with a size of output from the output end of the fifth eighth convolution block as F Q8 , and use F Q8 as the sixth intermediate layer feature map; the number of input channels of the third seventh convolution block is 512, and the number of output channels is 1024. The input end of the third seventh convolution block receives F Q8 , denote the feature map with a size of output from the output end of the third seventh convolution block as F Q9 ; the number of input channels of the sixth eighth convolution block is 1024, and the number of output channels is 1024. The input end of the sixth eighth convolution block receives F Q9 , denote the feature map with a size of output from the output end of the sixth eighth convolution block as F Q10 ; the number of input channels of the seventh eighth convolution block is 1024, and the number of output channels is 1024. The input end of the seventh eighth convolution block receives F Q10 , denote the feature map with a size of output from the output end of the seventh eighth convolution block as F Q11 ; the number of input channels of the eighth eighth convolution block is 1024, and the number of output channels is 1024. The input end of the eighth eighth convolution block receives F Q11 , denote the feature map with a size of output from the output end of the eighth eighth convolution block as F Q12 ; the number of input channels of the ninth eighth convolution block is 1024, and the number of output channels is 1024. The input end of the ninth eighth convolution block receives F Q12 , denote the feature map with a size of output from the output end of the ninth eighth convolution block as F Q13; The number of input channels of the 10th eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the 10th eighth convolutional block receives F Q13 , and the feature map with a size of output from the output end of the 10th eighth convolutional block is denoted as F Q14 , and F Q14 is used as the seventh intermediate layer feature map; The number of input channels of the 4th seventh convolutional block is 1024, and the number of output channels is 2048. The input end of the 4th seventh convolutional block receives F Q14 , and the feature map with a size of output from the output end of the 4th seventh convolutional block is denoted as F Q15 ; The number of input channels of the 11th eighth convolutional block is 2048, and the number of output channels is 2048. The input end of the 11th eighth convolutional block receives F Q15 , and the feature map with a size of output from the output end of the 11th eighth convolutional block is denoted as F Q16 ; The number of input channels of the 12th eighth convolutional block is 2048, and the number of output channels is 2048. The input end of the 12th eighth convolutional block receives F Q16 , and the feature map with a size of output from the output end of the 12th eighth convolutional block is denoted as F Q17 , and F Q17 is used as the eighth intermediate layer feature map; The number of input channels of the first input end of the 1st ninth convolutional block is 256, the number of input channels of the second input end is 64, and the number of output channels is 64. The first input end of the 1st ninth convolutional block receives F Q4 , the second input end of the 1st ninth convolutional block receives F DC4 , and the feature map with a size of output from the output end of the 1st ninth convolutional block is denoted as F DQ1 ; The number of input channels of the first input end of the 2nd ninth convolutional block is 512, the number of input channels of the second input end is 128, and the number of output channels is 128. The first input end of the 2nd ninth convolutional block receives F Q8 , the second input end of the 2nd ninth convolutional block receives F DC3 , and the feature map with a size of output from the output end of the 2nd ninth convolutional block is denoted as F DQ2 ; The number of input channels of the first input end of the 3rd ninth convolutional block is 1024, the number of input channels of the second input end is 256, and the number of output channels is 256. The first input end of the 3rd ninth convolutional block receives F Q14 , the second input end of the 3rd ninth convolutional block receives F DC2 , and the feature map with a size of output from the output end of the 3rd ninth convolutional block is denoted as FDQ3 ; The number of input channels at the first input end of the 4th ninth convolutional block is 2048, the number of input channels at the second input end is 512, and the number of output channels is 512. The first input end of the 4th ninth convolutional block receives F Q17 , and the second input end of the 4th ninth convolutional block receives F DC1 . Denote the feature map with a size of output from the output end of the 4th ninth convolutional block as F DQ4 ; The number of input channels of the 1st global average pooling model is 64, and the number of output channels is 64. The input end of the 1st global average pooling model receives F DQ1 , and the output end of the 1st global average pooling model outputs a feature vector with a size of 1×1×64; The number of input channels of the 2nd global average pooling model is 128, and the number of output channels is 128. The input end of the 2nd global average pooling model receives F DQ2 , and the output end of the 2nd global average pooling model outputs a feature vector with a size of 1×1×128; The number of input channels of the 3rd global average pooling model is 256, and the number of output channels is 256. The input end of the 3rd global average pooling model receives F DQ3 , and the output end of the 3rd global average pooling model outputs a feature vector with a size of 1×1×256; The number of input channels of the 4th global average pooling model is 512, and the number of output channels is 512. The input end of the 4th global average pooling model receives F DQ4 , and the output end of the 4th global average pooling model outputs a feature vector with a size of 1×1×512; Perform a concatenation operation on the feature vector with a size of 1×1×64, the feature vector with a size of 1×1×128, the feature vector with a size of 1×1×256, and the feature vector with a size of 1×1×512 to obtain a feature vector with a size of 1×1×960, denoted as F iqa1 ; The number of input channels of the fully connected layer is 960, and the number of output channels is 1. The input end of the fully connected layer receives F iqa1 , and the output end of the fully connected layer outputs a value, which represents the quality prediction score of the RGB image;
[0011] Step 2: Select N1 original underwater images in different scenarios and the corresponding pseudo-label images of each original underwater image to form the first training set; where N1≥800, and the sizes of the original underwater images and the pseudo-label images are H×W, that is, the heights of the original underwater images and the pseudo-label images are H and the widths are W;
[0012] Step 3: Input the R, G, and B channels of each original underwater image in the first training set into the image enhancement network for training. The image enhancement network outputs the enhanced result image corresponding to each original underwater image in the first training set. Then, for each pseudo-label image in the first training set, calculate the loss function value, denoted as Loss IE , where represents the mean squared error loss function value, "|| ||2" is the l2 norm operation symbol, I result represents the enhanced result image corresponding to the original underwater image I raw , I pseudo represents the original underwater image I raw corresponding pseudo-label image, represents the perceptual loss function value, represents the perceptual loss network, i.e., the VGG-16 network, 1 ≤ j ≤ 16, represents the j-th convolutional layer in the VGG-16 network, represents I result input into the feature map output by the j-th convolutional layer in the VGG-16 network, represents I pseudo input into the feature map output by the j-th convolutional layer in the VGG-16 network, H j ×W j ×C j represents and 's size, H j represents and 's height, W j represents and 's width, C j represents and 's number of channels;
[0013] Step 4: Use the first training set to train for more than 100 rounds according to the process in Step 3. Finally, train the image enhancement network training model and freeze the parameters in subsequent training;
[0014] Step 5: Select N2 original underwater images in different scenarios; then use N3 different underwater image enhancement methods to enhance each original underwater image to obtain N3 underwater enhanced images corresponding to each original underwater image; for each of the N3 underwater enhanced images corresponding to each original underwater image, arrange the N3 underwater enhanced images in a row, and combine each underwater enhanced image with each of the subsequent underwater enhanced images pairwise to form image pairs, obtaining a total of (N3 - 1) + (N3 - 2) + … + 1 pairs of image pairs; then form a second training set with N2×((N3 - 1) + (N3 - 2) + … + 1) pairs of image pairs; where N2≥100, N3≥10, and the size of the original underwater image is H×W, that is, the height of the original underwater image is H and the width is W;
[0015] Step 6: Input the R, G, and B channels of each underwater enhanced image in each pair of image pairs in the second training set into the image enhancement network training model at the same time. The image enhancement network training model outputs the first intermediate layer feature map, the second intermediate layer feature map, the third intermediate layer feature map, and the fourth intermediate layer feature map corresponding to each underwater enhanced image in each pair of image pairs in the second training set. Denote the first intermediate layer feature map, the second intermediate layer feature map, the third intermediate layer feature map, and the fourth intermediate layer feature map corresponding to any underwater enhanced image in any pair of image pairs in the second training set as F' DC1 、F' DC2 、F' DC3 、F' DC4 ;
[0016] At the same time, input the R, G, and B channels of each underwater enhanced image in each pair of image pairs in the second training set into the encoding network of the quality evaluation network. The encoding network of the quality evaluation network outputs the fifth intermediate layer feature map, the sixth intermediate layer feature map, the seventh intermediate layer feature map, and the eighth intermediate layer feature map corresponding to each underwater enhanced image in each pair of image pairs in the second training set. Denote the fifth intermediate layer feature map, the sixth intermediate layer feature map, the seventh intermediate layer feature map, and the eighth intermediate layer feature map corresponding to any underwater enhanced image in any pair of image pairs in the second training set as F' Q4 、F' Q8 、F' Q14 、F' Q17; Then, for each pair of underwater enhanced images in the second training set, the fourth intermediate layer feature maps and the fifth intermediate layer feature maps, the third intermediate layer feature maps and the sixth intermediate layer feature maps, the second intermediate layer feature maps and the seventh intermediate layer feature maps, and the first intermediate layer feature maps and the eighth intermediate layer feature maps corresponding to each underwater enhanced image are paired and input into the feature fusion module of the quality evaluation network. The quality evaluation network outputs the quality prediction scores corresponding to each underwater enhanced image in each pair of images in the second training set; then, for each pair of images in the second training set, the loss function value is calculated and denoted as Loss quality , Loss quality = max(0, -R × (Q1 - Q2) + margin), where max() is the maximum value function, Q1 represents the quality prediction score corresponding to the first underwater enhanced image in each pair of images, Q2 represents the quality prediction score corresponding to the second underwater enhanced image in each pair of images, margin is a constant, margin = 0.5, and R represents the subjective preference value between the first underwater enhanced image and the second underwater enhanced image in each pair of images. If the first underwater enhanced image is subjectively preferred, then R = 1; if the second underwater enhanced image is subjectively preferred, then R = -1;
[0017] Step 7: Use the second training set to train for more than 100 rounds according to the process in Step 6, and finally train to obtain the quality evaluation network training model;
[0018] Step 8: Arbitrarily select an underwater image with a size of H × W and denote it as T; then input the R, G, and B channels of T into the image enhancement network training model and the quality evaluation network training model at the same time. The quality evaluation network training model outputs the quality prediction score of T.
[0019] In the said Step 1, the first convolutional block consists of a first convolutional layer and a first ReLU activation layer connected in sequence. The input end of the first convolutional layer is the input end of the first convolutional block. The input end of the first ReLU activation layer receives the feature map output by the output end of the first convolutional layer. The output end of the first ReLU activation layer is the output end of the first convolutional block. Among them, the number of input channels of the first convolutional layer is 3, the number of output channels is 32, and the convolutional kernel size is 1 × 1;
[0020] The second convolutional block consists of a second convolutional layer, a second ReLU activation layer, a third convolutional layer, a fourth convolutional layer, a third ReLU activation layer, and a fifth convolutional layer. The input end of the second convolutional layer is the input end of the second convolutional block where it is located. The input end of the second ReLU activation layer receives the feature map output from the output end of the second convolutional layer. The input end of the third convolutional layer receives the feature map output from the output end of the second ReLU activation layer. The input end of the fourth convolutional layer receives the feature map obtained by element-wise addition of the feature map output from the output end of the third convolutional layer and the feature map received by the input end of the second convolutional layer. The input end of the third ReLU activation layer receives the feature map output from the output end of the fourth convolutional layer. The input end of the fifth convolutional layer receives the feature map output from the output end of the third ReLU activation layer. The output end of the second convolutional block outputs the feature map obtained by element-wise addition of the feature map output from the output end of the fifth convolutional layer and the feature map received by the input end of the fourth convolutional layer; among them, the number of input channels, output channels, and the size of the convolutional kernel of the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer in the first second convolutional block are 32, 32, and 3×3 respectively. The number of input channels of the second convolutional layer in the second second convolutional block is 32, the number of output channels is 64, and the size of the convolutional kernel is 3×3. The number of input channels, output channels, and the size of the convolutional kernel of the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer in the second second convolutional block are 64, 64, and 3×3 respectively. The number of input channels of the second convolutional layer in the third second convolutional block is 64, the number of output channels is 128, and the size of the convolutional kernel is 3×3. The number of input channels, output channels, and the size of the convolutional kernel of the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer in the third second convolutional block are 128, 128, and 3×3 respectively. The number of input channels of the second convolutional layer in the fourth second convolutional block is 128, the number of output channels is 256, and the size of the convolutional kernel is 3×3. The number of input channels, output channels, and the size of the convolutional kernel of the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer in the fourth second convolutional block are 256, 256, and 3×3 respectively.
[0021] In the said Step 1, the third convolution block consists of a first global average pooling layer, a sixth convolution layer, a fourth ReLU activation layer, a seventh convolution layer, and a first Sigmoid activation layer connected in sequence. The input end of the first global average pooling layer is the input end of the third convolution block where it is located. The input end of the sixth convolution layer receives the feature map output from the output end of the first global average pooling layer. The input end of the fourth ReLU activation layer receives the feature map output from the output end of the sixth convolution layer. The input end of the seventh convolution layer receives the feature map output from the output end of the fourth ReLU activation layer. The input end of the first Sigmoid activation layer receives the feature map output from the output end of the seventh convolution layer. The output end of the first Sigmoid activation layer is the output end of the third convolution block where it is located. Among them, the number of input channels of the sixth convolution layer in the first third convolution block is 32, the number of output channels is 2, and the convolution kernel size is 1×1. The number of input channels of the seventh convolution layer in the first third convolution block is 2, the number of output channels is 32, and the convolution kernel size is 1×1. The number of input channels of the sixth convolution layer in the second third convolution block is 64, the number of output channels is 4, and the convolution kernel size is 1×1. The number of input channels of the seventh convolution layer in the second third convolution block is 4, the number of output channels is 64, and the convolution kernel size is 1×1. The number of input channels of the sixth convolution layer in the third third convolution block is 128, the number of output channels is 8, and the convolution kernel size is 1×1. The number of input channels of the seventh convolution layer in the third third convolution block is 8, the number of output channels is 128, and the convolution kernel size is 1×1. The number of input channels of the sixth convolution layer in the fourth third convolution block is 256, the number of output channels is 16, and the convolution kernel size is 1×1. The number of input channels of the seventh convolution layer in the fourth third convolution block is 16, the number of output channels is 256, and the convolution kernel size is 1×1. The input size of the first global average pooling layer in the first third convolution block is The output size is 1×1×32. The input size of the first global average pooling layer in the second third convolution block is The output size is 1×1×64. The input size of the first global average pooling layer in the third third convolution block is The output size is 1×1×128. The input size of the first global average pooling layer in the fourth third convolution block is The output size is 1×1×256.
[0022] In the said Step 1, the fourth convolutional block consists of an eighth convolutional layer, a first upsampling layer, a ninth convolutional layer, a fifth ReLU activation layer, a tenth convolutional layer, an eleventh convolutional layer, a sixth ReLU activation layer, and a twelfth convolutional layer. The input end of the eighth convolutional layer is the input end of the fourth convolutional block where it is located. The input end of the first upsampling layer receives the feature map output from the output end of the eighth convolutional layer. The input end of the ninth convolutional layer receives the feature map output from the output end of the first upsampling layer. The input end of the fifth ReLU activation layer receives the feature map output from the output end of the ninth convolutional layer. The input end of the tenth convolutional layer receives the feature map output from the output end of the fifth ReLU activation layer. The input end of the eleventh convolutional layer receives the feature map obtained by element-wise addition of the feature map output from the output end of the tenth convolutional layer and the feature map received by the input end of the ninth convolutional layer. The input end of the sixth ReLU activation layer receives the feature map output from the output end of the eleventh convolutional layer. The input end of the twelfth convolutional layer receives the feature map output from the output end of the sixth ReLU activation layer. The output end of the fourth convolutional block outputs the feature map obtained by element-wise addition of the feature map output from the output end of the twelfth convolutional layer and the feature map received by the input end of the tenth convolutional layer; among them, the number of input channels of the eighth convolutional layer in the first fourth convolutional block is 256, the number of output channels is 256, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the first fourth convolutional block are 256, 256, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the second fourth convolutional block is 512, the number of output channels is 128, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the second fourth convolutional block are 128, 128, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the third fourth convolutional block is 256, the number of output channels is 64, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the third fourth convolutional block are 64, 64, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the fourth fourth convolutional block is 128, the number of output channels is 32, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the fourth fourth convolutional block are 32, 32, and 3×3 respectively. The sampling multiple of the first upsampling layer is 2 times;
[0023] The fifth convolutional block consists of a thirteenth convolutional layer and a seventh ReLU activation layer connected in sequence. The input end of the thirteenth convolutional layer is the input end of the fifth convolutional block. The input end of the seventh ReLU activation layer receives the feature map output from the output end of the thirteenth convolutional layer, and the output end of the seventh ReLU activation layer is the output end of the fifth convolutional block. Among them, the number of input channels of the thirteenth convolutional layer is 64, the number of output channels is 3, and the convolutional kernel size is 1×1.
[0024] In the said step 1, the sixth convolutional block consists of a fourteenth convolutional layer and a first downsampling layer connected in sequence. The input end of the fourteenth convolutional layer is the input end of the sixth convolutional block. The input end of the first downsampling layer receives the feature map output from the output end of the fourteenth convolutional layer, and the output end of the first downsampling layer is the output end of the sixth convolutional block. Among them, the number of input channels of the fourteenth convolutional layer is 3, the number of output channels is 64, the convolutional kernel size is 7×7, and the downsampling multiple of the first downsampling layer is 2 times.
[0025] The seventh convolutional block consists of the fifteenth convolutional layer, the sixteenth convolutional layer, the seventeenth convolutional layer, and the eighteenth convolutional layer. The common connection end of the input ends of the fifteenth convolutional layer and the eighteenth convolutional layer is the input end of the seventh convolutional block where it is located. The input end of the sixteenth convolutional layer receives the feature map output from the output end of the fifteenth convolutional layer. The input end of the seventeenth convolutional layer receives the feature map output from the output end of the sixteenth convolutional layer. The output end of the seventh convolutional block outputs the feature map obtained by element-wise adding the feature map output from the output end of the seventeenth convolutional layer and the feature map output from the output end of the eighteenth convolutional layer; among them, the number of input channels of the fifteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer and the eighteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the second seventh convolutional block is 256, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the second seventh convolutional block is 128, the number of output channels is 128, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer in the second seventh convolutional block is 128, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the second seventh convolutional block is 256, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the third seventh convolutional block is 512, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the third seventh convolutional block is 256, the number of output channels is 256, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer in the third seventh convolutional block is 256, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the third seventh convolutional block is 512, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the fourth seventh convolutional block is 1024, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the fourth seventh convolutional block is 512, the number of output channels is 512, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer in the fourth seventh convolutional block is 512, the number of output channels is 2048, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the fourth seventh convolutional block is 1024, the number of output channels is 2048, and the kernel size is 1×1;
[0026] The eighth convolutional block consists of the nineteenth convolutional layer, the twentieth convolutional layer, and the twenty-first convolutional layer. The input end of the nineteenth convolutional layer is the input end of the eighth convolutional block it belongs to. The input end of the twentieth convolutional layer receives the feature map output from the output end of the nineteenth convolutional layer. The input end of the twenty-first convolutional layer receives the feature map output from the output end of the twentieth convolutional layer. The output end of the eighth convolutional block outputs the feature map obtained by element-wise adding the feature map output from the output end of the twenty-first convolutional layer and the feature map received by the input end of the nineteenth convolutional layer. Among them, the number of input channels of the nineteenth convolutional layer in the first and second eighth convolutional blocks is 256, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the first and second eighth convolutional blocks is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the first and second eighth convolutional blocks is 64, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the third to fifth eighth convolutional blocks is 512, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the third to fifth eighth convolutional blocks is 128, the number of output channels is 128, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the third to fifth eighth convolutional blocks is 128, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the sixth to tenth eighth convolutional blocks is 1024, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the sixth to tenth eighth convolutional blocks is 256, the number of output channels is 256, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the sixth to tenth eighth convolutional blocks is 256, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the eleventh and twelfth eighth convolutional blocks is 2048, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the eleventh and twelfth eighth convolutional blocks is 512, the number of output channels is 512, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the eleventh and twelfth eighth convolutional blocks is 512, the number of output channels is 2048, and the kernel size is 1×1.
[0027] In the aforementioned Step 1, the ninth convolutional block consists of the twenty-second convolutional layer, the first batch normalization layer, the eighth ReLU activation layer, the second upsampling layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer, the first max pooling layer, and the second Sigmoid activation layer. The input end of the twenty-second convolutional layer is the first input end of the ninth convolutional block where it is located. The common connection end of the input ends of the twenty-fourth convolutional layer and the twenty-sixth convolutional layer is the second input end of the ninth convolutional block where they are located. The input end of the first batch normalization layer receives the feature map output from the output end of the twenty-second convolutional layer. The input end of the eighth ReLU activation layer receives the feature map output from the output end of the first batch normalization layer. The input end of the second upsampling layer receives the feature map output from the output end of the eighth ReLU activation layer. The input end of the twenty-third convolutional layer receives the feature map output from the output end of the second upsampling layer. The input end of the twenty-fifth convolutional layer receives the feature map obtained by concatenating the feature map output from the output end of the twenty-third convolutional layer and the feature map output from the output end of the twenty-fourth convolutional layer. The input end of the first max pooling layer receives the feature map output from the output end of the twenty-sixth convolutional layer. The input end of the second Sigmoid activation layer receives the feature map output from the output end of the first max pooling layer. The feature map output from the output end of the twenty-fifth convolutional layer and the feature map output from the output end of the second Sigmoid activation layer are multiplied pixel by pixel. The feature map obtained by pixel-by-pixel multiplication and the feature map output from the output end of the twenty-fourth convolutional layer are added pixel by pixel. The feature map obtained by pixel-by-pixel addition is output from the output end of the ninth convolutional block. Among them, in the first ninth convolutional block, the number of input channels of the twenty-second convolutional layer is 256, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the first batch normalization layer is 64, and the number of output channels is 64. The sampling multiple of the second upsampling layer is 4 times. The number of input channels of the twenty-third convolutional layer is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-fourth convolutional layer is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-fifth convolutional layer is 128, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-sixth convolutional layer is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the first max pooling layer is 64, and the number of output channels is 1;In the second ninth convolutional block, the input channels of the twenty-second convolutional layer are 512, the output channels are 128, and the convolutional kernel size is 3×3; the input channels of the first batch normalization layer are 128, the output channels are 128; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 128, the output channels are 128, and the convolutional kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 128, the output channels are 128, and the convolutional kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 256, the output channels are 128, and the convolutional kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 128, the output channels are 128, and the convolutional kernel size is 3×3; the input channels of the first max pooling layer are 128, and the output channels are 1; in the third ninth convolutional block, the input channels of the twenty-second convolutional layer are 1024, the output channels are 256, and the convolutional kernel size is 3×3; the input channels of the first batch normalization layer are 256, the output channels are 256; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 256, the output channels are 256, and the convolutional kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 256, the output channels are 256, and the convolutional kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 512, the output channels are 256, and the convolutional kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 256, the output channels are 256, and the convolutional kernel size is 3×3; the input channels of the first max pooling layer are 256, and the output channels are 1; in the fourth ninth convolutional block, the input channels of the twenty-second convolutional layer are 2048, the output channels are 512, and the convolutional kernel size is 3×3; the input channels of the first batch normalization layer are 512, the output channels are 512; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 512, the output channels are 512, and the convolutional kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 512, the output channels are 512, and the convolutional kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 1024, the output channels are 512, and the convolutional kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 512, the output channels are 512, and the convolutional kernel size is 3×3; the input channels of the first max pooling layer are 512, and the output channels are 1.;
[0028] Compared with the prior art, the advantages of the present invention are as follows:
[0029] 1) In order to be able to target the characteristics of underwater images, i.e., the degradation that occurs during the underwater imaging process, obtaining image degradation information becomes crucial. If the image degradation information can be extracted from the input image, then the loss of information such as the color and details of the underwater image can be evaluated based on the obtained image degradation information. Therefore, the method of the present invention discovers the image degradation information of underwater images, constructs an underwater image degradation information discovery network, i.e., an image enhancement network, for training, and extracts the image degradation content and information loss content caused by the underwater environment in the underwater image, which can guide the subsequent image quality evaluation process and obtain a more accurate quality prediction score.
[0030] 2) The method of the present invention selects the trained underwater image degradation information discovery network, i.e., the image enhancement network, to automatically extract the adaptive features related to the image degradation information. When constructing the underwater image degradation information discovery network, i.e., the image enhancement network, multiple residual block structures are used as the encoding network and the decoding network to make the output of the image enhancement network as consistent as possible with the pseudo-labeled images in the image database. During testing, the encoding network with frozen parameters after training is used as the feature extraction network to extract the adaptive features related to the image degradation information, and these features are used to guide the training of the quality evaluation network. Finally, the quality prediction score of a single underwater image is obtained. A large number of experiments on the underwater image database and the underwater enhanced image database have proved the effectiveness of the method of the present invention. In addition, the adaptive features related to the image degradation information can be automatically extracted, saving time and effort.
[0031] 3) The method of the present invention extracts the features related to the image degradation information through the trained underwater image degradation information discovery network, i.e., the image enhancement network, and uses this part of the features for the training of the quality evaluation network. The experimental results show that compared with the state-of-the-art no-reference underwater image quality evaluation method, the method of the present invention has a higher consistency with the perceived quality of underwater images by humans in evaluating the quality of underwater images.
[0032] 4) When training the overall network, in order to ensure that the underwater image degradation information discovery network, i.e., the image enhancement network, can extract the degradation information of the underwater image as accurately as possible, when the underwater image degradation information discovery network, i.e., the image enhancement network, is trained, its parameters are frozen and then added to the training of the quality evaluation network. The experimental results show that compared with the state-of-the-art no-reference underwater image quality evaluation method, the method of the present invention has a higher consistency with the perceived quality of underwater images by humans in evaluating the quality of underwater images. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is the overall implementation block diagram of the method of the present invention;
[0034] Figure 2Schematic diagram of the composition structure of the image enhancement network constructed by the method of the present invention;
[0035] Figure 3 Schematic diagram of the composition structure of the quality evaluation network constructed by the method of the present invention. Detailed implementation manners
[0036] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments.
[0037] An underwater image quality evaluation method based on degradation information guidance proposed by the present invention has an overall implementation block diagram as Figure 1 shown, and it includes the following steps:
[0038] Step 1: Construct two neural networks. The first neural network serves as an image enhancement network, and the second neural network serves as a quality evaluation network.
[0039] As Figure 2 shown, the image enhancement network includes 1 first convolutional block, 4 second convolutional blocks, 4 third convolutional blocks, 4 fourth convolutional blocks, and 1 fifth convolutional block. The first convolutional block, the 1st second convolutional block, the 2nd second convolutional block, the 3rd second convolutional block, and the 4th second convolutional block form an encoding network. The 1st third convolutional block, the 2nd third convolutional block, the 3rd third convolutional block, and the 4th third convolutional block form a channel attention module. The 1st fourth convolutional block, the 2nd fourth convolutional block, the 3rd fourth convolutional block, the 4th fourth convolutional block, and the fifth convolutional block form a decoding network. The input channel number of the first convolutional block is 3, and the output channel number is 32. The input end of the first convolutional block serves as the input end of the image enhancement network and simultaneously receives the R, G, and B channels of an RGB image with a size of H×W (height is H and width is W). Denote the feature map with a size of output from the output end of the first convolutional block as F E1 ; the input channel number of the 1st second convolutional block is 32, and the output channel number is 32. The input end of the 1st second convolutional block receives F E1 , and denote the feature map with a size of output from the output end of the 1st second convolutional block as F E2 ; the input channel number of the 2nd second convolutional block is 32, and the output channel number is 64. The input end of the 2nd second convolutional block receives F E2 , and denote the feature map with a size of output from the output end of the 2nd second convolutional block as F E3 ; the input channel number of the 3rd second convolutional block is 64, and the output channel number is 128. The input end of the 3rd second convolutional block receives F E3 , and denote the feature map with a size of output from the output end of the 3rd second convolutional block as FE4 ; The number of input channels of the 4th second convolutional block is 128, and the number of output channels is 256. The input end of the 4th second convolutional block receives F E4 , and the feature map with a size of output from the output end of the 4th second convolutional block is denoted as F E5 ; The number of input channels of the 1st third convolutional block is 32, and the number of output channels is 32. The input end of the 1st third convolutional block receives F E2 , and the feature map with a size of output from the output end of the 1st third convolutional block is denoted as F C1 ; The number of input channels of the 2nd third convolutional block is 64, and the number of output channels is 64. The input end of the 2nd third convolutional block receives F E3 , and the feature map with a size of output from the output end of the 2nd third convolutional block is denoted as F C2 ; The number of input channels of the 3rd third convolutional block is 128, and the number of output channels is 128. The input end of the 3rd third convolutional block receives F E4 , and the feature map with a size of output from the output end of the 3rd third convolutional block is denoted as F C3 ; The number of input channels of the 4th third convolutional block is 256, and the number of output channels is 256. The input end of the 4th third convolutional block receives F E5 , and the feature map with a size of output from the output end of the 4th third convolutional block is denoted as F C4 ; The number of input channels of the 1st fourth convolutional block is 256, and the number of output channels is 256. The input end of the 1st fourth convolutional block receives F E5 , and the feature map with a size of output from the output end of the 1st fourth convolutional block is denoted as F D1 ; The number of input channels of the 2nd fourth convolutional block is 512, and the number of output channels is 128. The input end of the 2nd fourth convolutional block receives the feature map F with a size of D1 after performing a concatenation operation on F C4 and F , and the feature map with a size of DC1 output from the output end of the 2nd fourth convolutional block is denoted as F , and at the same time, F D2 is used as the first intermediate layer feature map; The number of input channels of the 3rd fourth convolutional block is 256, and the number of output channels is 64. The input end of the 3rd fourth convolutional block receives the feature map F with a size of DC1 after performing a concatenation operation on F D2 and F C3 with a size of , and the feature map F DC2, denote the feature map with a size of output from the output end of the third fourth convolutional block as F D3 , and at the same time, regard F DC2 as the second intermediate layer feature map; the number of input channels of the fourth fourth convolutional block is 128, and the number of output channels is 32. The input end of the fourth fourth convolutional block receives the feature map F D3 and F C2 after concatenation operation, with a size of denoted as F DC3 . Denote the feature map with a size of output from the output end of the fourth fourth convolutional block as F D4 , and at the same time, regard F DC3 as the third intermediate layer feature map; the number of input channels of the fifth convolutional block is 64, and the number of output channels is 3. The input end of the fifth convolutional block receives the feature map F D4 and F C1 after concatenation operation, with a size of denoted as F DC4 . Denote the feature map with a size of H×W×3 output from the output end of the fifth convolutional block as F D5 , and at the same time, regard F DC4 as the fourth intermediate layer feature map, regard F D5 as the image degradation information corresponding to the RGB image, perform element-wise addition on the RGB image and its corresponding image degradation information, and regard the image obtained by element-wise addition as the enhanced result image output from the output end of the image enhancement network.
[0040] As Figure 3 shown, the quality evaluation network includes 1 sixth convolutional block, 4 seventh convolutional blocks, 12 eighth convolutional blocks, 4 ninth convolutional blocks, 4 global average pooling models, and 1 fully connected layer. The encoding network is composed of 1 sixth convolutional block, 4 seventh convolutional blocks, and 12 eighth convolutional blocks. The feature fusion module is composed of 4 ninth convolutional blocks. The regression network is composed of 4 global average pooling models and 1 fully connected layer; the number of input channels of the sixth convolutional block is 3, and the number of output channels is 64. The input end of the sixth convolutional block simultaneously receives the R, G, and B channels of an RGB image with a size of H×W. The RGB image received by the input end of the sixth convolutional block is the same as the RGB image received by the input end of the first convolutional block. Denote the feature map with a size of output from the output end of the sixth convolutional block as F Q1 ; the number of input channels of the first seventh convolutional block is 64, and the number of output channels is 256. The input end of the first seventh convolutional block receives F Q1 . Denote the feature map with a size of output from the output end of the first seventh convolutional block as F Q2; The number of input channels of the first eighth convolutional block is 256, and the number of output channels is 256. The input end of the first eighth convolutional block receives F Q2 , and the feature map with a size of output from the output end of the first eighth convolutional block is denoted as F Q3 ; The number of input channels of the second eighth convolutional block is 256, and the number of output channels is 256. The input end of the second eighth convolutional block receives F Q3 , and the feature map with a size of output from the output end of the second eighth convolutional block is denoted as F Q4 , and F Q4 is used as the fifth intermediate layer feature map; The number of input channels of the second seventh convolutional block is 256, and the number of output channels is 512. The input end of the second seventh convolutional block receives F Q4 , and the feature map with a size of output from the output end of the second seventh convolutional block is denoted as F Q5 ; The number of input channels of the third eighth convolutional block is 512, and the number of output channels is 512. The input end of the third eighth convolutional block receives F Q5 , and the feature map with a size of output from the output end of the third eighth convolutional block is denoted as F Q6 ; The number of input channels of the fourth eighth convolutional block is 512, and the number of output channels is 512. The input end of the fourth eighth convolutional block receives F Q6 , and the feature map with a size of output from the output end of the fourth eighth convolutional block is denoted as F Q7 ; The number of input channels of the fifth eighth convolutional block is 512, and the number of output channels is 512. The input end of the fifth eighth convolutional block receives F Q7 , and the feature map with a size of output from the output end of the fifth eighth convolutional block is denoted as F Q8 , and F Q8 is used as the sixth intermediate layer feature map; The number of input channels of the third seventh convolutional block is 512, and the number of output channels is 1024. The input end of the third seventh convolutional block receives F Q8 , and the feature map with a size of output from the output end of the third seventh convolutional block is denoted as F Q9 ; The number of input channels of the sixth eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the sixth eighth convolutional block receives F Q9 , and the feature map with a size of output from the output end of the sixth eighth convolutional block is denoted as F Q10 ; The number of input channels of the seventh eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the seventh eighth convolutional block receives FQ10 , denote the feature map with a size of output from the output end of the 7th eighth convolution block as F Q11 ; the input channel number of the 8th eighth convolution block is 1024, the output channel number is 1024, and the input end of the 8th eighth convolution block receives F Q11 , denote the feature map with a size of output from the output end of the 8th eighth convolution block as F Q12 ; the input channel number of the 9th eighth convolution block is 1024, the output channel number is 1024, and the input end of the 9th eighth convolution block receives F Q12 , denote the feature map with a size of output from the output end of the 9th eighth convolution block as F Q13 ; the input channel number of the 10th eighth convolution block is 1024, the output channel number is 1024, and the input end of the 10th eighth convolution block receives F Q13 , denote the feature map with a size of output from the output end of the 10th eighth convolution block as F Q14 , and regard F Q14 as the seventh intermediate layer feature map; the input channel number of the 4th seventh convolution block is 1024, the output channel number is 2048, and the input end of the 4th seventh convolution block receives F Q14 , denote the feature map with a size of output from the output end of the 4th seventh convolution block as F Q15 ; the input channel number of the 11th eighth convolution block is 2048, the output channel number is 2048, and the input end of the 11th eighth convolution block receives F Q15 , denote the feature map with a size of output from the output end of the 11th eighth convolution block as F Q16 ; the input channel number of the 12th eighth convolution block is 2048, the output channel number is 2048, and the input end of the 12th eighth convolution block receives F Q16 , denote the feature map with a size of output from the output end of the 12th eighth convolution block as F Q17 , and regard F Q17 as the eighth intermediate layer feature map; the input channel number of the first input end of the 1st ninth convolution block is 256, the input channel number of the second input end is 64, the output channel number is 64, the first input end of the 1st ninth convolution block receives F Q4 , the second input end of the 1st ninth convolution block receives F DC4 , denote the feature map with a size of output from the output end of the 1st ninth convolution block as F DQ1; The number of input channels of the first input end of the second ninth convolution block is 512, the number of input channels of the second input end is 128, and the number of output channels is 128. The first input end of the second ninth convolution block receives F Q8 , the second input end of the second ninth convolution block receives F DC3 , and the feature map with the size of output from the output end of the second ninth convolution block is denoted as F DQ2 ; The number of input channels of the first input end of the third ninth convolution block is 1024, the number of input channels of the second input end is 256, and the number of output channels is 256. The first input end of the third ninth convolution block receives F Q14 , the second input end of the third ninth convolution block receives F DC2 , and the feature map with the size of output from the output end of the third ninth convolution block is denoted as F DQ3 ; The number of input channels of the first input end of the fourth ninth convolution block is 2048, the number of input channels of the second input end is 512, and the number of output channels is 512. The first input end of the fourth ninth convolution block receives F Q17 , the second input end of the fourth ninth convolution block receives F DC1 , and the feature map with the size of output from the output end of the fourth ninth convolution block is denoted as F DQ4 ; The number of input channels of the first global average pooling model is 64, and the number of output channels is 64. The input end of the first global average pooling model receives F DQ1 , and the output end of the first global average pooling model outputs a feature vector with the size of 1×1×64; The number of input channels of the second global average pooling model is 128, and the number of output channels is 128. The input end of the second global average pooling model receives F DQ2 , and the output end of the second global average pooling model outputs a feature vector with the size of 1×1×128; The number of input channels of the third global average pooling model is 256, and the number of output channels is 256. The input end of the third global average pooling model receives F DQ3 , and the output end of the third global average pooling model outputs a feature vector with the size of 1×1×256; The number of input channels of the fourth global average pooling model is 512, and the number of output channels is 512. The input end of the fourth global average pooling model receives F DQ4 , and the output end of the fourth global average pooling model outputs a feature vector with the size of 1×1×512; The feature vectors with the sizes of 1×1×64, 1×1×128, 1×1×256, and 1×1×512 are concatenated to obtain a feature vector with the size of 1×1×960, which is denoted as F iqa1; The number of input channels of the fully connected layer is 960, and the number of output channels is 1. The input end of the fully connected layer receives F iqa1 , and the output end of the fully connected layer outputs a value, which represents the quality prediction score of the RGB image.
[0041] Step 2: Select N1 original underwater images in different scenes and the corresponding pseudo-label images of each original underwater image to form a first training set; where N1 ≥ 800, and the sizes of the original underwater images and the pseudo-label images are H×W, that is, the heights of the original underwater images and the pseudo-label images are H and the widths are W.
[0042] Step 3: Input the R, G, and B channels of each original underwater image in the first training set into the image enhancement network for training. The image enhancement network outputs the enhanced result image corresponding to each original underwater image in the first training set. Then, for each pseudo-label image in the first training set, calculate the loss function value, denoted as Loss IE , where, represents the mean square error loss function value, "|| ||2" is the l2 norm operation symbol, I result represents the enhanced result image corresponding to the original underwater image I raw corresponding to the original underwater image I pseudo represents the original underwater image I raw corresponding to the pseudo-label image, represents the perceptual loss function value, represents the perceptual loss network, that is, the VGG-16 network (Visual Geometry Group Network). The VGG-16 network has a total of 16 convolutional layers, 1 ≤ j ≤ 16, represents the j-th convolutional layer in the VGG-16 network, represents I result input into the feature map output by the j-th convolutional layer in the VGG-16 network, represents I pseudo input into the feature map output by the j-th convolutional layer in the VGG-16 network, H j ×W j ×C j represents and of the size, H j represents and of the height, W j represents and of the width, C j represents and The number of channels.
[0043] Step 4: Use the first training set to train for more than 100 rounds according to the process of Step 3, finally train to obtain an image enhancement network training model, and freeze the parameters in subsequent training.
[0044] Step 5: Select N2 original underwater images in different scenarios; then use N3 existing different underwater image enhancement methods to enhance each original underwater image to obtain N3 underwater enhanced images corresponding to each original underwater image; then for the N3 underwater enhanced images corresponding to each original underwater image, arrange the N3 underwater enhanced images in a row, and combine each underwater enhanced image with each of the subsequent underwater enhanced images pairwise to form image pairs, obtaining a total of (N3 - 1) + (N3 - 2) + … + 1 pairs of image pairs; then form the second training set with N2 × ((N3 - 1) + (N3 - 2) + … + 1) pairs of image pairs; where N2 ≥ 100, N3 ≥ 10, and the size of the original underwater image is H × W, that is, the height of the original underwater image is H and the width is W.
[0045] Step 6: Input the R, G, and B channels of each underwater enhanced image in each pair of image pairs in the second training set into the image enhancement network training model at the same time. The image enhancement network training model outputs the first intermediate layer feature map, the second intermediate layer feature map, the third intermediate layer feature map, and the fourth intermediate layer feature map corresponding to each underwater enhanced image in each pair of image pairs in the second training set. Denote the first intermediate layer feature map, the second intermediate layer feature map, the third intermediate layer feature map, and the fourth intermediate layer feature map corresponding to any underwater enhanced image in any pair of image pairs in the second training set as F' DC1 、F' DC2 、F' DC3 、F' DC4 。
[0046] At the same time, input the R, G, and B channels of each underwater enhanced image in each pair of image pairs in the second training set into the encoding network of the quality evaluation network at the same time. The encoding network of the quality evaluation network outputs the fifth intermediate layer feature map, the sixth intermediate layer feature map, the seventh intermediate layer feature map, and the eighth intermediate layer feature map corresponding to each underwater enhanced image in each pair of image pairs in the second training set. Denote the fifth intermediate layer feature map, the sixth intermediate layer feature map, the seventh intermediate layer feature map, and the eighth intermediate layer feature map corresponding to any underwater enhanced image in any pair of image pairs in the second training set as F' Q4 、F' Q8 、F' Q14 、F' Q17; Then, for each pair of underwater enhanced images in the second training set, the fourth intermediate layer feature map and the fifth intermediate layer feature map, the third intermediate layer feature map and the sixth intermediate layer feature map, the second intermediate layer feature map and the seventh intermediate layer feature map, and the first intermediate layer feature map and the eighth intermediate layer feature map corresponding to each underwater enhanced image are input into the feature fusion module of the quality evaluation network in pairs. The quality evaluation network outputs the quality prediction scores corresponding to each underwater enhanced image in each pair of images in the second training set; then, for each pair of images in the second training set, the loss function value is calculated and denoted as Loss quality , Loss quality = max(0, -R × (Q1 - Q2) + margin), where max() is the function to take the maximum value, Q1 represents the quality prediction score corresponding to the first underwater enhanced image in each pair of images, Q2 represents the quality prediction score corresponding to the second underwater enhanced image in each pair of images, margin is a constant, margin = 0.5, and R represents the subjective preference value between the first underwater enhanced image and the second underwater enhanced image in each pair of images. If the first underwater enhanced image is subjectively preferred, then R = 1; if the second underwater enhanced image is subjectively preferred, then R = -1.
[0047] Step 7: Use the second training set to train for more than 100 rounds according to the process of Step 6, and finally train to obtain the quality evaluation network training model.
[0048] Step 8: Arbitrarily select an underwater image with a size of H×W and denote it as T; then input the R, G, and B channels of T into the image enhancement network training model and the quality evaluation network training model at the same time. The quality evaluation network training model outputs the quality prediction score of T.
[0049] In this embodiment, in Step 1, the first convolutional block is composed of a first convolutional layer and a first ReLU activation layer connected in sequence. The input end of the first convolutional layer is the input end of the first convolutional block. The input end of the first ReLU activation layer receives the feature map output from the output end of the first convolutional layer. The output end of the first ReLU activation layer is the output end of the first convolutional block. Among them, the number of input channels of the first convolutional layer is 3, the number of output channels is 32, and the convolutional kernel size is 1×1.
[0050] In this embodiment, in step 1, the second convolutional block consists of a second convolutional layer, a second ReLU activation layer, a third convolutional layer, a fourth convolutional layer, a third ReLU activation layer, and a fifth convolutional layer. The input end of the second convolutional layer is the input end of the second convolutional block where it is located. The input end of the second ReLU activation layer receives the feature map output by the output end of the second convolutional layer. The input end of the third convolutional layer receives the feature map output by the output end of the second ReLU activation layer. The input end of the fourth convolutional layer receives the feature map obtained by element-wise adding the feature map output by the output end of the third convolutional layer and the feature map received by the input end of the second convolutional layer. The input end of the third ReLU activation layer receives the feature map output by the output end of the fourth convolutional layer. The input end of the fifth convolutional layer receives the feature map output by the output end of the third ReLU activation layer. The output end of the second convolutional block outputs the feature map obtained by element-wise adding the feature map output by the output end of the fifth convolutional layer and the feature map received by the input end of the fourth convolutional layer; where the number of input channels, output channels, and kernel size of the second convolutional layer, third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the first second convolutional block are 32, 32, and 3×3 respectively, the number of input channels of the second convolutional layer in the second second convolutional block is 32, the number of output channels is 64, and the kernel size is 3×3, the number of input channels, output channels, and kernel size of the third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the second second convolutional block are 64, 64, and 3×3 respectively, the number of input channels of the second convolutional layer in the third second convolutional block is 64, the number of output channels is 128, and the kernel size is 3×3, the number of input channels, output channels, and kernel size of the third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the third second convolutional block are 128, 128, and 3×3 respectively, the number of input channels of the second convolutional layer in the fourth second convolutional block is 128, the number of output channels is 256, and the kernel size is 3×3, the number of input channels, output channels, and kernel size of the third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the fourth second convolutional block are 256, 256, and 3×3 respectively.
[0051] In this embodiment, in step 1, the third convolutional block consists of a first global average pooling layer, a sixth convolutional layer, a fourth ReLU activation layer, a seventh convolutional layer, and a first Sigmoid activation layer connected in sequence. The input end of the first global average pooling layer is the input end of the third convolutional block where it is located. The input end of the sixth convolutional layer receives the feature map output from the output end of the first global average pooling layer. The input end of the fourth ReLU activation layer receives the feature map output from the output end of the sixth convolutional layer. The input end of the seventh convolutional layer receives the feature map output from the output end of the fourth ReLU activation layer. The input end of the first Sigmoid activation layer receives the feature map output from the output end of the seventh convolutional layer. The output end of the first Sigmoid activation layer is the output end of the third convolutional block where it is located. Among them, the number of input channels of the sixth convolutional layer in the first third convolutional block is 32, the number of output channels is 2, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the first third convolutional block is 2, the number of output channels is 32, and the kernel size is 1×1. The number of input channels of the sixth convolutional layer in the second third convolutional block is 64, the number of output channels is 4, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the second third convolutional block is 4, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the sixth convolutional layer in the third third convolutional block is 128, the number of output channels is 8, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the third third convolutional block is 8, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the sixth convolutional layer in the fourth third convolutional block is 256, the number of output channels is 16, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the fourth third convolutional block is 16, the number of output channels is 256, and the kernel size is 1×1. The input size of the first global average pooling layer in the first third convolutional block is The output size is 1×1×32. The input size of the first global average pooling layer in the second third convolutional block is The output size is 1×1×64. The input size of the first global average pooling layer in the third third convolutional block is The output size is 1×1×128. The input size of the first global average pooling layer in the fourth third convolutional block is The output size is 1×1×256.
[0052] In this embodiment, in step 1, the fourth convolutional block is composed of an eighth convolutional layer, a first upsampling layer, a ninth convolutional layer, a fifth ReLU activation layer, a tenth convolutional layer, an eleventh convolutional layer, a sixth ReLU activation layer, and a twelfth convolutional layer. The input end of the eighth convolutional layer is the input end of the fourth convolutional block where it is located. The input end of the first upsampling layer receives the feature map output from the output end of the eighth convolutional layer. The input end of the ninth convolutional layer receives the feature map output from the output end of the first upsampling layer. The input end of the fifth ReLU activation layer receives the feature map output from the output end of the ninth convolutional layer. The input end of the tenth convolutional layer receives the feature map output from the output end of the fifth ReLU activation layer. The input end of the eleventh convolutional layer receives the feature map obtained by element-wise addition of the feature map output from the output end of the tenth convolutional layer and the feature map received by the input end of the ninth convolutional layer. The input end of the sixth ReLU activation layer receives the feature map output from the output end of the eleventh convolutional layer. The input end of the twelfth convolutional layer receives the feature map output from the output end of the sixth ReLU activation layer. The output end of the fourth convolutional block outputs the feature map obtained by element-wise addition of the feature map output from the output end of the twelfth convolutional layer and the feature map received by the input end of the tenth convolutional layer; among them, the number of input channels of the eighth convolutional layer in the first fourth convolutional block is 256, the number of output channels is 256, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the first fourth convolutional block are 256, 256, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the second fourth convolutional block is 512, the number of output channels is 128, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the second fourth convolutional block are 128, 128, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the third fourth convolutional block is 256, the number of output channels is 64, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the third fourth convolutional block are 64, 64, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the fourth fourth convolutional block is 128, the number of output channels is 32, and the convolution kernel size is 1×1. The number of input channels, output channels, and convolution kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the fourth fourth convolutional block are 32, 32, and 3×3 respectively. The sampling multiple of the first upsampling layer is 2 times.
[0053] In this embodiment, in step 1, the fifth convolutional block consists of a thirteenth convolutional layer and a seventh ReLU activation layer connected in sequence. The input end of the thirteenth convolutional layer is the input end of the fifth convolutional block. The input end of the seventh ReLU activation layer receives the feature map output from the output end of the thirteenth convolutional layer, and the output end of the seventh ReLU activation layer is the output end of the fifth convolutional block. Among them, the number of input channels of the thirteenth convolutional layer is 64, the number of output channels is 3, and the convolutional kernel size is 1×1.
[0054] In this embodiment, in step 1, the sixth convolutional block consists of a fourteenth convolutional layer and a first downsampling layer connected in sequence. The input end of the fourteenth convolutional layer is the input end of the sixth convolutional block. The input end of the first downsampling layer receives the feature map output from the output end of the fourteenth convolutional layer, and the output end of the first downsampling layer is the output end of the sixth convolutional block. Among them, the number of input channels of the fourteenth convolutional layer is 3, the number of output channels is 64, the convolutional kernel size is 7×7, and the sampling multiple of the first downsampling layer is 2 times.
[0055] In this embodiment, in step 1, the seventh convolutional block is composed of the fifteenth convolutional layer, the sixteenth convolutional layer, the seventeenth convolutional layer, and the eighteenth convolutional layer. The common connection end of the input end of the fifteenth convolutional layer and the input end of the eighteenth convolutional layer is the input end of the seventh convolutional block where it is located. The input end of the sixteenth convolutional layer receives the feature map output from the output end of the fifteenth convolutional layer. The input end of the seventeenth convolutional layer receives the feature map output from the output end of the sixteenth convolutional layer. The output end of the seventh convolutional block outputs the feature map obtained by element-wise addition of the feature map output from the output end of the seventeenth convolutional layer and the feature map output from the output end of the eighteenth convolutional layer. Among them, the number of input channels of the fifteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels and output channels of the seventeenth and eighteenth convolutional layers in the first seventh convolutional block are both 64 and 256 respectively, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the second seventh convolutional block is 256, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the second seventh convolutional block is 128, the number of output channels is 128, and the kernel size is 3×3. The number of input channels and output channels of the seventeenth convolutional layer in the second seventh convolutional block are both 128 and 512 respectively, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the second seventh convolutional block is 256, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the third seventh convolutional block is 512, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the third seventh convolutional block is 256, the number of output channels is 256, and the kernel size is 3×3. The number of input channels and output channels of the seventeenth convolutional layer in the third seventh convolutional block are both 256 and 1024 respectively, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the third seventh convolutional block is 512, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the fourth seventh convolutional block is 1024, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the fourth seventh convolutional block is 512, the number of output channels is 512, and the kernel size is 3×3. The number of input channels and output channels of the seventeenth convolutional layer in the fourth seventh convolutional block are both 512 and 2048 respectively, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the fourth seventh convolutional block is 1024, the number of output channels is 2048, and the kernel size is 1×1.
[0056] In this embodiment, in step 1, the eighth convolutional block consists of the nineteenth convolutional layer, the twentieth convolutional layer, and the twenty-first convolutional layer. The input end of the nineteenth convolutional layer is the input end of the eighth convolutional block where it is located. The input end of the twentieth convolutional layer receives the feature map output from the output end of the nineteenth convolutional layer. The input end of the twenty-first convolutional layer receives the feature map output from the output end of the twentieth convolutional layer. The output end of the eighth convolutional block outputs the feature map obtained by element-wise adding the feature map output from the output end of the twenty-first convolutional layer and the feature map received by the input end of the nineteenth convolutional layer; among them, the number of input channels of the nineteenth convolutional layer in the first and second eighth convolutional blocks is 256, the number of output channels is 64, and the convolutional kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the first and second eighth convolutional blocks is 64, the number of output channels is 64, and the convolutional kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the first and second eighth convolutional blocks is 64, the number of output channels is 256, and the convolutional kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the third to fifth eighth convolutional blocks is 512, the number of output channels is 128, and the convolutional kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the third to fifth eighth convolutional blocks is 128, the number of output channels is 128, and the convolutional kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the third to fifth eighth convolutional blocks is 128, the number of output channels is 512, and the convolutional kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the sixth to tenth eighth convolutional blocks is 1024, the number of output channels is 256, and the convolutional kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the sixth to tenth eighth convolutional blocks is 256, the number of output channels is 256, and the convolutional kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the sixth to tenth eighth convolutional blocks is 256, the number of output channels is 1024, and the convolutional kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the eleventh and twelfth eighth convolutional blocks is 2048, the number of output channels is 512, and the convolutional kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the eleventh and twelfth eighth convolutional blocks is 512, the number of output channels is 512, and the convolutional kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the eleventh and twelfth eighth convolutional blocks is 512, the number of output channels is 2048, and the convolutional kernel size is 1×1.
[0057] In this embodiment, in step 1, the ninth convolutional block is composed of the twenty-second convolutional layer, the first batch normalization (BatchNormalization, BN) layer, the eighth ReLU activation layer, the second upsampling layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer, the first max pooling layer, and the second Sigmoid activation layer. The input end of the twenty-second convolutional layer is the first input end of the ninth convolutional block where it is located. The common connection end of the input ends of the twenty-fourth convolutional layer and the twenty-sixth convolutional layer is the second input end of the ninth convolutional block where they are located. The input end of the first batch normalization layer receives the feature map output by the output end of the twenty-second convolutional layer. The input end of the eighth ReLU activation layer receives the feature map output by the output end of the first batch normalization layer. The input end of the second upsampling layer receives the feature map output by the output end of the eighth ReLU activation layer. The input end of the twenty-third convolutional layer receives the feature map output by the output end of the second upsampling layer. The input end of the twenty-fifth convolutional layer receives the feature map obtained by performing a concatenation operation on the feature map output by the output end of the twenty-third convolutional layer and the feature map output by the output end of the twenty-fourth convolutional layer. The input end of the first max pooling layer receives the feature map output by the output end of the twenty-sixth convolutional layer. The input end of the second Sigmoid activation layer receives the feature map output by the output end of the first max pooling layer. The feature map output by the output end of the twenty-fifth convolutional layer and the feature map output by the output end of the second Sigmoid activation layer are multiplied pixel by pixel. The feature map obtained by pixel-by-pixel multiplication and the feature map output by the output end of the twenty-fourth convolutional layer are added pixel by pixel. The feature map obtained by pixel-by-pixel addition is output by the output end of the ninth convolutional block. Among them, in the first ninth convolutional block, the number of input channels of the twenty-second convolutional layer is 256, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the first batch normalization layer is 64, and the number of output channels is 64. The sampling multiple of the second upsampling layer is 4 times. The number of input channels of the twenty-third convolutional layer is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-fourth convolutional layer is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-fifth convolutional layer is 128, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-sixth convolutional layer is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the first max pooling layer is 64, and the number of output channels is 1;In the second ninth convolutional block, the input channels of the twenty-second convolutional layer are 512, the output channels are 128, and the kernel size is 3×3; the input channels of the first batch normalization layer are 128, the output channels are 128; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 128, the output channels are 128, and the kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 128, the output channels are 128, and the kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 256, the output channels are 128, and the kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 128, the output channels are 128, and the kernel size is 3×3; the input channels of the first max pooling layer are 128, the output channels are 1; in the third ninth convolutional block, the input channels of the twenty-second convolutional layer are 1024, the output channels are 256, and the kernel size is 3×3; the input channels of the first batch normalization layer are 256, the output channels are 256; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 256, the output channels are 256, and the kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 256, the output channels are 256, and the kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 512, the output channels are 256, and the kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 256, the output channels are 256, and the kernel size is 3×3; the input channels of the first max pooling layer are 256, the output channels are 1; in the fourth ninth convolutional block, the input channels of the twenty-second convolutional layer are 2048, the output channels are 512, and the kernel size is 3×3; the input channels of the first batch normalization layer are 512, the output channels are 512; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 512, the output channels are 512, and the kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 512, the output channels are 512, and the kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 1024, the output channels are 512, and the kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 512, the output channels are 512, and the kernel size is 3×3; the input channels of the first max pooling layer are 512, the output channels are 1.
[0058] As described above, the concatenation operation, element-wise addition, and element-wise multiplication are all prior arts; the global average pooling model is a prior art, which is from Lin M, Chen Q, Yan S. Network In Network[J]. Computer Science, 2013. (Network in Network, Computer Science).
[0059] To further illustrate the feasibility and effectiveness of the method of the present invention, experiments are conducted on the method of the present invention.
[0060] Experiments are carried out to verify the performance of the method of the present invention on the Underwater Image Quality Assessment (UWIQA) and the Underwater Image Enhancement Benchmark (UIEB), and the obtained results are compared with those obtained by existing image quality assessment methods. Since there are no high-quality images available as reference standards under underwater conditions, the comparison methods selected for the experiments only include common reference-free image quality assessment methods.
[0061] The research of the experiment first compared the performance of the method of the present invention and existing reference-free image quality assessment methods on the underwater image database, which contains 890 underwater images with different resolutions. Considering the requirements of the neural network for input images, the method of the present invention unified all underwater images into a resolution of 224×224 for the experiment. Each underwater image has an artificial subjective evaluation score as the MOS value, which is used as the label for training. Secondly, the method of the present invention also conducted further experiments on the underwater enhanced image database, which also has 100 original underwater images with different resolutions. These 100 original underwater images are respectively enhanced using 10 different underwater image enhancement methods to obtain 1000 underwater enhanced images. Each underwater enhanced image has an artificial subjective evaluation score as the MOS value, which is used as the label for training.
[0062] To accurately compare the performance of the method of the present invention and existing reference-free image quality assessment methods, four commonly used performance evaluation indicators are selected in the experiment to evaluate the performance differences between the method of the present invention and existing reference-free image quality assessment methods. They are the Pearson linear correlation coefficient (PLCC), the Spearman rank correlation coefficient (SRCC), the Kendall rank correlation coefficient (KRCC), and the root mean square error (RMSE). In the experiment, when calculating the PLCC, a 5-parameter non-linear fitting function is used to map the quality prediction score to the subjective score. where f(x) is the subjective score, x is the quality prediction score, and τ i(i = 1, 2, …, 5) are fitting parameters, e is the natural base, e = 2.71…, the closer the values of SRCC and PLCC are to 1, and the closer the value of RMSE is to 0, indicating a higher consistency between the quality prediction score and the MOS value. KRCC is used to evaluate the monotonicity of the prediction. 80% of the images in the underwater image database and the underwater enhanced image database are respectively used for training, and the remaining 20% of the images are used for testing. There are no duplicate images in the training set and the test set. A total of 5 trainings are conducted and the average value is taken as the final results of PLCC, SRCC, KRCC, and RMSE.
[0063] Then, the method of the present invention was compared with existing no-reference image quality assessment methods on the underwater image database and the underwater enhanced image database. A total of 16 existing no-reference image quality assessment methods were selected for the comparison experiment, including 6 general no-reference image quality assessment methods, 5 evaluation models developed based on convolutional neural networks, and 5 no-reference image quality assessment methods for underwater images. The 6 general no-reference image quality assessment methods are DIIVINE (A.K. Moorthy and A.C. Bovik, “Blind image quality assessment: From natural scene statistics to perceptual quality,” IEEE Transactions on Image Processing, vol. 20, no. 12, pp. 3350–3364, 2011. (No-reference image quality assessment: From natural scene statistics to perceptual quality, IEEE Transactions on Image Processing)), BRISQUE (A. Mittal, A.K. Moorthy, and A.C. Bovik, “No-reference image quality assessment in the spatial domain,” IEEE Transactions on Image Processing, vol. 21, no. 12, pp. 4695–4708, 2012. (Spatial domain no-reference image quality assessment, IEEE Transactions on Image Processing)), GM-LOG (W. Xue, X. Mou, L. Zhang, A.C. Bovik, and X. Feng, “Blind image quality assessment using joint statistics of gradient magnitude and Laplacian features,” IEEE Trans. Image Process., vol. 23, no. 11, pp. 4850–4862, Nov. 2014 (Using joint statistics of gradient magnitude and Laplacian features for blind image quality assessment, IEEE Transactions on Image Processing)), NFERM (K. Gu, G. Zhai, X. Yang, and W. Zhang, “Using free energy principle for blind image quality assessment,” IEEE Trans. Multimedia, vol. 17, no. 1, pp. 50–63, Jan.2015 ("Blind Image Quality Assessment Using Free Energy Principle," IEEE Transactions on Multimedia), NRSL (Q. Li, W. Lin, J. Xu, and Y. Fang, "Blind image quality assessment using statistical structural and luminance features," IEEE Trans. Multimedia, vol. 18, no. 12, pp. 2457–2469, Dec. 2016 ("Blind Image Quality Assessment Using Statistical Structural and Luminance Features," IEEE Transactions on Multimedia)), BIQME (K. Gu, D. Tao, J.-F. Qiao, and W. Lin, "Learning a no-reference quality assessment model of enhanced images with big data," IEEE Trans. Neural Netw. Learn. Syst., vol. 29, no. 4, pp. 1301–1313, Apr. 2018 ("Learning a No-Reference Quality Assessment Model of Enhanced Images with Big Data," IEEE Transactions on Neural Networks and Learning Systems)), these methods extract different image features through artificially designed feature extraction methods, and then fuse them through support vector machine (SVR) to obtain the final image quality. At the same time, considering the application of deep learning in the field of image quality assessment in recent years, the experiment also selected 5 methods that use convolutional neural network (CNN) to build image quality assessment models, namely CNN (L. Kang, P. Ye, Y. Li, and D. Doermann, "Convolutional neural networks for no-reference image quality assessment," in Proc. IEEE Conf. Comput. Vis. Pattern Recognit., pp. 1733–1740, Jun. 2014. ("Convolutional Neural Networks for No-Reference Image Quality Assessment," IEEE International Conference on Computer Vision and Pattern Recognition)), WaDIQaM (S. Bosse, D. Maniry, K.-R. Müller, T. Wiegand, and W. Samek, "Deep neural networks for no-reference and full-reference image quality assessment," IEEE Trans. Image Process., vol. 27, no. 1, pp.206–219, Jan. 2018 (Deep Neural Networks for No-Reference and Full-Reference Image Quality Assessment, IEEE Transactions on Image Processing), CaHDC (J. Wu, J. Ma, F. Liang, W. Dong, G. Shi and W. Lin, “End-to-End Blind Image Quality Prediction With Cascaded Deep Neural Network,” IEEE Trans. Image Process., vol. 29, pp. 7414-7426, 2020. (End-to-End Blind Image Quality Prediction With Cascaded Deep Neural Network, IEEE Transactions on Image Processing)), DBCNN (W. Zhang, K. Ma, J. Yan, D. Deng and Z. Wang, “Blind Image Quality Assessment Using a Deep Bilinear Convolutional Neural Network,” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, no. 1, pp. 36-47, Jan. 2020 (Blind Image Quality Assessment Using a Deep Bilinear Convolutional Neural Network, IEEE Transactions on Circuits and Systems for Video Technology)), VCRNet (Z. Pan, F. Yuan, J. Lei, Y. Fang, X. Shao and S. Kwong, “VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality Assessment,” IEEE Transactions on Image Processing, vol. 31, pp. 1613-1627, 2022. (VCRNet: Visual Compensation Restoration Network for No-Reference Image Quality Assessment, IEEE Transactions on Image Processing)). In addition, 5 no-reference image quality assessment methods for underwater images were selected for the experiment, namely UCIQE (M. Yang and A. Sowmya, “An underwater color image quality evaluation metric,” IEEE Transactions on Image Processing, vol. 24, no. 12, pp. 6062–6071, 2015. (An underwater color image quality evaluation metric, IEEE Transactions on Image Processing)), UIQM (K. Panetta, C.Gao, and S. Agaian, “Human-visual-system-inspired underwater image quality measures,” IEEE Journal of Oceanic Engineering, vol. 41, no. 3, pp. 541–551, 2016. (Based on human visual system for underwater image quality prediction, IEEE Journal of Oceanic Engineering), CCF (Y. Wang, N. Li, Z. Li, Z. Gu, H. Zheng, B. Zheng, M. Sun, An imaging-inspired no-reference underwater color image quality assessment metric, Comput. Electr. Eng. 70 (2017) 904–913. (An imaging-inspired no-reference underwater color image quality assessment method, Computers and Electrical Engineering), FDUM (N. Yang, Q. Zhong, K. Li, R. Cong, and S. Kwong, “A reference-free underwater image quality assessment metric in frequency domain,” Signal Processing Image Communication, vol. 94, no. 12, p. 116218, 2021. (A frequency-domain reference-free underwater image quality assessment method, Signal Processing Image Communication), NUIQ (Q. Jiang, Y. Gu, C. Li, R. Cong and F. Shao, "Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective Metric," IEEE Transactions on Circuits and Systems for Video Technology, 2022. (Underwater enhanced image quality assessment: Benchmark dataset and objective method, IEEE Transactions on Circuits and Systems for Video Technology).
[0064] Table 1 presents the performance comparison results of the method of the present invention and existing reference-free image quality assessment methods on the underwater image database.
[0065] Table 1 Performance comparison of the method of the present invention and existing reference-free image quality assessment methods on the underwater image database
[0066]
[0067] The following conclusions can be drawn from the data listed in Table 1. First, the evaluation model developed based on the convolutional neural network is superior in performance to the general no-reference image quality assessment method, which is reasonable. The application of deep learning in the field of image quality assessment is a very worthy research direction. Second, under the same training conditions, the no-reference image quality assessment method for underwater images is superior to the general no-reference image quality assessment methods, such as the FDUM and NUIQ methods, and the performance of the training-oriented method is significantly stronger than that of the untrained UCIQE and UIQM methods. Third, the effects of some image quality assessment methods based on the statistical characteristics of natural scenes are not particularly good on the underwater image database, such as the BRISQUE and NRSL methods. This may be because the statistical characteristics of natural scenes are not very applicable to underwater images and cannot fully represent the quality-related features of underwater images. Finally, compared with the existing no-reference image quality assessment methods, the method of the present invention has the best performance on the underwater image database. Its PLCC, SRCC, and KRCC values are all the highest, indicating that the method of the present invention has the highest correlation with the subjective scores of underwater images. This is because the method of the present invention fully considers the degradation characteristics of underwater images and guides the quality assessment process through the degradation information existing in the underwater images themselves.
[0068] In order to test whether the method of the present invention can be applied to other databases and verify the generality and robustness of the method of the present invention, a re-experiment was conducted on the underwater enhanced image database. Similarly, the method of the present invention was compared with 16 existing no-reference image quality assessment methods. The SRCC, KRCC, PLCC, and RMSE values of the experimental results are shown in Table 2.
[0069] Table 2 Performance comparison between the method of the present invention and existing no-reference image quality assessment methods on the underwater enhanced image database
[0070]
[0071] As can be seen from the results in Table 2, the method of the present invention still achieved relatively good results on the 1000 underwater enhanced images in the underwater enhanced image database, and the SRCC, KRCC, and PLCC are all the highest. Generally speaking, the method of the present invention is also effective on underwater enhanced images and has good database independence.
Claims
1. An underwater image quality evaluation method guided by degradation information, characterized in that It includes the following steps: Step 1: Construct two neural networks. The first neural network serves as an image enhancement network, and the second neural network serves as a quality evaluation network. The image enhancement network includes one first convolutional block, four second convolutional blocks, four third convolutional blocks, four fourth convolutional blocks, and one fifth convolutional block. The first convolutional block, the first second convolutional block, the second second convolutional block, the third second convolutional block, and the fourth second convolutional block form an encoding network. The first third convolutional block, the second third convolutional block, the third third convolutional block, and the fourth third convolutional block form a channel attention module. The first fourth convolutional block, the second fourth convolutional block, the third fourth convolutional block, the fourth fourth convolutional block, and the fifth convolutional block form a decoding network. The input channel number of the first convolutional block is 3, and the output channel number is 32. The input end of the first convolutional block serves as the input end of the image enhancement network and simultaneously receives the R, G, and B channels of an RGB image with a size of H×W. Denote the feature map with a size of output from the output end of the first convolutional block as F E1 ; The input channel number of the first second convolutional block is 32, and the output channel number is 32. The input end of the first second convolutional block receives F E1 , and denote the feature map with a size of output from the output end of the first second convolutional block as F E2 ; The input channel number of the second second convolutional block is 32, and the output channel number is 64. The input end of the second second convolutional block receives F E2 , and denote the feature map with a size of output from the output end of the second second convolutional block as F E3 ; The input channel number of the third second convolutional block is 64, and the output channel number is 128. The input end of the third second convolutional block receives F E3 , and denote the feature map with a size of output from the output end of the third second convolutional block as F E4 ; The input channel number of the fourth second convolutional block is 128, and the output channel number is 256. The input end of the fourth second convolutional block receives F E4 , and denote the feature map with a size of output from the output end of the fourth second convolutional block as F E5 ; The input channel number of the first third convolutional block is 32, and the output channel number is 32. The input end of the first third convolutional block receives F E2 , and denote the feature map with a size of output from the output end of the first third convolutional block as F C1 ; The input channel number of the second third convolutional block is 64, and the output channel number is 64. The input end of the second third convolutional block receives F E3 , and denote the feature map with a size of output from the output end of the second third convolutional block as F C2 ; The number of input channels of the third third convolutional block is 128, and the number of output channels is 128. The input end of the third third convolutional block receives F E4 , and the feature map with a size of output from the output end of the third third convolutional block is denoted as F C3 ; The number of input channels of the fourth third convolutional block is 256, and the number of output channels is 256. The input end of the fourth third convolutional block receives F E5 , and the feature map with a size of output from the output end of the fourth third convolutional block is denoted as F C4 ; The number of input channels of the first fourth convolutional block is 256, and the number of output channels is 256. The input end of the first fourth convolutional block receives F E5 , and the feature map with a size of output from the output end of the first fourth convolutional block is denoted as F D1 ; The number of input channels of the second fourth convolutional block is 512, and the number of output channels is 128. The input end of the second fourth convolutional block receives the feature map F D1 and F C4 obtained after a concatenation operation, with a size of , and the feature map with a size of DC1 output from the output end of the second fourth convolutional block is denoted as F , and at the same time, F D2 is used as the first intermediate layer feature map; The number of input channels of the third fourth convolutional block is 256, and the number of output channels is 64. The input end of the third fourth convolutional block receives the feature map F DC1 and F D2 obtained after a concatenation operation, with a size of C3 and F , and the feature map with a size of DC2 output from the output end of the third fourth convolutional block is denoted as F , and at the same time, F D3 is used as the second intermediate layer feature map; The number of input channels of the fourth fourth convolutional block is 128, and the number of output channels is 32. The input end of the fourth fourth convolutional block receives the feature map F DC2 obtained after a concatenation operation, with a size of D3 and F C2 and F , and the feature map with a size of DC3 output from the output end of the fourth fourth convolutional block is denoted as F , and at the same time, F D4 is used as the third intermediate layer feature map; The number of input channels of the fifth convolutional block is 64, and the number of output channels is 3. The input end of the fifth convolutional block receives the feature map F DC3 and F D4 and F C1 The size obtained after the splicing operation is feature map F DC4 , and the feature map with a size of H×W×3 output from the output end of the fifth convolutional block is denoted as F D5 . At the same time, F DC4 is used as the fourth intermediate layer feature map, and F D5 is used as the image degradation information corresponding to the RGB image. The RGB image and its corresponding image degradation information are added element by element, and the image obtained by the element-by-element addition is used as the enhanced result image output from the output end of the image enhancement network; The quality evaluation network includes one sixth convolutional block, four seventh convolutional blocks, twelve eighth convolutional blocks, four ninth convolutional blocks, four global average pooling models, and one fully connected layer. The encoding network is composed of one sixth convolutional block, four seventh convolutional blocks, and twelve eighth convolutional blocks. The feature fusion module is composed of four ninth convolutional blocks. The regression network is composed of four global average pooling models and one fully connected layer. The input channel number of the sixth convolutional block is 3, and the output channel number is 64. The input end of the sixth convolutional block simultaneously receives the R, G, and B channels of an RGB image with a size of H×W. The RGB image received at the input end of the sixth convolutional block is the same as the RGB image received at the input end of the first convolutional block. Denote the feature map with a size of output from the output end of the sixth convolutional block as F Q1 ; The input channel number of the first seventh convolutional block is 64, and the output channel number is 256. The input end of the first seventh convolutional block receives F Q1 , and denote the feature map with a size of output from the output end of the first seventh convolutional block as F Q2 ; The input channel number of the first eighth convolutional block is 256, and the output channel number is 256. The input end of the first eighth convolutional block receives F Q2 , and denote the feature map with a size of output from the output end of the first eighth convolutional block as F Q3 ; The input channel number of the second eighth convolutional block is 256, and the output channel number is 256. The input end of the second eighth convolutional block receives F Q3 , and denote the feature map with a size of output from the output end of the second eighth convolutional block as F Q4 , and use F Q4 as the fifth intermediate layer feature map; The input channel number of the second seventh convolutional block is 256, and the output channel number is 512. The input end of the second seventh convolutional block receives F Q4 , and denote the feature map with a size of output from the output end of the second seventh convolutional block as F Q5 ; The input channel number of the third eighth convolutional block is 512, and the output channel number is 512. The input end of the third eighth convolutional block receives F Q5 , and denote the feature map with a size of output from the output end of the third eighth convolutional block as F Q6 ; The input channel number of the fourth eighth convolutional block is 512, and the output channel number is 512. The input end of the fourth eighth convolutional block receives F Q6 , and denote the feature map with a size of output from the output end of the fourth eighth convolutional block as F Q7 ; The number of input channels of the 5th eighth convolutional block is 512, and the number of output channels is 512. The input end of the 5th eighth convolutional block receives F Q7 , and the feature map with a size of output from the output end of the 5th eighth convolutional block is denoted as F Q8 , and F Q8 is used as the sixth intermediate layer feature map; The number of input channels of the 3rd seventh convolutional block is 512, and the number of output channels is 1024. The input end of the 3rd seventh convolutional block receives F Q8 , and the feature map with a size of output from the output end of the 3rd seventh convolutional block is denoted as F Q9 ; The number of input channels of the 6th eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the 6th eighth convolutional block receives F Q9 , and the feature map with a size of output from the output end of the 6th eighth convolutional block is denoted as F Q10 ; The number of input channels of the 7th eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the 7th eighth convolutional block receives F Q10 , and the feature map with a size of output from the output end of the 7th eighth convolutional block is denoted as F Q11 ; The number of input channels of the 8th eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the 8th eighth convolutional block receives F Q11 , and the feature map with a size of output from the output end of the 8th eighth convolutional block is denoted as F Q12 ; The number of input channels of the 9th eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the 9th eighth convolutional block receives F Q12 , and the feature map with a size of output from the output end of the 9th eighth convolutional block is denoted as F Q13 ; The number of input channels of the 10th eighth convolutional block is 1024, and the number of output channels is 1024. The input end of the 10th eighth convolutional block receives F Q13 , and the feature map with a size of output from the output end of the 10th eighth convolutional block is denoted as F Q14 , and F Q14 is used as the seventh intermediate layer feature map; The number of input channels of the 4th seventh convolutional block is 1024, and the number of output channels is 2048. The input end of the 4th seventh convolutional block receives F Q14 , and the feature map with a size of output from the output end of the 4th seventh convolutional block is denoted as F Q15 ; The number of input channels of the 11th eighth convolutional block is 2048, and the number of output channels is 2048. The input end of the 11th eighth convolutional block receives F Q15 , and the feature map with a size of output from the output end of the 11th eighth convolutional block is denoted as F Q16 ; The number of input channels of the 12th eighth convolutional block is 2048, and the number of output channels is 2048. The input end of the 12th eighth convolutional block receives F Q16 , and the feature map with a size of output from the output end of the 12th eighth convolutional block is denoted as F Q17 , and F Q17 is used as the eighth intermediate layer feature map; The number of input channels of the first input end of the 1st ninth convolutional block is 256, the number of input channels of the second input end is 64, and the number of output channels is 64. The first input end of the 1st ninth convolutional block receives F Q4 , and the second input end of the 1st ninth convolutional block receives F DC4 , and the feature map with a size of output from the output end of the 1st ninth convolutional block is denoted as F DQ1 ; The number of input channels of the first input end of the 2nd ninth convolutional block is 512, the number of input channels of the second input end is 128, and the number of output channels is 128. The first input end of the 2nd ninth convolutional block receives F Q8 , and the second input end of the 2nd ninth convolutional block receives F DC3 , and the feature map with a size of output from the output end of the 2nd ninth convolutional block is denoted as F DQ2 ; The number of input channels of the first input end of the 3rd ninth convolutional block is 1024, the number of input channels of the second input end is 256, and the number of output channels is 256. The first input end of the 3rd ninth convolutional block receives F Q14 , and the second input end of the 3rd ninth convolutional block receives F DC2 , and the feature map with a size of output from the output end of the 3rd ninth convolutional block is denoted as F DQ3 ; The number of input channels of the first input end of the 4th ninth convolutional block is 2048, the number of input channels of the second input end is 512, and the number of output channels is 512. The first input end of the 4th ninth convolutional block receives F Q17 , and the second input end of the 4th ninth convolutional block receives F DC1 , and the feature map with a size of output from the output end of the 4th ninth convolutional block is denoted as F DQ4 ; The number of input channels of the 1st global average pooling model is 64, and the number of output channels is 64. The input end of the 1st global average pooling model receives F DQ1 , the output end of the first global average pooling model outputs a feature vector with a size of 1×1×64; the input channel number of the second global average pooling model is 128 and the output channel number is 128. The input end of the second global average pooling model receives F DQ2 , the output end of the second global average pooling model outputs a feature vector with a size of 1×1×128; the input channel number of the third global average pooling model is 256 and the output channel number is 256. The input end of the third global average pooling model receives F DQ3 , the output end of the third global average pooling model outputs a feature vector with a size of 1×1×256; the input channel number of the fourth global average pooling model is 512 and the output channel number is 512. The input end of the fourth global average pooling model receives F DQ4 , the output end of the fourth global average pooling model outputs a feature vector with a size of 1×1×512; perform a concatenation operation on the feature vector with a size of 1×1×64, the feature vector with a size of 1×1×128, the feature vector with a size of 1×1×256, and the feature vector with a size of 1×1×512 to obtain a feature vector with a size of 1×1×960, denoted as F iqa1 ; the input channel number of the fully connected layer is 960 and the output channel number is 1. The input end of the fully connected layer receives F iqa1 , the output end of the fully connected layer outputs a value, and this value represents the quality prediction score of the RGB image; Step 2: Select N1 original underwater images in different scenes and the corresponding pseudo-label images for each original underwater image to form a first training set. Among them, N1≥800, and the sizes of the original underwater images and the pseudo-label images are H×W, that is, the heights of the original underwater images and the pseudo-label images are H and the widths are W. Step 3: Input the R, G, and B channels of each original underwater image in the first training set into the image enhancement network for training. The image enhancement network outputs the enhanced result image corresponding to each original underwater image in the first training set. Then, for each pseudo-label image in the first training set, calculate the loss function value, denoted as Loss IE , where represents the mean squared error loss function value, "|| ||2” is the l2 norm operation symbol, I result represents the original underwater image I raw corresponding enhanced result image, I pseudo represents the original underwater image I raw corresponding pseudo-label image, represents the perceptual loss function value, represents the perceptual loss network, i.e., the VGG-16 network, 1 ≤ j ≤ 16, represents the j-th convolutional layer in the VGG-16 network, represents I result input into the feature map output by the j-th convolutional layer in the VGG-16 network, represents I pseudo input into the feature map output by the j-th convolutional layer in the VGG-16 network, H j ×W j ×C j represents and 's size, H j represents and 's height, W j represents and 's width, C j represents and 's number of channels; Step 4: Use the first training set to train for more than 100 rounds according to the process of Step 3. Finally, an image enhancement network training model is obtained, and the parameters are frozen in subsequent training. Step 5: Select N2 original underwater images in different scenes; then use N3 different underwater image enhancement methods to enhance each original underwater image to obtain N3 underwater enhanced images corresponding to each original underwater image; then for the N3 underwater enhanced images corresponding to each original underwater image, arrange the N3 underwater enhanced images in a row, and combine each underwater enhanced image with each of the subsequent underwater enhanced images in pairs to obtain a total of (N3 - 1)+(N3 - 2)+…+1 pairs of image pairs; then form a second training set with N2×((N3 - 1)+(N3 - 2)+…+1) pairs of image pairs. Among them, N2≥100, N3≥10, and the size of the original underwater image is H×W, that is, the height of the original underwater image is H and the width is W. Step 6: Input the R, G, and B channels of each underwater enhanced image in each pair of images in the second training set into the image enhancement network training model simultaneously. The image enhancement network training model outputs the first intermediate layer feature map, the second intermediate layer feature map, the third intermediate layer feature map, and the fourth intermediate layer feature map corresponding to each underwater enhanced image in each pair of images in the second training set. Denote the first intermediate layer feature map, the second intermediate layer feature map, the third intermediate layer feature map, and the fourth intermediate layer feature map corresponding to any underwater enhanced image in any pair of images in the second training set as F' DC1 、F' DC2 、F' DC3 、F' DC4 ; Meanwhile, the R, G, and B channels of each underwater enhanced image in each pair of image pairs in the second training set are simultaneously input into the encoding network of the quality evaluation network. The encoding network of the quality evaluation network outputs the fifth intermediate layer feature map, the sixth intermediate layer feature map, the seventh intermediate layer feature map, and the eighth intermediate layer feature map corresponding to each underwater enhanced image in each pair of image pairs in the second training set. The fifth intermediate layer feature map, the sixth intermediate layer feature map, the seventh intermediate layer feature map, and the eighth intermediate layer feature map corresponding to any one underwater enhanced image in any pair of image pairs in the second training set are denoted as F' Q4 , F' Q8 , F' Q14 , F' Q17 ; Then, the fourth intermediate layer feature map and the fifth intermediate layer feature map, the third intermediate layer feature map and the sixth intermediate layer feature map, the second intermediate layer feature map and the seventh intermediate layer feature map, and the first intermediate layer feature map and the eighth intermediate layer feature map corresponding to each underwater enhanced image in each pair of image pairs in the second training set are input into the feature fusion module of the quality evaluation network in pairs. The quality evaluation network outputs the quality prediction scores corresponding to each underwater enhanced image in each pair of image pairs in the second training set; Then, for each pair of image pairs in the second training set, calculate the loss function value, denoted as Loss quality , Loss quality = max(0, -R×(Q1 - Q2) + margin), where max() is the maximum value function, Q1 represents the quality prediction score corresponding to the first underwater enhanced image in each pair of image pairs, Q2 represents the quality prediction score corresponding to the second underwater enhanced image in each pair of image pairs, margin is a constant, margin = 0.5, and R represents the subjective preference value between the first underwater enhanced image and the second underwater enhanced image in each pair of image pairs. If the first underwater enhanced image is subjectively preferred, then R = 1; if the second underwater enhanced image is subjectively preferred, then R = -1; Step 7: Use the second training set to train for more than 100 rounds according to the process of Step 6. Finally, a quality evaluation network training model is obtained. Step 8: Arbitrarily select an underwater image with a size of H×W and denote it as T; then input the R, G, and B channels of T into the image enhancement network training model and the quality evaluation network training model at the same time, and the quality evaluation network training model outputs the quality prediction score of T.
2. The underwater image quality evaluation method based on degradation information guidance according to claim 1, wherein In the said Step 1, the first convolutional block is composed of a first convolutional layer and a first ReLU activation layer connected in sequence. The input end of the first convolutional layer is the input end of the first convolutional block. The input end of the first ReLU activation layer receives the feature map output from the output end of the first convolutional layer. The output end of the first ReLU activation layer is the output end of the first convolutional block. Among them, the number of input channels of the first convolutional layer is 3, the number of output channels is 32, and the convolution kernel size is 1×1. The second convolutional block consists of a second convolutional layer, a second ReLU activation layer, a third convolutional layer, a fourth convolutional layer, a third ReLU activation layer, and a fifth convolutional layer. The input end of the second convolutional layer is the input end of the second convolutional block where it is located. The input end of the second ReLU activation layer receives the feature map output by the output end of the second convolutional layer. The input end of the third convolutional layer receives the feature map output by the output end of the second ReLU activation layer. The input end of the fourth convolutional layer receives the feature map obtained by element-wise addition of the feature map output by the output end of the third convolutional layer and the feature map received by the input end of the second convolutional layer. The input end of the third ReLU activation layer receives the feature map output by the output end of the fourth convolutional layer. The input end of the fifth convolutional layer receives the feature map output by the output end of the third ReLU activation layer. The output end of the second convolutional block outputs the feature map obtained by element-wise addition of the feature map output by the output end of the fifth convolutional layer and the feature map received by the input end of the fourth convolutional layer; among them, the number of input channels, output channels, and kernel size of the second convolutional layer, third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the first second convolutional block are 32, 32, and 3×3 respectively. The number of input channels of the second convolutional layer in the second second convolutional block is 32, the number of output channels is 64, and the kernel size is 3×3. The number of input channels, output channels, and kernel size of the third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the second second convolutional block are 64, 64, and 3×3 respectively. The number of input channels of the second convolutional layer in the third second convolutional block is 64, the number of output channels is 128, and the kernel size is 3×3. The number of input channels, output channels, and kernel size of the third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the third second convolutional block are 128, 128, and 3×3 respectively. The number of input channels of the second convolutional layer in the fourth second convolutional block is 128, the number of output channels is 256, and the kernel size is 3×3. The number of input channels, output channels, and kernel size of the third convolutional layer, fourth convolutional layer, and fifth convolutional layer in the fourth second convolutional block are 256, 256, and 3×3 respectively.
3. The underwater image quality evaluation method based on degradation information guidance according to claim 1, characterized in that In the aforementioned Step 1, the third convolutional block consists of a first global average pooling layer, a sixth convolutional layer, a fourth ReLU activation layer, a seventh convolutional layer, and a first Sigmoid activation layer connected in sequence. The input end of the first global average pooling layer is the input end of the third convolutional block where it is located. The input end of the sixth convolutional layer receives the feature map output from the output end of the first global average pooling layer. The input end of the fourth ReLU activation layer receives the feature map output from the output end of the sixth convolutional layer. The input end of the seventh convolutional layer receives the feature map output from the output end of the fourth ReLU activation layer. The input end of the first Sigmoid activation layer receives the feature map output from the output end of the seventh convolutional layer. The output end of the first Sigmoid activation layer is the output end of the third convolutional block where it is located. Among them, the number of input channels of the sixth convolutional layer in the first third convolutional block is 32, the number of output channels is 2, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the first third convolutional block is 2, the number of output channels is 32, and the kernel size is 1×1. The number of input channels of the sixth convolutional layer in the second third convolutional block is 64, the number of output channels is 4, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the second third convolutional block is 4, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the sixth convolutional layer in the third third convolutional block is 128, the number of output channels is 8, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the third third convolutional block is 8, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the sixth convolutional layer in the fourth third convolutional block is 256, the number of output channels is 16, and the kernel size is 1×1. The number of input channels of the seventh convolutional layer in the fourth third convolutional block is 16, the number of output channels is 256, and the kernel size is 1×1. The input size of the first global average pooling layer in the first third convolutional block is The output size is 1×1×32. The input size of the first global average pooling layer in the second third convolutional block is The output size is 1×1×64. The input size of the first global average pooling layer in the third third convolutional block is The output size is 1×1×128. The input size of the first global average pooling layer in the fourth third convolutional block is The output size is 1×1×256.
4. The underwater image quality evaluation method based on degradation information guidance according to claim 1, wherein In the said step 1, the fourth convolutional block consists of an eighth convolutional layer, a first upsampling layer, a ninth convolutional layer, a fifth ReLU activation layer, a tenth convolutional layer, an eleventh convolutional layer, a sixth ReLU activation layer, and a twelfth convolutional layer. The input end of the eighth convolutional layer is the input end of the fourth convolutional block where it is located. The input end of the first upsampling layer receives the feature map output from the output end of the eighth convolutional layer. The input end of the ninth convolutional layer receives the feature map output from the output end of the first upsampling layer. The input end of the fifth ReLU activation layer receives the feature map output from the output end of the ninth convolutional layer. The input end of the tenth convolutional layer receives the feature map output from the output end of the fifth ReLU activation layer. The input end of the eleventh convolutional layer receives the feature map obtained by element-wise addition of the feature map output from the output end of the tenth convolutional layer and the feature map received by the input end of the ninth convolutional layer. The input end of the sixth ReLU activation layer receives the feature map output from the output end of the eleventh convolutional layer. The input end of the twelfth convolutional layer receives the feature map output from the output end of the sixth ReLU activation layer. The output end of the fourth convolutional block outputs the feature map obtained by element-wise addition of the feature map output from the output end of the twelfth convolutional layer and the feature map received by the input end of the tenth convolutional layer. Among them, the number of input channels of the eighth convolutional layer in the first fourth convolutional block is 256, the number of output channels is 256, and the kernel size is 1×1. The number of input channels, output channels, and kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the first fourth convolutional block are 256, 256, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the second fourth convolutional block is 512, the number of output channels is 128, and the kernel size is 1×1. The number of input channels, output channels, and kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the second fourth convolutional block are 128, 128, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the third fourth convolutional block is 256, the number of output channels is 64, and the kernel size is 1×1. The number of input channels, output channels, and kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the third fourth convolutional block are 64, 64, and 3×3 respectively. The number of input channels of the eighth convolutional layer in the fourth fourth convolutional block is 128, the number of output channels is 32, and the kernel size is 1×1. The number of input channels, output channels, and kernel size of the ninth, tenth, eleventh, and twelfth convolutional layers in the fourth fourth convolutional block are 32, 32, and 3×3 respectively. The sampling multiple of the first upsampling layer is 2 times; The fifth convolutional block consists of a thirteenth convolutional layer and a seventh ReLU activation layer connected in sequence. The input end of the thirteenth convolutional layer is the input end of the fifth convolutional block. The input end of the seventh ReLU activation layer receives the feature map output from the output end of the thirteenth convolutional layer, and the output end of the seventh ReLU activation layer is the output end of the fifth convolutional block. Among them, the number of input channels of the thirteenth convolutional layer is 64, the number of output channels is 3, and the convolutional kernel size is 1×1.
5. The underwater image quality evaluation method based on degradation information guidance according to claim 1, wherein In the step 1 described above, the sixth convolutional block consists of a fourteenth convolutional layer and a first downsampling layer connected in sequence. The input end of the fourteenth convolutional layer is the input end of the sixth convolutional block. The input end of the first downsampling layer receives the feature map output from the output end of the fourteenth convolutional layer, and the output end of the first downsampling layer is the output end of the sixth convolutional block. Among them, the number of input channels of the fourteenth convolutional layer is 3, the number of output channels is 64, the convolutional kernel size is 7×7, and the downsampling multiple of the first downsampling layer is 2 times. The seventh convolutional block consists of the fifteenth convolutional layer, the sixteenth convolutional layer, the seventeenth convolutional layer, and the eighteenth convolutional layer. The common connection end of the input ends of the fifteenth convolutional layer and the eighteenth convolutional layer is the input end of the seventh convolutional block where it is located. The input end of the sixteenth convolutional layer receives the feature map output from the output end of the fifteenth convolutional layer, and the input end of the seventeenth convolutional layer receives the feature map output from the output end of the sixteenth convolutional layer. The output end of the seventh convolutional block outputs the feature map obtained by element-wise adding the feature map output from the output end of the seventeenth convolutional layer and the feature map output from the output end of the eighteenth convolutional layer; among them, the number of input channels of the fifteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the first seventh convolutional block is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels and output channels of the seventeenth and eighteenth convolutional layers in the first seventh convolutional block are both 64 and 256 respectively, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the second seventh convolutional block is 256, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the second seventh convolutional block is 128, the number of output channels is 128, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer in the second seventh convolutional block is 128, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the second seventh convolutional block is 256, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the third seventh convolutional block is 512, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the third seventh convolutional block is 256, the number of output channels is 256, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer in the third seventh convolutional block is 256, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the third seventh convolutional block is 512, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the fifteenth convolutional layer in the fourth seventh convolutional block is 1024, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the sixteenth convolutional layer in the fourth seventh convolutional block is 512, the number of output channels is 512, and the kernel size is 3×3. The number of input channels of the seventeenth convolutional layer in the fourth seventh convolutional block is 512, the number of output channels is 2048, and the kernel size is 1×1. The number of input channels of the eighteenth convolutional layer in the fourth seventh convolutional block is 1024, the number of output channels is 2048, and the kernel size is 1×1; The eighth convolutional block consists of the nineteenth convolutional layer, the twentieth convolutional layer, and the twenty-first convolutional layer. The input end of the nineteenth convolutional layer is the input end of the eighth convolutional block where it is located. The input end of the twentieth convolutional layer receives the feature map output from the output end of the nineteenth convolutional layer. The input end of the twenty-first convolutional layer receives the feature map output from the output end of the twentieth convolutional layer. The output end of the eighth convolutional block outputs the feature map obtained by element-wise adding the feature map output from the output end of the twenty-first convolutional layer and the feature map received by the input end of the nineteenth convolutional layer; among them, the number of input channels of the nineteenth convolutional layer in the 1st and 2nd eighth convolutional blocks is 256, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the 1st and 2nd eighth convolutional blocks is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the 1st and 2nd eighth convolutional blocks is 64, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the 3rd to 5th eighth convolutional blocks is 512, the number of output channels is 128, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the 3rd to 5th eighth convolutional blocks is 128, the number of output channels is 128, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the 3rd to 5th eighth convolutional blocks is 128, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the 6th to 10th eighth convolutional blocks is 1024, the number of output channels is 256, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the 6th to 10th eighth convolutional blocks is 256, the number of output channels is 256, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the 6th to 10th eighth convolutional blocks is 256, the number of output channels is 1024, and the kernel size is 1×1. The number of input channels of the nineteenth convolutional layer in the 11th and 12th eighth convolutional blocks is 2048, the number of output channels is 512, and the kernel size is 1×1. The number of input channels of the twentieth convolutional layer in the 11th and 12th eighth convolutional blocks is 512, the number of output channels is 512, and the kernel size is 3×3. The number of input channels of the twenty-first convolutional layer in the 11th and 12th eighth convolutional blocks is 512, the number of output channels is 2048, and the kernel size is 1×1.
6. The underwater image quality evaluation method based on degradation information guidance according to claim 1, wherein In the aforementioned Step 1, the ninth convolutional block consists of the twenty-second convolutional layer, the first batch normalization layer, the eighth ReLU activation layer, the second upsampling layer, the twenty-third convolutional layer, the twenty-fourth convolutional layer, the twenty-fifth convolutional layer, the twenty-sixth convolutional layer, the first max pooling layer, and the second Sigmoid activation layer. The input end of the twenty-second convolutional layer is the first input end of the ninth convolutional block where it is located. The common connection end of the input ends of the twenty-fourth convolutional layer and the twenty-sixth convolutional layer is the second input end of the ninth convolutional block where they are located. The input end of the first batch normalization layer receives the feature map output from the output end of the twenty-second convolutional layer. The input end of the eighth ReLU activation layer receives the feature map output from the output end of the first batch normalization layer. The input end of the second upsampling layer receives the feature map output from the output end of the eighth ReLU activation layer. The input end of the twenty-third convolutional layer receives the feature map output from the output end of the second upsampling layer. The input end of the twenty-fifth convolutional layer receives the feature map obtained by concatenating the feature map output from the output end of the twenty-third convolutional layer and the feature map output from the output end of the twenty-fourth convolutional layer. The input end of the first max pooling layer receives the feature map output from the output end of the twenty-sixth convolutional layer. The input end of the second Sigmoid activation layer receives the feature map output from the output end of the first max pooling layer. The feature map output from the output end of the twenty-fifth convolutional layer and the feature map output from the output end of the second Sigmoid activation layer are multiplied pixel by pixel. The feature map obtained by pixel-by-pixel multiplication and the feature map output from the output end of the twenty-fourth convolutional layer are added pixel by pixel. The feature map obtained by pixel-by-pixel addition is output from the output end of the ninth convolutional block. Among them, in the first ninth convolutional block, the number of input channels of the twenty-second convolutional layer is 256, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the first batch normalization layer is 64, and the number of output channels is 64. The sampling multiple of the second upsampling layer is 4 times. The number of input channels of the twenty-third convolutional layer is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-fourth convolutional layer is 64, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-fifth convolutional layer is 128, the number of output channels is 64, and the kernel size is 1×1. The number of input channels of the twenty-sixth convolutional layer is 64, the number of output channels is 64, and the kernel size is 3×3. The number of input channels of the first max pooling layer is 64, and the number of output channels is 1;In the second ninth convolutional block, the input channels of the twenty-second convolutional layer are 512, the output channels are 128, and the kernel size is 3×3; the input channels of the first batch normalization layer are 128, the output channels are 128; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 128, the output channels are 128, and the kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 128, the output channels are 128, and the kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 256, the output channels are 128, and the kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 128, the output channels are 128, and the kernel size is 3×3; the input channels of the first max pooling layer are 128, the output channels are 1; in the third ninth convolutional block, the input channels of the twenty-second convolutional layer are 1024, the output channels are 256, and the kernel size is 3×3; the input channels of the first batch normalization layer are 256, the output channels are 256; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 256, the output channels are 256, and the kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 256, the output channels are 256, and the kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 512, the output channels are 256, and the kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 256, the output channels are 256, and the kernel size is 3×3; the input channels of the first max pooling layer are 256, the output channels are 1; in the fourth ninth convolutional block, the input channels of the twenty-second convolutional layer are 2048, the output channels are 512, and the kernel size is 3×3; the input channels of the first batch normalization layer are 512, the output channels are 512; the sampling multiple of the second upsampling layer is 4 times; the input channels of the twenty-third convolutional layer are 512, the output channels are 512, and the kernel size is 1×1; the input channels of the twenty-fourth convolutional layer are 512, the output channels are 512, and the kernel size is 1×1; the input channels of the twenty-fifth convolutional layer are 1024, the output channels are 512, and the kernel size is 1×1; the input channels of the twenty-sixth convolutional layer are 512, the output channels are 512, and the kernel size is 3×3; the input channels of the first max pooling layer are 512, the output channels are 1.;
Citation Information
Patent Citations
Asymmetric multi-modal fusion saliency detection method based on attention mechanism
CN111563418A
Image processing method and device, model training method and device, medium and equipment
CN111797855A