A method for underwater image quality assessment based on hierarchical feature fusion
By extracting the full-level features of underwater images and combining the parameter density model, a reference-free underwater image quality evaluation method is constructed, which solves the problem that the existing methods cannot be applied to the non-reference situation, and achieves the simplified decision-making process and the quality evaluation effect applicable to underwater tasks.
Patent Information
- Application Number
- CN202310396890.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-04-14
AI Technical Summary
The existing underwater image quality evaluation methods have problems such as not being applicable to no reference situations and requiring a large amount of database training. The quality evaluation methods of natural scene images are not applicable to underwater images.
The underwater image quality evaluation method based on hierarchical feature fusion is adopted. By extracting the full-level features of the underwater image, including low-level transformation domain information, medium-level contour information and high-level semantic information, combined with the parameter density model to capture distortion, a reference-free underwater image evaluation method is constructed.
It realizes underwater image quality evaluation without reference, does not require a large amount of data for training, simplifies the decision-making process, and is suitable for tasks such as underwater object detection and recognition.
Smart Images

Figure CN116486245B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of underwater image quality assessment, and in particular to an underwater image quality assessment method based on hierarchical feature fusion. Background Art
[0002] Underwater images are widely used in the fields of marine resource development, environmental monitoring, submarine archaeology, underwater robots, etc. However, due to the complexity of the underwater environment, the difficulty of underwater image acquisition, and the limited bandwidth of the underwater acoustic channel, the underwater images obtained at the receiving end often have complex distortions. Therefore, it is necessary to evaluate the quality of underwater images obtained at the receiving end. Accurate evaluation of underwater image quality is of great significance to ensure the successful completion of underwater tasks and the accuracy of data analysis.
[0003] At present, quality evaluation is mainly divided into subjective quality evaluation and objective quality evaluation. Subjective quality evaluation depends on the perception of the image by the human eye, while objective quality evaluation depends on the calculation of the image by the mathematical model. Subjective quality evaluation is time-consuming and expensive, and cannot be directly embedded in the actual system as an optimization indicator. Therefore, objective quality evaluation is often used in scientific research. There are large differences in visual features and statistical characteristics between underwater images and natural scene images, and underwater images are often task-oriented. Underwater image quality evaluation is generally based on tasks. Therefore, the quality evaluation method of natural scene images is not applicable to underwater images. For underwater images, the commonly used evaluation methods are peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). The principle of these two methods is to compare the pixel information and structural differences between the original image and the image to be tested to evaluate the quality of the image. These two methods require reference information, but underwater images often do not have original images to refer to in real scenes. In recent years, with the rise of deep learning, deep learning has also been applied to the field of image quality evaluation. The underwater image quality assessment method based on deep learning does not require the information of the original image. It only needs to extract the underwater image features and construct the mapping relationship between the underwater image and the quality. One of the drawbacks of the deep learning-based method is that it requires a large-scale database, and the existing database often cannot meet this requirement. In summary, there is still room for improvement in the existing underwater image quality assessment method. Summary of the invention
[0004] The purpose of the present invention is to solve the existing problems and provide an underwater image quality evaluation method based on hierarchical feature fusion. The method takes into account the task-oriented background of underwater images, namely underwater target detection and recognition, and combines the visual recognition principle of the human brain to extract the hierarchical features of underwater images. Several parameter density models are used to effectively capture distortion, and a reference-free underwater image evaluation method is comprehensively proposed. This method does not require a large amount of data for training, making the decision process simpler.
[0005] To achieve the above object, the technical solution of the present invention is: an underwater image quality evaluation method based on hierarchical feature fusion, comprising the following steps:
[0006] S1. Propose a full-level feature extraction method for underwater images;
[0007] S2. Propose a multi-feature fusion scheme and construct a quality evaluation method based on full-level features.
[0008] In one embodiment of the present invention, the step S1 specifically includes the following steps:
[0009] S11. Extract image information from a low level, analyze image clarity, and use underwater images to obtain information in the transform domain to evaluate image quality. Specifically, the local information volume law of the transform domain coefficients of underwater images is statistically analyzed. First, the transform domain coefficient matrix D of the m×m image block is calculated, and the transform domain coefficients are normalized to obtain a spectral probability map:
[0010]
[0011] Where 1<i≤m, 1<j≤m;
[0012] Secondly, the amount of information of each image block is calculated as:
[0013]
[0014] Each image is matched with an information matrix; Rayleigh distribution or Gaussian distribution f(x 1 ; θ) is fitted, where x 1 Represents the amount of local information in the transform domain to be fitted, θ represents the distribution parameter, and θ 2 As the clarity feature f 1 :
[0015] f 1 =θ 2
[0016] S12, extracting mid-level information from the image, using the peripheral inhibition of the non-classical receptive field and combining it with the classical edge detection method to obtain contour information; extracting the gradient amplitude M δ (x, y), and use the Gaussian difference to calculate the suppression weight W δ (x, y); Since the edges of underwater images are not rich, only the effect of distance on peripheral suppression is considered; isotropic suppression term t δ (x, y) is defined as the convolution of the gradient magnitude and the suppression weight, and the final contour operator C δ The calculation of (x, y) is:
[0017] Cδ (x, y) = max{[M δ (x, y)-γt δ (x, y)], 0}
[0018] Among them, γ is the influencing factor used to control the contour information;
[0019] Then calculate the MSCN coefficient of the suppressed image MSCN coefficient is defined as follows:
[0020]
[0021] Among them, I(i, j) corresponds to the intensity of the central pixel, μ(i, j) and σ(i, j) correspond to the mean and standard deviation of the current local area respectively, and C = 1 is a constant to prevent the denominator from being zero. According to the histogram, the generalized Gaussian distribution is selected to capture the statistical law of the underwater image coefficients. The definition of the generalized Gaussian distribution is as follows:
[0022]
[0023]
[0024] where Γ(·) is the gamma function, x 2 Corresponding to the MSCN coefficient, α corresponds to the mean, δ 2 Corresponding to the variance, take (α, δ 2 ) as the second set of features; since the mean and variance are independent of each other, they are combined into feature f 2 :
[0025]
[0026] S13, extract high-level semantic information from the image, extract the fully connected layer from the deep learning model and map it into a score, select three feature information of entropy, kurtosis and skewness to balance the amount of information and feature space dimension; since these three features are of the same dimension, they are directly combined together as the third set of features f 3 :
[0027] f 3 =entropy+skewness+kurtosis
[0028]
[0029]
[0030]
[0031] Among them, entropy is entropy, skewness is skewness, kurtosis is kurtosis; n is the dimension of the fully connected layer, P(i, j) is the probability of each value, x 3 and x 4 The eigenvectors corresponding to skewness and kurtosis respectively, α and δ correspond to mean and standard deviation, and E(·) represents the mean.
[0032] In one embodiment of the present invention, a Gaussian kernel function is used to blur an image, construct an image scale space, and obtain a multi-resolution image by multiple downsampling. The Gaussian kernel function is defined as follows:
[0033]
[0034] Among them, m′ is the center of the kernel function, ||mm′|| is the Euclidean distance between vector m and vector m′, and ε controls the range of the Gaussian kernel function. The optimal multi-scale fusion quality index is determined by nonlinearly combining multi-scale image features.
[0035]
[0036] Among them, s is the scale parameter, c is a constant, and f 1i 、f 2i and f 3i The corresponding feature is f 1 、f 2 and f 3 Values at different scale parameters.
[0037] Compared with the prior art, the present invention has the following beneficial effects: the method of the present invention takes into account the task-oriented background of underwater images, namely underwater target detection and recognition, combines the visual recognition principle of the human brain, extracts the hierarchical features of underwater images, and uses several parameter density models to effectively capture distortion, and comprehensively proposes a reference-free underwater image evaluation method. This method does not require a large amount of data for training, making the decision process simpler. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The present invention provides a model framework for an underwater image quality assessment method based on hierarchical feature fusion. DETAILED DESCRIPTION
[0039] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings.
[0040] Please refer to Figure 1 The present invention provides an underwater image quality evaluation method based on hierarchical feature fusion. In this example, sonar images are used as images to be tested. Figure 1 (a), specifically comprising the following steps:
[0041] Step S1, proposing a full-level feature extraction method for underwater images;
[0042] Step S2: propose a multi-feature fusion scheme and construct a quality evaluation method based on full-level features.
[0043] In this embodiment, the step S1 is specifically as follows:
[0044] Step S11, extract image information from the low level, analyze the clarity of the image, and use the underwater image to obtain information in the transform domain to evaluate the image quality. Specifically, the local information quantity law of the transform domain coefficients of the underwater image is statistically analyzed. First, the transform domain coefficient matrix D of the m×m image block is calculated, and the transform domain coefficients are normalized to obtain a spectral probability map:
[0045]
[0046] Where 1<i≤m, 1<j≤m. In this example, m=8. Then calculate the information amount of each block as:
[0047]
[0048] Each image is matched with an information matrix. Since most sonar images have black backgrounds, the frequency of 0 values will be very high when drawing the histogram, and this part does not contain useful information. Therefore, when fitting the histogram, the frequency of 0 values needs to be set to 0. The local information matrix of the transform domain coefficients is calculated by using f(x 1 ; θ) is fitted, where x 1 Represents the amount of local information in the transform domain to be fitted, θ represents the distribution parameter, and θ 2 As the clarity feature f 1 :
[0049] f 1 =θ 2
[0050] In this example, if Figure 1 As shown, the information of the image in the transform domain is analyzed at different resolutions. Figure 1 As shown in (b), the local information of the extracted image is Figure 1 (c) shows that its frequency domain histogram is Figure 1 (d) shows a Rayleigh distribution.
[0051] Step S12: extracting mid-level information from the image, using the peripheral inhibition of the non-classical receptive field and combining it with the classical edge detection method, thereby obtaining contour information. Extracting the gradient amplitude M δ (x, y), and use the Gaussian difference to calculate the suppression weight W δ(x, y). Since the edges of sonar images are not rich, only the effect of distance on peripheral suppression is considered. Isotropic suppression term t δ (x, y) is defined as the convolution of the gradient magnitude and the suppression weight, and the final contour operator C δ The calculation of (x, y) is:
[0052] C δ (x, y) = max{[M δ (x, y)-γt δ (x, y)], 0}
[0053] Among them, γ is used to control the influencing factor of contour information, and γ=1 is set by adjusting the parameter. Then calculate the MSCN coefficient of the suppressed image MSCN coefficient is defined as follows:
[0054]
[0055] Among them, I(i, j) corresponds to the intensity of the central pixel, μ(i, j) and σ(i, j) correspond to the mean and standard deviation of the current local area respectively, and C=1 is a constant to prevent the denominator from being zero. According to the histogram, the generalized Gaussian distribution is selected to capture the statistical law of the sonar image coefficients. The definition of the generalized Gaussian distribution is as follows:
[0056]
[0057]
[0058] where Γ(·) is the gamma function, x 2 Corresponding to the MSCN coefficient, α corresponds to the mean, δ 2 Corresponding to the variance, take (α, δ 2 ) as the second set of features. Since the mean and variance are independent of each other, they are combined into feature f 2 :
[0059]
[0060] In this example, the contour information extraction process is as follows Figure 1 As shown, Figure 1 (e) corresponds to the extracted gradient magnitude map, Figure 1 (f) corresponds to the suppressed contour map. Its contour probability density distribution is as follows Figure 1 As shown in (g), it presents a generalized Gaussian distribution.
[0061] Step S13: extract high-level semantic information from the image, extract the fully connected layer from the VGG16 model and map it into a score, and select three feature information, entropy, kurtosis, and skewness, to balance the amount of information and feature space dimension. Since these three features are of the same dimension, they are directly combined together as the third set of features f 3 :
[0062] f 3 =entropy+skewness+kurtosis
[0063] Among them, entropy is entropy, skewness is skewness, and kurtosis is kurtosis.
[0064] in
[0065]
[0066]
[0067]
[0068] Where n is the dimension of the fully connected layer, P(i, j) is the probability of each value, and x 3 and x 4 The eigenvectors corresponding to skewness and kurtosis respectively, α and δ correspond to mean and standard deviation, and E(·) represents the mean.
[0069] In this example, if Figure 1 As shown in (h), in different scale spaces, the selected deep learning network is the VGG16 network, which is used to extract high-level semantic information of the image and expand the extracted features into a feature vector by column as follows: Figure 1 As shown in (i), we finally select three features: entropy, kurtosis, and skewness, and their statistical distribution is as follows: Figure 1 (j) as shown.
[0070] In this embodiment, step S2 is specifically: using a Gaussian kernel function to blur the image, construct an image scale space, and obtain a multi-resolution image by multiple downsampling. The Gaussian kernel function is defined as follows:
[0071]
[0072] Where m′ is the center of the kernel function, ||mm′|| is the Euclidean distance between vector m and vector m′, and σ controls the range of the Gaussian kernel function. The optimal multi-scale fusion quality index is determined by nonlinearly combining multi-scale image features. Semantic information serves as a guide so that the index changes with the content:
[0073]
[0074] Where s is the scale parameter. After ablation experiments, it is found that the performance is optimal when s=3. c is a constant. By adjusting the parameters, c=0.01 is set. In this example, the scores obtained in step S11 are f 11 =0.2227, f 12 =0.2033, f 13 =0.1218; the scores obtained in step S12 are f 21 =2.4906, f 22 =2.4323, f 23 =2.4364; the scores obtained in step S13 are f 31 =14.4278, f 32 =14.5761,f 33 =14.7736. Finally, the above results are combined nonlinearly to get the final score of 0.5270.
[0075] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A method for underwater image quality assessment based on hierarchical feature fusion, characterized in that: The steps include: S1. A full-level feature extraction method for underwater images is proposed; specifically, the following steps are included: S11. Extract image information from a low level, analyze image clarity, and use underwater images to obtain information in the transform domain to evaluate image quality. Specifically, the local information volume law of the transform domain coefficients of underwater images is statistically analyzed. First, the transform domain coefficient matrix D of the m×m image block is calculated, and the transform domain coefficients are normalized to obtain a spectral probability map: 1 of them <i≤m,1<j≤m; Secondly, the amount of information of each image block is calculated as: Each image is matched with an information matrix; the local information matrix of the transform domain coefficient is fitted using Rayleigh distribution or Gaussian distribution f(x1;θ), where x1 represents the local information of the transform domain to be fitted, θ represents the distribution parameter, and θ is taken 2 As the clarity feature f1: f1=θ 2 S12, extracting mid-level information from the image, using the peripheral inhibition of the non-classical receptive field and combining it with the classical edge detection method to obtain contour information; extracting the gradient amplitude M δ (x, y), and use the Gaussian difference to calculate the suppression weight W δ (x, y); Since the edges of underwater images are not rich, only the effect of distance on peripheral suppression is considered; isotropic suppression term t δ (x, y) is defined as the convolution of the gradient magnitude and the suppression weight, and the final contour operator C δ The calculation of (x,y) is: C δ (x,y)=max{[M δ (x,y)-γt δ (x,y)],0} Among them, γ is the influencing factor used to control the contour information; Then calculate the MSCN coefficient of the suppressed image MSCN coefficient is defined as follows: Among them, I(i,j) corresponds to the intensity of the central pixel, μ(i,j) and σ(i,j) correspond to the mean and standard deviation of the current local area respectively, and C=1 is a constant to prevent the denominator from being zero. According to the histogram, the generalized Gaussian distribution is selected to capture the statistical law of the underwater image coefficients. The definition of the generalized Gaussian distribution is as follows: where Γ(·) is the gamma function, x2 corresponds to the MSCN coefficient, α corresponds to the mean, and δ 2 Corresponding to the variance, take (α,δ 2 ) as the second set of features; since the mean and variance are independent of each other, they are combined into feature f2: S13, extract high-level semantic information from the image, extract the fully connected layer from the deep learning model and map it into a score, select entropy, kurtosis, and skewness to balance the amount of information and feature space dimension; since these three features are of the same dimension, they are directly combined together as the third set of features f3: f3=entropy+skewness+kurtosis Where entropy is entropy, skewnes is skewness, kurtosis is kurtosis; n is the dimension of the fully connected layer, P(i,j) is the probability of each value, x3 and x4 correspond to the eigenvectors of skewness and kurtosis respectively, α and δ correspond to mean and standard deviation, and E(·) means to find the mean; S2. Propose a multi-feature fusion scheme and construct a quality evaluation method based on full-level features; The step S2 is specifically: using a Gaussian kernel function to blur the image, construct an image scale space, and obtain a multi-resolution image by multiple downsampling. The Gaussian kernel function is defined as follows: Among them, m' is the center of the kernel function, ||m-m'|| is the Euclidean distance between vector m and vector m', and ε controls the range of the Gaussian kernel function; the optimal multi-scale fusion quality index is determined by nonlinearly combining multi-scale image features; Among them, s is the scale parameter, c is a constant, and f 1i 、f 2i and f 3i The corresponding values are the features f1, f2 and f3 at different scale parameters.
Citation Information
Patent Citations
A full-reference and no-reference image quality evaluation method with a unified structure
CN109919920A
No-reference tone mapping image quality evaluation method based on multi-feature fusion
CN110046673A