Underwater image enhancement method for cross knowledge embedding learning

By combining deep learning and physics intersecting knowledge in underwater image enhancement, the cross-knowledge embedding model is designed, which solves the problem that existing technology is difficult to deal with complex underwater environments, and achieves better image quality and detail retention effects.

CN120125481AActive Publication Date: 2025-06-10HOHAI UNIV +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510180453.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-20
Filing Date
2025-02-19
Publication Date
2025-06-10
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods are difficult to deal with complex underwater environments, resulting in poor image quality, color decay, low contrast and blurred details.

Method used

Combining deep learning and physics cross-knowledge cross-knowledge embedding model, we design cross-knowledge embedding models, and obtain red channel, contrast and gradient knowledge by constructing underwater imaging models and inversion degradation models, and embed them into the encoder-decoder network, and optimize the model through cross-knowledge guidance module and gradient loss function.

Benefits of technology

Improves the enhancement ability of underwater images, improves color correction and detail retention effects, and significantly improves image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125481A_ABST
    Figure CN120125481A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater image enhancement method based on cross knowledge embedding learning, and the method comprises the steps: embedding the red channel knowledge, contrast knowledge and gradient knowledge of an underwater image into a deep learning model, and improving the adaptability of the deep learning model to an underwater optical scene based on the common guidance of a plurality of knowledge. The problems of wavelength selective attenuation and forward and backward scattering in the underwater imaging process are comprehensively solved, and the data quality of underwater images is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an underwater image enhancement method for cross - knowledge embedding learning, belonging to the technical field of image processing. Background Art

[0002] During the underwater imaging process, problems of attenuation and scattering often occur during the light propagation process, resulting in the decline of image quality. The attenuation characteristics of light in water are different from those on land. Compared with the best - absorbed red color, the absorption attenuation coefficients of short - wave blue and short - wave green are the smallest. Therefore, underwater images tend to be bluish - green. At the same time, there are large - molecule particulate matters in the water body. These particles have a reflection effect on light and will cause scattering during the imaging process. Therefore, underwater images generally present phenomena such as color decline, low contrast, and blurred details.

[0003] As a clear underwater image is crucial for underwater work, at first, the processing of underwater images started from the image itself. By means of image processing, the pixel points of the image were adjusted to eliminate image noise and improve image quality. This method is called an image enhancement method of a non - physical model. With the in - depth study of underwater imaging, considering the particularity of the underwater imaging process, many studies based on physical models began to invert the underwater degradation process, establish an underwater imaging degradation model, and restore the image to the non - degraded state by analyzing the model parameter information. With the development of deep learning, the underwater image enhancement method based on deep learning has become the mainstream direction. However, only considering improving the model structure is not enough. For underwater images, it is difficult for image enhancement methods to handle complex underwater environments. Summary of the Invention

[0004] Object of the Invention: Aiming at the problems and deficiencies in the prior art, the present invention combines deep learning with physical cross - knowledge, and adds physical prior knowledge guidance from network design to model training, and proposes a cross - knowledge embedding model. This method solves the problem that traditional underwater image enhancement methods are difficult to handle complex underwater environments and improves the underwater image enhancement ability.

[0005] Technical Solution: An underwater image enhancement method for cross - knowledge embedding learning includes the following steps:

[0006] S1. According to the process of underwater light propagation, construct an underwater imaging model and analyze the light attenuation mechanism of the target during the underwater imaging process;

[0007] S2. According to the light attenuation mechanism during the underwater imaging process, invert the underwater degradation model to obtain the red - channel knowledge, contrast knowledge, and gradient knowledge of the current underwater image;

[0008] S3. Design an encoder-decoder deep learning network with three types of cross-knowledge embeddings: red-channel knowledge, contrast knowledge, and gradient knowledge. Propose a cross-knowledge guidance module to embed cross-knowledge into the model, where red-channel knowledge is embedded into the encoder module, contrast knowledge is embedded into the decoder module, and gradient knowledge is used as the loss function to train the network model;

[0009] S4. Train the cross-knowledge embedding model designed in S3 through an underwater image database until convergence;

[0010] S5. Input the current original underwater image into the trained cross-knowledge embedding model and output the enhanced underwater image result.

[0011] Preferably, in step S1, an underwater imaging model is constructed according to the process of underwater light propagation, and the specific formula is:

[0012] I(x) = J(x)t(x) + A(1 - t(x)) (1)

[0013] where I(x) is the light reaching the imaging plane, J(x) is the direct component of the target reflected light reaching the imaging plane, A is the background light, t(x) is the transmission rate, and x is the pixel point.

[0014] According to the Lambert-Beer law, the light decreases exponentially with respect to the distance d(x) during transmission, β is the attenuation coefficient, and the specific formula for the transmission rate t(x) is:

[0015] t(x) = e -βd(x) (2). Preferably,

[0016] In step S2, according to the light attenuation mechanism that the red wave attenuates faster with the increase of distance during underwater imaging, an underwater red-channel degradation model is inverted, and the specific formula is:

[0017] 1 - I R = t(1 - J R ) + (1 - t)(1 - A R ) (3)

[0018] According to the red-channel degradation model, the red-channel knowledge of the underwater image is proposed, that is, the estimation of the transmission rate of the red channel of the clear image, and the specific formula is:

[0019]

[0020] where I R (y), IG(y), I B (y) are the intensities of the red, green, and blue channels of the underwater image respectively, and A CLet \(I_C\) be the intensity of the underwater background light for each color channel, where \(C\in\{R, G, B\}\) represents the red, green, and blue color channels. Let \(y\in\Omega(x)\) denote the pixel neighborhood of \(x\), \(\min()\) be the function for finding the minimum value, and \(\max()\) be the function for finding the maximum value.

[0021] Preferably, in step S2, assuming that the scene depth is similar for each block after underwater image segmentation, find the optimal transmission rate in each block and invert the degradation model of underwater contrast. The specific formula is:

[0022]

[0023] where \(J(y)\) is the clear image, \(t\) is the transmission rate, and \(A\) is the background light.

[0024] By adjusting the transmission rate of each small block, improve the image contrast at the pixel level and propose the contrast knowledge, that is, the estimation of the contrast transmission rate of the clear image. The specific formula is:

[0025]

[0026] (6) where \(I\) c (x)=\(\{I\) r (x), \(I\) g (x), \(I\) b (x)\} is the original image, \(A\) C =\(\{A\) R , \(A\) G , \(A\) B \} is the background light, \(y\in\Omega(x)\) is the pixel neighborhood of \(x\), \(\min()\) is the function for finding the minimum value, and \(\max()\) is the function for finding the maximum value.

[0027] Preferably, in step S2, based on the fact that the direct component of underwater imaging generally has more textures and significant image gradients, while the layer information of the backscattering component of the underwater image is relatively smooth and the gradient approaches zero, invert the degradation model of underwater gradient. The specific formula is:

[0028] \(I = L\) p (\lambda)e -α(λ)l +E b (\lambda) (7)

[0029] where \(L\) p (\lambda)\) is the light reflected by the target on the imaging plane, \(e\) -α(λ)l is the light attenuation coefficient, and \(E\) b (\lambda)\) is the backscattering component of the background light.

[0030] Model the gradient distribution of the underwater image. The specific formula is:

[0031]

[0032] Among them, are the probability density functions of the gradient distributions of the direct component J and the scattering component B, respectively.

[0033] In a clear image, the gradient of the direct component J is the largest, and the gradient of the scattering component B approaches zero, enabling the optimal separation of the clear layer and the noise layer. The gradient knowledge, that is, the gradient optimization function, is proposed. The specific formula is:

[0034]

[0035] Among them, min() is the minimum value solving function.

[0036] Preferably, the cross-knowledge guidance module proposed in step S3 combines the input image with the prior map and the transmission map. The prior knowledge, as a feature selector, can mark the features that need special attention in a weighted form. The specific steps are as follows: first, invert the transmission map; multiply the input feature map pixel by pixel with the transmission map based on the prior knowledge; add the result of the multiplication as a weight pixel by pixel to the input feature map.

[0037] Preferably, in step S3, according to the gradient knowledge, a gradient knowledge loss function is proposed.

[0038] Preferably, in step S3, a cross-knowledge embedding model is constructed, which specifically includes the following steps:

[0039] Step S301: Based on the encoder-decoder network as the basic framework and the residual module as the basic module, the encoder extracts features through downsampling, and the decoder upsamples to restore the original dimension; the input of the encoder-decoder network is the RGB image, the red channel transmission map, and the contrast transmission map;

[0040] Step S302: Add the prior knowledge to the network through the cross-knowledge guidance module. According to the characteristic that the red wave attenuates rapidly underwater, add the red channel knowledge when the encoder extracts features; to improve the contrast of the enhancement effect, add the contrast knowledge to the decoder;

[0041] Step S303: Add a channel attention mechanism to the encoder-decoder network, and mark the features containing more information by assigning different weights;

[0042] Step S304: The gradient knowledge is added to the model training as a loss function, and as a part of the model training loss function, it is linearly combined with the mean square error loss function and the perceptual loss function.

[0043] Through the data learning and training of the underwater image library in step S4, the accuracy of the cross-knowledge embedding network model designed in step S3 is improved.

[0044] Step S5: Input the original underwater image into the model trained in Step S4 to obtain a clear enhanced underwater image, thereby realizing the enhancement of the original underwater image.

[0045] Beneficial effects: Through comparative experiments with existing methods, the cross - knowledge embedding learning method proposed in the present invention can achieve better results in color correction and retaining image details. Description of the Drawings

[0046] Figure 1 is a flowchart of the underwater image enhancement method based on cross - knowledge embedded learning according to an embodiment of the present invention;

[0047] Figure 2 is a network structure diagram of the cross - knowledge embedded network according to an embodiment of the present invention;

[0048] Figure 3 is a structure diagram of the cross - knowledge guidance module according to an embodiment of the present invention;

[0049] Figure 4 is a comparison diagram of the underwater image enhancement effects in an embodiment of the present invention. Detailed Embodiments

[0050] The following further clarifies the present invention in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification by those skilled in the art fall within the scope defined by the appended claims of this application.

[0051] For the underwater image enhancement method based on cross - knowledge embedding learning, first, invert the underwater imaging degradation process and model the underwater imaging process. On this basis, for different degradation problems, deduce the prior knowledge of underwater imaging;

[0052] Secondly, according to the combination method of cross - knowledge and neural network, realize multi - knowledge embedding, and design a cross - knowledge guidance module and a prior - knowledge loss function;

[0053] Finally, based on the encoder - decoder framework, construct a cross - knowledge embedding model. This model combines prior knowledge and neural network and uses prior knowledge to guide the learning and training of the network. The trained model is used to process underwater images to realize the enhancement of underwater images.

[0054] As Figure 1 described, it specifically includes the following steps:

[0055] S1. According to the process of underwater light propagation, construct an underwater imaging model. The specific content is as follows:

[0056] First, according to the McGlamery-Jaffe underwater imaging model, the light energy obtained by the underwater imaging system is divided into three parts: the light energy transmitted directly, the light energy scattered forward, and the light energy scattered backward; the specific formula is:

[0057] I(x) = J(x)t(x) + A(1 - t(x)) (1)

[0058] Where I(x) is the light ray reaching the imaging plane, J(x) is the direct component of the target reflected light reaching the imaging plane, A is the background light, t(x) is the transmittance, and x is the pixel point.

[0059] According to the Lambert-Beer law, the light ray decreases exponentially with respect to the distance d(x) during transmission, β is the attenuation coefficient, and the specific formula for the transmittance t(x) is:

[0060] t(x) = e -βd(x) (2)

[0061] S2. According to the attenuation problem in the underwater imaging process, invert the underwater degradation model to obtain the prior knowledge of underwater imaging. The specific content is:

[0062] (1) Construct red channel knowledge

[0063] During the underwater imaging process, as the distance increases, the red wave attenuates faster. To reflect this characteristic, the red channel prior of the underwater image adjusts the red channel of the original model in the dark channel prior, and the other channels remain unchanged. It is proposed to restore the clear image through the transmittance of the red channel. The specific formula is:

[0064] 1 - I R = t(1 - J R ) + (1 - t)(1 - A R ) (3)

[0065]

[0066] Where the original image and the clear image are I = (I R , I G , I B ) and J = (J R , J G , J B ), the water light component is A = (A R , A G , A B ),, Ω(x) is the pixel neighborhood of x, t is the transmittance, is the transmittance of the red channel of the clear image, and min() is the minimum value solving function.

[0067] (2) Construct contrast knowledge

[0068] During the underwater imaging process, assuming that the scene depth of each block is similar after image segmentation, the optimal transmission rate is found in each block. The block-based transmission rate can improve the contrast of the image. Since the contrast is inversely proportional to the projection rate t, assuming the enhanced image is J(y), the transmission rate estimation based on contrast knowledge is The specific formula is:

[0069]

[0070] where J(y) is the clear image, t is the transmission rate, A is the background light, and I c (x) = {I r (x), I g (x), I b (x)} is the original image, and A C = {A R , A G , A B} is the background light. y ∈ Ω(x) is the pixel neighborhood of x, min() is the minimum value solving function, and max() is the maximum value solving function.

[0071] (3) Construct gradient knowledge

[0072] Analyze the scattering process of underwater light. The original image I can be divided into the target reflection component and the background light backscattering component. The direct component J of underwater imaging generally has more textures and significant image gradients. The layer information of the backscattering component of underwater imaging is relatively smooth, and the gradient approaches zero. Therefore, model the probability density function ε of the gradient distribution of the underwater image. According to the model, when ε reaches the minimum value, the gradient of the clear layer is the largest, and the noise layer approaches zero, and the optimal separation of the clear layer and the noise layer can be achieved. The specific formula is:

[0073] I = L p (λ)e -α(λ)l + E b (λ) (7)

[0074]

[0075] where L p (λ) is the target reflection light on the imaging surface, e -α(λ)l is the light attenuation coefficient, and E b (λ) is the backscattering component of the background light, are the probability density functions of the gradient distributions of the direct component J and the scattering component B respectively, and min() is the minimum value solving function.

[0076] S3. According to the method of combining cross knowledge and neural network, design a cross knowledge guidance module, through Figure 3The module shown combines the input image and prior knowledge; first, the transmission map is inverted. Then, the input feature map is multiplied pixel by pixel with the transmission map based on prior knowledge, and the resulting result is added pixel by pixel to the input feature map as weights. The specific formula is:

[0077]

[0078] where X represents the input feature map, T represents the transmission map, represents pixel-by-pixel multiplication, represents pixel-by-pixel addition, represents the feature map incorporating prior knowledge.

[0079] To achieve the feature of cross-knowledge embedding, a gradient loss function is designed. The original image is put into the cross-knowledge-guided network for training, and the results obtained in each round pass through the gradient loss function, which plays a role in optimizing the model. The specific formula is:

[0080]

[0081] where ||·|| F is the Frobenius norm, f 1 is the first-order horizontal gradient operator, f 2 is the transpose of the first-order gradient operator, f La is the second-order Laplacian operator, and λ and ρ are weight coefficients.

[0082] Construct a cross-knowledge embedding network, including the following sub-steps:

[0083] S301. The network is based on an encoder-decoder framework and consists of residual modules; the residual modules are as Figure 2 shown. Each residual module contains two residual blocks, and each residual block consists of three 3×3 convolutional kernels with ReLU activation functions and a single 3×3 convolutional kernel; after each residual block, it is connected through a residual structure.

[0084] S302. Add the cross-knowledge guidance module proposed in S3 to the network, as Figure 2 shown. The input of the network is an RGB image, a red-channel transmission map, and a contrast transmission map; each residual module is connected to the cross-knowledge guidance module through a concat operation; the red-channel transmission map and the RGB image are input into the encoder, and features are extracted through downsampling; in the decoder, the contrast transmission map combines with the RGB feature map and is upsampled to restore the size.

[0085] S303. Add Figure 2The SE attention mechanism in [it] serves as a channel attention module. After the encoder feature map is weighted by the channel attention module, it is combined with the decoder feature map to obtain the relevant features of the underwater image.

[0086] S304. Linearly combine the gradient loss function with the mean square error function and the perceptual loss function as the loss function for optimizing the model training to achieve a balance between the visual effect and the quantitative index; the mean square error loss function Calculate the mean of the squared difference between the predicted value f(x) and the true value y; the perceptual loss function Calculate the conversion from the space to the feature space, defined by the ReLU activation layer of the VGG-19 network pre-trained on ImageNet; during the process of training the model, the gradient optimization function As part of the objective function, it can achieve the optimal separation of the clear layer and the noise layer. The specific formula is:

[0087]

[0088] Among them, represents the predicted value of the pixel value x i and y i is the true value; represents the pre-trained VGG-19 network, represents the j-th convolutional layer, C j H j W j represents the size of the j-th feature map, ||·|| 2 is the 2-norm, J f and B f are the direct component and the scattering component of the predicted value f(x) respectively, ||·|| F is the Frobenius norm, f 1 is the first-order horizontal gradient operator, f 2 is the transpose of the first-order gradient operator, f La is the second-order Laplacian operator, and λ and ρ are weight coefficients; λ 1 and λ 2 are set to 0.01 and 0.05 respectively, indicating the proportion of each function in the total objective function.

[0089] In step S4, through learning and training on the data of the underwater image library, the accuracy of the cross-knowledge embedding network model designed in step S3 is improved.

[0090] In step S5, the original underwater image is input into the model obtained in step S4 to obtain a clear enhanced underwater image, realizing the enhancement of the original underwater image.

[0091] To verify the effectiveness of the underwater enhancement method proposed in the present invention in terms of color correction and detail preservation, representative images are selected from six underwater datasets, including green underwater images, blue underwater images, and underwater images with rich details, and compared with six different underwater image processing methods, including Figure 4 the underwater image enhancement method based on the underwater scene prior in the second row of Figure 4 the underwater image enhancement based on the gated fusion network in the third row, Figure 4 the underwater image enhancement based on the conditional generative adversarial network in the fourth row, Figure 4 the underwater image enhancement based on the multi-color space encoder-medium transmission guided decoder in the fifth row, Figure 4 the underwater image restoration based on the attenuation prior in the sixth row, Figure 4 the underwater image restoration based on the gradient prior in the seventh row. Figure 4 The comparison results between the method proposed in the present invention and other methods are shown. It can be seen that the present invention has good effects, with better color correction and detail restoration effects.

Claims

1. A method for underwater image enhancement based on cross-knowledge embedding learning, characterized in that: The steps include: S1. According to the underwater light propagation process, an underwater imaging model is constructed to analyze the light attenuation mechanism of the target during underwater imaging; S2. According to the light attenuation mechanism in the underwater imaging process, the underwater degradation model is inverted to obtain the red channel knowledge, contrast knowledge and gradient knowledge of the current underwater image; S3. Design an encoder-decoder deep learning network with three kinds of cross knowledge embedded: red channel knowledge, contrast knowledge, and gradient knowledge. Propose a cross knowledge guidance module to embed cross knowledge into the model. The red channel knowledge is embedded into the encoder module, the contrast knowledge is embedded into the decoder module, and the gradient knowledge is used as the loss function to train the network model. S4, training the encoder-decoder deep learning network designed in S3 through the underwater image database until convergence; S5. Input the current original underwater image and output the enhanced underwater image result.

2. The underwater image enhancement method based on cross-knowledge embedding learning according to claim 1, characterized in that: The step S1 constructs an underwater imaging model according to the underwater light propagation process, and the specific formula is: I(x)=J(x)t(x)+A(1-t(x)) (1) Among them, I(x) is the light reaching the imaging plane, J(x) is the direct component of the target reflected light reaching the imaging plane, A is the background light, t(x) is the transmission rate, and x is the pixel point.

3. The underwater image enhancement method based on cross-knowledge embedding learning according to claim 1, characterized in that: The step S2 inverts the degradation model of the underwater red channel according to the light attenuation mechanism that the red wave attenuates faster as the distance increases during underwater imaging. The specific formula is: 1-I R =t(1-J R )+(1-t)(1-A R ) (3) Based on the red channel degradation model, the red channel knowledge of underwater images is proposed, that is, the red channel transmission rate estimation of clear images. The specific formula is: Among them, I R (y), I G (y), I B (y) are the intensities of the red, green, and blue channels of the underwater image, respectively. C is the intensity of each color channel of the underwater background light, C∈{R,G,B} is the red, green and blue color channels, y∈Ω(x) is the pixel area based on x, min() is the minimum solution function, and max() is the maximum solution function.

4. The underwater image enhancement method based on cross-knowledge embedding learning according to claim 1, characterized in that: In step S2, assuming that the scene depth of each block of the underwater image is similar after the underwater image is divided into blocks, the optimal transmission rate is found in each block, and the degradation model of the underwater contrast is inverted. The specific formula is: Among them, J(y) is the clear image, t is the transmission rate, and A is the background light; By adjusting the transmission rate of each block, the image contrast is improved from the pixel, and the contrast knowledge is proposed, that is, the contrast transmission rate estimation of the clear image. The specific formula is: Among them, I c (x) = {I r (x),I g (x),I b (x)} is the original image, A C ={A R ,A G ,A B } is the background light, y∈Ω(x) is the pixel area based on x, min() is the minimum value solving function, and max() is the maximum value solving function.

5. The underwater image enhancement method based on cross-knowledge embedding learning according to claim 1, characterized in that: The degradation model of the underwater gradient is inverted, and the specific formula is: I=L p (l)e -α(λ)l +E b (l) (7) Among them, L p (λ) is the reflected light from the target on the imaging surface, e -α(λ)l is the light attenuation coefficient, E b (λ) is the backscattered component of the background light; The gradient distribution of underwater images is modeled, and the specific formula is: in, are the probability density functions of the gradient distribution of the direct component J and the scattered component B respectively; In the clear image, the gradient of the direct component J is the largest, and the gradient of the scattered component B is close to zero, achieving the optimal separation of the clear layer and the noise layer; the gradient knowledge, that is, the gradient optimization function, is proposed, and the specific formula is: Among them, min() is the minimum value solving function.

6. The underwater image enhancement method of cross-knowledge embedding learning according to claim 1, characterized in that: The cross-knowledge guidance module proposed in step S3 combines the input image with the prior map and the transmission map. The prior knowledge is used as a feature selector to mark the features that need special attention in a weighted form. The specific steps include: first negating the transmission map; multiplying the input feature map by the transmission map based on prior knowledge pixel by pixel; and adding the multiplication result as a weight to the input feature map pixel by pixel.

7. The underwater image enhancement method of cross-knowledge embedding learning according to claim 1, characterized in that: The step S3 proposes a gradient knowledge loss function based on the gradient knowledge.

8. The underwater image enhancement method of cross-knowledge embedding learning according to claim 1, characterized in that: The step S3, constructing a cross-knowledge embedding model, specifically includes the following steps: Step S301, using the encoder-decoder network as the basic framework and the residual module as the basic module, the encoder extracts features by downsampling, and the decoder restores the original dimension by upsampling; the input of the encoder-decoder network is the RGB image, the red channel transfer map and the contrast transfer map; Step S302, adding prior knowledge to the network through the cross knowledge guidance module, adding red channel knowledge when the encoder extracts features according to the characteristic that underwater red waves decay quickly; adding contrast knowledge in the decoder to improve the contrast of the enhancement effect; Step S303: The encoder-decoder network adds a channel attention mechanism to mark features containing more information by assigning different weights; Step S304: Gradient knowledge is added to the model training as a loss function, and as a part of the model training loss function, it is linearly combined with the mean square error loss function and the perceptual loss function.

Citation Information

Patent Citations

  • Restoration method for underwater image

    CN114549342A

  • Underwater image enhancement method combining physical prior and deep learning

    CN116309232A

  • Underwater image enhancement method based on knowledge conversion

    CN116823644A

  • Underwater image enhancement method based on physical perception transformer

    CN117522755A

  • Underwater target detection method based on edge prior

    CN118864270A