A Fishery Underwater Image Enhancement Method Based on a Multi-Branch Joint Correction Network

Through the multi-branch joint correction network method, the problems of color distortion and brightness degradation in underwater images are solved, the details of images and color calibration are realized, and the quality of underwater images is improved.

CN119991494BActive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510460234.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-04
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

The prior art lacks joint optimization of multiple degradation types in underwater images, especially the tight coupling of color distortion and brightness degradation in physical mechanisms, resulting in a decline in underwater image quality, especially in artificial light source scenarios.

Method used

A method based on multi-branch joint correction network is adopted, including dark channel priors, minimum graph cutting noise separator, brightness high sensitivity decomposition network, detail enhancement network and artificial light source lighting area mask, respectively, and joint optimization is achieved through a deep learning framework.

Benefits of technology

Effectively remove noise in underwater images, correct brightness inhomogeneity and color offset, improve the visual effect of the image, and provide reliable data support for advanced underwater tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991494B_ABST
    Figure CN119991494B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for enhancing underwater fishery images based on a multi-branch joint correction network. First, the original underwater image is denoised by a minimum graph cut noise separator to obtain the luminance and color channels. Secondly, the luminance channel is decoupled to obtain the intrinsic characteristic information and the luminance information, and a luminance correction network is designed to eliminate the luminance degradation of the image; multi-scale feature extraction is performed on the corrected luminance channel, and at the same time, the luminance channel is transformed to the frequency domain by wavelet transform to obtain the enhanced luminance channel. Then, the difference operation is performed on the two luminance channels before and after luminance correction and then binarized to obtain the artificial light source illumination area mask. Finally, color restoration guided by the artificial light source illumination area mask is performed on the color channel, the enhanced color channel is output, and it is fused with the enhanced luminance channel to output the enhancement result. The present invention can effectively realize the detail restoration and color calibration of underwater images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of underwater image enhancement, and particularly relates to a method for enhancing fishery underwater images based on a multi-branch joint correction network. Background Art

[0002] In recent years, with the continuous development of human exploration of marine resources by human beings, the underwater environment has become a research hotspot for researchers all over the world. As an important marine resource, marine fishery resources are an important part of China's economic development. It is of great positive significance to fully explore and develop them for the development of coastal areas. Underwater optical images can intuitively reflect complex underwater conditions and play a key role in underwater environment detection and investigation. With the vigorous development of computer vision technology, underwater fishery resource detection based on optical images has been widely applied, including fish biomass detection, fish recognition and analysis, etc., which is of great significance for fishery aquaculture.

[0003] However, the images obtained by underwater optical imaging systems generally have quality degradation phenomena such as spectral distortion, global contrast attenuation, and non-linear brightness deterioration. The reason for the degradation of underwater image quality is the color shift phenomenon caused by the selective attenuation effect of the non-uniform spectral absorption characteristics of water on the visible light band. At the same time, the discrete scattering formed by suspended particles in water will cause the disorder of the light field energy distribution and simultaneously form a background noise field, resulting in the problems of decreased sharpness of image edges, blurred detail features, and significantly reduced global contrast. In addition, the natural light conditions are severely weakened in the underwater environment, especially after leaving shallow water. To obtain relatively clear images, artificial light sources need to be used to improve illumination. Since the artificial light source is relatively close to the imaging device and is usually white light, the illuminated area will show higher brightness and less color distortion, while the unilluminated area tends to be dark, which makes the underwater images show serious and uneven degradation and further complicates the problem.

[0004] To solve the above problems, many researchers have carried out research work. In 2012, Chen et al. adopted a simplified atmospheric scattering model to reduce the system model parameters while improving the robustness and universality of the system. In 2019, Akkaynak et al. proposed a modified underwater imaging model, and proposed to separately process the direct and backscattering systems. This differential attenuation coefficient model realizes underwater image restoration through depth parameters. In 2019, Sun et al. proposed an underwater image enhancement structure based on a symmetric encoder-decoder for learning the mapping from low-quality underwater images to high-quality underwater images. In 2020, an image enhancement method based on underwater scene prior was proposed to solve the problem of lack of paired underwater data sets.

[0005] However, most current studies focus on the correction of a single degradation type and lack a systematic design for the joint optimization of multiple degradations. Color distortion and brightness degradation are tightly coupled in the physical mechanism. The selective absorption of water causes the rapid attenuation of red light (resulting in color deviation), while the scattering effect reduces the light intensity (causing a decrease in brightness). Most current studies regard the two as independent problems and lack research on their synergistic effects. Summary of the Invention

[0006] To solve the problems of excessive noise, brightness degradation, and color shift in underwater images in artificial light source scenarios, based on the above analysis, the present invention proposes a method for enhancing fishery underwater images based on a multi-branch joint correction network. Using a deep learning-based underwater image enhancement algorithm framework, noise removal, brightness correction, and color restoration of underwater images are carried out simultaneously to obtain clear optical images with good visual effects, providing reliable data support for advanced tasks based on underwater images.

[0007] The tracking scheme includes three aspects: an image denoising method based on an underwater imaging model, an image brightness correction and detail enhancement method in an artificial light source scenario, and an image color restoration method in an artificial light source scenario. The specific steps are as follows:

[0008] S1 Take the original degraded underwater image as the input. Through the method of dark channel prior, search for scattered pixels (i.e., background light and impurities) with different distributions in the depth information of the original underwater image to obtain the backscattering map; calculate the pixel difference between the backscattering map and the original underwater image to obtain the forward transmission component that enhances the contrast; finally, design a minimum graph cut noise separator to further remove the impurities existing in the forward transmission component, and finally obtain the denoised image.

[0009] S2 Convert the color space of the denoised underwater image from the RGB space to the LAB space, aiming to process the brightness (L) and color (AB) channels separately.

[0010] S3 Construct a brightness hypersensitive decomposition network to decouple the brightness channel L in S2 to obtain the intrinsic characteristic information and brightness information, and design a brightness correction network to eliminate the brightness degradation of the image.

[0011] S4 Construct a detail enhancement network to perform multi-scale feature extraction on the corrected brightness channel L to obtain the spatial domain information of the image. At the same time, the brightness channel is transformed to the frequency domain by wavelet transform to restore more high-frequency details. The spatial domain information and high-frequency details complement each other to achieve the detail enhancement of the underwater image and obtain the enhanced brightness channel L.

[0012] S5 Perform a difference operation on the two brightness channels L before and after brightness correction and then perform binarization processing to finally obtain the artificial light source illumination area mask M, which is used to guide the restoration of the color channel AB.

[0013] S6 performs color restoration on the color channels AB of the degraded image guided by the artificial light source illumination area mask M. First, the shallow features of the color channels AB and the mask M are extracted separately and fused to obtain the illumination guidance prior. Then, context features are extracted from the illumination guidance prior, and a dual-branch module combining convolution and Transformer is introduced in the deep layer of the feature extraction network to fully extract global and local color information to assist image understanding. Finally, the enhanced color channels are output.

[0014] S7 Re-fuses the enhanced luminance channel L and color channels AB and outputs the final enhancement result.

[0015] The step S1 further includes:

[0016] S1-1: The underwater optical image is mainly composed of two parts: forward transmitted light and backscattered light. According to the imaging principle, an approximate underwater image imaging model can be obtained.

[0017] I(x) = D(x)t(x) + B(1 - t(x))

[0018] Where x represents the pixel point coordinate, I(x) represents the picture taken by the camera in the underwater environment, D(x) represents the radiated light of the object itself, t(x) represents the transmittance field, and B represents the ambient light.

[0019] S1-2: Improve the image restoration process through the underwater imaging model. Process the separated depth image of the image, perform dark channel depth calculation, and search for backscattered images with different depth distributions, that is, the background light and impurities outside the target object.

[0020]

[0021] Where Ω(x) is a local window centered on pixel x. Since the illumination conditions are uneven in the underwater scene, regions at different depths may be scattered and attenuated to different degrees, and the traditional fixed background light cannot adapt to this change. Therefore, the present invention constructs an adaptive background light fitting. By stratifying the depth map, the dark channel values of each layer are statistically analyzed, and then curve fitting is performed to obtain the relationship between the background light and the depth, that is, the backscattered image. Specifically, first divide the depth map d(x) into K intervals {d1, d2,..., d K}, and then statistically analyze the dark channels of each layer. For each depth d K Calculate the statistical value of the dark channel in the corresponding region (the present invention takes the 90th percentile).

[0022] B K = Percentile({J dark (x)|d(x) ∈ d K}), 90%)

[0023] Among them, Percentile is an index used to represent the position of data distribution in statistics. Finally, using (d K , B K ), the backscattering diagram B(d) is obtained.

[0024] B(d) = B ∞ (1 - e -γd(x) )

[0025] Among them, B ∞ is the background light at infinity, and γ is the scattering coefficient.

[0026] S1-3: Calculate the pixel difference between the backscattering diagram and the original underwater image to obtain the forward transmission component that enhances the contrast. First, combine the depth information d(x) with the water attenuation characteristics to construct the transmittance field t(x).

[0027] t(x) = e -βd(x)

[0028] Among them, β is the water attenuation coefficient. Secondly, through pixel difference calculation and combining with the transmittance field, the forward transmission component that enhances the contrast, that is, the enhanced image J(x), is obtained.

[0029]

[0030] Among them is the numerical stability constant.

[0031] S1-4: The noise separator based on the minimum graph cut first defines a weighted undirected graph G = (V, E), where the vertex set V contains all the pixels of the enhanced image J(x) and two special nodes: the source point (S, representing the foreground) and the sink point (T, representing the background); the edge set E includes the edges between each pixel and its adjacent pixels and the edges between each pixel and the source point / sink point, which connect adjacent pixels and assign certain weights. At the same time, design an energy function, and as long as the energy function reaches the minimum value, the separation of background noise can be achieved. The energy function is divided into two parts: the data term and the smooth term. For the calculation of the data term, the present invention uses two center points F and B as the representative values of the target and background noise respectively, and then defines the weight w iS of each pixel and the weight w iT with the source point:

[0032]

[0033] Among them, I i is the gray value of pixel i, and σ controls the sensitivity to the distance from the center point. Therefore, the expression of the data term is as follows:

[0034]

[0035] For the calculation of the smooth term, it can be defined according to the color difference between adjacent pixels:

[0036]

[0037] where I i and I j are the grayscale values of adjacent pixels respectively, and γ is a constant used to control the smoothness. Therefore, the total energy function E(L):

[0038] E(L) = ∑ i∈V D(i) + λ∑ (i,j)∈N V(I i , I j )

[0039] where λ is the weight parameter. Finally, the minimum cut algorithm is used to divide the graph G into two subsets: the foreground V1 and the background V2, such that the source node S belongs to V1 and the sink node T belongs to V2, and at the same time the energy function E(L) reaches the minimum value. The foreground V1 is the denoised image.

[0040] The step S3 further includes:

[0041] S3-1: The high-brightness sensitivity decomposition network first processes the channel L image based on the color constancy of the object's reflected light:

[0042] L = R * I

[0043] where L represents the luminance channel, R represents the reflection map, I represents the illumination map, and * represents the per-pixel multiplication operation. The reflection map expresses the details and color information of the object itself, while the illumination map contains the global illumination information. The high-brightness sensitivity decomposition network adopts a classic encoder-decoder symmetric design, captures context semantics through the contraction path (encoding), and restores spatial details through the expansion path (decoding) to form a U-shaped data stream. The input L first obtains the feature representation of the input image through dilated convolution, expands the receptive field without increasing the number of parameters, and sequentially performs feature processing through the encoding part and the decoding part of the network. The encoding part obtains multi-scale image features through downsampling at different ratios, and uses the residual module (RM) to aggregate information of specific-scale features after each downsampling to fully obtain context information to assist image understanding. The decoder part gradually restores the original size of the image through upsampling, sets a skip connection structure between the encoder and the decoder to combine low-level spatial information and high-level semantic information, reduces the information lost due to downsampling, and then adjusts the number of feature channels through 1×1 convolution, and finally outputs the illumination map and the reflection map.

[0044] S3-2: The brightness correction network is constructed based on the symmetric feature pyramid of U-Net, including an encoding path and a decoding path. That is, it extracts multi-scale features through downsampling, gradually reducing the spatial size of the feature map while increasing the number of channels. Then, it gradually restores the resolution through upsampling, combines the features of the encoding path for fine segmentation or correction, and uses skip connections to combine spatial and semantic features. The real image has severe and uneven brightness degradation. Convolution equally focuses on different channels and spatial positions of features, allowing unimportant information to occupy processing resources and hindering the transmission of important information. Therefore, the present invention constructs a group attention method based on spatial correlation and channel statistics. The group attention includes two parts, namely intra-group attention and inter-group attention. Intra-group attention is the concatenation of spatial attention and channel attention, and inter-group attention is self-attention. Intra-group attention is applied to the encoding path, pooling the features at each scale in the spatial and channel dimensions to generate attention maps in two dimensions for calibrating the original features, thereby making full use of multi-scale context information. Inter-group attention is used in the decoding path to fuse low-level spatial features and high-level semantic features, selectively performing multi-level information aggregation to enhance the expression ability of the network, and finally outputting the corrected brightness map.

[0045] The step S4 further includes: The detail reconstruction network consists of two parts, namely a symmetric encoder-decoder network for extracting spatial domain features and a spatial-frequency domain information fusion module. The symmetric encoder-decoder network is composed of convolution modules, responsible for multi-scale feature extraction, realizing the perception of image details under multiple receptive fields. The detail features are stitched with the input image through skip connections to form the output. The spatial-frequency domain information fusion module consists of two branches, a convolution branch and a wavelet reconstruction module, which respectively process the output of the symmetric encoder-decoder network and the luminance channel L, aiming to supplement spatial information with frequency domain information and restore more high-frequency details of the image. The wavelet reconstruction module mainly converts the information in the spatial domain to the wavelet domain through wavelet transform for restoration. After wavelet transform, the input feature is first transformed into four different sub-bands:

[0046] {F LL ,F LH ,F HL ,F HH} = DWT(F input )

[0047] where F LL ,F LH ,F HL ,F HH are the four sub-bands of the feature respectively, and DWT(.) represents the wavelet transform operation. Then, the four sub-bands share a convolutional layer for restoration to prevent interference between bands. Finally, the restored sub-bands are reconstructed into the output feature through inverse wavelet transform operation:

[0048] F out= IDWT(F LL , F LH , F HL , F HH )

[0049] Where IDWT(.) represents the inverse wavelet transform operation. The output of the wavelet reconstruction module and the output of the convolutional branch are added pixel by pixel to obtain the final detail enhancement result.

[0050] The step S5 further includes: calculating the artificial light source illumination area mask M. Considering that the artificial light source illumination area has less color distortion compared to other areas, it is therefore instructive for the color restoration of underwater images. By performing a difference operation on the illumination maps before and after brightness correction (S = I delight - I) and performing binarization to obtain the artificial light source illumination area mask, where the pixels in the area greater than the threshold are set to 1 and the rest are set to 0. In addition, the parts greater than the threshold are aggregated because the artificial light source tends to affect the entire area. The specific formula is as follows:

[0051] S = I delight - I

[0052] therefold = (max(S) - min(S)) * 0.6

[0053] mask = 1, S ≥ therefold

[0054] mask = 0, S < therefold

[0055] Where therefold is the threshold, S is the brightness residual, max(.) is the operation of taking the maximum pixel value, and min(.) is the operation of taking the minimum pixel value. Thus, the artificial light source illumination area mask M is obtained, which is used to guide the restoration of the color channels AB.

[0056] The step S6 further includes:

[0057] S6 - 1: Extract the shallow geometric and texture features of the color channels AB through a lightweight convolutional network:

[0058] F AB = ConvBlock(AB)

[0059] Where ConvBlock contains a 3×3 convolution, batch normalization (BN), and ReLU activation layer, and the number of output channels C = 32. Encode the mask M as the spatial weight map W M , and enhance the local features of the illuminated area:

[0060] W M = σ(Conv 1×1 (M))

[0061] where σ is the Sigmoid function, and Conv 1×1 is a convolutional layer with a 1×1 convolutional kernel. Finally, cross-modal feature fusion is performed, and the color feature and the mask weight are cascaded at the channel level and spatially modulated:

[0062] F fuse = Concat(F AB ) * W M

[0063] where Concat is an operation of concatenating by dimension.

[0064] S6-2: Design a color restoration network guided by the mask M. First, extract the features of the AB channels and the mask M and perform feature concatenation to reduce the influence of the blank part in the mask, and use the auxiliary reference information provided by the mask M features in the artificial light source illumination area to guide the color restoration of the network. Use the dense residual block as the basic feature extraction module to make full use of the feature information at each level and help the training of the network. In particular, use the CT module to improve the deepest layer of the network. The CT module is a dual-branch parallel structure, namely the Transformer feature extraction branch and the convolutional feature extraction branch. Relying on the powerful global modeling ability of the Transformer, the module can fully extract the global color information. At the same time, the convolutional branch makes up for its lack of attention to local information. After passing through the color restoration network, the color channels with eliminated color distortion are finally output.

[0065] Advantages of the present invention:

[0066] The method proposed by the present invention can effectively solve the problems of many image noise, brightness degradation, and color shift in the artificial light source scene. Among them, the dark channel prior and the graph cut noise separator proposed by the present invention can remove background noise and improve the image purity; by constructing a high-brightness sensitivity decomposition network and a brightness correction network to correct the image brightness, the problem of uneven brightness degradation is solved. Finally, by calculating the mask M to guide color restoration, the obvious color degradation caused by the absorption of light by water is eliminated. The image enhancement framework proposed by the present invention can effectively realize the detail restoration and color calibration of underwater images. Description of the drawings

[0067] Figure 1 is the overall framework of underwater image enhancement of the present invention;

[0068] Figure 2 is the noise cancellation network of the present invention;

[0069] Figure 3 is the subjective experimental comparison diagram of the present invention. Detailed implementation manners

[0070] This section further describes the present invention in combination with specific implementation processes. The following examples are listed and described in detail with the accompanying drawings.

[0071] As Figure 1 shown, a fishery underwater image enhancement method based on a multi-branch joint correction network can effectively improve the color and brightness degradation of underwater images, including the following implementation steps:

[0072] S1 Take the degraded original underwater image as input, and obtain depth information using monocular depth estimation; secondly, search for scattering pixels with different depth distributions through the dark channel prior method; through inverse operation, obtain the forward transmission component that enhances the contrast, as Figure 2 shown; finally, design a minimum graph cut noise separator to further remove the image background noise and obtain a denoised image.

[0073] S2 After denoising, convert the color space of the underwater image from the RGB space to the LAB space, aiming to process the luminance (L) and color (AB) channels respectively.

[0074] S3 Construct a luminance hypersensitive decomposition network to decouple the intrinsic characteristic information and luminance information of the luminance L channel of the image, and eliminate the luminance degradation of the image by designing a luminance correction network.

[0075] S4 Use the corrected luminance channel to construct a detail enhancement network, fully extract the spatial domain information of the image through a multi-scale feature extraction method, and at the same time input it to the frequency domain through wavelet transform to restore more high-frequency details. The two complement each other to enhance the details of the underwater image.

[0076] S5 Perform a difference operation on the two luminance channels L before and after luminance correction and then perform binarization processing to finally obtain the artificial light source illumination area mask M, which is used to guide the restoration of the color channels AB.

[0077] S6 Perform color restoration of the color channels AB of the degraded image guided by the artificial light source illumination area mask M. First, extract the shallow features of the color channels AB and the mask M respectively and fuse them to obtain the illumination guidance prior, construct a multi-scale color restoration network to fully extract the context features of the illumination guidance prior, and introduce a double-branch module combining convolution and Transformer in the deep layer of the network to fully extract global and local color information to help image understanding, and finally output the enhanced color channels.

[0078] S7 Re-fuse the enhanced L and AB channels and output the final enhancement result.

[0079] Furthermore, the denoised image in the above step S1 is specifically implemented by the following steps:

[0080] S1-1 Monocular depth estimation to obtain depth information: Obtain the depth information d(x) of the scene by estimating the contrast attenuation from the monocular color image I RGB Estimate the scene depth d(x) through contrast attenuation.

[0081] S1-2 Dark channel prior: Calculate the dark channel depth through and take the mean of the top 0.1% pixels with the highest intensity as the global background light, and combine the depth information d(x) with the water attenuation characteristics to construct the transmittance field t(x) = e -βd(x) Finally, through inverse operation, obtain the forward transmission component that enhances the contrast.

[0082] S1-3 Minimum graph cut noise separator: Define a weighted undirected graph G = (V, E) and an energy function E(L), where V is the set of vertices, including all pixels of the image and two special nodes: the source node (S, representing the foreground target) and the sink node (T, representing the background noise), E is the set of edges, and E(L) consists of a data term and a smooth term. To effectively separate the background noise, use the minimum cut algorithm to divide the graph G into two subsets: the foreground V1 and the background V2, such that the source node S belongs to V1 and the sink node T belongs to V2, while the energy function E(L) reaches the minimum value.

[0083] Furthermore, the correction of brightness degradation in step S3 above is specifically implemented by the following steps:

[0084] S3-1 Brightness hypersensitive decomposition network: The brightness hypersensitive decomposition network adopts a U-shaped encoding-decoding structure. First, obtain the feature representation of the input image through dilated convolution, and perform feature processing through the encoding part and the decoding part of the network in sequence. The encoding part obtains multi-scale image features through downsampling with different ratios, and aggregates the information of the features at a specific scale using a residual module after each downsampling, thereby expanding the receptive field and fully obtaining the context information to help image understanding. At the same time, the residual connection within the module can help network training. The decoder part gradually restores the original size of the image through upsampling, and sets a skip connection structure between the encoder and decoder to combine the low-level spatial information and the high-level semantic information, reduce the information lost due to downsampling, and then adjust the number of feature channels through 1×1 convolution, and finally output the illumination map and the reflection map.

[0085] S3-2 Brightness Correction Network: The brightness correction network follows the main principle of the symmetric feature pyramid done by U-net, including two paths of encoding and decoding. That is, first, multi-scale features of the image are extracted through downsampling with different ratios (encoding), and then the image resolution is gradually restored through upsampling (decoding). Using skip connections, the 256×256 feature map of the first layer is concatenated with the 256×256 feature map of the fourth layer to fuse spatial details and semantic information. The present invention constructs a group attention method based on spatial correlation and channel statistics. The group attention includes two parts, namely intra-group attention and inter-group attention. Intra-group attention is applied to the encoding path to generate attention maps for features at each scale in the spatial dimension and channel dimension, calibrating the importance of features, so as to make full use of multi-scale context information. Inter-group attention is used to fuse the low-level spatial features and high-level semantic features of the encoding path, selectively aggregate information, enhance the expression ability of the network, and finally output the corrected brightness map.

[0086] Furthermore, the detail reconstruction network in the above step S4 is specifically implemented by the following steps:

[0087] S4 Detail Reconstruction Network: It consists of two parts, namely a symmetric encoder-decoder network for extracting spatial domain features and a spatial-frequency domain information fusion module. The former consists of a convolutional module and a multi-scale feature extraction module. Among them, the multi-scale feature extraction module consists of convolutional kernels of different sizes, perceives image details under multiple receptive fields, and the output features are concatenated with the input image through skip connections to form the output. The spatial-frequency domain information fusion module consists of two branches of convolution and wavelet reconstruction module, aiming to use frequency domain information to supplement spatial domain information and restore more high-frequency details of the image. The wavelet reconstruction module mainly converts the information in the spatial domain to the wavelet domain through wavelet transform for restoration. After wavelet transform, the input features are first transformed into four different sub-bands, and the four sub-bands share convolutional layers for restoration to prevent interference between bands. Finally, the restored sub-bands are reconstructed into output features through inverse wavelet transform operation.

[0088] Furthermore, the calculation of the artificial light source illumination area mask M in the above step S5 is specifically implemented by the following steps:

[0089] S5 Mask M Calculation: By performing a difference operation on the illumination maps before and after brightness correction (S = I delight -I) and performing binarization processing, the artificial light source illumination area mask is obtained, where the pixels in the area greater than the threshold are set to 1, and the rest are set to 0.

[0090] Furthermore, the color restoration network guided by the artificial light source illumination mask M in the above step S6 is specifically implemented by the following steps:

[0091] S6 Mask-Guided Color Restoration Method: First, extract the features of the AB channels and the mask M and perform feature splicing to reduce the influence of the blank parts in the mask, and use the auxiliary reference information provided by the features of the mask M in the artificial light source illumination area to guide the color restoration of the network. Use the dense residual block as the basic feature extraction module to make full use of the feature information at all levels and help the training of the network. In particular, use the CT module to improve the deepest layer of the network. The CT module is a dual-branch parallel structure, namely the Transformer feature extraction branch and the convolutional feature extraction branch. Relying on the powerful global modeling ability of the Transformer, the module can fully extract global color information. At the same time, the convolutional branch makes up for its lack of attention to local information. After passing through the color restoration network, finally output the color channels that eliminate color distortion.

[0092] Evaluation Metrics:

[0093] To illustrate the effectiveness of the method of the present invention for underwater image enhancement, an underwater image quality assessment metric (UIQM) similar to the visual system is used for evaluation. UIQM aims to objectively evaluate the quality of underwater images by quantifying the degradation characteristics of underwater images (such as non-uniform color cast, low contrast, blur, etc.). It combines three key components, namely color measurement (UICM), sharpness measurement (UISM), and contrast measurement (UIConM). Color measurement is used to evaluate color balance and saturation. Sharpness measurement is based on the RGB color space and evaluates sharpness through the local contrast of the image. Contrast measurement evaluates the uniformity of the brightness distribution through the overall contrast of the image. To comprehensively evaluate the role of each index, UIQM weights the above three key components:

[0094] UIQM = c1 × UICM + c2 × UISM + c3 × UIConM

[0095] where c1, c2, and c3 are all weight parameters.

[0096] Use the enhancement effect on the underwater captured images to evaluate the method of the present invention. Figure 3 A subjective experiment showing the enhancement effect comparison of the present invention with other enhancement methods can be observed that the present invention can achieve better color correction and detail restoration for underwater degraded images. To fully illustrate the effectiveness of the enhancement framework constructed by the present invention, compare the enhancement results of this method with the commonly used models for underwater image enhancement and evaluate using the above evaluation metrics. The comparison results are shown in Table 1. It can be seen from Table 1 that the image enhancement structure proposed by the present invention has certain advantages in the underwater image enhancement task.

[0097] Table 1 Comparison of Image Enhancement Effects of Different Models

[0098]

Claims

1. A fishery underwater image enhancement method based on a multi-branch joint correction network, characterized in that It includes the following steps: S1: Take the original underwater image with degradation as the input and denoise it through the minimum graph cut noise separator; S2: Convert the color space of the denoised underwater image from the RGB space to the LAB space to obtain the luminance L and the color AB channels; S3: Construct a luminance hypersensitive decomposition network to decouple the luminance channel L to obtain the intrinsic characteristic information and the luminance information, and design a luminance correction network to eliminate the luminance degradation of the image; S4: Construct a detail enhancement network to perform multi-scale feature extraction on the corrected luminance channel L to obtain the spatial domain information. At the same time, the luminance channel is transformed to the frequency domain through wavelet transform to obtain the enhanced luminance channel L; S5: Perform a difference operation on the two luminance channels L before and after luminance correction and then perform binarization processing to obtain the artificial light source illumination area mask M; S6: Perform color restoration on the color channels AB guided by the artificial light source illumination area mask M, output the enhanced color channels, and fuse them with the enhanced luminance channel L to output the final enhancement result.

2. The method for enhancing fishery underwater images based on a multi-branch joint correction network according to claim 1, characterized in that The specific implementation process of S1 is as follows: S1-1: The underwater optical image consists of two parts, the forward transmitted light and the backscattered light. According to the imaging principle, an approximate underwater image imaging model is obtained; S1-2: Based on the approximate underwater image imaging model, perform the calculation of the dark channel depth J of the depth image separated from the image, obtain the depth map according to the dark channel depth, divide the depth map d(x) into K intervals {d1, d2,..., dK}, then statistically calculate the dark channels of each layer, and calculate the statistical value B of the dark channel within the corresponding region for each depth d; use (d, B) to obtain the backscattering map B(d) = B(1 - e^(-γd)), where B is the background light at infinity and γ is the scattering coefficient; dark (x) Calculate, obtain the depth map according to the dark channel depth, divide the depth map d(x) into K intervals {d1, d2,..., d K K}, then statistically calculate the dark channels of each layer, and for each depth d K calculate the statistical value B of the dark channel within the corresponding region respectively K ; use (d K , B K ) to obtain the backscattering map B(d) = B ∞ (1 - e -γd(x) )), where B ∞ is the background light at infinity and γ is the scattering coefficient; S1-3: Combine the depth information d(x) with the water body attenuation characteristics to construct the transmittance field t(x); through pixel difference calculation, combine the transmittance field to obtain the forward transmission component that enhances the contrast, that is, the enhanced image J(x): Among them is the numerical stability constant, I(x) represents the picture taken by the camera in the underwater environment, and β is the water body attenuation coefficient; S1-4: The noise separator based on the minimum graph cut first defines a weighted undirected graph G=(V, E), where the vertex set V contains all the pixels of the enhanced image J(x) as well as the source point S and the sink point T; the edge set E includes the edges between each pixel and its adjacent pixels and the edges between each pixel and the source point / sink point; design an energy function to perform the minimum cut algorithm to divide the graph G into the foreground V1 and the background V2, so that the source node S belongs to V1 and the sink node T belongs to V2. At the same time, the energy function E(L) reaches the minimum value, and the foreground V1 is the denoised image.

3. The method for enhancing fishery underwater images based on a multi-branch joint correction network according to claim 2, wherein, The energy function is divided into two parts: a data term and a smoothness term; for the calculation of the data term, taking two center points F and B as representative values of the target and background noise respectively, the weight w of each pixel is defined iS and the weight w of the source point iT , I i is the gray value of pixel i, and σ controls the sensitivity to the center point: Smoothing term V(I i ,I j ) is defined according to the color difference between adjacent pixels, where I i and I j are the gray values of adjacent pixels respectively, and the energy function E(L) is the sum of the smoothing terms added to the sum of the data terms through a coefficient in the middle.

4. The method for enhancing underwater fishery images based on a multi-branch joint correction network according to claim 3, wherein, The specific implementation process of step S3 is as follows: S3-1: The luminance hypersensitive decomposition network adopts an encoder-decoder symmetric design. The input L first passes through dilated convolution to obtain the feature representation of the input image, and then performs feature processing through the encoding part and the decoding part of the network in sequence; the encoding part obtains multi-scale image features through downsampling with different ratios, and uses the residual module RM to aggregate the features after each downsampling; The decoder part gradually restores the original size of the image through upsampling, sets a skip connection structure between the encoder and the decoder to realize the combination of semantic information, and then adjusts the number of feature channels through convolution to output the illumination map and the reflection map; S3-2: The luminance correction network is constructed based on the symmetric feature pyramid of U-Net, including two paths of encoding and decoding, and uses skip connections to combine spatial and semantic features; construct a group attention method based on spatial correlation and channel statistics; The group attention includes two parts, namely intra-group attention and inter-group attention. Intra-group attention is the concatenation of spatial attention and channel attention, and inter-group attention is self-attention; The intra-group attention is applied to the encoding path, and the features at each scale are pooled in the spatial dimension and the channel dimension to generate attention maps in two dimensions for calibrating the original features; The inter-group attention is used for decoding path semantic features, performing multi-level information aggregation, and finally outputting the corrected luminance map.

5. The method for enhancing underwater fishery images based on a multi-branch joint correction network according to claim 4, wherein The specific implementation of step S4 is as follows: The detail reconstruction network consists of two parts, namely a symmetric encoder-decoder network for extracting spatial domain features and a spatial-frequency domain information fusion module; The symmetric encoder-decoder network consists of convolutional modules, which are responsible for multi-scale feature extraction, realizing the perception of image details under multiple receptive fields, and the detail features are spliced with the input image through skip connections to form an output; The spatial-frequency domain information fusion module consists of two branches, a convolutional branch and a wavelet reconstruction module, which respectively process the output of the symmetric encoder-decoder network and the luminance channel L; The wavelet reconstruction module converts the information in the spatial domain to the wavelet domain for restoration through wavelet transform. After wavelet transform, the input features are first transformed into four different sub-bands; The four sub-bands share a convolutional layer for restoration, and finally the restored sub-bands are reconstructed into output features through inverse wavelet transform operation; The output of the wavelet reconstruction module is pixel-wise added to the output of the convolutional branch to obtain the final detail enhancement result.

6. The method for enhancing underwater fishery images based on a multi-branch joint correction network according to claim 5, characterized in that, The artificial light source illumination area mask M is obtained by performing a difference operation on the illumination maps before and after brightness correction and then performing binarization processing. The pixels in the area greater than the threshold are set to 1, and the rest are set to 0; In addition, the parts greater than the threshold are aggregated to obtain the artificial light source illumination area mask M.

7. The method for enhancing underwater fishery images based on a multi-branch joint correction network according to claim 6, wherein The specific implementation process of the color restoration of the color channels AB guided by the artificial light source illumination area mask M and outputting the enhanced color channels is as follows: S6-1: Extract the geometric and texture features of color channels AB through a lightweight convolutional network. The mask M is encoded as a spatial weight map W through a convolutional network and a sigmoid function M , enhancing the local features of the illuminated area; Finally, cross-modal feature fusion is performed, and the color features and the spatial weight map are cascaded in channels and spatially modulated; S6-2: Design a color restoration network guided by the mask M. First, extract the features of the AB channels and the mask M for feature splicing, and use the auxiliary reference information provided by the features of the artificial light source illumination area mask M to guide color restoration; Use the residual block as the basic feature extraction module and help the training of the network; Improve the deepest layer of the network using the CT module. The CT module is a dual-branch parallel structure, namely a Transformer feature extraction branch and a convolutional feature extraction branch. After passing through the color restoration network, the color channels with color distortion eliminated are output.

Citation Information

Patent Citations

  • Laminated light guide using optical anisotropic film and planar light source device using same

    CN110214286A

  • Underwater image enhancement method of multi-attention mechanism guided by brightness mask

    CN116402715A