Fishery underwater image enhancement method based on multi-branch joint correction network
Through the multi-branch joint correction network, the noise, brightness and color problems of underwater images are solved, and the joint optimization of multiple degradation types in underwater images is achieved, the clarity and color recovery of the image are achieved, and high-quality data support is provided.
Patent Information
- Application Number
- CN202510460234.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The prior art lacks joint optimization of multiple degradation types in underwater images, especially the tight coupling of color distortion and brightness degradation in physical mechanisms, resulting in a decline in underwater image quality, especially in artificial light source scenarios.
A method based on multi-branch joint correction network is adopted, including dark channel priors, minimum image cutting noise separator, brightness high sensitivity decomposition network, brightness correction network, detail enhancement network and color restoration network, respectively, and image enhancement is achieved through multi-scale feature extraction and frequency domain information fusion.
Effectively remove noise in underwater images, correct brightness inhomogeneity and color shifts, restore clear visual effects, and provide reliable data support for advanced underwater tasks.
Smart Images

Figure CN119991494A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of underwater image enhancement, and in particular relates to a fishery underwater image enhancement method based on a multi-branch joint correction network. Background Art
[0002] In recent years, with the continuous development of marine resources by humans, the underwater environment has become a hot topic for researchers around the world. Marine fishery resources, as an important marine resource, are an important part of my country's economic development. Full exploration and development of them has a very positive significance for the development of coastal areas. Underwater optical images can intuitively reflect complex underwater conditions and play a key role in underwater environmental detection and investigation. With the vigorous development of computer vision technology, underwater fishery resource detection based on optical images has been widely used, including fish biomass detection, fish identification and analysis, etc., which is of great significance to aquaculture.
[0003] However, images acquired by underwater optical imaging systems generally suffer from quality degradation phenomena such as spectral distortion, global contrast attenuation, and nonlinear brightness degradation. The reason for the degradation of underwater image quality is the color shift caused by the non-uniform spectral absorption characteristics of water bodies that produce a selective attenuation effect on the visible light band. At the same time, the discrete scattering of suspended particles in the water body will cause the energy distribution of the light field to be disordered and form a background noise field at the same time, resulting in the reduction of image edge sharpness, the blurring of detail features, and the significant reduction of global contrast. In addition, the natural light conditions in the underwater environment, especially after leaving shallow water, are severely weakened. In order to obtain clearer images, artificial light sources need to be used to improve lighting. Since artificial light sources are relatively close to imaging equipment and are usually white light, the illuminated areas will show higher brightness and less color distortion, while the unilluminated areas tend to be dark, which causes serious and uneven degradation of underwater images, further complicating the problem.
[0004] To solve the above problems, many researchers have carried out research on this. In 2012, Chen et al. adopted a simplified atmospheric scattering model to reduce the system model parameters while improving the robustness and universality of the system. In 2019, Akkaynak et al. proposed a modified underwater imaging model, proposing to process direct and backscattered systems separately. The differential attenuation coefficient model realizes underwater image restoration through depth parameters. In 2019, Sun et al. proposed a symmetric encoder-decoder based underwater image enhancement structure for learning the mapping from low-quality underwater images to high-quality underwater images. In 2020, an image enhancement method based on underwater scene prior was proposed to solve the problem of lack of paired underwater data sets.
[0005] However, most current studies focus on the correction of a single type of degradation, lacking a systematic design for the joint optimization of multiple degradations. Color distortion and brightness degradation are closely coupled in physical mechanism. The selective absorption of water causes rapid attenuation of red light (causing color shift), while the scattering effect reduces the light intensity (causing brightness reduction). Most current studies regard the two as independent issues, lacking research on their synergistic effects. Summary of the invention
[0006] In order to solve the problems of high noise, brightness degradation and color shift in underwater images under artificial light source scenes, based on the above analysis, the present invention proposes a method for underwater image enhancement in fisheries based on a multi-branch joint correction network. With a deep learning-based underwater image enhancement algorithm framework, underwater image noise removal, brightness correction and color restoration are carried out simultaneously to obtain clear optical images with good visual effects, providing reliable data support for advanced tasks based on underwater images.
[0007] The tracking solution includes three aspects: image denoising method based on underwater imaging model, image brightness correction and detail enhancement method under artificial light source scene, and image color restoration method under artificial light source scene. The specific steps are as follows:
[0008] S1 takes the degraded original underwater image as input, and searches for the scattered pixels (i.e., background light and impurities) with different distributions in the depth information of the original underwater image through the dark channel prior method to obtain the backscattering map; it calculates the pixel difference between the backscattering map and the original underwater image to obtain the forward transmission component with enhanced contrast; finally, it designs a minimum graph cut noise separator to further remove the impurities in the forward transmission component to finally obtain a denoised image.
[0009] S2 converts the color space of the denoised underwater image from RGB space to LAB space, aiming to process the brightness (L) and color (AB) channels separately.
[0010] S3 constructs a brightness high-sensitivity decomposition network to decouple the brightness channel L in S2, obtains intrinsic characteristic information and brightness information, and designs a brightness correction network to eliminate the brightness degradation of the image.
[0011] S4 constructs a detail enhancement network and performs multi-scale feature extraction on the corrected brightness channel L to obtain the spatial information of the image. At the same time, the brightness channel is transformed to the frequency domain through wavelet transform to restore more high-frequency details. The spatial information and high-frequency details complement each other to achieve underwater image detail enhancement and obtain the enhanced brightness channel L.
[0012] S5 performs a difference operation on the two brightness channels L before and after brightness correction and then performs a binarization process, and finally obtains an artificial light source illumination area mask M, which is used to guide the restoration of the color channel AB.
[0013] S6 performs color restoration on the color channel AB of the degraded image guided by the mask M of the artificial light source illumination area. First, the shallow features of the color channel AB and the mask M are extracted and fused to obtain the illumination guidance prior. Then, the context feature extraction is performed on the illumination guidance prior. In the deep layer of the feature extraction network, a dual-branch module combining convolution and Transformer is introduced to fully extract global and local color information to help image understanding, and finally the enhanced color channel is output.
[0014] S7 re-integrates the enhanced brightness channel L and color channel AB and outputs the final enhanced result.
[0015] The step S1 further comprises:
[0016] S1-1: Underwater optical images are mainly composed of two parts: forward transmitted light and backscattered light. According to the imaging principle, an approximate underwater image imaging model can be obtained.
[0017]
[0018] Where x represents the pixel coordinates, Represents a picture taken by a camera in an underwater environment. represents the radiated light of the object itself, represents the transmittance field and B represents the ambient light.
[0019] S1-2: Improve the image restoration process through the underwater imaging model. Process the depth image separated from the image, perform dark channel depth calculation, and search for backscattering patterns distributed at different depths, that is, background light and impurities outside the target object.
[0020]
[0021] Where Ω(x) is a local window centered at pixel x. Due to the uneven lighting conditions in underwater scenes, areas at different depths may be subject to different degrees of scattering and attenuation. The traditional fixed background light cannot adapt to such changes. Therefore, the present invention constructs an adaptive background light fitting method. By dividing the depth map into layers, counting the dark channel values of each layer, and then performing curve fitting, the relationship between background light and depth, i.e., the backscattering map, is obtained. Specifically, the depth map is first Divide into K intervals , and then count the dark channels of each layer, for each depth The statistical values of the dark channels in the corresponding areas are calculated respectively (the present invention takes the 90% quantile).
[0022]
[0023] Percentile is an indicator used in statistics to indicate the distribution of data. Get the backscatter map .
[0024]
[0025] in is the background light at infinity, and γ is the scattering coefficient.
[0026] S1-3: Calculate the pixel difference between the backscattered image and the original underwater image to obtain the forward transmission component with improved contrast. First, combine the depth information d(x) with the water body attenuation characteristics to construct the transmittance field t(x).
[0027]
[0028] in is the water attenuation coefficient. Secondly, after pixel difference calculation and combined with the transmittance field, the forward transmission component with improved contrast is obtained, that is, the enhanced image .
[0029]
[0030] in is a numerical stability constant.
[0031] S1-4: Noise Separator Based on Minimum Graph Cut First, define a weighted undirected graph G=(V,E), where the vertex set V contains the enhanced image All pixels of the image and two special nodes: the source point (S, representing the foreground) and the sink point (T, representing the background); the edge set E includes the edges between each pixel and its adjacent pixels and the edges between each pixel and the source point / sink point, which connect the adjacent pixels and give them certain weights. At the same time, an energy function is designed. As long as the energy function reaches the minimum value, the background noise can be separated. The energy function is divided into two parts: the data term and the smooth term. For the calculation of the data term, the present invention uses two center points F and B as the representative values of the target and background noise respectively, and then defines the weight of each pixel The weight of the source point :
[0032]
[0033] in is the gray value of pixel i, and σ controls the sensitivity of the distance to the center point. Therefore, the expression of the data term is as follows:
[0034]
[0035] For the calculation of the smoothness term, it can be defined based on the color difference between adjacent pixels:
[0036]
[0037] in, and are the grayscale values of adjacent pixels, is a constant used to control the degree of smoothness. Therefore, the total energy function E(L):
[0038]
[0039] in is the weight parameter. Finally, the minimum cut calculation method is used to divide the graph G into two subsets: foreground and background , so that the source node S belongs to , the sink node T belongs to , and the energy function E(L) reaches its minimum value, the foreground This is the denoised image.
[0040] The step S3 further comprises:
[0041] S3-1: The brightness high-sensitivity decomposition network first processes the channel L image based on the color constancy of the object's reflected light:
[0042]
[0043] in represents the brightness channel, represents the reflection map, represents the lighting diagram, Represents pixel-by-pixel multiplication operations. The reflectance map expresses the details and color information of the object itself, and the illumination map contains global illumination information. The brightness high-sensitivity decomposition network adopts the classic encoder-decoder symmetric design, capturing contextual semantics by contracting the path (encoding) and restoring spatial details by expanding the path (decoding), forming a U-shaped data flow. Input First, the feature representation of the input image is obtained through the dilated convolution, and the receptive field is expanded without increasing the number of parameters. The feature processing is then performed in turn through the encoding and decoding parts of the network. The encoding part obtains multi-scale image features through downsampling at different ratios, and uses the residual module (RM) to aggregate information on features of a specific scale after each downsampling, fully obtaining contextual information to help image understanding. The decoder part gradually restores the original size of the image through upsampling, and sets a jump connection structure between the encoder and decoder to achieve the combination of low-level spatial information and high-level semantic information, reducing the information lost due to downsampling. Convolution adjusts the number of feature channels and finally outputs the illumination map and the reflection map.
[0044] S3-2: The brightness correction network is constructed based on the symmetric feature pyramid of U-Net, including two paths of encoding and decoding, that is, multi-scale features are extracted by downsampling, the spatial size of the feature map is gradually reduced, and the number of channels is increased at the same time, and then the resolution is gradually restored by upsampling, and the features of the encoding path are combined for fine segmentation or correction, and the spatial and semantic features are combined by skip connection. The brightness of the real image is severely degraded and uneven, and the convolution pays equal attention to the different channels and spatial positions of the features, so that unimportant information occupies processing resources and hinders the transmission of important information. Therefore, the present invention constructs a group attention method based on spatial correlation and channel statistics. Group attention includes two parts, namely, intra-group attention and inter-group attention. The intra-group attention is the series connection of spatial attention and channel attention, and the inter-group attention is self-attention. The intra-group attention is applied to the encoding path, and the features of each scale are pooled in the spatial dimension and the channel dimension to generate an attention map in two dimensions for calibrating the original features, thereby making full use of multi-scale context information. Inter-group attention is used to fuse low-level spatial features and high-level semantic features in the decoding path, selectively perform multi-level information aggregation, enhance the expression ability of the network, and finally output the corrected brightness map.
[0045] The step S4 further includes: the detail reconstruction network consists of two parts, namely a symmetric codec network for extracting spatial domain features and a spatial-frequency domain information fusion module. The symmetric codec network consists of a convolution module, which is responsible for multi-scale feature extraction and realizes the perception of image details under multiple receptive fields. The detail features are spliced with the input image through jump connections to form an output. The spatial-frequency domain information fusion module consists of two branches, a convolution module and a wavelet reconstruction module, which respectively process the output of the symmetric codec network and the brightness channel L, aiming to use frequency domain information to supplement spatial domain information and restore more high-frequency details of the image. The wavelet reconstruction module mainly converts the information in the spatial domain to the wavelet domain for restoration through wavelet transform. After wavelet transform, the input features are first transformed into four different sub-bands:
[0046]
[0047] in are the four characteristic sub-bands, Represents the wavelet transform operation. After that, the four sub-bands share the convolution layer for restoration to prevent interference between bands, and finally the restored sub-bands are reconstructed into output features through the inverse wavelet transform operation:
[0048]
[0049] in Represents the inverse wavelet transform operation. The output of the wavelet reconstruction module is added to the output of the convolution branch at the pixel level to obtain the final detail enhancement result.
[0050] The step S5 further includes: calculating the mask M of the artificial light source illumination area. Considering that the artificial light source illumination area has less color distortion than other areas, it has guiding significance for the color restoration of underwater images. By performing a difference operation on the illumination map before and after brightness correction ( ) and performs binarization to obtain the artificial light illumination area mask, where the pixels in the area greater than the threshold are set to 1 and the rest are set to 0. In addition, the part greater than the threshold is clustered because the artificial light source tends to affect the entire area. The specific formula is as follows:
[0051]
[0052] in is the threshold value, is the brightness residual, To obtain the maximum pixel value operation, The operation of taking the minimum pixel value is performed. Thus, the artificial light source illumination area mask M is obtained, which is used to guide the restoration of the color channel AB.
[0053] The step S6 further comprises:
[0054] S6-1: Extract shallow geometric and texture features of color channel AB through lightweight convolutional network:
[0055]
[0056] ConvBlock contains 3×3 convolution, batch normalization (BN) and ReLU activation layers, and the number of output channels C = 32. The mask M is encoded as a spatial weight map , strengthen the local features of the illuminated area:
[0057]
[0058] Where σ is the Sigmoid function, is a convolutional layer with a 1×1 convolution kernel. Finally, cross-modal feature fusion is performed, and the color features and mask weights are channel-concatenated and spatially modulated:
[0059]
[0060] Concat is a dimension-based concatenation operation.
[0061] S6-2: Design a color restoration network under the guidance of mask M. First, extract the features of AB channels and mask M and perform feature splicing to reduce the impact of the blank part in the mask, and use the auxiliary reference information provided by the mask M features in the artificial light illumination area to guide the color restoration of the network. Use the dense residual block as the basic feature extraction module to make full use of the feature information at all levels and help the training of the network. In particular, the CT module is used to improve the deepest layer of the network. The CT module is a dual-branch parallel structure, which is the Transformer feature extraction branch and the convolution feature extraction branch. Relying on the powerful global modeling ability of Transformer, the module can fully extract global color information. At the same time, the convolution branch makes up for its lack of attention to local information. After the color restoration network, the color channel with color distortion eliminated is finally output.
[0062] Beneficial effects of the present invention:
[0063] The method proposed in the present invention can effectively solve the problems of high image noise, brightness degradation and color shift in artificial light source scenes. Among them, the dark channel prior and graph cut noise separator proposed in the present invention can remove background noise and improve image purity; by constructing a brightness high-sensitivity decomposition network and a brightness correction network, the image brightness is corrected to solve the problem of uneven brightness degradation. Finally, by calculating the mask M to guide color restoration, the obvious color degradation caused by the absorption of light by water is eliminated. The image enhancement framework proposed in the present invention can effectively realize the detail restoration and color calibration of underwater images. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is the overall framework of underwater image enhancement of the present invention;
[0065] Figure 2 A noise cancellation network of the present invention;
[0066] Figure 3 This is a subjective experiment comparison chart of the present invention. DETAILED DESCRIPTION
[0067] This section further explains the present invention in conjunction with a specific implementation process, and the following examples are listed herewith, and described in detail with reference to the accompanying drawings.
[0068] like Figure 1 The method for enhancing underwater fishery images based on a multi-branch joint correction network is shown, which can effectively improve the color and brightness degradation of underwater images, and includes the following implementation steps:
[0069] S1 takes the degraded original underwater image as input and uses monocular depth estimation to obtain depth information. Secondly, it searches for scattered pixels distributed at different depths through the dark channel prior method. After inverse operation, it obtains the forward transmission component with improved contrast, such as Figure 2 As shown; finally, a minimum graph cut noise separator is designed to further remove the image background noise and obtain a denoised image.
[0070] S2 converts the color space of underwater images from RGB space to LAB space after denoising, aiming to process the brightness (L) and color (AB) channels separately.
[0071] S3 constructs a brightness high-sensitivity decomposition network to decouple the intrinsic characteristic information of the brightness L channel of the image from the brightness information, and eliminates the brightness degradation of the image by designing a brightness correction network.
[0072] S4 uses the corrected brightness channel to construct a detail enhancement network, fully extracts the spatial information of the image through a multi-scale feature extraction method, and simultaneously inputs the wavelet transformed image to the frequency domain to restore more high-frequency details. The two complement each other to enhance the details of underwater images.
[0073] S5 performs a difference operation on the two brightness channels L before and after brightness correction and then performs a binarization process, and finally obtains an artificial light source illumination area mask M, which is used to guide the restoration of the color channel AB.
[0074] S6 performs color restoration on the color channel AB of the degraded image guided by the mask M of the artificial light source illumination area. First, the shallow features of the color channel AB and the mask M are extracted and fused to obtain the illumination guidance prior. A multi-scale color restoration network is constructed to fully extract contextual features of the illumination guidance prior. A dual-branch module combining convolution and Transformer is introduced in the deep layer of the network to fully extract global and local color information to help image understanding, and finally the enhanced color channel is output.
[0075] S7 re-integrates the enhanced L and AB channels and outputs the final enhanced result.
[0076] Furthermore, the denoised image in the above step S1 is specifically implemented by the following steps:
[0077] S1-1 Monocular Depth Estimation Obtaining Depth Information: From Monocular Color Image Estimating scene depth via contrast falloff .
[0078] S1-2 Dark channel prior: Pass Calculate the dark channel depth, take the average of the top 0.1% pixels with the highest intensity as the global background light, and combine the depth information d(x) with the water attenuation characteristics to construct the transmittance field , and finally after inverse operation, the forward transmission component with improved contrast is obtained.
[0079] S1-3 Minimum graph cut noise separator: Define a weighted undirected graph G = (V, E) and an energy function E (L), where V is a vertex set, including all pixels of the image and two special nodes: source points (S, representing foreground targets) and sink points (T, representing background noise), E is an edge set, and E (L) consists of two parts: data items and smooth items. In order to effectively separate the background noise, the minimum cut algorithm is used to split the graph G into two subsets: foreground and background , so that the source node S belongs to , the sink node T belongs to , and the energy function E(L) reaches its minimum value.
[0080] Furthermore, the correction of brightness degradation in the above step S3 is specifically implemented by the following steps:
[0081] S3-1 Brightness High-Sensitivity Decomposition Network: The brightness high-sensitivity decomposition network adopts a U-type encoding-decoding structure. It first obtains the feature representation of the input image through hole convolution, and then performs feature processing through the encoding part and decoding part of the network in turn. The encoding part obtains multi-scale image features through downsampling at different ratios, and uses the residual module to aggregate information on features of a specific scale after each downsampling, thereby expanding the receptive field and fully obtaining contextual information to help image understanding. At the same time, the residual connection within the module can help network training. The decoder part gradually restores the original size of the image through upsampling, and sets a jump connection structure between the codec to realize the combination of low-level spatial information and high-level semantic information, reducing the information lost due to downsampling, and then through Convolution adjusts the number of feature channels and finally outputs the illumination map and the reflection map.
[0082] S3-2 Brightness Correction Network: The brightness correction network follows the main principle of the symmetrical feature pyramid made by U-net, and includes two paths, encoding and decoding, that is, first extracting the multi-scale features of the image by downsampling at different ratios (encoding), and then gradually restoring the image resolution by upsampling (decoding), and using jump connections to splice the 256×256 feature map of the first layer with the 256×256 feature map of the fourth layer to fuse spatial details and semantic information. The present invention constructs a group attention method based on spatial correlation and channel statistics. Group attention consists of two parts, namely, intra-group attention and inter-group attention. Intra-group attention is applied to the encoding path, and attention maps are generated for each scale feature in the spatial dimension and channel dimension to calibrate the importance of the features, thereby making full use of multi-scale contextual information. Inter-group attention is used to fuse the low-level spatial features and high-level semantic features of the encoding path, selectively aggregate information, enhance the expression ability of the network, and finally output the corrected brightness map.
[0083] Furthermore, the detailed reconstruction network in the above step S4 is specifically implemented by the following steps:
[0084] S4 detail reconstruction network: It consists of two parts, namely the symmetric encoding and decoding network for extracting spatial domain features and the spatial-frequency domain information fusion module. The former consists of a convolution module and a multi-scale feature extraction module, wherein the multi-scale feature extraction module consists of convolution kernels of different sizes, perceives image details under multiple receptive fields, and the output features are spliced with the input image through jump connections to form output. The spatial-frequency domain information fusion module consists of two branches, the convolution and wavelet reconstruction modules, which aim to supplement the spatial domain information with frequency domain information and restore more high-frequency details of the image. The wavelet reconstruction module mainly converts the information in the spatial domain to the wavelet domain for restoration through wavelet transform. After wavelet transform, the input features are first transformed into four different sub-bands, and the four sub-bands share the convolution layer for restoration to prevent interference between bands. Finally, the restored sub-bands are reconstructed into output features through inverse wavelet transform.
[0085] Furthermore, the calculation of the artificial light source illumination area mask M in the above step S5 is specifically implemented by the following steps:
[0086] S5 Mask M calculation: The illumination map before and after brightness correction is calculated by difference operation ( ) and performs binarization processing to obtain the artificial light illumination area mask, where the pixels in the area greater than the threshold are set to 1 and the rest are set to 0.
[0087] Furthermore, the color restoration network guided by the artificial light source illumination mask M in the above step S6 is specifically implemented by the following steps:
[0088] S6 Mask-guided color restoration method: First, extract the features of AB channels and mask M and perform feature splicing to reduce the influence of the blank part in the mask, and use the auxiliary reference information provided by the mask M features of the artificial light illumination area to guide the color restoration of the network. The dense residual block is used as the basic feature extraction module to make full use of the feature information at all levels and help the training of the network. In particular, the CT module is used to improve the deepest layer of the network. The CT module is a dual-branch parallel structure, which is the Transformer feature extraction branch and the convolution feature extraction branch. Relying on the powerful global modeling ability of the Transformer, the module can fully extract global color information. At the same time, the convolution branch makes up for its lack of attention to local information. After the color restoration network, the color channel with color distortion eliminated is finally output.
[0089] Evaluation indicators:
[0090] In order to illustrate the effectiveness of the method of the present invention for underwater image enhancement, the underwater image quality evaluation metric (UIQM) of the visual system is used for evaluation. UIQM aims to objectively evaluate the quality of underwater images by quantifying their degradation characteristics (such as non-uniform color cast, low contrast, blur, etc.). It combines three key components, namely color measurement (UICM), clarity measurement (UISM) and contrast measurement (UIConM). Color measurement is used to evaluate color balance and saturation. Clarity measurement is based on the RGB color space and evaluates clarity through the local contrast of the image. Contrast measurement evaluates the uniformity of brightness distribution through the overall contrast of the image. In order to comprehensively evaluate the role of each indicator, UIQM weights the above three key components:
[0091]
[0092] in , , are all weight parameters.
[0093] The method of the present invention is evaluated by the enhancement effect on underwater collected images. Figure 3 The subjective experiment showing the enhancement effect of the present invention compared with other enhancement methods shows that the present invention can achieve better color correction and detail restoration for underwater degraded images. In order to fully illustrate the effectiveness of the enhancement framework constructed by the present invention, the enhancement results of this method are compared with the commonly used model for underwater image enhancement and evaluated using the above-mentioned evaluation indicators. The comparison results are shown in Table 1. It can be seen from Table 1 that the image enhancement structure proposed by the present invention has certain advantages in underwater image enhancement tasks.
[0094] Table 1 Comparison of image enhancement effects of different models
[0095]
Claims
1. A method for fishery underwater image enhancement based on a multi-branch joint correction network, characterized in that: The following steps are involved: S1 takes the degraded original underwater image as input and denoises it through the minimum graph cut noise separator; S2 converts the color space of the denoised underwater image from RGB space to LAB space to obtain the brightness L and color AB channels; S3 constructs a brightness high-sensitivity decomposition network to decouple the brightness channel L, obtains intrinsic characteristic information and brightness information, and designs a brightness correction network to eliminate the brightness degradation of the image; S4 builds a detail enhancement network, performs multi-scale feature extraction on the corrected brightness channel L, obtains spatial domain information, and at the same time, transforms the brightness channel to the frequency domain through wavelet transformation to obtain the enhanced brightness channel L; S5 performs a difference operation on the two brightness channels L before and after brightness correction and then performs a binarization process to obtain an artificial light source illumination area mask M; S6 performs color restoration on the color channel AB under the guidance of the artificial light source illumination area mask M, outputs an enhanced color channel, merges it with the enhanced brightness channel L, and outputs a final enhancement result.
2. The method for fishery underwater image enhancement based on a multi-branch joint correction network according to claim 1 is characterized in that: The specific implementation process of S1 is as follows: S1-1: Underwater optical images are composed of forward transmitted light and backscattered light. An approximate underwater image imaging model is obtained based on the imaging principle. S1-2: Based on the approximate underwater image imaging model, the dark channel depth of the depth image separated from the image is analyzed. Calculate and get the depth map according to the dark channel depth. Divide into K intervals , and then count the dark channels of each layer, for each depth Calculate the statistical values of the dark channels in the corresponding areas respectively ;use Get the backscatter map ,in is the background light at infinity, γ is the scattering coefficient; S1-3: Combine the depth information d(x) with the water attenuation characteristics to construct the transmittance field t(x); after pixel difference calculation and combining with the transmittance field, obtain the forward transmission component with improved contrast, i.e. the enhanced image : ; in is a numerical stability constant; S1-4: Noise Separator Based on Minimum Graph Cut First, define a weighted undirected graph G=(V,E), where the vertex set V contains the enhanced image All pixels of and source point S and sink point T; the edge set E includes the edges between each pixel and its neighboring pixels and the edges between each pixel and the source point / sink point; Design energy function to perform minimum cut calculation to divide graph G into foreground and background , so that the source node S belongs to , the sink node T belongs to , and the energy function E(L) reaches its minimum value, the foreground This is the denoised image.
3. The method for fishery underwater image enhancement based on a multi-branch joint correction network according to claim 2 is characterized in that: The energy function is divided into two parts: data term and smooth term. For the calculation of data term, two center points F and B are used as representative values of target and background noise respectively, and the weight of each pixel is defined as The weight of the source point , is the grayscale value of pixel i, and σ controls the sensitivity of the distance to the center point: ; Smooth Item The energy function E(L) is defined as the sum of the smoothness terms added to the sum of the data terms by the intermediate coefficients, based on the color difference between adjacent pixels.
4. The method for fishery underwater image enhancement based on a multi-branch joint correction network according to claim 3 is characterized in that: The specific implementation process of step S3 is as follows: S3-1: The brightness high-sensitivity decomposition network adopts an encoder-decoder symmetrical design, with input First, the feature representation of the input image is obtained through the dilated convolution, and then the feature processing is performed in turn through the encoding part and the decoding part of the network; the encoding part obtains multi-scale image features through downsampling at different ratios, and the residual module RM is used to aggregate the feature information after each downsampling; The decoder part gradually restores the original size of the image by upsampling, and sets a skip connection structure between the encoder and decoder to achieve the combination of semantic information. It then adjusts the number of feature channels through convolution and outputs the illumination map and the reflection map. S3-2: The brightness correction network is built based on the symmetric feature pyramid of U-Net, including two paths, encoding and decoding, and uses skip connections to combine spatial and semantic features; a group attention method based on spatial correlation and channel statistics is constructed; Group attention consists of two parts: intra-group attention and inter-group attention. Intra-group attention is the concatenation of spatial attention and channel attention, while inter-group attention is self-attention. Intra-group attention is applied to the encoding path, and the features of each scale are pooled in the spatial dimension and channel dimension to generate attention maps in two dimensions for calibrating the original features; The inter-group attention is used to decode path semantic features, perform multi-level information aggregation, and finally output the corrected brightness map.
5. The method for fishery underwater image enhancement based on a multi-branch joint correction network according to claim 4 is characterized in that: The step S4 is specifically implemented as follows: the detail reconstruction network consists of two parts, namely, a symmetric coding and decoding network for extracting spatial domain features and a spatial-frequency domain information fusion module; the symmetric coding and decoding network consists of a convolution module, which is responsible for multi-scale feature extraction, and realizes the perception of image details under multiple receptive fields, and the detail features are spliced with the input image through jump connections to form an output; the spatial-frequency domain information fusion module consists of two branches, a convolution module and a wavelet reconstruction module, which respectively process the output of the symmetric coding and decoding network and the brightness channel L; the wavelet reconstruction module converts the information in the spatial domain to the wavelet domain for restoration through wavelet transform, and after wavelet transform, the input features are first transformed into four different sub-bands; The four sub-bands share the convolution layer for restoration, and finally the restored sub-bands are reconstructed into output features through inverse wavelet transform operation; The output of the wavelet reconstruction module is added to the output of the convolution branch at the pixel level to obtain the final detail enhancement result.
6. The method for fishery underwater image enhancement based on a multi-branch joint correction network according to claim 5, characterized in that: The artificial light source illumination area mask M is obtained by performing a difference operation on the illumination image before and after brightness correction and performing a binarization process, wherein pixels in an area greater than a threshold are set to 1, and the rest are set to 0; in addition, the part greater than the threshold is gathered to obtain the artificial light source illumination area mask M.
7. The method for fishery underwater image enhancement based on a multi-branch joint correction network according to claim 6, characterized in that: The specific implementation process of performing color restoration on the color channel AB under the guidance of the artificial light source illumination area mask M and outputting the enhanced color channel is as follows: S6-1: Extract the geometric and texture features of color channel AB through a lightweight convolutional network, and encode the mask M into a spatial weight map through a convolutional network and a Sigmoid function. , strengthen the local features of the illuminated area; Finally, cross-modal feature fusion is performed to perform channel cascade and spatial modulation on the color features and spatial weight map; S6-2: Design a color restoration network under the guidance of mask M. First, extract the features of AB channels and mask M for feature splicing, and use the auxiliary reference information provided by the features of mask M in the artificial light illumination area to guide color restoration; use the residual block as the basic feature extraction module to help network training; use the CT module to improve the deepest layer of the network. The CT module is a dual-branch parallel structure, which consists of the Transformer feature extraction branch and the convolutional feature extraction branch. After passing through the color restoration network, it outputs a color channel with color distortion eliminated.
Citation Information
Patent Citations
Laminated light guide using optical anisotropic film and planar light source device using same
CN110214286A
Underwater image enhancement method of multi-attention mechanism guided by brightness mask
CN116402715A
Two-stage network underwater image enhancement method based on color correction and multicolor space stretching
CN119168894A
Cited By
Multi-branch underwater image enhancement system for non-uniform illumination and color distortion
CN120953154A
Multi-branch underwater image enhancement system for non-uniform illumination and color distortion
CN120953154B
Method and system for identifying aquatic animals in culture water body, storage medium and equipment
CN121121811A
Electric power operation element segmentation method and device, equipment and storage medium
CN121190760A
Underwater image enhancement method based on brightness channel replacement and dual-scale fusion
CN121837095A