Underwater image enhancement method based on frequency-space cross-domain Transform and mixed cooperative representation
Through the frequency-space cross-domain Transformer and hybrid collaborative representation underwater image enhancement methods, the problems of underwater image blur and color deviation are solved, efficient image enhancement effect is achieved, and the visual quality and detail recovery of underwater images are improved.
Patent Information
- Application Number
- CN202510366616.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively deal with the problems of blur and color deviation in underwater images. The traditional methods have problems with overcorrection and oversaturation, while deep learning-based methods have limitations in global and local feature processing.
The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation is adopted. Through the U-shaped multi-scale learning module and the dual-stage hybrid collaborative representation module, the advantages of CNN and Transformer are combined to extract global and local features, and iteratively optimized, and trained using Charbonnier loss and perceptual loss.
It significantly improves the color casting problem of image, enhances the visual effect of underwater images, improves clarity and texture details, balances the processing of global and local information, and reduces the loss of feature information.
Smart Images

Figure CN120198306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater image processing, and in particular to an underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation. Background Art
[0002] When acquiring underwater image data, high-quality visual image data faces severe challenges. The main challenges include the severe constraints of water body scattering effect and absorption effect. Among them, the scattering effect causes the light propagation path to deviate, resulting in the underwater image becoming blurred, while the absorption effect gradually weakens the light with the increase of water depth, further reducing the image brightness and contrast. These adverse factors cause problems such as color deviation, low contrast, and blurred details, which widely affect the practical application of underwater images.
[0003] In the past decade, researchers have deeply explored a variety of underwater image enhancement strategies, which are roughly divided into two categories: traditional methods and deep learning-based methods. Among traditional methods, the method based on physical imaging models dominates, but its applicability to various underwater degradations is not ideal. Another widely studied traditional method is to use prior heuristic information, such as color balance, histogram stretching, and contrast improvement, to directly adjust pixel values to improve the visual quality of the image. However, the problems of overcorrection and over-saturation often plague these methods and may even further damage the already poor visual quality.
[0004] With the rapid development of deep learning technology, deep learning-based methods have opened up new paths for underwater image enhancement research. Underwater image enhancement methods based on CNN and Transformer have emerged continuously, but each has its limitations. Among them, CNN methods are often limited by the limited receptive field and are difficult to comprehensively capture global features and multi-scale features; while Transformer methods, although attracting much attention in the field of vision, its self-attention mechanism is good at capturing global features, but performs mediocrely in local feature processing, which to a certain extent limits the ability to process local details of underwater images. Although Transformer methods have shown certain potential in underwater applications, they mainly focus on the spatial domain.
[0005] Therefore, how to use new and efficient deep learning technologies for underwater image enhancement processing has become an urgent technical problem to be solved. Summary of the Invention
[0006] To solve the defect that the existing technology is difficult to enhance underwater images, the purpose of the present invention is to provide an underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation, which can significantly improve the problem of image color cast, thereby realizing the enhancement of the image and significantly enhancing the visual effect of underwater images.
[0007] To achieve the above object, the present invention adopts the following technical solutions: An underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation, the method comprising the following steps in sequence:
[0008] (1) Obtain the underwater images to form a training set and perform preprocessing;
[0009] (2) Based on the frequency-space cross-domain Transformer and hybrid collaborative representation, construct an underwater image enhancement model;
[0010] (3) Input the preprocessed underwater images into the underwater image enhancement model for training;
[0011] (4) Obtain the underwater image to be enhanced and perform preprocessing;
[0012] (5) Input the preprocessed underwater image to be enhanced into the trained underwater image enhancement model to obtain the underwater image enhancement result.
[0013] Step (1) specifically refers to: Select 4300 degraded underwater images and corresponding clear underwater images from three publicly available underwater datasets UIEB, UFO, and EUVP as the training set, and scale the images in the training set to a size of 256×256.
[0014] Step (2) specifically refers to: The underwater image enhancement model includes a U-shaped multi-scale learning module and a processing unit. The U-shaped multi-scale learning module is composed of 3 HRB modules connected in parallel, and the processing unit is composed of 3 HRB modules connected in series. The HRB module includes:
[0015] A hybrid feature extraction module for extracting global features and local features in the underwater image;
[0016] A two-stage hybrid collaborative representation module for fusing the global features and local features extracted by the hybrid feature extraction module;
[0017] The preprocessed underwater image first uses PatchEmbed patch embedding with multiple two-dimensional convolutions, and then embeds the features into a specific feature space through two-dimensional convolution for image preprocessing and extraction of initial features. Then, the extracted features are fed into the U-shaped multi-scale learning module. Using bilinear interpolation, the two HRB modules in the U-shaped multi-scale learning module scale the initial features to 1 / 2 and 1 / 4 of the original scale respectively, while the initial features of the third HRB module remain unchanged at the original scale. Then, the features at three different scales are upsampled, concatenated, and pointwise convolved. Then, a processing unit is used to continuously iteratively optimize the features. Finally, after a 3×3 convolution, the total loss of the underwater image enhancement model is calculated by combining the Charbonnier loss and the perceptual loss. The underwater image enhancement model is trained with error backpropagation and parameter updates, and finally the enhanced image is output.
[0018] Step (3) specifically refers to: The loss function of the underwater image enhancement model uses the Charbonnier loss and the perceptual loss. The weighting factor of the Charbonnier loss is set to 0.5, and the weighting factor of the perceptual loss is set to 2.5 to ensure the synchronous convergence speed.
[0019] 4300 images are used for training. 8 images are used in each batch, and all the images are used for training once in each epoch, for a total of 538 batches. The Adam optimizer is used to optimize the network, with β1 = 0.9, β2 = 0.999, and the initial learning rate is 4e-4. The learning rate is set to decay to 0.8 every 60 epochs, and the training ends after 300 epochs.
[0020] The hybrid feature extraction module includes a first branch and a second branch. The first branch includes four spatio-frequency cross-domain Transformer modules, a first feature attention module, a strided convolution, a 1×1 pointwise convolution, and a UP transposed convolution. The second branch uses a second feature attention module. The input of the hybrid feature extraction module is the original image, and the output of the hybrid feature extraction module is the global feature map and the local feature map.
[0021] The structures of the first feature attention module and the second feature attention module are the same. The first feature attention module includes two depthwise separable convolutions, a PReLU activation function, a pointwise convolution, and a CA channel attention. The first feature attention module first decouples and extracts the input features in the spatial and channel dimensions through a depthwise separable convolution, then applies the PReLU activation function, and further refines and abstracts the features through a depthwise separable convolution. Finally, before entering the channel attention, a skip connection containing a pointwise convolution is introduced.
[0022] The two-stage hybrid collaborative representation module includes a hybrid attention refinement module and a hybrid fusion processing module.
[0023] The frequency-space cross-domain Transformer module includes a frequency-domain feature extraction module, a frequency-space fusion attention module, a FFN feed-forward neural network, and two LN layer normalizations. The input of the frequency-space cross-domain Transformer module is the initial spatial-domain feature map, and the output is the cross-domain feature map.
[0024] The frequency-domain feature extraction module includes a two-dimensional real fast Fourier transform module, a two-dimensional real inverse fast Fourier transform module, and two CIP modules. The CIP module consists of a 1×1Conv convolution, IN instance normalization, and a PReLU activation function. First, the input feature is transformed from the intuitive spatial domain to the frequency domain using the two-dimensional real fast Fourier transform to obtain the initial frequency-domain feature map. Then, through the first CIP module and using skip connections, the initial frequency-domain feature map is added element-wise to the feature map extracted by the first CIP module. Subsequently, through the second CIP module, and then the features refined in the frequency domain are transformed back to the spatial domain using the two-dimensional real inverse fast Fourier transform, and the frequency-domain feature map is output.
[0025] The frequency-space fusion attention module includes three depthwise separable convolutions, three CA channel attentions, a PReLU activation function, a pointwise convolution, and a DOB dynamic optimization module. The input of the frequency-space fusion attention module is the frequency-domain feature map and the spatial-domain feature map normalized by one LN layer, and the output is the frequency-space feature map. The input spatial-domain feature map is processed in parallel through three depthwise separable convolutions to obtain three corresponding spatial-domain feature maps Q, K, and V, and then through three CA channel attentions. The frequency-domain feature map is passed through a PReLU activation function and multiplied pointwise with the spatial-domain feature maps Q, K, and V to adaptively adjust the frequency-domain feature map and the spatial-domain feature maps Q, K, and V to fit each other, obtaining the frequency-space feature maps Q G , frequency-space feature map K G and frequency-space feature map V G ; Subsequently, the frequency-space feature map Q G and the frequency-space feature map K G are matrix-multiplied, and through the DOB dynamic optimization module, redundant features are filtered and important features are retained to obtain the optimized feature map. Finally, after matrix-multiplying the optimized feature map and the frequency-space feature map V G and passing through a pointwise convolution, the final frequency-space feature map is obtained.
[0026] The hybrid attention refinement module includes a first gating component, a second gating component, and an addition operation; the first gating component consists of a GAP global average pooling, an AKC adaptive convolution, an IN instance normalization, and a Sigmoid function, and the second gating component consists of a 1×1 convolution and a Sigmoid function;
[0027] Input the global feature map J s into the first gating component for deep refinement extraction to obtain a deep global feature map; input the local feature map J c into the second gating component for cross-channel shallow refinement extraction to obtain a shallow local feature map, and perform an addition operation on the deep global feature map and the shallow local feature map to obtain a refined global feature map T c ;
[0028] Similarly, input the local feature map J c into the first gating component for deep refinement extraction to obtain a deep local feature map; input the global feature map J s into the second gating component for cross-channel shallow refinement extraction to obtain a shallow global feature map, and perform an addition operation on the deep local feature map and the shallow global feature map to obtain a refined local feature map T s 。
[0029] The hybrid fusion processing module includes two subtraction operations and two addition operations;
[0030] First, the global feature map J s is subtracted from the refined global feature map T c through the first subtraction operation, and then through the first addition operation, an addition is performed with the refined local feature map T s to obtain a hybrid global feature map Z s ; similarly, the local feature map J c is subtracted from the refined local feature map T s through the second subtraction operation, and then through the second addition operation, an addition is performed with the refined global feature map T c to obtain a hybrid local feature map Z c 。
[0031] The DOB dynamic optimization module includes a squared PReLU activation function, a SoftMax activation function, and two normalized weights W;
[0032] First, the DOB dynamic optimization module adopts a parallel connection method. The input feature map passes through an SQB sparse lower branch composed of a squared PReLU activation function and a normalized weight W, and a DB dense upper branch composed of a SoftMax activation function and a normalized weight W respectively. The normalized weight is used to adaptively modulate and fuse these two branches to obtain an optimized feature map.
[0033] As can be seen from the above technical solutions, the beneficial effects of the present invention are as follows: First, the present invention can make more effective use of global content with the help of the frequency-space cross-domain Transformer module. By superimposing the information in the spatial domain and the frequency domain and putting it into the self-attention layer of the Transformer, it can make full use of the features of the image in different domains, capture these two cross-domain features through the self-attention mechanism, and significantly improve the problem of image color cast, thereby realizing the enhancement of the image. Second, the present invention constructs a two-stage hybrid collaborative representation module. This module extracts the features of the underwater image through the interaction mechanism of the upper and lower branches. For the output features of the hybrid feature extraction module, a two-stage strategy is adopted. First, fine-grained hybrid processing is performed, and then hybrid fusion is carried out. This process effectively improves the image texture details and significantly enhances the visual effect of the underwater image. Third, the present invention helps to reduce the loss of feature information, improve the clarity, improve the image texture details, and correct the color deviation. Fourth, the present invention combines the convolutional neural network and the Transformer part. Since the Transformer shows excellent capabilities in remote interaction modeling and related feature learning, and the feature extraction module composed of CNN is good at extracting local features, focusing on local representation, and noise removal, it can capture the non-local distribution of the main structure and learn the local details of the image at the same time, balancing the processing of global and local information. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is the flowchart of the method of the present invention;
[0035] Figure 2 is the structural schematic diagram of the underwater image enhancement model of the present invention;
[0036] Figure 3 is the structural schematic diagram of the hybrid feature extraction module of the present invention;
[0037] Figure 4 is the structural schematic diagram of the first feature attention module of the present invention;
[0038] Figure 5 is the structural schematic diagram of the frequency-space cross-domain Transformer module of the present invention;
[0039] Figure 6 is the structural schematic diagram of the frequency domain feature extraction module of the present invention;
[0040] Figure 7 It is a schematic structural diagram of the frequency-space fusion attention module in the present invention;
[0041] Figure 8 It is a schematic structural diagram of the DOB dynamic optimization module in the present invention;
[0042] Figure 9 It is a schematic structural diagram of the two-stage hybrid collaborative representation module in the present invention;
[0043] Figure 10 It is a schematic structural diagram of the hybrid attention refinement module in the present invention;
[0044] Figure 11 It is a histogram comparison diagram of the HLS color space of the present invention and other methods on a paired dataset;
[0045] Figure 12 It is a comparison diagram of the present invention and other methods on an unpaired dataset;
[0046] Figure 13 It is a comparison diagram of the present invention and other methods on underwater color card images
[0047] Figure 14 It is a result demonstration diagram of different methods for underwater target detection using YOLOv5 in the present invention;
[0048] Figure 15 It is a result demonstration diagram of different methods for salient object detection in the present invention;
[0049] Figure 16 It is an application test data display diagram of the present invention;
[0050] Figure 17 It is an enhanced visual effect diagram of the present invention for haze, low light, sandstorm and remote sensing images. Detailed implementation manners
[0051] As Figure 1 shown, an underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation, the method includes the following steps in sequence:
[0052] (1) Obtain underwater images to form a training set and perform preprocessing;
[0053] (2) Based on the frequency-space cross-domain Transformer and hybrid collaborative representation, construct an underwater image enhancement model;
[0054] (3) Input the preprocessed underwater images into the underwater image enhancement model for training;
[0055] (4) Obtain the underwater images to be enhanced and perform preprocessing;
[0056] (5) Input the pre - processed underwater image to be enhanced into the trained underwater image enhancement model to obtain the underwater image enhancement result.
[0057] Specifically, step (1) means: Select 4300 degraded underwater images and their corresponding clear underwater images from three publicly available underwater datasets UIEB, UFO, and EUVP as the training set, and scale the images in the training set to a size of 256×256.
[0058] As Figure 2 shown, specifically, step (2) means: The underwater image enhancement model includes a U - shaped multi - scale learning module and a processing unit. The U - shaped multi - scale learning module is composed of 3 HRB modules in parallel, and the processing unit is composed of 3 HRB modules in series. The HRB module includes:
[0059] A hybrid feature extraction module for extracting global features and local features in the underwater image;
[0060] A two - stage hybrid collaborative representation module for fusing the global features and local features extracted by the hybrid feature extraction module;
[0061] The pre - processed underwater image first uses PatchEmbed patches of multiple two - dimensional convolutions, and then embeds the features into a specific feature space through two - dimensional convolutions for image pre - processing and extraction of initial features. Then, the extracted features are fed into the U - shaped multi - scale learning module. Using bilinear interpolation, two HRB modules in the U - shaped multi - scale learning module scale the initial features to 1 / 2 and 1 / 4 of the original scale respectively, while the initial features of the third HRB module remain unchanged at the original scale. Then, the features at three different scales are upsampled, concatenated, and point - wise convoluted. Then, the processing unit is used to continuously iteratively optimize the features. Finally, after 3×3 convolution, the total loss of the underwater image enhancement model is calculated by combining the Charbonnier loss and the perceptual loss. The underwater image enhancement model is trained by error backpropagation and parameter updated, and finally the enhanced image is output.
[0062] Specifically, step (3) means: The loss function of the underwater image enhancement model uses the Charbonnier loss and the perceptual loss. The weighting factor of the Charbonnier loss is set to 0.5, and the weighting factor of the perceptual loss is set to 2.5 to ensure the synchronous convergence speed. The Charbonnier loss can reduce the blurring phenomenon that occurs during image reconstruction. It can help reduce the difference between the reconstructed image and the reference image while minimizing the image blurring effect. The perceptual loss can better capture the high - level features of the image and improve the quality of the generated image.
[0063] The network is trained using 4300 images. 8 images are used in each batch, and all the images are used for one pass of training in each epoch, resulting in a total of 538 batches. The Adam optimizer is used to optimize the network, with β1 = 0.9, β2 = 0.999, and an initial learning rate of 4e-4. The learning rate is set to decay to 0.8 every 60 epochs, and the training ends after 300 epochs.
[0064] As Figure 3 shown, the hybrid feature extraction module includes a first branch and a second branch. The first branch includes four spatio-temporal cross-domain Transformer modules, a first feature attention module, a strided convolution, a 1×1 pointwise convolution, and a UP transposed convolution. The second branch uses a second feature attention module. The input of the hybrid feature extraction module is the original image, and the outputs are the global feature map and the local feature map.
[0065] The hybrid feature extraction module is constructed as a hybrid two-branch network to learn the representations of the main structure and local details of the image, and to explore the mutual cooperative representations between the main structure and local details by leveraging the advantages of Transformer and CNN. In the first branch of the hybrid feature extraction module, four cascaded FSCDT spatio-temporal cross-domain Transformer modules are mainly used to extract the main structure features of the underwater image, and a FAB feature attention module is used to extract the local features of the underwater image. Specifically, a Down strided convolution is used for downsampling before entering the four cascaded spatio-temporal cross-domain Transformer modules. The downsampled feature map is concatenated with the feature map after passing through the four cascaded FSCDT spatio-temporal cross-domain Transformer modules. Then, a 1×1 pointwise convolution is used to adjust the number of channels. Finally, a UP transposed convolution is used for upsampling to form an efficient U-shaped structure, and element-wise addition is used to add it to the skip connection containing the FAB feature attention module, and the output is the global feature map. In the second branch of the hybrid feature extraction module, only a FAB feature attention module, i.e., the second feature attention module, is used to extract the local features of the underwater image, and the output is the local feature map. The structures of the first and second branches enable the hybrid feature extraction module to capture the non-local distribution of the main structure while learning to enhance the local details of the image, balancing the processing of global and local information.
[0066] The structures of the first feature attention module and the second feature attention module are the same, as Figure 4As shown, the first feature attention module includes two depthwise separable convolutions, a PReLU activation function, a pointwise convolution, and a CA channel attention. The first feature attention module first decouples and extracts the input features in the spatial and channel dimensions through a depthwise separable convolution, then applies the PReLU activation function, and further refines and abstracts the features through another depthwise separable convolution, not only deepening the depth of feature extraction but also improving the robustness of the features. Finally, before entering the channel attention, a skip connection containing a pointwise convolution is introduced.
[0067] As Figure 9 shown, the two-stage hybrid collaborative representation module includes a hybrid attention refinement module and a hybrid fusion processing module.
[0068] As Figure 5 shown, the frequency-space cross-domain Transformer module includes a frequency-domain feature extraction module, a frequency-space fusion attention module, an FFN feed-forward neural network, and two LN layer normalizations. The input of the frequency-space cross-domain Transformer module is the initial spatial-domain feature map, and the output of the frequency-space cross-domain Transformer module is the cross-domain feature map.
[0069] As Figure 6 shown, the frequency-domain feature extraction module includes a two-dimensional real fast Fourier transform module, a two-dimensional real inverse fast Fourier transform module, and two CIP modules. The CIP module consists of a 1×1Conv convolution, IN instance normalization, and a PReLU activation function. First, the input features are transformed from the intuitive spatial domain to the frequency domain using the two-dimensional real fast Fourier transform to obtain the initial frequency-domain feature map. Then, through the first CIP module and using a skip connection, the initial frequency-domain feature map is added to the feature map extracted by the first CIP module through element-wise addition. Subsequently, through the second CIP module, and then the features refined in the frequency domain are transformed back to the spatial domain using the two-dimensional real inverse fast Fourier transform, and the frequency-domain feature map is output. The frequency-domain feature extraction module can capture complex structures and subtle differences that are difficult to directly observe in the original spatial domain, thus greatly enriching the dimension and depth of feature extraction.
[0070] As Figure 7As shown, the frequency-space fusion attention module includes three depthwise separable convolutions, three CA channel attentions, a PReLU activation function, a pointwise convolution, and a DOB dynamic optimization module. The input of the frequency-space fusion attention module is a frequency-domain feature map and a space-domain feature map normalized by an LN layer, and the output is a frequency-space feature map. The input space-domain feature map is processed by three depthwise separable convolutions in parallel to obtain three corresponding space-domain feature maps Q, K, and V, and then passed through three CA channel attentions. The frequency-domain feature map is passed through a PReLU activation function and multiplied pointwise with the space-domain feature maps Q, K, and V to adaptively adjust the frequency-domain feature map and the space-domain feature maps Q, K, and V to fit each other, obtaining the frequency-space feature maps Q G , the frequency-space feature map K G and the frequency-space feature map V G ; Subsequently, the frequency-space feature map Q G and the frequency-space feature map K G are subjected to matrix multiplication, and through the DOB dynamic optimization module, redundant features are filtered and important features are retained to obtain an optimized feature map. Finally, after the optimized feature map and the frequency-space feature map V G are subjected to matrix multiplication, a pointwise convolution is performed to obtain the final frequency-space feature map. The depthwise separable convolution operation is used to aggregate spatial context information. At the same time, in order to better extract the feature information of the space domain, CA channel attention is added, so that the model can adaptively focus on more important channel information during the process of feature extraction in the space domain.
[0071] As Figure 10 shown, the hybrid attention refinement module includes a first gating component, a second gating component, and an addition operation; the first gating component consists of a GAP global average pooling, an AKC adaptive convolution, an IN instance normalization, and a Sigmoid function, and the second gating component consists of a 1×1 convolution and a Sigmoid function;
[0072] The global feature map J s is input into the first gating component for deep refinement extraction to obtain a deep global feature map; the local feature map J c is input into the second gating component for cross-channel shallow refinement extraction to obtain a shallow local feature map. The deep global feature map and the shallow local feature map are added to obtain a refined global feature map T c ;
[0073] Similarly, the local feature map J c is input into the first gating component for deep refinement extraction to obtain a deep local feature map; the global feature map J sInput the second gating component to perform shallow refinement extraction across channels to obtain a shallow global feature map, add the deep local feature map and the shallow global feature map to obtain a refined local feature map T s .
[0074] The hybrid fusion processing module includes two subtraction operations and two addition operations;
[0075] First, the global feature map J s The refined global feature map T is subtracted by the first subtraction operation c , and then through the first addition operation, and refine the local feature map T s Do addition to get the mixed global feature map Z s ; Similarly, the local feature map J c The refined local feature map T is subtracted by the second subtraction operation s , and then through the second addition operation, and refine the global feature map T c Do addition to get the mixed local feature map Z c .
[0076] like Figure 8 As shown, the DOB dynamic optimization module includes a square PReLU activation function, a SoftMax activation function and two normalized weights W;
[0077] First, the DOB dynamic optimization module adopts a parallel approach. The input feature map passes through the SQB sparse lower branch composed of a squared PReLU activation function and a normalized weight W, and the DB dense upper branch composed of a SoftMax activation function and a normalized weight W. The normalized weight is used to adaptively modulate and fuse the two branches to obtain the optimized feature map.
[0078] from Figure 11 It can be seen that, with the ground truth as a reference, some methods used for comparison have limited quality improvement, and some methods have obvious quality improvement, but will cause over-enhancement or wrong color correction, and the image enhanced by the model proposed in the present invention has color balance, high contrast and better visual effect. In the paired image test, the present invention shows very good performance, almost the same as the ground truth. From Table 1, it can be seen that the PSNR, LPIPS and TM indicators of the present invention on the paired test groups of Test-U90, Test-UFO120 and Test-S185 of the test set are the best, and compared with other deep learning comparison methods, the comprehensive indicators of all paired image test groups are the best, which is due to the superiority of the frequency-space cross-domain Transformer module in the present invention in extracting global and local information and texture features of images.
[0079] It should be noted that traditional methods often lead to undesired appearance changes, such as additional distortions (as shown by MMLE in Figure 11 ) and over-enhancement (as shown by CBLA in Figure 11 ). This is mainly because they adopt inappropriate priors to simulate mismatched degradation processes and lack the constraints imposed by the learning process. Although these over-enhanced effects do not conform to human visual preferences, their high contrast, detail enhancement, and channel value balance usually favor non-reference metrics, resulting in higher scores. To deeply explore this phenomenon, three traditional methods from different years were specifically selected for comparison, and it was found that with the update of the years, the non-reference metrics of traditional methods showed an upward trend. However, compared with deep learning methods, the full-reference metrics of traditional methods are still relatively low. Traditional methods include ULAP, MMLE, and CBLA, and deep learning methods include UIEC2Net, STSC, Ushape, PUGAN, WPFNet, CCMSRNet, and CDCRNet. Table 1 shows the test results of CDCRNet and other methods using reference metrics on the paired image test set in the present invention. The test results include four full-reference metric-based measures and six non-reference metric-based measures. The four full-reference metrics include SSIM, PSNR, LPIPS, and MSE (×10 2 ), and the six non-reference metrics include NIQE, UCIQE, U, FDUM, TM, and URanker.
[0080] Table 1
[0081]
[0082]
[0083] From Figure 12 and Figure 13It can be seen that the present invention achieves the best effect both in terms of color correction and blurring elimination, making the enhanced image have richer colors, higher contrast, and obvious details. The comparison results of numerical indicators are shown in Table 2. As can be seen from Table 2, the present invention has better performance. From the average results of the four test groups, it can be seen that in the Test-C60 and Test-UCCS test sets, the present invention ranks first in three non-reference indicators, and the number of non-reference indicators ranking first accounts for half of the total number of non-reference indicators. Compared with other deep learning methods, the comprehensive indicators are all the best. In the Test-UIQS and Test-S16 test sets, the present invention ranks first in two non-reference indicators. Compared with other deep learning methods, the comprehensive indicators are all the best. This benefits from the role of the two-stage hybrid collaborative representation module, which extracts color features and texture features of different attention regions, thereby achieving information complementarity between color features and detailed texture features. It shows strong performance in the underwater environment with severe color differences, not only being able to provide robust color correction, avoiding unnecessary color deviation, but also enhancing texture clarity while restoring colors. Table 2 shows the test results of CDCRNet in the present invention and other methods using non-reference indicators on the unpaired image test set.
[0084] Table 2
[0085]
[0086]
[0087] Table 3 shows the PSNR, FLOPs, Params, and Runtimes for comparing different methods on Test-S185. As shown in Table 3, CDCRNet in the present invention compares its performance in terms of the number of floating-point operations, the number of parameters, and the running time. CDCRNet shows the highest peak signal-to-noise ratio (PSNR: 27.691), which indicates that it has excellent reconstruction advantages in underwater image enhancement. Although CDCRNet is not the best in terms of FLOPs, Parameters, and Runtimes, it is not considered that this hinders the feasibility and practicality of CDCRNet. Since the underwater image enhancement model in the present invention, namely CDCRNet, adopts multiple self-attention mechanisms, it is higher than some comparison methods in terms of FLOPs. However, among all the comparison methods, CDCRNet still has an obvious advantage in terms of the number of parameters. CDCRNet has significant competitiveness in balancing the number of parameters and quantitative indicators. In addition, the comprehensive experimental results show that, measured according to image quality indicators and subjective perception, CDCRNet is superior to other methods in overall performance.
[0088] Table 3
[0089]
[0090] UIE is often used to mitigate the impact of harsh and complex underwater conditions on captured image data and ensure the effective execution of subsequent downstream tasks. Therefore, several representative methods are selected to verify the enhanced performance of downstream vision tasks such as target detection, salient object detection, and other low-level vision tasks.
[0091] For target detection, Yolov 5 is trained on the underwater object detection dataset URPC. This dataset contains four types of targets: starfish, sea urchins, scallops, and sea cucumbers. Figure 14 As shown in the figure, the images enhanced by CDCRNet in the present invention show significant advantages in detection effect, and can identify more sea urchins and starfish. In terms of confidence, compared with the images enhanced by CCMSRNet, the images enhanced by CDCRNet maintain a similar confidence level, but more importantly, the number of targets detected in the enhanced images of the present invention is greater.
[0092] For salient object detection, the present invention is tested on the benchmark USOD dataset. Figure 15 As shown in FIG. 1 , compared with other methods, the saliency map generated by CDCRNet in the present invention is more complete in structure and has more precise boundaries. Figure 16 The left figure in the figure shows the evaluation results of CDCRNet in terms of F-measure, MaxF-measure and S-measure. Obviously, CDCRNet outperforms other representative methods in these three indicators.
[0093] In order to further evaluate the enhanced performance of CDCRNet in this invention on other low-level visual tasks, we try to enhance Figure 17 Four types of low-visibility images. It can be clearly seen that CDCRNet shows excellent enhancement effects on these images. It is particularly worth mentioning that when facing the challenge of low-light images, the subjective enhancement visual effect of CDCRNet is particularly significant. It not only successfully weakens the overall dim effect of the image caused by insufficient light, but also restores the color information and detail features that were originally obscured in the image. This improvement not only makes the image visually brighter and clearer, but also provides a more reliable and rich information basis for subsequent image processing or analysis tasks. In addition, for other types of low-visibility images, such as haze images, sandstorm images, and remote sensing images with low contrast, CDCRNet also shows satisfactory enhancement effects. It effectively reduces the blur and noise in the image, improves the contrast and overall visual effect of the image, and further verifies the wide applicability and strong strength of CDCRNet in enhancing the performance of low-level visual tasks.
[0094] likeFigure 16 As shown in the right figure, the present invention uses the natural image quality evaluator (NIQE) as an objective evaluation index. It can be found that the CDCRNet of the present invention has improved the indexes of these four kinds of low visibility images to varying degrees. Among them, the low light and remote sensing images are the most prominent, with improvements of 15.47% and 19.95% respectively, which is consistent with the improvement effect of the subjective effect image. It can be seen that whether it is applied to object detection, salient object detection and other low-level vision tasks, the present invention has been improved, and the improvement effect is better than other methods.
[0095] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.
Claims
1. An underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation, characterized by: The method comprises the following steps in order: (1) Obtain underwater images to form a training set and perform preprocessing; (2) Construct an underwater image enhancement model based on frequency-space cross-domain Transformer and hybrid collaborative representation; (3) Inputting the preprocessed underwater image into the underwater image enhancement model for training; (4) Acquire the underwater image to be enhanced and perform preprocessing; (5) The preprocessed underwater image to be enhanced is input into the trained underwater image enhancement model to obtain the underwater image enhancement result.
2. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 1 is characterized in that: Step (1) specifically refers to: selecting 4300 degraded underwater images and corresponding clear underwater images from three public underwater datasets UIEB, UFO and EUVP as training sets, and scaling the images in the training sets to a size of 256×256.
3. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 1 is characterized in that: Step (2) specifically means that the underwater image enhancement model includes a U-shaped multi-scale learning module and a processing unit, the U-shaped multi-scale learning module is composed of three HRB modules connected in parallel, the processing unit is composed of three HRB modules connected in series, and the HRB module includes: A hybrid feature extraction module for extracting global and local features from underwater images; A two-stage hybrid collaborative representation module is used to fuse the global features and local features extracted by the hybrid feature extraction module; The preprocessed underwater image is first embedded using PatchEmbed patches composed of multiple two-dimensional convolutions, and then the features are embedded into a specific feature space through two-dimensional convolution. The image is preprocessed and the initial features are extracted. The extracted features are then sent to the U-shaped multi-scale learning module. Using bilinear interpolation, the two HRB modules in the U-shaped multi-scale learning module scale the initial features to 1 / 2 and 1 / 4 of the original scale respectively, while the initial features of the third HRB module remain unchanged at the original scale. The features of three different scales are then upsampled, concatenated and convolved point by point, and the processing unit is used to continuously iteratively optimize the features. Finally, after 3×3 convolution, the total loss of the underwater image enhancement model is calculated by combining the Charbonnier loss and the perceptual loss. The underwater image enhancement model is trained with error back propagation and parameter updates, and the enhanced image is output.
4. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 1 is characterized in that: Step (3) specifically means: the loss function of the underwater image enhancement model uses Charbonnier loss and perceptual loss, the weighting factor of Charbonnier loss is set to 0.5, and the weighting factor of perceptual loss is set to 2.5 to ensure synchronous convergence speed; 4300 images were used for training, 8 images were used in each batch, and all images were used for training in each epoch, for a total of 538 batches. The network was optimized using the Adam optimizer, with β1=0.9, β2=0.999, an initial learning rate of 4e-4, and the learning rate was set to decay to 0.8 every 60 epochs. The training ended after 300 epochs.
5. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 3 is characterized by: The hybrid feature extraction module includes a first branch and a second branch, the first branch includes four frequency-space cross-domain Transformer modules, a first feature attention module, a strided convolution, a 1×1 point-by-point convolution and a UP transposed convolution, and the second branch uses a second feature attention module; The input of the hybrid feature extraction module is the original image, and the output of the hybrid feature extraction module is a global feature map and a local feature map; The structure of the first feature attention module and the second feature attention module is the same. The first feature attention module includes two depth-wise separable convolutions, a PReLU activation function, a point-wise convolution, and a CA channel attention. The first feature attention module first performs a preliminary decoupling and extraction of the input features in terms of spatial and channel dimensions through a depthwise separable convolution, then applies the PReLU activation function, and further refines and abstracts the features through a depthwise separable convolution. Finally, before entering the channel attention, a skip connection with pointwise convolution is introduced. The two-stage hybrid collaborative representation module includes a hybrid attention refinement module and a hybrid fusion processing module.
6. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 5 is characterized in that: The frequency-space cross-domain Transformer module includes a frequency-domain feature extraction module, a frequency-space fusion attention module, an FFN feedforward neural network and two LN layer normalizations. The input of the frequency-space cross-domain Transformer module is the initial spatial domain feature map, and the output of the frequency-space cross-domain Transformer module is the cross-domain feature map. The frequency domain feature extraction module includes a two-dimensional real fast Fourier transform module, a two-dimensional real inverse fast Fourier transform module, and two CIP modules. The CIP module consists of 1×1Conv convolution, IN instance normalization, and PReLU activation function. First, the input features are converted from the intuitive spatial domain to the frequency domain by using the two-dimensional real fast Fourier transform to obtain an initial frequency domain feature map. Then, through the first CIP module, a jump connection is used to add the initial frequency domain feature map to the feature map extracted by the first CIP module by element-by-element addition. Subsequently, through the second CIP module, the features extracted in the frequency domain are converted back to the spatial domain by using the two-dimensional real inverse fast Fourier transform to output the frequency domain feature map. The frequency-space fusion attention module includes three depth-separable convolutions, three CA channel attentions, a PReLU activation function, a point-by-point convolution and a DOB dynamic optimization module. The input of the frequency-space fusion attention module is a frequency domain feature map and a spatial domain feature map normalized by an LN layer, and the output is a frequency-space feature map. The input spatial domain feature map is processed by three depth-separable convolutions in parallel to obtain three corresponding spatial domain feature maps Q, K and V, respectively, followed by three CA channel attentions. The frequency domain feature map passes through a PReLU activation function and is point-by-point multiplied with the spatial domain feature maps Q, K, and V, and the frequency domain feature map and the spatial domain feature maps Q, K, and V are adaptively adjusted to fit each other to obtain the frequency-space feature map Q G , frequency-space characteristic diagram K G Sum frequency space characteristic diagram V G ; Then, the frequency-space characteristic map Q G Sum frequency space characteristic map K G Perform matrix multiplication, filter redundant features and retain important features through the DOB dynamic optimization module, and obtain the optimized feature map. Finally, the optimized feature map and the frequency-space feature map V G After matrix multiplication, a point-by-point convolution is performed to obtain the final frequency-space feature map.
7. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 5 is characterized by: The hybrid attention refinement module includes a first gating component, a second gating component and an addition operation; the first gating component consists of a GAP global average pooling, an AKC adaptive convolution, an IN instance normalization and a Sigmoid function, and the second gating component consists of a 1×1 convolution and a Sigmoid function; The global feature map J s Input the first gating component, perform deep refinement extraction, and obtain a deep global feature map; convert the local feature map J c Input the second gating component to perform shallow refinement extraction across channels to obtain a shallow local feature map. Add the deep global feature map and the shallow local feature map to obtain a refined global feature map T c ; Similarly, the local feature map J c Input the first gating component, perform deep refinement extraction, and obtain a deep local feature map; convert the global feature map J s Input the second gating component to perform shallow refinement extraction across channels to obtain a shallow global feature map, add the deep local feature map and the shallow global feature map to obtain a refined local feature map T s .
8. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 5, characterized in that: The hybrid fusion processing module includes two subtraction operations and two addition operations; First, the global feature map J s The refined global feature map T is subtracted by the first subtraction operation c , and then through the first addition operation, and refine the local feature map T s Do addition to get the mixed global feature map Z s ; Similarly, the local feature map J c The refined local feature map T is subtracted by the second subtraction operation s , and then through the second addition operation, and refine the global feature map T c Do addition to get the mixed local feature map Z c .
9. The underwater image enhancement method based on frequency-space cross-domain Transformer and hybrid collaborative representation according to claim 6, characterized in that: The DOB dynamic optimization module includes a squared PReLU activation function, a SoftMax activation function and two normalized weights W; First, the DOB dynamic optimization module adopts a parallel approach. The input feature map passes through the SQB sparse lower branch composed of a squared PReLU activation function and a normalized weight W, and the DB dense upper branch composed of a SoftMax activation function and a normalized weight W. The normalized weight is used to adaptively modulate and fuse the two branches to obtain the optimized feature map.