Underwater image enhancement method based on space-frequency feature interaction

Through the underwater image enhancement method of space-frequency feature interaction, combined with the amplitude subnet and the phase subnet, the problems of global consistency and detail fidelity in underwater image enhancement are solved, and high-quality underwater images are generated to adapt to complex environments.

CN120278894APending Publication Date: 2025-07-08DALIAN MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510250431.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing underwater image enhancement methods are difficult to take into account global consistency, detail fidelity and color accuracy in complex and changeable underwater environments. The effects of traditional methods are unstable, and deep learning methods lack global feature capture and frequency domain information utilization.

Method used

The underwater image enhancement method based on space-frequency feature interaction is adopted, and the amplitude subnet and phase subnet combine with the degraded feature extraction module, and the air-domain and frequency-domain feature interaction module is used to introduce the frequency domain loss function for model constraints, and a two-stage network model is constructed.

Benefits of technology

Generate images with bright colors, clear textures, and high contrast, with high generalization ability and robustness, and adapt to diverse and complex underwater environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278894A_ABST
    Figure CN120278894A_ABST
Patent Text Reader

Abstract

The invention discloses an underwater image enhancement method based on space-frequency feature interaction, and the method achieves the high-quality enhancement of an underwater image through the deep learning technology. The underwater image enhancement network mainly comprises an amplitude sub-network, a phase sub-network and a degradation feature extraction module, and a space-frequency feature interaction module is introduced. The amplitude sub-network is used for improving color and contrast features of an underwater image, the phase sub-network is used for recovering texture details and structure information of the image, and the degradation feature extraction module is used for pertinently guiding an enhancement process of degradation features. The space-frequency feature interaction module realizes complementary learning of global and local features, and solves the problem of lack of global consistency caused by convolution operation locality. Besides, a special frequency domain loss function is designed in the method, frequency components are constrained, and image artifacts caused by frequency deviation are suppressed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater image enhancement and restoration. Specifically, it particularly relates to an underwater image enhancement method based on spatio-frequency feature interaction. Background Art

[0002] The unknown and mysterious underwater environment and rich marine resources attract people to continuously explore. However, due to the absorption and scattering of light by water, the quality of underwater images is often severely degraded, manifested as problems such as color shift, reduced contrast, and blurred details. These degradation phenomena not only affect people's research on marine organisms and geological structures but also restrict the application effects of autonomous underwater vehicles (such as AUVs, ROVs) in tasks such as ship maintenance, ocean exploration, and underwater search and rescue. Therefore, developing effective underwater image enhancement methods is of great significance for improving underwater visual perception ability and promoting the development of marine science.

[0003] Currently, underwater image enhancement methods are mainly divided into traditional methods and deep learning-based methods. Traditional methods include non-physical model methods and physical model-based methods.

[0004] Non-physical model methods usually directly process the pixel values of underwater images, such as techniques like histogram equalization, color correction, and contrast stretching. These methods are simple to implement and have high computational efficiency, and can improve the visual quality of images to a certain extent. However, due to the lack of physical constraints on the underwater imaging process, the enhancement effects of such methods are often unstable, prone to problems such as over-enhancement, detail loss, or color distortion, and have poor generalization ability in different underwater environments.

[0005] Physical model-based methods rely on specific underwater imaging models and reverse restore the image before degradation by estimating parameters such as illumination, scattering, and medium absorption. These methods can theoretically restore the original scene more accurately. However, the actual underwater environment is complex and variable, and it is difficult to accurately obtain the parameters in the imaging model, and the prior assumptions often do not hold, resulting in new distortions in the enhancement results, such as over-sharpening, color deviation, or artifacts.

[0006] With the rapid development of deep learning technology, underwater image enhancement methods based on deep learning have gradually become a research hotspot. These methods use a large amount of data to train neural networks, automatically learn the complex features of underwater images, and achieve end-to-end enhancement processing. To a certain extent, deep learning methods overcome the dependence on prior assumptions and handcrafted features of traditional methods, showing strong adaptability and stability. However, most existing deep learning methods are based on convolutional neural networks (CNNs). Limited by the local characteristics of convolutional operations, it is difficult to effectively capture the global features and long-range dependencies of images, resulting in the enhancement results may lack global consistency. In addition, many methods mainly focus on the extraction of spatial domain features, ignoring the role of frequency domain information, which is prone to image artifacts and detail loss.

[0007] In summary, existing underwater image enhancement methods still face many challenges when dealing with complex and variable underwater environments, and it is difficult to balance global consistency, detail fidelity, and color accuracy simultaneously. Therefore, there is an urgent need for an underwater image enhancement method that can effectively combine spatial and frequency domain features, taking into account both global and local information, to improve the overall quality of underwater images and meet the needs of practical applications. Summary of the Invention

[0008] In view of the above technical problem that existing convolutional neural networks are difficult to accurately model long-range dependencies and global feature distributions of images, an underwater image enhancement method based on spatio-frequency feature interaction is provided. The network of the present invention includes three parts: an amplitude subnet, a phase subnet, and a degradation feature extraction module. The amplitude subnet and the phase subnet correspond to two subtasks of image style restoration and semantic enhancement, respectively. The degradation feature extraction module is used to specifically guide the enhancement process of degradation features. To achieve the comprehensive extraction of global and local features, this patent designs two formats of spatio-frequency feature interaction modules as the basic building blocks of the two subnets. In addition, to suppress the deviation of frequency components caused by the inherent bias of neural networks, an additional frequency domain loss function is introduced to constrain the model learning. The excellent performance of the proposed method has been fully verified on multiple underwater datasets and shows good practicality in practical applications.

[0009] The technical means adopted by the present invention are as follows:

[0010] An underwater image enhancement method based on spatio-frequency feature interaction, comprising the following steps:

[0011] Step 1: Acquisition and construction of a dataset; randomly select images from the publicly available underwater image dataset UIEB, and divide the images into a training set and a test set according to a certain ratio;

[0012] Step 2: Construction of the network model; the network model adopts a two-stage design, including: an amplitude subnet and a phase subnet; a reconstruction module is arranged between the amplitude subnet and the phase subnet; the reconstruction module is used to output the rough enhancement result of the network; the output of the reconstruction module is the input of the phase subnet;

[0013] Establish a degradation feature extraction module, and use the difference between the rough enhanced image and the original input underwater image as the input of the degradation feature extraction module to obtain the degradation features of the image.

[0014] Furthermore, the amplitude subnet adopts an encoder-decoder structure and is composed of 5 amplitude-feature interaction modules. The 5 amplitude-feature interaction modules are arranged in a "U" shape to extract features of different scales, and skip connections are established between the symmetric amplitude-feature interaction modules before and after to promote the propagation of feature information.

[0015] Furthermore, the amplitude-feature interaction module includes: a frequency domain branch that extracts global frequency features by means of Fourier transform, a spatial domain branch that extracts local spatial features, and a global-guided local structure that interactively processes global and local features;

[0016] The mathematical expression of the amplitude-feature interaction module is:

[0017] a f1 =F -1 (LR(PW(LR(PW(A(a i ))))),P(a i )) (1);

[0018]

[0019] where, a i represents the input feature of the module; A(·) and P(·) respectively represent the amplitude and phase component extraction operations, and F -1 represents the inverse Fourier transform; LR and Re respectively represent the LeakyReLU and ReLU activation functions; represents the element addition operation; PW represents the pointwise convolution; DW 3×3 represents the depth convolution with a convolution kernel size of 3×3;

[0020]

[0021] where, a f1 and a s1 respectively represent the global and local features extracted by the frequency domain branch and the spatial domain branch of the first half of the module;

[0022] For the global feature a f1Perform global average pooling Avg(·) and softmax function activation processing to obtain corresponding global guidance weights, and then multiply the obtained channel weights with the local feature a s1 Perform element-wise multiplication a m Denote the intermediate feature obtained via the global guidance local structure;

[0023] a f2 = F -1 (LR(PW(LE(PW(A(a m ))))), P(a m )) (4);

[0024]

[0025] a o = PW(concat(a f2 , a s2 )) (6);

[0026] Perform global and local feature extraction on the intermediate feature a m , and then concatenate the global feature a f2 extracted in the frequency domain with the local feature a s2 extracted in the spatial domain, and apply a pointwise convolution operation to the concatenated features to obtain the final output feature a o .

[0027] Furthermore, the phase subnet includes: 4 phase-feature interaction modules and 2 adaptive guidance modules.

[0028] Furthermore, the phase-feature interaction module introduces a multi-scale feature extraction structure in the spatial domain branch;

[0029] Introduce a multi-scale feature extraction structure in the spatial domain branch of the phase-feature interaction module. The mathematical expression of the phase-feature interaction module is:

[0030] p f1 = F -1 (LR(PW(LR(PW(P(p i ))))), A(p i )) (7);

[0031] p s1 = LR(concat(Re(PW(p i )), DW 3×3 (Re(DW 3×3 (p i ))),... Re(DW 3×3 (Re(DW3×3 (p i )))))) (8);

[0032]

[0033] p f2 = F -1 (LR(PW(LR(PW(P(p m ))))), A(p m )) (10);

[0034] p s2 = LR(concat(Re(PW(p m ))), DW 3×3 (Re(DW 3×3 (p m ))),... Re(DW 3×3 (Re(DW 3×3 (p m )))))) (11);

[0035] p o = PW(concat(p f2 , p s2 )) (12).

[0036] Furthermore, the reconstruction module extracts the amplitude component from the amplitude subnet output I a , extracts the phase component from the original underwater image I, and then reconstructs the coarse enhancement result J through inverse fast Fourier transform coarse ; the reconstruction process of the reconstruction module is as follows:

[0037] J coarse = F -1 (A(I a ), P(I)) (13).

[0038] Furthermore, the degradation feature extraction module includes: an encoder and a multi-layer perceptron; the degradation feature extraction module is used to extract the degradation feature vector D from the coarse enhancement image and the residual degradation map of the input image;

[0039] The encoder includes: 6 stacked convolutional modules and 1 two-dimensional adaptive average pooling layer; the multi-layer perceptron includes: 2 fully connected layers.

[0040] Furthermore, the adaptive guidance module adaptively guides the input features in two dimensions of channels and pixels by using the degradation feature vector D.

[0041] Furthermore, during the construction of the network model in step 2, the design of the loss function is also included. To better constrain the training process of the constructed model and effectively learn the mapping relationship between the degraded underwater images and the corresponding high-quality images, a linear combination of multiple loss functions is introduced to form the loss function for training the disclosed method; the multiple loss functions include: the traditional L1 loss, the multi-scale structural similarity loss, the perceptual loss, and the frequency domain loss; the loss function is defined as follows:

[0042]

[0043] Among them, α, β, λ, and ψ represent hyperparameters used to balance the contribution ratios of different losses.

[0044] The traditional L1 loss is used to measure the difference between the enhancement result and the reference image at the pixel level and is defined as:

[0045]

[0046] Among them, and J c (i) respectively represent the pixel intensities corresponding to the enhancement result and the reference image at the i-th pixel point in the c-th channel;

[0047] The multi-scale structural similarity loss evaluates the image quality by calculating the structural similarity of the image at different scales and is defined as:

[0048]

[0049] Among them, l(i) and cs(i) represent the measurements of brightness, contrast, and structural similarity between images within the window range centered on pixel i, and α and β m both represent default parameters;

[0050] The perceptual loss L perceptual evaluates the similarity degree of two images by comparing the Euclidean distance in the feature space of the pre-trained CNN network and is defined as:

[0051]

[0052] Among them, φ j represents the feature map corresponding to the j-th layer in the VGG backbone network, and C j , W j and H j respectively represent the number of channels, width, and height of the feature map;

[0053] The frequency domain loss L frequency includes: the amplitude loss L amp and the phase loss L pha;

[0054] L amp =‖A(J)-A(J coarse )‖ (18)

[0055] L pha =‖P(J)-P(J out )‖ (19)

[0056] Where J represents the reference image, and A(·) and P(·) are responsible for extracting the amplitude component and phase component of the corresponding image; the frequency-domain loss L frequency is expressed as:

[0057] L frequency =α f L amp +β f L pha (20)

[0058] Where α f and β f respectively represent the hyperparameter weights corresponding to the two-component losses; α f =0.2, β f =1.

[0059] Compared with the prior art, the present invention has the following advantages:

[0060] The method of the present invention can effectively utilize spatial-domain information and frequency-domain information to enhance degraded underwater images, alleviate the under-enhancement or over-enhancement phenomena that occur during the enhancement process, and generate images with distinct colors, clear textures, and high contrast. In addition, the method of the present invention also has high generalization ability and robustness, and can better adapt to diverse and complex underwater environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0062] Figure 1 is the underwater enhancement network structure diagram proposed by the present invention;

[0063] Figure 2 is the detailed structure of the amplitude-feature interaction module used in the underwater enhancement network proposed by the present invention;

[0064] Figure 3 is the detailed structure of the phase-feature interaction module used in the underwater enhancement network proposed by the present invention;

[0065] Figure 4 It is the detailed structure of the degradation feature extraction module used in the underwater enhancement network proposed by the present invention;

[0066] Figure 5 It is the detailed structure of the adaptive guidance module used in the underwater enhancement network proposed by the present invention;

[0067] Figure 6 It is the subjective evaluation of the method of the present invention and the comparative methods on the Test90 test set, where (a) represents the original degraded underwater image, (b) represents the result of the UDCP method, (c) represents the result of the RB method, (d) represents the result of the ULAP method, (e) represents the result of the HLRP method, (f) represents the result of the FUnIE method, (g) represents the result of the UWCNN method, (h) represents the result of the Dual-stream method, (i) represents the result of the UGAN method, (j) represents the result of the U-shape method, (k) represents the result of the present invention, and (l) represents the reference image.

[0068] Figure 7 It is the subjective evaluation of the method of the present invention and the comparative methods on the UFO-120 test set, where (a) is the original degraded underwater image, (b) represents the result of the UDCP method, (c) represents the result of the RB method, (d) represents the result of the ULAP method, (e) represents the result of the HLRP method, (f) represents the result of the FUnIE method, (g) represents the result of the UWCNN method, (h) represents the result of the Dual-stream method, (i) represents the result of the UGAN method, (j) represents the result of the U-shape method, (k) represents the result of the present invention, and (l) represents the reference image. Detailed implementation manners

[0069] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0070] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0071] As Figure 1 shown, the present invention provides an underwater image enhancement method based on spatio-frequency feature interaction, including the following steps:

[0072] Step 1: Acquisition and construction of the data set; randomly select images from the publicly available underwater image data set UIEB, and divide the images into a training set and a test set according to a certain ratio;

[0073] Step 2: Construction of the network model; the network model adopts a two-stage design, including: an amplitude subnet and a phase subnet; a reconstruction module is arranged between the amplitude subnet and the phase subnet; the reconstruction module is used to output the rough enhancement result of the network; the output of the reconstruction module is the input of the phase subnet; a degradation feature extraction module is established, and the difference between the roughly enhanced image and the original input underwater image is used as the input of the degradation feature extraction module to obtain the degradation features of the image. During the construction of the network model, the design of the loss function is also included. In order to better constrain the training process of the constructed model and effectively learn the mapping relationship between the degraded underwater image and the corresponding high-quality image, a linear combination of multiple loss functions is introduced to form the loss function for training the disclosed method; the multiple loss functions include: the traditional L1 loss, the multi-scale structural similarity loss, the perceptual loss, and the frequency domain loss; the loss function is defined as follows:

[0074]

[0075] where α, β, λ, and ψ represent hyperparameters used to balance the contribution ratios of different losses.

[0076] The traditional L1 loss is used to measure the difference between the enhancement result and the reference image at the pixel level, and is defined as:

[0077]

[0078] where and Jc (i) represents the pixel intensity corresponding to the enhanced result and the reference image at the i-th pixel in the c-th channel;

[0079] The multi-scale structural similarity loss evaluates the image quality by calculating the structural similarity of the image at different scales and is defined as:

[0080]

[0081] where l(i) and cs(i) represent the measurements of luminance, contrast, and structural similarity between images within a window centered on pixel i, and α and β m both represent default parameters;

[0082] The perceptual loss L perceptual evaluates the similarity degree of two images by comparing the Euclidean distance in the feature space of a pre-trained CNN network and is defined as:

[0083]

[0084] where φ j represents the feature map corresponding to the j-th layer in the VGG backbone network, and C j , W j and H j represent the number of channels, width, and height of the feature map respectively;

[0085] The frequency-domain loss L frequency includes: the amplitude loss L amp and the phase loss L pha ;

[0086] L amp =‖A(J)-A(J coarse )‖ (18);

[0087] L pha =‖P(J)-P(J out )‖ (19);

[0088] where J represents the reference image, and A(·) and P(·) are responsible for extracting the amplitude component and phase component of the corresponding image; the frequency-domain loss L frequency is expressed as:

[0089] L frequency =α f L amp +β f L pha (20);

[0090] where α f and β frespectively represent the hyperparameter weights corresponding to the two-component losses; α f = 0.2, β f = 1.

[0091] In the present application, as a preferred implementation, the amplitude subnet adopts an encoder-decoder structure and is composed of 5 amplitude-feature interaction modules. The 5 amplitude-feature interaction modules are arranged in a "U" shape to extract features of different scales, and skip connections are established between the amplitude-feature interaction modules that are symmetric before and after to promote the propagation of feature information.

[0092] Preferably, the amplitude-feature interaction module includes: a frequency domain branch that extracts global frequency features by means of Fourier transform, a spatial domain branch that extracts local spatial features, and a global-guided local structure that interactively processes global and local features;

[0093] The mathematical expression of the amplitude-feature interaction module is:

[0094] a f1 = F -1 (LR(PW(LR(PW(A(a i ))))), P(a i )) (1);

[0095]

[0096] where a i represents the input feature of the module; A(·) and P(·) respectively represent the amplitude and phase component extraction operations, and F -1 represents the inverse Fourier transform; LR and Re respectively represent the LeakyReLU and ReLU activation functions; represents the element-wise addition operation; PW represents the pointwise convolution; DW 3×3 represents the depth convolution with a convolution kernel size of 3×3d;

[0097]

[0098] where a f1 and a s1 respectively represent the global and local features extracted by the frequency domain branch and the spatial domain branch of the first half of the module;

[0099] The global feature a f1 is subjected to global average pooling Avg(·) and softmax function activation processing to obtain the corresponding global-guided weight, and then the obtained channel weight is multiplied element-wise with the local feature a s1 ; a m represents the intermediate feature obtained through the global-guided local structure;

[0100] a f2 = F -1 (LR(PW(LR(PW(A(a m ))))),P(a m )) (4);

[0101]

[0102] α o = PW(concat(a f2 ,a s2 )) (6);

[0103] For the intermediate feature a m , global and local feature extraction is performed, and then the global feature a f2 extracted in the frequency domain is concatenated with the local feature a s2 extracted in the spatial domain, and a pointwise convolution operation is applied to the concatenated features to obtain the final output feature a o of the amplitude - feature interaction module.

[0104] As a preferred embodiment, in this application, the phase subnet includes: 4 phase - feature interaction modules and 2 adaptive guidance modules. The phase - feature interaction module introduces a multi - scale feature extraction structure in the spatial domain branch;

[0105] A multi - scale feature extraction structure is introduced in the spatial domain branch of the phase - feature interaction module. The mathematical expression of the phase - feature interaction module is:

[0106] p f1 = F -1 (LR(PW(LR(PW(P(p i ))))),A(p i )) (7);

[0107] p s1 = LR(concat(Re(PW(p i )),DW 3×3 (Re(DW 3×3 (p i ))),...Re(DW 3×3 (Re(DW 3×3 (p i )))))) (8);

[0108]

[0109] p f2 = F -1 (LR(PW(LR(PW(P(p m))))), A(p m )) (10);

[0110] p s2 = LR(concat(Re(PW(p m ))), DW 3×3 (Re(DW 3×3 (p m ))),... Re(DW 3×3 (Re(DW 3×3 (p m )))))) (11);

[0111] p o = PW(concat(p f2 , p s2 )) (12).

[0112] Furthermore, the reconstruction module extracts the amplitude component from the output I of the amplitude subnet, extracts the phase component from the original underwater image I, and then reconstructs the coarsely enhanced result J through the inverse fast Fourier transform a ; the reconstruction process of the reconstruction module is as follows: coarse ; the reconstruction process of the reconstruction module is:

[0113] J coarse = F -1 (A(I a ), P(I)) (13).

[0114] In this application, the degradation feature extraction module includes: an encoder and a multi-layer perceptron; the degradation feature extraction module is used to extract the degradation feature vector D from the coarsely enhanced image and the residual degradation map of the input image;

[0115] The encoder includes: 6 stacked convolutional modules and 1 two-dimensional adaptive average pooling layer; the multi-layer perceptron includes: 2 fully connected layers.

[0116] As a preferred implementation manner, the adaptive guidance module adaptively guides the input features in two dimensions of channels and pixels by using the degradation feature vector D.

[0117] Example:

[0118] An underwater image enhancement method based on spatio-temporal feature interaction includes the following steps:

[0119] Step 1: As described in Step 1 of the Invention Content section, construct a dataset and divide it into a training set and a test set. The training set includes: 800 pairs of images randomly selected from the UIEB-1 subset of the UIEB dataset, and 1500 pairs of synthetic underwater images from the UFO-120 dataset. The test set includes: Test90 (the remaining 90 pairs of images from the UIEB-1 subset), and 120 pairs of synthetic images from the UFO-120 dataset.

[0120] Step 2: Parameter setting. The specific training parameters are set as follows: Use the PyTorch deep learning framework to build the proposed network model; Train on a computer equipped with an Intel Core i5-10400F CPU @ 2.90GHz processor and an Nvidia GTX 2060 graphics card (14GB video memory); The optimizer uses the Adam algorithm, and the initial learning rate is set to 0.0001; The batch size is 4; The total number of training epochs is 100.

[0121] Step 2-1: Build an underwater image enhancement network based on spatio-frequency feature interaction as described in Step 2 of the Invention Content section, including an amplitude subnet, a phase subnet, a reconstruction module, a degradation feature extraction module, and an adaptive guidance module, as shown in the appendix Figure 1 as follows.

[0122] Step 2-2: Use the training set in Step 1 to train the constructed network model. During the training process, the batch size is 4, the total number of iterations is 100, the optimizer uses Adam, and the initial learning rate is set to 0.0001.

[0123] Step 3: Conduct subjective and objective evaluations on the test set. The present invention selects 9 representative underwater image enhancement methods for comparative experiments to prove the effectiveness of the present invention. The selected comparison methods are: UDCP, ULAP, RB, HLRP, FUnIE, Dual-stream, UGAN, U-Shape, and UWCNN.

[0124] Step 3-1: Subjective evaluation. The subjective evaluations of the method of the present invention and other comparison methods on the Test90 and UFO-120 datasets are respectively as shown in Figure 6 and Figure 7 as follows. According to Figure 6 and Figure 7 it can be seen that traditional underwater image enhancement methods all have problems of insufficient processing effects to varying degrees. Deep learning-based methods have achieved good effects in color restoration, contrast enhancement, and detail enhancement, but there are still deficiencies such as color distortion and detail loss. In contrast, the method proposed in this paper has achieved excellent effects in color restoration, contrast enhancement, detail enhancement, etc. The generated enhanced images are closer to the real scene, showing better visual quality and generalization ability.

[0125] Step 3-2: Objective evaluation. In addition to subjective evaluation, the underwater images enhanced by the method in this paper are also used for quantitative analysis. Four objective evaluation indexes are selected in this paper to quantitatively analyze the results of the method in this paper and all comparative methods, namely: mean square error (MSE), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and patch-based contrast quality index (PCQI). These four full-reference evaluation indexes are used to evaluate the results of each method. Among them, a lower MSE value and higher PSNR, SSIM, and PCQI values indicate that the enhanced image is closer to the reference image in terms of content, texture, and contrast, reflecting better restoration performance. The quantitative analysis results of the method in this paper and each method on different datasets are shown in Table 1 and Table 2. According to Table 1 and Table 2, the method in this paper has obtained the optimal evaluation index values on the reference datasets, indicating that the method of the present invention has significant advantages in aspects such as noise suppression, brightness and contrast enhancement, structure strengthening, and color correction, which is consistent with the conclusion of subjective evaluation.

[0126] Table 1 Objective evaluation of the method of the present invention and comparative methods on the Test90 test set

[0127]

[0128] Table 2 Objective evaluation of the method of the present invention and comparative methods on the UFO-120 test set

[0129]

[0130]

[0131] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0132] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0133] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0134] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0135] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0136] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of various embodiments of the present invention.

Claims

1. An underwater image enhancement method based on spatio-frequency feature interaction, characterized in that It includes the following steps: Step 1: Acquisition and construction of the dataset; randomly select images from the publicly available underwater image dataset UIEB, and divide the images into a training set and a test set according to a certain ratio; Step 2: Construction of the network model; The network model adopts a two-stage design, including: an amplitude subnet and a phase subnet; a reconstruction module is arranged between the amplitude subnet and the phase subnet; the reconstruction module is used to output the rough enhancement result of the network; The output of the reconstruction module is the input of the phase subnet; A degradation feature extraction module is established, and the difference between the roughly enhanced image and the original input underwater image is used as the input of the degradation feature extraction module to obtain the degradation features of the image, which are used to specifically guide the enhancement process of the degradation features.

2. The underwater image enhancement method based on spatio-frequency feature interaction according to claim 1, wherein The amplitude subnet adopts an encoder-decoder structure and is composed of 5 amplitude-feature interaction modules. The 5 amplitude-feature interaction modules are arranged in a "U" shape to extract features of different scales, and skip connections are established between the symmetric amplitude-feature interaction modules before and after to promote the propagation of feature information.

3. The underwater image enhancement method based on spatio-frequency feature interaction according to claim 2, wherein, The amplitude-feature interaction module includes: a frequency domain branch that extracts global frequency features by means of Fourier transform, a spatial domain branch that extracts local spatial features, and a global-guided local structure that interactively processes global and local features; The mathematical expression of the amplitude-feature interaction module is: a f1 = F -1 (LR(PW(LR(PW(A(a i ))))), P(a i )) (1); Among them, a i represents the input feature of the module; A(·) and P(·) respectively represent the amplitude and phase component extraction operations, and F -1 represents the inverse Fourier transform; LR and Re respectively represent the LeakyReLU and ReLU activation functions; represents the element addition operation; PW represents the pointwise convolution; DW 3×3 represents the depth convolution with a convolution kernel size of 3×3; Among them, a f1 and a s1 respectively represent the global and local features extracted from the frequency-domain branch and the spatial-domain branch of the first half of the module; For the global feature a f1 Perform global average pooling Avg(·) and softmax function activation processing to obtain the corresponding global guidance weights, and then multiply the obtained channel weights with the local feature a s1 Element-wise multiplication a m a represents the intermediate feature obtained via the global guidance local structure; a f2 = a -1 (LR(PW(LR(PW(A(a m ))))), P(a m )) (4); a o = PW(concat(a f2 , a s2 )) (6); Extract global and local features for the intermediate feature a m Then, concatenate the global feature a extracted in the frequency domain f2 with the local feature a extracted in the spatial domain s2 Perform concat feature concatenation and apply a pointwise convolution operation to the concatenated features to obtain the final output feature a of the amplitude-feature interaction module o .

4. An underwater image enhancement method based on spatio-frequency feature interaction according to claim 2, characterized in that The phase subnet includes: 4 phase-feature interaction modules and 2 adaptive guidance modules.

5. A method for underwater image enhancement based on spatio-frequency feature interaction according to claim 4, characterized in that, The phase-feature interaction module introduces a multi-scale feature extraction structure in the spatial domain branch; A multi-scale feature extraction structure is introduced in the spatial domain branch of the phase-feature interaction module, and the mathematical expression of the phase-feature interaction module is: p f1 = F -1 (LR(PW(LR(PW(P(p i ))))), A(p i )) (7); p s1 = LR(concat(Re(PW(p i )),DW 3×3 (Re(DW 3×3 (p i ))),...Re(DW 3×3 (Re(DW 3×3 (p i ))))))(8); p f2 = F -1 (LR(PW(LR(PW(P(p m ))))), A(p m )) (10); p s2 = LR(concat(Re(PW(p m )),DW 3×3 (Re(DW 3×3 (p m ))),...Re(DW 3×3 (Re(DW 3×3 (p m ))))))(11); p o = PW(concat(p f2 , p s2 )) (12).

6. The underwater image enhancement method based on spatio-frequency feature interaction according to claim 1, wherein The reconstruction module extracts the amplitude component from the output I of the amplitude subnet, extracts the phase component from the original underwater image I, and then reconstructs the coarsely enhanced result J through inverse fast Fourier transform. a ; The reconstruction process of the reconstruction module is as follows: coarse ; J coarse = F -1 (A(I a ), P(I)) (13).

7. An underwater image enhancement method based on spatio-frequency feature interaction according to claim 1, characterized in that The degradation feature extraction module includes: an encoder and a multi-layer perceptron; the degradation feature extraction module is used to extract the degradation feature vector D from the residual degradation map of the roughly enhanced image and the input image; The encoder includes: 6 stacked convolutional modules and 1 two-dimensional adaptive average pooling layer; the multi-layer perceptron includes: 2 fully connected layers.

8. A method for underwater image enhancement based on spatio-frequency feature interaction according to claim 4, characterized in that, The adaptive guidance module uses the degradation feature vector D to adaptively guide the input features in two dimensions: channel and pixel.

9. A method for underwater image enhancement based on spatio - frequency feature interaction according to claim 1, wherein During the construction process of the network model in Step 2, the design of the loss function is also included. In order to better constrain the training process of the constructed model and learn the mapping relationship between the degraded underwater images and the corresponding high-quality images, a linear combination of multiple loss functions is introduced to form the loss function for training the disclosed method; the multiple loss functions include: the traditional L1 loss, the multi-scale structural similarity loss, the perceptual loss, and the frequency domain loss; the loss function is defined as follows: Among them, α, β, λ, and ψ represent hyperparameters used to balance the contribution ratios of different losses. The traditional L1 loss is used to measure the difference between the enhancement result and the reference image at the pixel level, and is defined as: Among them, and J C (i) respectively represent the pixel intensities corresponding to the enhanced result and the reference image at the i-th pixel point in the c-th channel. The multi-scale structural similarity loss evaluates the image quality by calculating the structural similarity of the image at different scales and is defined as: where l(i) and cs(i) represent the measurements of between-image luminance, contrast, and structural similarity within the window centered at pixel i, and α and β m both represent default parameters; The perceptual loss L perceptual is used to evaluate the similarity degree of images by comparing the Euclidean distance of two images in the feature space of a pre-trained CNN network, and is defined as: Among them, φ j represents the feature map corresponding to the j-th layer in the VGG backbone network, and C j , W j and H j respectively represent the number of channels, width, and height of the feature map; The frequency-domain loss L frequency includes: the amplitude loss L amp and the phase loss L pha ; L amp = ‖A(J) - A(J coarse )‖ (18) L pha = ‖P(J) - P(J oit )‖ (19) Among them, J represents the reference image, and A(·) and P(·) are responsible for extracting the amplitude component and the phase component of the corresponding image; the frequency domain loss L frequency is expressed as: L frequency = α f L amp + β f L pha (20) Among them, α f and β f respectively represent the hyperparameter weights corresponding to the two-component losses; α f = 0.2, β f = 1.

Citation Information

Cited By

  • Underwater image enhancement method and system based on double-domain collaboration

    CN121053048A