An underwater image enhancement method, device, storage medium and electronic equipment

CN118982471BActive Publication Date: 2026-09-29HANGZHOU ZHUOXI INST OF BRAIN & INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410987837.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-09-29
Estimated Expiration
2044-07-23

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的在于提供一种水下图像增强方法、装置、存储介质及电子设备,用以解决现有技术中优化方案无法有效平衡增强性能与计算效率的问题

Benefits of technology

[0009]本申请实施例的有益效果在于:通过结合图像自适应的三维查找表和轻量级卷积神经网络,构建全局增强分支网络和局部增强分支网络,在保证性能的同时降低模型参数量;同时在不同分支内都引入提示学习策略,强化了模型对不同退化类型图像的适应性,使水下图像增强效果得到提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118982471B_ABST
    Figure CN118982471B_ABST
Patent Text Reader

Abstract

The application provides an underwater image enhancement method and device, a storage medium and an electronic device. The method comprises the following steps: establishing a training set; constructing a global enhancement branch network, outputting a global enhancement branch result of a degraded underwater image through the global enhancement branch network; constructing a local enhancement branch network, outputting a local enhancement branch result of the degraded underwater image through the local enhancement branch network; superimposing the global enhancement branch result and the local enhancement branch result to obtain an enhanced image; training the global enhancement branch network and the local enhancement branch network according to the enhanced image and a reference image until convergence; and inputting an underwater image to be enhanced into the trained global enhancement branch network and the local enhancement branch network respectively to obtain an enhanced image of the underwater image to be enhanced. The application reduces the model parameter quantity while ensuring the performance, strengthens the adaptability of the model to images of different degradation types, and improves the underwater image enhancement effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an underwater image enhancement method, apparatus, storage medium, and electronic device. Background Technology

[0002] Underwater optical imaging plays a crucial role in numerous scientific fields, including marine ecology, marine geology, and underwater resource discovery. However, due to light absorption and scattering effects in the underwater environment, acquired images often suffer from various quality degradation problems, such as color shift, reduced contrast, blurred details, and limited visibility. These issues not only affect human visual perception but also significantly hinder the performance of machine learning-based algorithms in scene understanding and analysis, especially in resource-constrained underwater mobile terminal applications. Therefore, designing lightweight and efficient underwater optical image enhancement methods to address these challenges is extremely urgent. In recent years, deep learning-based underwater image enhancement methods have received widespread attention, but these methods present a significant contradiction between enhancement performance and computational cost. For example, high-performance methods, such as Ucolor, employ deep and complex network architectures to achieve high-quality image enhancement. However, these methods typically have a high number of model parameters and high floating-point computation requirements, leading to low computational efficiency and making them difficult to integrate into underwater mobile platforms that require real-time feedback. Conversely, lightweight methods like Shallow-UWnet, while offering some control over model parameter count and computational cost, often perform poorly when dealing with complex and diverse underwater degradation types, failing to provide satisfactory enhancements. These lightweight methods, due to their limited adaptability to different degradation scenarios, struggle to match the performance of deep and complex network models.

[0003] Therefore, existing optimization schemes cannot effectively balance performance enhancement and computational efficiency. There is an urgent need for an enhancement method that can address diverse underwater degradation problems and operate efficiently under resource-constrained conditions to fill this gap in the existing technology. Summary of the Invention

[0004] The purpose of this application is to provide an underwater image enhancement method, apparatus, storage medium, and electronic device to solve the problem that existing optimization schemes cannot effectively balance enhancement performance and computational efficiency.

[0005] The embodiments of this application adopt the following technical solution: an underwater image enhancement method, comprising: establishing a training set composed of a degraded underwater image and a reference image; constructing a global enhancement branch network and outputting the global enhancement branch result of the degraded underwater image through the global enhancement branch network; constructing a local enhancement branch network and outputting the local enhancement branch result of the degraded underwater image through the local enhancement branch network; superimposing the global enhancement branch result and the local enhancement branch result to obtain an enhanced image of the degraded underwater image; training the global enhancement branch network and the local enhancement branch network according to the enhanced image and the reference image until convergence; inputting the underwater image to be enhanced into the trained global enhancement branch network and the local enhancement branch network respectively to obtain the enhanced image of the underwater image to be enhanced.

[0006] This application also provides an underwater image enhancement device, comprising: a training set establishment module for establishing a training set composed of a degraded underwater image and a reference image; a global network construction module for constructing a global enhancement branch network and outputting the global enhancement branch result of the degraded underwater image through the global enhancement branch network; a local network construction module for constructing a local enhancement branch network and outputting the local enhancement branch result of the degraded underwater image through the local enhancement branch network; an overlay processing module for overlaying the global enhancement branch result and the local enhancement branch result to obtain an enhanced image of the degraded underwater image; a training module for training the global enhancement branch network and the local enhancement branch network based on the enhanced image and the reference image until convergence; and an enhancement module for inputting the underwater image to be enhanced into the trained global enhancement branch network and the local enhancement branch network respectively to obtain the enhanced image of the underwater image to be enhanced.

[0007] This application embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the underwater image enhancement method described above.

[0008] This application also provides an electronic device, including at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the underwater image enhancement method described above when executing the computer program in the memory.

[0009] The beneficial effects of this application's embodiments are as follows: by combining an image-adaptive 3D lookup table and a lightweight convolutional neural network, a global enhancement branch network and a local enhancement branch network are constructed, which reduces the number of model parameters while ensuring performance; at the same time, a cue learning strategy is introduced in different branches, which enhances the model's adaptability to images of different degradation types and improves the underwater image enhancement effect. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart of the underwater image enhancement method in the first embodiment of this application;

[0012] Figure 2 This is a schematic diagram of the structure of the global enhanced branch network in the first embodiment of this application;

[0013] Figure 3 This is a diagram illustrating the specific network architecture of the prompt generation module in the first embodiment of this application;

[0014] Figure 4 This is a network architecture diagram of the prompting and interaction module in the first embodiment of this application;

[0015] Figure 5 This is a schematic diagram of the structure of the locally enhanced branch network in the first embodiment of this application;

[0016] Figure 6 This is a schematic diagram of the underwater image enhancement device in the second embodiment of this application. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this document.

[0018] Underwater optical imaging plays a crucial role in numerous scientific fields, including marine ecology, marine geology, and underwater resource discovery. However, due to light absorption and scattering effects in the underwater environment, acquired images often suffer from various quality degradation problems, such as color shift, reduced contrast, blurred details, and limited visibility. These issues not only affect human visual perception but also significantly hinder the performance of machine learning-based algorithms in scene understanding and analysis, especially in resource-constrained underwater mobile terminal applications. Therefore, designing lightweight and efficient underwater optical image enhancement methods to address these challenges is extremely urgent. In recent years, deep learning-based underwater image enhancement methods have received widespread attention, but these methods present a significant contradiction between enhancement performance and computational cost. For example, high-performance methods, such as Ucolor, employ deep and complex network architectures to achieve high-quality image enhancement. However, these methods typically have a high number of model parameters and high floating-point computation requirements, leading to low computational efficiency and making them difficult to integrate into underwater mobile platforms that require real-time feedback. Conversely, lightweight methods like Shallow-UWnet, while offering some control over model parameter count and computational cost, often perform poorly when dealing with complex and diverse underwater degradation types, failing to provide satisfactory enhancements. These lightweight methods, due to their limited adaptability to different degradation scenarios, struggle to match the performance of deep and complex network models.

[0019] To address the problem that existing technologies cannot effectively balance performance enhancement and computational efficiency, the first embodiment of this application provides an underwater image enhancement method, the flowchart of which is shown below. Figure 1 As shown, it mainly includes steps S10 to S60:

[0020] S10, establish a training set consisting of degraded underwater images and reference images.

[0021] In this embodiment, the degraded underwater image X∈R H×W×3 The underwater optical image directly acquired by the underwater acquisition equipment, with reference image Z∈R H×W×3 To obtain high-quality underwater images after enhancement of degraded underwater images, a training set {X,Z} is formed by selecting various types of degraded underwater images X and reference images Z. When selecting degraded underwater images and reference images, multiple types and sizes of images should be selected. However, before integrating the training set, all images of different sizes should be normalized.

[0022] S20, construct a global enhancement branch network, and output the global enhancement branch results of the degraded underwater image through the global enhancement branch network.

[0023] Based on the requirements for efficiency and effectiveness in underwater image enhancement, this embodiment uses a global enhancement branch to represent the global context of degraded underwater images. Subsequently, an adaptive three-dimensional lookup table is generated to form a lightweight global enhancement branch network, which is used to output the global enhancement branch results of degraded underwater images. Figure 2 The diagram shows the structure of the global enhancement branch network. The following will combine... Figure 2 The content shown describes the process of constructing a global augmented branch network and outputting the results of the global augmented branch network.

[0024] First, bilinear interpolation is performed on the degraded underwater image to obtain a low-resolution degraded underwater image. To improve the network's processing of highly degraded low-resolution images, this embodiment performs bilinear interpolation on the degraded underwater image X in the training set to obtain its corresponding low-resolution version X′∈R. 256×256×3 .

[0025] The low-resolution degraded underwater image X′ is then input into a lookup table backbone network for feature extraction, yielding the global contextual features E∈r of the degraded underwater image. h×w×M The lookup table backbone network consists of multiple convolutional layers, LeakyReLU activation functions, and instance normalization layers. Figure 2 The lookup table backbone network shown includes four feature extraction layers: a 3×3 convolutional layer, a LeakyReLU activation function, and an instance normalization layer; and a single feature extraction layer: a 1×1 convolutional layer, a LeakyReLU activation function, and an instance normalization layer. After processing by the above feature extraction layers, the global context features E of the degraded underwater image are obtained.

[0026] Simultaneously, the degraded underwater image X is input into the shallow feature extraction network of the global enhancement branch to obtain the shallow features H∈R of the degraded underwater image. H×W×M Among them, combining Figure 2 The shallow feature extraction network shown consists of three 1×1 convolutional layers, which extract the corresponding shallow features H from the degraded underwater image X through three 1×1 convolutional layers.

[0027] Then, the shallow feature H and the randomly initialized cue component P0∈R are... S×D The input is fed into the prompt generation module to obtain the prompt feature P carrying degradation information. g ∈R S×D Specifically, the network architecture of the prompt generation module is as follows: Figure 3 As shown, the main processing methods include global average pooling, 1×1 convolution and Softmax function processing, dot product processing, etc.; in the processing of cue features P carrying degradation information... gDuring generation, global average pooling is first performed on the spatial dimension of the shallow features H. Then, the shallow features H after global average pooling are processed through a 1×1 convolutional layer and a softmax function to generate weights w∈R based on the input. D Finally, the randomly initialized cue component P0 is multiplied by the weight w to obtain the cue feature P carrying degradation information. g .

[0028] In the prompt feature P carrying degradation information g After obtaining this, it is input into the prompting interaction module along with the global contextual features E of the degraded underwater image for dynamic feature calibration, resulting in the global contextual features E′ for degradation perception. Specifically, Figure 4 The diagram shows the specific network architecture of the prompting interaction module in this embodiment. It mainly includes a cross-attention module and a multi-branch gating mechanism. When calculating and outputting the global context feature E′ for degradation perception, the global context feature E of the degraded underwater image is first processed by 1×1 convolution and 3×3 depthwise separable convolution to obtain compressed features. Then compress feature E c and the cue feature p carrying degradation information g Input is fed into the cross-attention module to generate cue-aware features.

[0029] In some embodiments, the specific implementation process of the cross-attention module is as follows:

[0030] First, two layers of 1×1 convolutional encoding are used to encode the requested component p:

[0031] P = W2(W1(E) c (1)

[0032] Here, W1(·) and W2(·) represent 1×1 convolutions. Simultaneously, the mapping between key K and value V is learned using 1×1 convolutions:

[0033] K = W3(P g (2)

[0034] V = W4(P) g (3)

[0035] Here, W3(·) and W4(·) both represent 1×1 convolutions. Next, matrix multiplication is performed between P and the transpose K to obtain the attention matrix A:

[0036]

[0037] Here, α represents a learnable parameter to control the distribution of the attention matrix. Finally, cue-aware features E′ are generated based on A, V, and the residual policy. c:

[0038] E′ c =E c +W5(A·V); (5)

[0039] Where W5(·) represents a 1×1 convolution.

[0040] Generate cue-aware features E′ c Next, a multi-branch gating mechanism is used for feature interaction. Specifically, the global context feature E is first uniformly divided into N segmentation features along the channel dimension, where N is the total number of branches in the gating mechanism. Figure 4 The value of N is 4, meaning that the global context feature E is segmented to obtain four segmentation features. Next, the segmentation feature E is divided in each branch. i With cue-perceived features E′ c Adaptive interaction is performed to obtain the channel-dimensional interaction features E′ of each branch. i Among them, combining Figure 4 The architecture of the gating mechanism shown has the following channel-dimensional interaction feature E′. i The calculation formula is:

[0041]

[0042] Here, ⊙ represents element-wise multiplication. W represents the GELU non-linear activation function, and W6(·) and W7(·) represent 1×1 convolutions. Finally, the interaction features E′ of the trace dimension of each branch are... i The features are concatenated to obtain the global context features E′ of degradation perception, i.e., E′=[E′1;E′2;E′3;E′4], where [;] indicates the feature concatenation operation along the channel dimension.

[0043] After obtaining the degradation-aware global context feature E′, the degradation-aware global context feature E′ is average pooled and flattened to obtain... Then The input is fed into two fully connected layers to construct a three-dimensional lookup table T∈[0,1]. 3×t×t×t As a global enhancement branch network, t represents the number of elements in each color dimension of the 3D lookup table.

[0044] After the three-dimensional lookup table T is constructed, the global enhancement branch result of the degraded underwater image based on the three-dimensional lookup table can be output. In this embodiment, the global enhancement branch result Y1 is obtained by adaptive enhancement of the degraded underwater image in two steps. Specifically, for the RGB value of any pixel in the degraded underwater image X... First, determine the corresponding position (x, y, z) and index (i, j, k) in the three-dimensional lookup table. The calculation formula is as follows:

[0045]

[0046] Where d = λ max / t, λ max The maximum color value, This indicates a floor function. Then, based on the 8 elements located by the index (i,j,k) in the 3D lookup table, the color value corresponding to the output pixel is calculated using trilinear interpolation.

[0047]

[0048] Q x =xi,Q y =yj,Q z =zk; (10)

[0049] Among them, O c Let the values ​​stored in the three-dimensional lookup table T be c∈{r,g,b}, where r,g,b represent the red, green, and blue channels, respectively, and Q represent the deviation between the precise position (x,y,z) and index (i,j,k) of the input color value in the lookup table. After calculating the color value corresponding to each pixel in the degraded underwater image X, the global enhancement branch result Y1 can be formed based on the color value corresponding to each pixel in the degraded underwater image.

[0050] S30, construct a local enhancement branch network, and output the local enhancement branch results of the degraded underwater image through the local enhancement branch network.

[0051] The local enhancement branch network in this embodiment is used to capture details of the input image. At the same time, the local enhancement branch also introduces a cue learning strategy to improve adaptability to different degradation types. Figure 5 This diagram illustrates the structure of the local enhancement branch network in this embodiment. The degraded underwater image X is first input into the local feature extraction network to extract local features F∈R. H×W×N Extraction, where the structure of the local feature extraction network is as follows: Figure 5 As shown in the dashed box, it mainly includes three 3×3 convolutional layers, global average pooling, 1×1 convolutional layers and ReLU activation function, 1×1 convolutional layers and Sigmoid activation function, and dot product operation.

[0052] The local feature F and the randomly initialized cue component P0 are then input into the cue generation module to obtain the cue feature P carrying local information. l ∈R S×DThe structure of the generated module is also as shown here. Figure 3 As shown, the process by which the prompt generation module processes the local feature F and the randomly initialized prompt component P0 is the same as the process in step S20 for processing the global context feature E and the randomly initialized prompt component P0, and will not be repeated here. Subsequently, the prompt feature P carrying local information... l The local feature F is input into the prompting interaction module to obtain the local feature F′ of degradation perception. The structure of the prompting interaction module here is also as follows... Figure 4 As shown, the prompting interaction module responds to the prompt feature P carrying local information. l The processing of local features F, and the processing of global context features E and degenerate information-carrying feature P by the prompting interaction module in step S20. g The steps are the same and will not be repeated here. After obtaining the local features F′ of degradation perception, a 1×1 convolutional layer is applied to them to obtain the local enhancement branch result Y2 of the degraded underwater image.

[0053] S40, the global enhancement branch results and the local enhancement branch results are superimposed to obtain the enhanced image of the degraded underwater image.

[0054] For any degraded underwater image, after enhancement processing by the global enhancement branch network and the local enhancement branch network respectively, we can obtain the global enhancement branch result Y1 and the local enhancement branch result Y2. By superimposing the two, we can obtain the finely enhanced predicted image Y corresponding to the degraded underwater image.

[0055] S50, train the global enhancement branch network and the local enhancement branch network based on the enhanced image and the reference image until convergence.

[0056] During training, this embodiment can construct a loss function as the basis for training and convergence of the two branch networks. For example, the loss function can be constructed based on the predicted image Y formed by the global enhancement branch network and the local enhancement branch network and the reference image Z corresponding to the degraded underwater image. as follows:

[0057]

[0058] Based on the loss function between the predicted image and the reference image for each degraded underwater image, the parameters in the two branch networks are continuously optimized and trained until the loss function is minimized, at which point convergence is achieved, and the training of the global enhancement branch network and the local enhancement branch network is considered complete. It should be noted that the loss function formula shown in the above formula (11) is only one of the loss function formulas that can be used in image enhancement training. In actual processing, other loss function formulas can also be selected. This embodiment does not impose specific restrictions. In some embodiments, the number of iterations can also be set as a convergence condition. For example, after the network has undergone 500 iterations, the model is considered to have converged and the training is complete.

[0059] S60, the underwater image to be enhanced is input into the trained global enhancement branch network and local enhancement branch network respectively to obtain the enhanced image of the underwater image to be enhanced.

[0060] For the unenhanced underwater image to be processed, it is input into the trained global enhancement branch network and local enhancement branch network respectively to obtain the global enhancement branch result and the local enhancement branch result. After integration, a high-resolution enhanced image can be obtained.

[0061] In practical implementation, the trained network can be validated by establishing a test set, which includes degraded underwater images different from those in the training set, along with their reference images. During the testing phase, the test images are input into the trained network to generate enhanced images.

[0062] To verify the performance of this embodiment, the inventors conducted performance tests on two publicly available datasets. The evaluation results for the UIEBD test set are shown in Table 1, where "-" indicates exceeding the maximum memory limit. Table 1 shows that the proposed method outperforms existing methods on mainstream image quality evaluation metrics. On an RTX 2080Ti GPU, when enhancing a 1920×1080 resolution underwater image, the proposed method achieved a runtime of 30 frames per second, indicating that it better balances enhancement effect and computational complexity. Table 2 shows the performance evaluation on the LSUI test set. Table 2 shows that this method also exhibits good generalization performance.

[0063] Table 1

[0064]

[0065]

[0066] Table 2

[0067]

[0068] This embodiment proposes a lightweight two-branch underwater optical image enhancement method based on cue learning and lookup tables, aiming to efficiently and robustly improve the quality of underwater images. To adapt to various underwater degradation types, cue learning strategies are introduced in both branches. First, for global enhancement, a lightweight lookup table backbone network is used to learn the global context representation, thereby guiding the generation of the color transformation function. Simultaneously, a shallow convolutional neural network is used to optimize randomly initialized cue features. These generated cue components carry specific degradation context information. The global context representation is recalibrated through dynamic features to generate adaptive 3D lookup tables (3D LUTs) for the input image. The global enhancement result is obtained by performing lookup and interpolation operations between the degraded image and the generated 3D LUTs. Second, for local enhancement, multiple convolutional blocks are stacked to capture the rich texture and details of the input image. Similar to the global enhancement branch, the local enhancement branch utilizes generated input conditional cues and feature interaction mechanisms to improve adaptability to different degradation types. Finally, the enhancement results of the global and local branches are fused to obtain a finely enhanced image.

[0069] The lightweight design of this embodiment allows for real-time image enhancement at 1920×1080 resolution on a single RTX 2080Ti GPU, while maintaining enhanced performance, keeping the model parameter size to approximately 0.03M. This significantly reduces computational resource consumption and makes it suitable for various resource-constrained underwater applications. The cue learning strategy enhances the model's adaptability to different types of image degradation. Underwater scenes exhibit complex and diverse degradation types, making it crucial to improve the model's generalization performance. This embodiment introduces a cue learning strategy, using cue words to encode degradation information in the input image and dynamically calibrating features through a multi-branch cue interaction module, thereby enhancing the model's applicability to different types of image degradation.

[0070] Based on the same inventive concept, a second embodiment of this application provides an underwater image enhancement device, the structural schematic of which is shown below. Figure 6As shown, the system mainly includes the following sequentially coupled modules: a training set establishment module 10 for establishing a training set consisting of a degraded underwater image and a reference image; a global network construction module 20 for constructing a global enhancement branch network and outputting the global enhancement branch result of the degraded underwater image through the global enhancement branch network; a local network construction module 30 for constructing a local enhancement branch network and outputting the local enhancement branch result of the degraded underwater image through the local enhancement branch network; an overlay processing module 40 for overlaying the global enhancement branch result and the local enhancement branch result to obtain an enhanced image of the degraded underwater image; a training module 50 for training the global enhancement branch network and the local enhancement branch network based on the enhanced image and the reference image until convergence; and an enhancement module 60 for inputting the underwater image to be enhanced into the trained global enhancement branch network and the local enhancement branch network respectively to obtain the enhanced image of the underwater image to be enhanced.

[0071] In some embodiments, the global network construction module 20 is specifically used to perform bilinear interpolation on the degraded underwater image to obtain a low-resolution degraded underwater image; input the low-resolution degraded underwater image into a lookup table backbone network to obtain global context features of the degraded underwater image; input the degraded underwater image into a shallow feature extraction network to obtain shallow features of the degraded underwater image; input the shallow features and randomly initialized prompt components into a prompt generation module to obtain prompt features carrying degradation information; input the global context features and the prompt features into a prompt interaction module for dynamic feature calibration to obtain degradation-aware global context features; and perform average pooling, flattening, and two fully connected layers on the degradation-aware global context features to construct a three-dimensional lookup table as the global enhancement branch network.

[0072] In some embodiments, the global network construction module 20 is further configured to determine the position and index corresponding to the RGB value of any pixel in the degraded underwater image in the three-dimensional lookup table; calculate the color value corresponding to the pixel by a trilinear interdigitation based on the eight elements located by the index in the three-dimensional lookup table; and form the global enhancement branch result according to the color value corresponding to each pixel in the degraded underwater image.

[0073] In some embodiments, the global network construction module 20 is further specifically used to perform global average pooling on the spatial dimension of the shallow features; process the shallow features after global average pooling through a 1×1 convolutional layer and a Softmax function to generate input-based weights; and multiply the randomly initialized cue component with the weights to obtain the cue feature carrying degradation information.

[0074] In some embodiments, the prompting interaction module includes at least a cross-attention module and a multi-branch gating mechanism; the global network construction module 20 is further specifically used to perform 1×1 convolution and 3×3 depthwise separable convolution on the global context features to obtain compressed features; input the compressed features and the prompting features carrying degradation information into the cross-attention module to obtain prompting perception features; uniformly divide the global context features into N segmentation features along the channel dimension, where N is the total number of branches of the gating mechanism; adaptively interact the segmentation features with the prompting perception features in each branch to obtain the channel dimension interaction features of each branch; and concatenate the channel dimension interaction features of each branch to obtain the prompting features carrying degradation information.

[0075] In some embodiments, the local network construction module 30 is specifically used to input the degraded underwater image into a local feature extraction network to obtain local features; input the local features and a randomly initialized cue component into a cue generation module to obtain cue features carrying local information; input the cue features carrying local information and the local features into a cue interaction module to obtain degraded perception local features; and input the degraded perception local features into a 1×1 convolutional layer for processing to obtain the local enhancement branch result of the degraded underwater image.

[0076] In some embodiments, the training module 50 is specifically used to establish a loss function based on the enhanced image and the reference image; train the global enhancement branch network and the local enhancement branch network until the function value of the loss function is minimized, thus completing convergence.

[0077] This embodiment combines an image-adaptive 3D lookup table with a lightweight convolutional neural network to construct a global enhancement branch network and a local enhancement branch network, reducing the number of model parameters while ensuring performance. At the same time, a cue learning strategy is introduced in different branches to enhance the model's adaptability to images with different degradation types, thereby improving the underwater image enhancement effect.

[0078] Based on the same inventive concept, the third embodiment of this application provides a storage medium storing a computer program, which, when executed by a processor, implements the underwater image enhancement method described in the first embodiment of this application.

[0079] Based on the same inventive concept, the fourth embodiment of this application provides an electronic device, which includes at least a memory and a processor. The memory stores a computer program, and the processor implements the underwater image enhancement method described in the first embodiment of this application when executing the computer program in the memory.

[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An underwater image enhancement method, characterized in that, include: Establish a training set consisting of degraded underwater images and reference images; A global enhancement branch network is constructed, and the global enhancement branch results of the degraded underwater image are output through the global enhancement branch network; A local enhancement branch network is constructed, and the local enhancement branch results of the degraded underwater image are output through the local enhancement branch network; The global enhancement branch result and the local enhancement branch result are superimposed to obtain the enhanced image of the degraded underwater image; The global enhancement branch network and the local enhancement branch network are trained based on the enhanced image and the reference image until convergence; The underwater image to be enhanced is input into the trained global enhancement branch network and the local enhancement branch network respectively to obtain the enhanced image of the underwater image to be enhanced; The construction of the globally enhanced branch network includes: The degraded underwater image is subjected to bilinear interpolation to obtain a low-resolution degraded underwater image; The low-resolution degraded underwater image is input into a lookup table backbone network to obtain the global contextual features of the degraded underwater image; The degraded underwater image is input into a shallow feature extraction network to obtain the shallow features of the degraded underwater image; The shallow features and the randomly initialized cue components are input into the cue generation module to obtain cue features carrying degradation information; The global context features and the prompt features are input into the prompt interaction module for dynamic feature calibration to obtain the global context features for degradation perception. The global context features of the degradation perception are processed by average pooling, flattening, and two fully connected layers to construct a three-dimensional lookup table as the global enhancement branch network.

2. The underwater image enhancement method according to claim 1, characterized in that, The global enhancement branch results of the global enhancement branch network outputting the degraded underwater image include: The position and index of the RGB value of any pixel in the degraded underwater image are determined in the three-dimensional lookup table. Based on the eight elements located by the index in the three-dimensional lookup table, the color value corresponding to the pixel is calculated by a trilinear interdigitation. The global enhancement branch result is formed based on the color value corresponding to each pixel in the degraded underwater image.

3. The underwater image enhancement method according to claim 1, characterized in that, The step of inputting the shallow features and randomly initialized cue components into the cue generation module to obtain cue features carrying degradation information includes: Perform global average pooling on the spatial dimension of the shallow features; pass Convolutional layers and the Softmax function process the shallow features after global average pooling to generate input-based weights; The prompt feature carrying degradation information is obtained by multiplying the random initialization prompt component with the weighted key.

4. The underwater image enhancement method according to claim 3, characterized in that, The prompt interaction module includes at least a cross-attention module and a multi-branch gating mechanism; the step of inputting the shallow features and randomly initialized prompt components into the prompt generation module to obtain prompt features carrying degradation information includes: Perform global context features Convolution and Depthwise separable convolutions are used to obtain compressed features; The compressed features and the cue features carrying degradation information are input into the cross-attention module to obtain the cue perception features; The global context features are uniformly divided into N segmentation features along the channel dimension, where N is the total number of gating mechanism branches; In each branch, the segmentation features and the cue-aware features are adaptively interacted to obtain the channel-dimensional interaction features of each branch; The interaction features of the channel dimension of each branch are concatenated to obtain the prompt features carrying degradation information.

5. The underwater image enhancement method according to claim 1, characterized in that, The process of constructing a local enhancement branch network and outputting the local enhancement branch result of the degraded underwater image through the local enhancement branch network includes: The degraded underwater image is input into a local feature extraction network to obtain local features; The local features and the randomly initialized prompt components are input into the prompt generation module to obtain prompt features carrying local information; The prompt features carrying local information and the local features are input into the prompt interaction module to obtain the local features of degradation perception; Input the local features of the degradation perception The convolutional layer process yields the local enhancement branch results of the degraded underwater image.

6. The underwater image enhancement method according to any one of claims 1 to 5, characterized in that, The step of training the global enhancement branch network and the local enhancement branch network based on the enhanced image and the reference image until convergence includes: A loss function is established based on the enhanced image and the reference image; The global enhancement branch network and the local enhancement branch network are trained until the loss function value is minimized, thus achieving convergence.

7. An underwater image enhancement device, characterized in that, include: The training set creation module is used to create a training set consisting of degraded underwater images and reference images; A global network construction module is used to construct a global enhancement branch network and output the global enhancement branch result of the degraded underwater image through the global enhancement branch network; The construction of the global enhancement branch network includes: performing bilinear interpolation on the degraded underwater image to obtain a low-resolution degraded underwater image; inputting the low-resolution degraded underwater image into a lookup table backbone network to obtain global context features of the degraded underwater image; inputting the degraded underwater image into a shallow feature extraction network to obtain shallow features of the degraded underwater image; inputting the shallow features and randomly initialized prompt components into a prompt generation module to obtain prompt features carrying degradation information; inputting the global context features and the prompt features into a prompt interaction module for dynamic feature calibration to obtain degradation-aware global context features; and performing average pooling, flattening, and two fully connected layers on the degradation-aware global context features to construct a three-dimensional lookup table as the global enhancement branch network. A local network construction module is used to construct a local enhancement branch network and output the local enhancement branch result of the degraded underwater image through the local enhancement branch network; The overlay processing module is used to overlay the global enhancement branch result and the local enhancement branch result to obtain the enhanced image of the degraded underwater image; The training module is used to train the global enhancement branch network and the local enhancement branch network based on the enhanced image and the reference image until convergence; An enhancement module is used to input the underwater image to be enhanced into the trained global enhancement branch network and the local enhancement branch network, respectively, to obtain the enhanced image of the underwater image to be enhanced.

8. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater image enhancement method according to any one of claims 1 to 6.

9. An electronic device comprising at least a memory and a processor, wherein the memory stores a computer program, characterized in that, The processor implements the steps of the underwater image enhancement method according to any one of claims 1 to 6 when executing the computer program on the memory.

Citation Information

Patent Citations

  • Underwater image enhancement method combining physical prior and deep learning

    CN116309232A