A real-time underwater image enhancement method of cross-wave attention guidance and precise color regulation
Patent Information
- Application Number
- CN202610987216.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-07-03
AI Technical Summary
[0006]针对现有深度学习模型在处理复杂图像退化问题时,往往难以在计算效率、实时性与增强质量之间取得有效平衡的技术问题,本发明提供一种基于跨波注意力引导与精准色彩调控的实时水下图像增强方法
[0075] (1) The present invention introduces a color adaptive enhancement mechanism (CAL) at the network front end to perform preliminary color correction on the input image at the optical level. By adjusting the ratio between pixels, it eliminates illumination interference and standardizes the color space, providing high-quality input features for subsequent enhancement stages and laying the foundation for color fidelity.
Smart Images

Figure CN122530040B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image data processing technology, specifically relating to a real-time underwater image enhancement method with cross-wave attention guidance and precise color control. Background Technology
[0002] Because light is severely affected by absorption and scattering effects when propagating in water, the acquired images typically suffer from multiple quality degradation problems: different wavelengths of light attenuate at different rates, resulting in severe color imbalance and a predominantly blue-green hue; suspended particles in the water cause forward scattering, leading to blurred image details and reduced contrast; simultaneously, as water depth increases, ambient light rapidly diminishes, resulting in insufficient overall image brightness and noise. These degradation factors, determined by physical laws, are coupled together, making the acquisition of underwater visual information extremely difficult.
[0003] To address these issues, research methods have evolved from physics-based approaches to deep learning-based approaches. Early physics-based methods attempted to construct underwater imaging models and then solve them using various prior knowledge. However, due to the complexity and variability of the real underwater environment, model parameters are difficult to estimate accurately, leading to inconsistent enhancement effects and limited generalization capabilities across different scenarios.
[0004] In recent years, deep learning-based methods have made significant progress in enhancing images by constructing deep neural networks to directly learn the nonlinear mapping relationship between degraded and sharp images. These data-driven methods eliminate the dependence on precise physical parameters, demonstrating powerful feature extraction and recovery capabilities. However, this leads to a new bottleneck: in pursuit of higher peak signal-to-noise ratios or better visual effects, existing deep learning models are often designed to be extremely complex, resulting in a large number of model parameters and high computational resource consumption. This high complexity sharply contradicts the stringent requirements of underwater mobile devices (such as autonomous underwater vehicles and submersible vision systems) for real-time processing, low latency, and low power consumption.
[0005] Therefore, a core research focus in current underwater image enhancement technology is finding an efficient balance between enhancement performance, model lightweighting, and processing speed. Developing an algorithm that maintains enhancement effectiveness while possessing low computational overhead and high operating efficiency is not only a theoretical exploration at the algorithm design level, but also a crucial step in promoting the real-time engineering deployment of underwater vision technology in complex and dynamic environments, and has significant practical implications for improving the intelligence level of underwater operations. Summary of the Invention
[0006] To address the challenge of existing deep learning models struggling to achieve a balance between computational efficiency, real-time performance, and enhancement quality when dealing with complex image degradation, this invention provides a real-time underwater image enhancement method based on cross-wavelength attention guidance and precise color control. This method introduces a Color Adaptive Enhancement (CAL) module to pre-correct the input image at the optical level, reducing the difficulty of subsequent processing while ensuring accurate color restoration. To achieve efficient and high-quality feature extraction, this invention designs a cross-wavelength attention module (CWA): utilizing learnable discrete wavelet transform to decouple features into low-frequency and high-frequency components, channel weights generated from high-frequency features emphasize key semantic channels in low-frequency features, and spatial weights generated from low-frequency features locate effective edge regions in high-frequency features, making structural information more focused and detail information more precise. This achieves a stronger feature enhancement effect than a single attention mechanism. The entire process is completed at half the image size, reducing computational load while accurately capturing the texture and structural details of the image. In addition, to address the issues of detail loss and contrast imbalance that easily occur in real-time processing, this invention constructs a Feature Correction and Enhancement (FNE) module, which dynamically compensates for lost global details through a local dynamic range perception mechanism, ensuring that an enhanced image with natural visual appearance and rich details is output while meeting high real-time requirements.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] A real-time underwater image enhancement method with cross-wavelength attention guidance and precise color control includes the following steps:
[0009] S1. Global color adaptive preprocessing: Input the original underwater RGB image, use the color transformation matrix to decouple the brightness and chromaticity components, analyze the chromaticity distribution through the feature extractor and guide the proportional recombination, and adjust the color gain and bias through parameterized nonlinear mapping while maintaining the brightness illumination structure. Output the separated and normalized feature map.
[0010] S2. Encoder Feature Extraction: After the preprocessed features are fed into the encoder, spatial details are extracted first through a combination of group convolution and axial convolution. Then, learnable discrete wavelet transform is used to separate frequency features. Cross-wavelet attention is used to achieve complementary information between high and low frequencies. Uncertainty-guided gating and channel attention are used to filter key features. Batch normalization is then performed to output deep semantic features and skip connection features.
[0011] S3. Bottleneck Feature Enhancement: The deep semantic features output by the encoder in step S2 are combined with depthwise convolution and dilated depthwise convolution to form an equivalent large kernel convolution, which expands the receptive field to capture global contextual information while preserving local details, and outputs bottleneck features rich in semantic information.
[0012] S4. Decoder Feature Reconstruction: The bottleneck features are input into the decoder, and after pixel rearrangement to restore the spatial resolution, they are fused with the skip connection features output by the encoder in step S2. High and low layer features are integrated through the attention mechanism, and then the channel dimension is restored through projection convolution to output the initially reconstructed image features.
[0013] S5. Local Adaptive Dynamic Range Mapping: Gaussian blur smoothing is first applied to the image features initially reconstructed in step S4, then local contrast is calculated, and spatial adaptive weights are generated by combining the brightness coefficient calculated based on the global mean. The weighted features are nonlinearly adjusted and mapped to the standard range through local adaptive normalization, and finally the enhanced underwater image is output.
[0014] Further, in step S1, a depthwise separable convolution operator is first defined, with the following formula:
[0015] ,
[0016] in, This represents a k×k depth separable convolution. The kernel size is the side length. For the input feature map, Represents a k×k convolution. This represents a 1×1 convolution, and group=dim indicates that the convolution is grouped according to the channel dimension;
[0017] Then, based on the depthwise separable convolution operator, a chromaticity shift feature extractor is constructed, with the following formula:
[0018] ,
[0019] This represents a chroma shift feature extractor. Represents the ReLU activation function. This represents a 3×3 depth-separable convolution;
[0020] Global color adaptive preprocessing specifically includes:
[0021] (1) Color transformation: The input raw underwater image The luminance and chrominance are decoupled using a color transformation matrix T and transformed to I. Space obtained ;
[0022] (2) Chromaticity shift prediction: using the feature extractor to predict the chromaticity shift. Perform offset prediction to obtain the chromaticity offset prediction value. ;
[0023] (3) Nonlinear mapping: The original luminance component, learnable parameters and chromaticity shift prediction value are used to perform a nonlinear transformation to obtain the final chromaticity shift value. After the overall chromaticity adjustment is completed, the image is converted back by the inverse transformation of the color transformation matrix T. ;
[0024] The specific formula is as follows:
[0025] , ,
[0026] , ,
[0027] , ,
[0028] ,
[0029] in, This represents the original underwater RGB image. Represents the color transformation matrix. Indicates transformation to I Feature map of space, This represents the raw luminance component obtained by summing the RGB three-channel pixel values of a single pixel. , , These are the three-channel pixel values of a single pixel in the original underwater RGB image. This represents the predicted chromaticity shift value for the first path. This represents the second-path chromaticity shift prediction value. This represents the first output branch of the chroma shift feature extractor. This represents the second output branch of the chroma shift feature extractor. This represents the chromaticity shift correction value calculated from the first-path chromaticity shift prediction value. This represents the chromaticity shift correction value calculated from the second-path chromaticity shift prediction value. This represents the corrected feature map output after global color adaptive preprocessing. Represents the color transformation matrix The inverse matrix is represented as carrying the original luminance components. chromaticity components , I Spatial features, , Indicate I Two chromaticity components in space.
[0030] Furthermore, in step S2, the learnable discrete wavelet transform specifically involves defining four 2×2 learnable convolution kernel parameters. and initialized to The input is separated by grouped convolution with a stride of 2. The decomposition yields low-frequency LL and high-frequency components: horizontal LH, vertical HL, and diagonal HH. The calculation formulas for each frequency component are as follows:
[0031] ,
[0032] ,
[0033] in, This represents the learnable low-frequency characteristic components of the discrete wavelet transform. Represents the horizontal high-frequency characteristic components. Represents the vertical high-frequency characteristic components. Represents the diagonal high-frequency characteristic components. This indicates that the axial-convolutional hybrid mechanism outputs a fused feature map. , , , This represents four 2×2 learnable convolutional kernel parameters, with initial values of... .
[0034] Furthermore, in step S2, cross-wave attention guidance specifically includes: splicing the high-frequency information LH, HL, and HH on the channel to obtain... The high-frequency features are obtained by sequentially processing the data through 1×1 convolution to adjust the channels, 3×3 depthwise separable convolution, batch normalization, and ReLU function. The low-frequency information LL obtained by the learnable discrete wavelet transform is used as the low-frequency feature. Construct channel attention (CA) and spatial attention (SA) for low-frequency features. Spatial attention operation and high-frequency features Element-wise multiplication yields High-frequency characteristics Channel attention computation and low-frequency features Element-wise multiplication yields The specific formula is as follows:
[0035] ,
[0036] ,
[0037] ,
[0038] ,
[0039] in, Indicates high-frequency characteristics, Represents a 1×1 convolution. This represents a 3×3 depthwise separable convolution. Indicates batch normalization, Represents the ReLU activation function. This represents the high-frequency splicing input characteristics obtained after channel splicing of LH, HL, and HH. For the input feature map, This represents channel attention operations. Indicates average pooling. This represents the Sigmoid activation function. This represents spatial attention operations. Represents a 7×7 convolution. This indicates element-wise multiplication; `max` indicates taking the maximum value; `mean` indicates taking the average value; and `Concat` indicates concatenating channels. Indicates low-frequency characteristics. High-frequency features that are influenced by low-frequency spatial attention. This indicates low-frequency features affected by the attention of high-frequency channels.
[0040] Furthermore, in step S2, the uncertainty-guided gating specifically includes: concatenating the high and low frequency features obtained from cross-wave attention guidance, and obtaining the initial gating features through 1×1 convolution and the Sigmoid function. Calculate the local variance of the original low-frequency information and take the inverse ratio to obtain the confidence level. Learnable parameters will be introduced. The confidence level is added element-wise to the initial gated features, then the uncertainty-guided gate is generated by the Sigmoid function, and finally multiplied element-wise with the high- and low-frequency features to obtain the enhanced features after uncertainty-gated fusion. The specific formula is as follows:
[0041] ,
[0042] ,
[0043] ,
[0044] ,
[0045] in, Indicates the initial gating feature. Indicates the confidence level. This represents the local variance of the original low-frequency information, where k is a fixed coefficient of 5. This indicates uncertainty-guided gating. Indicates learnable parameters, This indicates element-wise addition. This indicates the enhanced features after uncertainty-gated fusion.
[0046] Further, in step S2, the preprocessing extraction using a hybrid mechanism of group convolution and axial-convolution is as follows: the separated and normalized feature maps output from step S1 are converted to 8 channels via 1×1 pointwise convolution, and then concatenated with 3×1 and 1×3 convolutions to obtain the first axial convolution feature; the first spatial convolution feature is obtained through 3×3 group convolution; the first axial convolution feature and the first spatial convolution feature are fused with learnable parameters and weighted, and then converted to 32 channels via 3×3 group convolution to obtain the fused feature output by the hybrid mechanism of axial-convolution. The specific formula is as follows:
[0047] ,
[0048] ,
[0049] ,
[0050] in, Represents the first axis convolution feature. This represents the corrected feature map output after global color adaptive preprocessing. This represents a 3×1 convolution. This represents a 1×3 convolution. Represents the first spatial convolution feature. This represents a 3×3 group convolution. , Represents the learnable parameters. This indicates the output fused features of the axial-convolution mixing mechanism.
[0051] Furthermore, in step S3, the bottleneck feature enhancement specifically refers to: using the deep features output by the encoder. The original input features are sequentially processed through 3×3 grouped convolution, GeLU function, batch normalization, and 3×3 grouped convolution with a dilation factor of 2 to form an equivalent 7×7 convolutional receptive field. The convolutional features are then residually concatenated with the original input features to obtain the bottleneck enhancement features. The specific formula is as follows:
[0052] ,
[0053] in, This represents the deep features output by the encoder. This represents a 3×3 grouped convolution. This represents batch normalization, and GeLU represents the GeLU activation function. This indicates a bottleneck enhancement feature.
[0054] Further, in step S4, the decoder feature reconstruction specifically involves: upsampling the bottleneck enhancement features through 1×1 convolution and pixel rearrangement, doubling the feature map size to obtain upsampled features; adding the upsampled features element-wise with the encoder skip connection features; then converting the number of channels to 8 through 1×1 convolution and activating with ReLU to obtain fused features; extracting multi-scale spatial features from the fused features using a hybrid mechanism of axial convolution and group convolution; and finally processing them through 3×3 group convolution, channel attention, and batch normalization to output the decoder features, with the specific formula as follows:
[0055] ,
[0056] ,
[0057] ,
[0058] ,
[0059] ,
[0060] in, Indicates upsampling features, This represents the pixel rearrangement function. Represents a 1×1 convolution. This indicates a bottleneck enhancement feature. This indicates the encoder's skip connection feature. This indicates element-wise addition. Represents the ReLU activation function. Indicates fusion characteristics, This represents the second axial convolution feature. This represents a 3×1 convolution. This represents a 1×3 convolution. This represents the second spatial convolution feature. This represents a 3×3 group convolution. , Represents the learnable parameters. Indicates channel attention. Indicates batch normalization, This represents the decoder features.
[0061] Furthermore, in step S5, the local adaptive dynamic range mapping includes local dynamic range calculation and adaptive weight calculation, specifically as follows:
[0062] (a) Local dynamic range calculation: for decoder features Gaussian blur Processing yields fuzzy feature maps An 8×8 window average pooling layer is used to extract the local window mean. The difference between the maximum and minimum values of all window means is calculated to obtain the local dynamic range d-range of the corresponding channel.
[0063] (b) Adaptive weight calculation: Calculate the fuzzy feature map The mean value is used as the global brightness I_global. The deviation between the global brightness I_global and the neutral gray value of 0.5 is calculated, and this is combined with the learnable sensitivity parameter. With local dynamic range d-range modulation, adaptive weights w are generated through the Tanh function, and these adaptive weights w are then combined with the decoder input features. Element-wise multiplication and residual concatenation yield contrast correction features. The specific formula is:
[0064] ,
[0065] ,
[0066] ,
[0067] ,
[0068] ,
[0069] in, express Gaussian blur calculation with window radius. Indicates feature map in channel high width The pixel value corresponding to the position. Indicates the window radius. Represents the natural coefficient. Represents the Gaussian kernel variance. Indicates decoder features, Represents a fuzzy feature map. Indicates the local dynamic range. This indicates average pooling within a t×t window. Indicates the maximum value. This represents the minimum value. Indicates global brightness. This represents the learnable sensitivity parameter. This represents the hyperbolic tangent activation function. This represents the adaptive weight correction feature.
[0070] Furthermore, step S5 also includes local adaptive normalization processing, specifically: dividing the fuzzy feature map into a 6×6 grid and performing local region statistics through average pooling; calculating the maximum value ma and minimum value mi of all window statistics as normalization boundaries; and mapping the features to the standard interval using a pixel-wise normalization method to obtain the locally adaptive normalized output enhanced features. The specific formula is as follows:
[0071] ,
[0072] ,
[0073] in, This indicates the maximum value of the local window statistics. This represents the minimum value of a local window. This indicates that the local adaptive normalization output enhances the features. This represents the zero constant, with a value of 0.001.
[0074] The beneficial effects of this invention are as follows:
[0075] (1) The present invention introduces a color adaptive enhancement mechanism (CAL) at the network front end to perform preliminary color correction on the input image at the optical level. By adjusting the ratio between pixels, it eliminates illumination interference and standardizes the color space, providing high-quality input features for subsequent enhancement stages and laying the foundation for color fidelity.
[0076] (2) This invention uses the Cross-Wave Attention (CWA) mechanism to decouple the learnable wavelet transform features into low-frequency structural components and high-frequency texture components. It uses the channel weights generated by the high-frequency features to highlight the key semantic channels in the low-frequency features, and at the same time uses the spatial weights generated by the low-frequency features to locate the effective edge regions in the high-frequency features. This makes the structural information more focused and the detail information more accurate, achieving a stronger feature enhancement effect than single attention. The entire process is completed at half the size, which reduces the amount of computation and can accurately capture the texture and structural details of the image.
[0077] (3) This invention integrates adaptive normalization and contrast-perceived compensation strategies through the Local Dynamic Range Mapping (FNE) mechanism. By leveraging Gaussian blur-guided local dynamic range analysis, noise interference is effectively suppressed, the contrast distribution characteristics of the image region are quantified, and the contrast dynamic compensation module is driven to generate spatially sensitive weights. While preserving high-frequency details, the normalization mechanism stabilizes the feature distribution, improving the model's adaptability to complex lighting conditions and the consistency of the enhancement effect.
[0078] (4) This invention has excellent real-time deployment capabilities. The overall structure of this invention mainly uses depthwise separable convolution, grouped convolution, 1×1 convolution, average pooling, pixel rearrangement, and element-wise operations, avoiding large-scale global attention calculations and complex iterative optimization processes, thus reducing the number of model parameters, computational load, and GPU memory usage. Simultaneously, the cross-wavelength attention module performs high-low frequency interaction on the low-resolution frequency domain features after wavelet decomposition, further reducing spatial dimension computational overhead. Therefore, this invention can be adapted to resource-constrained platforms such as underwater robots, unmanned underwater vehicles, and edge computing devices, meeting the engineering application requirements of real-time underwater image enhancement tasks. Attached Figure Description
[0079] Figure 1 Overall framework diagram of the invention;
[0080] Figure 2 The color adaptive enhancement mechanism CAL in step S1 of the present invention;
[0081] Figure 3 The cross-wave attention structure CWA in step S2 of the present invention;
[0082] Figure 4 The local dynamic range mapping mechanism FNE in step S5 of the present invention. Detailed Implementation
[0083] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, but the present invention is not limited thereto.
[0084] The overall framework diagram of the present invention is as follows: Figure 1 As shown.
[0085] In practical implementation, the input image can be an RGB image or video frame obtained from an underwater robot, submersible, fixed underwater camera, or underwater acquisition device. For video input, each frame can be fed into the network of this invention for enhancement processing; for single image input, it can be directly used as network input. The input image is first uniformly converted to RGB format, and the pixel values are normalized to ensure the consistency of numerical scale between images acquired by different devices. The network output is still an RGB image, which can be restored to an 8-bit image format after denormalization for real-time display, storage, or input for subsequent visual tasks such as target detection, target recognition, and navigation obstacle avoidance.
[0086] The present invention provides a real-time underwater image enhancement method with cross-wave attention guidance and precise color control, which specifically includes the following steps:
[0087] Step S1: Input through CAL, which contains the following three parts, such as Figure 2 As shown: Color transformation: for the input 3-channel RGB image The luminance and chrominance are decoupled using a color transformation matrix T and transformed to I. Space obtained ;
[0088] Chromaticity shift prediction: using the feature extractor... Perform offset prediction to obtain the chromaticity offset prediction value. .
[0089] Nonlinear mapping: through the original luminance component I, learnable parameters and chromaticity shift prediction value A non-linear transformation is performed to obtain the final color shift, the overall color is adjusted, and then converted back to an RGB image. .
[0090] ,
[0091] ,
[0092] , ,
[0093] ,
[0094] ,
[0095] ,
[0096] Where DSC stands for Depthwise Separable Convolution, and dim is the number of channels in the input feature. Indicates the final chromaticity offset value This represents a k×k convolution kernel.
[0097] Step S2: Convert the preprocessed image to 8 channels using a 1×1 pointwise convolution, and then use a mixed mechanism of group convolution and axial convolution, such as... Figure 1 As shown in the lower left, multi-scale spatial features are extracted and the number of channels is converted to 32 using a 3×3 DSC to obtain the features. The specific formula is as follows:
[0098] ,
[0099] ,
[0100] ,
[0101] in, and These are learnable parameters.
[0102] Then it is processed through CWA, which includes the following three parts, such as Figure 3 As shown:
[0103] Learnable Discrete Wavelet Transform: Defining four 2×2 learnable convolution kernel parameters And initialize each parameter to The input is processed by grouped convolution with a stride of 2. The feature map is separated and its size is halved to obtain the low-frequency LL and the high-frequency components: horizontal LH, vertical HL, and diagonal HH.
[0104] In engineering implementation, four 2×2 wavelet convolutional kernels can be initialized using Haar wavelet filters, enabling the model to possess basic low-frequency structure decomposition and high-frequency texture extraction capabilities from the early stages of training. Subsequently, these kernels are set as learnable parameters, automatically adapting to degradation features in different water bodies, such as light attenuation, color shift, scattering blur, and suspended particle noise, based on training data. This wavelet decomposition process can be implemented using grouped convolutions within a deep learning framework, eliminating the need for additional complex wavelet transform code and facilitating deployment in PyTorch, TensorFlow, ONNX, or TensorRT environments.
[0105] Cross-wave attention guidance: This involves stitching together high-frequency information (LH, HL, HH) across channels. High-frequency features are obtained through preliminary processing using 1×1 convolution to adjust channels, 3×3 depthwise separable convolution, batch normalization, and the ReLU function. The low-frequency information LL obtained by the learnable discrete wavelet transform is used as the low-frequency feature. Construct channel attention (CA) and spatial attention (SA) for low-frequency features. Spatial attention operation and high-frequency features Element-wise multiplication yields High-frequency characteristics Channel attention computation and low-frequency features Element-wise multiplication yields The specific formula is:
[0106] ,
[0107] ,
[0108] ,
[0109] ,
[0110] ,
[0111] ,
[0112] in This indicates element-wise multiplication, max indicates taking the maximum value, mean indicates taking the average value, and Concat indicates concatenating channels.
[0113] Uncertainty-guided gating: for high and low frequency features , Channel concatenation is performed, and an initial gate `base_gate` is obtained through 1×1 convolution and the Sigmoid function. The local variance `var` of the original low-frequency information is calculated, and its inverse ratio is taken to obtain the confidence score. The confidence score is element-wise added to the learnable parameter `base_gate`, and then the uncertainty-guided gate `dynamic_gate` is generated through the Sigmoid function. This gate is then element-wise multiplied with the high- and low-frequency features to obtain the enhanced features. In real-world underwater scenarios, suspended particles, bubbles, and localized reflective areas typically exhibit localized high-frequency disturbances. Directly amplifying high-frequency information can easily lead to the simultaneous amplification of noise, particles, and artifacts. Therefore, this invention estimates the uncertainty of image regions through local variance, reducing enhancement weights in areas with high local variance and retaining higher enhancement strengths in areas with structural edges and stable textures. This process relies solely on average pooling, element-wise operations, and the sigmoid function, resulting in low computational overhead and suitability for real-time inference. Furthermore, this invention reduces the feature map size to half its original size using learnable wavelet decomposition with a stride of 2, and then performs high-low frequency interaction and attention calculations on low-resolution frequency domain features. Compared to directly performing global feature modeling at the original resolution, this approach significantly reduces computational load and memory usage in the spatial dimension, thereby improving model inference speed. The specific formulas and processes are as follows:
[0114] ,
[0115] ,
[0116] ,
[0117] ,
[0118] in This indicates element-wise addition, where k is a fixed coefficient of 5.
[0119] Finally, regarding the output Channel attention is performed sequentially. After filtering key channel information and performing batch normalization, the encoder output is obtained. .
[0120] ,
[0121] ,
[0122] Step S3: Input to the encoder The input is sequentially processed through 3×3 grouped convolution, batch normalization, GeLU activation, and then another 3×3 grouped convolution with a dilation factor of 2 (this combination is equivalent to a 7×7 convolution), followed by... Perform residual connection to obtain the result. The specific formula and process are as follows:
[0123] ,
[0124] Step S4: Transform the input using 1×1 convolution and pixel shuffle. Double the size of the feature map to obtain the feature Then, a jump connection with the encoder. Element-wise addition is performed; then, a 1×1 convolution is used to convert the number of channels to 8, and the ReLU activation function is used to introduce non-linear features to obtain the features. Next, multi-scale spatial features are extracted using an axial fusion block composed of group convolutions and axial convolutions. Finally, the decoder output is obtained through 3×3 depthwise separable convolution (DSC), channel attention (CAT), and batch normalization. .
[0125] ,
[0126] ,
[0127] ,
[0128] ,
[0129] ,
[0130] Step S5: Map the decoder output through a Local Dynamic Range (FNE), which consists of the following three parts, such as... Figure 4 As shown:
[0131] Calculated using local dynamic range: input to the decoder Gaussian blur Processing, then targeting the blurred feature map The mean value of a local window with a size of 8 is extracted using an average pooling layer. Then, the difference between the maximum and minimum values of all window means is calculated to obtain the dynamic range d-range of the corresponding channel.
[0132] Adaptive weight calculation: by calculating feature maps The mean value is used to estimate the global brightness I_global, then the deviation between I_global and the neutral gray value (0.5) is calculated, and then combined with the learnable sensitivity parameter. Generate basic weights, modulate them using a local dynamic range (d-range), obtain adaptive weights w using the Tanh function, and finally combine w with... Element-wise multiplication and union of feature maps Residual connection obtained .
[0133] Local adaptive normalization: for input Local region statistics are performed by dividing the image into a 6×6 grid using average pooling and calculating the mean of each local window. Then, the maximum value *ma* and minimum value *mi* are found from all window statistics as normalization boundaries, and pixel-wise normalization is performed to obtain the locally adaptive normalization output. .
[0134] Different regions may simultaneously exhibit issues such as low light, localized overexposure, color attenuation, and scattering blur. A single global enhancement approach can easily lead to over-enhancement or under-enhancement in certain areas. Therefore, this invention utilizes local window statistics to determine the dynamic range, enabling the network to adaptively adjust the enhancement intensity based on the brightness distribution of different regions. The Gaussian blur, average pooling, maximum value statistics, minimum value statistics, and pixel-wise normalization operations in this module are all conventional image processing or neural network basic operators, easily implemented on edge devices and embedded platforms.
[0135] The specific formula and process are as follows:
[0136] ,
[0137] ,
[0138] ,
[0139] ,
[0140] ,
[0141] ,
[0142] ,
[0143] in, express Gaussian blur calculation with window radius. Indicates feature map in channel high width The pixel value corresponding to the position. Indicates the window radius. Represents the natural coefficient. Represents the Gaussian kernel variance. Indicates decoder features, Representing fuzzy feature maps Indicates the local dynamic range. This indicates average pooling within a t×t window. Indicates the maximum value. This represents the minimum value. Indicates global brightness. This represents the learnable sensitivity parameter. This represents the hyperbolic tangent activation function. This represents the adaptive weight correction feature. Indicates to Average pooling is performed on the window. It is 0.001 (to prevent division by zero).
[0144] This invention uses depthwise separable convolution as its core foundation, significantly reducing the computational cost and number of parameters in the model. By introducing global color adaptive preprocessing, it effectively eliminates illumination interference and lays the foundation for color fidelity. Through the introduction of a cross-wavelength attention mechanism, it achieves focused structural information and accurate extraction of detailed features. Furthermore, by introducing local dynamic range mapping, it enhances the model's adaptability under complex lighting conditions and the stability of the enhancement effect. The real-time performance of this invention does not solely rely on improved hardware performance, but is achieved through a combination of lightweight network structure design, frequency domain downsampling computation, and parallelizable basic operators, thereby reducing the computational burden of single-frame image enhancement at the algorithmic level.
[0145] For continuous video input, this invention can input the video stream captured by an underwater camera frame by frame into a network for enhancement processing. Each frame of the image is normalized and then enters a color adaptive enhancement module, followed by an encoder, a cross-wave attention module, a bottleneck enhancement module, a decoder, and a local dynamic range mapping module, outputting an enhanced RGB image frame. The output result, after inverse normalization, can be directly used for real-time display, video storage, or downstream target detection, target recognition, navigation, obstacle avoidance, and other tasks.
[0146] During engineering deployment, FP32, FP16, or INT8 precision can be selected for inference based on the equipment's computing power, and inference frameworks such as ONNX and TensorRT can be combined for model conversion and acceleration. Since this invention primarily employs easily parallelizable operations such as convolution, pooling, normalization, and element-wise operations, it can fully utilize hardware parallel computing capabilities, reduce single-frame image processing latency, and meet the requirements of underwater real-time image enhancement scenarios for low latency and continuous processing capabilities.
Claims
1. A real-time underwater image enhancement method with cross-wave attention guidance and precise color control, characterized in that, Includes the following steps: S1. Global color adaptive preprocessing: Input the original underwater RGB image, use the color transformation matrix to decouple the brightness and chromaticity components, analyze the chromaticity distribution through the feature extractor and guide the proportional recombination, and adjust the color gain and bias through parameterized nonlinear mapping while maintaining the brightness illumination structure. Output the separated and normalized feature map. S2. Encoder Feature Extraction: After inputting the separated and normalized feature map obtained in step S1 into the encoder, spatial details are first extracted through a combination of group convolution and axial convolution. Then, learnable discrete wavelet transform is used to separate frequency features. Cross-wavelet attention is used to achieve complementary information between high and low frequencies. Uncertainty-guided gating and channel attention are used to filter key features. Batch normalization is performed to output deep semantic features and skip connection features. S3. Bottleneck Feature Enhancement: The deep semantic features output by the encoder in step S2 are combined with depthwise convolution and dilated depthwise convolution to form an equivalent large kernel convolution, which expands the receptive field to capture global contextual information while preserving local details, and outputs bottleneck features rich in semantic information. S4. Decoder Feature Reconstruction: The bottleneck features are input into the decoder, and after pixel rearrangement to restore the spatial resolution, they are fused with the skip connection features output by the encoder in step S2. High and low layer features are integrated through the attention mechanism, and then the channel dimension is restored through projection convolution to output the initially reconstructed image features. S5. Local adaptive dynamic range mapping: Gaussian blur smoothing is first applied to the image features initially reconstructed in step S4, then local contrast is calculated, and spatial adaptive weights are generated by combining the brightness coefficient calculated based on the global mean. The weighted features are nonlinearly adjusted and mapped to the standard range through local adaptive normalization, and finally the enhanced underwater image is output. In step S2, cross-wave attention guidance specifically includes: splicing the high-frequency information LH, HL, and HH on the channel to obtain... The high-frequency features are obtained by sequentially processing the data through 1×1 convolution to adjust the channels, 3×3 depthwise separable convolution, batch normalization, and ReLU function. The low-frequency information LL obtained by the learnable discrete wavelet transform is used as the low-frequency feature. Construct channel attention (CA) and spatial attention (SA) for low-frequency features. Spatial attention operation and high-frequency features Element-wise multiplication yields High-frequency characteristics Channel attention operation and low-frequency features Element-wise multiplication yields The specific formula is as follows: , , , , in, Indicates high-frequency characteristics, Represents a 1×1 convolution. This represents a 3×3 depthwise separable convolution. Indicates batch normalization, Represents the ReLU activation function. This represents the high-frequency splicing input characteristics obtained after channel splicing of LH, HL, and HH. For the input feature map, This represents channel attention operations. Indicates average pooling. This represents the Sigmoid activation function. This represents spatial attention operations. Represents a 7×7 convolution. This indicates element-wise multiplication; `max` indicates taking the maximum value; `mean` indicates taking the average value; and `Concat` indicates concatenating channels. Indicates low-frequency characteristics. High-frequency features that are influenced by low-frequency spatial attention. This indicates low-frequency characteristics affected by the attention of high-frequency channels; In step S2, the uncertainty-guided gating specifically includes: concatenating the high and low frequency features obtained from cross-wave attention guidance, and then obtaining the initial gating features through 1×1 convolution and the Sigmoid function. Calculate the local variance of the original low-frequency information and take the inverse ratio to obtain the confidence level. Learnable parameters will be introduced. The confidence level is added element-wise to the initial gated features, then the uncertainty-guided gate is generated by the Sigmoid function, and finally multiplied element-wise with the high- and low-frequency features to obtain the enhanced features after uncertainty-gated fusion. The specific formula is as follows: , , , , in, Indicates the initial gating feature. Indicates the confidence level. This represents the local variance of the original low-frequency information. The fixed coefficient is 5. This indicates uncertainty-guided gating. Indicates learnable parameters, This indicates element-wise addition. This indicates the enhanced features resulting from uncertainty-gated fusion. In step S5, the local adaptive dynamic range mapping includes local dynamic range calculation and adaptive weight calculation, specifically as follows: (a) Local dynamic range calculation: for decoder features Gaussian blur Processing yields fuzzy feature maps An 8×8 window average pooling layer is used to extract the local window mean. The difference between the maximum and minimum values of all window means is calculated to obtain the local dynamic range d-range of the corresponding channel. (b) Adaptive weight calculation: Calculate the fuzzy feature map The mean value is used as the global brightness I_global. The deviation between the global brightness I_global and the neutral gray value of 0.5 is calculated, and this is combined with the learnable sensitivity parameter. With local dynamic range d-range modulation, adaptive weights w are generated through the Tanh function, and these adaptive weights w are then combined with the decoder input features. Element-wise multiplication and residual concatenation yield contrast correction features. The specific formula is: , , , , , in, express Gaussian blur calculation with window radius. Indicates feature map in channel high width The pixel value corresponding to the position. Indicates the window radius. Represents the natural coefficient. Represents the Gaussian kernel variance. Indicates decoder features, Represents a fuzzy feature map. Indicates the local dynamic range. This indicates average pooling within a t×t window. Indicates the maximum value. This represents the minimum value. Indicates global brightness. This represents the learnable sensitivity parameter. This represents the hyperbolic tangent activation function. This represents the adaptive weight correction feature.
2. The real-time underwater image enhancement method with cross-wave attention guidance and precise color control according to claim 1, characterized in that, In step S1, the depthwise separable convolution operator is first defined, with the following formula: , in, express × Depthwise separable convolution, The kernel size is the side length. For the input feature map, express × convolution, This represents a 1×1 convolution, and group=dim indicates that the convolution is grouped according to the channel dimension; Then, based on the depthwise separable convolution operator, a chroma shift feature extractor is constructed, with the following formula: , This represents a chroma shift feature extractor. Represents the ReLU activation function. This represents a 3×3 depthwise separable convolution; The global color adaptive preprocessing specifically includes: (1) Color transformation: The input raw underwater image The luminance and chrominance are decoupled using a color transformation matrix T and transformed to I. Space obtained ; (2) Chromaticity shift prediction: using the feature extractor to predict the chromaticity shift. Perform offset prediction to obtain the chromaticity offset prediction value. ; (3) Nonlinear mapping: using the original luminance component and learnable parameters and chromaticity shift prediction value The final chromaticity shift value is obtained by performing a nonlinear transformation. After the overall color adjustment is completed, the image is converted back by the inverse transformation of the color transformation matrix T. ; The specific formula is as follows: 、 , , , , , , in, This represents the original underwater RGB image. Represents the color transformation matrix. Indicates transformation to I Feature map of space, This represents the raw luminance component obtained by summing the RGB three-channel pixel values of a single pixel. , , These are the three-channel pixel values of a single pixel in the original underwater RGB image. This represents the predicted chromaticity shift value for the first path. This represents the second-path chromaticity shift prediction value. This represents the first output branch of the chroma shift feature extractor. This represents the second output branch of the chroma shift feature extractor. This represents the chromaticity shift correction value calculated from the first-path chromaticity shift prediction value. This represents the chromaticity shift correction value calculated from the second-path chromaticity shift prediction value. Indicates learnable parameters, This represents the corrected feature map output after global color adaptive preprocessing. Represents the color transformation matrix The inverse matrix is represented as carrying the original luminance components. chromaticity components , I Spatial features, , Indicate I Two chromaticity components in space.
3. The real-time underwater image enhancement method with cross-wave attention guidance and precise color control according to claim 1, characterized in that, In step S2, the learnable discrete wavelet transform specifically involves defining four 2×2 learnable convolution kernel parameters. and initialized to The input is separated by grouped convolution with a stride of 2. The decomposition yields low-frequency LL and high-frequency components: horizontal LH, vertical HL, and diagonal HH. The calculation formulas for each frequency component are as follows: , , in, This represents the learnable low-frequency characteristic components of the discrete wavelet transform. Represents the horizontal high-frequency characteristic components. Represents the vertical high-frequency characteristic components. Represents the diagonal high-frequency characteristic components. This indicates that the axial-convolutional hybrid mechanism outputs a fused feature map. , , , This represents four 2×2 learnable convolutional kernel parameters, with initial values of... .
4. The real-time underwater image enhancement method with cross-wave attention guidance and precise color control according to claim 2, characterized in that, In step S2, the preprocessing is extracted using a hybrid mechanism of group convolution and axial convolution. Specifically, the separated and normalized feature maps output in step S1 are converted into 8 channels by 1×1 pointwise convolution, and then the first axial convolution feature is obtained by concatenating 3×1 convolution and 1×3 convolution respectively. The first spatial convolutional feature is obtained through 3×3 convolutions; the first axial convolutional feature and the first spatial convolutional feature are then fused with learnable parameters and weighted, and then transformed into 32 channels through 3×3 convolutions to obtain the fused feature output by the axial-convolution hybrid mechanism. The specific formula is as follows: , , , in, Indicates the first axis convolution feature, This represents the corrected feature map output after global color adaptive preprocessing. This represents a 3×1 convolution. This represents a 1×3 convolution. Represents the first spatial convolution feature. This represents a 3×3 group convolution. , Represents the learnable parameters. This indicates that the axial-convolutional hybrid mechanism outputs a fused feature map.
5. The real-time underwater image enhancement method with cross-wave attention guidance and precise color control according to claim 1, characterized in that, In step S3, the bottleneck feature enhancement specifically refers to: using the deep features output by the encoder. The original input features are sequentially processed through 3×3 grouped convolution, batch normalization, GeLU function, and 3×3 grouped convolution with a dilation factor of 2 to form an equivalent 7×7 convolutional receptive field. The convolutional features are then residually concatenated with the original input features to obtain the bottleneck enhancement features. The specific formula is as follows: , in, This represents the deep features output by the encoder. This represents a 3×3 grouped convolution. This represents batch normalization, and GeLU represents the GeLU activation function. This indicates a bottleneck enhancement feature.
6. The real-time underwater image enhancement method with cross-wave attention guidance and precise color control according to claim 1, characterized in that, In step S4, the decoder feature reconstruction specifically involves: upsampling the bottleneck enhancement features through 1×1 convolution and pixel rearrangement, doubling the feature map size to obtain upsampled features; adding the upsampled features element-wise with the encoder skip connection features; then converting the number of channels to 8 through 1×1 convolution and activating with ReLU to obtain fused features; extracting multi-scale spatial features from the fused features using a hybrid mechanism of axial convolution and group convolution; and finally processing them through 3×3 group convolution, channel attention, and batch normalization to output the decoder features. The specific formula is as follows: , , , , , in, Indicates upsampling features, This represents the pixel rearrangement function. Represents a 1×1 convolution. This indicates a bottleneck enhancement feature. This indicates the encoder's skip connection feature. This indicates element-wise addition. Represents the ReLU activation function. Indicates fusion features, This represents the second axial convolution feature. This represents a 3×1 convolution. This represents a 1×3 convolution. This represents the second spatial convolution feature. This represents a 3×3 group convolution. , Represents the learnable parameters. Indicates channel attention. Indicates batch normalization, This represents the decoder features.
7. The real-time underwater image enhancement method with cross-wave attention guidance and precise color control according to claim 1, characterized in that, Step S5 also includes local adaptive normalization processing, specifically: dividing the blurred feature map into a 6×6 grid and performing local region statistics through average pooling; calculating the maximum value ma and minimum value mi of all window statistics as the normalization boundary; and mapping the features to the standard interval using a pixel-wise normalization method to obtain the locally adaptive normalized output enhanced features. The specific formula is as follows: , , in, This indicates the maximum value of the local window statistics. This represents the minimum value of a local window. This indicates that the local adaptive normalization output enhances the features. Indicates to Average pooling is performed on the window. This represents the zero constant, with a value of 0.001.
Citation Information
Patent Citations
Underwater image enhancement method based on wavelet Mama
CN121353107A