An underwater image enhancement method based on tensor learnable prior

By using a learnable prior model based on the Tensor Train kernel, structured priors are dynamically generated for underwater image enhancement, solving the problem of multimodal degradation adaptation and achieving stable and natural image restoration results, which are suitable for underwater robots and ocean exploration.

CN121582127BActive Publication Date: 2026-04-28DALIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2026-01-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods struggle to adapt to multi-scale scattering, spectral absorption differences, and regional non-uniform degradation simultaneously, resulting in inconsistent colors, over-enhancement, or insufficient detail recovery in the enhancement results.

Method used

We employ a learnable prior model based on the Tensor Train kernel. By dynamically generating the Tensor Train kernel through multi-scale feature extraction and kernel prediction sub-network, we construct a structured prior, perform spectral compensation and scattering suppression, and combine reconstruction loss, color consistency loss and structure preservation loss for training.

Benefits of technology

It achieves stable and natural image enhancement in different underwater environments, avoiding over-enhancement and color drift, and improving contrast and detail recovery, making it suitable for underwater robots and ocean exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582127B_ABST
    Figure CN121582127B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of deep learning and image processing, and discloses an underwater image enhancement method based on a tensor learnable prior. By designing a learnable prior based on a Tensor Train core, explicit modeling of underwater degradation rules is realized. Four Tensor Train cores correspond to height, width, channel and image block modalities, so that the system simultaneously describes vertical structural changes, horizontal scattering, spectral absorption differences and regional scale attenuation, thereby significantly enhancing the description ability of underwater multi-modal degradation. Relying on the core predictor network to dynamically generate the Tensor Train core, the prior structure of the application can be adaptively adjusted according to the input image, improving the stability and generalization ability of the enhancement result. While maintaining the structure edge, the application effectively corrects the color offset and suppresses the scattering blur, so that the enhancement process is changed from the traditional black box mapping to the structured reasoning of physical consistency, and more natural color, higher contrast and better detail recovery effect are achieved in a complex turbid environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of deep learning and image processing technology, and relates to an underwater image enhancement method based on tensor learnable priors. Background Technology

[0002] Underwater imaging environments are characterized by complex factors such as spectral absorption, scattering interference, and ambient light superposition, leading to a general decline in the visual quality of acquired images, including color shift, decreased contrast, blurred details, and an overall bluish or greenish tint. Because different wavelengths of light exhibit significantly different propagation and attenuation patterns in water, imaging results are affected not only by depth but also by water turbidity, dissolved particle concentration, and illumination conditions, resulting in a pronounced multimodal coupling characteristic in the degradation model of underwater images.

[0003] To address underwater image degradation, existing technologies mainly include physically based modeling methods, traditional image enhancement methods, and end-to-end restoration methods based on deep learning. Physically based algorithms rely on pre-defined or estimated water scattering models, but accurately obtaining relevant parameters in real, complex water bodies limits their applicability. Traditional image enhancement methods improve images through histogram equalization, color correction, or filtering, but cannot simultaneously achieve color consistency and structural detail restoration. In recent years, deep learning methods have achieved some success by constructing end-to-end networks for automatic underwater image enhancement. However, most of these methods directly perform convolutional mapping on degraded images, relying on large amounts of training data to learn implicit mapping relationships, lacking explicit modeling of multi-modal underwater degradation patterns. Furthermore, existing deep models generally represent features as unstructured vectors or matrices, making it difficult to effectively capture the coupling relationship between color attenuation, scattering distribution, and spatial structure. This leads to problems such as inconsistent color, over-enhancement, or insufficient detail restoration in enhancement results under different water quality environments.

[0004] Therefore, how to construct a structured prior model that can explicitly express underwater multimodal degradation characteristics and can adaptively learn in deep networks has become a key technical challenge for improving the quality of underwater image enhancement. Summary of the Invention

[0005] This invention aims to address the problem that existing underwater image enhancement methods are unable to simultaneously adapt to multi-scale scattering, spectral absorption differences, and regional non-uniform degradation, and proposes an underwater image enhancement method based on tensor learnable priors.

[0006] To achieve the above objectives, the technical solution proposed in this invention includes: normalizing and color preprocessing the input underwater image to obtain a standardized image for subsequent processing; acquiring initial features such as local texture, scattering distribution, and illumination changes through a multi-scale feature extraction network, and constructing a high-order feature tensor that can simultaneously characterize spatial structure, spectral attenuation, and regional degradation; performing statistical fusion on the high-order feature tensor along modes such as height, width, spectral channels, and region index to obtain a modal statistical vector; dynamically generating a Tensor Train (TT) kernel based on the modal statistical vector of the feature tensor using a kernel prediction sub-network (CPN), thereby constructing a learnable prior based on the Tensor Train kernel; performing structured modeling of the correlation and degradation intensity among the four modes based on this prior; and performing prior constraints, spectral compensation, and scattering suppression on the feature tensor accordingly; and feeding the prior-enhanced features into a decoding and reconstruction network, restoring the image resolution through convolutional layers and residual connections to obtain an enhanced underwater image. During the training phase, the model is able to stably learn the multi-modal degradation patterns of different underwater scenarios through joint constraints of reconstruction loss, color consistency loss, structure preservation loss, and prior regularization loss.

[0007] This invention differs from existing technologies in that:

[0008] Compared to traditional techniques, existing methods mainly rely on end-to-end convolutional networks or scattering correction algorithms based on physical models, lacking a structured representation of underwater degradation at different spatial, spectral, and regional scales. Traditional tensor decomposition methods only use the decomposition results as a dimensionality reduction tool and do not use their structure to describe the degradation process, thus making it difficult to model the non-uniform degradation laws of underwater multi-mode coupling.

[0009] This invention differs from the methods described above by not treating Tensor Train as a static decomposition algorithm, but instead embedding it as a learnable multimodal prior structure within a deep network for the first time. More importantly, the Tensor Train kernels in this invention are not obtained through explicit numerical decomposition, but are dynamically generated based on the modal statistical vectors of the feature tensors via an adaptive kernel prediction subnetwork (CPN). This allows the prior structure to adjust in real-time according to the input scenario, establishing a structured reasoning process for underwater degradation patterns within the network. The prior modeling method, Tensor Train kernel generation mechanism, and modal structure design proposed in this invention are all undisclosed solutions in existing technologies, exhibiting significant non-obviousness and innovation.

[0010] The technical solution of the present invention:

[0011] An underwater image enhancement method based on tensor learnable priors, comprising the following steps:

[0012] Step 1: Normalize and preprocess the input underwater color image to obtain a preprocessed image for subsequent processing;

[0013] Step 2: Process the preprocessed image obtained in Step 1 through a multi-scale feature extraction network to obtain a fused multi-scale feature map, which includes local texture, scattering distribution and illumination changes;

[0014] Step 3: Based on the fusion of multi-scale feature maps, construct a fourth-order feature tensor that can simultaneously characterize spatial structure, spectral attenuation, and image patch degradation;

[0015] Step 4: Perform average pooling or max pooling on the fourth-order feature tensor along the four modalities to obtain the modal statistical vector; and use the kernel prediction sub-network to dynamically generate the Tensor Train kernel based on the modal statistical vector of the fourth-order feature tensor to construct a learnable prior based on the Tensor Train kernel; wherein, the four modalities are height, width, channel and image patch index;

[0016] Step 5: Based on the learnable prior of the Tensor-Train kernel, the correlation and degradation intensity among the four modes are structurally modeled, and prior constraints, spectral compensation and scattering suppression are performed on the fourth-order feature tensor accordingly to obtain the enhanced feature map;

[0017] Step 6: Feed the enhanced feature map into the decoding and reconstruction network to reconstruct the underwater enhanced image through convolutional layers and residual connections;

[0018] Step 7: Joint training is performed using the joint constraints of reconstruction loss, color consistency loss, structure preservation loss, and prior regularization loss to enable the model to stably learn the multi-modal degradation rules of different underwater scenarios; the training phase is used for model parameter learning and is not part of the inference process.

[0019] This invention achieves explicit modeling of underwater degradation patterns by designing learnable priors based on Tensor Train (TT) kernels. Four Tensor Train kernels correspond to four modalities: height, width, channel, and image patch, enabling the system to simultaneously characterize key factors such as vertical structural changes, horizontal scattering, spectral absorption differences, and regional scale attenuation, thus significantly enhancing the ability to describe underwater multimodal degradation. Relying on a kernel prediction subnetwork (CPN) to dynamically generate Tensor Train kernels, the prior structure of this invention can adaptively adjust with the input image. Compared to traditional static decomposition methods, it can better adapt to different water conditions, turbidity levels, and spectral attenuation, improving the stability and generalization ability of the enhancement results. By using learnable prior tensors to perform directional modulation of the feature space at the position, channel, and regional scale, this invention effectively corrects color shifts and suppresses scattering blur while preserving structural edges. This shifts the enhancement process from traditional black-box mapping to physically consistent structured reasoning, achieving more natural colors, higher contrast, and better detail recovery in complex turbid environments.

[0020] The beneficial effects of this invention are:

[0021] This invention introduces learnable priors based on Tensor Train kernels, which can perform structured modeling of multimodal degradation features of underwater images without relying on explicit physical models or fixed convolution operators, making the enhancement process more stable, natural, and closer to the real underwater imaging patterns.

[0022] (1) The overall model is lightweight, robust, and suitable for various computational platforms.

[0023] This invention combines the physical mechanism of underwater imaging with the multi-mode structural representation capability of TT decomposition to achieve adaptive modeling of complex water quality conditions. It has the advantages of lightweight model, high robustness and easy deployment.

[0024] (2) The prior structure has dynamic self-adaptive capability.

[0025] The kernel prediction subnetwork (CPN) dynamically generates the TT kernel based on the modal statistical vector of the input image, enabling the system to automatically adjust prior rules for different depths, water qualities and spectral attenuation conditions, significantly reducing cross-scene enhancement distortion.

[0026] (3) Effectively avoid over-enhancement, color drift and local structural damage.

[0027] The TT prior acts as a modulation term on the feature tensor, which can suppress or enhance local regions, while compensating for high-attenuation spectra while preserving edge details, making the enhancement results more stable and natural.

[0028] (4) The enhancement process has higher interpretability and controllability.

[0029] The four-modal prior structure of this invention corresponds to a clear physical meaning (height, width, spectral channels, and regional scale). The enhancement behavior no longer relies on the network's implicit memory. The adjustment path and enhancement logic are clear and controllable, making it more suitable for verification and debugging in engineering applications.

[0030] In summary, this invention can effectively compensate for underwater multimodal degradation and restore structural consistency while maintaining low model complexity, thereby enhancing the natural stability of the results. It is applicable to various practical scenarios such as underwater robots, marine exploration, and aquatic product monitoring. Attached Figure Description

[0031] Figure 1 This is a diagram illustrating the method steps of an embodiment of the present invention;

[0032] Figure 2 This is a network structure diagram of an embodiment of the present invention;

[0033] Figure 3 This is a schematic diagram of the kernel prediction subnetwork (CPN) structure according to an embodiment of the present invention. Detailed Implementation

[0034] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and technical solutions.

[0035] like Figure 1 As shown, this invention provides an underwater image enhancement method based on tensor learnable priors, the steps of which are as follows:

[0036] Step 1: Normalize and preprocess the input underwater color image to obtain a preprocessed image for subsequent processing;

[0037] Step 2: Process the preprocessed image obtained in Step 1 through a multi-scale feature extraction network to obtain a fused multi-scale feature map, which includes local texture, scattering distribution and illumination changes;

[0038] Step 3: Based on the fusion of multi-scale feature maps, construct a fourth-order feature tensor that can simultaneously characterize spatial structure, spectral attenuation, and image patch degradation;

[0039] Step 4: Perform average pooling or max pooling on the fourth-order feature tensor along the four modalities to obtain the modal statistical vector; and use the kernel prediction sub-network to dynamically generate the Tensor Train kernel based on the modal statistical vector of the fourth-order feature tensor to construct a learnable prior based on the Tensor Train kernel; wherein, the four modalities are height, width, channel and image patch index;

[0040] Step 5: Based on the learnable prior of the Tensor-Train kernel, the correlation and degradation intensity among the four modes are structurally modeled, and prior constraints, spectral compensation and scattering suppression are performed on the fourth-order feature tensor accordingly to obtain the enhanced feature map;

[0041] Step 6: Feed the enhanced feature map into the decoding and reconstruction network to reconstruct the underwater enhanced image through convolutional layers and residual connections;

[0042] Step 7: Joint training is performed using the joint constraints of reconstruction loss, color consistency loss, structure preservation loss, and prior regularization loss to enable the model to stably learn the multi-modal degradation rules of different underwater scenarios; the training phase is used for model parameter learning and is not part of the inference process.

[0043] like Figure 2 As shown, the specific implementation process is as follows:

[0044] Step 1 is as follows:

[0045] First, acquire an underwater color image to be enhanced from the underwater imaging device, and represent it as a size of [size missing]. tensor ,in, and These represent the height and width of the underwater color image, respectively, with three color channels corresponding to R, G, and B. Normalization preprocessing is performed on the underwater color image, linearly mapping the value of each pixel to... Interval, for any spatial location and color channels According to the formula:

[0046]

[0047] Normalization is performed to obtain a normalized image. ;

[0048] To address the issue of dark underwater color images, a gain coefficient is adaptively introduced based on global brightness. ,Will Multiply by This is used to enhance the contrast of dark areas; in addition, the adjusted image is converted from the RGB color space to the YCbCr or Lab color space as needed to obtain a pre-processed image. .

[0049] Step 2 is as follows:

[0050] Preprocessed image The input is fed into the front-end convolutional layer of a multi-scale feature extraction network, where the preprocessed image with three color channels is processed through one or more convolutional operations with non-linear activation. Mapped to a feature representation with a higher number of channels; using a convolutional kernel size of A two-dimensional convolution operator with a stride of 1 and edge padding of 1 is used to process the preprocessed image. Perform convolution operations to obtain the initial feature map. Its size is

[0051]

[0052] in, The number of channels in the initial feature map;

[0053] In the initial feature map Based on this, a multi-branch, multi-scale feature extraction structure is constructed, containing three parallel branches that use convolution operators with different receptive fields to extract the same initial input feature map. Processing: The first branch uses Convolutional and residual convolutional blocks capture the initial feature map. The first branch extracts local texture details and edge contours; the second branch employs a dilated convolution with a dilation rate of 2 to expand the receptive field to medium-range spatial image patches in order to capture the initial feature map. Medium-scale brightness variations and local scattering distribution; the third branch uses a convolutional layer with a 5×5 kernel to perceive the initial feature map. The overall color cast trend and slowly changing background components;

[0054] Let the feature maps obtained from the three parallel branches be as follows: , and They are identical in spatial dimensions but may differ in channel dimensions; the feature maps output by the three parallel branches are concatenated along the channel dimension and then processed through a... The convolutional layers are used for channel fusion and compression to obtain fused multi-scale feature maps. The calculation process is expressed as follows:

[0055]

[0056] in, This indicates a splicing operation in the channel dimension. The size is , This represents the number of channels after merging.

[0057] Step 3 is as follows:

[0058] In fusing multi-scale feature maps Based on this, a fourth-order feature tensor is constructed by partitioning and rearranging image blocks; according to a fixed size Integrate multi-scale feature maps The image is divided into several non-overlapping patches in the spatial dimension, assuming a multi-scale fusion feature map. Divisible by in the height direction Image patches, divisible by a factor of 1 in the width direction. If there are 10 image patches, then the total number of image patches is:

[0059]

[0060] For the There are 1 image patch, and the corresponding image patch is denoted as . Stack all image patches on a new modality and construct a fourth-order feature tensor through reshape and permutation operations. Its dimensions are:

[0061]

[0062] in, These represent the vertical and horizontal spatial positions within the image patch, respectively. Indicates the number of channels. This indicates that different image patches are fused into multi-scale feature maps. Image block index in the image.

[0063] Step 4 is as follows:

[0064] 4.1 Construction of Modal Statistical Vectors;

[0065] The fourth-order feature tensor is subjected to average pooling or max pooling along the four modalities of height, width, channel, and image patch index, respectively, to obtain the four modal statistical vectors. , respectively The channel modal statistics vector is defined as follows:

[0066]

[0067] Indicates the first The first channel in the Location within each image patch The characteristic response value at; where, These represent the vertical position index, horizontal position index, channel index, and image block index within the image block, respectively.

[0068] 4.2 Constructing the kernel prediction subnetwork;

[0069] The kernel prediction subnetwork is designed as follows:

[0070] The kernel prediction subnetwork includes a set of dual-path feature transformation units for processing the input modality statistics vector. Perform parallel feature mapping; convert the input modality statistics vector Two independent feature transformation paths are input; one of these is the main mapping path, used to extract modal statistical vectors. The dominant direction of change in the parameter space; another feature transformation path is the structure modulation path, which is used to characterize the structural change trend inside the modal statistical vector;

[0071] The outputs of the two feature transformation paths described above are respectively expressed as:

[0072]

[0073]

[0074] in, Indicates the first The statistical description vector corresponding to each mode =1, 2, 3, 4; and These are the learnable weight matrices in the main mapping path and the structure modulation path, respectively; and The bias parameters in the main mapping path and the structure modulation path; the superscript (1) indicates the first layer mapping operation in the kernel prediction subnetwork. It is a non-linear activation function. To represent the hyperbolic tangent nonlinear activation function;

[0075] The principal mapping path and the structural modulation path interact through the structural coupling mapping module, enabling the principal modal components and modulation components to influence each other in the parameter space. Their mapping representation is as follows:

[0076]

[0077] in, The first one obtained after mapping by the structure coupling mapping module The fusion feature representation of each modality in the first layer; Indicates the first The modal principal direction feature vectors extracted from the first-layer principal mapping path of each modality are used to characterize the dominant response of the modality in the overall degradation modeling. Indicates the first The structural change feature vectors extracted from the modulation path of the first layer structure of each mode are used to characterize the internal change trend and local modulation information of the mode. In order to be with the first The trainable structural coupling matrix corresponding to each modality is used to map the structural modulation features to a parameter space consistent with the main mapping features, so as to establish the structural correlation between the principal components of the modality and the modulation components.

[0078] Further, morphological reshaping units are introduced to represent the fusion features. Rearrangement and nonlinear expansion are performed; the morphological reshaping unit uses a set of morphological transformation weights. Fusion feature representation of input The expression for fragmentation, chunking, and recombination is as follows:

[0079]

[0080] in, GELU is a nonlinear function;

[0081] Subsequently, a recursive self-modulation mechanism is introduced to process the renormalized representation. A recursive update is performed layer by layer, ensuring that the output of each layer is simultaneously influenced by the current modal information and the modulation result of the previous layer. The update rule is defined as follows:

[0082]

[0083] in, and To assign sovereign weights and recursively modulated weights to different layers;

[0084] Finally, after multiple layers of recursive self-modulation, the fused feature representation of this mode is obtained. The output layer of the kernel prediction subnetwork linearly maps it to the parameter space of the corresponding Tensor Train kernel, resulting in a one-dimensional kernel parameter sequence:

[0085]

[0086] in, For the first The output mapping weight matrix corresponding to each modality is used to map the fused features to the parameter space of the corresponding Tensor Train kernel. This is the bias vector corresponding to the weight matrix;

[0087] And then restore it to a third-order Tensor Train kernel that satisfies the preset rank constraint through a reshape operation:

[0088]

[0089] Obtain the set of Tensor Train kernels that satisfy the preset rank constraint. ;

[0090] 4.3 Learnable priors based on Tensor Train kernels;

[0091] Learnable Prior Tensors Its size is consistent with that of the fourth-order feature tensor. This can be viewed as the result of a chained multiplication of four Tensor Train kernels, for any combination of indices. The components of the learnable prior tensor are represented as:

[0092]

[0093] in, , , , , The intermediate TT rank is used; each Tensor Train kernel is associated with a specific mode of underwater degradation. The prior local structure in the vertical direction of the corresponding image patch is used to describe the texture and edge continuity of the underwater scene in the height direction; Corresponding horizontal structural and scattering characteristics; Corresponding channel modes are used to model the spectral coupling relationship between red, green, and blue channels and the absorption differences of different wavelengths of light in water. The degradation of different spatial image patches is used to reflect the differences in depth, suspended particle concentration, and turbidity among different image patches;

[0094] The learnable prior tensor is reconstructed based on the Tensor-Train kernel chain multiplication relationship. Furthermore, its numerical range is constrained by a monotonic nonlinear function:

[0095]

[0096] Obtain learnable prior tensors based on Tensor Train kernels This learnable prior tensor It corresponds in size to the fourth-order feature tensor. Each element in the graph is interpreted as a magnification or suppression weight corresponding to the height, width, channel, and image patch index.

[0097] Step 5 is as follows:

[0098] Having obtained a learnable tensor prior Subsequently, the learnable prior is directly applied to the fourth-order feature tensor to achieve learnable prior-driven feature enhancement. Scattering suppression is then achieved by suppressing scattering-dominant image patches through prior weights. The scattering-dominant image patches are determined by the weight distribution of the learnable prior on the image patch index mode; image patches with smaller prior weights correspond to regions with stronger scattering influence. Suppression weights are applied to these regions to reduce brightness diffusion and detail blurring caused by water scattering. Specifically, an element-wise multiplication method is used to apply the learnable prior... Multiplying it by the fourth-order feature tensor yields the enhanced fourth-order feature tensor. This can be expressed as a formula:

[0099]

[0100] in, This represents element-wise multiplication;

[0101] Channel modes from the Tensor Train kernel corresponding to the channel The spectral compensation coefficients of each color channel are explicitly extracted to achieve color restoration with clearer physical meaning; channel modes In the third dimension, different color channels correspond to... To aggregate them, use a function. Mapped to spectral compensation coefficients :

[0102]

[0103] right The slicing operation is represented in a fixed color channel. Under these conditions, the channel modes from the Tensor Train kernel Extract the two-dimensional subtensor corresponding to the channel to characterize the degenerate weight distribution of the color channel in the prior constraints; function For two-dimensional subtensors The function performing the convergence mapping employs the method of averaging and norming the slice elements, or is implemented through at least one layer of linear transformation combined with a nonlinear activation function, to map the two-dimensional subtensor into scalar spectral compensation coefficients; for fusing multi-scale feature maps Scale along the channel dimension to obtain the enhanced feature map. That is, execute for each channel:

[0104] .

[0105] Step 6 is as follows:

[0106] After completing the prior constraints, spectral compensation, and scattering suppression, the enhanced fourth-order characteristic tensor will be... Reverting to a two-dimensional convolutional feature map format yields the enhanced feature map. Subsequently, a decoding and reconstruction network consisting of several layers of convolutions and nonlinear activations maps the enhanced feature map back to the image space, generating the final underwater enhanced image. The decoding and reconstruction network includes several residual convolutional blocks to further integrate the structural information of each channel and smooth out detail noise, and finally passes it through a... The convolutional layer compresses the number of channels to 3, and the decoded output reconstructed image. The calculation process can be abstractly represented as follows:

[0107]

[0108] To enhance the stability of the decoding and reconstruction network and preserve the preprocessed image Based on the overall structure, a residual reconstruction method is introduced to output the underwater enhanced image. Represented as a preprocessed image Sum of the network-predicted residual signal:

[0109]

[0110] in, The output from the decoding and reconstruction network is used to compensate for the loss of detail and color caused by underwater degradation.

[0111] Step 7 is as follows:

[0112] During the training phase, paired underwater degraded images and their reference images are used as training data. End-to-end joint optimization is performed on the parameters in the preprocessing, multi-scale feature extraction network, kernel prediction sub-network, learnable priors based on Tensor Train kernels, and decoding / reconstruction network. The underwater degraded images are underwater color images to be enhanced; the reference images are manually collected underwater color images, or approximate reference results obtained through image calibration or image fusion. It is assumed that each pair of samples in the training set is... ,in, This represents the input underwater degraded image. Indicates and The corresponding reference image is used; a comprehensive loss function 𝓛 is constructed, which includes reconstruction error, color constraints, structural constraints, and prior regularization terms, as follows:

[0113]

[0114] in, This represents the L1 norm, used to measure the pixel-level reconstruction error between the output underwater augmented image and its corresponding reference image. The weight coefficient of the loss term is used to balance the impact of different constraints on the training process. Its value is set according to the dataset and application requirements. This represents the color consistency loss, used to constrain the color shift of the enhancement result; and After conversion to the Lab color space, its channel differences are calculated to measure color difference and suppress color cast. The structure preservation loss is used to constrain the structural details and edge consistency of the enhancement result; structural similarity index loss or gradient difference-based loss is used to preserve edge and texture structure. For prior regularization terms; Indicates the first One Tensor Train core; express The squared Frobenius norm is used to regularize the Tensor Train kernel.

Claims

1. An underwater image enhancement method based on tensor learnable priors, characterized in that, The steps are as follows: Step 1: Normalize and preprocess the input underwater color image to obtain a preprocessed image for subsequent processing; Step 2: Process the preprocessed image obtained in Step 1 through a multi-scale feature extraction network to obtain a fused multi-scale feature map, which includes local texture, scattering distribution and illumination changes; Step 3: Based on the fusion of multi-scale feature maps, construct a fourth-order feature tensor that can simultaneously characterize spatial structure, spectral attenuation, and image patch degradation; Step 4: Perform average pooling or max pooling on the fourth-order feature tensor along the four modes to obtain the modal statistics vector; and use the kernel prediction sub-network to dynamically generate the Tensor Train kernel based on the modal statistics vector of the fourth-order feature tensor to construct a learnable prior based on the Tensor Train kernel; wherein, the four modes are height, width, channel and image patch index; Step 5: Based on the learnable prior of the Tensor-Train kernel, the correlation and degradation intensity among the four modes are structurally modeled, and prior constraints, spectral compensation and scattering suppression are performed on the fourth-order feature tensor accordingly to obtain the enhanced feature map; Step 6: Feed the enhanced feature map into the decoding and reconstruction network to reconstruct the underwater enhanced image through convolutional layers and residual connections; Step 7: Joint training is performed using the joint constraints of reconstruction loss, color consistency loss, structure preservation loss, and prior regularization loss to enable the model to stably learn the multi-modal degradation rules of different underwater scenarios; the training phase is used for model parameter learning and is not part of the inference process.

2. The underwater image enhancement method based on tensor learnable priors according to claim 1, characterized in that, Step 1 is as follows: First, acquire an underwater color image to be enhanced from the underwater imaging device, and represent it as a size of [size missing]. tensor ,in, and These represent the height and width of the underwater color image, respectively, with three color channels corresponding to R, G, and B. Normalization preprocessing is performed on the underwater color image, linearly mapping the value of each pixel to... Interval, for any spatial location and color channels According to the formula: Normalization is performed to obtain a normalized image. ; A gain coefficient is adaptively introduced based on global brightness. ,Will Multiply by This is used to enhance the contrast of dark areas; in addition, the adjusted image is converted from the RGB color space to the YCbCr or Lab color space as needed to obtain a pre-processed image. .

3. The underwater image enhancement method based on tensor learnable priors according to claim 2, characterized in that, Step 2 is as follows: Preprocessed image The input is fed into the front-end convolutional layer of a multi-scale feature extraction network, where the preprocessed image with three color channels is processed through one or more convolutional operations with non-linear activation. Mapped to a feature representation with a higher number of channels; using a convolutional kernel size of A two-dimensional convolution operator with a stride of 1 and edge padding of 1 is used to process the preprocessed image. Perform convolution operations to obtain the initial feature map. Its size is in, The number of channels in the initial feature map; In the initial feature map Based on this, a multi-branch, multi-scale feature extraction structure is constructed, containing three parallel branches that use convolution operators with different receptive fields to extract the same initial input feature map. Processing: The first branch uses Convolutional and residual convolutional blocks capture the initial feature map. The first branch extracts local texture details and edge contours; the second branch employs a dilated convolution with a dilation rate of 2 to expand the receptive field to medium-range spatial image patches in order to capture the initial feature map. Medium-scale brightness variations and local scattering distribution; the third branch uses a convolutional layer with a 5×5 kernel to perceive the initial feature map. The overall color cast trend and slowly changing background components; Let the feature maps obtained from the three parallel branches be as follows: , and They are identical in spatial dimensions but may differ in channel dimensions; the feature maps output by the three parallel branches are concatenated along the channel dimension and then processed through a... The convolutional layers are used for channel fusion and compression to obtain fused multi-scale feature maps. The calculation process is expressed as follows: in, This indicates a splicing operation in the channel dimension. The size is , This represents the number of channels after merging.

4. The underwater image enhancement method based on tensor learnable priors according to claim 3, characterized in that, Step 3 is as follows: In fusing multi-scale feature maps Based on this, a fourth-order feature tensor is constructed by partitioning and rearranging image blocks; according to a fixed size Integrate multi-scale feature maps The image is divided into several non-overlapping patches in the spatial dimension, assuming a multi-scale feature map fusion. Divisible by in the height direction Image patches, divisible by a factor of 1 in the width direction. If there are 10 image patches, then the total number of image patches is: For the There are 1 image patch, and the corresponding image patch is denoted as . Stack all image patches on a new modality and construct a fourth-order feature tensor through reshape and permutation operations. Its dimensions are: in, These represent the vertical and horizontal spatial positions within the image patch, respectively. Indicates the number of channels. This indicates that different image patches are fused into multi-scale feature maps. Image block index in the image.

5. The underwater image enhancement method based on tensor learnable priors according to claim 4, characterized in that, Step 4 is as follows: 4.1 Construction of Modal Statistical Vectors; The fourth-order feature tensor is subjected to average pooling or max pooling along the four modalities of height, width, channel, and image patch index, respectively, to obtain four modal statistical vectors. , respectively The channel modal statistics vector is defined as follows: Indicates the first The first channel in the Location within each image patch The characteristic response value at; where, These represent the vertical position index, horizontal position index, channel index, and image block index within the image block, respectively. 4.2 Constructing the kernel prediction subnetwork; The kernel prediction subnetwork is designed as follows: The kernel prediction subnetwork includes a set of dual-path feature transformation units for processing the input modality statistics vector. Perform parallel feature mapping; convert the input modality statistics vector Two independent feature transformation paths are input; one of these is the main mapping path, used to extract modal statistical vectors. The dominant direction of change in the parameter space; another feature transformation path is the structure modulation path, which is used to characterize the structural change trend inside the modal statistical vector; The outputs of the two feature transformation paths described above are respectively expressed as: in, Indicates the first The statistical description vector corresponding to each mode =1, 2, 3, 4; and These are the learnable weight matrices in the main mapping path and the structure modulation path, respectively; and These are the bias parameters in the main mapping path and the structure modulation path; the superscript (1) indicates the first-layer mapping operation in the kernel prediction subnetwork. It is a non-linear activation function. To represent the hyperbolic tangent nonlinear activation function; The principal mapping path and the structural modulation path interact through the structural coupling mapping module, enabling the principal modal components and modulation components to influence each other in the parameter space. Their mapping representation is as follows: in, The first one obtained after mapping by the structure coupling mapping module The fusion feature representation of each modality in the first layer; Indicates the first The modal principal direction feature vectors extracted from the first layer principal mapping path; Indicates the first The structural change feature vectors extracted from the modulation path of the first layer structure of each modality; In order to be with the first The trainable structural coupling matrix corresponding to each mode; Further, morphological reshaping units are introduced to represent the fusion features. Rearrangement and nonlinear expansion are performed; the morphological remodeling unit uses a set of morphological transformation weights. Fusion feature representation of input The expression for fragmentation, chunking, and recombination is: in, GELU is a nonlinear function; Subsequently, a recursive self-modulation mechanism is introduced to process the renormalized representation. A recursive update is performed layer by layer, ensuring that the output of each layer is simultaneously influenced by the current modal information and the modulation result of the previous layer. The update rule is defined as follows: in, and To assign sovereign weights and recursively modulated weights to different layers; Finally, after multiple layers of recursive self-modulation, the fused feature representation of this mode is obtained. The output layer of the kernel prediction sub-network linearly maps it to the parameter space of the corresponding Tensor Train kernel, resulting in a one-dimensional kernel parameter sequence: in, For the first The output mapping weight matrix corresponding to each modality is used to map the fused features to the parameter space of the corresponding Tensor Train kernel. This is the bias vector corresponding to the weight matrix; And then restore it to a third-order Tensor Train kernel that satisfies the preset rank constraint through a reshape operation: Obtain the Tensor Train kernel set that satisfies the preset rank constraint. ; 4.3 Learnable priors based on Tensor Train kernels; Learnable Prior Tensors Its size is consistent with that of the fourth-order feature tensor. This can be viewed as the result of a chained multiplication of four Tensor Train kernels, for any combination of indices. The components of the learnable prior tensor are represented as: in, , , , , The intermediate TT rank is used; each Tensor Train kernel is associated with a specific mode of underwater degradation. The prior knowledge of the local structure in the vertical direction of the corresponding image patch; Corresponding horizontal structural and scattering characteristics; Corresponding channel modes; Degradation corresponding to different spatial image patches; The learnable prior tensor is reconstructed based on the Tensor-Train kernel chain multiplication relationship. Furthermore, its numerical range is constrained by a monotonic nonlinear function: Obtain learnable prior tensors based on Tensor Train kernels This learnable prior tensor It corresponds in size to the fourth-order feature tensor. Each element in the graph is interpreted as a magnification or suppression weight corresponding to the height, width, channel, and image patch index.

6. The underwater image enhancement method based on tensor learnable priors according to claim 5, characterized in that, Step 5 is as follows: Having obtained a learnable tensor prior Subsequently, the learnable prior is directly applied to the fourth-order feature tensor to achieve learnable prior-driven feature enhancement. Scattering suppression is then achieved by suppressing scattering-dominant image patches through prior weights. The scattering-dominant image patches are determined by the weight distribution of the learnable prior on the image patch index mode; image patches with smaller prior weights correspond to regions with stronger scattering influence. Suppression weights are applied to these regions to reduce brightness diffusion and detail blurring caused by water scattering. Specifically, an element-wise multiplication method is used to apply the learnable prior... Multiplying it by the fourth-order feature tensor yields the enhanced fourth-order feature tensor. This can be expressed as a formula: in, This represents element-wise multiplication; Channel modes from the Tensor Train kernel corresponding to the channel The spectral compensation coefficients of each color channel are explicitly extracted to achieve color restoration with clearer physical meaning; channel modes In the third dimension, different color channels correspond to... To aggregate them, use a function. Mapped to spectral compensation coefficients : right The slicing operation is represented in a fixed color channel. Under these conditions, the channel modes from the Tensor Train kernel Extract the two-dimensional subtensor corresponding to the channel to characterize the degenerate weight distribution of the color channel in the prior constraints; function For two-dimensional subtensors The function performing the convergence mapping employs the method of averaging and norming the slice elements, or is implemented through at least one layer of linear transformation combined with a nonlinear activation function, to map the two-dimensional subtensor into scalar spectral compensation coefficients; for fusion of multi-scale feature maps Scale along the channel dimension to obtain the enhanced feature map. That is, execute for each channel: 。 7. The underwater image enhancement method based on tensor learnable priors according to claim 6, characterized in that, Step 6 is as follows: After completing the prior constraints, spectral compensation, and scattering suppression, the enhanced fourth-order characteristic tensor will be... Reverting to a two-dimensional convolutional feature map format yields the enhanced feature map. ; Subsequently, a decoding and reconstruction network consisting of several layers of convolutions and nonlinear activations maps the enhanced feature map back to the image space, generating the final underwater enhanced image. The decoding and reconstruction network includes several residual convolutional blocks to further integrate the structural information of each channel and smooth out detail noise, and finally passes it through a... The convolutional layer compresses the number of channels to 3, and the decoded output reconstructed image. The calculation process can be abstractly represented as follows: To enhance the stability of the decoding and reconstruction network and preserve the preprocessed image Based on the overall structure, a residual reconstruction method is introduced to output the underwater enhanced image. Represented as a preprocessed image Sum of the network-predicted residual signal: in, The output from the decoding and reconstruction network is used to compensate for the loss of detail and color caused by underwater degradation.

8. The underwater image enhancement method based on tensor learnable priors according to claim 7, characterized in that, Step 7 is as follows: During the training phase, paired underwater degraded images and their reference images are used as training data. End-to-end joint optimization is performed on the parameters in the preprocessing, multi-scale feature extraction network, kernel prediction sub-network, learnable priors based on Tensor Train kernels, and decoding / reconstruction network. The underwater degraded images are underwater color images to be enhanced; the reference images are manually collected underwater color images, or approximate reference results obtained through image calibration or image fusion. It is assumed that each pair of samples in the training set is... ,in, This represents the input underwater degraded image. Indicates and The corresponding reference image is used; a comprehensive loss function 𝓛 is constructed, which includes reconstruction error, color constraints, structural constraints, and prior regularization terms, as follows: in, This represents the L1 norm, used to measure the pixel-level reconstruction error between the output underwater augmented image and its corresponding reference image. The weight coefficient of the loss term is used to balance the impact of different constraints on the training process. Its value is set according to the dataset and application requirements. This represents the color consistency loss, used to constrain the color shift of the enhancement result; and After conversion to the Lab color space, its channel differences are calculated to measure color difference and suppress color cast. This represents the structure preservation loss, used to constrain the structural details and edge consistency of the augmentation result; For prior regularization terms; Indicates the first One Tensor Train core; express The squared Frobenius norm is used to regularize the Tensor Train kernel.

Citation Information

Patent Citations

  • Underwater image enhancement method combining physical prior and deep learning

    CN116309232A

  • Remote sensing sewage area identification method and system based on graph structure and multi-stage enhancement

    CN120726484A