Degradation-aware self-supervised underwater image super-resolution reconstruction method

Through a self-supervised degradation perception method, the multimodal degradation features of underwater images are extracted and low-quality reconstructed images are generated. Combined with self-supervised training to optimize the super-resolution network, the limitations of underwater image super-resolution reconstruction in existing technologies are solved, and efficient image reconstruction is achieved in complex underwater environments.

CN120450962BActive Publication Date: 2025-09-19SHANDONG UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510953654.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-19
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing underwater image super-resolution reconstruction methods have limitations in training data construction, network structure adaptability and loss function design, making it difficult to effectively improve the quality of underwater images, especially in complex underwater environments.

Method used

A degradation-aware self-supervised method is adopted to extract multimodal degradation features through an underwater degradation-aware pseudo-feature extraction network to generate low-quality reconstructed images. The super-resolution reconstruction network is trained by self-supervised fine-tuning, and a fine-tuning loss function is constructed to optimize the bottleneck layer and decoding layer, realizing self-supervised training without paired labels.

Benefits of technology

It improves the super-resolution reconstruction quality of underwater images, enhances the adaptability and image reconstruction capability of the network in real underwater environments, and can better restore details and structures, improving image clarity and visual quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450962B_ABST
    Figure CN120450962B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing technology, and in particular to a self-supervised underwater image super-resolution reconstruction method based on degradation perception, which obtains a real underwater image, extracts multimodal degradation features of the real underwater image through an underwater degradation-aware pseudo-feature extraction network, and then generates a low-quality reconstructed image of the real underwater image based on the multimodal degradation features; freezes the underwater degradation-aware pseudo-feature extraction network and the super-resolution reconstruction network, unfreezes the bottleneck layer and the decoding layer in the super-resolution reconstruction network for fine-tuning training, and uses the fine-tuned super-resolution reconstruction network to restore the image structure of the real underwater image; during the fine-tuning training process, obtains a super-reconstructed image obtained by reconstructing the real underwater image through the super-resolution reconstruction network, and fine-tunes the bottleneck layer and the decoding layer based on the error calculated based on the super-reconstructed image and the fine-tuning loss function, so as to improve the adaptability and image reconstruction capability in a real underwater environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a self-supervised underwater image super-resolution reconstruction method based on degradation perception. Background Art

[0002] In the fields of computer vision and image processing, image restoration and super-resolution reconstruction technologies have made significant progress with the development of deep learning. Common methods often build on standard technical pathways established in natural image processing tasks, constructing deep neural network models to enhance image clarity and detail. However, direct application of these methods to underwater image restoration and super-resolution tasks faces numerous challenges and struggles to meet the demands for improved image quality in complex underwater environments.

[0003] First, there are fundamental problems with the construction of training data. Most current methods downsample high-resolution images to generate low-resolution samples using artificial rules (such as bicubic interpolation). This approach can only simulate spatial degradation and fails to truly reflect multimodal degradation phenomena such as non-uniform illumination, color shift, blur, and scattering in underwater images. Furthermore, paired high- and low-resolution image pairs are difficult to obtain in real underwater scenes. Existing methods often rely on artificially constructed paired training data for supervised learning, but this data struggles to accurately simulate underwater degradation characteristics. The difficulty in obtaining true high-resolution images leads to poor model performance in real-world scenarios.

[0004] Secondly, in terms of network structure, existing mainstream deep networks typically adopt a fixed structure and lack the ability to dynamically adapt to different degradation characteristics. When the input image exhibits multiple complex degradations or scene changes, the model struggles to make structural adjustments, prone to insufficient stability and poor transferability, which affects the model's effectiveness in complex underwater environments.

[0005] Finally, in terms of loss function design, current methods generally use loss functions that primarily focus on pixel error, ignoring the critical impact of high-frequency image information (such as texture and edges) on image quality. This is particularly true in underwater imaging, where high-frequency components are more susceptible to loss due to optical degradation. Without targeted frequency-domain modeling or perceptual guidance mechanisms, the model is prone to generating blurry and unrealistic images.

[0006] In summary, existing underwater image restoration and super-resolution methods have significant limitations in training data construction, network structure adaptability and loss function design. There is an urgent need to propose more targeted and robust technical solutions for complex underwater environments. Summary of the Invention

[0007] The embodiments of the present application provide a self-supervised underwater image super-resolution reconstruction method based on degradation perception, so as to at least improve the adaptability and image reconstruction capability of the network in a real underwater environment without the participation of high-resolution labeled data.

[0008] To achieve the above objectives, the present invention provides a degradation-aware self-supervised underwater image super-resolution reconstruction method, comprising:

[0009] Degradation modeling steps to obtain real underwater images , extracting the real underwater image through an Underwater Degradation-aware pseudo-feature extraction network (UWDNet) Then, a real underwater image is generated based on the multimodal degradation features. Low-quality reconstructed image ;

[0010] In the self-supervised fine-tuning step, the underwater degradation-aware pseudo-feature extraction network and a super-resolution reconstruction network are frozen, the bottleneck layer and decoding layer in the super-resolution reconstruction network are unfrozen for fine-tuning training, and the fine-tuned super-resolution reconstruction network is used to reconstruct the real underwater image. Perform image structure restoration;

[0011] During the fine-tuning training process, the super-resolution reconstruction network is obtained for real underwater images. Super-resolution reconstruction image obtained by super-resolution reconstruction , combined with real underwater images , low-quality reconstructed images and super-resolution reconstructed images Construct a fine-tuning loss function, calculate the error based on the fine-tuning loss function, fine-tune the bottleneck layer and decoding layer, and optimize the super-resolution reconstructed image after fine-tuning training .

[0012] Based on the above steps, this application effectively integrates the degradation modeling and image reconstruction processes, and improves the modeling capability of underwater image degradation characteristics by extracting multimodal degradation features through UWDNet, so that the generated low-quality images are closer to the actual degradation situation, and the stability and effectiveness of self-supervised training are enhanced. At the same time, by only fine-tuning the key structures in the super-resolution reconstruction network, the network stability is maintained and the reconstruction accuracy is improved, which can better restore the details and structures of underwater images and improve the quality of super-resolution reconstruction. In the absence of real high-resolution paired images, low-quality reconstructed images are used as pseudo-reference samples for fine-tuning training to construct supervision signals without paired labels. End-to-end self-supervised training can be performed directly in real underwater scenes to guide the network to achieve structural restoration and perception optimization on the target domain image, effectively improving the clarity, detail fidelity and overall visual quality of underwater images.

[0013] In some embodiments, the degradation modeling step further comprises:

[0014] Degradation analysis step, extracting real underwater images through degradation encoder The multimodal degradation features of are encoded and the output is the potential degradation vector ,Among them, the multimodal degradation features include: color features, structural features, and frequency ,domain features;

[0015] Degraded reconstruction step, based on the latent degradation vector and a reference high-resolution image Generating realistic underwater images in a modulated image reconstructor Low-quality reconstructed image .

[0016] Based on the above-mentioned degradation analysis step and degradation reconstruction step, the embodiment of the present application extracts color features, structural features, and frequency domain features representing color offset, structural blur, and spectral distortion from real underwater images. Using the degradation vector as a guide, a low-quality reconstructed image with consistent style and structural alignment with the real underwater image is formed in the modulated reconstructor as a pseudo-supervisory signal to replace the real label for subsequent target domain self-supervised training, thereby improving the credibility of the training process and the adaptability to the target domain, solving the problem of degradation modeling distortion caused by the lack of real paired data, enhancing the network's adaptability and generalization capabilities, and improving the robustness and reconstruction effect of the overall model in the target domain.

[0017] In some embodiments, the degradation analysis step includes:

[0018] Extracting real underwater images through a residual backbone network Specifically, the residual backbone network adopts the MSRResNet architecture, uses MSRResNet to extract local structural degradation information in the image, retains texture and edge weakening features, and obtains structural features;

[0019] Real underwater images Perform Fourier transform in RGB space to extract amplitude features and phase features in the spectrum to form the frequency domain features;

[0020] Real underwater images Mapped to HSV color space and YCbCr color space respectively, combined with real underwater images The color features are constructed by using the RGB channels of

[0021] In some embodiments, the degradation analysis step further comprises:

[0022] After the structural features, frequency domain features, and color features are concatenated, subjected to convolution processing, feature fusion, and downsampling, the sampling results are input into the Global Context Attention (GCA) module to extract spatial-channel joint perception features;

[0023] The spatial-channel joint perception feature is output through the average pooling layer and the fully connected layer to obtain the potential degradation vector .

[0024] In some embodiments, the degenerate reconstruction step includes:

[0025] Feature extraction step, through the initial convolution layer on the reference high-resolution image Perform shallow feature extraction to obtain its basic structural color features;

[0026] The style modulation step is to transform the latent degradation vector Input multiple style modulation residual blocks as style modulation parameters, process the basic structural color features through a group of style modulation residual blocks, and compare the processing results with the reference high-resolution image After channel splicing, another set of style modulation residual blocks are input for secondary processing to output the modulated features. The style modulation residual blocks are used to normalize, scale, and offset the feature channels to achieve degradation-consistent modulation of image features.

[0027] A color enhancement step is performed to process the modulated features through a color enhancement module to output a low-quality reconstructed image. , the color enhancement module includes multiple convolutional layers and activation functions.

[0028] Based on the above steps, the style modulation mechanism driven by the potential degradation vector is used to effectively improve the adaptability to different degradation styles, making the image reconstruction process highly adaptable and target consistent, and overall enhancing the accuracy of degradation modeling and the quality of generated images.

[0029] In some embodiments, the loss function of the underwater degradation-aware pseudo feature extraction network is based on pixel reconstruction loss , perceptual consistency loss , color loss Weighted calculation.

[0030] Specifically, the loss function is expressed as: , They are pixel reconstruction losses , perceptual consistency loss , color loss The weight of

[0031] In the above expression, the pixel reconstruction loss The point-by-point difference between the low-quality reconstructed image and the real underwater image in pixel space; perceptual consistency loss To construct a structural similarity metric in the perceptual space by extracting image features using the intermediate layers of the pre-trained VGG16 network.

[0032] In some embodiments, the color loss Calculated based on the following calculation model:

[0033] ,

[0034] in:

[0035] , ;

[0036] is the Lab color space error, is the statistical distribution difference of the red channel, and They are and The weight of is the total number of pixels in the image, the real underwater image For low-quality reconstructed images pixels, The first pixels, 、 Represent the pixel mean values ​​of the real underwater image and the low-quality reconstructed image on the red channel, 、 Represent the standard deviation of the real underwater image and the low-quality reconstructed image on the red channel.

[0037] In some embodiments, the fine-tuning loss function is expressed as:

[0038] ,

[0039] in, is the high-frequency weighted reconstruction loss, For high-frequency perception loss, is the color consistency loss in the Lab color space, is the total variation loss, , , , are the weight hyperparameters of high-frequency weighted reconstruction loss, high-frequency perception loss, color consistency loss, and total variation loss;

[0040] Extract the high-frequency response map of the input image through discrete wavelet transform (DWT) , ;

[0041] The high frequency response diagram As attention weights for pixel reconstruction loss Perform spatial weighting to obtain high-frequency weighted reconstruction loss ,Right now , Represents pixel-by-pixel weighting, guiding the network to more accurately align textures and edges, improving the network's attention to image details;

[0042] The high frequency response diagram Perceptual consistency loss Weighted to obtain high-frequency perception loss ,Right now , guiding the model to pay more attention to detail texture and edge consistency in the perceptual space, thereby improving the subjective restoration quality of the image structure;

[0043] Loss of color consistency Can be based on Lab color space error Calculated, constrained image color distribution in the perceptual space and low-quality reconstructed image Consistent, suppressing color shift;

[0044] Total variation loss It is used to penalize image spatial gradient changes, suppress artifacts and local noise in the reconstructed image, and improve edge continuity and visual smoothness.

[0045] Based on the above-mentioned fine-tuning loss function, the embodiment of the present application uses low-quality reconstructed images as pseudo-supervisory signals to construct a composite optimization mechanism that includes multi-source losses such as high-frequency perception, color consistency and structural smoothness, and fine-tunes the bottleneck layer and decoding layer for target domain adaptability without relying on real high-resolution labels. Through pseudo-feature-driven self-supervision, the super-resolution reconstruction network can effectively improve the structural restoration ability, color fidelity and texture expression of real underwater images, and improve the model's reconstruction ability in texture, edges and color consistency.

[0046] In some embodiments, the total variation loss Calculated based on the following calculation model:

[0047] ,

[0048] in, , Represents the pixel index of the image in the vertical and horizontal directions respectively.

[0049] In some embodiments, the super-resolution reconstruction network is constructed based on Swin Transformer, including: an encoding layer, a bottleneck layer, and a decoding layer, wherein the encoding layer is used to extract initial features and output feature maps at multiple levels;

[0050] The bottleneck layer is composed of a stack of multiple enhanced Swin Transformer blocks, which is used to enhance the global perception and local texture response of key areas in the deepest feature space. The bottleneck layer is equipped with an inter-layer attention mechanism module for cross-scale feature interaction and fusion, so as to facilitate the adaptive compression of features in the bottleneck layer.

[0051] The decoding layer is configured as a multi-scale decoder.

[0052] The decoder starts by gradually upsampling the features from the deepest layer. At each upsampling layer, the current decoder features are fused with the features of the corresponding scale provided by the bottleneck layer to improve the quality of edge details and structure restoration, and output a super-resolved reconstructed image. ‌.

[0053] Compared to related technologies, the degradation-aware self-supervised underwater image super-resolution reconstruction method provided in the embodiments of this application can extract multimodal degradation features from real underwater images and encode them into potential degradation vectors. This is used as a modulation condition to guide the reference high-resolution image, generating a low-quality reconstructed image that is consistent with the input image degradation, thus achieving an organic unification of degradation modeling and pseudo-sample construction. On this basis, by constructing a self-supervised training mechanism, using low-quality reconstructed images as supervisory signals, and through high-frequency response guidance and color structure consistency loss, the super-resolution reconstruction network is adaptively fine-tuned under unlabeled conditions, effectively improving the structural restoration capability of underwater images and the target domain generalization performance.

[0054] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0056] Figure 1 is a flowchart of a method for super-resolution reconstruction of underwater images according to an embodiment of the present application;

[0057] Figure 2 is a step-by-step flow chart of a method for super-resolution reconstruction of underwater images according to an embodiment of the present application;

[0058] Figure 3 is another step-by-step flowchart of the underwater image super-resolution reconstruction method according to an embodiment of the present application;

[0059] Figure 4 2 is a schematic diagram of the principle of an underwater degradation-aware pseudo-feature extraction network according to an embodiment of the present application;

[0060] Figure 5 is a network structure diagram of a degenerate encoder according to an embodiment of the present application;

[0061] Figure 6 Schematic diagram of the principle of the self-supervised fine-tuning step according to an embodiment of the present application;

[0062] Figure 7 is a network structure diagram of an image reconstructor according to an embodiment of the present application;

[0063] Figure 8 2 is a schematic diagram comparing the effects of an underwater degradation-aware pseudo-feature extraction network according to a preferred embodiment of the present application;

[0064] Figure 9This is a schematic diagram comparing the reconstruction effects before and after fine-tuning. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is described and illustrated below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without making any creative efforts are within the scope of protection of this application.

[0066] Obviously, the drawings described below are merely examples or embodiments of the present application. Those skilled in the art can, without inventive effort, apply the present application to other similar scenarios based on these drawings. Furthermore, it is also understood that, although the effort involved in such a development process may be complex and lengthy, for those skilled in the art related to the content disclosed in this application, changes in design, manufacturing, or production based on the technical content disclosed in this application are merely conventional technical means and should not be construed as an insufficiency of the content disclosed in this application.

[0067] References to "embodiments" in this application mean that a particular feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described in this application may be combined with other embodiments unless there is a conflict.

[0068] A color space is a mathematical description of a set of colors.

[0069] The HSV color space is also known as the Hexcone Model. Each color is represented by hue (H), saturation (S), and value (V). The color parameters in this model are: hue (H), saturation (S), and value (V).

[0070] YCbCr color space: Y' represents the brightness (luma) component of a color, while CB and CR represent the blue and red density offsets. Y' and Y are different, with Y representing luminance, which indicates light intensity and is nonlinear, using gamma correction encoding.

[0071] This embodiment provides a self-supervised underwater image super-resolution reconstruction method based on degradation perception. Figures 1 to 3 FIG. 1 is a flow chart of a method for super-resolution reconstruction of underwater images according to an embodiment of the present application. Figures 1 to 3 As shown, the process includes the following steps:

[0072] Degradation modeling step S1, obtaining real underwater images , extracting the real underwater image through an underwater degradation-aware pseudo feature extraction network UWDNet Then, a real underwater image is generated based on the multimodal degradation features. Low-quality reconstructed image ;

[0073] In the self-supervised fine-tuning step S2, the underwater degradation-aware pseudo-feature extraction network UWDNet and a super-resolution reconstruction network are frozen, and the bottleneck layer and decoding layer in the super-resolution reconstruction network are unfrozen for fine-tuning training. The fine-tuned super-resolution reconstruction network is used to reconstruct the real underwater image. To restore the image structure, the bottleneck layer and decoding layer are as follows Figure 6 The blue part in the middle shows the trainable part, that is, the weight parameters of the fixed UWDNet and the weight parameters of other structures of the super-resolution reconstruction network except the bottleneck layer and the decoding layer, such as Figure 6 Shown in yellow.

[0074] During the fine-tuning training process, the super-resolution reconstruction network is obtained for real underwater images. Super-resolution reconstruction image obtained by super-resolution reconstruction , combined with real underwater images , low-quality reconstructed images and super-resolution reconstructed images Constructing fine-tuning loss function , based on the fine-tuning loss function The bottleneck layer and decoding layer are fine-tuned based on the calculation error, and the super-resolution reconstructed image is optimized after fine-tuning training. The super-resolution reconstruction network URSCT of the embodiment of the present application is based on the paper “Reinforced swin-convs transformer for simultaneous underwater sensing sceneimage enhancement and super-resolution”. The network is constructed based on Swin Transformer and includes: an encoding layer, a bottleneck layer and a decoding layer. The encoding layer is used to extract the initial features and output feature maps of multiple levels; the bottleneck layer is composed of a stack of multiple enhanced Swin Transformer blocks, which is used to enhance the global perception and local texture response of key areas in the deepest feature space. An inter-layer attention mechanism module is provided at the bottleneck layer for cross-scale feature interaction and fusion, so as to facilitate adaptive compression features of the bottleneck layer. The decoding layer is set as a multi-scale decoder. The decoder starts to upsample gradually from the deepest layer features. At each level of upsampling layer, the current decoder features are fused with the features of the corresponding scale provided by the bottleneck layer to improve the edge details and structure restoration quality, and output the super-resolution reconstructed image. ,The training strategy adopts a combination of back propagation and gradient descent.

[0075] Based on the above steps, this application effectively integrates the degradation modeling and image reconstruction processes, and improves the modeling capability of underwater image degradation characteristics by extracting multimodal degradation features through UWDNet, so that the generated low-quality images are closer to the actual degradation situation, and the stability and effectiveness of self-supervised training are enhanced. At the same time, by only fine-tuning the key structures in the super-resolution reconstruction network, the network stability is maintained and the reconstruction accuracy is improved, which can better restore the details and structures of underwater images and improve the quality of super-resolution reconstruction. In the absence of real high-resolution paired images, low-quality reconstructed images are used as pseudo-reference samples for fine-tuning training to construct supervision signals without paired labels. End-to-end self-supervised training can be performed directly in real underwater scenes to guide the network to achieve structural restoration and perception optimization on the target domain image, effectively improving the clarity, detail fidelity and overall visual quality of underwater images.

[0076] Combine Figure 2 、 Figure 4 As shown, the underwater degradation-aware pseudo-feature extraction network UWDNet includes a degradation encoder and an image reconstructor, and the degradation modeling step S1 further includes:

[0077] Degradation analysis step S11, extracting the real underwater image through the degradation encoder The multimodal degradation features of are encoded and the output is the potential degradation vector ,Among them, the multimodal degradation features include: color features, structural features, and frequency ,domain features;

[0078] Degraded reconstruction step S12, based on the potential degradation vector and a reference high-resolution image Generating realistic underwater images in a modulated image reconstructor Low-quality reconstructed image .

[0079] Based on the above-mentioned degradation analysis step S11 and degradation reconstruction step S12, the embodiment of the present application extracts color features, structural features, and frequency domain features representing color offset, structural blur, and spectral distortion from real underwater images, and uses the degradation vector as a guide to form a low-quality reconstructed image with consistent style and structural alignment with the real underwater image in the modulated reconstructor as a pseudo-supervisory signal to replace the real label for subsequent target domain self-supervised training, thereby improving the credibility of the training process and the adaptability to the target domain, solving the problem of degradation modeling distortion caused by the lack of real paired data, enhancing the network's adaptability and generalization capabilities, and improving the robustness and reconstruction effect of the overall model in the target domain.

[0080] Figure 5 is a network structure diagram of a degenerate encoder according to an embodiment of the present application, with reference to Figure 5 As shown, the degradation analysis step S11 includes:

[0081] S111: Extracting Real Underwater Images via a Residual Backbone Network Specifically, the residual backbone network adopts the MSRResNet architecture, uses MSRResNet to extract local structural degradation information in the image, retains texture and edge weakening features, and obtains structural features;

[0082] S112: Real underwater images Perform Fourier transform (FFT) in the RGB space to extract amplitude and phase features from the spectrum to form the frequency domain features;

[0083] S113: Real underwater images Mapped to HSV color space and YCbCr color space respectively, combined with real underwater images The RGB channels are fused to construct the color feature; steps S112 and S113 can be executed in parallel.

[0084] S114: The frequency domain features and color features are concatenated and processed through a 1×1 convolution layer and a channel attention mechanism layer to output color frequency domain features, which are then concatenated with the structural features. Then, after 1×1 convolution processing, feature fusion, and three layers of 4×4 downsampling, the sampling results are input into the global context attention module GCA to extract spatial-channel joint perception features, and the 1×1 convolution layer is used to compress the dimension and enhance the channel representation capability, and three layers of 4×4 downsampling are used to form a compact representation.

[0085] S115: The spatial-channel joint perception feature is output through the average pooling layer Pooling and the fully connected layer FC to output the potential degradation vector , with a length of 512.

[0086] Based on the above steps, this application analyzes the structure, frequency and color information of the image by fusing them. This method achieves fine modeling of the complex degradation mechanism of underwater images. The MSRResNet structure can effectively capture weak edges and texture degradation, Fourier frequency domain analysis improves the model's sensitivity to frequency distortion, and multi-color space analysis supplements the perception of color changes. Feature fusion and attention mechanisms enhance the contextual relevance and discriminability of feature representation. The potential degradation vector finally extracted is used as a high-dimensional feature expression, which can significantly improve the performance of subsequent degradation modeling and image reconstruction tasks. Through this multi-dimensional, multi-stage modeling strategy, the ability to fully understand image degradation is improved, thereby enhancing the accuracy and adaptability of the image restoration process.

[0087] Figure 7 This is a network structure diagram of an image reconstructor according to an embodiment of the present application, with reference to Figure 3 、 Figure 7 As shown, the degenerate reconstruction step S12 in the above embodiment includes:

[0088] Feature extraction step S121, through the 3×3 initial convolution layer to the reference high-resolution image Perform shallow feature extraction to obtain its basic structural color features. The initial convolution layer uses LeakyReLU as the activation function, and the reference high-resolution image is a high-resolution image that is paired with the target real underwater image.

[0089] Style modulation step S122, the potential degradation vector Input multiple style-modulated residual blocks as style modulation parameters to guide the adjustment of feature mapping. Each two style-modulated residuals form a group. The basic structural color features are processed by a group of style-modulated residual blocks, and the processed results are compared with the reference high-resolution image. After channel splicing and feature fusion, another set of style modulation residual blocks are input for secondary processing to output the modulated features. The style modulation residual block is used to normalize, scale and offset the feature channels to achieve degradation consistency modulation of image features. In the process, the processing results are compared with the reference high-resolution image. Stitching, to achieve enhanced color consistency and texture restoration capabilities, to achieve adaptive restoration of a variety of degraded style images, to improve the robustness and versatility of image reconstruction in complex changing environments, to introduce the StyleGAN2_Mod module into the style modulation residual block, to use the potential degradation vector As a modulation factor, dynamic modeling is achieved through scaling and offset operations after channel normalization, allowing the network to adjust the feature mapping method according to different degradation distributions, thereby achieving degradation-driven adaptive reconstruction, ensuring that the output image is highly consistent with the original image in terms of structural details and degradation style, thereby effectively enhancing the perceptual authenticity and structural matching of the generated image;

[0090] In the color enhancement step S123, the modulated features are processed by a color enhancement module to adjust the dynamic range and color expression, and output a low-quality reconstructed image. The color enhancement module includes multiple convolutional layers and activation functions. Multiple convolutional layers are used to extract local nonlinear features of color, brightness and texture levels, which helps to enhance the visual vividness of the image. The activation function uses nonlinear activation functions such as ReLU, LeakyReLU or Swish to enhance the contrast between dark and bright parts of the image, making the image have richer tones and visual levels.

[0091] Based on the above steps, the style modulation mechanism driven by the potential degradation vector is used to effectively improve the adaptability to different degradation styles, making the image reconstruction process highly adaptable and target consistent, and overall enhancing the accuracy of degradation modeling and the quality of generated images.

[0092] In some embodiments, the loss function of the underwater degradation-aware pseudo feature extraction network UWDNet is Pixel-based reconstruction loss , perceptual consistency loss , color loss Weighted calculation.

[0093] Specifically, the loss function is expressed as: , They are pixel reconstruction losses , perceptual consistency loss , color loss The weight of

[0094] In the above expression, the pixel reconstruction loss It is the point-by-point difference between the low-quality reconstructed image and the real underwater image in the pixel space, which is used as the basic supervision signal of the image reconstruction process. Optionally, the embodiment of the present application uses absolute difference to calculate the difference between the pixel value of each pixel position of the low-quality reconstructed image and the real underwater image; Perceptual consistency loss In order to extract image features by using the middle layer of the pre-trained VGG16 network, construct a structural similarity metric in the perceptual space, guide the network to more accurately reconstruct the detailed texture in the real image, and improve the subjective visual quality; color loss Calculated based on the following calculation model:

[0095] ,

[0096] in:

[0097] , ;

[0098] is the Lab color space error, is the statistical distribution difference of the red channel, and They are and The weight of is the total number of pixels in the image, the real underwater image For low-quality reconstructed images pixels, The first pixels, 、 Represent the pixel mean values ​​of the real underwater image and the low-quality reconstructed image on the red channel, 、 Represent the standard deviation of the real underwater image and the low-quality reconstructed image on the red channel.

[0099] The above model is based on the Lab color space error and red channel statistical distribution difference , collaboratively modeling the common color shift and red light attenuation phenomena in underwater images, effectively improving the model's color restoration ability and regional color consistency.

[0100] In some embodiments, fine-tuning the loss function Expressed as:

[0101] ,

[0102] in, is the high-frequency weighted reconstruction loss, is the high-frequency perception loss, is the color consistency loss in the Lab color space, is the total variation loss, , , , are the weight hyperparameters of high-frequency weighted reconstruction loss, high-frequency perception loss, color consistency loss, and total variation loss;

[0103] Extract the high-frequency response map of the input image through discrete wavelet transform (DWT) , ;

[0104] The high frequency response diagram As attention weights for pixel reconstruction loss Perform spatial weighting to obtain high-frequency weighted reconstruction loss ,Right now , Represents pixel-by-pixel weighting, guiding the network to more accurately align textures and edges, improving the network's attention to image details;

[0105] The high frequency response diagram Perceptual consistency loss Weighted to obtain high-frequency perception loss , , guiding the model to pay more attention to detail texture and edge consistency in the perceptual space, thereby improving the subjective restoration quality of the image structure;

[0106] Loss of color consistency Constraining image color distribution and low-quality image reconstruction in Lab color space Consistent, suppressing color shift;

[0107] Total variation loss Used to penalize image spatial gradient changes, suppress artifacts and local noise in reconstructed images, improve edge continuity and visual smoothness, and total variation loss Calculated based on the following calculation model:

[0108] ,

[0109] in, , Represents the pixel index of the image in the vertical and horizontal directions respectively.

[0110] Based on the above-mentioned fine-tuning loss function, the embodiment of the present application uses low-quality reconstructed images as pseudo-supervisory signals to construct a composite optimization mechanism that includes multi-source losses such as high-frequency perception, color consistency and structural smoothness, and fine-tunes the bottleneck layer and decoding layer for target domain adaptability without relying on real high-resolution labels. Through pseudo-feature-driven self-supervision, the super-resolution reconstruction network can effectively improve the structural restoration ability, color fidelity and texture expression of real underwater images, and improve the model's reconstruction ability in texture, edges and color consistency.

[0111] The embodiments of the present application are described and illustrated below through preferred embodiments.

[0112] S301: Data preparation and data preprocessing steps

[0113] During the data preparation phase, image data containing a variety of real underwater scenes was selected as experimental subjects. Standardized preprocessing was performed to ensure effective training. Image samples were resized (high-resolution images were cropped to 256×256 pixels, and low-resolution images were scaled to 64×64 pixels), color normalized, and data augmented (including random horizontal and vertical flips and slight perturbations to brightness, contrast, saturation, and hue) to enhance the network's robustness to diverse image styles. Furthermore, the dataset was split into an 8:1:1 ratio based on the training, validation, and testing tasks to ensure objectivity and comprehensiveness in network training and generalization assessment.

[0114] Construction and training of underwater degradation perception pseudo feature extraction network UWDNet: During the network structure initialization process, the degradation encoder is first built, such as Figure 5 As shown. This module includes three key components: the first is the MSRResNet residual backbone network, which is used to extract the structural features of the input image; the second is the color and frequency domain feature extraction module, which contains a frequency domain branch and a color branch. The former performs Fourier transform on the RGB image to extract amplitude features and phase features, and the latter maps the image to RGB, HSV and YCbCr space to extract color features. The two features are fused in the channel dimension after convolution extraction, and the color frequency domain features are constructed through convolution compression; the third is feature fusion and downsampling compression, which splices the above structural features and color frequency domain features at the channel level, and fused and compressed through 1×1 convolution to achieve compact representation of multi-source degradation features. Subsequently, the context dependency perception is enhanced by the GCA module, and finally a degradation vector with a length of 512 is output through average pooling and a fully connected layer. , as a modulation prior for the reconstruction process.

[0115] Then build the modulated image reconstructor module, such as Figure 7 This module is As input, The vector is used as a modulation condition. First, the shallow structural features are extracted through the convolution layer and sent to multiple style modulation residual blocks for nonlinear modeling. Each residual block fuses the degradation information through channel normalization and style vector scaling and offset to achieve degradation consistency reconstruction. The original image is spliced ​​with the current feature to enhance the color restoration ability, and finally the output structure and color style of the color enhancement module are consistent with the original Figure 1 Consistent low-quality reconstructed image .

[0116] During the training phase, multiple loss functions are used to jointly optimize the model performance. As previously mentioned, the training process dynamically adjusts the loss weight coefficient to balance initial pixel restoration with later enhancement of perceptual detail. During training, the system uses a backpropagation algorithm to calculate the gradient of the loss function with respect to network parameters and iteratively updates the gradient using the Adam optimizer. The initial learning rate is set to 1e-4 (0.0001), and the total number of training epochs is 100. During the training process, metrics such as SSIM and PSNR are continuously monitored to evaluate the quality of the artifacts, and the optimal performing model is saved after each epoch.

[0117] S302: Self-supervised fine-tuning step of super-resolution reconstruction network

[0118] like Figure 6 As shown, in this stage, the UWDNet network weights are fixed, and only the bottleneck layer and decoding layer of the super-resolution reconstruction network are unfrozen. During the training process of this stage, the main network reconstructs the image with low quality. is the supervision signal, for the input image The reconstruction results Optimize. High-frequency weighted graph As a significant region guidance factor, and The training process continues with a combination of backpropagation and gradient descent, propagating error signals back to all learnable modules in the main network, further improving the network's reconstruction accuracy and perceptual consistency in the real target domain.

[0119] In order to evaluate the reconstruction effect of this application on underwater image datasets of multiple scenes, quantitative analysis was performed using indicators such as PSNR, SSIM, UIQM and UCIQE. Visual comparison was also used to verify the performance of the proposed method and the original URSCT network (i.e., the super-resolution reconstruction network without fine-tuning) in detail restoration, color correction and perceptual consistency.

[0120] Figure 8FIG. 1 is a schematic diagram showing a comparison of the effects of the underwater degradation-aware pseudo-feature extraction network according to a preferred embodiment of the present application. Figure 8 As shown, the real underwater images of the input in multiple underwater scenes are shown. Corresponding low-quality reconstructed images generated by the UWDNet network From the visualization results, it can be observed that UWDNet can generate reconstruction results that are highly close to the input image in terms of structure, texture, and color distribution in various complex scenes, showing good degradation style modeling and alignment capabilities. The experimental results show that the UWDNet constructed by this invention can effectively capture and reconstruct the multimodal degradation features in real underwater images, providing training guidance images with consistent structure and stable style for subsequent self-supervised training, laying the foundation for accurate optimization.

[0121] In order to evaluate the adaptability and enhancement effect of the super-resolution reconstruction network on the target domain image before and after fine-tuning, multiple indicators such as PSNR (peak signal-to-noise ratio), SSIM (structural similarity), UIQM (underwater image quality index) and UCIQE (underwater image color quality evaluation) were used to quantitatively evaluate the model performance. The specific experimental settings are as follows:

[0122] First, the URSCT, a representative super-resolution reconstruction network, was selected as the backbone network and pre-trained on the publicly available USR248 dataset training set. Its performance was then directly tested on the UFO120 dataset test set as a baseline without fine-tuning. Subsequently, based on steps S1 and S2, 15 real low-quality underwater images were randomly selected from the UFO120 training set. The URSCT model was fine-tuned and optimized without paired high-resolution images. The entire fine-tuning process was performed on a single RTX3090 GPU, with a total training time of approximately 2 minutes for 100 epochs.

[0123] As shown in Table 1, the fine-tuned model performs significantly better than the pre-fine-tuning model in terms of PSNR, SSIM, UIQM, and UCIQE on the UFO120 test set, fully demonstrating that the method of the present invention has significant effects in improving the quality of image reconstruction.

[0124] Table 1 - Performance comparison before and after fine-tuning

[0125]

[0126] like Figure 9As shown in the figure, from top to bottom (a)-(d) are the input underwater low-resolution image, the original URSCT network output, the fine-tuned output image, and the real high-resolution image. The fine-tuned model has achieved significant improvements in edge clarity, texture details, color restoration and overall visual perception, further verifying that this method has good migration capabilities and reconstruction effects in small sample and unlabeled scenarios.

[0127] To further validate the system's adaptability to diverse underwater environments, we conducted scene-specific fine-tuning experiments on typical green and blue water images. We selected a small number of representative green and blue scene images, briefly fine-tuned the pre-trained URSCT model, and performed visual comparisons on the corresponding test images. The results showed that the fine-tuned network output performed particularly well in restoring red areas, removing the residual green cast in the original images to a certain extent and achieving clearer and more natural-looking images. It also outperformed the original model in texture clarity and structural restoration, demonstrating stronger style adaptability and environmental transfer capabilities.

[0128] It should be noted that the steps shown in the above process or the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0129] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0130] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A degradation-aware self-supervised underwater image super-resolution reconstruction method, characterized in that: include: Degradation modeling steps to obtain real underwater images , extracting the real underwater image through an underwater degradation-aware pseudo feature extraction network Then, a real underwater image is generated based on the multimodal degradation features. Low-quality reconstructed image ; In the self-supervised fine-tuning step, the underwater degradation-aware pseudo-feature extraction network and a super-resolution reconstruction network are frozen, the bottleneck layer and decoding layer in the super-resolution reconstruction network are unfrozen for fine-tuning training, and the fine-tuned super-resolution reconstruction network is used to reconstruct the real underwater image. Perform image structure restoration; During the fine-tuning training process, the super-resolution reconstruction network is obtained for real underwater images. Super-resolution reconstruction image obtained by super-resolution reconstruction , combined with real underwater images , low-quality reconstructed images and super-resolution reconstructed images Construct a fine-tuning loss function, calculate the error based on the fine-tuning loss function, fine-tune the bottleneck layer and decoding layer, and optimize the super-resolution reconstructed image after fine-tuning training .

2. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 1, characterized in that: The degradation modeling step further comprises: Degradation analysis step, extracting real underwater images through degradation encoder The multimodal degradation features of are encoded and the output is the potential degradation vector ,Among them, the multimodal degradation features include: color features, structural features, and frequency ,domain features; Degraded reconstruction step, based on the latent degradation vector and a reference high-resolution image Generating realistic underwater images in a modulated image reconstructor Low-quality reconstructed image .

3. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 2, characterized in that: The degradation analysis step comprises: Extracting real underwater images through a residual backbone network Structural features; Real underwater images Perform Fourier transform in RGB space to extract amplitude features and phase features in the spectrum to form the frequency domain features; Real underwater images Mapped to HSV color space and YCbCr color space respectively, combined with real underwater images The color features are constructed by using the RGB channels of 4. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 3, characterized in that: The degradation analysis step further includes: After the structural features, frequency domain features, and color features are concatenated, subjected to convolution processing, feature fusion, and downsampling, the sampling results are input into the global context attention module to extract spatial-channel joint perception features; The spatial-channel joint perception feature is output through the average pooling layer and the fully connected layer to obtain the potential degradation vector .

5. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 2, characterized in that: The degenerate reconstruction step includes: Feature extraction step, through the initial convolution layer on the reference high-resolution image Perform shallow feature extraction to obtain its basic structural color features; The style modulation step is to transform the latent degradation vector Input multiple style modulation residual blocks as style modulation parameters, process the basic structural color features through a group of style modulation residual blocks, and compare the processing results with the reference high-resolution image After channel splicing, another set of style modulation residual blocks are input for secondary processing to output the modulated features; A color enhancement step is performed to process the modulated features through a color enhancement module to output a low-quality reconstructed image. .

6. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to any one of claims 1 to 5, characterized in that: The loss function of the underwater degradation-aware pseudo feature extraction network is based on pixel reconstruction loss , perceptual consistency loss , color loss Weighted calculation.

7. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 6, characterized in that: The color loss Calculated based on the following calculation model: , in: , ; is the Lab color space error, is the statistical distribution difference of the red channel, and They are and The weight of is the total number of pixels in the image, the real underwater image For low-quality reconstructed images pixels, The first pixels, 、 Represent the pixel mean values ​​of the real underwater image and the low-quality reconstructed image on the red channel, 、 Represent the standard deviation of the real underwater image and the low-quality reconstructed image on the red channel.

8. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 7, characterized in that: The fine-tuning loss function is expressed as: , in, is the high-frequency weighted reconstruction loss, For high-frequency perception loss, is the color consistency loss in the Lab color space, is the total variation loss, , , , are the weight hyperparameters of high-frequency weighted reconstruction loss, high-frequency perception loss, color consistency loss, and total variation loss; Extract the high-frequency response map of the input image through discrete wavelet transform , ; The high frequency response diagram As attention weights for pixel reconstruction loss Perform spatial weighting to obtain high-frequency weighted reconstruction loss ,Right now ; The high frequency response diagram Perceptual consistency loss Weighted to obtain high-frequency perception loss ,Right now .

9. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 8, characterized in that: Total variation loss Calculated based on the following calculation model: , in, , Represents the pixel index of the image in the vertical and horizontal directions respectively.

10. The degradation-aware self-supervised underwater image super-resolution reconstruction method according to claim 9, wherein: The super-resolution reconstruction network is constructed based on Swin Transformer, including: an encoding layer, a bottleneck layer and a decoding layer. The bottleneck layer is provided with an inter-layer attention mechanism module, and the decoding layer is configured as a multi-scale decoder.

Citation Information

Patent Citations

  • Unsupervised implicit modeling blind super-resolution reconstruction method and device

    CN116523739A

  • Infrared image super-resolution reconstruction method based on contrast learning

    CN119624777A