A synchronous feature fusion underwater image enhancement method and system
Patent Information
- Application Number
- CN202411034034.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-07-30
AI Technical Summary
[0004]本发明的目的在于解决现有技术中在进行图像增强时图像细节无法被保留的问题,提供一种同步式特征融合的水下图像增强方法及系统
[0035] This invention proposes a synchronous feature fusion underwater image enhancement method. By designing an adaptive detail-weighted loss, it combines Transformer and CNN to establish an underwater image enhancement network model. The model is trained on an open dataset of real underwater scenes and using the adaptive detail-weighted loss. The distorted image is input into the trained model, and the enhanced underwater image is output, thus achieving underwater image enhancement. The synchronous feature fusion underwater image enhancement network model not only leverages the structural advantages of each feature but also integrates local and global features at different scales, overcoming the problem of missing effective information in existing enhancement techniques. Simultaneously, the adaptive detail-weighted loss dynamically allocates loss weights based on the magnitude of loss between different pixels, allowing the model to focus more on the more severely degraded areas, effectively achieving underwater image enhancement.
Smart Images

Figure CN119067863B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology and relates to a method and system for underwater image enhancement using synchronous feature fusion. Background Technology
[0002] Underwater images are affected by optical phenomena such as water scattering, resulting in visual quality problems such as color distortion, reduced contrast, and image blur. These optical challenges limit the accurate acquisition and clear display of image information in the underwater environment, which is not conducive to underwater exploration, underwater robot vision, marine biological research and other tasks. Therefore, underwater image enhancement technology is of great significance for improving image quality and obtaining clear underwater images.
[0003] Existing underwater image enhancement methods can be categorized into physical model-based, visual prior-based, and deep learning-based methods. Physical model-based methods estimate environmental and transmission parameters through conditional assumptions and prior knowledge, then deduce the underwater image degradation process in reverse. This approach relies on manual priors, is time-consuming, and highly dependent on parameters, making it unsuitable for complex underwater scenarios. Visual prior-based methods, on the other hand, do not consider optical imaging models; they enhance image quality simply by adjusting pixel values. While easy to implement, this approach ignores imaging variations under different aquatic environments, resulting in limitations and often exhibiting under-enhancement or over-enhancement. Typical deep learning networks include CNNs, GANs, and Transformers. These networks improve the efficiency of underwater image enhancement by constructing neural networks to learn the end-to-end mapping between distorted and reference images on a dataset. CNNs, through convolutional kernel sliding filters, can accurately extract local image features, exhibiting better generalization ability and faster convergence speed. However, this method still has obvious shortcomings: (1) Since the convolution kernel has a limited receptive field, it focuses more on the local features of the image, thus making CNN lack the ability to perceive global information; (2) The convolution filter has static weights during the calculation process, and cannot adaptively enhance regions with inconsistent attenuation in the input image. GAN can learn and understand the global features and overall distribution of underwater image data to a certain extent by combining the generator and the discriminator. However, the training of GAN is usually more complex and unstable, and may face the problem of mode collapse, which means that the image quality generated by the generator has a certain degree of degradation. Transformer has performed well on several vision tasks with its advanced global modeling capabilities. However, when faced with underwater images with color distortion, blurriness and low contrast, Transformer's focus on long-distance dependencies often makes it impossible to preserve the details of the enhanced image. Summary of the Invention
[0004] The purpose of this invention is to solve the problem that image details cannot be preserved when performing image enhancement in the prior art, and to provide an underwater image enhancement method and system with synchronous feature fusion.
[0005] To achieve the above objectives, the present invention employs the following technical solution:
[0006] The present invention proposes a synchronous feature fusion underwater image enhancement method, comprising the following steps:
[0007] Obtain an open dataset of real-world underwater scenarios and design an adaptive detail-weighted loss;
[0008] An underwater image enhancement network model was established and trained based on an open dataset of real underwater scenes and adaptive detail-weighted loss.
[0009] The distorted image is input into the trained underwater image enhancement network model, and the enhanced underwater image is output, thus achieving underwater image enhancement.
[0010] Preferably, the underwater image enhancement network model includes a global feature extraction module and a detail feature extraction module.
[0011] Preferably, the global feature extraction module is specifically as follows:
[0012] enter Divide the input X into non-overlapping patch blocks. Each patch is treated as a "token," and after layer normalization, it enters two convolutional layers to obtain the spatial information of the original features. The feature map is then projected onto a feature representation with three times the number of channels of the input X. Then, the obtained feature X′ Patch The relative positional deviation is calculated using window-based multi-head self-attention; finally, the global feature map is obtained through layer normalization and MLP. in, It is a symbol for representing image data as a tensor, where H represents the number of rows, W represents the number of columns, and C represents the number of channels.
[0013] Preferably, the global feature extraction module consists of an even number of modules connected in series. Two consecutive global feature extraction modules use a conventional and a shifted window multi-head self-attention structure, respectively, and the output tensor dimension always remains unchanged from the input.
[0014] Preferably, the detail feature extraction module is specifically as follows:
[0015] For input A 1×1 convolution kernel is used to perform feature transformation on the input and expand the channels to obtain the first convolution result;
[0016] Based on the first convolution result, three types of convolution kernels, 1×3, 3×3 and 3×1, are used to perform convolution operations on the same input in different directions and sizes, and batch normalization is performed. The outputs of the three layers are directly added together to obtain the detailed features of the image.
[0017] The final detail features are obtained by aligning the channel dimensions with the detail feature map of the image through a 1×1 convolution operation.
[0018] Preferably, the adaptive detail-weighted loss is specifically:
[0019]
[0020] Where α, γ, and β are all hyperparameters. For adaptive detail-weighted loss function, For the perceptual loss function, This is the gradient difference loss function.
[0021] Preferably, the adaptive detail-weighted loss function as follows:
[0022]
[0023] The perceptual loss function as follows:
[0024]
[0025] The gradient difference loss function as follows:
[0026]
[0027] in, and y i These are the enhanced image and the true ground image, α i It is the loss value after the i-th pixel, where ∈ is a constant e. -6 N is the number of scales at which the image is decomposed. This represents the SSIM value at each scale, β. i These are weight parameters at different scales; and These represent the gradients of the enhanced image and the reference image at each pixel, respectively.
[0028] This invention proposes a synchronous feature fusion underwater image enhancement system, comprising:
[0029] The parameter acquisition module is used to acquire open datasets of real underwater scenes and design an adaptive detail-weighted loss.
[0030] The model building module is used to build an underwater image enhancement network model, which is trained based on an open dataset of real underwater scenes and adaptive detail-weighted loss.
[0031] The data processing module is used to input distorted images into a trained underwater image enhancement network model and output enhanced underwater images to achieve underwater image enhancement.
[0032] A terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of a synchronous feature fusion underwater image enhancement method.
[0033] A computer-readable storage medium storing a computer program, characterized in that, when executed by a processor, the computer program implements the steps of a synchronous feature fusion underwater image enhancement method.
[0034] Compared with the prior art, the present invention has the following beneficial effects:
[0035] This invention proposes a synchronous feature fusion underwater image enhancement method. By designing an adaptive detail-weighted loss, it combines Transformer and CNN to establish an underwater image enhancement network model. The model is trained on an open dataset of real underwater scenes and using the adaptive detail-weighted loss. The distorted image is input into the trained model, and the enhanced underwater image is output, thus achieving underwater image enhancement. The synchronous feature fusion underwater image enhancement network model not only leverages the structural advantages of each feature but also integrates local and global features at different scales, overcoming the problem of missing effective information in existing enhancement techniques. Simultaneously, the adaptive detail-weighted loss dynamically allocates loss weights based on the magnitude of loss between different pixels, allowing the model to focus more on the more severely degraded areas, effectively achieving underwater image enhancement.
[0036] This invention proposes an underwater image enhancement system based on synchronous feature fusion. By dividing the system into a parameter acquisition module, a model construction module, and a data processing module, underwater image enhancement is ultimately achieved. The modular approach ensures that each module is independent, facilitating unified management of all modules. Attached Figure Description
[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0038] Figure 1 This is a flowchart of the underwater image enhancement method with synchronous feature fusion according to the present invention.
[0039] Figure 2 This is the overall network architecture diagram of the present invention.
[0040] Figure 3 This is a structural diagram of the global feature extraction module of the present invention.
[0041] Figure 4 This is a structural diagram of the detailed feature extraction module of the present invention.
[0042] Figure 5 This is a diagram of the underwater image enhancement system with synchronous feature fusion according to the present invention.
[0043] Figure 6 This is a schematic diagram of the structure of an electronic device according to the present invention. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0045] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0046] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0047] The present invention will now be described in further detail with reference to the accompanying drawings:
[0048] This invention proposes a synchronous feature fusion underwater image enhancement method, such as... Figure 1As shown, it includes the following steps:
[0049] S1. Obtain an open dataset of real underwater scenes and design an adaptive detail-weighted loss;
[0050] The adaptive detail-weighted loss is specifically as follows:
[0051]
[0052] Where α, γ, and β are all hyperparameters. For adaptive detail-weighted loss function, For the perceptual loss function, This is the gradient difference loss function.
[0053] The adaptive detail-weighted loss function as follows:
[0054]
[0055] The perceptual loss function as follows:
[0056]
[0057] The gradient difference loss function as follows:
[0058]
[0059] in, and y i These are the enhanced image and the true ground image, α i It is the loss value after the i-th pixel, where ∈ is a constant e. -6 N is the number of scales at which the image is decomposed. This represents the SSIM value at each scale, β. i These are weight parameters at different scales; and These represent the gradients of the enhanced image and the reference image at each pixel, respectively.
[0060] S2. Establish an underwater image enhancement network model, and train the underwater image enhancement network model based on an open dataset of real underwater scenes and adaptive detail-weighted loss.
[0061] The underwater image enhancement network model specifically includes a global feature extraction module and a detail feature extraction module.
[0062] The global feature extraction module is as follows:
[0063] enter Divide the input X into non-overlapping patch blocks. Each patch is treated as a "token," and after layer normalization, it enters two convolutional layers to obtain the spatial information of the original features. The feature map is then projected onto a feature representation with three times the number of channels of the input X. Then, the obtained feature X′ Patch The relative positional deviation is calculated using window-based multi-head self-attention; finally, the global feature map is obtained through layer normalization and MLP. in, It is a symbol for representing image data as a tensor, where H represents the number of rows, W represents the number of columns, and C represents the number of channels.
[0064] The global feature extraction module consists of an even number of modules connected in series. Two consecutive global feature extraction modules use a multi-head self-attention structure with a regular window and a shifted window, respectively, and the output tensor dimension always remains unchanged from the input.
[0065] The detailed feature extraction module is as follows:
[0066] For input A 1×1 convolution kernel is used to perform feature transformation on the input and expand the channels to obtain the first convolution result;
[0067] Based on the first convolution result, three types of convolution kernels, 1×3, 3×3 and 3×1, are used to perform convolution operations on the same input in different directions and sizes, and batch normalization is performed. The outputs of the three layers are directly added together to obtain the detailed features of the image.
[0068] The final detail features are obtained by aligning the channel dimensions with the detail feature map of the image through a 1×1 convolution operation.
[0069] S3. Input the distorted image into the trained underwater image enhancement network model, and output the enhanced underwater image to achieve underwater image enhancement.
[0070] The following is combined with Figures 2 to 4 This underwater image enhancement method is described in detail:
[0071] We acquired the large-scale underwater dataset EUVP, containing over 16,000 image pairs used to train the network model, and 515 corresponding image pairs used as a full-reference test set to evaluate the model. To test the model's generalization ability, we also used the UIEB and U45 datasets as non-reference test sets.
[0072] Transformers excel in long-range dependency and global feature extraction, while CNNs demonstrate powerful performance in extracting local detail features. To better leverage the strengths of both, a synchronous feature fusion underwater image enhancement network framework was designed. The same input is fed to both modules. Feature map spatial alignment is achieved by using window reset and residual connections in the Transformer block, and by adjusting convolutional kernel parameters in the CNN block. Furthermore, 1×1 convolutions are used in the CNN to ensure that the output feature map maintains consistency with the Transformer feature map in the channel dimension, facilitating feature fusion.
[0073] like Figure 2 This is the overall network architecture of the present invention. The network follows the U-Net principle and is an encoder-decoder structure, where the output of each encoder is fed into its corresponding mirror decoder via skip connections. Given an underwater distorted image... Low-level features of the original image are obtained through a synchronous feature fusion module. Where 256×256 represents the spatial dimension of the feature map, and 3 represents the number of channels. Specifically, the input image... The data is transmitted to two separate modules. The global feature extraction module processes the input image I and outputs the global features. The detail feature extraction module processes the input image I and outputs the detail features. After fusing the two outputs, the low-level features of the original image are obtained. Next, the spatial size is reduced by downsampling while the number of channels is increased, and then a synchronous feature fusion module is used to obtain deeper features. Finally, the latent features of the image are obtained by continuing to use downsampling operations and a synchronous feature fusion module. The high-quality ground image is then progressively reconstructed using this as input to the decoder. The decoder consists solely of a global feature extraction module. This is because the decoder needs to transform abstract encoded features into the spatial structure of the image. The Transformer, through self-attention and positional encoding, can more naturally generate sequence data corresponding to the image space. To assist the reconstruction process, low-level image features from the encoder are fused with high-level image features from the symmetric decoder. This allows the network to preserve the fine structure and texture features of the original image during the enhancement process, ultimately generating a high-quality visually perceptual enhanced image.
[0074] like Figure 3This is the structure diagram of the global feature extraction module of this invention. It is an improvement upon the Swin Transformer, embedding stacked 1×1 and 3×3 convolutions into the Transformer block, allowing the model to better preserve the spatial information of the image. This module inherits the original basic architecture, employs a windowing strategy to process images, and achieves efficient processing of large images through block-based and localized attention mechanisms. This enables the model to flexibly model at various spatial scales and possesses complex linear computation methods suitable for images of different sizes. For input... First, divide it into non-overlapping patch blocks. Each patch is treated as a "token," an abstract representation of the original RGB image pixels. After layer normalization, it enters two convolutional layers to obtain the spatial information of the original features, and then projects the feature map onto its three-dimensional feature representation. For feature X′ Patch A window-based multi-head self-attention method is used to calculate the relative positional deviation. Subsequently, layer normalization and an MLP are applied to obtain the global feature map.
[0075] The global feature extraction layer consists of an even number of the above-mentioned single modules cascaded together. Two consecutive global feature extraction modules use conventional and shifted window multi-head self-attention structures respectively, ensuring that the output tensor dimension remains constant with the input. Specifically, the output of the previous module... As the input to the current module, the input undergoes layer normalization and a multi-head attention structure based on a regular window to obtain a feature map. This feature map is then added to the input to obtain the output feature of the multi-head attention. After layer normalization and MLP, a new feature map is obtained. This new feature map is... The output X of the current module is obtained by adding them together. l ;X l As input to the next module, a feature map is obtained after layer normalization and a multi-head attention structure based on a shift window. This feature map is then added to the input to obtain the output feature of the multi-head attention. After layer normalization and MLP, a new feature map is obtained. This new feature map is... The output X of the current module is obtained by adding them together. l+1 The calculation mechanism is as follows:
[0076]
[0077] in, and X lrepresents the output features of the (S)W-MSA and MLP of the l-th module, respectively; W-MSA and SW-MSA represent multi-head self-attention based on regular windows and shifted windows, respectively; LN represents layer normalization.
[0078] For a given relative position of the input, the output is calculated using a multi-head self-attention mechanism:
[0079]
[0080] Where Q, K, and V represent the query, key, and value matrices, respectively. QK T This indicates that the query matrix and the key matrix are subjected to a dot product operation. This indicates the scaling factor, and softmax represents the softmax function.
[0081] Figure 4 This is a structural diagram of the detail feature extraction module of the present invention. The key mechanism of this module lies in utilizing the characteristics of convolutional kernels of different sizes: a 3×3 convolutional kernel provides a wider receptive field, which helps to capture medium-scale features such as texture and shape; 3×1 and 1×3 convolutional kernels operate in the horizontal and vertical directions respectively, enabling more refined capture of edge and directional features. Especially in underwater image enhancement tasks, the degree of attenuation of degraded images varies between different local regions. This diversity of convolutional kernel sizes allows the network to more flexibly adapt to image details of different scales and orientations, helping to improve the model's ability to capture subtle image features. For input... A 1×1 convolution kernel is used to transform the features and expand the channels to obtain the features.
[0082]
[0083] Among them, F 1×1 The features are obtained using a 1×1 convolution kernel, * represents a two-dimensional convolution operator, X i It is the i-th channel of the input feature matrix, δ(·) is the ReLU activation function, and W 4C,C,1,1 It is a single weight within a 1×1 convolution kernel, used to expand the number of channels of the input X from C to 4C.
[0084] Then, three parallel layers are used, employing 1×3, 3×3, and 3×1 convolutional kernels to perform convolution operations on the same input in different directions and sizes, respectively. Batch normalization is applied after each layer. Finally, the outputs of the three layers are directly added together to obtain the detailed features of the image.
[0085] F d =δ(BN(F0*W) :,:,1,3)+BN(F0*W :,:,3,3 )+BN(F0*W :,:,3,1 ))
[0086] Among them, F d The detailed features are obtained using three types of convolution kernels, W :,:,1,3 W :,:,3,3 and W :,:,3,1 These represent convolutional kernels of sizes 1×3, 3×3, and 3×1, respectively, with the output having the same number of input feature channels.
[0087] Finally, a 1×1 convolution operation is performed to align the channel dimensions with the output feature map of the global feature extraction module, resulting in the final output feature map.
[0088] F out =F d *W 4C,C,1,1
[0089] Among them, F out It is the output feature obtained by using a 1×1 convolution operation.
[0090] Furthermore, this invention designs three loss functions to collaboratively guide network training. The adaptive detail-weighted loss uses Softmax normalization to dynamically allocate loss weights based on the magnitude of the loss between different pixels, increasing the penalty for outliers and minimizing the error between the enhanced image and the reference image at the pixel level. The loss between them is defined as follows:
[0091]
[0092] in, and y i These are the enhanced image and the true ground image, α i It is the loss value after the i-th pixel, where ∈ is a small constant, empirically set to e. -6 .
[0093] The perceptual loss function, building upon SSIM, introduces the concept of multi-scale analysis. By decomposing and comparing images at different scales, it better considers the structural and content information of the image. The loss is defined as follows:
[0094]
[0095] Where N is the number of scales at which the image is decomposed. and y i These represent sub-images of the enhanced image and the reference image at different scales. This represents the SSIM value at each scale, β. i These are weight parameters at different scales.
[0096] Gradient difference loss evaluates the difference between the enhanced image and the reference image by comparing their gradient information. Minimizing gradient difference loss helps improve the model's understanding and preservation of image structure, edges, and details, and can better mitigate the blurring and distortion caused by the underwater environment. We express this loss as follows:
[0097]
[0098] in, and These represent the gradients of the enhanced image and the reference image at each pixel, respectively.
[0099] The final loss function assigns weights to the three loss functions mentioned above and combines them linearly as follows:
[0100]
[0101] in, This is the final loss function. α, γ, and β are all hyperparameters. After extensive experiments, they were set to 0.2, 0.3, and 0.5 respectively to balance the impact of different losses on the training process.
[0102] Example 2
[0103] This invention proposes a synchronous feature fusion underwater image enhancement system, such as... Figure 5 As shown, it includes:
[0104] The parameter acquisition module is used to acquire open datasets of real underwater scenes and design an adaptive detail-weighted loss.
[0105] The model building module is used to build an underwater image enhancement network model, which is trained based on an open dataset of real underwater scenes and adaptive detail-weighted loss.
[0106] The data processing module is used to input distorted images into a trained underwater image enhancement network model and output enhanced underwater images to achieve underwater image enhancement.
[0107] Example 3
[0108] Please see Figure 6 As shown, the present invention also provides an electronic device 100 for a synchronous feature fusion underwater image enhancement method; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0109] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the underwater image enhancement method with synchronous feature fusion described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0110] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.
[0111] The memory 101 in the electronic device 100 stores multiple instructions to implement a synchronous feature fusion underwater image enhancement method, and the processor 102 can execute the multiple instructions to achieve the following:
[0112] Obtain an open dataset of real-world underwater scenarios and design an adaptive detail-weighted loss;
[0113] An underwater image enhancement network model was established and trained based on an open dataset of real underwater scenes and adaptive detail-weighted loss.
[0114] The distorted image is input into the trained underwater image enhancement network model, and the enhanced underwater image is output, thus achieving underwater image enhancement.
[0115] Example 4
[0116] If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0117] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0118] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0119] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.
[0120] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0121] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A synchronous feature fusion underwater image enhancement method, characterized in that, Includes the following steps: Obtain an open dataset of real-world underwater scenarios and design an adaptive detail-weighted loss; An underwater image enhancement network model was established and trained based on an open dataset of real underwater scenes and adaptive detail-weighted loss. The distorted image is input into the trained underwater image enhancement network model, and the enhanced underwater image is output to achieve underwater image enhancement; the underwater image enhancement network model specifically includes a global feature extraction module and a detail feature extraction module. The detailed feature extraction module is as follows: For input The first convolution result is obtained by using a 1×1 convolution kernel to perform feature transformation on the input and expand the channels. Based on the first convolution result, 1×3, 3×3 and 3×1 convolution kernels are used to perform convolution operations on the same input in different directions and sizes, and batch normalization is performed. The outputs of the three layers are directly added to obtain the detailed features of the image. The channel dimensions are aligned with the image's detail feature map through a 1×1 convolution operation to obtain the final detail features. The adaptive detail-weighted loss is specifically as follows: The adaptive detail-weighted loss function Perceptual loss function Gradient difference loss function in, and These are augmented images and true ground images. It is after the first i Each pixel loss value It is a constant e -6 ; It represents the number of scales at which the image is decomposed. This represents the SSIM value at each scale. These are weight parameters at different scales; and These represent the gradients of the enhanced image and the reference image at each pixel, respectively. , and These are all hyperparameters. For adaptive detail-weighted loss function, For the perceptual loss function, This is the gradient difference loss function.
2. The underwater image enhancement method with synchronous feature fusion according to claim 1, characterized in that, The global feature extraction module is as follows: enter , will input Divide into non-overlapping patch blocks Each patch is treated as a "token," and after layer normalization, it enters two convolutional layers to obtain the spatial information of the original features. The feature map is then projected onto the input. The characteristic representation of three times the number of channels Then, the obtained features The relative positional deviation is calculated using window-based multi-head self-attention; finally, the global feature map is obtained through layer normalization and MLP. ;in, It is a symbol for representing image data as a tensor, where H represents the number of rows, W represents the number of columns, and C represents the number of channels.
3. The underwater image enhancement method with synchronous feature fusion according to claim 2, characterized in that, The global feature extraction module consists of an even number of modules connected in series. Two consecutive global feature extraction modules use a multi-head self-attention structure with a regular window and a shifted window, respectively, and the output tensor dimension always remains unchanged from the input.
4. A synchronous feature fusion underwater image enhancement system, characterized in that, The underwater image enhancement method using synchronous feature fusion as described in claim 1 includes: The parameter acquisition module is used to acquire open datasets of real underwater scenes and design an adaptive detail-weighted loss. The model building module is used to build an underwater image enhancement network model, which is trained based on an open dataset of real underwater scenes and adaptive detail-weighted loss. The data processing module is used to input distorted images into a trained underwater image enhancement network model and output enhanced underwater images to achieve underwater image enhancement.
5. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the underwater image enhancement method with synchronous feature fusion as described in any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the underwater image enhancement method with synchronous feature fusion as described in any one of claims 1 to 3.