Self-supervised underwater image enhancement method and device, equipment and storage medium

By building an underwater image enhancement model including a global lighting estimation module, a transmission loss estimation network and a scene radiation estimation network, and using self-supervised learning and proxy tasks to train the model, the problem of requiring a large amount of reference data and lack of constraints in the prior art is solved, and efficient underwater image enhancement is achieved, removing noise, color shift and overexposure.

CN120198335APending Publication Date: 2025-06-24YUNYANG ZHIHAI IND TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223599.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art requires a large amount of reference data for model training in underwater image enhancement, and lacks constraints on agent tasks, so it is impossible to effectively remove the same subnoise effects brought by the feature extraction module, resulting in the inability to effectively remove the bad imaging effects such as noise, color shift and overexposure in underwater images.

Method used

Underwater image enhancement model including global lighting estimation module, transmission loss estimation network and scene radiation estimation network is constructed. The model is trained using agent tasks through self-supervised learning methods, including synthesizing the original image and enhanced image into mixed images, and feature extraction and loss calculation are performed through multi-layer convolutional neural networks, and iterative training is performed to optimize the model.

Benefits of technology

The training of the underwater image enhancement model is carried out without reference to the picture, which simplifies the processing flow, improves the processing efficiency, and effectively removes the influence of noise caused by light scattering in the underwater image, eliminates color shifts, and suppresses the overexposure and over-enhancement effects generated in image enhancement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198335A_ABST
    Figure CN120198335A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image enhancement, in particular to a self-supervised underwater image enhancement method and device, equipment and a storage medium. The self-supervised underwater image enhancement method comprises the following steps: constructing an underwater image enhancement model, and inputting an original image into the underwater image enhancement model to obtain a global illumination image, a first transmission loss image and a first enhanced image; synthesizing the first enhanced image and the original image to obtain a mixed image, and obtaining a second transmission loss image and a second enhanced image based on the mixed image; calculating combination loss based on the plurality of images, and obtaining an optimized underwater image enhancement model through iterative training; and obtaining a final enhanced image of the original image based on the optimized underwater image enhancement model. According to the method provided by the invention, the noise influence caused by light scattering in the underwater image can be effectively removed, the color cast is eliminated, and the overexposure and the overexposure effect generated in the image enhancement are inhibited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image enhancement, and particularly to a self-supervised underwater image enhancement method, device, equipment and storage medium. Background Art

[0002] With the increase of human underwater activities, underwater vision has become an important means for humans to explore the underwater world. However, due to the scattering and absorption of light, underwater images have serious color cast and degradation, resulting in poor imaging effects and affecting the development of other downstream activities, such as underwater target recognition. When traditional image processing methods are applied to underwater image enhancement, overcompensation effects often occur, such as overexposure and overenhancement, resulting in the loss of picture information. Existing deep learning-based underwater image enhancement methods usually require a large amount of reference data for model training, and at the same time lack constraints on proxy tasks, and cannot effectively remove the influence of the same sub-noise brought by the feature extraction module, and cannot meet the purpose of removing noise, color cast and overexposure and other bad imaging effects in underwater images. Summary of the Invention

[0003] Embodiments of the present application aim to provide a self-supervised underwater image enhancement method, device, equipment and storage medium, so as to solve the problems in the prior art that a large amount of reference data is required for model training, and there is a lack of constraints on proxy tasks, and the influence of the same sub-noise brought by the feature extraction module cannot be effectively removed.

[0004] To solve the above technical problems, the embodiments of the present application provide the following technical solutions:

[0005] According to a first aspect of the present application, there is provided a self-supervised underwater image enhancement method, the method comprising:

[0006] Construct an underwater image enhancement model, the underwater image enhancement model comprising a global illumination estimation module, a transmission loss estimation network and a scene radiance estimation network;

[0007] Input an original image into the underwater image enhancement model, obtain a global illumination image through the global illumination estimation module, obtain a first transmission loss image through the transmission loss estimation network, and obtain a first enhanced image through the scene radiance estimation network;

[0008] Synthesize the first enhanced image and the original image to obtain a mixed image as a proxy task, and input the mixed image into the transmission loss estimation network and the scene radiance estimation network respectively to obtain a second transmission loss image and a second enhanced image;

[0009] Calculate a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and the reconstructed image, and iteratively train the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model, where the reconstructed image is synthesized based on the first enhanced image, the first transmission loss image, and the global illumination image;

[0010] Based on the optimized underwater image enhancement model, obtain the final enhanced image of the original image.

[0011] Optionally, the calculating the combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, the preset maximum transmission loss image, the original image, and the reconstructed image includes:

[0012] Calculate a homology loss based on the similarity between the first enhanced image and the second enhanced image;

[0013] Calculate a transmission loss based on the distances between the first transmission loss image and the second transmission loss image and the preset maximum transmission loss image respectively;

[0014] Calculate a reconstruction loss based on the similarity between the original image and the reconstructed image;

[0015] Calculate the combined loss based on the homology loss, the transmission loss, and the reconstruction loss. Optionally, the formula for the homology loss is:

[0016] L h =‖J′ - stopgrad(J)‖2 2

[0017] where L h is the homology loss, J′ is the second enhanced image, J is the first enhanced image, and stopgrad(J) means to stop gradient transfer when calculating the homology loss L h ;

[0018] The formula for the reconstruction loss is:

[0019] L r =‖I rec - I‖2 2

[0020] where L r is the reconstruction loss, I rec is the reconstructed image, and I is the original image.

[0021] Optionally, the formula for the transmission loss is:

[0022]

[0023] Among them, L p is the transmission loss, Dist1 represents the distance between the first transmission loss image T and the maximum transmission loss image T m , Dist2 represents the distance between the first transmission loss image T and the second transmission loss image T′, and margin is a preset boundary value.

[0024] Optionally, the combined loss further includes a color balance loss, and the calculation formula of the color balance loss is:

[0025]

[0026] Among them, L G is the color balance loss, Ω represents the set of color channels, J c is the component of the first enhanced image J on the color channel c, and μ(J c ) represents the average gray value of the first enhanced image J on the color channel c.

[0027] Optionally, the calculation formula of the mixed image is:

[0028] I′ = cI + (1 - c)J

[0029] Among them, I′ is the mixed image, I is the original image, J is the first enhanced image, and c is a random number between (0, 1).

[0030] Optionally, the transmission loss estimation network and the scene radiation estimation network include a feature extraction module, a structure preservation module, and a spatial group cross-channel attention module composed of a multi-layer convolutional neural network.

[0031] Optionally, the spatial group cross-channel attention module groups the features of the input original feature map in the channel dimension to obtain multiple grouped feature maps. Based on each grouped feature map, channel attention is calculated using average pooling and max pooling to obtain two channel attention maps. After adding the two channel attention maps element-wise and passing through a one-dimensional convolutional operation, the channel attention of the grouped feature map is obtained. Then, the channel attention of the grouped feature map is multiplied element-wise with the grouped feature map to obtain the weighted feature map of the grouped feature map. The weighted feature maps of each group are fused to obtain the enhanced feature map of the original feature map.

[0032] According to the second aspect of the present application, a self-supervised underwater image enhancement device is provided, and the device includes:

[0033] A model construction unit for constructing an underwater image enhancement model, where the underwater image enhancement model includes a global illumination estimation module, a transmission loss estimation network, and a scene radiation estimation network;

[0034] An image estimation unit for inputting an original image into the underwater image enhancement model, obtaining a global illumination image through the global illumination estimation module, obtaining a first transmission loss image through the transmission loss estimation network, and obtaining a first enhanced image through the scene radiation estimation network;

[0035] A proxy task unit for synthesizing the first enhanced image and the original image to obtain a mixed image as a proxy task, and inputting the mixed image into the transmission loss estimation network and the scene radiation estimation network respectively to obtain a second transmission loss image and a second enhanced image;

[0036] A model training unit for calculating a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and a reconstructed image, and iteratively training the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model, where the reconstructed image is synthesized based on the first enhanced image, the first transmission loss image, and the global illumination image;

[0037] A result output unit for obtaining a final enhanced image of the original image based on the optimized underwater image enhancement model.

[0038] According to a third aspect of the present application, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the self-supervised underwater image enhancement method described above.

[0039] According to a fourth aspect of the present application, there is provided a computer storage medium storing instructions or programs, and when the instructions or programs are executed by at least one processor, the at least one processor is enabled to execute the self-supervised underwater image enhancement method described above.

[0040] The beneficial effects of the embodiments of the present application are as follows: Different from the prior art, in the embodiments of the present application, a self-supervised underwater image enhancement method is provided, and an underwater image enhancement model including a global illumination estimation module, a transmission loss estimation network, and a scene radiance estimation network is constructed; the original image is input into the underwater image enhancement model to obtain a global illumination image, a first transmission loss image, and a first enhanced image; the first enhanced image and the original image are synthesized to obtain a mixed image as a proxy task, and the mixed image is input into the transmission loss estimation network and the scene radiance estimation network to obtain a second transmission loss image and a second enhanced image; a combined loss is calculated based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and the reconstructed image, and the underwater image enhancement model is iteratively trained based on the combined loss to obtain an optimized underwater image enhancement model; based on the optimized underwater image enhancement model, the final enhanced image of the original image is obtained. The method of the present application does not require reference pictures, trains the underwater image enhancement model in a step-by-step manner based on the proxy task, and constrains the proxy task from multiple aspects, which not only simplifies the processing flow of underwater image enhancement and improves the processing efficiency of underwater image enhancement, but also can effectively remove the noise influence caused by light scattering in underwater images, eliminate color cast, and suppress overexposure and over-enhancement effects generated in image enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.

[0042] Figure 1 It is a schematic framework diagram of a self-supervised underwater image enhancement method provided by an embodiment of the present application;

[0043] Figure 2 It is a network structure diagram of T-Net and J-Net provided by an embodiment of the present application;

[0044] Figure 3 It is a structure diagram of a structure preservation module provided by an embodiment of the present application;

[0045] Figure 4 It is a structure diagram of a spatial group cross-channel attention module provided by an embodiment of the present application;

[0046] Figure 5 It is a schematic flowchart of a self-supervised underwater image enhancement method provided by an embodiment of the present application;

[0047] Figure 6It is a schematic structural diagram of a self-supervised underwater image enhancement device provided by an embodiment of the present application;

[0048] Figure 7 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0050] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0051] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0052] Please refer to Figure 1 , Figure 1 It is a framework schematic diagram of a self-supervised underwater image enhancement method provided by an embodiment of the present application. As Figure 1 shown, the self-supervised underwater image enhancement method of the present application trains the underwater image enhancement model in a step-by-step manner. Among them, the underwater image enhancement model is designed based on the Koschmieder light scattering model and includes a global illumination estimation module GB, a transmission loss estimation network T-Net, and a scene radiance estimation network J-Net. After the input image is input into the underwater image enhancement model, the three modules independently calculate the enhancement distribution of the input image.

[0053] Specifically, the global illumination estimation module GB uses Gaussian blur to estimate the global illumination of the picture, and the calculation formula is as follows:

[0054]

[0055] Among them, k is the radius of the Gaussian kernel, I(x-i,y-j) is the pixel value of the image I, and G(i,j) is the value of the Gaussian kernel at (x,y).

[0056] Furthermore, the calculation formula of the Gaussian kernel function is:

[0057]

[0058] Among them, σ is the standard deviation controlling the degree of fuzziness, and exp is the exponential function.

[0059] In one embodiment, the Gaussian kernel size is where w and h respectively represent the width and height (pixel values) of the image, representing rounding up.

[0060] The transmission loss estimation network T-Net is used to estimate the transmission loss of the original image. The first transmission loss image obtained after the original image passes through the transmission loss estimation network has the same size as the original image, and each pixel corresponds to a transmission loss coefficient.

[0061] The scene radiation estimation network J-Net is used to output a preliminary estimated enhanced image corresponding to the low-quality original image.

[0062] In one embodiment, the transmission loss estimation network T-Net and the scene radiation estimation network J-Net have the same network architecture.

[0063] Please refer to Figure 2 , Figure 2 which is the network structure diagram of T-Net and J-Net provided by the embodiments of the present application. As Figure 2 shown, T-Net and J-Net include a feature extraction module composed of a multi-layer convolutional neural network, a structure-preserving block (SPB), and a spatial group cross-channel attention module (SGCA).

[0064] Among them, the feature extraction module is composed of four sequentially arranged convolutional neural networks, and each sub-layer adopts the same structure, including a convolutional layer, a normalization layer, and a ReLu activation layer.

[0065] Please refer to Figure 3 , Figure 3 which is the structure diagram of the structure-preserving module provided by the embodiments of the present application. As Figure 3 shown, the structure-preserving module uses two asymmetric filters to construct cross-convolutions, extracts features from both the horizontal and vertical directions simultaneously, can better capture the structural information such as edges and textures of the image, and at the same time maintain the integrity of the structure.

[0066] Since the features between different sub-layers and channels usually have similar patterns, containing useful recovery information and interference from the same noise. When these sub-features are processed in the same way, they will receive overlapping interference from the same noise between sub-layers. To reduce this interference, the present application proposes a spatial group cross-channel attention mechanism.

[0067] Please refer to Figure 4 , Figure 4 which is the structural diagram of the spatial group cross-channel attention module provided by the embodiments of the present application. As Figure 4 shown, the working process of the spatial group cross-channel attention module is as follows:

[0068] (1) Channel grouping

[0069] Perform a grouping operation on the input original feature map F (F ∈ R C×H×W , where C represents the number of channels, and H, W represent the spatial dimensions) in the channel dimension to obtain n grouped feature maps, that is, F = {F1, ……, F k , ……, F n},

[0070] (2) Calculate channel attention

[0071] Based on each grouped feature map, use average pooling (Average Pooling, AP) and max pooling (MaxPooling, MP) to calculate the channel attention, obtain two channel attention maps, and perform an element-wise addition on the two channel attention maps and then obtain the channel attention of the grouped feature maps through a one-dimensional convolution operation.

[0072] Specifically, for each grouped feature map F k , the calculation formula for the channel attention can be expressed as:

[0073] F′ K = conv 1×1 (pool ave (F k ) + pool max (F k ))

[0074] where pool ave (·) and pool max (·) respectively represent average pooling and max pooling in the spatial dimension, and conv 1×1 represents a 1×1 point convolution. For the grouped feature map F k , the calculation formulas for average pooling and max pooling can be expressed as:

[0075]

[0076] where F k (i, j) represents the value of the pixel (i, j) in the grouped feature map F k .

[0077] (3) Calculate spatial attention

[0078] Based on channel attention, SGCA further introduces a spatial attention mechanism, multiplies the channel attention of the grouped feature maps element-wise with the grouped feature maps to obtain the weighted feature maps of the grouped feature maps.

[0079] (4) Feature fusion

[0080] After obtaining the weighted feature maps of the grouped feature maps based on the foregoing steps, connect the weighted feature maps of each group to restore the original channel dimension and obtain the enhanced feature maps of the original feature maps.

[0081] In this application, using SGCA to group channels can reduce the overlapping interference of the same noise between sub-layers, thereby improving the robustness of features. At the same time, the grouping operation enables each group of features to independently learn different attention patterns, enhancing the diversity of features. In addition, SGCA considers both channel attention and spatial attention, and can comprehensively capture important information in the feature maps.

[0082] Please refer to Figure 5 , Figure 5 which is a schematic flowchart of a self-supervised underwater image enhancement method provided by an embodiment of this application. The method includes:

[0083] Step S501, construct an underwater image enhancement model, where the underwater image enhancement model includes a global illumination estimation module, a transmission loss estimation network, and a scene radiance estimation network.

[0084] In one embodiment, this underwater image enhancement model is designed based on the Koschmieder light scattering model, and includes a global illumination estimation module, a transmission loss estimation network, and a scene radiance estimation network. For the detailed design methods of each module, please refer to the foregoing description.

[0085] Step S502, input the original image into the underwater image enhancement model, obtain the global illumination image through the global illumination estimation module, obtain the first transmission loss image through the transmission loss estimation network, and obtain the first enhanced image through the scene radiance estimation network.

[0086] Step S503, synthesize the first enhanced image and the original image to obtain a mixed image as a proxy task, and input the mixed image into the transmission loss estimation network and the scene radiance estimation network respectively to obtain the second transmission loss image and the second enhanced image.

[0087] Self-supervised learning generates pseudo-labels from unlabeled data by designing proxy tasks, so that the model training can be completed without referring to reference pictures. In the embodiment of this application, the first enhanced image and the original image are synthesized to obtain a mixed image as a proxy task, and its calculation formula is:

[0088] I′ = cI+(1 - c)J

[0089] Wherein, I′ is the mixed image, I is the original image, J is the first enhanced image, and c is a random number between (0, 1).

[0090] The mixed image is respectively input into a transmission loss estimation network and a scene radiation estimation network to obtain a second transmission loss image and a second enhanced image. It should be noted that the transmission loss estimation network and the scene radiation estimation network have the same parameters as those in step S502.

[0091] Step S504, calculate a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and the reconstructed image, and iteratively train the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model.

[0092] Wherein, the reconstructed image is synthesized based on the first enhanced image, the first transmission loss image, and the global illumination image, and its calculation formula is:

[0093] I rec = TJ+(1 - T)A

[0094] Wherein, I rec is the reconstructed image, T is the first transmission loss image, J is the first enhanced image, and A is the global illumination image.

[0095] In one embodiment, the combined loss includes a homology loss between the first enhanced image and the second enhanced image, a transmission loss of the first transmission loss image respectively with the second transmission loss image and the preset maximum transmission loss image, and a reconstruction loss between the original image and the reconstructed image.

[0096] Specifically, the homology loss is a mean square error loss, calculated based on the similarity between the first enhanced image and the second enhanced image, and its calculation formula is:

[0097] L h = ‖J′ - stopgrad(J)‖2 2

[0098] Wherein, L h is the homology loss, J′ is the second enhanced image, J is the first enhanced image, stopgrad(J) represents stopping gradient transfer when calculating the homology loss L h , and ‖·‖2 is the L2 norm, used to calculate the distance between two numerical values.

[0099] The reconstruction loss is also a mean square error loss, calculated based on the similarity between the original image and the reconstructed image, and its calculation formula is:

[0100] L r =‖I rec -I‖2 2

[0101] where L r is the reconstruction loss, I rec is the reconstructed image, and I is the original image.

[0102] The transmission loss is a triplet loss, calculated based on the distances between the first transmission loss image and the second transmission loss image and a preset maximum transmission loss image respectively. Its calculation formula is:

[0103]

[0104] where L p is the transmission loss, Dist1 represents the distance between the first transmission loss image T and the maximum transmission loss image T m , Dist2 represents the distance between the first transmission loss image T and the second transmission loss image T′, and margin is a preset boundary value. In one embodiment, all pixel values of the maximum transmission loss image T m are 1.

[0105] In one embodiment, the combined loss further includes a color balance loss, which is calculated based on the gray world assumption of the first enhanced image J. Its calculation formula is:

[0106]

[0107] where L G is the color balance loss, Ω represents the set of color channels, J c is the component of the first enhanced image J on the color channel c, and μ(J c ) represents the average gray value of the first enhanced image J on the color channel c.

[0108] In one embodiment, the combined loss L can be expressed as:

[0109] L = ω1L h + ω2L p + ω3L r + ω4L G

[0110] where ω1, ω2, ω3, and ω4 are weight factors.

[0111] Based on this combined loss, the underwater image enhancement model is iteratively trained to continuously optimize the parameters of the underwater image enhancement model, so that the combined loss meets the preset conditions, thereby obtaining an optimized underwater image enhancement model.

[0112] Step S505: Obtain the final enhanced image of the original image based on the optimized underwater image enhancement model.

[0113] Specifically, use the image output by the scene radiance estimation network of the optimized underwater image enhancement model for the original image as the final enhanced image of the original image.

[0114] The self-supervised underwater image enhancement method provided in this application constructs an underwater image enhancement model including a global illumination estimation module, a transmission loss estimation network, and a scene radiance estimation network; input the original image into the underwater image enhancement model to obtain a global illumination image, a first transmission loss image, and a first enhanced image; synthesize the first enhanced image and the original image to obtain a mixed image as a proxy task, input the mixed image into the transmission loss estimation network and the scene radiance estimation network to obtain a second transmission loss image and a second enhanced image; calculate a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and the reconstructed image, and iteratively train the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model; based on the optimized underwater image enhancement model, obtain the final enhanced image of the original image. The method of this application does not require reference pictures, trains the underwater image enhancement model in a step-by-step manner based on a proxy task, and constrains the proxy task from multiple aspects, which not only simplifies the processing flow of underwater image enhancement and improves the processing efficiency of underwater image enhancement, but also can effectively remove the noise influence caused by light scattering in underwater images, eliminate color cast, and suppress overexposure and over-enhancement effects generated during image enhancement.

[0115] According to an embodiment of this application, a self-supervised underwater image enhancement device is provided. As Figure 6 shown, it is a schematic structural diagram of a self-supervised underwater image enhancement device provided by an embodiment of this application. The device includes:

[0116] A model construction unit 601, configured to construct an underwater image enhancement model, where the underwater image enhancement model includes a global illumination estimation module, a transmission loss estimation network, and a scene radiance estimation network;

[0117] An image estimation unit 602, configured to input the original image into the underwater image enhancement model, obtain a global illumination image through the global illumination estimation module, obtain a first transmission loss image through the transmission loss estimation network, and obtain a first enhanced image through the scene radiance estimation network;

[0118] The proxy task unit 603 is used to synthesize the first enhanced image and the original image to obtain a mixed image as a proxy task, and input the mixed image into the transmission loss estimation network and the scene radiation estimation network respectively to obtain a second transmission loss image and a second enhanced image;

[0119] The model training unit 604 is used to calculate a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and the reconstructed image, and iteratively train the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model, where the reconstructed image is synthesized based on the first enhanced image, the first transmission loss image, and the global illumination image;

[0120] The result output unit 605 is used to obtain a final enhanced image of the original image based on the optimized underwater image enhancement model.

[0121] In one embodiment, the specific steps for the model training unit 604 to calculate the combined loss are as follows: calculate the homology loss based on the similarity between the first enhanced image and the second enhanced image; calculate the transmission loss based on the distances between the first transmission loss image and the second transmission loss image and the preset maximum transmission loss image respectively; calculate the reconstruction loss based on the similarity between the original image and the reconstructed image; calculate the combined loss based on the homology loss, the transmission loss, and the reconstruction loss.

[0122] In one embodiment, the calculation formula for the homology loss is:

[0123] L h =‖J′ - stopgrad(J)‖2 2

[0124] where L h is the homology loss, J′ is the second enhanced image, J is the first enhanced image, and stopgrad(J) means to stop gradient transfer when calculating the homology loss L h ;

[0125] The calculation formula for the reconstruction loss is:

[0126] L r =‖I rec - I‖2 2

[0127] where L r is the reconstruction loss, I rec is the reconstructed image, and I is the original image.

[0128] In one embodiment, the calculation formula for the transmission loss is:

[0129]

[0130] Among them, L p is the transmission loss, Dist1 represents the distance between the first transmission loss image T and the maximum transmission loss image T m , and Dist2 represents the distance between the first transmission loss image T and the second transmission loss image T′, and margin is a preset boundary value.

[0131] In one embodiment, the combined loss further includes a color balance loss, and the calculation formula of the color balance loss is:

[0132]

[0133] Among them, L G is the color balance loss, Ω represents the set of color channels, J c is the component of the first enhanced image J on the color channel c, and μ(J c ) represents the average gray value of the first enhanced image J on the color channel c.

[0134] In one embodiment, the proxy task unit 603 synthesizes the first enhanced image and the original image based on the following formula to obtain a mixed image as the proxy task:

[0135] I′ = cI+(1 - c)J

[0136] Among them, I′ is the mixed image, I is the original image, J is the first enhanced image, and c is a random number between (0,1).

[0137] In one embodiment, the transmission loss estimation network and the scene radiation estimation network of the underwater image enhancement model constructed by the model construction unit 601 include a feature extraction module, a structure preservation module, and a spatial group cross-channel attention module composed of a multi-layer convolutional neural network.

[0138] In one embodiment, the spatial group cross-channel attention module groups the features of the input original feature map in the channel dimension to obtain a plurality of grouped feature maps. Based on each grouped feature map, channel attention is calculated using average pooling and max pooling to obtain two channel attention maps. After adding the two channel attention maps element by element and passing through a one-dimensional convolution operation to obtain the channel attention of the grouped feature maps, the channel attention of the grouped feature maps is multiplied by the grouped feature maps element by element to obtain the weighted feature maps of the grouped feature maps. The weighted feature maps of each group are fused to obtain the enhanced feature map of the original feature map.

[0139] According to the embodiments of the present application, an electronic device is provided, such as Figure 7As shown in the figure, it is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. The electronic device 100 includes a processor 10, a memory 20, and a communication interface 30. The processor 10, the memory 20, and the communication interface 30 are connected by lines. In Figure 7 In the shown embodiment, the processor 10, the memory 20, and the communication interface 30 are communicatively connected to each other through a bus.

[0140] The memory 20 is used to store software programs, computer-executable program instructions, etc. The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc.

[0141] The memory 20 may be a read-only memory (ROM), or may be other types of static storage devices that can store static information and instructions, or may be a random access memory (RAM), or may be other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM). Specifically, it is not limited here.

[0142] Exemplarily, the aforementioned memory 20 may be a double data rate synchronous dynamic random access memory DDR SDRAM (abbreviation: DDR). The memory 20 may exist independently, but is connected to the processor 10. Optionally, the memory 20 may also be integrated with the processor 10. For example, integrated within one or more chips.

[0143] In some embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10. These remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0144] The processor 10 uses various interfaces and lines to connect all parts of the entire electronic device 100. By running or executing the software program stored in the memory 20, and by calling the data stored in the memory 20, it executes various functions of the electronic device and processes data, such as implementing the method described in any embodiment of the present application.

[0145] The processor 10 can be a field programmable gate array (FPGA), a digital signal processor (DSP), a central processing unit (CPU), etc.

[0146] The processor 10 can be a single-core processor or a multi-core processor. For example, the processor 10 can be composed of multiple FPGAs or multiple DSPs. In addition, the processor 10 can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The processor 10 can be a separate semiconductor chip or integrated with other circuits into a semiconductor chip. For example, it can form a system on a chip (SoC) with other circuits (such as codec circuits, hardware acceleration circuits, or various bus and interface circuits), or it can also be integrated as a built-in processor of an application specific integrated circuit (ASIC) into the ASIC. The ASIC integrated with the processor can be separately packaged or packaged together with other circuits.

[0147] The communication interface 30 can use a transceiver device such as a transceiver to implement communication between the electronic device and other devices or communication networks.

[0148] The embodiments of the present application also provide a computer storage medium. The computer storage medium stores instructions or programs, and these instructions or programs are executed by one or more processors. For example Figure 7 one of the processors 10, can enable the above one or more processors to execute the self-supervised underwater image enhancement method in any of the above method embodiments.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions, or the part that contributes to the related technologies, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0150] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions. Therefore, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A self-supervised underwater image enhancement method, characterized in that: The method comprises: Constructing an underwater image enhancement model, wherein the underwater image enhancement model includes a global illumination estimation module, a transmission loss estimation network, and a scene radiation estimation network; Inputting the original image into the underwater image enhancement model, obtaining a global illumination image through the global illumination estimation module, obtaining a first transmission loss image through the transmission loss estimation network, and obtaining a first enhanced image through the scene radiation estimation network; The first enhanced image and the original image are synthesized to obtain a mixed image as a proxy task, and the mixed image is respectively input into the transmission loss estimation network and the scene radiation estimation network to obtain a second transmission loss image and a second enhanced image; Calculating a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and a reconstructed image, and iteratively training the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model, wherein the reconstructed image is synthesized based on the first enhanced image, the first transmission loss image, and the global illumination image; Based on the optimized underwater image enhancement model, a final enhanced image of the original image is obtained.

2. The method according to claim 1, characterized in that The calculation based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, the preset maximum transmission loss image, the original image and the reconstructed image comprises: Calculating homology loss based on the similarity between the first enhanced image and the second enhanced image; Calculating transmission loss based on the distances between the first transmission loss image and the second transmission loss image and a preset maximum transmission loss image; Calculating a reconstruction loss based on the similarity between the original image and the reconstructed image; A combined loss is calculated based on the homology loss, the transmission loss, and the reconstruction loss.

3. The method according to claim 2, characterized in that The calculation formula of the homology loss is: THE h =||J′-stopgrad(J)||2 2 Among them, L h is the homology loss, J′ is the second enhanced image, J is the first enhanced image, stopgrad(J) represents the calculation of the homology loss L h Stop gradient propagation when The calculation formula of the reconstruction loss is: L r =‖I rec -I‖2 2 Among them, L r is the reconstruction loss, I rec is the reconstructed image, and I is the original image.

4. The method according to claim 2, characterized in that: The calculation formula of the transmission loss is: Among them, L p is the transmission loss, Dist1 represents the first transmission loss image T and the maximum transmission loss image T m Dist2 represents the distance between the first transmission loss image T and the second transmission loss image T′, and margin is a preset boundary value.

5. The method according to claim 2, characterized in that: The combination loss also includes color balance loss, and the calculation formula of the color balance loss is: Among them, L G is the color balance loss, Ω represents the color channel set, and J c is the component of the first enhanced image J on the color channel c, μ(J c ) represents the average gray value of the first enhanced image J in color channel c.

6. The method according to claim 1, characterized in that The calculation formula of the mixed image is: I′=cI+(1-c)J Wherein, I′ is the mixed image, I is the original image, J is the first enhanced image, and C is a random number between (0, 1).

7. The method according to any one of claims 1 to 6, characterized in that: The transmission loss estimation network and the scene radiation estimation network include a feature extraction module composed of a multi-layer convolutional neural network, a structure preservation module and a spatial group cross-channel attention module.

8. The method according to claim 7, characterized in that The spatial group cross-channel attention module groups the features of the input original feature map on the channel dimension to obtain multiple grouped feature maps. Based on each grouped feature map, the channel attention is calculated using average pooling and maximum pooling to obtain two channel attention maps. The two channel attention maps are added element by element and then a one-dimensional convolution operation is performed to obtain the channel attention of the grouped feature map. The channel attention of the grouped feature map is then multiplied element by element with the grouped feature map to obtain a weighted feature map of the grouped feature map. The weighted feature maps of each group are fused to obtain an enhanced feature map of the original feature map.

9. A self-supervised underwater image enhancement device, characterized in that: The device comprises: A model building unit, used to build an underwater image enhancement model, wherein the underwater image enhancement model includes a global illumination estimation module, a transmission loss estimation network, and a scene radiation estimation network; An image estimation unit, configured to input an original image into the underwater image enhancement model, obtain a global illumination image through the global illumination estimation module, obtain a first transmission loss image through the transmission loss estimation network, and obtain a first enhanced image through the scene radiation estimation network; A proxy task unit, configured to synthesize the first enhanced image and the original image to obtain a mixed image as a proxy task, and input the mixed image into the transmission loss estimation network and the scene radiation estimation network respectively to obtain a second transmission loss image and a second enhanced image; a model training unit, configured to calculate a combined loss based on the first enhanced image, the second enhanced image, the first transmission loss image, the second transmission loss image, a preset maximum transmission loss image, the original image, and a reconstructed image, and iteratively train the underwater image enhancement model based on the combined loss to obtain an optimized underwater image enhancement model, wherein the reconstructed image is synthesized based on the first enhanced image, the first transmission loss image, and the global illumination image; The result output unit is used to obtain a final enhanced image of the original image based on the optimized underwater image enhancement model.

10. An electronic device, characterized in that: It includes at least one processor and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the self-supervised underwater image enhancement method as described in any one of claims 1 to 8.

11. A computer storage medium, characterized in that: The computer storage medium stores instructions or programs, which, when executed by at least one processor, enable the at least one processor to execute the self-supervised underwater image enhancement method as described in any one of claims 1 to 8.