An underwater image enhancement model training method and enhancement method, device and medium

By employing underwater image enhancement model training methods and utilizing techniques such as cycle consistency loss and semantic awareness contrast modules, the problem that underwater image enhancement methods cannot improve the performance of machine vision tasks in complex environments is solved, and significant improvements are achieved in target detection and segmentation tasks in complex underwater environments.

CN119399576BActive Publication Date: 2025-12-26SHANGHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411421720.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-12-26
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods cannot effectively improve the performance of machine vision tasks, especially in complex and ever-changing underwater environments. They lack specific consideration for the needs of machine vision tasks, resulting in problems such as inconsistent data features, loss of key information, and mismatch between enhancement objectives.

Method used

An underwater image enhancement model training method is adopted, including an underwater content encoder, a distortion encoder, an aerial content encoder, a synthetic image generator, an enhanced image generator, a segmentation encoder-decoder network, and a semantic awareness contrast module. The model is trained by using cycle consistency loss, identity mapping loss, adversarial loss, contrast loss, and task loss. The deentanglement representation and semantic awareness contrast module are introduced to ensure information transmission and preservation of task-related features.

Benefits of technology

It improves the performance of underwater machine vision tasks, enhances the preservation of sufficient semantic information in images in complex environments, and improves the accuracy and generalization ability of object detection, semantic segmentation, and salient object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399576B_ABST
    Figure CN119399576B_ABST
Patent Text Reader

Abstract

The application relates to an underwater image enhancement model training method and an enhancement method, equipment and a medium, wherein the model comprises an underwater content encoder, an underwater distortion encoder, an aerial content encoder, a synthetic image generator, an enhanced image generator, a segmentation encoding decoding network, a segmentation decoder and a semantic perception contrast module; the model training method obtains relevant data for calculating model loss through training image preprocessing, forward cross conversion, backward cross conversion, semantic perception and double-task guiding steps, and the underwater image enhancement model is trained by using the model loss; the enhancement method uses the model trained by the training method to realize enhancement processing on the underwater image to be enhanced. Compared with the prior art, the application provides a model training method based on disentangled representation, so that the trained model can focus on enhancing features beneficial to machine vision tasks, thereby achieving the purpose of improving the performance of underwater machine vision tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of underwater optical image processing, and in particular to an underwater image enhancement model training method and enhancement method, device and medium. BACKGROUND

[0002] Due to the uniqueness of the underwater environment, underwater machine vision tasks are limited in practical application in the marine environment. First, factors such as light attenuation, scattering and dispersion in the underwater environment cause underwater images to often have problems such as color deviation, low contrast and blurring. This distortion makes the target boundary in the image unclear and the details lost, so that the machine vision algorithm is difficult to accurately identify the target or extract the features, such as target detection, tracking and identification algorithms. Second, factors such as water flow, particulate matter and biological tissue in the underwater environment may introduce noise and interference in the image, further reducing the accuracy and reliability of the machine vision task.

[0003] In recent years, a large number of underwater image enhancement methods have been proposed to solve the problem of underwater image distortion. These methods can be generally classified into three categories: pixel-based methods, physical model-based methods and deep learning-based methods. Pixel-based methods mainly improve the visual quality of images by modifying pixel values, but due to the inability to capture higher-level semantic information, the ability to recover details in complex scenes is limited. Physical model-based methods consider the physical properties of underwater optical transmission and scattering, but when the required model parameters are not accurately estimated, the reliability of the method is affected, which indirectly affects the development of underwater detection technology and other technologies. Deep learning-based methods enhance images by learning a large amount of underwater image data, however, it is difficult and expensive to obtain large-scale labeled underwater data. In addition, due to the complexity and variability of the underwater environment, deep learning models may exhibit poor generalization ability. However, these three types of methods mainly focus on optimizing human visual perception quality, lack of targeted consideration for the characteristics and needs of machine vision tasks, resulting in limited performance improvement of downstream machine vision tasks. The specific reasons are as follows:

[0004] Inconsistent data features: Underwater image enhancement algorithms usually rely on existing underwater data for design. However, there are large differences in data features in different underwater environments, lighting conditions and water quality. When the features of the training data are inconsistent with the target task data, the enhancement algorithm may not effectively adapt to the features of the target task data, and thus cannot improve the performance of the target task.

[0005] Loss of key information: In the process of underwater image enhancement, operations such as denoising and image smoothing may cause the loss of key information required by certain tasks, thereby affecting the accuracy of machine vision tasks.

[0006] Enhancement purpose mismatch: human visual driven enhancement methods aim to improve image quality and visibility, such as denoising, contrast enhancement and color correction. However, the purpose of machine vision tasks is not completely consistent with the purpose of human visual perception. If the enhancement method only focuses on image quality and ignores task-related features, the enhanced result may not directly provide substantial help for machine vision tasks.

[0007] As the exploration of underwater vision field is increasingly in-depth, in addition to improving the visual quality of underwater images, the influence of enhancement methods on target detection and other machine vision tasks is increasingly concerned. Some methods combine hybrid synthesis model, enhancement model and detection perception into a cycle-consistent adversarial network, and use two cycle-consistency paths to realize the conversion of unpaired images between underwater and aerial domains. The detection perception provides feedback information in the form of gradient to guide the enhancement model to generate patch-level images that are visually pleasing or conducive to detection. Some methods propose a bidirectional constraint closed-loop adversarial enhancement module to reduce the demand for paired data in an unsupervised manner and retain more information features. In addition, in order to make the enhanced image more realistic, a contrast learning strategy is adopted in the training stage. Finally, a task-aware feedback module is embedded in the enhancement process to incorporate the gradient information of the detector to guide the enhancement to develop in a direction conducive to detection. Although these methods improve the visual effect and target detection ability of underwater image enhancement to some extent, these methods mainly focus on assisting machine vision tasks through visual quality improvement, ignoring the close combination of task requirements and image enhancement.

[0008] Therefore, it is a problem to be solved to provide a method with strong generalization ability to meet the performance requirements of multiple tasks when facing complex and variable underwater environments, and a model suitable for the method. SUMMARY

[0009] The purpose of the present application is to provide a method for training an underwater image enhancement model and an enhancement method, a device and a medium to overcome the problem that general underwater image enhancement methods cannot effectively improve the performance of underwater machine vision tasks.

[0010] The purpose of the present application can be achieved by the following technical solutions:

[0011] According to a first aspect of the present application, a method for training an underwater image enhancement model is provided, wherein the underwater image enhancement model comprises an underwater content encoder, an underwater distortion encoder, an aerial content encoder, a synthetic image generator, an enhanced image generator, a segmentation encoding-decoding network, a segmentation decoder and a semantic perception contrast module.

[0012] The training method comprises:

[0013] Training image preprocessing: the underwater dataset and the aerial dataset are randomly combined to form an unpaired data set, and the unpaired data set is preprocessed to obtain an input image pair; the input image pair includes an underwater image and an aerial image;

[0014] Forward cross conversion: the underwater content encoder, the underwater distortion encoder, the aerial content encoder, the synthetic image generator, and the enhanced image generator perform forward cross conversion on the input image pair based on the disentangled representation, perform first feature extraction on the underwater image and the aerial image respectively, and generate an enhanced image and a synthetic image based on the first feature; the first feature includes underwater content features, underwater distortion features, and aerial content features;

[0015] Backward cross conversion: the underwater content encoder, the underwater distortion encoder, the aerial content encoder, the synthetic image generator, and the enhanced image generator perform backward cross conversion on the enhanced image and the synthetic image based on the disentangled representation to output a reconstructed image, the reconstructed image includes an aerial reconstructed image and an underwater reconstructed image;

[0016] Semantic perception: input the input image pair, the enhanced image, and the synthetic image into the semantic perception comparison module to extract second features, and form two triplets according to a preset rule; process the triplets to output a comparison loss;

[0017] Dual-task guidance: input the underwater content features into the segmentation decoder to output a feature-level result, and input the enhanced image into the segmentation encoder-decoder network to output an image-level result;

[0018] Obtain the model loss based on the input image, the reconstructed image, the first feature, the enhanced image, the synthetic image, the comparison loss, the image-level result, and the feature-level result;

[0019] Model training: use the model loss to train the underwater image enhancement model.

[0020] As a preferred technical solution, the network structures of the underwater content encoder, the aerial content encoder, and the underwater distortion encoder are the same, and a dense residual block is used instead of a normal convolution in the network structure; the network structure includes a convolution normalization ReLu activation layer and a plurality of dense residual blocks, and a single dense residual block is connected by a plurality of convolution ReLu activation layers in a channel splicing manner; the structures of the enhanced image generator and the synthetic image generator are the same, and both include a convolution normalization ReLu activation layer, a convolution Tanh activation layer, and a plurality of dense residual blocks.

[0021] As a preferred technical solution, the specific steps of the forward cross conversion include:

[0022] inputting the underwater image into an underwater content encoder to output underwater content features, and inputting the aerial image into an aerial content encoder to output aerial content features;

[0023] inputting the underwater content features into an enhanced image generator to output an enhanced image;

[0024] inputting the underwater image into an underwater distortion encoder, and undergoing at least four times of down-sampling in the encoding process to output a first down-sampling result and underwater distortion features;

[0025] adding the first down-sampling result and the aerial content features pixel by pixel, and inputting the pixel addition result, the underwater distortion features and the aerial content features into a composite image generator to output a composite image.

[0026] As a preferred technical solution, the specific steps of the backward cross conversion include:

[0027] inputting the composite image into the underwater content encoder to output composite content features, and inputting the enhanced image into the aerial content encoder to output enhanced content features;

[0028] inputting the composite content features into the enhanced image generator to output an aerial reconstructed image;

[0029] inputting the composite image into the underwater distortion encoder, and undergoing at least four times of down-sampling in the encoding process to output a second down-sampling result and composite distortion features;

[0030] adding the second down-sampling result and the enhanced content features pixel by pixel, and inputting the pixel addition result, the composite distortion features and the enhanced content features into the composite image generator to output an underwater reconstructed image.

[0031] As a preferred technical solution, the second features include enhanced image features, composite image features, underwater image features and aerial image features.

[0032] The preset rule refers to dividing the second features into enhanced triplets and composite triplets by using a contrast learning strategy, and specifically, the enhanced triplets take the enhanced image features as anchor points, the underwater features of the input image as positive samples, and the aerial features of the input image as negative samples; and the composite triplets take the composite image features as anchor points, the aerial features of the input image as positive samples, and the underwater features of the input image as negative samples.

[0033] As a preferred technical solution, the model loss is obtained by weighted summation of a cycle consistency loss, an identity mapping loss, an adversarial loss, a contrast loss and a task loss.

[0034] As a preferred technical solution, the expression of the cycle consistency loss is:

[0035]

[0036] wherein L cc represents a cycle consistency loss, represents a water reconstruction image, X represents a water image in an input image, represents an air reconstruction image, Y represents an air image in an input image;

[0037] The expression of the identity mapping loss is:

[0038]

[0039] wherein L idt represents an identity mapping loss, is a first distortion feature, is an air content feature, G syn is a synthetic image generator;

[0040] The expression of the adversarial loss is:

[0041] L adv =-logD(X)-log(1-D(X Y ))

[0042] wherein L adv represents an adversarial loss, X Y represents a synthetic image, and D(·) represents a discriminator;

[0043] The expression of the contrastive loss is:

[0044]

[0045] wherein L con represents a contrastive loss, G(Y X ) represents an enhanced image feature, G(X) represents a water image feature, G(Y) represents an air image feature, and G(X Y ) represents a synthetic image feature;

[0046] The expression of the task loss is:

[0047]

[0048] wherein L task represents a task loss, CE(·) represents a cross-entropy loss of a semantic segmentation task, is a task-level result, mask represents a semantic label, and N seg (X Y ) represents an image-level result.

[0049] According to a second aspect of the present application, there is provided an underwater image enhancement method, which inputs an underwater image to be enhanced into an underwater image enhancement model trained by the above method for enhancement processing, and specifically comprises the following steps:

[0050] The underwater image to be enhanced is input into the underwater content encoder to output underwater content features;

[0051] The enhanced image generator receives the underwater content features and outputs an enhanced underwater image.

[0052] According to a third aspect of the present application, there is provided an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method when executing the program.

[0053] According to a fourth aspect of the present application, there is provided a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the method.

[0054] Compared with the prior art, the present application has the following advantages:

[0055] 1) The present application trains the enhancement model from cycle consistency loss, identity mapping loss, adversarial loss, contrastive loss and task loss, respectively, so that the enhancement model is more suitable for machine vision tasks, wherein the cycle consistency loss ensures effective information transmission during enhancement and synthesis; the identity mapping loss improves the quality of the converted image and stabilizes the training process; the contrastive loss fully retains the task-related content structure information; and the adversarial loss and the task loss avoid information loss or distortion during the training process;

[0056] 2) The present application introduces a semantic perception contrast module, which effectively avoids the loss of semantic information during underwater image enhancement, especially in scenarios where common operations such as denoising and contrast enhancement may cause information loss. The module can retain key details related to the task and ensure that the enhanced image still has sufficient semantic information in complex underwater environments;

[0057] 3) The present application designs a task-guided branch for image and feature levels, respectively, so that the underwater content encoder extracts features friendly to machine vision tasks, and further optimizes the enhanced image generator, which is convenient for generating images beneficial to the task and improves the performance of underwater machine vision tasks;

[0058] 4)、The application introduces disentangled representation learning, clearly distinguishes the content features conducive to the task and the underwater distortion features not conducive to the task, so that the image enhancement can more effectively focus on the preservation and enhancement of task-related features, not only improves the interpretability and generalization ability of the enhancement effect of the enhancement model, but also ensures that the enhanced image can better serve subsequent machine vision tasks. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 The method framework of the application is shown in the figure;

[0060] Figure 2 The forward cross conversion design of the application is shown in the figure;

[0061] Figure 3 The semantic perception design of the application is shown in the figure;

[0062] Figure 4 The double-task guidance design of the application is shown in the figure;

[0063] Figure 5 The target detection task visualization result comparison of the application is shown in the figure;

[0064] Figure 6 The semantic perception visualization result comparison of the application is shown in the figure;

[0065] Figure 7 The saliency target detection task visualization result comparison of the application is shown in the figure. DETAILED DESCRIPTION

[0066] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the application.

[0067] In this application, the phrase "embodiment" means that the specific features, structures or characteristics described in conjunction with the embodiment can be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor is it an independent or alternative embodiment that is not mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in the application can be combined with other embodiments without conflict.

[0068] Unless otherwise defined, technical terms and scientific terms used in the present application shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. Unless otherwise defined, the terms "one" and "a" or "an" shall not be construed to mean "at least one" or "one or more". Unless otherwise defined, the terms "including," "comprising," "attached," "coupled," "connected," and the like are not limited to direct connections, but can include indirect connections unless otherwise contextually implied. The term "plurality" means two or more. The term "and / or" describes associated objects in association relationships, which means that there are three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. The terms "first", "second", "third", and the like only distinguish similar objects, and do not represent a specific order for the objects.

[0069] Embodiment 1:

[0070] The embodiment provides a method for training an underwater image enhancement model for a machine vision task. Due to the lack of real underwater paired data sets, the training method adopts an unsupervised enhancement method, which maps non-paired underwater real images and lossless aerial images by using cyclic consistent mapping. However, the ordinary mapping method may not accurately capture the intrinsic characteristics of the data, and lacks interpretability, making it difficult for the trained model to generalize to tasks with different data characteristics in complex underwater environments. To overcome this limitation, the training method introduces an entangled representation to learn the latent features and representations of underwater images. These latent representations can capture important information and semantic content in images, exhibiting higher-level abstraction and expression capabilities. To address the problem of loss of key semantic information during the enhancement process, the training method proposes a semantic-aware feature extractor to extract high-resolution semantic-aware content features that are independent of underwater distortion. By introducing a contrastive loss during training, the network is encouraged to retain as much important detail as possible that is relevant to the task. To more closely integrate the relationship between enhancement and machine vision tasks, the training method also introduces a task-guided branch from both the feature and image perspectives. At the feature level, it encourages the underwater content encoder to extract more features that are beneficial to the machine vision task. At the image level, it further guides the enhancer to generate task-friendly images.

[0071] The above-mentioned underwater image enhancement model includes an underwater content encoder, an underwater distortion encoder, an aerial content encoder, a synthetic image generator, an enhanced image generator, a segmentation encoding-decoding network, a segmentation decoder, and a semantic-aware contrast module, and specific network structures are referred to Figure 1 .

[0072] The training method includes:

[0073] S1, image preprocessing:

[0074] S11, the underwater data set SUIM includes 1634 underwater images with semantic labels, and 1634 images with a resolution of 512*512 are randomly selected and cropped from the corresponding high-resolution data set Flickr2K in the air;

[0075] S12, randomly combine the underwater data set with the space data set to form a non-paired data set, and adjust the size to 256*256 to obtain an input image pair, wherein the input image pair includes an underwater image and an aerial image, and the input image pair is used as the input of the enhanced image.

[0076] S2, forward cross conversion:

[0077] Because the data sets in different underwater environments usually exhibit different data characteristics, it is difficult for general underwater image enhancement algorithms to effectively generalize to different underwater scenes. In order to improve the explainability and generalization ability of the network, the present application introduces disentangled representation learning. Specifically, the network details of the forward cross conversion are as shown in Figure 2 For underwater images, a underwater content encoder and a underwater distortion encoder are used to separate task-friendly content features and task-unfriendly distortion features. The underwater content encoder is responsible for extracting underwater content features, and the underwater distortion encoder is used to extract underwater distortion features; for aerial images, since there is no underwater distortion, only an aerial content encoder is needed to extract its content features, and the specific steps include:

[0078] S21, input the underwater image into the underwater content encoder to output the underwater content feature, and input the aerial image into the aerial content encoder to output the aerial content feature;

[0079] S22, input the underwater content feature into the enhanced image generator G enh to output the enhanced image, and its expression is:

[0080]

[0081] Y X Y represents an enhanced image, and X represents an underwater image;

[0082] S23, input the underwater image into the underwater distortion encoder and undergo at least four times of down-sampling in the encoding process, output the first down-sampling result and the underwater distortion feature;

[0083] S24, add the first down-sampling result to the aerial content feature pixel, input the pixel addition result, the underwater distortion feature and the aerial content feature into the synthetic image generator G syn , output the synthetic image, whose expression is:

[0084]

[0085] X Y Y represents a synthetic image, and X represents an aerial image.

[0086] Moreover, the network structures of the underwater content encoder, the aerial content encoder and the underwater distortion encoder are the same, and a U-Net-based encoder-decoder structure is used to replace the ordinary convolution in the network structure, which includes a convolution normalization ReLu activation layer and a plurality of dense residual blocks, and a single dense residual block is connected by a plurality of convolution ReLu activation layers in a channel splicing manner; the structures of the enhanced image generator and the synthetic image generator are the same, and both include a convolution normalization ReLu activation layer, a convolution Tanh activation layer and a plurality of dense residual blocks.

[0087] S3, backward cross conversion:

[0088] The purpose of the forward cross conversion is to generate an enhanced image Y X and a synthetic image X Y , and the purpose of the backward cross conversion is to use the converted result together with the original image to provide pixel-level constraints for network training, and the specific steps include:

[0089] S31, input the synthetic image into the underwater content encoder to output a synthetic content feature, and input the enhanced image into the aerial content encoder to output an enhanced content feature;

[0090] S32, input the synthetic content feature into the enhanced image generator to output an aerial reconstruction image, whose expression is:

[0091]

[0092] X Y represents a synthetic image;

[0093] S33, input the synthesized image into the underwater distortion encoder, and at least four times of down-sampling are experienced in the encoding process, output the second down-sampling result and the synthesized distortion feature;

[0094] S34, add the second down-sampling result and the enhanced content feature pixel, input the pixel addition result, the synthesized distortion feature and the enhanced content feature into the synthesized image generator, and output the underwater reconstruction image, the expression of which is:

[0095]

[0096] Y represents the underwater reconstruction image, Y X represents the enhanced image.

[0097] S4, semantic perception:

[0098] Underwater image enhancement involves a series of processing methods such as noise removal, image smoothing and contrast enhancement, etc., however, these processing methods may introduce some errors or change the characteristics of the original image when performing local repair; although these methods can improve the visual quality of underwater images, some important details related to machine vision tasks may be lost; in order to minimize the loss of details, the idea of contrast learning is introduced into the proposed model during the training of the enhancement model; accordingly, the application proposes a semantic perception contrast module, which guides the semantic perception feature extractor to extract task-related semantic information through contrast loss, ensuring that the enhanced underwater image not only reduces the distortion caused by underwater conditions, but also maintains the integrity of the original structural details, and the structure of the semantic perception contrast module is as shown in Figure 3 The specific steps are:

[0099] S41, using the semantic perception feature extractor G(·) to extract the features of the underwater image X, the aerial image Y, the enhanced image Y X and the synthesized image X Y , the above features maintain the same resolution as the input image to ensure that rich details are retained; wherein, the enhanced image feature G(Y X ), the underwater image feature G(X), the aerial image feature G(Y) and the synthesized image feature G(X Y ) are collectively referred to as the second feature;

[0100] S42, using a contrast learning strategy to divide the second feature into an enhanced triplet and a synthesized triplet, which is specifically: the enhanced triplet takes the enhanced image feature as the anchor point, the underwater feature of the input image as the positive sample, and the aerial feature of the input image as the negative sample;

[0101] The synthesized triplet takes the synthesized image feature as the anchor point, the aerial feature of the input image as the positive sample, and the underwater feature of the input image as the negative sample.

[0102] S43, the anchor points are close to the positive samples with the same content and far away from the negative samples with different content, ensuring that the enhancement process does not lose the semantic detail information in the original image that is potentially beneficial to the machine vision task.

[0103] S5, double-task guidance:

[0104] In order to establish a closer relationship between enhancement and machine vision tasks, the present application sets up a task guidance branch at the feature level and the image level respectively; at the feature level, the underwater content encoder is made more proficient in capturing the key information required for the subsequent machine vision task by combining the task-specific constraint mechanism, and the task-favorable features are extracted; at the image level, the enhanced image generator generates a three-channel image that is favorable to the task, and the detailed structure of the double-task guidance branch is as shown in Figure 4 , which specifically includes the following steps:

[0105] S51, input the underwater content features into the segmentation decoder to output the feature level results, and input the enhanced image into the segmentation encoder-decoder network to output the image level results;

[0106] S6, obtain the model loss:

[0107] S61, calculate the cycle consistency loss: since the proposed unsupervised framework contains two cycles, in order to ensure the effective transmission of information in the enhancement and synthesis process, the present application introduces the cycle consistency loss, which serves as a kind of supervision signal to guide the enhancement and synthesis network to retain the content of the input image, and the cycle consistency loss expression is:

[0108]

[0109] wherein, L cc represents the cycle consistency loss, represents the underwater reconstructed image, X represents the underwater image in the input image, represents the aerial reconstructed image, Y represents the aerial image in the input image;

[0110] S62, calculate the identity mapping loss: the present training method proposes to apply the identity mapping loss to improve the quality of the converted image and stabilize the training process, and the expression is:

[0111]

[0112] wherein, L idt represents the identity mapping loss, is the first distorted feature, is the aerial content feature, G syn is the synthesized image generator;

[0113] S63, calculate the adversarial loss: for distinguishing real underwater images X and synthetic underwater images X Y , the expression is:

[0114] L adv = -logD(X) - log(1-D(X Y ))

[0115] where L adv adversarial loss, X Y indicates the synthetic image, D(·) indicates the discriminator;

[0116] S64, calculate the contrast loss: in order to fully retain the content structure information related to the task, the semantic perception feature is extracted using the feature extractor G(·), and the L1 regularization is used to strengthen the contrast prior, the expression is:

[0117]

[0118] where L con contrast loss, G(Y X ) indicates the enhanced image feature, G(X) indicates the underwater image feature, G(Y) indicates the aerial image feature, and G(X Y ) indicates the synthetic image feature;

[0119] S65, calculate the task loss: in the dual-task guided branch, the underwater semantic segmentation task is trained jointly with the enhancement network, so the task loss is composed of two parts, which are constrained from the feature level and the image level respectively, the expression is:

[0120]

[0121] where L task task loss, CE(·) indicates the cross-entropy loss of the semantic segmentation task, is the task level result, mask indicates the semantic label, N seg (X Y ) indicates the image level result.

[0122] S66, calculate the model loss: for the joint training process of the enhancement model, the model loss is obtained by weighted sum of the cycle consistency loss, the identity mapping loss, the adversarial loss, the contrast loss and the task loss, the expression is:

[0123] L total = λ1L cc + λ2L idt + λ3L adv + λ4L con + λ5L task ,

[0124] wherein, L total denotes the total loss, λ1=10, λ2=5, λ3=1, λ4=5 and λ5=0.2 are hyperparameters.

[0125] S7, model training: training the underwater image enhancement model using the model loss.

[0126] The embodiment also provides an underwater image enhancement method, which uses the underwater image enhancement model trained by the training method to achieve the purpose of enhancing the to-be-enhanced underwater image, and specifically includes the following steps:

[0127] A1: inputting the to-be-enhanced underwater image into the underwater content encoder to output underwater content features;

[0128] A2: the enhanced image generator receives the underwater content features and outputs the enhanced underwater image.

[0129] Embodiment 2:

[0130] In order to verify that the model trained by the underwater image enhancement model training method provided in the above embodiment has significant advantages, the trained enhancement model is evaluated on machine vision tasks in this embodiment, including underwater target detection, underwater semantic segmentation and underwater saliency target detection; some advanced underwater image enhancement methods are compared, the Adam optimizer used by the enhancement model is set to 200 cycles of end-to-end training, and the learning rate is fixed at 2×e -4 -4 during the first 100 cycles, and gradually decays to 0 during the last 100 cycles.

[0131] For underwater target detection:

[0132] The ChinaMM dataset is used for verification, which has a total of 2071 training images and 676 test images, contains three types of objects: sea cucumber, sea urchin and scallop, and uses mAP@0.5:0.95% as an index to evaluate the target detection accuracy. In order to study the influence of underwater image enhancement on the target detection task, the enhancement results of different models are applied to various detection algorithms in this embodiment, including two one-stage detection networks RetinaNet and SSD, one two-stage detection network Faster-RCNN, and one network FERNet specially designed for underwater target detection. The comparison results of the detection performance of the four networks are shown in Table 1.

[0133] The experimental results show that the model significantly improves the mAP of the four detectors, and compared with the original image, the mAP is increased by 4.6%, 5.7%, 3.8% and 1.6% respectively, while the detection performance of other models decreases in all four detection networks, which shows that the enhanced model trained by the underwater image enhancement model training method provided by the application has better performance in the underwater target detection task, because the features beneficial to the task are focused on during the training process. Figure 5 The detection visualization results of different enhancement algorithms on the RetinaNet network are given; it can be seen from the figure that not only the detection effect of the difficult area is excellent, but also the detected target shows higher confidence.

[0134] Table 1 Comparison of underwater target detection performance

[0135]

[0136] For underwater semantic segmentation:

[0137] The performance in mIoU and mPA is evaluated on two commonly used segmentation network models DeepLabv3+ and U-Net, and Table 2 gives the segmentation performance comparison of different methods on DeepLabv3+ and U-Net networks. It can be seen that the segmentation performance of the model is excellent on the two networks, and compared with the original image, the mIoU is increased by 2.8% and 2.6% respectively on the two segmentation networks.

[0138] Figure 6 The segmentation visualization results on the SUIM test set are given, and it can be seen from the figure that the segmentation results obtained by the model are most similar to the segmentation label, and the classification of the pixels in the blurred background area also has good performance, which shows that the enhanced model trained by the underwater image enhancement model training method provided by the application can effectively remove the interference of underwater distortion on the image semantics.

[0139] Table 2 Comparison of underwater semantic segmentation performance

[0140]

[0141] For underwater salient target detection:

[0142] The performance comparison is performed on the U2NET network, and Table 3 gives quantitative comparison results of three evaluation indexes on the UFO-120 and USOD10K data sets. As can be seen from the table, the present model is always superior to other methods in all three indexes, and the performance improvement of most comparison methods relative to the original image is relatively limited, and even decreases. And on the UFO-120 data set, compared with the original image, the F-measure of the present model is improved by 3.4%, the E-measure is improved by 2.0%, and the MAE is reduced by 0.008. On the USOD10K data set, the F-measure is improved by 2.5%, the E-measure is improved by 1.7%, and the MAE is reduced by 0.006; it can be seen that the introduced disentangled representation separates the underwater content and distortion, and even if only the segmentation task loss is introduced to optimize the enhancement network in the training process, it still has great advantages in the target detection and salient target detection tasks, effectively verifying that the present method has strong generalization ability.

[0143] The salient target detection visualization results on the UFO-120 data set are as shown in Figure 7 As can be seen from the figure, the UDCP and Fusion algorithms incorrectly identify the background as a salient object, and in contrast, the enhanced model trained by the underwater image enhancement model training method provided by the present application shows more accurate edge detection, and is even subjectively superior to the salient label.

[0144] Table 3 Comparison of underwater salient target detection performance

[0145]

[0146] In order to further verify the effectiveness of the present model in multiple machine vision tasks, three key factors are considered in the present embodiment, including: disentangled representation, semantic perception contrast module and double-task guided branch.

[0147] The benchmark truth is the CycleGAN network without adding any module, for the target detection task, the disentangled representation is added to the benchmark truth, on the four detection networks, the mAP is improved by 1.6% to 4.7%; the addition of the double-task guided branch further improves the performance by 3.1% to 7.9%; the addition of the semantic-aware contrast module improves the detection performance by 2.8% to 7.2%. For the semantic segmentation task, compared with the benchmark truth, the addition of the disentangled representation improves the mIoU by 4.8% and 1.4% on the DeepLabv3+ and U-Net networks respectively; the addition of the double-task guided branch further improves the mIoU by 2.0% and 2.4% respectively; the addition of the semantic-aware contrast module improves the mIoU by 1.1% and 0.5% respectively. For the salient object detection task, compared with the benchmark truth, the model after introducing the disentangled representation improves the F-measure by 2.1% and 0.9% respectively on two datasets; the F-measure is improved by 3.1% and 0.6% respectively after adding the double-task guided branch; the addition of the semantic-aware contrast module improves the F-measure by 4.0% and 0.3% respectively.

[0148] In summary, the training method of the underwater image enhancement model provided in the above embodiments has significant advantages in various different underwater machine vision tasks, effectively improving the performance of target detection, semantic segmentation and salient object detection tasks in complex underwater environments.

[0149] Embodiment 3

[0150] The electronic device provided in this embodiment includes a central processing unit (CPU) which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or computer program instructions loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for device operation can also be stored. The CPU, ROM and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.

[0151] A plurality of components in the device are connected to the I / O interface, including: an input unit such as a keyboard, a mouse, etc.; an output unit such as various types of displays, a loudspeaker, etc.; a storage unit such as a magnetic disk, an optical disk, etc.; and a communication unit such as a network card, a modem, a wireless communication transceiver, etc. The communication unit allows the device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0152] The processing units perform the various methods and processes described above, such as methods S1-S7 and A1-A2. For example, in some embodiments, methods S1-S7 and A1-A2 can be implemented as a computer software program tangibly embodied in a machine readable medium, such as a storage unit. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the device via the ROM and / or the communication unit. When the computer program is loaded onto the RAM and executed by the CPU, one or more of the steps of methods S1-S7 and A1-A2 described above can be performed. Alternatively, in other embodiments, the CPU can be configured to perform methods S1-S7 and A1-A2 by way of other suitable means, such as by way of firmware.

[0153] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0154] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, causes the machine to perform the functions / acts specified in the flow diagrams and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0155] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include one or more lines of a system, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] The above merely illustrates the specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements shall be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. An underwater image enhancement model training method, characterized in that, The underwater image enhancement model comprises an underwater content encoder, an underwater distortion encoder, an aerial content encoder, a synthetic image generator, an enhanced image generator, a segmentation encoding-decoding network, a segmentation decoder, and a semantic perception comparison module. The training method comprises: Training image preprocessing: the underwater dataset and the aerial dataset are randomly combined to form an unpaired data group, and the unpaired data group is preprocessed to obtain an input image pair; the input image pair comprises an underwater image and an aerial image; Forward cross conversion: the underwater content encoder, the underwater distortion encoder, the aerial content encoder, the synthetic image generator, and the enhanced image generator are used to perform forward cross conversion on the input image pair based on the disentangled representation, first feature extraction is performed on the underwater image and the aerial image respectively, and an enhanced image and a synthetic image are generated based on the first feature; the first feature comprises underwater content feature, underwater distortion feature, and aerial content feature; Backward cross conversion: the underwater content encoder, the underwater distortion encoder, the aerial content encoder, the synthetic image generator, and the enhanced image generator are used to perform backward cross conversion on the enhanced image and the synthetic image based on the disentangled representation to output a reconstructed image; the reconstructed image comprises an aerial reconstructed image and an underwater reconstructed image; Semantic perception: the input image pair, the enhanced image, and the synthetic image are input into the semantic perception comparison module to extract second feature, the second feature is combined into two triplets according to a preset rule, and the triplets are processed to output a comparison loss; Dual-task guidance: the underwater content feature is input into the segmentation decoder to output a feature-level result, and the enhanced image is input into the segmentation encoding-decoding network to output an image-level result; Model loss acquisition: a model loss is acquired based on the input image, the reconstructed image, the first feature, the enhanced image, the synthetic image, the comparison loss, the image-level result, and the feature-level result; Model training: the underwater image enhancement model is trained using the model loss.

2. The method of claim 1, wherein, The network structures of the underwater content encoder, the aerial content encoder, and the underwater distortion encoder are the same, and a dense residual block is used instead of an ordinary convolution in the network structure; the network structure comprises a convolution normalization ReLu activation layer and a plurality of dense residual blocks; a single dense residual block is connected by a plurality of convolution ReLu activation layers in a channel splicing manner; the structures of the enhanced image generator and the synthetic image generator are the same, and each comprises a convolution normalization ReLu activation layer, a convolution Tanh activation layer, and a plurality of dense residual blocks.

3. The method of claim 1, wherein, The specific steps of the forward cross conversion comprise: The underwater image is input into the underwater content encoder to output underwater content feature, and the aerial image is input into the aerial content encoder to output aerial content feature; The underwater content feature is input into the enhanced image generator to output an enhanced image; The underwater image is input into the underwater distortion encoder, and at least four down-sampling processes are performed in the encoding process to output a first down-sampling result and underwater distortion feature; The first down-sampling result is added to the aerial content feature pixel, and the pixel addition result, the underwater distortion feature, and the aerial content feature are input into the synthetic image generator to output a synthetic image.

4. The method of claim 1, wherein, The specific steps of the backward cross conversion include: inputting the synthetic image into an underwater content encoder to output synthetic content features, and inputting the enhanced image into an aerial content encoder to output enhanced content features; inputting the synthetic content features into an enhanced image generator to output an aerial reconstructed image; inputting the synthetic image into an underwater distortion encoder, and undergoing at least four times of down-sampling in the encoding process to output a second down-sampling result and synthetic distortion features; adding the second down-sampling result and the enhanced content features pixel by pixel, and inputting the pixel addition result, the synthetic distortion features and the enhanced content features into a synthetic image generator to output an underwater reconstructed image.

5. The method of claim 1, wherein, The second features include enhanced image features, synthetic image features, underwater image features and aerial image features. The preset rule refers to dividing the second features into enhanced triplets and synthetic triplets by using a contrast learning strategy, and specifically, the enhanced triplets take the enhanced image features as anchor points, the underwater features of the input image as positive samples, and the aerial features of the input image as negative samples; and the synthetic triplets take the synthetic image features as anchor points, the aerial features of the input image as positive samples, and the underwater features of the input image as negative samples.

6. The method of claim 1, wherein, The model loss is obtained by weighted summation of a cycle consistency loss, an identity mapping loss, an adversarial loss, a contrast loss and a task loss.

7. The method of claim 6, wherein, The expression of the cycle consistency loss is: wherein L cc denotes the cycle-consistency loss, denotes the underwater reconstructed image, X denotes the underwater image in the input image, denotes the aerial reconstructed image, Y denotes the aerial image in the input image; The expression of the identity mapping loss is: wherein L idt denotes an identity mapping loss, is a first distortion feature, is an aerial content feature, G syn is a synthetic image generator; The expression of the adversarial loss is: L adv = -logD(X) - log(1 - D(X Y )) wherein L adv represents the adversarial loss, X Y represents the synthetic image, D(·) represents the discriminator; The expression of the contrast loss is: Wherein, L con represents the contrast loss, G(Y X ) represents the enhanced image feature, G(X) represents the underwater image feature, G(Y) represents the aerial image feature, and G(X Y ) represents the synthetic image feature. The expression of the task loss is: wherein L task represents the task loss, CE(·) represents the cross-entropy loss of the semantic segmentation task, is the task-level result, mask represents the semantic label, N seg (X y ) represents the image-level result.

8. An underwater image enhancement method characterized by, The underwater image enhancement method inputs a to-be-enhanced underwater image into an underwater image enhancement model trained by the training method of any one of claims 1-7 for enhancement processing, and specifically includes the following steps: inputting the to-be-enhanced underwater image into an underwater content encoder to output underwater content features; an enhanced image generator receives the underwater content features and outputs an enhanced underwater image.

9. An electronic device comprising a memory and a processor, said memory having stored thereon a computer program, characterized in that, The processor implements the method of any one of claims 1-7 when executing the program.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program implements the method of any one of claims 1-7 when executed by the processor.

Citation Information

Patent Citations

  • Image semantic segmentation model training method and system

    CN111832570A

  • Underwater image enhancement method based on brightness-mask-guided multi-attention mechanism

    WO2024208188A1