Underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance
Through the methods of cross-domain feature representation and semantic segmentation guidance, the problem of insufficient dependence on paired data sets and semantic information in underwater image enhancement is solved, and efficient underwater image enhancement effect is achieved, improving image quality and detail recovery ability.
Patent Information
- Application Number
- CN202510419520.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
Existing underwater image enhancement methods rely on paired datasets to be difficult to obtain, and lack of utilization of image semantic information, resulting in limited detail recovery effect in complex underwater environments.
Using cross-domain feature representation and semantic segmentation guidance methods, a high-quality underwater image enhancement network is generated through paired data set training, combining contrast learning loss and semantic segmentation loss.
It realizes accurate enhancement and detail recovery of underwater images, improves color, contrast and detail quality, reduces data acquisition costs, and has strong generalization capabilities.
Smart Images

Figure CN120259111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of underwater image processing, and particularly relates to an underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance. Background Art
[0002] Due to the influence of scattering and absorption of light during the propagation process in water, underwater images often suffer from problems such as color distortion, reduced contrast, and blurred details. These problems not only affect the visual quality of underwater images but also pose challenges to the recognition and analysis capabilities of automated systems.
[0003] Currently, for underwater image enhancement based on deep learning, with the rapid development of deep learning, data-driven methods have made remarkable progress in the field of underwater image enhancement. However, these methods usually rely on paired degraded images and clear images for supervised training, and in practical applications, it is very difficult to obtain a large-scale and high-quality paired dataset, especially in the underwater environment.
[0004] In addition, most existing methods lack the utilization of image semantic information, which leads to a lack of pertinence in the processing of different semantic regions during the enhancement process and makes it difficult to comprehensively improve the structure and detail quality of the images.
[0005] Han Junlin et al. published "Underwater image restoration via contrastive learning and a real-world dataset" (Han, Junlin, et al. "Underwater image restoration via contrastive learning and a real-world dataset." Remote Sensing 14.17 (2022): 4297.). This method optimizes the feature distributions in the degraded image domain and the clear image domain through contrastive learning, improving the model's generalization ability and restoration effect. However, since this method mainly relies on the contrast of global features and does not combine semantic segmentation information for local optimization, it is difficult to precisely enhance the semantic regions in complex underwater scenes. Huang Shirui et al. published "Contrastive semi-supervised learning for underwater image restoration via reliable bank" (Huang, Shirui, et al. "Contrastive semi-supervised learning for underwater image restoration via reliable bank." Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2023.). This literature proposed an underwater image restoration method based on contrastive learning and semi-supervised learning, which optimizes the feature representation by constructing a reliable bank and uses the semi-supervised mechanism to improve the model performance during training. However, since this method mainly focuses on improving the global visual quality and does not fully utilize the high-level semantic information of the image, it lacks pertinence in enhancing different semantic regions, especially the detail restoration effect in complex underwater environments is limited. Summary of the Invention
[0006] To overcome the above-mentioned drawbacks of the prior art, the object of the present invention is to provide an underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance, which realizes the efficient enhancement of degraded underwater images by jointly combining the contrastive learning loss and the semantic segmentation loss.
[0007] The core of this method lies in using unpaired high-quality images and degraded images for training, avoiding the dependence on paired data in traditional methods. By designing the degraded image domain and the clear image domain, a mapping relationship of cross-domain feature distributions is constructed.
[0008] To achieve the above object, the technical solution adopted by the present invention is:
[0009] An underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance, comprising the following steps:
[0010] Step 1, construct an unpaired underwater image sample data set;
[0011] Step 2, design and initialize an image enhancement network;
[0012] Step 3, divide the data set constructed in Step 1 into positive samples and negative samples, randomly select an image I H from the positive samples, randomly select an image I L from the negative samples, input the image I L into the image enhancement network designed in Step 2 to obtain an enhanced image I E , and input {I L , I H , I E} into the image feature extraction model to calculate the contrastive learning loss;
[0013] Step 4, pre-train the semantic segmentation model on the semantic segmentation data set, and then fine-tune it on the semantic segmentation data set of underwater images. Input the image I H and the image I E obtained in Step 3 into the fine-tuned semantic segmentation model respectively to generate corresponding semantic prediction maps, and calculate the semantic segmentation loss;
[0014] Step 5, jointly train the image enhancement network designed in Step 2 with the contrastive learning loss calculated in Step 3 and the semantic segmentation loss calculated in Step 4, and update the parameters of the image enhancement network through end-to-end training to generate high-quality enhanced underwater images.
[0015] The data set constructed in Step 1 includes degraded underwater images and high-quality underwater images:
[0016] The degraded underwater images serve as the degraded image domain D L , and the high-quality underwater images and high-definition land images serve as the clear image domain D H .
[0017] The image enhancement network described in Step 2 adopts a U-Net structure. The U-Net structure includes an encoder and a decoder part. The input is a degraded underwater image, and the output is an enhanced image.
[0018] Step 3 specifically includes the following steps:
[0019] Step 3.1: Select a model F for image feature extraction and keep the parameters fixed;
[0020] Step 3.2: Use the degraded image domain D in the unpaired underwater image sample dataset L as negative samples, and the clear image domain D H as positive samples. Randomly sample the degraded underwater image I L and the clear image I H from the negative and positive samples respectively;
[0021] Step 3.3: Input the degraded underwater image I L into the image enhancement network to generate the enhanced image I E , and input the degraded underwater image I L , the enhanced image I E and the clear image I H into the model F with fixed parameters selected in Step 3.1 to extract its multi-layer features. Calculate the Gram matrix for the feature map of each layer and convert it into the latent feature vectors v U , v E , v H to represent the cross-domain feature distribution of the image in the latent feature space;
[0022] Gram(F k (I)) = F k (I) · F k (I) T
[0023] where F k (I) represents the feature map output by the k-th layer in the feature extraction model F. The Gram matrix is used to capture the relationship between features, and the extracted latent feature vectors are respectively denoted as the feature vector v U of the degraded image, the feature vector v E of the enhanced image, and the feature vector v H of the clear image;
[0024] Step 3.4: Construct a contrastive learning loss function. By minimizing the feature distance between the feature vector v E and the feature vector v H , and maximizing the feature distance between the feature vector v E and the feature vector v U , make the features of the enhanced image approach the features of the clear image and move away from the features of the degraded image;
[0025] Minimize the feature distance between the feature vector v E and the feature vector v HThe characteristic distance, the positive sample loss in the contrastive learning loss function is calculated as follows:
[0026]
[0027] Maximize the feature vector v E With the feature vector v U The negative sample loss in the contrastive learning loss function is calculated as follows:
[0028]
[0029] Where τ is the temperature parameter, used to adjust the gradient range of contrastive learning;
[0030] Combining the above positive and negative sample loss functions, the final contrastive learning loss function is defined as: L uns-cons = L pulI + L push .
[0031] The process of calculating the semantic segmentation loss in step 4 includes:
[0032] Step 4.1, select a model for semantic segmentation, pre-train it on the semantic segmentation dataset, and then fine-tune it on the semantic segmentation dataset of underwater images;
[0033] Step 4.2, use the semantic segmentation model fine-tuned in step 4.1 to segment the enhanced image I E And the high-quality image I H To generate an enhanced semantic prediction map S(I E ) and a high-quality semantic prediction map S(I H );
[0034] Step 4.3, utilize the semantic prediction map S(I E ) generated in step 4.2, and calculate the semantic probability p of each class of object o through the following formula o :
[0035]
[0036] Where, (S(I E )) o Represents the semantic prediction value of the enhanced image I E On the class o, I S Represents the pixel intensity value, Represents the element-wise multiplication operation, H, W are the height and width of the image;
[0037] Step 4.4, introduce the semantic segmentation loss function L sup-sem Is defined as follows:
[0038]
[0039] Among them, and respectively represent the semantic probability distributions of the enhanced image I E and the high-quality image I H on channels α and β. ∈ represents the color channel combination. is a dynamic adaptive coefficient used to measure the channel difference and is calculated by the following formula:
[0040]
[0041] In this process, CONV is a convolution operation used to extract the features of channel differences; BN (Batch Normalization) is used to normalize the features and stabilize the model training; ReLU is an activation function used to introduce non-linearity; Sigmoid finally maps the weight coefficient to the interval [0, 1].
[0042] Step 5 specifically includes the following steps:
[0043] Combine the contrast learning loss L uns-cons and the semantic segmentation loss L sup-sem to form the total loss function L total Train the enhancement network in an end-to-end manner to obtain the trained image enhancement network. The total loss function L total is defined as follows:
[0044] L total = αL uns-cons + βL sup-sem
[0045] Among them, α and β are weight parameters used to balance the contributions of the contrast learning loss and the semantic segmentation loss to the network training.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] 1. The present invention adopts a method combining semantic segmentation guidance and cross-domain feature contrast learning, achieving precise enhancement and detail restoration, and improving the color, contrast, and detail quality of underwater images.
[0048] 2. The present invention adopts a cross-domain contrast learning mechanism, realizing training without paired data, reducing the dependence on paired data sets, lowering the data acquisition cost, and having strong generalization ability.
[0049] In summary, the present invention has the effects of precise enhancement of underwater images, detail restoration, strong generalization ability, and low data acquisition cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the method flow of the present invention.
[0051] Figure 2 Schematic diagram of the overall structure of the model of the present invention.
[0052] Figure 3 Performance comparison chart of the present invention and existing methods.
[0053] Figure 4 Shows a comparison chart of the effects of the enhanced images generated by the present invention and other methods. Detailed implementation manners
[0054] The following will describe the present invention in detail with reference to the accompanying drawings.
[0055] As Figure 1 shown, the implementation of an underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance according to the present invention includes dataset construction, training of a cross-domain contrastive learning model, fine-tuning of a semantic segmentation module, and joint optimization of an image enhancement network. The specific implementation steps are as follows:
[0056] Step 1: Construct an unpaired underwater image sample dataset, including degraded underwater images and high-quality underwater images;
[0057] The degraded underwater images serve as the degraded image domain D L , and the high-quality underwater images and high-definition land images serve as the clear image domain D H .
[0058] Select 2,000 degraded underwater images from the EVUP dataset as the degraded image domain D L , and 890 high-quality images from the UIEB dataset and 800 high-definition land images from the DIV2K dataset as the clear image domain D H . Uniformly crop all images to 384×384 and perform normalization processing.
[0059] Step 2: Design and initialize an image enhancement network for generating clear images from degraded underwater images; the image enhancement network adopts a U-Net structure, which includes an encoder and a decoder part. The input is a degraded underwater image, and the output is an enhanced image. Its encoder part captures low-level and high-level features of the image through multi-scale feature extraction; its decoder part fuses the multi-scale features output by the encoder through skip connections, thereby retaining image details and improving the reconstruction ability.
[0060] Initialize the model parameters to ensure the gradient stability in the initial stage of training. The initialized image enhancement network will be jointly optimized by the contrastive learning loss and the semantic segmentation loss in the subsequent steps.
[0061] As Figure 2 shown, in step 3, the dataset constructed in step 1 is divided into positive samples and negative samples, and an image I is randomly selected from the positive samples H , and an image I is randomly selected from the negative samples L . The image I L is input into the image enhancement network designed in step 2 to obtain the enhanced image I E . Then, {I L , I H , I E} is input into the image feature extraction model to calculate the contrastive learning loss, thereby optimizing the cross-domain features and improving the feature learning ability of the underwater image enhancement network;
[0062] Step 3.1: Select a model F for image feature extraction and keep the parameters fixed to extract the latent feature representation of the image to improve the calculation efficiency. The feature extraction model F uses the pre-trained VGG16 network;
[0063] Step 3.2: Use the degraded image domain D in the unpaired underwater image sample dataset L as the negative sample and the clear image domain D H as the positive sample. Randomly sample the degraded underwater image I L and the clear image I H from the negative and positive samples respectively;
[0064] Step 3.3: Input the degraded underwater image I L into the image enhancement network to generate the enhanced image I E . Then, input the degraded underwater image I L , the enhanced image I E and the clear image I H into the model F with fixed parameters selected in step 3.1 to extract its multi-layer features. Extract the feature maps of the 1st, 3rd, and 5th layers of VGG16, which correspond to shallow, middle, and deep features respectively. Calculate the Gram matrix for each layer of the feature maps and convert them into the latent feature vectors v U , v E , v H to represent the cross-domain feature distribution of the image in the latent feature space;
[0065] Gram(F k (I)) = F k (I) · F k (I) T
[0066] where F k (I) represents the feature map output by the k-th layer in the feature extraction model F. The Gram matrix is used to capture the relationship between features, and the extracted latent feature vectors are respectively denoted as the feature vectors v of the degraded imageU , enhance the feature vector v of the image E , the feature vector v of the clear image H ;
[0067] Step 3.4, construct a contrastive learning loss function, by minimizing the feature distance between the feature vector v E and the feature vector v H , and maximizing the feature distance between the feature vector v E and the feature vector v U , make the features of the enhanced image approach those of the clear image and move away from those of the degraded image;
[0068] Minimize the feature distance between the feature vector v E and the feature vector v H , the positive sample loss in the contrastive learning loss function is calculated as follows:
[0069]
[0070] Maximize the feature distance between the feature vector v E and the feature vector v U , the negative sample loss in the contrastive learning loss function is calculated as follows:
[0071]
[0072] where τ is the temperature parameter, used to adjust the gradient range of contrastive learning;
[0073] Combining the above positive and negative sample loss functions, the final contrastive learning loss function is defined as: L uns-cons = L pulI + L push .
[0074] Step 4, pre-train the semantic segmentation model on the semantic segmentation dataset, and then fine-tune it on the semantic segmentation dataset of underwater images. Input the image I H and the image I E obtained in Step 3 into the fine-tuned semantic segmentation model respectively to generate corresponding semantic prediction maps, and calculate the semantic segmentation loss;
[0075] Step 4.1: Select a model for semantic segmentation. The semantic segmentation model adopts the DeepLabv3 network structure and is pre-trained on the Cityscapes dataset to initialize the network weights. The semantic segmentation dataset for underwater images uses the SUIM dataset. The SUIM dataset contains 1,640 semantically annotated images, which are used to fine-tune the semantic segmentation module. The dataset annotates 8 types of semantic targets (such as fish, coral, marine plants, etc.), and the annotation information for the blurred areas is manually corrected. To enhance the robustness of the model, data augmentation processes such as random rotation, cropping, and brightness adjustment are performed on the SUIM dataset to adapt to the semantic segmentation task of underwater images. All images are uniformly cropped to 384×384 and normalized;
[0076] Step 4.2: Use the semantic segmentation model fine-tuned in Step 4.1 to segment the enhanced image I E and the high-quality image I H to generate an enhanced semantic prediction map S(I E ) and a high-quality semantic prediction map S(I H ). These prediction maps serve as the visual guidance of the model to capture the subtle semantic differences in underwater images, thereby more accurately restoring the structure and details of the images;
[0077] Step 4.3: Utilize the semantic prediction map S(I E ) generated in Step 4.2 and calculate the semantic probability p o of each class of object o through the following formula:
[0078]
[0079] where (S(I E )) o represents the semantic prediction value of the enhanced image I E on the class o, I S represents the pixel intensity value, represents the element-wise multiplication operation, and H, W are the height and width of the image;
[0080] Step 4.4: To ensure that the enhanced image can be consistent with the high-quality image I H at the semantic level, a semantic segmentation loss function is introduced. This loss function not only evaluates the semantic differences between images but also, through semantic consistency constraints, promotes the structure and content of the enhanced image to tend to be real. The semantic segmentation loss function L sup-sem is defined as follows:
[0081]
[0082] where, and respectively represent the enhanced image IE and high-quality image I H Semantic probability distributions on channels α and β, ∈ represents color channel combinations, such as (R,G), (G,B), etc., is the dynamic adaptive coefficient, used to measure the channel difference, calculated by the following formula:
[0083]
[0084] In this process, CONV is the convolution operation, used to extract the features of channel difference; BN (Batch Normalization) is used to normalize the features and stabilize the model training; ReLU is an activation function, used to introduce non-linearity; Sigmoid finally maps the weight coefficient to the interval [0,1].
[0085] Step 5: Jointly train the image enhancement network designed in Step 2 with the contrast learning loss calculated in Step 3 and the semantic segmentation loss calculated in Step 4. Update the parameters of the image enhancement network through end-to-end training to generate high-quality enhanced underwater images. Use the images in the UIEB, EVUP, and RUIE datasets as the test datasets for the enhancement network model for performance evaluation;
[0086] The contrast learning loss L uns-cons and the semantic segmentation loss L sup-sem are combined to form the total loss function L total The enhancement network is trained in an end-to-end manner to obtain the trained image enhancement network. The total loss function L total is defined as follows:
[0087] L total = αL uns-cons + βL sup-sem
[0088] where α and β are weight parameters, used to balance the contributions of the contrast learning loss and the semantic segmentation loss to the network training;
[0089] The Adam optimizer is used for parameter optimization. The initial learning rate is set to 10 -4 , the batch size is set to 4, the number of training epochs is 100, and each epoch completes one pass through the entire training set.
[0090] The settings of the weight coefficients α and β in the joint optimization process are based on experimental tuning. The recommended default values are α = 0.5, β = 0.5.
[0091] Simulation experiment
[0092] Such as Figure 3As shown, in the embodiment, three metrics, UIQM, UCIQE, and MUSIQ, are adopted to evaluate the quality of underwater images, and tests are conducted on three publicly available datasets, UIEB, EUVP, and RUIE.
[0093] These metrics measure the performance of underwater image enhancement methods from different perspectives. Among them, UIQM (Underwater Image Quality Measure) mainly evaluates the brightness, color balance, and contrast of images. The higher the value, the better the overall image quality.
[0094] UCIQE (Underwater Color Image Quality Evaluation) measures image quality through chromaticity deviation, saturation, and contrast, mainly reflecting the color and clarity of the image.
[0095] MUSIQ is a deep learning-based perceptual quality metric that evaluates the perceptual quality of images by combining subjective and objective criteria. The higher the value, the closer the image is to human visual perception.
[0096] The experimental results on the three datasets, UIEB, EUVP, and RUIE, show that the method of the present invention performs excellently. Figure 3 The data of Ours in the table is the result of this embodiment.
[0097] In the UIQM metric, the present invention achieved the highest score of 3.208 on the RUIE dataset, indicating that the method of the present invention can significantly improve the brightness and color balance of underwater images.
[0098] At the same time, the present invention also achieved good performances of 2.661 and 3.112 on the UIEB and EUVP datasets respectively.
[0099] In the UCIQE metric, the present invention achieved 0.617 on the UIEB dataset, showing strong color restoration and contrast optimization capabilities.
[0100] In terms of the MUSIQ metric, the present invention performed most prominently on the EUVP dataset, obtaining the highest score of 48.14, indicating that the present invention is superior to other methods in perceptual quality evaluation.
[0101] At the same time, the present invention also achieved scores of 43.89 and 34.68 on the UIEB and RUIE datasets respectively.
[0102] The present invention performs excellently on multiple datasets and metrics, especially achieving the best results on the UIQM metric of the RUIE dataset and the MUSIQ metric of the EUVP dataset.
[0103] Such asFigure 4 As shown Figure 4 Column (a) in [reference] shows the effects of the present invention on the RUIE, EVUP, and UIEB datasets, Figure 4 and columns (b)–(g) in [reference] show the visual comparison of the effects of the GDCP, MMLE, WaterNet, FUNIE, CWR, and SEMUIR methods on the RUIE, EVUP, and UIEB datasets. It can be seen Figure 4 that the method of the present invention is superior to the prior art in multiple key aspects. In terms of color restoration, the present invention can more accurately restore the true color of underwater images, avoiding the common color cast problem in traditional methods. For example, compared with the GDCP and MMLE methods, the present invention can more effectively remove the color distortion caused by underwater illumination, making the enhanced image more natural and realistic. In terms of clarity improvement, the present invention significantly improves the detail representation ability of the image through an innovative image enhancement method. Compared with the WaterNet and FUNIE methods, the images generated by the present invention are clearer in terms of edge and texture details, and can better retain the original structural information of the underwater scene. In terms of defogging and contrast enhancement, compared with the CWR and SEMUIR methods, the present invention can more effectively remove the blur caused by underwater suspended particles and light scattering, making the enhanced image have a higher contrast and more obvious object contours.
[0104] This fully demonstrates that the present invention can effectively improve the quality, color, and clarity of images in underwater image enhancement tasks, and also has high advantages in visual perception. These results indicate that the present invention has strong generalization ability and robustness, and can adapt to different underwater environments and complex scenes.
Claims
1. An underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance, characterized in that It includes the following steps: Step 1: Construct an unpaired underwater image sample dataset; Step 2: Design and initialize an image enhancement network; Step 3: Divide the dataset constructed in Step 1 into positive and negative samples, and randomly select an image I from the positive samples H , and randomly select an image I from the negative samples L . Input the image I L into the image enhancement network designed in Step 2 to obtain the enhanced image I E . Input {I L , I H , I E} into the image feature extraction model to calculate the contrastive learning loss; Step 4: Pre-train the semantic segmentation model on the semantic segmentation dataset, and then fine-tune it on the semantic segmentation dataset of underwater images. Input the images I H and the image I E into the fine-tuned semantic segmentation model respectively to generate corresponding semantic prediction maps, and calculate the semantic segmentation loss; Step 5: Jointly train the image enhancement network designed in Step 2 with the contrastive learning loss calculated in Step 3 and the semantic segmentation loss calculated in Step 4, and update the parameters of the image enhancement network through end-to-end training to generate high-quality enhanced underwater images.
2. The underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance according to claim 1, wherein The dataset constructed in Step 1 includes degraded underwater images and high-quality underwater images; The degraded underwater image serves as the degraded image domain D L , and the high-quality underwater image and the high-definition land image serve as the clear image domain D H .
3. The underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance according to claim 1, wherein In Step 2, the image enhancement network adopts a U-Net structure, which includes an encoder and a decoder part. The input is a degraded underwater image, and the output is an enhanced image.
4. The underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance according to claim 1, wherein Step 3 specifically includes the following steps: Step 3.1: Select a model F for image feature extraction and keep the parameters fixed; Step 3.2, using the degraded image domain D in the unpaired underwater image sample dataset L as negative samples and the clear image domain D H as positive samples, randomly sampling the degraded underwater image I L and the clear image I H ; Step 3.3, input the degraded underwater image I L into the image enhancement network to generate the enhanced image I E , and input the degraded underwater image I L , the enhanced image I E and the clear image I H into the model F with fixed parameters selected in Step 3.1, extract its multi-layer features, calculate the Gram matrix for the feature map of each layer, and convert it into the latent feature vector v U , v E , v H to represent the cross-domain feature distribution of the image in the latent feature space; Gram(F k (I)) = F k (I)·F k (I) T Among them, F k (I) represents the feature map output by the k-th layer in the feature extraction model F. The Gram matrix is used to capture the relationship between features. The extracted latent feature vectors are respectively denoted as the feature vector v of the degraded image U , the feature vector v of the enhanced image E , and the feature vector v of the clear image H ; Step 3.4, construct a contrastive learning loss function, and by minimizing the feature distance between the feature vector v E and the feature vector v H , and maximizing the feature distance between the feature vector v E and the feature vector v U , make the features of the enhanced image approach the features of the clear image and move away from the features of the degraded image; Minimize the eigenvector v E With the eigenvector v H The feature distance of, the positive sample loss in the contrastive learning loss function is calculated as follows: Maximize the eigenvector v E For the eigenvector v U The feature distance, and the negative sample loss in the contrastive learning loss function is calculated as follows: where τ is a temperature parameter used to adjust the gradient range of contrastive learning; Combining the above positive and negative sample loss functions, the final contrastive learning loss function is defined as: L uns-cons = L pulI + L push .
5. The underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance according to claim 1, wherein The process of calculating the semantic segmentation loss in Step 4 includes: Step 4.1: Select a model for semantic segmentation, pre-train it on a semantic segmentation dataset, and then fine-tune it on the semantic segmentation dataset of underwater images; Step 4.2, use the semantic segmentation model fine-tuned in Step 4.1 to segment the enhanced image I E and the high-quality image I H to generate an enhanced semantic prediction map S(I E ) and a high-quality semantic prediction map S(I H ); Step 4.3, using the semantic prediction map S(I E ) generated in Step 4.2, and calculating the semantic probability p of each type of object o through the following formula o : Among them, (S(I E )) o represents the semantic prediction value of the enhanced image I E on the category i, and I S represents the pixel intensity value, represents the element-wise multiplication operation, and H and W are the height and width of the image; Step 4.4, introduce the semantic segmentation loss function L sup-sem It is defined as follows: Among them, and respectively represent the semantic probability distributions of the enhanced image I E and the high-quality image I H on channels α and β. ∈ represents the color channel combination. is a dynamic adaptive coefficient used to measure the channel difference and is calculated by the following formula: In this process, CONV is a convolution operation used to extract features of channel differences; BN (Batch Normalization) is used to normalize features and stabilize model training; ReLU is an activation function used to introduce non-linearity; Sigmoid finally maps the weight coefficients to the interval [0,1].
6. The underwater image enhancement method based on cross-domain feature representation and semantic segmentation guidance according to claim 1, wherein Step 5 specifically includes the following steps: Combine the contrastive learning loss $L$ uns-cons and the semantic segmentation loss $L$ sup-sem to form the total loss function $L$ total Train the enhancement network in an end-to-end manner to obtain the trained image enhancement network, and the total loss function $L$ total is defined as follows: L total = αL uns-cons + βL sup-sem where α and β are weight parameters used to balance the contributions of the contrastive learning loss and the semantic segmentation loss to network training.
Citation Information
Cited By
No-reference underwater image quality evaluation method based on subject and scene semantic prompt
CN120451763A
Semantic-driven frequency consistency underwater image enhancement method
CN121961897A
A semantic-driven frequency-consistent underwater image enhancement method
CN121961897B