Underwater image enhancement method and system based on visual-text fusion

Through the underwater image enhancement method of visual-text fusion, the differentiable color difference between natural images and underwater images is used to construct an enhancement network, which solves the problem of over- or under-enhancement in underwater image enhancement and achieves efficient image quality improvement under zero-sample conditions.

CN120634934BActive Publication Date: 2025-10-17崂山国家实验室
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511127097.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-17
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

Existing underwater image enhancement methods suffer from problems of over-enhancement or under-enhancement and scarce paired data, resulting in limited generalization capabilities in diverse and complex underwater scenes.

Method used

A method based on vision-text fusion is adopted, which utilizes the differentiable underwater color difference between natural images and underwater images. Through a high-speed image generation diffusion model and a contrast language-image pre-training model, an enhancement network is constructed. The image is enhanced using an unpaired training strategy and text prompts. The color, detail loss function and adversarial loss function are combined to generate enhanced images with natural colors and fine textures.

Benefits of technology

It achieves efficient processing of underwater images under zero-sample conditions, reduces color deviation, improves visual quality, overcomes the defects of relying on artificial reference images, and has stable and efficient zero-sample generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634934B_ABST
    Figure CN120634934B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of underwater image enhancement, and relates to an underwater image enhancement method and system based on visual-text fusion, the steps of the method being: constructing an enhancement network based on visual-text fusion; the enhancement network comprising a generator for processing an image to generate an enhanced image and a discriminator for determining whether the image is a real image or a generator-generated image, the generator comprising a high-speed image generation diffusion model and a text encoder based on a contrastive language-image pre-training model, the encoder, a U-net module and the decoder in the high-speed image generation diffusion model being sequentially connected in order, and the text encoder generating text embedding for U-net module adjustment; inputting a to-be-enhanced underwater image into the enhancement network to obtain an underwater enhanced image. The application realizes visual-text fusion by using the inference ability and strong prior knowledge of the high-speed image generation diffusion model, so that efficient processing and zero-shot generalization are achieved, the stability is high, and the visual quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of underwater image enhancement, and particularly relates to an underwater image enhancement method and system based on visual-text fusion. BACKGROUND

[0002] Underwater vision systems play an irreplaceable role in the fields of underwater robots, ocean observation, resource exploration, and ecological monitoring. However, due to the selective absorption and scattering effects of water medium on light, underwater images generally exhibit color distortion (mainly blue-green bars), reduced contrast, blurred details, and enhanced noise, which severely limit the accuracy of subsequent visual tasks. In order to improve the quality of underwater images, the industry has proposed a large number of underwater image enhancement methods. Existing underwater image enhancement methods can be roughly divided into three categories: physical model-based methods, non-physical model-based methods, and deep learning-based methods.

[0003] Physical model-based methods model the propagation of light in water. This method studies the propagation characteristics of light in water and constructs a mathematical model to describe the scattering, absorption, and attenuation behavior of water medium. This method estimates the propagation path and attenuation characteristics of light using the parameters of the underwater scene, and then restores the true color and details of the degraded image through inverse calculation. This method can better restore the true color and structure of the degraded image under certain conditions (laboratory or single water quality scene). However, this method often relies on simplified assumptions about the underwater scene, and the parameters in practice are difficult to accurately estimate, resulting in color drift or loss of details in the enhancement results. This makes the method less adaptable to diverse and complex scenes and has limited generalization ability.

[0004] Non-physical model-based methods ignore the physical process of underwater imaging and convert the enhancement task into a general image processing problem. This method does not focus on the physical degradation mechanism of underwater images, but instead uses existing image processing techniques to optimize pixel intensity values to improve visual quality. This method mainly enhances underwater images by adjusting color, contrast, and texture details, without relying on complex physical modeling, thereby generating more natural and satisfactory results. However, since they do not consider the underlying mechanism of underwater imaging, their adaptability to different scenes is often limited, which can easily lead to over-enhancement (local oversaturation, halos, etc.) or insufficient enhancement (greenish color, blurred details, etc.) of images.

[0005] The deep learning-based method utilizes the powerful representation ability of a neural network to learn the feature distribution from a large-scale underwater image dataset, so as to realize accurate reconstruction and quality improvement of the degraded image. Such a method constructs an enhancement task as an end-to-end nonlinear mapping, models the complex degradation process through network training, and thus generates an enhanced image with better visual quality. However, such a method is restricted by the following factors: paired data composed of real underwater images and their corresponding real non-underwater images (benchmark images) are scarce. In addition, the dependence on artificially selected reference images further limits the optimal performance. The lack of high-quality paired datasets hinders the effective training of the model, and thus restricts the generalization ability of the model in diverse underwater scenes. SUMMARY

[0006] To solve the above problems in the prior art, such as over-enhancement or insufficient enhancement of underwater images, scarcity of paired data, etc., the present application provides a method and system for underwater image enhancement based on visual-text fusion, which reduces color deviation and improves visual quality by using the differentiable underwater color difference between natural images and underwater images; realizes visual-text fusion by using the inference ability and strong prior knowledge of the high-speed image generation diffusion model, so as to achieve efficient processing and zero-shot generalization, strong stability, and ensure visual quality; overcomes the defects caused by the scarcity of paired data through a pair-free training strategy; and realizes effective underwater image enhancement by using only text prompts and any natural image under non-water conditions as a reference, overcoming the defects caused by the dependence on artificially selected reference images.

[0007] To achieve the above purpose, in a first aspect, the present application provides a method for underwater image enhancement based on visual-text fusion, comprising the following steps:

[0008] A network construction step: constructing an enhancement network based on visual-text fusion; the enhancement network comprises a generator and a discriminator, the generator is used to process images to generate enhanced images, and the discriminator is used to determine whether the image is a real image or an image generated by the generator; the generator comprises a high-speed image generation diffusion model and a text encoder based on a contrastive language-image pre-trained model, the encoder, the U-net module and the decoder in the high-speed image generation diffusion model are sequentially connected in order, and the text encoder generates a text embedding used for adjusting the U-net module;

[0009] An image enhancement step: inputting an underwater image to be enhanced into the enhancement network to obtain an underwater enhanced image.

[0010] In some embodiments, the generator and the discriminator are both provided with two; the method for constructing the enhancement network based on visual-text fusion is as follows:

[0011] An image acquisition step: acquiring any underwater image and any natural image under waterless condition ;

[0012] image processing step: transforming underwater image inputting the first generator with the first text prompt inputting the first generator as condition generating enhanced image ; transforming natural image inputting the second generator with the second text prompt inputting the second generator as condition generating virtual underwater image ; the first text prompt indicates that this is a waterless image, the second text prompt indicates that this is an underwater image

[0013] image reconstruction step: transforming enhanced image inputting the second generator with the second text prompt inputting the second generator as condition obtaining reconstructed underwater image ; transforming virtual underwater image inputting the first generator with the first text prompt inputting the first generator as condition obtaining reconstructed natural image ;

[0014] wavelet transform step: decomposing enhanced image and virtual underwater image into low-frequency components that preserve image color distribution and overall structure information and high-frequency components that capture image details and edge textures, respectively, by discrete wavelet transform

[0015] constraint step: constraining the low-frequency components by color loss function, constraining the high-frequency components by detail loss function, distinguishing real image from generated image by adversarial loss function, constraining consistency of image and reconstructed image by cycle consistency function, constraining matching degree of generator input and output by identity loss function, until reaching set constraint condition

[0016] In some embodiments, the total loss function of the generator is represented as:

[0017]

[0018] wherein, a weight coefficient for balancing the color loss function, a weight coefficient for balancing the color loss function, a color loss function, a weight coefficient for balancing the detail loss function, a detail loss function, a weight coefficient for balancing the adversarial loss function, an adversarial loss function, a weight coefficient for balancing the cycle consistency loss function, a cycle consistency loss function, a weight coefficient for balancing the identity loss function, an identity loss function;

[0019] The color loss function is expressed as:

[0020]

[0021] wherein:

[0022]

[0023]

[0024] In the formula, is a maximum function that selects a large value from two inputs, is a sum of soft assignment color histogram average similarities between red-green channels, red-blue channels and green-blue channels of the natural image, is a soft assignment color histogram average similarity between red-green channels of the natural image, is a soft assignment color histogram average similarity between red-blue channels of the natural image, is a soft assignment color histogram average similarity between green-blue channels of the natural image; is a sum of soft assignment color histogram average similarities between red-green channels, red-blue channels and green-blue channels of the underwater image, is a soft assignment color histogram average similarity between red-green channels of the underwater image, is a soft assignment color histogram average similarity between red-blue channels of the underwater image, is a soft assignment color histogram average similarity between green-blue channels of the underwater image;

[0025] The detail loss function is expressed as:

[0026]

[0027] wherein:

[0028]

[0029] Where, 、 、 are weight coefficients, which are used to balance the contributions of enhanced image detail loss, virtual underwater image loss, and total variation regularization; is the predefined target value for the enhanced image, is a predefined target value of a virtual underwater image; is the total number of pixels in the high-frequency component; is the total variation regularization term;

[0030] The adversarial loss function is expressed as:

[0031]

[0032] Where, is the first discriminator, is the second discriminator;

[0033] The cycle consistency loss function is expressed as:

[0034]

[0035] in:

[0036]

[0037] Where, represents the reconstruction loss function, is the processed image, is the target image, To balance the difference loss The weight coefficient of To balance learning perception of image patch similarity loss The weight coefficient of To balance the structural similarity index loss The weight coefficient of

[0038] The identity loss function is expressed as:

[0039]

[0040] The total loss function of the discriminator is expressed as:

[0041] .

[0042] In some embodiments, a skip connection is established between the encoder and the decoder via a zero convolutional layer.

[0043] In some embodiments, the generator further comprises a low-rank adaptation module connected to the encoder, the U-net module, and the decoder, respectively, to keep all original weights of the frozen high-speed image generation diffusion model unchanged and only update the low-rank adaptation module parameters, the initial convolutional layer of the U-net module, and the skip-connection convolutional layer in the encoder and the decoder during the fine-tuning process.

[0044] In some embodiments, the discriminator comprises a contrastive language-image pre-trained model serving as a frozen backbone network to extract a high-dimensional feature representation from the image of each input discriminator.

[0045] In some embodiments, the discriminator comprises a contrastive language-image pre-trained model serving as a frozen backbone network to extract a high-dimensional feature representation from the image of each input discriminator.

[0046] In some embodiments, the discriminator further comprises:

[0047] a convolutional down-sampling unit comprising two consecutive convolutional down-sampling modules, an input of the first convolutional down-sampling module being connected to an output of the contrastive language-image pre-trained model;

[0048] a feed-forward module having an input connected to an output of the second convolutional down-sampling module, for performing linear mapping and cross-channel fusion on the features output by the second convolutional down-sampling module;

[0049] an output layer having an input connected to an output of the feed-forward module, for mapping the features output by the feed-forward module into discriminator scores.

[0050] In some embodiments, the convolutional down-sampling module comprises a first convolutional layer, a first activation function layer, a blur pooling layer, and a second convolutional layer connected in sequence, the first convolutional layer being configured to perform channel projection, the first activation function layer being configured to introduce nonlinearity and anti-gradient disappearance, the blur pooling layer being configured to perform anti-aliasing down-sampling, and the second convolutional layer being configured to reduce spatial resolution.

[0051] the feed-forward module comprises a fully connected layer and a second activation function layer connected in sequence, the fully connected layer being configured to perform linear mapping, and the second activation function layer being configured to introduce nonlinearity and anti-gradient disappearance.

[0052] In a second aspect, the present application provides an underwater image enhancement system based on visual-text fusion, configured to implement the underwater image enhancement method based on visual-text fusion according to the first aspect of the present application, comprising:

[0053] a model construction module configured to construct an enhancement network based on visual-text fusion;

[0054] An image enhancement module inputs an underwater image to be enhanced into the enhancement network to obtain an underwater enhanced image.

[0055] In some embodiments, the model construction module comprises:

[0056] An image acquisition module acquires any underwater image and any natural image under water-free conditions;

[0057] An image processing module inputs the underwater image into a first generator to generate an enhanced image by inputting a first text prompt into the first generator; and inputs the natural image into a second generator to generate a virtual underwater image by inputting a second text prompt into the second generator.

[0058] An image reconstruction module inputs the enhanced image into the second generator to obtain a reconstructed underwater image by inputting the second text prompt into the second generator; and inputs the virtual underwater image into the first generator to obtain a reconstructed natural image by inputting the first text prompt into the first generator.

[0059] A wavelet transform module decomposes the enhanced image and the virtual underwater image into low-frequency components that retain image color distribution and overall structural information and high-frequency components that capture image details and edge textures through discrete wavelet transform.

[0060] A constraint module constrains the low-frequency components through a color loss function, constrains the high-frequency components through a detail loss function, distinguishes real images from generated images through an adversarial loss function, constrains the consistency of images and reconstructed images through a cycle consistency function, and constrains the matching degree of generator input and output through an identity loss function until a set constraint condition is reached.

[0061] Compared with the prior art, the advantages and positive effects of the present application are that:

[0062] (1) The underwater image enhancement method and system based on visual-text fusion provided by the present application construct an enhancement network based on visual-text fusion, utilize the inference ability and strong prior knowledge of a high-speed image generation diffusion model to realize visual-text fusion, thereby achieving efficient processing and zero-shot generalization, strong stability, ensuring visual quality, taking differentiable underwater color difference and detail intensity regularization as a clue to guide low-frequency color restoration and high-frequency detail optimization, and effectively generating enhanced images that have natural color appearance and fine texture reconstruction.

[0063] (2) The underwater image enhancement method and system based on visual-text fusion provided by the present application construct an enhancement network that, on the one hand, utilizes differentiable underwater color difference between natural images and underwater images to reduce color deviation and improve visual quality; and on the other hand, overcomes defects caused by paired data scarcity through an unpaired training strategy.

[0064] (3) The underwater image enhancement method and system based on visual-text fusion provided in this application can achieve effective underwater image enhancement by only using text prompts and any natural image under water-free conditions as a reference, overcoming the defects caused by relying on manual selection of reference images. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a flowchart of the underwater image enhancement method based on vision-text fusion described in an embodiment of the present application;

[0066] Figure 2 This is a structural block diagram of the enhanced network described in an embodiment of the present application;

[0067] Figure 3 This is a structural block diagram of the generator described in an embodiment of the present application;

[0068] Figure 4 A flowchart of a method for constructing an enhanced network based on visual-text fusion as described in an embodiment of the present application;

[0069] Figure 5 A flowchart of a method for constructing an enhanced network based on visual-text fusion as described in an embodiment of the present application;

[0070] Figure 6 This is a structural block diagram of the discriminator described in an embodiment of the present application;

[0071] Figure 7 This is a structural block diagram of the convolution downsampling module described in an embodiment of the present application;

[0072] Figure 8 This is a structural block diagram of the feedforward module described in an embodiment of the present application;

[0073] Figure 9 The natural image described in the embodiment of this application;

[0074] Figure 10 A hard-assigned color histogram of the natural image described in the embodiment of the present application;

[0075] Figure 11 The underwater image described in the embodiment of the present application;

[0076] Figure 12 A hard-assigned color histogram of the underwater image described in the embodiment of the present application;

[0077] Figure 13 A soft-assigned color histogram of the natural image described in the embodiment of the present application;

[0078] Figure 14 A soft-assigned color histogram of the underwater image described in an embodiment of the present application;

[0079] Figure 15 A structural block diagram of the underwater image enhancement system based on visual-text fusion according to the embodiments of the present application is shown in the figure.

[0080] Figure 16 A result diagram of underwater image enhancement using the underwater image enhancement system based on visual-text fusion according to the embodiments of the present application and the existing image enhancement method is shown in the figure.

[0081] Figure 17 A result diagram of underwater image detail enhancement using the underwater image enhancement system based on visual-text fusion according to the embodiments of the present application and the existing image enhancement method is shown in the figure.

[0082] In the figure, 1, generator, 101, text encoder, 102, encoder, 103, U-net module, 104, decoder, 105, zero convolution layer, 106, low-rank adaptation module, 107, input image, 108, generated image, 2, discriminator, 21, contrastive language-image pre-training model, 22, convolution down-sampling unit, 221, first convolution down-sampling module, 222, second convolution down-sampling module, 2201, first convolution layer, 2202, first activation function layer, 2203, blur pooling layer, 2204, second convolution layer, 23, feedforward module, 231, fully connected layer, 232, second activation function layer, 24, output layer, 3, model construction module, 301, image acquisition module, 302, image processing module, 303, image reconstruction module, 304, wavelet transform module, 305, constraint module, 4, image enhancement module. DETAILED DESCRIPTION

[0083] The present application will be described in detail below through exemplary embodiments. However, it should be understood that the elements, structures and features in one embodiment can also be beneficially combined into other embodiments without further description.

[0084] In view of the problems of excessive enhancement or insufficient enhancement, lack of paired data and the like existing in the existing underwater image enhancement method, the present application provides an underwater image enhancement method and system based on visual-text fusion, which reduces color deviation and improves visual quality by using the differentiable underwater color difference between natural images and underwater images, and realizes visual-text fusion by using the inference ability and strong prior knowledge of the diffusion model generated by high-speed images, so as to achieve efficient processing and zero-shot generalization, and overcome the defects caused by the lack of paired data through the unpaired training strategy, and effectively enhance the underwater image by using only the text prompt and any natural image under the condition of no water as a reference, thereby overcoming the defects caused by the dependence on artificial selection of reference images. The underwater image enhancement method and system based on visual-text fusion provided by the present application will be described in detail below with reference to the accompanying drawings.

[0085] Referring to Figures 1 to 3 , the first aspect embodiment of the present application provides an underwater image enhancement method based on visual-text fusion, the steps of which are:

[0086] S1, network construction step: constructing an enhancement network based on visual-text fusion; the enhancement network comprises a generator 1 and a discriminator 2, the generator 1 is used to process images to generate enhanced images, and the discriminator 2 is used to determine whether the image is a real image or an image generated by the generator; the generator 1 comprises a high-speed image generation diffusion model and a text encoder 101 based on a contrastive language-image pre-trained model, an encoder 102, a U-net module 103, and a decoder 104 in the high-speed image generation diffusion model are sequentially connected in order, and the text encoder 101 generates text embedding used for adjusting the U-net module 103.

[0087] Referring to Figure 4 , Figure 5 , the method for constructing the enhancement network based on visual-text fusion is:

[0088] S11, image acquisition step: acquiring any underwater image and any natural image under water-free conditions ;

[0089] S12, image processing step: inputting the underwater image into a first generator , inputting a first text prompt into the first generator as a condition to generate an enhanced image ; inputting the natural image into a second generator , inputting a second text prompt into the second generator as a condition to generate a virtual underwater image ; the first text prompt indicates that this is a water-free image, and the second text prompt indicates that this is an underwater image;

[0090] S13, image reconstruction step: inputting the enhanced image into the second generator , inputting the second text prompt into the second generator as a condition to obtain a reconstructed underwater image ; inputting the virtual underwater image into the first generator , inputting the first text prompt into the first generator as a condition Get reconstructed natural image ;

[0091] S14, wavelet transform step: enhance the image by discrete wavelet transform Decompose into low-frequency components that retain image color distribution and overall structural information and high-frequency components that capture image details and edge textures ; The virtual underwater image is transformed by discrete wavelet transform Decompose into low-frequency components that retain image color distribution and overall structural information and high-frequency components that capture image details and edge textures ;

[0092] S15, constraint step: constraining the low-frequency component by a color loss function , constraining the high-frequency components through the detail loss function , the adversarial loss function is used to distinguish the real image from the generated image, the cycle consistency function is used to constrain the consistency between the image and the reconstructed image, and the identity loss function is used to constrain the matching degree between the generator input and output until the set constraints are met.

[0093] It should be noted that the order of the image reconstruction step S13 and the wavelet transform step S14 can be changed, that is, S13, wavelet transform step: the enhanced image is transformed by discrete wavelet transform Decompose into low-frequency components that retain image color distribution and overall structural information and high-frequency components that capture image details and edge textures ; The virtual underwater image is transformed by discrete wavelet transform Decompose into low-frequency components that retain image color distribution and overall structural information and high-frequency components that capture image details and edge textures ; S14, image reconstruction step: enhance the image Enter the second generator , with the second text prompt Enter the second generator as a condition Get the reconstructed underwater image ; The virtual underwater image Input the first generator , prompt with the first text Enter the first generator as a condition Get reconstructed natural image The image reconstruction step S13 and the wavelet transform step S14 may also be performed simultaneously.

[0094] It should be noted that after the enhancement network training is completed, only the generator can complete the underwater image enhancement task. Specifically, given an underwater image and a text prompt, the generator outputs an enhanced underwater image (i.e., an underwater enhanced image).

[0095] In an embodiment of the present application, the total loss function of the generator is represented as:

[0096]

[0097] In the formula, is the total loss function of the generator, is the weight coefficient of the balance color loss function, is the color loss function, is the weight coefficient of the balance detail loss function, is the detail loss function, is the weight coefficient of the balance adversarial loss function, is the adversarial loss function, is the weight coefficient of the balance cycle consistency loss function, is the cycle consistency loss function, is the weight coefficient of the balance identity loss function, is the identity loss function.

[0098] Specifically, the color loss function is represented as:

[0099]

[0100] Wherein:

[0101]

[0102]

[0103] In the formula, is the maximum value function of selecting the maximum value from two inputs, is the sum of the average similarity of the soft assignment color histogram between the red-green channel, the red-blue channel and the green-blue channel of the natural image, is the average similarity of the soft assignment color histogram between the red-green channel of the natural image, is the average similarity of the soft assignment color histogram between the red-blue channel of the natural image, is the average similarity of the soft assignment color histogram between the green-blue channel of the natural image; is the sum of the average similarity of the soft assignment color histogram between the red-green channel, the red-blue channel and the green-blue channel of the underwater image, is the average similarity of the soft assignment color histogram between the red-green channel of the underwater image, is the average similarity of the soft-assigned color histogram between the red and blue channels of the underwater image, The average similarity of the soft-assigned color histogram between the green and blue channels of the underwater image.

[0104] It should be noted that in the RGB color space, an image consists of three channels: red, green, and blue. Figure 9 Natural images typically exhibit a certain degree of hard distribution between any two color channels in their color histograms (see Figure 10 ) similarity, that is, dividing the intensity range of each channel into a fixed space and counting the number of pixels in each interval, the hard-assigned histogram similarity between any two channels of natural images is high. However, see Figure 11 Hard-assigned color histogram of the underwater image shown (see Figure 12 ), the hard-assigned histogram similarity between any two channels of underwater images is low, which shows that the phenomenon of high hard-assigned color histogram similarity between any two channels does not apply to underwater images. To quantitatively analyze this phenomenon, we first construct a hard-assigned color histogram for each channel and then calculate the color channel as follows and color channels Hard-assigned color histogram similarity between :

[0105]

[0106] Where, For color channels The hard-assigned color histogram of For color channels The hard-assigned color histogram of is the average value over all histogram bins, For color channels The hard-assigned color histogram of For color channels The hard-assigned color histogram of is the average value over all histogram bins, Indicates the sum of all histogram bins.

[0107] The above formula is used to calculate the average similarity of the hard-assigned color histograms between the red-green channel, red-blue channel, and green-blue channel in natural images and underwater images. It is found that the average similarity of the hard-assigned color square maps of these channels in natural images is significantly higher than that in underwater images. This phenomenon is called underwater color difference. Based on this result, the underwater color difference index is constructed. , used to quantitatively evaluate images Underwater color differences are defined as follows:

[0108]

[0109] wherein, is the hard-assignment color histogram average similarity between red and green channels, is the hard-assignment color histogram average similarity between red and blue channels, is the hard-assignment color histogram average similarity between green and blue channels.

[0110] It is also worth mentioning that although the underwater color difference index can effectively quantify the color inconsistency between underwater images and natural images, this index is not suitable as a loss function in network training, the main reason being that this index involves non-differentiable operations such as discrete binning and counting, which will hinder the gradient backpropagation required for the optimization process. To enable its application in the enhancement network in the present application, a differentiable underwater color difference index is constructed, a soft-assignment color histogram is adopted, and a Gaussian kernel function is introduced to ensure that each pixel value produces continuous contributions to all histogram intervals.

[0111] Specifically, let denote the i-th pixel value, denote the center of the j-th soft-assignment color histogram interval, then the soft-assignment weight of pixel to the j-th soft-assignment color histogram interval is

[0112]

[0113] wherein, is the total number of pixels, is the total number of intervals, is the smoothing degree of control assignment.

[0114] The soft-assignment color histogram of the i-th interval is obtained by summing the soft-assignment weights of all pixels and normalizing by the total weight of all intervals.

[0115] wherein,

[0116] is all intervals in the normalization process, is the soft-assignment weight of the j-th soft-assignment color histogram interval,

[0117] is a small constant set to ensure numerical stability. The complete soft-assignment color histogram is represented as: ​​​​​​

[0118]

[0119] According to the complete soft assignment color histogram The formula for calculating the soft assignment color histogram of natural images (see Figure 9 ) is shown in Figure 13 The formula for calculating the soft assignment color histogram of underwater images (see Figure 11 ) is shown in Figure 14 As can be seen from Figure 13 , Figure 14 , any two color channels of natural images exhibit a certain degree of soft assignment color histogram similarity, while underwater images do not. In order to further quantitatively analyze this linearity, according to the hard assignment color histogram similarity calculation formula and the complete soft assignment color histogram calculation formula, the soft assignment color histogram similarity between color channel and color channel is calculated according to the following formula:

[0120]

[0121] In the formula, is the soft assignment color histogram of color channel , is the average value of the soft assignment color histogram of color channel over all histogram intervals, is the soft assignment color histogram of color channel , is the average value of the soft assignment color histogram of color channel over all histogram intervals.

[0122] Specifically, according to the soft assignment color histogram similarity calculation formula, the soft assignment color histogram average similarity between the red-green channel, red-blue channel and green-blue channel of the 10 natural image datasets is calculated, denoted as , , respectively, and the soft assignment color histogram average similarity between the red-green channel, red-blue channel and green-blue channel of the 10 underwater image datasets is calculated, denoted as , , respectively. The calculation results are shown in Table 1.

[0123] Table 1

[0124]

[0125] As shown in Table 1, the average similarity of the soft-assigned color histograms of natural images is 、 、 , and always maintain a high level. In comparison, the average similarities of the soft-assigned color histograms of underwater images are 、 、 , significantly lower. These results show that both soft-assignment and hard-assignment color histogram similarity can effectively capture the difference between natural images and underwater images. Unlike hard-assignment color histogram similarity, which relies on discrete binning and non-differentiable computational operations, soft-assignment color histogram similarity distributes the contribution of each pixel to all histogram bins in a continuous and differentiable manner through a Gaussian function. Due to this differentiable property, the observed color difference between natural images and underwater images quantified by soft-assignment color histogram similarity is called differentiable underwater color difference. Based on this result, a differentiable underwater color difference index is constructed. , used to quantitatively evaluate images The differentiable underwater color difference is defined as follows:

[0126]

[0127] Where, is the average similarity of the soft-assigned color histogram between the red and green channels, is the average similarity of the soft-assigned color histogram between the red and blue channels, is the average similarity of the soft-assigned color histogram between the green and blue channels.

[0128] When enhancing underwater images to resemble natural images, the soft-assigned color histogram similarity needs to be increased to match the natural image's similarity level. Conversely, when degrading natural images to simulate the underwater image's appearance, the soft-assigned color histogram similarity needs to be reduced to match the underwater image's similarity level. This application constructs the color loss function based on the average similarity of the soft-assigned color histograms of natural and underwater images. This color loss function, designed in this application, can guide the optimization of low-frequency color restoration and effectively generate enhanced images with the appearance of natural images.

[0129] Specifically, the detail loss function is expressed as:

[0130]

[0131] in:

[0132]

[0133] wherein, , , are weight coefficients, respectively used to balance the contribution of the enhanced image detail loss, the virtual underwater image loss and the total variation regularization; is a predefined target value of the enhanced image, is a predefined target value of the virtual underwater image; is the total number of pixels in the high-frequency component; is the total variation regularization term.

[0134] It should be noted that, since the high-frequency component captures the fine details and edge textures of the image, the detail loss function described in the present application designs a detail intensity regularization to constrain the high-frequency component. For the enhanced image, the detail intensity regularization promotes the detail intensity to the predefined target value to achieve moderate detail enhancement. To suppress excessive noise, total variation regularization is introduced to ensure that the enhanced image presents natural and delicate details without introducing artificial textures or noise artifacts. For the virtual underwater image, the detail intensity regularization can promote the detail intensity to the predefined target value. Since the goal in this case is to suppress details, total variation regularization should not be applied. The detail loss function designed in the present application can guide the optimization of high-frequency details, effectively generating enhanced images with fine texture reconstruction.

[0135] Specifically, the adversarial loss function is represented as:

[0136]

[0137] wherein, is the first discriminator, is the second discriminator.

[0138] The present application adopts the adversarial loss function to ensure that the discriminators cannot distinguish between generated images and real images. The present application constructs two discriminators, the first discriminator is used to distinguish between enhanced images and natural images, and the second discriminator is used to distinguish between virtual underwater images and underwater images. The adversarial loss function designed in the present application can make the first discriminator unable to distinguish between enhanced images and natural images, and make the second discriminator unable to distinguish between virtual underwater images and underwater images.

[0139] Specifically, the cycle consistency loss function is represented as:

[0140]

[0141] wherein:

[0142]

[0143] In the formula, denotes a reconstruction loss function, is a processed image, is a target image, is a weight coefficient for balancing the difference loss is a weight coefficient for balancing the learning perceptual image block similarity loss is a weight coefficient for balancing the structural similarity index loss . In the cyclic consistency loss function of the present application, the difference loss emphasizes point-by-point accuracy at the pixel level, which helps to preserve overall color fidelity and fine details. The learning perceptual image block similarity loss

[0144] addresses the problem of similar pixel values but different visual perception, effectively addressing high-level semantics and perceptual quality. The structural similarity index loss constrains local structure and texture consistency, ensuring that the generated image maintains coherent structure in local regions. These three losses collectively form balanced constraints on image quality. The cyclic consistency loss function used in the present application ensures that the image can be mapped back to its original appearance after continuous conversion. The cyclic consistency loss function designed in the present application forces the forward mapping and the reverse mapping to be consistent, thereby improving the stability of the enhancement network training process and the reliability of the generated results.

[0145] Specifically, the identity loss function is represented as:

[0146]

[0147] The present application introduces an identity loss function to ensure that the color and content of the image already belonging to the target domain are preserved. Specifically, when the input image comes from the target domain, this function encourages the generator to output a result that highly matches the input, helping to avoid unnecessary modifications and color shifts.

[0148] The total loss function of the discriminator is represented as:

[0149]

[0150] .

[0151] In an embodiment of the present application, continuing to refer to Figure 3 , the encoder 102 and the decoder 104 are connected through a zero convolution layer 105. By establishing a jump connection through a zero convolution layer, these jump paths enable the key spatial features of the input image to be preserved in the output, effectively preserving the detailed information of the image.

[0152] In an embodiment of the present application, continuing to refer to​​Figure 3 The generator 1 further comprises a low-rank adaptation module 106 connected to the encoder 102, the U-net module 103 and the decoder 104, respectively, to keep all original weights of the frozen high-speed image generation diffusion model unchanged and only update the parameters of the low-rank adaptation module 106, the initial convolutional layers of the U-net module 103, and the skip-connection convolutional layers in the encoder 102 and the decoder 104 during the fine-tuning process. The low-rank adaptation module can achieve efficient adaptation with minimal overhead. This selective training strategy only requires a small number of trainable parameters while effectively preserving the prior knowledge embedded in the enhanced network.

[0153] In an embodiment of the present application, referring to Figure 6 The discriminator 2 comprises a contrastive language-image pre-trained model 21 serving as a frozen backbone network to extract high-dimensional feature representations from each input image of the discriminator 2. The use of the contrastive language-image pre-trained model as the backbone network can obtain high-quality and semantically rich high-dimensional feature representations without retraining, thereby reducing the overall training cost.

[0154] In some embodiments, continuing to refer to Figure 6 The discriminator 2 further comprises:

[0155] a convolutional down-sampling unit 22 comprising two consecutive convolutional down-sampling modules, the input of a first convolutional down-sampling module 221 being connected to the output of the contrastive language-image pre-trained model 21;

[0156] a feedforward module 23 having an input connected to the output of a second convolutional down-sampling module 222 and being configured to perform linear mapping and cross-channel fusion on the features output by the second convolutional down-sampling module 222;

[0157] an output layer 24 having an input connected to the output of the feedforward module 23 and being configured to map the features output by the feedforward module 23 into discriminator scores.

[0158] The convolutional down-sampling unit and the backbone network form a "coarse-fine" structure, which preserves global semantics while introducing local texture information, thereby enhancing the sensitivity of the discriminator to subtle fake traces.

[0159] In an embodiment of the present application, referring to Figure 7The convolution downsampling module comprises, in sequence, a first convolution layer 2201, a first activation function layer 2202, a blur pooling layer 2203, and a second convolution layer 2204. The first convolution layer 2201 is configured to perform channel projection. The first activation function layer 2202 is configured to introduce nonlinearity and gradient disappearance resistance. The blur pooling layer 2203 is configured to perform anti-aliasing downsampling. The second convolution layer 2204 is configured to reduce spatial resolution. The use of two convolution downsampling modules with the same structure can simplify the training process and reduce the parameter quantity.

[0160] In an embodiment of the present application, referring to Figure 8 The feedforward module 23 comprises, in sequence, a fully connected layer 231 and a second activation function layer 232. The fully connected layer 231 is configured to perform linear mapping. The second activation function layer 232 is configured to introduce nonlinearity (to enhance the fitting capability of the network for complex decision boundaries) and gradient disappearance resistance.

[0161] S2, image enhancement step: inputting the underwater image to be enhanced into the enhancement network to obtain an underwater enhanced image.

[0162] Referring to Figure 15 The second aspect embodiment of the present application provides an underwater image enhancement system based on visual-text fusion, which is configured to implement the underwater image enhancement method based on visual-text fusion according to the first aspect of the present application, and comprises:

[0163] A model construction module 3 is configured to construct an enhancement network based on visual-text fusion.

[0164] An image enhancement module 4 is configured to input an underwater image to be enhanced into the enhancement network to obtain an underwater enhanced image.

[0165] In some embodiments, referring to Figure 15 The model construction module 3 comprises:

[0166] An image acquisition module 301 is configured to acquire any underwater image and any natural image under water-free conditions.

[0167] An image processing module 302 is configured to input the underwater image into a first generator to generate an enhanced image by inputting a first text prompt into the first generator as a condition; and input the natural image into a second generator to generate a virtual underwater image by inputting a second text prompt into the second generator as a condition.

[0168] An image reconstruction module 303 is configured to input the enhanced image into the second generator to obtain a reconstructed underwater image by inputting the second text prompt into the second generator as a condition; and input the virtual underwater image into the first generator to obtain a reconstructed natural image by inputting the first text prompt into the first generator as a condition.

[0169] The wavelet transform module 304 decomposes the enhanced image and the virtual underwater image into a low-frequency component that retains the image color distribution and overall structure information and a high-frequency component that captures the image details and edge texture, respectively, through discrete wavelet transform;

[0170] The constraint module 305 constrains the low-frequency component through the color loss function, constrains the high-frequency component through the detail loss function, distinguishes the real image from the generated image through the adversarial loss function, constrains the consistency of the image and the reconstructed image through the cycle consistency function, and constrains the matching degree of the generator input and output through the identity loss function until the set constraint conditions are met.

[0171] The following, combined with specific examples, compares the aforementioned underwater image enhancement method and system based on visual-text fusion (hereinafter referred to as the present application) with nine other methods to verify its effectiveness. The other nine methods are GIFM, C3HLM, CRUHL, MSPE, MFCP, CBAF, PUGAN, LANET, and RLUIE. GIFM, C3HLM, and CRUHL are categorized as physical model-based methods; MSPE, MFCP, and CBAF are categorized as non-physical model methods; and PUGAN, LANET, and RLUIE are categorized as deep learning methods.

[0172] Example: The commonly used Underwater Image Enhancement Benchmark (UIEB) dataset is used as a verification dataset. The dataset contains 890 real underwater images collected under various environmental conditions.

[0173] The quality of underwater images was assessed using six metrics: UIQM, CCF, URanker, AG, EI, and BCRA. UIQM combines three sub-metrics: image color, sharpness, and contrast, to comprehensively characterize the perceptual quality of an image. CCF is a weighted sum of color, contrast, and haze. URanker utilizes an efficient convolutional-attention Transformer architecture to provide reliable visual quality ranking for underwater images enhanced by various methods. Higher scores for UIQM, CCF, and URanker indicate better visual quality of the enhanced image. AG reflects image clarity, while EI quantifies the strength of image edge structure. Higher scores for these two metrics indicate better visual quality. BCRA assesses the improvement in detail before and after image enhancement. It consists of two complementary components: one measures the improvement in edge visibility, and the other quantifies the increase in the gradient amplitude of edge pixels. Higher BCRA scores indicate better image enhancement in terms of edge fidelity and detail preservation.

[0174] The comprehensive enhancement performance of various methods is evaluated and compared on the UIEB dataset. As shown in Figure 16 Figure 16 In (a), (b), (c), (d), (e), (f), (g), (h), (i), (j), and (k), (a) corresponds to the original image, (b) corresponds to GIFM, (c) corresponds to C3HLM, (d) corresponds to CRUHL, (e) corresponds to MSPE, (f) corresponds to MFCP, (g) corresponds to CBAF, (h) corresponds to PUGAN, (i) corresponds to LANET, (j) corresponds to RLUIE, and (k) corresponds to the present application. For the image Img1, GIFM, MSPE, PUGAN, LANET, and RLUIE cannot effectively eliminate the color deviation; although C3HLM, CRUHL, MFCP, and CBAF alleviate the color deviation to some extent, there are still obvious residual color deviations in the background. In addition, these methods have limited effect on improving the contrast of the image, and MFCP introduces obvious loss of structural details.

[0175] For the image Img2, GIFM, CRUHL, MSPE, PUGAN, LANET, and RLUIE cannot sufficiently improve the contrast or eliminate the color deviation, and CRUHL even introduces additional red tones. Although C3HLM, MFCP, and CBAF correct the color deviation to some extent, they fail to effectively improve the contrast and cannot fully restore the structural detail information.

[0176] For the image Img3, although GIFM, C3HLM, CRUHL, MSPE, MFCP, CBAF, PUGAN, LANET, and RLUIE correct the color deviation to some extent, they all fail to completely remove the blue deviation in the background.

[0177] In contrast, the present application can successfully eliminate the color deviation, enhance the image contrast, and enrich the image details.

[0178] Figure 17 A comparative analysis of the detail enhancement results of the present application and different methods is given. Figure 17 In (a), (b), (c), (d), (e), (f), (g), (h), (i), (j), and (k), (a) corresponds to the original image, (b) corresponds to GIFM, (c) corresponds to C3HLM, (d) corresponds to CRUHL, (e) corresponds to MSPE, (f) corresponds to MFCP, (g) corresponds to CBAF, (h) corresponds to PUGAN, (i) corresponds to LANET, (j) corresponds to RLUIE, and (k) corresponds to the present application. In the entire image area, the present application shows excellent ability in solving common underwater degradation problems, including color deviation, low contrast, and blurred details. In addition, in the local area highlighted by the red box, the present application produces clearer structures and richer textures compared to other methods.

[0179] ​Quantitative evaluation results are computed on the UIEB dataset, as shown in Table 2, where boldface indicates the best performance and blue font indicates the second best performance. The present application achieves the best performance in all six evaluation metrics, including UIQM, CCF, AG, EI, BCRA, and URanker. This demonstrates the effectiveness of the present application in enhancing underwater images captured in a wide range of real-world scenarios.

[0180] Table 2

[0181]

[0182] The above embodiments are used to explain the present application, rather than limit the present application, and any modifications and changes made to the present application within the spirit and scope of the present application and the claims fall within the protection scope of the present application.

Claims

1. An underwater image enhancement method based on vision-text fusion, characterized in that: The steps are: Network construction steps: constructing an enhancement network based on visual-text fusion; the enhancement network includes a generator and a discriminator, the generator is used to process the image to generate an enhanced image, and the discriminator is used to determine whether the image is a real image or an image generated by the generator; the generator includes a high-speed image generation diffusion model and a text encoder based on a contrastive language-image pre-trained model, the encoder, U-net module, and decoder in the high-speed image generation diffusion model are connected in sequence, and the text encoder generates a text embedding for adjustment by the U-net module; Image enhancement step: inputting the underwater image to be enhanced into the enhancement network to obtain an underwater enhanced image; The generator and the discriminator are both provided with two; the method for constructing an enhanced network based on visual-text fusion is: Image acquisition steps: Acquire any underwater image and any natural image without water ; Image processing steps: underwater images Input the first generator , prompt with the first text Enter the first generator as a condition Generate enhanced images ; natural images Enter the second generator , prompt with the second text Enter the second generator as a condition Generate virtual underwater images ; The first text prompt Indicates that this is a waterless image. The second text prompt Indicates that this is an underwater image; Image reconstruction step: enhance the image Enter the second generator , with the second text prompt Enter the second generator as a condition Get the reconstructed underwater image ; The virtual underwater image Enter the first generator , prompt with the first text Enter the first generator as a condition Get reconstructed natural image ; Wavelet transform step: Enhance the image by discrete wavelet transform and virtual underwater images Decomposed into low-frequency components that retain image color distribution and overall structural information and high-frequency components that capture image details and edge textures; Constraint step: constrain the low-frequency components through the color loss function, constrain the high-frequency components through the detail loss function, distinguish the real image from the generated image through the adversarial loss function, constrain the consistency of the image and the reconstructed image through the cycle consistency function, and constrain the matching degree between the generator input and output through the identity loss function until the set constraints are met; The total loss function of the generator is expressed as: Where, is the total loss function of the generator, To balance the weight coefficient of the color loss function, is the color loss function, To balance the weight coefficient of the detail loss function, is the detail loss function, To balance the weight coefficient of the adversarial loss function, To counter the loss function, is the weight coefficient of the balanced cycle consistency loss function, is the cycle consistency loss function, is the weight coefficient to balance the identity loss function, is the identity loss function; The color loss function is expressed as: in: Where, is a maximum function that selects the larger value from two inputs, is the sum of the average similarities of the soft-assigned color histograms between the red-green channel, red-blue channel, and green-blue channel of the natural image, is the average similarity of the soft-assigned color histogram between the red and green channels of natural images, is the average similarity of the soft-assigned color histogram between the red and blue channels of natural images, is the average similarity of the soft-assigned color histogram between the green and blue channels of natural images; is the sum of the average similarities of the soft-assigned color histograms between the red-green channel, red-blue channel, and green-blue channel of the underwater image, is the average similarity of the soft-assigned color histogram between the red and green channels of the underwater image, is the average similarity of the soft-assigned color histogram between the red and blue channels of the underwater image, is the average similarity of the soft-assigned color histogram between the green and blue channels of the underwater image; The detail loss function is expressed as: in: Where, 、 、 are weight coefficients, which are used to balance the contributions of enhanced image detail loss, virtual underwater image loss, and total variation regularization; is the predefined target value for the enhanced image, is a predefined target value of a virtual underwater image; is the total number of pixels in the high-frequency component; is the total variation regularization term; The adversarial loss function is expressed as: Where, is the first discriminator, is the second discriminator; The cycle consistency loss function is expressed as: in: Where, represents the reconstruction loss function, is the processed image, is the target image, To balance the difference loss The weight coefficient of To balance learning perception of image patch similarity loss The weight coefficient of To balance the structural similarity index loss The weight coefficient of The identity loss function is expressed as: The total loss function of the discriminator is expressed as: 。 2. The underwater image enhancement method based on vision-text fusion according to claim 1, characterized in that: A skip connection is established between the encoder and the decoder via a zero convolution layer.

3. The underwater image enhancement method based on vision-text fusion according to claim 2, characterized in that: The generator also includes a low-rank adaptation module, which is respectively connected to the encoder, U-net module, and decoder to keep all the original weights of the high-speed image generation diffusion model frozen during the fine-tuning process, and only update the low-rank adaptation module parameters, the initial convolutional layer of the U-net module, and the jump connection convolutional layer in the encoder and decoder.

4. The underwater image enhancement method based on vision-text fusion according to claim 1, characterized in that: The discriminator includes a contrastive language-image pre-trained model, which serves as a frozen backbone network to extract high-dimensional feature representations from each image input to the discriminator.

5. The underwater image enhancement method based on vision-text fusion according to claim 4, characterized in that: The discriminator further includes: A convolutional downsampling unit, comprising two consecutive convolutional downsampling modules, wherein the input of the first convolutional downsampling module is connected to the output of the contrastive language-image pre-trained model; A feedforward module, whose input is connected to the output of the second convolution downsampling module, and is used to perform linear mapping and cross-channel fusion on the features output by the second convolution downsampling module; The output layer, whose input is connected to the output of the feedforward module, is used to map the features output by the feedforward module to the discriminator scores.

6. The underwater image enhancement method based on vision-text fusion according to claim 5, characterized in that: The convolution downsampling module includes a first convolution layer, a first activation function layer, a fuzzy pooling layer, and a second convolution layer connected in sequence, wherein the first convolution layer is used for channel projection, the activation function layer is used to introduce nonlinearity and anti-gradient disappearance, the fuzzy pooling layer is used for anti-aliasing downsampling, and the second convolution layer is used to reduce spatial resolution; The feedforward module includes a fully connected layer and a second activation function layer connected in sequence, the fully connected layer is used to perform linear mapping, and the second activation function layer is used to introduce nonlinearity and resist gradient disappearance.

7. An underwater image enhancement system based on vision-text fusion, used to implement the underwater image enhancement method based on vision-text fusion according to any one of claims 1 to 6, characterized in that: include: Model building module, used to build an enhanced network based on visual-text fusion; The image enhancement module inputs the underwater image to be enhanced into the enhancement network to obtain an underwater enhanced image.

8. The underwater image enhancement system based on vision-text fusion according to claim 7, characterized in that: The model building module includes: Image acquisition module, which acquires any underwater image and any natural image under water conditions; The image processing module inputs the underwater image into a first generator and generates an enhanced image by inputting the first text prompt as a condition; inputs the natural image into a second generator and generates a virtual underwater image by inputting the second text prompt as a condition; An image reconstruction module inputs the enhanced image into a second generator, and inputs the second generator with the second text prompt as a condition to obtain a reconstructed underwater image; and inputs the virtual underwater image into the first generator, and inputs the first text prompt as a condition to obtain a reconstructed natural image; Wavelet transform module, which decomposes the enhanced image and virtual underwater image into low-frequency components that retain the image color distribution and overall structure information and high-frequency components that capture image details and edge textures through discrete wavelet transform; The constraint module constrains the low-frequency components through the color loss function, constrains the high-frequency components through the detail loss function, distinguishes the real image from the generated image through the adversarial loss function, constrains the consistency of the image and the reconstructed image through the cycle consistency function, and constrains the matching degree of the generator input and output through the identity loss function until the set constraints are met.

Citation Information

Patent Citations

  • Multi-type ocean data-oriented cross-modal retrieval method and system

    CN110909181A

  • Underwater image processing method and system

    CN116320344A