Deep-sea underwater sequential image enhancement method and device, electronic equipment and storage medium
By employing a deep-sea underwater sequence image enhancement method that combines physical models and deep learning networks, the problems of low image clarity, color cast, and regional degradation in deep-sea environments have been solved, achieving improvements in color and clarity, and ensuring image consistency and the accuracy of scientific research.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT MARINE DATA & INFORMATION SERVICE
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-01
AI Technical Summary
Underwater image acquisition in deep-sea environments suffers from scattering, absorption, and overexposure, resulting in low image clarity, color cast, and regional degradation, which affects the accuracy of scientific research and target detection.
The target object is segmented by extracting strong prior information, and degraded areas are repaired by estimating global background light and color line prior theory using a physical model. The image is enhanced by frequency domain decomposition and fusion using the FUnIE-GAN network and the suirSIR network. Finally, color consistency is adjusted by perceptual hashing algorithm.
It enhances the color, illumination, and sharpness of deep-sea underwater sequence images, maintains color consistency and temporal smoothness of the sequence images, and improves the accuracy of scientific research and target detection.
Smart Images

Figure CN121685335B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of deep-sea exploration technology, and in particular relates to a method, device, electronic device and storage medium for enhancing deep-sea underwater sequence images. Background Technology
[0002] Deep-sea exploration is an important direction for scientific research and resource exploration. Underwater optical camera systems mounted on platforms such as ROV (Remotely Operated Vehicle) and AUV (Autonomous Underwater Vehicle) are the main tools for obtaining intuitive visual information about the deep sea. However, due to the special nature of the deep-sea environment, the images collected generally have serious quality problems, mainly manifested as: (1) Scattering effect: A large number of suspended particles in the water scatter light, resulting in foggy and blurry images with a significant decrease in contrast; (2) Absorption effect: Water selectively absorbs light wavelengths, with the red light band attenuating the fastest, resulting in severe blue-green bias and color distortion in the images; (3) Overexposure: Artificial light sources are used to compensate for insufficient underwater light, and the light source is too close to the foreground object, causing the foreground object to be overexposed and the details and colors to degrade.
[0003] Low resolution, color cast, and regional degradation in deep-sea underwater image sequences significantly impact the accuracy of subsequent scientific research, target detection, and biometric applications. Existing underwater image enhancement methods can be divided into two categories:
[0004] One approach is based on physical models. While the principles of this approach are clear, it relies on the accurate estimation of model parameters. In complex deep-sea scenarios, it is prone to estimation errors, leading to overexposure or insufficient enhancement of colors.
[0005] The second is data-driven deep learning methods. These methods can learn complex mappings, but require a large amount of paired data for training, and for sequential images, they tend to ignore temporal consistency. Summary of the Invention
[0006] In view of this, this application aims to provide a method, apparatus, electronic device and storage medium for enhancing deep-sea underwater sequence images, in order to solve at least one of the above problems.
[0007] To achieve the above objectives, the technical solution of this application is implemented as follows:
[0008] Firstly, this application provides a method for enhancing deep-sea underwater sequence images, including:
[0009] Strong prior information is extracted from the acquired original deep-sea underwater sequence images and the target body is segmented. Degraded regions are detected based on the segmented target body regions.
[0010] The global background light is estimated from the edge region of the degraded area using a physical model, and the accuracy of the transmittance estimation is improved by using the color line prior theory. The degraded area is repaired based on the optimized transmittance to obtain the restored deep-sea underwater sequence image.
[0011] Deep-sea underwater image enhancement is performed on the restored deep-sea underwater sequence images by combining the FUnIE-GAN network and the suirSIR network to obtain enhanced deep-sea underwater sequence images; the deep-sea underwater image enhancement is based on frequency domain decomposition and fusion of the output results of the FUnIE-GAN network and the suirSIR network.
[0012] By evaluating the color richness of the enhanced deep-sea underwater sequence images, the image with the highest color richness is selected as the seed image. A similarity matching algorithm is used to obtain a set of images that meet the similarity matching conditions. The seed image is used as the source image, and the set of images is used as the target image for color transfer, so as to output the deep-sea underwater sequence images with color consistency adjustment.
[0013] Secondly, based on the same inventive concept, this application also provides a deep-sea underwater sequence image enhancement device, comprising:
[0014] The degradation region detection and segmentation module is configured to extract strong prior information based on the acquired original deep-sea underwater sequence images and perform target body segmentation, and perform degradation region detection based on the segmented target body region;
[0015] The degradation region restoration module is configured to estimate the global background light from the edge region of the degradation region through a physical model, improve the estimation accuracy of transmittance using color line prior theory, and restore the degradation region based on the optimized transmittance to obtain the restored deep-sea underwater sequence image.
[0016] The image enhancement module is configured to perform deep-sea underwater image enhancement on the restored deep-sea underwater sequence image by jointly using the FUnIE-GAN network and the suirSIR network to obtain the enhanced deep-sea underwater sequence image; wherein, the deep-sea underwater image enhancement is based on frequency domain decomposition and fusion of the output results of the FUnIE-GAN network and the suirSIR network.
[0017] The color consistency optimization module is configured to evaluate the color richness of the enhanced deep-sea underwater sequence images, select the image with the highest color richness as the seed image, perform similarity matching through a perceptual hash algorithm to obtain a set of images that meet the similarity matching conditions, use the seed image as the source image, and use the set of images as the target image for color transfer, so as to output the deep-sea underwater sequence images with adjusted color consistency.
[0018] Thirdly, based on the same inventive concept, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect.
[0019] Fourthly, based on the same inventive concept, this application also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method as described in the first aspect.
[0020] Compared with existing technologies, the deep-sea underwater sequence image enhancement method, apparatus, electronic device, and storage medium described in this application have the following advantages:
[0021] The deep-sea underwater sequence image enhancement method described in this application detects and restores degraded regions in the sequence with high quality. It also combines the FUnIE-GAN network and the suirSIR network to enhance the color, illumination, and sharpness of deep-sea underwater sequence images, while maintaining the color consistency of the sequence images and ensuring the temporal smoothness of the output sequence. Attached Figure Description
[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0023] Figure 1 This is a flowchart of a deep-sea underwater sequence image enhancement method according to an embodiment of this application;
[0024] Figure 2 This is a schematic diagram of the suirSIR network structure described in the embodiments of this application;
[0025] Figure 3 This is a schematic diagram of the temporal image color consistency optimization process described in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of the structure of a deep-sea underwater sequence image enhancement device according to an embodiment of this application;
[0027] Figure 5 This is a schematic diagram of the hardware structure of the electronic device described in an embodiment of this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0029] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0030] The embodiments of this application are described in detail below with reference to the accompanying drawings.
[0031] Please see Figure 1 As shown in the figure, this embodiment provides a method for enhancing deep-sea underwater sequence images, which specifically includes the following steps:
[0032] Step S1: Extract strong prior information from the acquired original deep-sea underwater sequence image and segment the target body, and perform degradation region detection based on the segmented target body region.
[0033] Specifically, in this embodiment, this step segments the foreground and background regions by combining strong prior guidance and zero-shot deep learning methods, with the foreground region being the target object's body region. The specific steps are as follows:
[0034] Step S11: Strong Prior Information Generation. Background subtraction is performed on consecutive frame images using optical flow to obtain weak foreground regions. Morphological erosion is then performed on these weak foreground regions to obtain a binary image of the strong foreground region; where the foreground region is set to 1 and all others to 0.
[0035] Step S12: Foreground and background segmentation. Input the image into U... 2 -Net, in the model trained on the DUTS dataset, yields U 2 The predicted probability map of -Net is used to construct univariate and binary potentials based on the predicted probability map and strong foreground regions. The CRF algorithm is then used to obtain the final foreground segmentation result.
[0036] Univariate potential energy primarily represents the probability that a pixel belongs to the foreground or background, and the formula is:
[0037] ;
[0038] In the formula, One pixel in the image; This is a priori diagram; For U 2 -Net's predicted probability graph; It is a univariate potential energy; and This is a regularization parameter used to balance the influence of the prediction probability map and the strong foreground region.
[0039] Binary potential energy primarily represents the relationship between pixels and is used to smooth segmentation results. The formula is:
[0040] ;
[0041] In the formula, It is a binary potential energy; , The first The and the first 1 pixel; This is a regularization parameter used to control the weights of the binary potential energy; and These are the standard deviations of color and spatial distance, respectively.
[0042] The total energy function is the weighted sum of the univariate potential energy and the bivariate potential energy, and the formula is:
[0043] ;
[0044] In the formula, It is the set of segmentation labels for all pixels.
[0045] In this embodiment, the degradation area mainly refers to the overexposed area on the target object's body. Overexposed areas are prone to appear when the target object's body occupies a small proportion of the entire image. The overexposed area detection method is as follows:
[0046] Step S13: Select seed point. Select the front of the target object within the target object's body region. The pixel with the highest brightness is selected as the seed point to obtain its RGB value.
[0047] Step S14: Adaptive tolerance calculation. Construct a color complexity index based on the number and diversity of colors, and determine the tolerance based on the color complexity index.
[0048]
[0049] ;
[0050] ;
[0051] In the formula, For color complexity; and These are weighting coefficients, satisfying... ; This represents the total number of pixels within the target object's body region. For the number of colors; For color diversity; This represents the total number of color types. For color sequence numbers; For the first The probability of a color appearing; For tolerance; and These are the minimum and maximum tolerance boundaries for an 8-bit image (channel values 0-255). and They are 5 and 50 respectively.
[0052] Step S15: Neighborhood color similarity evaluation. For each seed point, evaluate the color similarity of its 8 neighboring pixels using the following formula. Pixels that meet the criteria are marked as selected.
[0053] ;
[0054] In the formula, , , These are the RGB values of the seed point, , , These are the RGB values of the neighboring pixels.
[0055] Neighborhood color similarity is evaluated for all seed points until all reachable, color-similar pixels originating from the seed points are detected, thus obtaining the degradation region.
[0056] Step S2: Estimate the global background light from the edge region of the degraded area using a physical model, and improve the estimation accuracy of transmittance using color line prior theory. Repair the degraded area based on the optimized transmittance to obtain the restored deep-sea underwater sequence image.
[0057] Specifically, in this embodiment, a quadtree-based search method is used to estimate the global background light from the edge region of the degradation region, and color line prior theory is used to improve the accuracy of transmittance estimation, thus performing preliminary restoration of the degradation region. The specific steps are as follows:
[0058] Step S21: Background light estimation. Using the degradation region as the search area, perform quadtree decomposition. Each time, divide the region into 4 equal blocks. Calculate the average brightness and color standard deviation for each block. Calculate the score based on the average brightness and average variance. The highest score is used as the new search area for iteration until the block size is less than 10 pixels. Then, calculate the background light in the final block.
[0059] ;
[0060] ;
[0061] In the formula, S represents the score; This is the mean calculation function; The image is divided into blocks; As a weighting factor; , , These represent the standard deviations in the red, green, and blue bands, respectively. As background light; , , These are the red, green, and blue bands of the final block, respectively.
[0062] Step S22: Transmittance estimation. Combining the dark channel prior and the color line prior, the transmittance is initially estimated. Then, guided filtering is used to smooth and optimize the estimated transmittance map while preserving the edges.
[0063] First, determine the dark channel image, and then calculate the dark channel blurred transmittance based on the dark channel image:
[0064] ;
[0065] In the formula, Dark channel image; For dark channel fuzzy transmittance. To retain the factor, this embodiment uses 0.95.
[0066] The prior blur projection rate of the color lines is obtained by calculating the projection of the observed values onto the vector formed by the background light using the color line prior.
[0067] ;
[0068] In the formula, For the a priori fuzzy transmittance of the color line; Projecting to the center pixel ε is the maximum projection, and ε is the regularization parameter, which is set to 0.0001 here;
[0069] The estimated transmittance is obtained by linearly fusing the dark channel fuzzy transmittance and the color line prior fuzzy transmittance:
[0070] ;
[0071] In the formula, To estimate transmittance, For experience weights.
[0072] Using the original image as the guide image and the estimated transmittance as the input image, guided filtering is used to optimize the estimated transmittance image to obtain the optimized transmittance. .
[0073] Step S23: Degraded Area Repair. Based on the optimized transmittance, the degraded area is repaired to obtain the restored image.
[0074] ;
[0075] In the formula, The restored image; For the original image, To calculate the maximum value function; This is the amplitude limiting function.
[0076] Step S3: Deep-sea underwater image enhancement is performed on the restored deep-sea underwater sequence image by combining the FUnIE-GAN network and the suirSIR network to obtain the enhanced deep-sea underwater sequence image; wherein, the deep-sea underwater image enhancement is based on frequency domain decomposition and fusion of the output results of the FUnIE-GAN network and the suirSIR network.
[0077] Specifically, in this embodiment, deep-sea images are taken in an extreme environment with almost no natural light, relying entirely on artificial light sources. Due to the characteristics of artificial light sources, the images exhibit obvious non-uniform lighting, with the center of the image being bright and the surrounding area being dark, and bright spots are easily formed. In addition, due to severe red light attenuation, the color distortion of the images is more severe compared to shallow sea areas.
[0078] To address image color cast and exposure issues, the FUnIE-GAN network offers good color enhancement for deep-water images but suffers from low sharpness, while the suirSIR network enhances sharpness but introduces noise. This embodiment combines the FUnIE-GAN and suirSIR networks for underwater image enhancement, fusing their outputs in the frequency domain to achieve effective underwater image enhancement. The FUnIE-GAN network is an adversarial generative underwater image enhancement network. Its core structure consists of a generator, a discriminator, and a loss function. The generator is a fully convolutional network with a U-Net structure, the discriminator is a PatchGAN classifier, and the loss function is a combination of adversarial loss, pixel-level loss, and perceptual loss. It boasts fast inference speed and can effectively correct color distortion and uniform illumination attenuation in underwater images. The suirSIR network is a self-guided underwater image restoration network based on Retinex decomposition. Its core structure consists of three parts: a decomposition network (SIR-Net), a self-guided high-frequency enhancement subnet (SGE), and a hybrid loss function. The decomposition network uses a dual-branch U-Net to estimate illumination and reflectivity separately. The SGE subnet injects local details through gradient self-guidance. The loss function combines reconstruction loss, gradient loss, Retinex consistency, and weak color regularization, with a moderate number of parameters. It can explicitly improve the texture sharpness of underwater images while simultaneously suppressing chromatic aberration. It has good scalability for 4K high-resolution scenes. See the attached diagram for the network structure. Figure 2 .
[0079] Because the FUnIE-GAN network's loss function is only L1 / VGG, and VGG deep features are insensitive to "fine textures," resulting in missing gradient terms; furthermore, the decoder's upsampling kernel cannot synthesize high-frequency information from high-resolution images, and the model structure causes weak texture contrast to be treated as statistical noise and erased, while the generator lacks the motivation to preserve large-scale edges. These issues result in ideal global color enhancement of images when using the FUnIE-GAN network for inference, but with blurred details, loss of foreground edge details, and insufficient image clarity. Additionally, due to the weak gray-world regularization of suirSIR, the failure of the Retinex hypothesis, and the lack of a mask in the SGE subnet, these issues lead to ideal image clarity and detail enhancement when using suirSIR for inference, but with artificial textures injected into uniform water areas, resulting in green / purple spot noise. This embodiment addresses this by using frequency domain fusion to segment the model's output image into the frequency domain and then fusing them to complement each other's strengths, resulting in an enhanced underwater image.
[0080] Step S31: Model training and inference based on open-source underwater data. Due to the lack of real and ideal deep-sea underwater image sample data, this embodiment uses the EUVP and UIEB open-source underwater datasets to train the network to obtain the best model, and uses the best model to infer the results of deep-sea underwater sequence images.
[0081] Step S311: Model Training. The EUVP and UIEB open-source underwater samples are divided into training and validation sets at an 8:1 ratio. These sets are then input into the FUnIE-GAN and suirSIR network models, respectively, for training. The Adam optimizer is used, with an initial learning rate of 2e-4 and 100 epochs. The optimal models for both networks are obtained through repeated training.
[0082] Step S312, Model Inference. The input image is cropped into 512×512 blocks with a 50% overlap between blocks. The position of each block on the entire image is recorded. The best model is then used for inference. The output results are then stitched together based on the position of each block to obtain the output images of the FUnIE-GAN network model and the suirSIR network model.
[0083] Step S313, Spatial Transformation. The image output by inference is converted to Lab space. Since the human eye is most sensitive to brightness details, frequency domain fusion is performed on the brightness channel to avoid mismatches and artifacts caused by directly fusing the color channels.
[0084] Step S32: Frequency Domain Decomposition and Fusion. The luminance channel is converted to the frequency domain using Fourier transform. A low-pass filter is created to extract the low-frequency components of the FUnIE-GAN network output image and the high-frequency components of the suirSIR network output image. The two frequency domain components are then fused. After fusion, the frequency domain components are converted back to the spatial domain to obtain the fused luminance channel. The specific formula is as follows:
[0085] ;
[0086] in, This is the merged luminance channel. and These are the Fast Fourier Transform and the Inverse Fast Fourier Transform, respectively. and These are the brightness channels of the output images from the FUnIE-GAN network and the suirSIR network, respectively. For transfer functions, and These represent frequencies in the horizontal and vertical directions, respectively. This is the pixel-level multiplication symbol.
[0087] Step S33: Brightness channel optimization based on smooth gradient map. Noise is introduced when using the high-frequency part of the output image of the suirSIR network. Therefore, it is necessary to use the details of the texture-rich areas to suppress the noise in the smooth areas.
[0088] Step S331: Calculate the gradient magnitude of the brightness channel of the suirSIR output image using the Sobel operator to obtain the gradient map, and then apply Gaussian blur to the gradient map to obtain a smooth gradient map.
[0089] ;
[0090] In the formula, For smooth gradient plots, and These are the Soble gradient operators for the horizontal and vertical directions, respectively. For Gaussian blur operator, , These represent the size and standard deviation of the Gaussian blur kernel, respectively.
[0091] Step S332, Luminance Channel Optimization. The fused luminance channel and the smooth gradient map are fused using pixel-level weighted fusion to suppress noise in the smooth areas.
[0092] ;
[0093] in, For the optimized luminance channel, This is the pixel-level multiplication symbol.
[0094] Step S34: Enhance the output result. Because the color effect of the FUnIE-GAN network output image is good, the a and b channels of the final enhanced image are the same as those of the FUnIE-GAN network output image, and the L channel is the optimized luminance channel. The image is then converted back to the RGB color space to obtain the enhanced deep-sea underwater sequence image.
[0095] Step S4: By evaluating the color richness of the enhanced deep-sea underwater sequence images, the image with the highest color richness is selected as the seed image. A similarity matching is performed using a perceptual hash algorithm to obtain a set of images that meet the similarity matching conditions. The seed image is used as the source image, and the set of images is used as the target image for color transfer, so as to output the deep-sea underwater sequence images with color consistency adjustment.
[0096] Specifically, in this embodiment, since the enhancement is performed on single images, there may be color differences between different scenes, which can generate noise during 3D modeling. Therefore, the color consistency of the time-series images needs to be considered. First, the processed sequence of images is color-evaluated, and the image with the richest color is selected as the seed image. Then, the remaining images are compared with the seed image for similarity matching. Using the seed image as a template image, color transfer is performed on the set of images that meet the similarity matching conditions. The image with the lowest similarity in the sequence of images is selected as the new seed image, and a new round of color transfer is performed until all images have been processed, thus unifying the color effect of the time-series images. The entire process is shown in the appendix. Figure 3 .
[0097] Step S41: Seed Image Selection. The color richness of the repaired and enhanced time-series images is evaluated, and the image with the highest color richness value is selected as the seed image.
[0098]
[0099] ;
[0100]
[0101] ;
[0102] ;
[0103] In the formula, For color richness; , , These are the red, green, and blue bands of the enhanced image, respectively. , and , They are respectively and The variance and mean, To account for the degree of dispersion, The overall average intensity.
[0104] Step S42, Similarity Matching. When detecting image similarity, content similarity is more important than color similarity. Therefore, this embodiment uses a perceptual hashing algorithm to perform similarity matching calculations and obtain the image sequence that is color-transferred to the seed image.
[0105] Step S421: Convert other images to grayscale images, resample them to a size of 32×32 pixels, and then convert the images to grayscale.
[0106] Step S422: Perform Discrete Cosine Transform (DCT) on the resampled image. The operation is completed by calling the cv2.dct() function in OpenCV to obtain the DCT coefficient matrix.
[0107] Step S423: Take the 8×8 region in the upper left corner of the DCT coefficient matrix, calculate the mean, and use it as the hash value baseline. Compare the 64 coefficients in the 8×8 region with the hash value baseline. If the value is less than the hash value baseline, record it as 0, and the rest as 1, to obtain a 64-bit binary hash value.
[0108] Step S424: Similarity Matching. Calculate the Hamming distance between the hash values of the remaining images and the seed image to measure their similarity to the seed image. If the Hamming distance is less than or equal to 5, the image with the matching hash is considered to be similar to the seed image. Repeat this step to obtain a set of images that meet the similarity matching criteria. The Hamming distance formula is as follows:
[0109] ;
[0110] In the formula, The Hamming distance between a given image and the seed image from the remaining images. and These are the hash values of a given image and the seed image, respectively. The 64 coefficient indices are located at the top left corner of the DCT coefficient matrix.
[0111] Step S43, Color Transfer. Using the seed image as the source image, the set of images that meet the similarity matching conditions are used as the target images, and color transfer is performed. The fuzzy C-means clustering (FCM) algorithm is used to complete the color transfer. The image with the largest D in the set of images that meet the similarity matching conditions is used as the seed image for the new round. The similarity matching is repeated to complete the color transfer of all images, ensuring the color consistency of the time-series images.
[0112] The deep-sea underwater sequence image enhancement method described in this embodiment detects and restores degraded regions in the sequence with high quality. It also combines the FUnIE-GAN network and the suirSIR network to enhance the color, illumination, and sharpness of deep-sea underwater sequence images, while maintaining the color consistency of the sequence images and ensuring the temporal smoothness of the output sequence.
[0113] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0114] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, the embodiments of this application also provide a deep-sea underwater sequence image enhancement device.
[0115] like Figure 4 As shown, the deep-sea underwater sequence image enhancement device includes:
[0116] The degradation region detection and segmentation module 11 is configured to extract strong prior information based on the acquired original deep-sea underwater sequence image and perform target body segmentation, and perform degradation region detection based on the segmented target body region.
[0117] The degradation region repair module 12 is configured to estimate the global background light from the edge region of the degradation region through a physical model, improve the estimation accuracy of transmittance using color line prior theory, and repair the degradation region based on the optimized transmittance to obtain the restored deep-sea underwater sequence image.
[0118] The image enhancement module 13 is configured to perform deep-sea underwater image enhancement on the restored deep-sea underwater sequence image by jointly using the FUnIE-GAN network and the suirSIR network to obtain the enhanced deep-sea underwater sequence image; wherein, the deep-sea underwater image enhancement is based on frequency domain decomposition and fusion of the output results of the FUnIE-GAN network and the suirSIR network.
[0119] The color consistency optimization module 14 is configured to evaluate the color richness of the enhanced deep-sea underwater sequence images, select the image with the highest color richness as the seed image, perform similarity matching through a perceptual hash algorithm to obtain a set of images that meet the similarity matching conditions, use the seed image as the source image and the set of images as the target image for color transfer, and output the deep-sea underwater sequence images with adjusted color consistency.
[0120] For ease of description, the above apparatus is described in terms of its functions, divided into various modules. Of course, in implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.
[0121] The apparatus of the above embodiments is used to implement the corresponding method in any of the foregoing embodiments and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0122] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the methods described in any of the above embodiments.
[0123] Figure 5 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0124] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0125] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0126] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices may include displays, speakers, vibrators, indicator lights, etc.
[0127] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0128] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0129] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0130] The electronic devices described above are used to implement the corresponding methods in any of the foregoing embodiments and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0131] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to perform the methods described in any of the above embodiments.
[0132] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0133] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the methods described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0134] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0135] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0136] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A method for enhancing deep-sea underwater sequence images, characterized in that, include: Strong prior information is extracted from the acquired original deep-sea underwater sequence images and the target body is segmented. Degraded regions are detected based on the segmented target body regions. The global background light is estimated from the edge region of the degraded area using a physical model, and the accuracy of the transmittance estimation is improved by using the color line prior theory. The degraded area is repaired based on the optimized transmittance to obtain the restored deep-sea underwater sequence image. Deep-sea underwater image enhancement is performed on the restored deep-sea underwater sequence images by combining the FUnIE-GAN network and the suirSIR network to obtain enhanced deep-sea underwater sequence images; the deep-sea underwater image enhancement is based on frequency domain decomposition and fusion of the output results of the FUnIE-GAN network and the suirSIR network. By evaluating the color richness of the enhanced deep-sea underwater sequence images, the image with the highest color richness is selected as the seed image. A similarity matching algorithm is used to obtain a set of images that meet the similarity matching conditions. The seed image is used as the source image, and the set of images is used as the target image for color transfer, so as to output the deep-sea underwater sequence images with color consistency adjustment.
2. The method according to claim 1, characterized in that, The target object body segmentation includes: Background subtraction is performed on consecutive frames of images to obtain weak foreground regions, and morphological erosion is performed on the weak foreground regions to obtain a binary image of strong foreground regions. The image is input into a zero-shot deep learning model to obtain a predicted probability map. Based on the predicted probability map and the binary map of the strong foreground region, a univariate potential and a binary potential are constructed. The foreground segmentation result is obtained based on the univariate potential and the binary potential.
3. The method according to claim 1, characterized in that, The degradation region detection based on the segmented target body region includes: Seed points are selected in the target object body region, and a color complexity index is constructed by the number and diversity of colors. The tolerance is determined based on the color complexity index. The degradation region is obtained by evaluating the neighborhood color similarity of each seed point based on the determined tolerance.
4. The method according to claim 3, characterized in that: The degraded region is used as the search region for quadtree decomposition to divide it into multiple image blocks. A score is determined based on the average brightness and color standard deviation of each image block. The search region is iterated based on the score to obtain the background light of the final image block. Transmittance is initially estimated using dark channel priors and color line priors, and then smoothed and optimized using guided filtering to obtain the optimized transmittance. The degraded areas are repaired based on the optimized transmittance to obtain restored deep-sea underwater sequence images.
5. The method according to claim 1, characterized in that: The network was trained using an open-source underwater dataset to obtain the best model, and the best model was used to infer and output images from the restored deep-sea underwater sequence images. The image output by inference is transformed into Lab space, and the luminance channel is transformed into the frequency domain through Fourier transform. The low-frequency part of the FUnIE-GAN network output image and the high-frequency part of the suirSIR network output image are extracted respectively. The two frequency domain parts are fused, and the fused frequency domain part is transformed back into the spatial domain to obtain the fused luminance channel. Brightness channel optimization based on smooth gradient map.
6. The method according to claim 5, characterized in that: The gradient magnitude of the brightness channel of the suirSIR output image is determined using the Sobel operator to obtain a gradient map, and the gradient map is then Gaussian blurred to obtain a smooth gradient map. The fused luminance channel and the smooth gradient map are then subjected to pixel-level weighted fusion to obtain the optimized luminance channel.
7. The method according to claim 6, characterized in that: The a and b channels of the final enhanced image are obtained by using the a and b channels of the FUnIE-GAN network output image, and the L channel is the optimized luminance channel. The image is then converted back to the RGB color space to obtain the enhanced deep-sea underwater sequence image.
8. A deep-sea underwater sequence image enhancement device, characterized in that, include: The degradation region detection and segmentation module is configured to extract strong prior information based on the acquired original deep-sea underwater sequence images and perform target body segmentation, and perform degradation region detection based on the segmented target body region; The degradation region restoration module is configured to estimate the global background light from the edge region of the degradation region through a physical model, improve the estimation accuracy of transmittance using color line prior theory, and restore the degradation region based on the optimized transmittance to obtain the restored deep-sea underwater sequence image. The image enhancement module is configured to perform deep-sea underwater image enhancement on the restored deep-sea underwater sequence image by jointly using the FUnIE-GAN network and the suirSIR network to obtain the enhanced deep-sea underwater sequence image; wherein, the deep-sea underwater image enhancement is based on frequency domain decomposition and fusion of the output results of the FUnIE-GAN network and the suirSIR network. The color consistency optimization module is configured to evaluate the color richness of the enhanced deep-sea underwater sequence images, select the image with the highest color richness as the seed image, perform similarity matching through a perceptual hash algorithm to obtain a set of images that meet the similarity matching conditions, use the seed image as the source image, and use the set of images as the target image for color transfer, so as to output the deep-sea underwater sequence images with adjusted color consistency.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-7.
10. A non-transitory computer-readable storage medium, characterized in that, in, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to perform the method described in any one of claims 1-7.
Citation Information
Patent Citations
Super-resolution graph recovery method for simultaneously enhancing underwater images
CN111882489A
Underwater scene reconstruction and water medium separation method based on 3D Gaussian model
CN119359906A