Lightweight embedded real-time image enhancement system and method

By combining FA+Net and CycleGAN network models, the problems of color shift and texture loss in underwater image enhancement are solved, and efficient underwater image enhancement is achieved, suitable for underwater equipment with resource-constrained resources.

CN120163751APending Publication Date: 2025-06-17GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510309780.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the color shift and texture loss problems in underwater image enhancement, and traditional deep learning models are difficult to efficiently realize image enhancement on underwater equipment with resource-constrained resources.

Method used

Combining the FA+Net and CycleGAN network models, FA+Net decomposes the underwater image enhancement problem into a sub-problem, correcting color distortion and restoring details respectively, while the CycleGAN network is used for style conversion and detail optimization to solve the problem of poor visual effects.

Benefits of technology

It realizes efficient color recovery and detail texture recovery of underwater images, improves the visual effect and performance indicators of the image, and is suitable for underwater equipment with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163751A_ABST
    Figure CN120163751A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight embedded real-time image enhancement system and method, and relates to the technical field of image processing. The lightweight embedded real-time image enhancement system comprises an FA + Net network model which is a lightweight convolutional neural network and comprises a strong prior stage and a fine granularity stage; the strong prior stage utilizes a physical model and prior knowledge to correct light scattering and absorption of an underwater image, and the fine granularity stage enhances details and colors of the underwater image through a multi-branch structure; and the CycleGAN network model is a loop consistent antagonism network and comprises a generator and a discriminator, the generator is used for generating an output image of the FA + Net network model into an image with a required style, and the discriminator is used for performing detail optimization on the generated image. The problem that the visual effect is poor after an existing underwater image is enhanced is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a lightweight embedded real-time image enhancement system and method. Background Art

[0002] Underwater images are often plagued by severe blurring and color distortion problems, making it difficult for them to meet the requirements of practical applications. With the successful application of deep learning in advanced computer vision tasks, more and more researchers have started to apply it to underwater image enhancement, such as the WaterNet model proposed by Li et al. in 2019. Although this model has achieved a certain degree of success in terms of performance, this method fails to establish a dedicated module to address the color bias and texture loss problems of degraded images. Subsequently, in 2021, Huo et al. proposed using a wavelet enhanced learning unit to decompose hierarchical features into high-frequency and low-frequency components, and then enhancing them through normalization and attention mechanization. Although this method shows excellent visual effects, its extensive network parameters (6.30M) and computational requirements (223.37G) make it unsuitable for existing underwater devices. Moreover, it cannot effectively solve the problem of color distortion. Limited by the resources of underwater robots and cameras, traditional deep learning models face challenges in improving the efficiency of underwater image enhancement on these devices. Although the Shallow-uwnet proposed by Ankita Naik et al. in 2021 constructs a streamlined network architecture by integrating lightweight network components and residual convolutional blocks, this resource-optimization-driven method does not always ensure a reduction in computational complexity. At the same time, in the field of underwater image enhancement, there is a lack of specialized designs for specific degradation problems, which limits the performance of restored images in terms of visual effects and performance metrics.

[0003] The present invention proposes a lightweight embedded real-time image enhancement system and method, specifically an underwater image enhancement method that combines the network model FA+Net with the network model CycleGAN. The network model FA+Net solves the underwater image enhancement problem by decomposing it into sub-problems, and effectively solves the mixed degradation problem by correcting color distortion and restoring the details of degraded images respectively. The CycleGAN network model is mainly used for style conversion and detail optimization to address the problem of poor visual effects after existing underwater image enhancement. Summary of the Invention

[0004] The purpose of the present invention is to provide a lightweight embedded real-time image enhancement system and method to solve the problems such as poor visual effects after underwater image enhancement in the prior art as described in the above background art.

[0005] To achieve the above purpose, the present invention is implemented by adopting the following technical solutions:

[0006] In a first aspect, the present invention proposes a lightweight embedded real-time image enhancement system, including:

[0007] An input module for acquiring underwater images;

[0008] The FA+Net network model, which is a lightweight convolutional neural network, including a strong prior stage and a fine-grained stage; in the strong prior stage, color distortion correction and detail restoration are performed on the underwater image, and light scattering and absorption of the underwater image are corrected using physical models and prior knowledge. In the fine-grained stage, details and colors of the underwater image are enhanced through a multi-branch structure;

[0009] The CycleGAN network model, which is a cycle-consistent adversarial network, including a generator and a discriminator. The generator is used to generate an image with a desired style from the output image of the FA+Net network model, and the discriminator is used to optimize the details of the generated image;

[0010] An output module for outputting the enhanced underwater image to a display or other devices.

[0011] Preferably, the strong prior stage includes a multi-branch color enhancement module and a multi-scale pyramid module;

[0012] The multi-branch color enhancement module is used to capture the color feature distribution of the R, G, and B channels by adopting a branch enhancement strategy for color restoration of the underwater image;

[0013] The multi-scale pyramid module is used to capture detail information at different scales to enhance the perception of underwater image details.

[0014] Preferably, the multi-branch color enhancement module divides the input feature map into four feature maps of different sizes, then uses 1×1 convolution and normalization layers to continuously process individual image pixels, and the four processed feature maps are merged back into one feature map through fusion splicing.

[0015] Preferably, the multi-scale pyramid module first extracts local features from the input feature map through a 3×3 convolutional layer. The convolutional feature map is divided into three feature maps of different scales through downsampling. Each scale of the feature map passes through a point convolution. The feature map after the point convolution is upsampled back to the original size, and the upsampled feature maps are fused through an addition operation; the fused feature map passes through a 3×3 convolutional layer again, and finally passes through a Tanh activation function to output the feature map.

[0016] Preferably, the FA+Net network model further includes a pixel convolutional layer and a spatial frequency domain interaction module;

[0017] The pixel convolutional layer is arranged before the strong prior stage and is used to extract basic features of the underwater image;

[0018] The spatial frequency domain interaction module is set after the strong prior stage and is used to screen valuable feature information in the feature map and obtain global context features;

[0019] The feature maps obtained by the multi-scale pyramid module and the feature maps obtained by the multi-branch color enhancement module are both used as inputs to the spatial frequency domain feature interaction module for feature addition and fusion to obtain the original fused feature map. Apply the fast Fourier transform to the original fused feature map to obtain its representation in the frequency domain, then perform frequency domain convolution on the amplitude spectrum and phase spectrum respectively. The processed amplitude spectrum and phase spectrum are converted back to the spatial domain through the inverse fast Fourier transform to obtain the processed feature map; finally, the processed feature map and the original fused feature map are combined through weighted fusion to obtain the final output feature map.

[0020] Preferably, the fine-grained stage includes a multi-branch color enhancement module and a pixel attention mechanism;

[0021] The multi-branch color enhancement module is used to adjust the color balance of the image and enhance color contrast and saturation;

[0022] The pixel attention mechanism is used to improve the ability to restore image details and enhancement effect by adaptively adjusting the degree of attention to different pixel regions;

[0023] The output image of the spatial frequency domain interaction module is successively processed by point convolution, the multi-branch color enhancement module, and the pixel attention mechanism, and finally the output image of the FA+Net network model is obtained.

[0024] Preferably, the generator includes an encoder and a decoder. The encoder encodes the input image into a latent space representation to capture the high-level features of the image; the decoder decodes the latent space representation back to the image space to generate an output image with enhanced details and style;

[0025] The discriminator includes an input layer, multiple convolutional layers, a fully connected layer, and an output layer; the input layer is used to receive the input image, and the multiple convolutional layers are used to process the input image through a series of convolution, batch normalization, and Leaky ReLU activation operations to gradually extract image features to obtain a feature map; the fully connected layer is used to map the feature map to a one-dimensional vector through the fully connected layer after dimensionality reduction of the feature map; the output layer is used to output a probability value through the sigmoid activation function to judge whether the image comes from the true data distribution.

[0026] Preferably, through the adversarial training of the generator and the discriminator, CycleGAN can generate high-quality image conversion results.

[0027] Preferably, there are two generators and discriminators, namely generator G, generator F, discriminator Dx, and discriminator Dy respectively;

[0028] The output image x of the FA+Net network model is used as the dataset X of the CycleGAN network model, and the clear underwater image y is used as the dataset Y of the CycleGAN network model. The content of dataset X and the content of dataset Y are not related;

[0029] Generator G generates an image y' with a style similar to y from the output image x, and the image x belongs to dataset X, that is, G(x)=y', x∈X; Generator F generates an image x' with a style similar to x from the clear underwater real image y, and the image y belongs to dataset Y, that is, F(y)=x', y∈Y;

[0030] If the content gap between the image y' generated by generator G and the image y in dataset Y is large or the image edge is blurred, discriminator Dy outputs a low score; conversely, if the image x' generated by generator F is similar in content and clear compared to the image x in dataset X, discriminator Dx outputs a high score;

[0031] After n times of training, the CycleGAN network model outputs an image with a style close to dataset Y and a content close to dataset X.

[0032] In the second aspect, the present invention discloses a lightweight embedded real-time image enhancement method, including the following steps:

[0033] S1. Construct an FA+Net network model and use the FA+Net network to preliminarily enhance the underwater image;

[0034] S2. Construct a CycleGAN network model and use the CycleGAN network to perform underwater image style conversion and detail optimization;

[0035] S3. Design a loss function to optimize the network model to obtain the lightweight embedded real-time image enhancement system as described in any one of claims 1-7;

[0036] S4. Use the lightweight embedded real-time image enhancement system to perform image enhancement on underwater image data for style conversion and detail optimization.

[0037] Preferably, the optimization of the network model is specifically as follows:

[0038] The FA+Net network model optimizes the feature aggregation and image enhancement parts through image quality evaluation metrics. The number of parameters of the FA+Net network model is optimized to adapt to resource-constrained mobile platforms, ensuring efficient operation in environments with low power consumption and limited computing capabilities. It is suitable for real-time operation in resource-constrained devices such as underwater robots and unmanned underwater vehicles.

[0039] The loss function of the CycleGAN network model consists of adversarial loss and cycle consistency loss, and the network parameters are optimized based on the loss function. The CycleGAN network model is trained to adapt to specific underwater environmental conditions, such as different water depths, water qualities, and lighting conditions, to improve the adaptability and generalization ability of the enhancement effect.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] (1) The present invention can improve the color restoration ability of underwater images. FA+Net is good at extracting multi-scale features of images, which is very effective for color correction and information extraction of underwater images, and can restore the red, green, and blue spectra absorbed by water bodies. CycleGAN has the ability of generative adversarial learning, which can transform underwater images through unsupervised learning, reduce color distortion and lighting problems in images, and improve the color restoration degree of images. The combination of these two networks can more accurately restore the true colors and details of underwater images, especially in underwater environments with severe light attenuation and color shift.

[0042] (2) The present invention can enhance image detail and texture restoration. The feature aggregation strategy of FA+Net can effectively fuse features from different scales and restore details in images, especially the blur and noise problems common in underwater images. CycleGAN can effectively remove noise in images through generative adversarial training, especially the particles and blur in underwater images, and restore detailed textures. Through the feature extraction of FA+Net and the adversarial training of CycleGAN, the global structure and local details of the image can be restored simultaneously, making the image clearer and more hierarchical.

[0043] (3) The present invention can automatically repair geometric distortions and background noise in underwater images. When aggregating features, FA+Net can pay attention to the geometric shapes and structures in the image, thereby repairing geometric distortions caused by factors such as water flow and light refraction. CycleGAN can generate a more natural background that conforms to the underwater environment through training a generative adversarial network, reducing background noise caused by suspended matter or lighting conditions. After the combination of the two, not only can geometric distortions be repaired, but also the interference caused by environmental noise can be reduced, significantly improving the quality of the image.

[0044] (4) The present invention has high efficiency and real-time performance. Through effective feature aggregation and processing strategies, FA+Net can maintain a low computational complexity while ensuring high-quality enhancement. CycleGAN reduces the demand for computing resources by generating adversarial learning without directly processing each pixel. This combination can achieve efficient underwater image enhancement with limited computing resources and is applicable to real-time image processing scenarios such as underwater robots and unmanned underwater vehicles. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 It is a structural block diagram of the lightweight embedded real-time image enhancement system in the present invention;

[0046] Figure 2 It is a structural block diagram of the multi-branch color enhancement module in the present invention;

[0047] Figure 3 It is a structural block diagram of the multi-scale pyramid module in the present invention;

[0048] Figure 4 It is a structural block diagram of the spatial frequency feature interaction module in the present invention;

[0049] Figure 5 It is a structural block diagram of the CycleGAN network model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0051] Embodiment 1:

[0052] The lightweight embedded real-time image enhancement method in the present invention adopts an image processing model, and the network architecture diagram of this model is as Figure 1 shown. This model consists of two main parts: FA+Net and CycleGAN. The present invention improves the original FA+Net network: adding a cyclic-consistent adversarial network CycleGAN part to obtain better-quality images. The underwater image enhancement method that combines FA+Net and CycleGAN can integrate the advantages of both, restore the color, details, and texture of underwater images, and solve problems such as blurring, noise, and color difference in underwater images. Through unsupervised learning, generative adversarial training, and feature aggregation, the quality of underwater images can be improved efficiently and accurately, and it is applicable to multiple fields such as underwater archaeology, marine ecological monitoring, and automated underwater navigation.

[0053] The lightweight embedded real-time image enhancement method includes the following steps:

[0054] Step 1: Construct the FA+Net network and use the FA+Net network to preliminarily enhance the image.

[0055] Receive the underwater image, obtain the original underwater image to be enhanced, and use the FA+Net network model to preliminarily enhance the image. The FA+Net network model solves the underwater image enhancement problem by decomposing it into sub-problems, and effectively solves the mixed degradation problem by correcting color distortion and restoring the details of the degraded image respectively. As Figure 2 shown, the overall architecture of FA+Net consists of a powerful early stage and a refinement stage, supplemented by an efficient spatial frequency domain interaction module.

[0056] For the powerful early stage, two complementary components are proposed: the multi-branch color enhancement module (MCEM) and the multi-scale pyramid module (MPM). MCEM is an effective module for solving the serious color distortion of underwater images. MCEM processes individual image pixels continuously, enabling the network to achieve precise color restoration. MPM supports processing the input feature map at multiple scales, thereby capturing detailed information at different scales to enhance the model's perception of image details. The spatial frequency domain interaction module (SDFIM) filters valuable feature information from the outputs of different components and obtains global context features. In this way, the powerful prior-based design endows the network with efficient underwater degradation recovery ability. Then, a fine-grained stage is added to the model, and the MCEM and Pixel Attention mechanism are incorporated to enhance image details and colors, thereby improving the performance and generalization ability of the model.

[0057] The multi-scale pyramid module MPM in the FA+Net network model. MPM supports processing the input feature map at multiple scales, thereby capturing detailed information at different scales to enhance the model's perception of image details. As Figure 3 shown, it extracts and integrates features at different scales in the image processing network. This module accepts an input feature map. The input feature map first passes through a 3×3 convolutional layer to extract local features. The feature map after convolution is segmented into three different-scale feature maps through downsampling. Each scale's feature map passes through a point convolution. The feature map after point convolution is upsampled back to the original size. The upsampled feature maps are fused through an addition operation. The fused feature map passes through a 3×3 convolutional layer again to further extract and integrate features. Finally, through the Tanh activation function, the output feature map is obtained.

[0058] The multi-branch color enhancement module MCEM adopts a branch enhancement strategy to better capture the color feature distributions of the R, G, and B channels, as Figure 2 shown. The module receives an input feature map, which is segmented into four feature maps of different sizes, and then the single image pixels are continuously processed using 1×1 convolutions and normalization layers. The four processed feature maps are merged back into one feature map through a fusion concatenation.

[0059] The feature map X obtained by the multi-scale pyramid module MPM and the feature map Y obtained by the multi-branch color enhancement module MCEM are both used as inputs to the spatial frequency domain feature interaction module SDFIM for feature addition. The fast Fourier transform is applied to the fused feature map to obtain its representation in the frequency domain, and then frequency domain convolutions are performed on the amplitude spectrum and phase spectrum respectively. The processed amplitude spectrum and phase spectrum are converted back to the spatial domain through the inverse fast Fourier transform to obtain the processed feature map. Finally, the processed feature map is combined with the original fused feature map through weighted fusion to obtain the final output feature map. As Figure 4 shown.

[0060] The output feature map passes through point convolutions, and then the multi-branch color enhancement module MCEM is used again to further enhance the details of the image. Finally, the feature map is input into the Pixel Attention mechanism to pay more attention to the high-frequency parts of the image, further improving the image quality.

[0061] Step 2: Construct a CycleGAN network and use the CycleGAN network to achieve underwater image style conversion and detail optimization.

[0062] To obtain images with better perceptual quality, higher image information entropy, and less noise, clearer images can be obtained by combining the cyclic consistent adversarial network CycleGAN with FA+Net. Figure 5 shows the whole process. CycleGAN consists of two generators and two discriminators.

[0063] CycleGAN performs unsupervised learning through a generative adversarial network to generate enhanced versions of underwater images. Its basic architecture includes a generator and a discriminator. The generator consists of two parts: a downsampling encoder and an upsampling decoder. The downsampling encoder converts the input underwater image into a low-dimensional feature representation. The upsampling decoder restores the details of the underwater image by generating an image. Skip connections directly pass the low-level features of the encoder to the decoder through skip connections, helping to generate clearer images. The goal of the discriminator is to determine whether the input image comes from a real clear image or an enhanced image generated by the generator. Through adversarial training, the discriminator makes the images generated by the generator increasingly close to real underwater images. The key feature of CycleGAN is cycle consistency, that is, converting the enhanced image back to the original underwater image and ensuring its consistency with the original image. This cycle consistency enables the model to preserve the structural and semantic information of the image.

[0064] The main purpose of the Cycle-Consistent Adversarial Network is to achieve style transfer. The image x processed by FA+Net is used as the dataset X, and the clear underwater image y is used as the dataset Y. The dataset X can generate a picture M1 with the style of y and the content of x through the generator G, that is, G(x) = y', x ∈ X; similarly, the dataset Y can generate a photo M2 with the style of x and the content of y through the generator F, that is, F(y) = x', y ∈ Y. To achieve this goal, two discriminators Dx and Dy need to be trained to respectively judge the quality of the pictures generated by the two generators. After n iterations, until the style output by the generator G is close to the real underwater image and the content output by the generator F is close to the image given by the FA+Net network model. After training, a picture with a clear style and a given content is obtained.

[0065] The generator of the CycleGAN network model includes an encoder and a decoder. The encoder encodes the input image into a latent space representation, capturing the high-level features of the image. The decoder decodes the latent space representation back into the image space, generating an output image with enhanced details and style.

[0066] The discriminator in CycleGAN includes an input layer that accepts an input image; multiple convolutional layers, that is, through a series of convolutional, batch normalization, and Leaky ReLU activation processes, gradually extracting image features; a fully connected layer, that is, after reducing the dimensionality of the feature map, mapping the features to a one-dimensional vector through the fully connected layer; and an output layer, that is, the final sigmoid activation function will output a probability value to determine whether the image comes from the real data distribution. Through the adversarial training of the generator and the discriminator, CycleGAN can generate high-quality image conversion results.

[0067] The CycleGAN of the cycle-consistent adversarial network combines two different network structures, and the specific steps are as follows: The output image x obtained from the model FA+Net is used as the dataset X of the module CycleGAN, and the clear underwater image y is used as the dataset Y. The content of dataset X and the content of dataset Y are not related; the generator G uses the output image x obtained from FA+Net to generate an image y' with a style similar to y, and the image x belongs to dataset X; that is, G(x)=y', x∈X; the generator F uses the clear underwater real image y to generate an image x' with a style similar to x, and the image y belongs to dataset Y, that is, F(y)=x', y∈Y; if the image y' generated by the generator G has a large content gap with the image y in the dataset Y, or the image edge is blurred, then the discriminator Dy outputs a low score at this time; on the contrary, if the image x' is similar in content and clear compared with the image x in the dataset X, then the discriminator Dx outputs a high score at this time; after n times of training, a dataset Y with a style infinitely close to the real image and an image of dataset X with a content infinitely close to the output of the model FA+Net will be obtained.

[0068] Step 3: Design a loss function to optimize the image processing model.

[0069] During the training process, the loss function of CycleGAN consists of an adversarial loss (the adversarial training of the generator and the discriminator) and a cycle-consistency loss (maintaining the integrity of the image structure). FA+Net optimizes its feature aggregation and image enhancement parts through traditional image quality evaluation metrics (such as SSIM, PSNR).

[0070] The dataset used is UIEB, which contains 890 high-resolution original underwater images and corresponding high-quality reference images, as well as 60 challenge images C60 without corresponding reference images. It is implemented using the PyTorch framework with a single NVIDIA GTX A100 GPU (40GB). During training, the training epoch is set to 400, the total batch size is 72, and Adam is used for optimization; the learning rate starts from 10 -4 −4, the default values of β1 and β2 are 0.5 and 0.999 respectively, CyclicLR is used to adjust the learning rate, the initial momentum is 0.9 and 0.999 respectively, 800 pairs of original images and clear images are extracted from UIEB to train the model, and the remaining 90 images are used to test the impact of this method on degraded images.

[0071] Input the degraded underwater image data in the test set into the network of the present invention that has been trained and verified. The output of the network is the underwater image data enhanced by the network of the present invention. The flowchart of the underwater image enhancement data is as Figure 1 shown.

[0072] The above is only used to help understand the method of the present invention and its core concept. However, the protection scope of the present invention is not limited thereto. For those of ordinary skill in the art in the technical scope disclosed by the present invention, any equivalent substitution or change made according to the technical solution and inventive concept of the present invention should be covered within the protection scope of the present invention. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A lightweight embedded real-time image enhancement system, characterized in that: include: An input module for acquiring underwater images; The FA+Net network model is a lightweight convolutional neural network, which includes a strong priori stage and a fine-grained stage; the strong priori stage corrects color distortion and restores details of underwater images, and the fine-grained stage enhances the details and colors of underwater images through a multi-branch structure; The CycleGAN network model is a cycle-consistent adversarial network, which includes a generator and a discriminator. The generator is used to generate the output image of the FA+Net network model into an image with the desired style, and the discriminator is used to optimize the details of the generated image. The output module is used to output the enhanced underwater image.

2. A lightweight embedded real-time image enhancement system according to claim 1, characterized in that: The strong prior stage includes a multi-branch color enhancement module and a multi-scale pyramid module; The multi-branch color enhancement module is used to capture the color feature distribution of R, G and B channels using a branch enhancement strategy to perform color restoration of underwater images; The multi-scale pyramid module is used to capture detail information at different scales to enhance the perception of underwater image details.

3. A lightweight embedded real-time image enhancement system according to claim 2, characterized in that: The multi-branch color enhancement module divides the input feature map into four feature maps of different sizes, and then uses 1×1 convolution and normalization layers to continuously process a single image pixel. The processed four feature maps are merged back into one feature map through fusion and splicing.

4. The lightweight embedded real-time image enhancement system according to claim 2, characterized in that: The multi-scale pyramid module first extracts local features from the input feature map through a 3×3 convolution layer, and the feature map after convolution is divided into feature maps of three different scales through downsampling. The feature map of each scale passes through a point convolution, and the feature map after point convolution is upsampled back to the original size. The upsampled feature maps are fused through addition operations; the fused feature map passes through a 3×3 convolution layer again, and finally passes through the Tanh activation function to output the feature map.

5. The lightweight embedded real-time image enhancement system according to claim 3, characterized in that: The FA+Net network model also includes a pixel convolution layer and a spatial frequency domain interaction module; The pixel convolution layer is arranged before the strong prior stage to extract the basic features of the underwater image; The spatial frequency domain interaction module is arranged after the strong prior stage, and is used to filter valuable feature information in the feature map and obtain global context features; The feature map obtained by the multi-scale pyramid module and the feature map obtained by the multi-branch color enhancement module are both used as the input of the spatial-frequency domain feature interaction module for feature addition and fusion to obtain the original fused feature map. Fast Fourier transform is applied to the original fused feature map to obtain its representation in the frequency domain. The amplitude spectrum and phase spectrum are then convolved in the frequency domain respectively. The processed amplitude spectrum and phase spectrum are converted back to the spatial domain by inverse fast Fourier transform to obtain the processed feature map. Finally, the processed feature map is combined with the original fused feature map by weighted fusion to obtain the final output feature map.

6. A lightweight embedded real-time image enhancement system according to claim 5, characterized in that: The fine-grained stage includes a multi-branch color enhancement module and a pixel attention mechanism; The multi-branch color enhancement module is used to adjust the color balance of the image and enhance the color contrast and saturation; The pixel attention mechanism is used to improve the image detail restoration capability and enhancement effect by adaptively adjusting the degree of attention paid to different pixel regions; The output image of the spatial-frequency domain interaction module is processed successively through point convolution, multi-branch color enhancement module, and pixel attention mechanism, and finally the output image of the FA+Net network model is obtained.

7. A lightweight embedded real-time image enhancement system according to claim 6, characterized in that: The generator includes an encoder and a decoder, the encoder encodes the input image into a latent space representation, capturing high-level features of the image; The decoder decodes the latent space representation back to the image space, generating an output image with enhanced details and style; The discriminator includes an input layer, a multi-layer convolution layer, a fully connected layer and an output layer; the input layer is used to receive an input image, the multi-layer convolution layer is used to process the input image through a series of convolutions, batch normalization, and Leaky ReLU activation, and gradually extract image features to obtain a feature map; the fully connected layer is used to map the features to a one-dimensional vector through the fully connected layer after the feature map is reduced in dimension; the output layer is used to output a probability value through a sigmoid activation function to determine whether the image comes from a real data distribution.

8. A lightweight embedded real-time image enhancement system according to claim 7, characterized in that: The generator and discriminator are each provided with two, namely, generator G and generator F, and discriminator Dx and discriminator Dy; The output image x of the FA+Net network model is used as the dataset X of the CycleGAN network model, and the clear underwater image y is used as the dataset Y of the CycleGAN network model. The content of dataset X is irrelevant to the content of dataset Y. The generator G generates the output image x into an image y' with a style similar to y, and the image x belongs to the dataset X, that is, G(x) = y', x∈X; the generator F generates the clear underwater real image y into an image x' with a style similar to x, and the image y belongs to the dataset Y, that is, F(y) = x', y∈Y; If the image y' generated by the generator G has a large content difference with the image y in the dataset Y or the image edge is blurred, the discriminator Dy outputs a low score; conversely, if the image x' generated by the generator F has similar content and a clear image compared to the image x in the dataset X, the discriminator Dx outputs a high score; After n trainings, the CycleGAN network model outputs images with a style close to that of dataset Y and a content close to that of dataset X.

9. A lightweight embedded real-time image enhancement method, characterized in that: The steps include: S1. Construct FA+Net network model and use FA+Net network to perform preliminary enhancement on underwater images; S2. Build a CycleGAN network model and use the CycleGAN network to perform underwater image style transfer and detail optimization; S3. Training the two network models and designing a loss function to optimize the network models to obtain a lightweight embedded real-time image enhancement system as described in any one of claims 1 to 7; S4, uses a lightweight embedded real-time image enhancement system to perform style transfer and detail optimization on underwater image data.

10. A lightweight embedded real-time image enhancement method according to claim 9, characterized in that: The optimization of the network model is as follows: The FA+Net network model optimizes feature aggregation and image enhancement parts through image quality evaluation indicators; The loss function of the CycleGAN network model consists of adversarial loss and cycle consistency loss, and the network parameters are optimized based on the loss function.

Citation Information

Cited By

  • Underwater image enhancement method and system based on vision-text fusion

    CN120634934A

  • 4K real-time rendering method and system for efficient cloud plug flow

    CN121037619A