Non-paired high-definition image enhancement method based on frequency decoupling and adaptive fusion

By adopting the method of frequency decoupling and adaptive fusion in the unpaired image enhancement technology, the problems of rough frequency information processing and insufficient contrast learning in the prior art are solved, and efficient restoration of image details and edges and significant improvement in enhancement quality are achieved.

CN120182099APending Publication Date: 2025-06-20CHONGQING UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510409574.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing unpaired image enhancement technology is rough in frequency information processing, resulting in excessive smoothing of edges or loss of detail, and lacks an effective discrimination mechanism to constrain the enhancement quality.

Method used

The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion is adopted. By constructing an enhanced image generator, the input image is decoupled with high and low frequency, and is enhanced using the global enhancement module and the high-frequency detail recovery module respectively. Finally, the authenticity and false discrimination and loss optimization are performed through the global-local dual-branch discriminator.

Benefits of technology

It significantly improves the restoration ability of image details and edges, avoids artifacts and blurring, and has a more stable and efficient enhancement quality, which is suitable for real-time enhancement tasks of high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182099A_ABST
    Figure CN120182099A_ABST
Patent Text Reader

Abstract

The invention relates to a non-paired high-definition image enhancement method based on frequency decoupling and adaptive fusion, and the method comprises the steps: constructing an enhanced image generator to process an input image into a complete enhanced image, and then carrying out the authenticity discrimination of the enhanced image through a global-local dual-branch discriminator; constructing an enhanced image generator total loss function, a global discriminator loss function and a local discriminator loss function, training the enhanced image generator and a global-local double-branch discriminator by adopting an alternating optimization strategy until the three losses do not change any more to obtain a trained generator, inputting a new image into the trained generator, and obtaining a new image; and outputting an enhanced image corresponding to the instant new picture. Through frequency decoupling and modular processing path design, different frequency information is optimized in a targeted mode, the original structure and details of the image are effectively reserved, and the enhanced image has better visual effects in the aspects of brightness, color, edge and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and computer vision, and in particular to a non-paired high-definition image enhancement method based on frequency decoupling and adaptive fusion applicable under reference-free conditions. Background Art

[0002] In the actual shooting process, problems such as image quality degradation due to low light, insufficient exposure, etc. widely exist. Traditionally, improving image quality involves professional personnel manually adjusting parameters such as brightness, contrast, and color. Although it can improve the visual effect and subsequent processing effect, this method has limited efficiency and inconsistent results, and is difficult to apply to large-scale scenarios. With the progress of the computer vision field, automated image enhancement technologies have increasingly attracted attention. Such methods can reduce manual intervention while providing more stable and consistent processing results, and are more suitable for modern image processing requirements.

[0003] Non-paired image enhancement technology learns the conversion relationship between a set of low-quality images and a set of high-quality images, getting rid of the dependence on strictly paired training data. It not only significantly reduces the data acquisition cost, but also extends the enhancement application to scenarios where paired data is difficult to obtain, such as historical image restoration and cross-device image processing. By learning the data distribution characteristics rather than the direct mapping relationship, this method can produce more natural results that conform to the real image distribution, providing a more practical and generalizable solution for automated image enhancement.

[0004] However, despite the significant progress made by non-paired image enhancement technologies in recent years, the existing methods generally have the following deficiencies:

[0005] 1. Coarse frequency information processing: Many methods use single-scale networks to process images, failing to distinguish between high-frequency and low-frequency content, resulting in over-smoothed edges or loss of details.

[0006] 2. Insufficient contrast learning: Due to the lack of supervision of the target image, the model often focuses on overall brightness improvement while ignoring the preservation and restoration of structural details.

[0007] 3. Poor artifacts and generalization ability: Non-paired methods are prone to introducing artifacts due to inconsistent distributions and lack an effective discrimination mechanism to constrain the enhancement quality.

[0008] Therefore, developing a non-paired high-definition image enhancement framework that integrates frequency domain decoupling modeling, brightness-guided enhancement, high-frequency detail restoration, and discriminative adversarial learning has important research value and application prospects. Summary of the Invention

[0009] Aiming at the above problems existing in the prior art, the technical problem to be solved by the present invention is: how to improve the ability to restore image details and edges and avoid artifacts and blurring phenomena.

[0010] To solve the above technical problems, the present invention adopts the following technical solutions: A non-paired high-definition image enhancement method based on frequency decoupling and adaptive fusion, comprising the following steps:

[0011] S1: Construct an enhanced image generator, and process the input image through the enhanced image generator into a complete enhanced image I, the steps are as follows: Processed into a complete enhanced image I, the steps are as follows:

[0012] S1-1: Image high-low frequency decoupling processing, decompose the input image I0 into a low-frequency component And a high-frequency component Where h and w respectively represent the height and width of the original image, and H0, H1, and H2 respectively correspond to the high-frequency detail information extracted at different resolutions.

[0013] S1-2: Use a global enhancement module to enhance H3, the global enhancement module includes a Transformer-based brightness perception module and a Mamba-based global information modeling module;

[0014] The L3 is input into the Transformer-based brightness perception module to obtain a brightness-enhanced low-frequency component Input into the Mamba-based global information modeling module to obtain the enhanced

[0015] S1-3: Upsample L3 and To 2 times the size and splice the high-frequency component H2 of the smallest size in the channel dimension, then input it into the adaptive modulation mask generation module to generate a mask graph M1, multiply the high-frequency component H2 of this layer by M1, add the residual, and then fine-tune the high-frequency component through a small network composed of two 1×1 convolutions and LeakyReLU to finally adjust and obtain the enhanced high-frequency component of this layer

[0016] Upsample M1 to 2 times its size to obtain a mask graph M2, then multiply it by the high-frequency component H1 of this layer and add the residual to obtain the result modulated by M2 The After upsampling to 2 times the size, multiply by the weight coefficient 0.05 and then add it to the Add, and then adjust the high-frequency component through a small network composed of two 1×1 convolutions and LeakyReLU to obtain the enhanced high-frequency component of this layer

[0017] Upsample M2 to 2 times its size to obtain a mask graph M3, then multiply it by the high-frequency component H0 of this layer and add the residual to obtain the result modulated by M3 The Multiply by the weight coefficient 0.05 and then with add them up, and then adjust the high-frequency components through a small network composed of two 1×1 convolutions and LeakyReLU to obtain the enhanced high-frequency components of this layer

[0018] S1-4: Finally, the enhanced high-frequency components of each layer and the enhanced are reconstructed through the Laplacian pyramid into a complete enhanced image I.

[0019] S2: Use the global-local dual-branch discriminator to judge the authenticity of I:

[0020] In the global discriminator branch, input I and the real normal light image I high into the global discriminator, and finally output a single-channel discrimination result representing the authenticity of I high and I through the convolutional layer.

[0021] In the local discriminator branch, crop I high and I into several small blocks with fixed sizes. Each small block is regarded as an independent part and input into the local discriminator. Use convolution to output the final single value of the authenticity of each small block, and take the average of the final single values of all small blocks to obtain the single-channel discrimination result representing the authenticity of I high and I.

[0022] S3: Construct the total loss of the enhanced image generator The loss function of the global discriminator and the local discriminator loss Adopt an alternating optimization strategy for training. Each training cycle includes the following three steps:

[0023] 1) Update the generator: Fix the global discriminator and the local full discriminator, and use to update the parameters of the enhanced image generator;

[0024] 2) Update the global discriminator: Fix the enhanced image generator and use to optimize the global discriminator;

[0025] 3) Update the local discriminator: Fix the enhanced image generator and use to optimize the local discriminator.

[0026] When and no longer decrease, the training is completed, and the trained enhanced image generator is obtained.

[0027] S4: For a new image, input the new image into the trained enhanced image generator, and the output is the corresponding enhanced image.

[0028] Preferably, in S1-1, the Laplace pyramid technique is used to perform high-low frequency decoupling processing on the input image to obtain a low-frequency component and a high-frequency component.

[0029] Preferably, in S1-2 is obtained as follows:

[0030] The L3 is obtained by initial feature extraction, then using a 3×3 convolutional layer to expand the number of channels from 3 to 16, followed by instance normalization and LeakyReLU activation, and then further expanding the number of channels to 48 through a 3×3 convolutional layer to obtain a set of features corresponding to L3.

[0031] The set of features corresponding to L3 passes through multiple serially stacked TransformerBlock modules, and then the number of channels is gradually reduced from 48 back to 16 through two convolutional layers, and finally back to 3. After applying a residual connection with L3 and passing through the tanh activation function, a low-frequency map with enhanced brightness is obtained

[0032] Preferably, in S1-2 is obtained as follows:

[0033] The expands the number of channels from 3 to 48 through a 3×3 convolutional layer, then extracts high-level semantic features through a structure of 5 attention state space modules stacked in series, and then maps the number of channels from 48 back to 3 through a 3×3 convolutional layer, and performs a residual connection with to obtain

[0034] Preferably, the process of generating the mask map M1 by the adaptive modulation mask generation module in S1-3 is as follows:

[0035] First, L3 and are upsampled to twice the size and concatenated with the high-frequency component H2 of the smallest size in the channel dimension to obtain a 9-channel concatenated feature. The 9-channel concatenated feature passes through the first convolutional layer and the LeakyReLU activation function to expand the 9-channel concatenated feature into a 64-channel feature map.

[0036] The 64-channel feature map flows through 3 residual blocks ResBlock in sequence and is then processed by the convolutional block attention module CBAM, and finally passes through the final convolutional layer to map the number of channels from 64 back to 3, obtaining the mask map M1.

[0037] Preferably, in S3 is constructed as follows:

[0038] The loss function of the enhanced image generator globally

[0039]

[0040] Among them, is the expected operation, X~P(X) represents sampling from the true distribution of the input image I0. G(X) is the enhanced image G output by the enhanced image generator, and D g (·) is the global discriminator.

[0041] The loss function of the enhanced image generator locally:

[0042]

[0043] Among them, G(X) (k) represents the k-th small block cropped from the enhanced image I output by the enhanced image generator, G(X) represents the enhanced image I, K is the number of small blocks, and D l (·) is the local discriminator.

[0044] Reconstruction loss: The reconstruction loss is calculated using the MSE loss for I and I0:

[0045]

[0046] Total loss of the enhanced image generator:

[0047]

[0048] Among them, δ, ε, and γ are all hyperparameters.

[0049] Preferably, in the S3 the construction process is as follows:

[0050]

[0051] Among them, Y represents I high , D g (·) is the global discriminator, is the sample linearly interpolated between G(X) and Y, λ gp is the weight of the gradient penalty term, represents the gradient operator for .

[0052] Preferably, in the S3 the construction process is as follows:

[0053]

[0054] Among them, Y (k) represents the k-th small block cropped from I high , Dl (·) is the local discriminator, G(X) (k) represents the k-th small block cropped from I, represents in Y (k) and G(X) (k) the linear interpolation sample between, K is the number of small blocks, represents the gradient operator of.

[0055] Compared with the prior art, the present invention has at least the following advantages:

[0056] 1. The enhancement quality is significantly improved: Through frequency decoupling and modular processing path design, targeted optimization of different frequency information is achieved, effectively retaining the original structure and details of the image, making the enhanced image have better visual effects in terms of brightness, color, and edges.

[0057] 2. The enhancement process is more stable and efficient: Adopting lightweight module design and introducing the Mamba module to optimize the calculation efficiency of the low-frequency path, combined with Transformer to improve the modeling depth, significantly reducing the consumption of computing resources while improving the enhancement quality, and being applicable to real-time enhancement tasks of high-resolution images.

[0058] 3. The adaptability of unpaired training is enhanced: By introducing a joint local-global discriminator and combining perceptual loss and frequency consistency loss for training, the problems of serious artifacts and color drift in traditional unpaired enhancement are overcome, and the enhancement results are more natural and conform to the real distribution, with good generalization ability and scene adaptability. Description of the Drawings

[0059] Figure 1 is a partial experimental result display diagram on the FiveK dataset.

[0060] Figure 2 is a partial experimental result display diagram on the private RCA dataset.

[0061] Figure 3 is a visual comparison diagram of the baseline method on the FiveK dataset.

[0062] Figure 4 is a flow diagram of the method of the present invention. Detailed Embodiment

[0063] The present invention will be further described in detail below.

[0064] The present invention proposes a non - paired high - definition image enhancement method based on multi - scale frequency - domain feature optimization and dynamic fusion. It innovatively introduces a frequency decoupling strategy, a brightness guidance mechanism, a dynamic fusion module, and a non - paired discriminant optimization strategy, effectively solving the following problems existing in the existing non - paired high - definition image enhancement technology:

[0065] (1) Solve the problem of blurred enhancement targets in non - paired training: Decompose the image into low - frequency and high - frequency components through the Laplacian pyramid, and construct enhancement paths respectively. Let the low - frequency be responsible for the reconstruction of brightness and color, and the high - frequency be responsible for detail restoration, so as to achieve a clearer and more definite enhancement target, effectively avoiding interference and misguidance between frequency - domain information.

[0066] (2) Solve the problem of insufficient low - frequency enhancement ability: Design a low - frequency modeling module that combines the Mamba attention mechanism, integrating the advantages of Transformer and state - space network to capture long - term dependencies such as global brightness and color in the image, significantly improving the reconstruction ability of low - frequency structures.

[0067] (3) Solve the problem of insufficient high - frequency detail restoration: The high - frequency enhancement module introduces a spatio - temporal attention mechanism to strengthen the edge texture response, and at the same time combines a residual structure for preliminary detail restoration; subsequently, the dynamic fusion module guides the high - frequency fusion process to further enhance the local contrast and fine structure of the image.

[0068] (4) Solve the problem of lack of alignment supervision in non - paired training: This method designs a local - global discriminator structure. The discriminator is trained in the perceptual domain. By making authenticity judgments on the overall structure of the global image and local detail regions respectively, it guides the generator to still generate realistic images under non - paired conditions, improving the stability and quality of non - paired enhancement.

[0069] The present invention decouples and processes the image in frequency, optimizes the low - frequency component by using brightness - guided enhancement and an attention state - space model, effectively enhancing the brightness and contrast of the image while maintaining the overall structure of the image. For the high - frequency component, an adaptive mask modulation mechanism is used for modulation, effectively promoting the adaptive fusion between frequencies at different scales and high - frequency detail restoration, ensuring the clarity of the generated image. At the same time, the overall architecture adopts a discriminative adversarial learning architecture, which can effectively handle non - paired image enhancement tasks, getting rid of the dependence on paired data and having good engineering significance.

[0070] See Figure 4 , the non - paired high - definition image enhancement method based on frequency decoupling and adaptive fusion, includes the following steps:

[0071] S1: Construct an enhanced image generator, and through the enhanced image generator, the input image Process it into a complete enhanced image I, the steps are as follows:

[0072] S1-1: Perform image high-low frequency decoupling processing, decompose the input image I0 into a low-frequency component and a high-frequency component where h and w respectively represent the height and width of the original image, and H0, H1, and H2 respectively correspond to the high-frequency detail information extracted at different resolutions.

[0073] Their sizes are h×w, h / 2×w / 2, and h / 4×w / 4 in sequence, mainly used to capture fine-grained structural features such as edges and textures in the image; while the low-frequency component L3 represents the global brightness, color, and structural information of the image, with the lowest spatial resolution but the richest semantics. Enhance them separately through different processing paths to avoid interference between high and low frequency information. The low-frequency component mainly preserves the global information such as the overall brightness and color of the image, while the high-frequency component contains the detail information such as edges and textures.

[0074] Decomposition: For the input image First, use the Laplacian pyramid to decompose the high and low frequency components to achieve effective decoupling of frequency information, and obtain a set of high-frequency components with gradually changing scales and a low-frequency component L3 with the smallest scale. The low-frequency component mainly preserves the global information such as the overall brightness and color of the image, while the high-frequency component contains the detail information such as edges and textures.

[0075] S1-2: Use the global enhancement module to enhance L3. The global enhancement module includes a Transformer-based brightness perception module and a Mamba-based global information modeling module;

[0076] The L3 is input into the Transformer-based brightness perception module to obtain the low-frequency component with enhanced brightness Input into the Mamba-based global information modeling module to obtain the enhanced

[0077] S1-3: Upsample L3 and the smallest-sized high-frequency component H2 to 2 times the size and splice them in the channel dimension, then input them into the adaptive modulation mask generation module to generate the mask graph M1. Multiply the high-frequency component H2 of this layer by M1, add the residual, and then finely adjust the high-frequency component through a small network composed of two 1×1 convolutions and LeakyReLU to finally obtain the enhanced high-frequency component of this layer

[0078] Upsample M1 to 2 times its size to obtain the mask graph M2, then multiply it by the high-frequency component H1 of this layer and add the residual to obtain the result modulated by M2 Up-sample to 2 times the size, multiply by the weight coefficient 0.05, and then add it to After adding, the high-frequency components are finely adjusted through a small network composed of two 1×1 convolutions and LeakyReLU, and finally the enhanced high-frequency components of this layer are obtained.

[0079] Up-sample M2 to 2 times its size to obtain the mask image M3, then multiply it by the high-frequency components H0 of this layer and add the residual to obtain the result modulated by M3. Up-sample Multiply by the weight coefficient 0.05 and then add it to After adding, the high-frequency components are finely adjusted through a small network composed of two 1×1 convolutions and LeakyReLU, and finally the enhanced high-frequency components of this layer are obtained.

[0080] S1-4: Finally, the enhanced high-frequency components of each layer are combined with the enhanced After being reconstructed by the Laplacian pyramid, it becomes the complete enhanced image I.

[0081] S2: Use the global-local dual-branch discriminator to distinguish the authenticity of I:

[0082] In the global discriminator branch, input I and the real normal light image I high into the global discriminator, and finally output a single-channel discrimination result representing the authenticity of I high and I through the convolutional layer.

[0083] First, perform feature extraction through a convolutional layer with a convolutional kernel size of 3 and a stride of 2, and at the same time expand the feature channels to 16 dimensions. Subsequently, apply the Leaky ReLU activation function to the feature map after instance normalization. By repeating the above convolutional operations, gradually increase the feature dimension to 128 channels, and finally output a single-channel discrimination result representing the authenticity of the image through the convolutional layer.

[0084] In the local discriminator branch, crop I high and I into several small blocks with fixed sizes. Each small block is regarded as an independent part and input into the local discriminator. Use convolution to output the final single value of the authenticity of each small block, and take the average of the final single values of all small blocks to obtain the single-channel discrimination result representing the authenticity of I high and I output by the local discriminator.

[0085] S3: Construct the total loss of the enhanced image generator Global discriminator loss function and local discriminator loss Training is carried out using an alternating optimization strategy. Each training cycle consists of the following three steps:

[0086] 1) Update the generator: Fix the global discriminator and the local full discriminator, and use to update the parameters of the enhanced image generator;

[0087] 2) Update the global discriminator: Fix the enhanced image generator, and use to optimize the global discriminator;

[0088] 3) Update the local discriminator: Crop each image into 6 small pieces of size 64×64, perform true / false discrimination separately, and after taking the average, fix the enhanced image generator, and use to optimize the local discriminator.

[0089] When and no longer decrease, the training is completed, and the trained enhanced image generator is obtained.

[0090] S4: For a new image, input the new image into the trained enhanced image generator, and the output is the corresponding enhanced image.

[0091] Specifically, in S1-1, the Laplace pyramid technique is used to perform high-low frequency decoupling processing on the input image to obtain the low-frequency component and the high-frequency component.

[0092] Specifically, in S1-2 is obtained as follows:

[0093] The L3 is obtained through initial feature extraction, then using a 3×3 convolutional layer to expand the number of channels from 3 to 16, followed by instance normalization and LeakyReLU activation, and then further expanding the number of channels to 48 through a 3×3 convolutional layer to obtain a set of features corresponding to L3, enhancing the richness of feature representation.

[0094] The L3 corresponding to a set of features passes through multiple cascaded TransformerBlock modules. Each TransformerBlock module contains two key components: 1) Dynamic Range Histogram Self-Attention (DHSA). This mechanism first sorts the set of features to generate a histogram feature representation, and then captures the relationships between different brightness levels through self-attention operations, effectively processing the image dynamic range and enhancing the contrast; 2) Dual-Scale Gated Feed-Forward Network (DGFF), which enhances feature expression through multi-scale convolution and gating mechanisms while retaining details. The features processed by DHSA undergo dual LayerNorm normalization to ensure training stability. Then, through two layers of convolution, the number of channels is gradually reduced from 48 back to 16 and finally to 3. After applying a residual connection with L3 and passing through the tanh activation function, a low-frequency map with enhanced brightness is obtained. This comprehensive processing flow effectively enhances the global performance of the low-frequency components, improving details, contrast, and dynamic range.

[0095] Specifically, the L3 first maps the image from 3 channels to 16 channels and then to 48 channels through two layers of convolutional networks, and then is processed by 2 cascaded Transformer blocks. Each Transformer block contains a Dynamic Histogram Self-Attention (DHSA) mechanism and a Dual-Scale Gated Feed-Forward Network (DGFF), which can effectively capture long-range dependencies and local details in the image. Finally, through two layers of convolution, the features are mapped back to 3 channels, and after applying a residual connection with the original low-frequency image L3 and passing through the tanh activation function, an enhanced low-frequency component is generated.

[0096] Specifically, in the S1-2 The acquisition process is as follows:

[0097] The is fed into the Attention State Space Model (Mamba) for further enhancement. The number of channels is expanded from 3 to 48 through a 3×3 convolutional layer, providing a richer feature representation space. These features are then reshaped into a sequence form, and then high-level semantic features are extracted through a structure of 5 cascaded Attention State Space modules. Then, the number of channels is mapped back from 48 to 3 through a 3×3 convolutional layer, and After applying a residual connection, we get The dual enhancement process of the low-frequency components is completed. This Transformer+Mamba hybrid architecture fully combines the advantages of the two models, achieving comprehensive enhancement of the low-frequency components.

[0098] Specifically, the process of generating the mask map M1 through the Adaptive Modulation Mask Generation Module in S1-3 is as follows:

[0099] First, L3 and are upsampled to twice the size and concatenated with the high-frequency component H2 of the smallest size in the channel dimension to obtain a 9-channel concatenated feature. The 9-channel concatenated feature passes through the first convolutional layer (Conv2d 9→64) and the LeakyReLU activation function, expanding the 9-channel concatenated feature into a 64-channel feature map. This step enables the network to extract richer feature information.

[0100] The 64-channel feature map flows through 3 residual blocks ResBlock in sequence and is then processed by the convolutional block attention module CBAM. Finally, it passes through the final convolutional layer (Conv2d 64→3), mapping the number of channels back from 64 to 3, resulting in the mask map M1. Each residual block ResBlock contains two 3×3 convolutional layers and the LeakyReLU activation function, maintaining gradient fluidity and enhancing network depth through skip connections. These residual blocks keep the number of channels at 64 but can extract deeper feature representations.

[0101] This CBAM contains two key components:

[0102] Channel Attention: It captures the relationships between channels using global average pooling and max pooling, generates channel weights through a multi-layer perceptron and the sigmoid function, and highlights important feature channels.

[0103] Spatial Attention: By generating a spatial attention map, it enables the network to focus on visually important regions in the image and enhances the perception ability of local details.

[0104] Specifically, the construction process of S3 in is as follows:

[0105] The loss function of the enhanced image generator globally

[0106]

[0107] where, is the expectation operation, X~P(X) represents sampling from the true distribution of the input image I0. G(X) is the enhanced image I output by the enhanced image generator, and D g (·) is the global discriminator.

[0108] The loss function of the enhanced image generator locally:

[0109]

[0110] where, G(X) (k)The k-th patch cropped from the enhanced image I output by the enhanced image generator, G(X) represents the enhanced image I, K is the number of patches, and D l (·) is the local discriminator.

[0111] Reconstruction loss: The reconstruction loss is calculated using the MSE loss for I and I0:

[0112]

[0113] Total loss of the enhanced image generator:

[0114]

[0115] Among them, δ, ε, and γ are all hyperparameters, δ = 1, ε = 1, γ = 50.

[0116] Specifically, in S3 The construction process is as follows:

[0117]

[0118] Among them, Y represents I high , D g (·) is the global discriminator, is the sample linearly interpolated between G(X) and Y, λ gp is the weight of the gradient penalty term, and its value is set to 50, represents the gradient operator for the interpolated sample of.

[0119] Specifically, in S3 The construction process is as follows:

[0120]

[0121] Among them, Y (k) represents the k-th patch cropped from I high , D l (·) is the local discriminator, G(X) (k) represents the k-th patch cropped from I, represents the linearly interpolated sample between Y (k) and G(X) (k) , K is the number of patches, represents the gradient operator for of.

[0122] Experimental comparison proves the effect

[0123] To verify the effectiveness of the present invention in the unpaired high-definition image enhancement task, we conducted systematic experiments on multiple public and private standard datasets and carried out a comparative analysis with the current mainstream unpaired image enhancement methods:

[0124] 1. Experimental Setup

[0125] Datasets: The MIT-Adobe FiveK and RCA datasets were selected for the experiments. Different from the supervision method of the paired enhancement task, an unsupervised learning strategy was adopted in this experiment. Each dataset was divided into two subsets: the first 2225 images were used as the low-quality domain, and the latter 2225 images were used as the high-quality domain to achieve unpaired image enhancement training.

[0126] Training Strategy: During the training process, a cyclic alternating optimization mechanism was adopted, combined with the collaborative optimization of the generator and the global-local discriminator. The global discriminator focuses on the overall style consistency of the image, while the local discriminator ensures the realism and consistency of the enhanced image at the local detail level by cropping multiple 64×64 image patches and scoring them independently.

[0127] Evaluation Metrics: Considering that the standard reference image cannot be obtained in the unpaired task, in addition to using PSNR and SSIM as auxiliary metrics, NIQE (Naturalness Image Quality Evaluator) was mainly used as the no-reference image quality assessment standard to comprehensively measure the naturalness and realism of the image.

[0128] Comparison Methods: Current mainstream unpaired image enhancement methods such as EnlightenGAN, LPTN, SCI, and Zero-DCE were selected for comparison.

[0129] 2. Experimental Results

[0130] On the MIT-Adobe FiveK and RCA datasets, the method of the present invention achieved the best performance in each index, significantly outperforming other comparison methods. The comparison results are shown in Table 1.

[0131] Table 1 Average PSNR / SSIM / NIQE Metrics of Test Set Images

[0132]

[0133]

[0134] From Figure 1 and Figure 2 it can be observed that the overall quality of the enhanced image is good, the color is relatively harmonious, the human subject and the environment are natural and coordinated, and it has a good texture.

[0135] From Figure 3It can be seen that the original input image (the first column) has typical low-illumination problems. Although EnlightenGAN (the second column) improves the overall brightness, it causes color distortion and loss of texture details, which is related to its design that overemphasizes brightness enhancement while neglecting feature preservation. LPTN (the third column) shows obvious color shift problems and generates significant noise in the dark areas. Zero-DCE and SCI (the fourth and fifth columns) have the problem of excessive brightness enhancement, resulting in loss of details in the highlight areas. In contrast, the method proposed in the present invention (the sixth column) shows significant advantages: while maintaining the natural color balance, it effectively improves the image brightness, clearly retains the detailed texture, and the overall visual effect is the most harmonious.

[0136] As can be seen from Table 1, the method proposed in the present invention is superior to the baseline methods such as EnlightenGAN, LPTN, and Zero-DCE in terms of PSNR, SSIM, and NIQE metrics, showing better image fidelity and structural similarity.

[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion, characterized by: The steps include: S1: Construct an enhanced image generator, through which the input image is transformed Processing into a complete enhanced image I, the steps are as follows: S1-1: Image high- and low-frequency decoupling processing, decomposing the input image I0 into low-frequency components and high frequency components Among them, h and w represent the height and width of the original image respectively, and H0, H1, and H2 correspond to the high-frequency detail information extracted at different resolutions respectively; S1-2: L3 is enhanced using a global enhancement module, wherein the global enhancement module includes a Transformer-based brightness perception module and a Mamba-based global information modeling module; The L3 input is based on the brightness perception module of Transformer to obtain the low-frequency component after brightness enhancement Input based on Mamba's global information modeling module has been enhanced S1-3: Connect L3 and The high-frequency component H2 upsampled to 2 times the size is concatenated with the minimum size in the channel dimension, and then input into the adaptive modulation mask generation module to generate the mask map M1. The high-frequency component H2 of this layer is multiplied by M1, and the residual is added. Then, the high-frequency component is fine-tuned through a small network composed of two 1×1 convolutions and LeakyReLU to finally obtain the enhanced high-frequency component of this layer. Upsample M1 to twice its size to obtain the mask image M2, then multiply it with the high-frequency component H1 of this layer and add the residual to obtain the result after M2 modulation. Will After upsampling to 2 times the size, multiply by the weight factor 0.05 and then Add them together, and then adjust the high-frequency components through a small network consisting of two 1×1 convolutions and LeakyReLU to obtain the enhanced high-frequency components of this layer. Upsample M2 to twice its size to obtain the mask image M3, then multiply it with the high-frequency component H0 of this layer and add the residual to obtain the result after M3 modulation. Will Multiply by the weight factor 0.05 and then Add them together, and then adjust the high-frequency components through a small network consisting of two 1×1 convolutions and LeakyReLU to obtain the enhanced high-frequency components of this layer. S1-4: Finally, the high-frequency components of each layer are enhanced With the enhanced Reconstructed into a complete enhanced image I through Laplacian pyramid; S2: Use the global-local two-branch discriminator to distinguish true from false: In the global discriminator branch, I and the true normal light image I high Input to the global discriminator, and finally output through the convolutional layer representation I high The single channel discrimination result of the authenticity of I; In the local discriminator branch, I high I is cut into several small blocks of fixed size, each small block is regarded as an independent part of the input local discriminator, and the final single value of the authenticity of each small block is output by convolution. The final single value of all small blocks is averaged to obtain the representation I output by the local discriminator high The single channel discrimination result of the authenticity of I; S3: Constructing the total loss of the enhanced image generator Global Discriminator Loss Function and the local discriminator loss The training is performed using an alternating optimization strategy. Each training cycle consists of the following three steps: 1) Update the generator: fix the global discriminator and the local full discriminator, and use Update enhanced image generator parameters; 2) Update the global discriminator: fix the enhanced image generator and use Optimize the global discriminator; 3) Update the local discriminator: fix the enhanced image generator and use Optimize local discriminator; when and If it no longer decreases, the training is completed and the trained enhanced image generator is obtained; S4: For a new image, the new image is input into the trained enhanced image generator, and the output is the corresponding enhanced image.

2. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion according to claim 1, characterized in that: In the S1-1, the Laplace pyramid technology is used to perform high-low frequency decoupling processing on the input image to obtain low-frequency components and high-frequency components.

3. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion as claimed in claim 2, characterized in that: In S1-2 The process of obtaining is: The L3 is extracted through initial features, and then a 3×3 convolution layer is used to expand the number of channels from 3 to 16, and then instance normalization and LeakyReLU activation are performed, and then a 3×3 convolution layer is used to further expand the number of channels to 48 to obtain a set of features corresponding to L3; The L3 corresponds to a set of features that pass through multiple series-stacked TransformerBlock modules, and then through two layers of convolution to gradually reduce the number of channels from 48 to 16, and finally to 3, and then pass through the tanh activation function after applying the residual connection with L3 to obtain the low-frequency image with brightness enhancement.

4. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion as claimed in claim 3, characterized in that: In S1-2 The process of obtaining is: Said The number of channels is expanded from 3 to 48 through a 3×3 convolutional layer, and then the high-level semantic features are extracted through a structure in which five attention state space modules are stacked in series. Then, the number of channels is mapped back from 48 to 3 through a 3×3 convolutional layer, which is similar to Perform residual connection to obtain 5. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion as claimed in claim 4, characterized in that: The process of generating the mask image M1 by the adaptive modulation mask generation module in S1-3 is as follows: First, L3 and After upsampling to twice the size and concatenating the high-frequency component H2 of the minimum size in the channel dimension, a 9-channel concatenated feature is obtained. The 9-channel concatenated feature is expanded to a 64-channel feature map after the first convolution layer and the LeakyReLU activation function. The 64-channel feature map flows through the three residual blocks ResBlock in sequence and then is processed by the convolution block attention module CBAM. Finally, it passes through the final convolution layer to map the number of channels from 64 back to 3, mask map M1.

6. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion according to claim 1, characterized in that: The S3 The construction process is as follows: Enhanced image generator global loss function in, is the expected operation, X~P(X) represents sampling from the real distribution of the input image I0; G(X) is the enhanced image I output by the enhanced image generator, D g (·) is the global discriminator; The local loss function of the enhanced image generator is: Among them, G(X) (k) represents the kth small block cropped from the enhanced image I output by the enhanced image generator, G(X) represents the enhanced image I, K is the number of small blocks, D l (·) is the local discriminator; Reconstruction loss: The reconstruction loss is calculated using MSE loss for I and I0: Total loss of enhanced image generator: Among them, δ, ε and γ are all hyperparameters.

7. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion according to claim 1, characterized in that: The S3 The construction process is as follows: Where Y represents I high , D g (·) is the global discriminator, is a sample linearly interpolated between G(X) and Y, λ gp is the gradient penalty weight, Express The gradient operator.

8. The unpaired high-definition image enhancement method based on frequency decoupling and adaptive fusion according to claim 1, characterized in that: The S3 The construction process is as follows: Among them, Y (k) Indicates I high The kth small piece cut out, D l (·) is the local discriminator, G(X) (k) represents the kth small block cut out by I, Indicates that in Y (k) and G(X) (k) Linear interpolation samples between, K is the number of small blocks, Express The gradient operator.

Citation Information

Cited By

  • Unmanned aerial vehicle image target detection method and device based on dynamic deformable convolution feature fusion

    CN120783172A

  • Night scene lane line detection method

    CN121564677A

  • High-magnification face super-resolution method based on enhanced Mama and adversarial learning

    CN122347503A

  • High-resolution face super-resolution method based on enhanced mamba and adversarial learning

    CN122347503B