Ultrahigh-definition image restoration method based on frequency-enhanced variational auto-encoder
Through the combination of frequency-enhanced variational autoencoder and wavelet transform adapter, the problems of information loss and high computing cost in ultra-high-definition image recovery are solved, high-quality and low-cost image repair effects are achieved, and real-time processing of consumer-grade devices is supported.
Patent Information
- Application Number
- CN202510376187.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-11
AI Technical Summary
The existing ultra-high-definition image recovery methods are difficult to achieve efficient real-time processing due to downsampling due to information loss, inconsistency in details and high computing costs, especially on consumer-grade devices.
Frequency-enhanced variational autoencoder (FE-VAE) is used to combine wavelet transform adapter (WTA) and potential spatial repair network (IRNet), and through space-frequency adaptive decomposition and high-frequency information injection, a two-stage training framework is built to improve image reconstruction quality and reduce computational costs.
It significantly improves the quality and detail fidelity of ultra-high-definition image repair, reduces computing costs, supports real-time full-resolution inference of 4K ultra-high-definition images on consumer-grade GPUs, and significantly improves PSNR and SSIM indicators.
Smart Images

Figure CN120298230A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of ultra - high - definition image processing, and particularly to an ultra - high - definition image restoration method based on a frequency - enhanced variational auto - encoder, which is used to improve the restoration effect of ultra - high - definition images in tasks such as de - blurring, de - fogging, and low - light enhancement. Background Art
[0002] With the rapid development of imaging sensors and display technologies, significant progress has been made in ultra - high - definition (UHD) imaging technology. However, ultra - high - definition images captured under harsh conditions such as low light, fast motion, or haze usually suffer from quality degradation problems. This not only seriously affects the visual quality but also limits the application of ultra - high - definition images in advanced visual tasks.
[0003] Existing learning - based image restoration algorithms are difficult to effectively process ultra - high - definition images. Traditional restoration methods usually reduce the computational complexity by downsampling and adopt a processing paradigm of "downsampling - enhancement - upsampling". However, such methods inevitably lose detail information during the downsampling process, resulting in inconsistent restoration results. In addition, directly performing high - resolution inference in the pixel space requires too much computational resources and is difficult to perform real - time inference on consumer - grade GPUs.
[0004] Recently, variational auto - encoders (VAEs) have received extensive attention due to their potential space representation ability and efficient image reconstruction ability. However, directly applying VAEs to ultra - high - definition image restoration faces the following challenges:
[0005] 1. High - performance VAEs have a large number of parameters and a large amount of computation, making it difficult to run on consumer - grade devices;
[0006] 2. Traditional VAEs have an inter - domain gap problem when encoding degraded images, resulting in a decline in restoration quality;
[0007] 3. High - frequency information in the latent space is easily lost, affecting the restoration of details.
[0008] Therefore, how to effectively combine the latent space representation ability of VAEs with frequency priors to solve the above problems in ultra - high - definition image restoration is an important issue that urgently needs to be solved. Summary of the Invention
[0009] The present invention is proposed to solve the above - mentioned deficiencies of the existing technologies, and provides an ultra - high - definition image restoration method based on a frequency - enhanced variational auto - encoder, aiming to solve the consistency problem caused by downsampling in existing UHD image restoration methods, overcome the problems of large number of parameters, domain gap, and high - frequency information loss in VAE for UHD image restoration, so as to achieve a higher - quality image restoration effect.
[0010] To achieve the above object, the present invention adopts the following technical solutions:
[0011] A method for super high-definition image restoration based on a frequency-enhanced variational autoencoder according to the present invention is characterized in that it is carried out according to the following steps:
[0012] Step 1: Obtain a degraded image in any degraded scenario and its corresponding clear image , where represents the length of the image, represents the width of the image, represents the number of channels of the image;
[0013] Step 2: Construct a frequency-enhanced variational autoencoder FE-VAE, including: an image input embedding layer, a frequency domain enhancement processing module, and an image output embedding layer, and process to obtain a reconstruction result , thereby constructing the total loss of the first stage, and performing the first stage training on the FE-VAE to obtain the frequency-enhanced variational autoencoder after the first stage training;
[0014] Step 3: Construct a frequency-enhanced VAE super high-definition image restoration network FEVAE-UHD, including: an encoder frequency domain replay module, a latent space image restoration network, and a decoder high-frequency information injection module; and process to obtain a restored image , thereby constructing the total loss of the second stage; and freeze the parameters of the frequency-enhanced variational autoencoder after the first stage training, perform the second stage training on the frequency-enhanced VAE super high-definition image restoration network, and obtain the frequency-enhanced variational autoencoder after the second stage training for processing the degraded image to obtain a restored image .
[0015] The method for super high-definition image restoration based on a frequency-enhanced variational autoencoder according to the present invention is also characterized in that the step 2 includes:
[0016] Step 2.1: The image input embedding layer extracts features from to obtain clear encoded features ; where represents the length of the encoded feature, represents the width of the encoded feature, represents the number of channels of the encoded feature;
[0017] Step 2.2: After the encoded feature is processed by the frequency domain enhancement processing module, the output feature is obtained ;
[0018] Step 2.3. Output features After being processed by the output image embedding layer, the reconstruction result is obtained ;
[0019] Step 2.4. Use Equations (3) - (5) to construct the reconstruction loss function, distribution loss function, and frequency loss function of the frequency-enhanced variational autoencoder FE-VAE respectively, so as to obtain the total loss of the first stage by using Equation (6), and perform the first-stage training on FE-VAE to obtain the frequency-enhanced variational autoencoder after the first-stage training: , distribution loss function , frequency loss function , so as to obtain the total loss of the first stage by using Equation (6) , and perform the first-stage training on FE-VAE to obtain the frequency-enhanced variational autoencoder after the first-stage training:
[0020] (3)
[0021] (4)
[0022] (5)
[0023] (6)
[0024] In Equations (3) - (6), represents the KL divergence, is the approximate posterior distribution of the encoded latent variable of the given input image , is the prior distribution of the encoded latent variable , and follows a standard Gaussian distribution, , , are three weighting coefficients.
[0025] Furthermore, the frequency domain enhancement processing module in Step 2.2 consists of an N-layer encoder and an N-layer decoder. Each layer of the encoder and decoder includes: a spatial frequency adaptive decomposition module SFAD, a frequency-aware feature extraction module FAFE, and a spatial frequency interaction module SFIM;
[0026] Step 2.2.1. The spatial-frequency adaptive decomposition module SFAD in the first-layer encoder processes to obtain the frequency domain feature and the spatial domain feature ;
[0027] Step 2.2.2. The Fourier-aware feature extraction module FAFE in the first-layer encoder processes and Process it to obtain the output feature of the frequency domain branch and obtain the output feature of the spatial domain branch ;
[0028] Step 2.2.3. The spatial-frequency interaction module SFIM in the encoder of the first layer processes and to obtain the fused feature ;
[0029] Step 2.2.4. According to the process of Step 2.2.1 - Step 2.2.3, input into the encoder of the second layer for processing, until the fused feature output by the encoder of the (N - 1)-th layer is input into the encoder of the N-th layer for processing, to obtain the encoded fused feature , map it to a distribution to obtain the encoded latent variable , so that after sampling , obtain the decoded fused feature , then input it into the decoder of the first layer, and process it according to the process of Step 2.2.1 - Step 2.2.3 until the final fused feature is output by the decoder of the N-th layer;
[0030] Furthermore, Step 2.2.1 is carried out as follows:
[0031] Step 2.2.1.1. After performing a channel Fourier transform on , obtain the amplitude spectrum and the phase spectrum ;
[0032] Step 2.2.1.2. After performing a pointwise convolution on the amplitude spectrum , and then performing an inverse Fourier transform together with , obtain the global perception feature ;
[0033] Step 2.2.1.3. respectively pass through two pointwise convolution operations with non-shared parameters, and correspondingly obtain the frequency domain feature and the spatial domain feature .
[0034] Furthermore, Step 2.2.2 is carried out as follows:
[0035] Step 2.2.2.1. Perform a Fourier transform on in the spatial dimension to obtain the frequency domain amplitude spectrum and the frequency domain phase spectrum ;
[0036] Step 2.2.2.2, only amplitude spectrum After performing point-by-point convolution, and then with together perform inverse Fourier transform to obtain the frequency-domain branch output feature ;
[0037] Step 2.2.2.3, use the residual network module composed of depthwise separable convolution to process to obtain the spatial-domain branch output feature .
[0038] Furthermore, Step 2.2.3 is carried out according to the following steps:
[0039] Step 2.2.3.1, respectively perform cross gating on and according to Equation (1) and Equation (2) to obtain the frequency-domain gating output feature and the spatial-domain gating output feature :
[0040] (1)
[0041] (2)
[0042] In Equation (1) and Equation (2), represents the Sigmoid activation function, represents element-wise multiplication;
[0043] Step 2.2.3.2, after concatenating and in the channel dimension, then perform point-by-point convolution fusion processing to obtain the fused feature .
[0044] Furthermore, the encoder frequency-domain replay module in Step 3 is composed of extracting the image input embedding layer, the N-layer encoder of the frequency-domain enhancement processing module in the pre-trained frequency enhancement variational autoencoder, and respectively adding an adapter encoding insertion module after each layer of the encoder;
[0045] The decoder high-frequency information injection module is composed of extracting the N-layer decoder of the frequency-domain enhancement processing module in the pre-trained frequency enhancement variational autoencoder and the image output embedding layer, and respectively adding an adapter decoding insertion module after each layer of the decoder;
[0046] Step 3.1, the encoder frequency-domain replay module processes to obtain the Nth-layer fused enhancement feature , the Nth-layer prior output feature and the Nth-layer high-frequency component feature ;
[0047] Step 3.2: Construct a latent space restoration network, and encode into a degraded latent variable and then process it to obtain a clean latent variable , so that after sampling in , the input features of the decoder are obtained ;
[0048] Step 3.3: The decoder high-frequency information injection module processes , and to obtain the final restoration result ;
[0049] Step 3.4: Use Equation (10) and Equation (11) to construct the restoration loss and the frequency loss , respectively, and then use Equation (12) to construct the total loss in the second stage:
[0050] (10)
[0051] (11)
[0052] (12), where and are two weighting coefficients.
[0053] Furthermore, Step 3.1 is carried out as follows:
[0054] Step 3.1.1: The image input embedding layer of the pre-trained frequency enhancement variational autoencoder in the encoder frequency domain replay module extracts features from to obtain the degraded encoded features ;
[0055] Step 3.1.2: When i = 1, the i-th layer encoder in the frequency domain enhancement processing module of the pre-trained frequency enhancement variational autoencoder in the encoder frequency domain replay module processes to obtain the i-th layer fusion features ;
[0056] Step 3.1.3: and are input into the i-th layer adapter encoding insertion module for processing to obtain the i-th layer fusion enhanced features ;
[0057] Step 3.1.4: When i = 2, 3,..., N, the (i - 1)-th layer fusion enhanced features As the input of the i-th layer encoder, insert the i-1-th layer prior output by the i-1-th layer adapter encoding insertion module As the i-th layer prior of the i-th layer adapter encoding insertion module, and process it according to the process of steps 3.1.2 - 3.1.3 and until the N-1-th layer fusion feature output by the N-1-th layer encoder and the N-1-th layer prior output by the N-1-th layer adapter encoding insertion module are input into the N-th layer adapter encoding insertion module for processing, and the N-th layer fusion enhanced feature , the N-th layer prior output feature and the N-th layer high-frequency component feature are output.
[0058] Furthermore, step 3.1.3 is carried out as follows:
[0059] Step 3.1.3.1: When i = 1, take as the i-1-th layer prior of the i-th layer adapter encoding insertion module, denoted as ; the i-th layer adapter encoding insertion module performs wavelet transform on to obtain the i-th layer subbands , and includes: low-frequency component , high-frequency component in the horizontal direction , high-frequency component in the vertical direction , high-frequency component in the diagonal direction ;
[0060] Step 3.1.3.2: After fusing with the four frequency components in respectively, four fusion features of the i-th layer are obtained correspondingly, and the four fusion features of the i-th layer are respectively mapped to the four weights of the filter kernel through a convolutional layer and a sigmoid activation function, so as to modulate the spectrum of using the four weights of the filter kernel respectively, and four modulated spectrum features are obtained correspondingly and pointwise convolution operation is performed, so as to obtain the i-th layer frequency fusion feature ;
[0061] Step 3.1.3.2: Use Equation (7) to obtain the modulated subband of the i-th layer , and after concatenating and on the channel, use three pointwise convolutions to decompose the concatenated features to obtain the i-th layer fusion enhanced feature , the i-th layer prior and the high-frequency encoded features of the i-th layer ;
[0062] (7)
[0063] In formula (7), represents average pooling, represents the Sigmoid activation function, represents element-wise multiplication;
[0064] Step 3.1.3.3, after performing zero convolution processing on and adding it to the i-th layer of fusion enhanced features is obtained.
[0065] Furthermore, Step 3.3 is carried out as follows:
[0066] Step 3.3.1, when i = 1, the i-th layer decoder of the pre-trained frequency enhanced variational autoencoder in the decoder high-frequency information injection module processes to obtain the i-th layer decoded original feature ;
[0067] When i = 1, the i-th layer adapter decoding insertion module in the decoder high-frequency information injection module performs inverse wavelet transform on and to obtain the i-th layer wavelet reconstruction feature , and then the i-th layer enhanced decoded feature is obtained using formula (8) and formula (9);
[0068] (8)
[0069] (9)
[0070] In formula (8) and formula (9), represents inverse wavelet transform, represents zero convolution;
[0071] Step 3.3.2, when i = 2, 3,..., N, the (i - 1)-th layer enhanced decoded feature is input into the i-th layer decoder of the pre-trained frequency enhanced variational autoencoder in the decoder high-frequency information injection module for processing to obtain the i-th layer decoded original feature ;
[0072] The i-th layer adapter decoding insertion module in the decoder high-frequency information injection module processes the high-frequency component features output by the i-th layer adapter encoding insertion module and the (i - 1)-th layer wavelet reconstruction feature Perform inverse wavelet transform to obtain the wavelet reconstruction features of the i-th layer , and thus obtain the enhanced decoding features of the i-th layer by using Equations (8) and (9) ; until the enhanced decoding features of the (N - 1)-th layer output by the (N - 1)-th layer decoder are input into the N-th layer decoder of the pre-trained frequency-enhanced variational autoencoder in the decoder high-frequency information injection module for processing, and the enhanced decoding features of the N-th layer are output ; and the N-th layer adapter decoding insertion module performs inverse wavelet transform on and the wavelet reconstruction features of the (N - 1)-th layer to obtain the wavelet reconstruction features of the N-th layer , and thus obtain the output enhanced decoding features of the N-th layer by using Equations (8) and (9) ;
[0073] Step 3.3.3, input into the image output embedding layer of the decoder high-frequency information injection module to obtain the restored image .
[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0075] 1. The present invention significantly improves the restoration quality and detail fidelity: Through the multi-level high-frequency information injection mechanism of the wavelet transform adapter (WTA), the problem of high-frequency detail loss caused by downsampling in traditional methods is effectively solved. Experiments show that in the ultra-high-definition defogging task, the PSNR reaches 24.35dB, nearly 3dB higher than the baseline method, and the SSIM index is improved from 0.942 to 0.945.
[0076] 2. The present invention reduces the computational cost and the number of parameters: By adopting the frequency-enhanced variational autoencoder (FE-VAE), through spatial-frequency adaptive decomposition (SFAD) and lightweight frequency-aware feature extraction (FAFE), the number of parameters of the traditional VAE is compressed from 83.6M to 1.06M, and the computational amount is reduced from 445.3 GFLOPs to 3.4 GFLOPs. At the same time, the latent space repair network (IRNet) only needs 4 layers of cubic hybrid blocks to achieve efficient optimization, supporting full-resolution real-time inference of 4K ultra-high-definition images on consumer-grade GPUs.
[0077] 3. The present invention enhances the spatial-frequency domain consistency of the restoration results: By combining the FFT loss function and KL divergence regularization, the spatial domain reconstruction error and frequency domain spectrum matching are simultaneously constrained during training, and it shows better performance than the prior art in ultra-high-definition image restoration tasks such as low-light enhancement, deblurring, and defogging. Description of the Drawings
[0078] Figure 1 It is a flow chart of the present invention;
[0079] Figure 2 It is an overall framework diagram of the frequency enhancement VAE ultra-high definition image restoration network FEVAE-UHD proposed by the present invention;
[0080] Figure 3 It is a structural diagram of the frequency enhancement variational autoencoder FE-VAE of the present invention. Detailed implementation manners
[0081] In this embodiment, an ultra-high definition image (UHD) restoration method based on a frequency enhancement variational autoencoder (FEVAE) is proposed for the problems of information loss, inconsistent details, and high computational cost caused by downsampling in ultra-high definition image restoration. The algorithm network mainly consists of a frequency enhancement variational autoencoder (FE-VAE), a wavelet transform adapter (WTA), and a latent space restoration network (IRNet). FE-VAE improves parameter efficiency and retains global spectral perception ability through spatial-frequency adaptive decomposition and frequency-aware feature extraction modules; WTA bridges the distribution difference between the degradation domain and the clean domain and injects high-frequency details through multi-level wavelet decomposition and frequency-aware adaptive modulation.
[0082] The working process of this method is divided into two stages:
[0083] The first stage: Train FE-VAE based on the image reconstruction task, input a high-definition image , and output a reconstructed image ;
[0084] The second stage: Freeze the parameters of FE-VAE, construct a complete restoration framework, input a degraded image , and output a restored image .
[0085] Specifically, as Figure 1 shown, this method is carried out according to the following steps:
[0086] Step 1: Obtain a degraded image under any degradation scenario and its corresponding clear image , where represents the length of the image, represents the width of the image, and represents the number of channels of the image.
[0087] Step 2: As Figure 3As shown in the figure, a frequency-enhanced variational autoencoder FE-VAE is constructed to extract a compact image latent space representation. This process includes an image input embedding layer, a frequency domain enhancement processing module, and an image output embedding layer. Among them, the frequency domain enhancement processing module consists of an N-layer encoder and an N-layer decoder. Each layer of the encoder and decoder includes a spatial-frequency adaptive decomposition module SFAD, a frequency-aware feature extraction module FAFE, and a spatial-frequency interaction module SFIM;
[0088] Step 2.1, the image input embedding layer performs feature extraction to obtain clear encoded features ; among them, represents the length of the encoded feature, represents the width of the encoded feature, represents the number of channels of the encoded feature.
[0089] Step 2.2, after the encoded feature is processed by the frequency domain enhancement processing module, the output feature is obtained;
[0090] Step 2.2.1, the spatial-frequency adaptive decomposition module SFAD in the encoder of the first layer processes to obtain the frequency domain feature and the spatial domain feature ;
[0091] Step 2.2.1.1, after performing a channel Fourier transform on , it is converted from the spatial domain to the frequency domain to obtain the amplitude spectrum and the phase spectrum ;
[0092] Step 2.2.1.2, since for a degraded image, its degradation components mainly gather in the amplitude spectrum, so the amplitude spectrum is subjected to pointwise convolution for global channel perception, and then together with is subjected to inverse Fourier transform to obtain the global perception feature ;
[0093] Step 2.2.1.3, are respectively separated into frequency domain and spatial domain features through two pointwise convolution operations with non-shared parameters, and the corresponding frequency domain feature and the spatial domain feature are obtained.
[0094] Step 2.2.2, the Fourier-aware feature extraction module FAFE in the encoder of the first layer processes and to obtain the frequency domain branch output feature and obtain the output feature of the spatial domain branch ;
[0095] Step 2.2.2.1: Perform Fourier transform on in the spatial dimension to obtain the frequency-domain amplitude spectrum and the frequency-domain phase spectrum ;
[0096] Step 2.2.2.2: After performing pointwise convolution on only the amplitude spectrum , and then performing inverse Fourier transform together with to obtain the output feature of the frequency domain branch ;
[0097] Step 2.2.2.3: Use the residual network module composed of lightweight and efficient depthwise separable convolution to process to obtain the output feature of the spatial domain branch .
[0098] Step 2.2.3: The spatial-frequency interaction module SFIM in the encoder of the first layer processes and to obtain the fused feature ;
[0099] Step 2.2.3.1: Perform cross gating on and respectively according to equations (1) and (2) to ensure effective spatial-frequency interaction between these two branches, promoting information exchange between the spatial domain and the frequency domain. Obtain the frequency-domain gated output feature and the spatial-domain gated output feature :
[0100] (1)
[0101] (2)
[0102] In equations (1) and (2), represents the Sigmoid activation function, represents element-wise multiplication;
[0103] Step 2.2.3.2: After concatenating and in the channel dimension and then performing pointwise convolution fusion processing, obtain the fused feature .
[0104] Step 2.2.4: According to the process of Step 2.2.1 - Step 2.2.3, It is processed in the encoder of the second layer until the fused features output by the encoder of the (N - 1)-th layer are input into the encoder of the N-th layer to obtain encoded fused features which are mapped to a distribution to obtain encoded latent variables , and after is sampled, decoded fused features are obtained , which are then input into the decoder of the first layer and processed according to the process of steps 2.2.1 - 2.2.3 until the final fused features are output by the decoder of the N-th layer .
[0105] Step 2.3, Output features After being processed by the output image embedding layer, a reconstruction result is obtained .
[0106] Step 2.4, Use equations (3) - (5) to construct the reconstruction loss function, distribution loss function , and frequency loss function of the frequency-enhanced variational autoencoder FE-VAE respectively. Among them, the reconstruction loss measures the difference between the decoder output and the original input to encourage the output to be consistent with the input. The KL divergence loss regularizes the latent space to ensure that the representation conforms to the prior distribution. The frequency loss function further maintains the frequency-domain consistency of the reconstruction result. This regularization enhances the coherence and continuity of the structure, thus generating consistent reconstruction results for similar inputs, and the total loss of the first stage is obtained using equation (6) , and the FE-VAE is trained in the first stage to obtain the frequency-enhanced variational autoencoder after the first-stage training:
[0107] (3)
[0108] (4)
[0109] (5)
[0110] (6)
[0111] In equations (3) - (6), represents the KL divergence, is the approximate posterior distribution of the encoded latent variable of the given input image , is the prior distribution of the encoded latent variable , and follows a standard Gaussian distribution, , There are 3 weighting coefficients.
[0112] Step 3, as Figure 2 shown, construct a frequency-enhanced VAE super high-definition image restoration network, including: an encoder frequency-domain replay module, a latent space image restoration network, and a decoder high-frequency information injection module; the encoder frequency-domain replay module and the decoder high-frequency information injection module together constitute a wavelet transform adapter (WTA) for reducing high-frequency loss during encoding and reducing the inter-domain gap of the encoder in the degradation domain. This adapter allows for efficient fine-tuning while keeping the pre-trained FE-VAE frozen during the restoration process. By only updating the parameters of the adapter, the FE-VAE can adapt to unknown degradation domains without changing the original FE-VAE.
[0113] Among them, the encoder frequency-domain replay module is composed of extracting the image input embedding layer, the N-layer encoder of the frequency-domain enhancement processing module in the pre-trained frequency-enhanced variational autoencoder, and adding an adapter encoding insertion module after each layer of the encoder.
[0114] The decoder high-frequency information injection module is composed of extracting the N-layer decoder of the frequency-domain enhancement processing module in the pre-trained frequency-enhanced variational autoencoder and the image output embedding layer, and adding an adapter decoding insertion module after each layer of the decoder.
[0115] Step 3.1 The encoder frequency-domain replay module processes to obtain the Nth-layer fused enhanced feature , the Nth-layer prior output feature and the Nth-layer high-frequency component feature ;
[0116] Step 3.1.1 The image input embedding layer of the pre-trained frequency-enhanced variational autoencoder in the encoder frequency-domain replay module extracts features from to obtain the degraded encoded feature ;
[0117] Step 3.1.2 When i = 1, the ith-layer encoder in the frequency-domain enhancement processing module of the pre-trained frequency-enhanced variational autoencoder in the encoder frequency-domain replay module processes to obtain the ith-layer fused feature .
[0118] Step 3.1.3 and are input into the ith-layer adapter encoding insertion module for processing to obtain the ith-layer fused enhanced feature ;
[0119] Step 3.1.3.1 When i = 1, Denote the (i - 1)-th layer prior as the input of the i-th layer adapter encoding insertion module, ; The i-th layer adapter encoding insertion module performs wavelet transform on to obtain the i-th layer sub-bands , and includes: low-frequency components , high-frequency components in the horizontal direction , high-frequency components in the vertical direction , and high-frequency components in the diagonal direction ;
[0120] Step 3.1.3.2 Construct the Frequency-Aware Adaptive Modulation Module (FAAM). After fusing with the four frequency components in respectively, four fused features of the i-th layer are obtained to capture multi-frequency feature representations. Then, through the convolutional layer and the sigmoid activation function, the four fused features of the i-th layer are respectively mapped to the four weights of the filter kernel. Thus, using the four weights of the filter kernel to modulate the spectrum of respectively, four modulated spectrum features are obtained and pointwise convolution operations are performed, thereby obtaining the i-th layer frequency fusion feature .
[0121] Step 3.1.3.2 Use Equation (7) to obtain the modulated sub-bands of the i-th layer , and after concatenating and on the channel, use three pointwise convolutions to decompose the concatenated features, obtaining the i-th layer fusion enhancement feature , the i-th layer prior and the i-th layer high-frequency coding feature ;
[0122] (7)
[0123] In Equation (7), represents average pooling, represents the Sigmoid activation function, represents element-wise multiplication;
[0124] Step 3.1.3.3 After performing zero convolution processing on , add it to to obtain the i-th layer fusion enhancement feature ; Among them, the zero convolution layer is initialized with zeros to ensure that the dimensions and characteristics of the original coding feature are retained.
[0125] Step 3.1.4 When i = 2, 3, …, N, use the i - 1-th layer fusion enhancement feature As the input of the i-th layer encoder, insert the i-1 layer prior output by the i-1 layer adapter encoding insertion module As the i-th layer prior of the i-th layer adapter encoding insertion module, and process according to the process of steps 3.1.2 - 3.1.3 and until the N-1 layer fusion feature output by the N-1 layer encoder and the N-1 layer prior output by the N-1 layer adapter encoding insertion module are input into the N-th layer adapter encoding insertion module for processing, and the N-th layer fusion enhanced feature , the N-th layer prior output feature and the N-th layer high-frequency component feature are output
[0126] Step 3.2 Construct a latent space restoration network to map from the degraded distribution to the clean distribution, and encode into the degraded latent variable and then process it to obtain the clean latent variable , so that after sampling in , the input feature of the decoder is obtained; when the ultra-high-definition input image is mapped from the pixel space to the latent space, its spatial size is reduced, and the feature distance between the clear image and the degraded image becomes closer. This helps to use a simple image restoration network in the latent space. Since our VAE already provides rich feature representations, high-quality restoration effects can be achieved without additional complex designs. Therefore, our latent space restoration network (IRNet) only uses a small number of efficient Cubic-Mixer modules
[0127] Step 3.3 The decoder high-frequency information injection module processes , and to obtain the final restoration result ;
[0128] Step 3.3.1 When i = 1, the i-th layer decoder of the pre-trained frequency-enhanced variational autoencoder in the decoder high-frequency information injection module processes to obtain the i-th layer decoded original feature ;
[0129] When i = 1, the i-th layer adapter decoding insertion module in the decoder high-frequency information injection module performs inverse wavelet transform on and to obtain the i-th layer wavelet reconstruction feature , and thus use equations (8) and (9) to obtain the i-th layer enhanced decoded feature ;
[0130] (8)
[0131] (9)
[0132] In equations (8) and (9), represents the inverse wavelet transform, represents zero convolution.
[0133] Step 3.3.2 When i = 2, 3, …, N, the enhanced decoded feature of the (i - 1)-th layer is input into the i-th layer decoder of the pre-trained frequency-enhanced variational autoencoder in the decoder high-frequency information injection module for processing, to obtain the decoded original feature of the i-th layer ;
[0134] The i-th layer adapter decoding insertion module in the decoder high-frequency information injection module performs inverse wavelet transform on the high-frequency component features output by the layer adapter encoding insertion module and to obtain the wavelet reconstruction feature of the i-th layer , reintroducing the high-frequency details into the spatial domain, so as to obtain the enhanced decoded feature of the i-th layer using equations (6) and (7); until the enhanced decoded feature of the (N - 1)-th layer output by the (N - 1)-th layer decoder is input into the N-th layer decoder of the pre-trained frequency-enhanced variational autoencoder in the decoder high-frequency information injection module for processing, and the enhanced decoded feature of the N-th layer is output; and the N-th layer adapter decoding insertion module performs inverse wavelet transform on and to obtain the wavelet reconstruction feature of the N-th layer , so as to obtain the output enhanced decoded feature of the N-th layer .
[0135] Step 3.3.3 Input into the image output embedding layer of the decoder high-frequency information injection module to obtain the restored image .
[0136] Step 3.4 Use equations (10) and (11) to construct the restoration loss and the frequency loss respectively, so as to construct the total loss of the second stage using equation (12):
[0137] (10)
[0138] (11)
[0139] (12), where and are two weighting coefficients.
[0140] Step 3.5: Freeze the parameters of the frequency-enhanced variational autoencoder after the first-stage training, and perform the second-stage training on the frequency-enhanced VAE super-high-definition image restoration network to obtain the frequency-enhanced variational autoencoder after the second-stage training, which is used to process the degraded image to obtain the restored image .
[0141] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.
[0142] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.
[0143] Embodiment
[0144] To verify the effectiveness of the method of the present invention, in this embodiment, commonly used ultra-high-definition blurred data (UHD-bulr) is selected for training and testing. In the experiment, ultra-high-definition images with a resolution of 4K (3840×2160) are used, and the peak signal-to-noise ratio (PSNR) and the structural similarity index (SSIM) are used as evaluation indicators. The quantitative results are shown in Table 1.
[0145] Table Quantitative results of image deblurring for the UHD-bulr dataset
[0146]
[0147] The results show that the proposed design of the frequency-aware feature extraction module FAFE, the spatial frequency adaptive decomposition module SFAD, the spatial frequency interaction module SFIM, and the frequency-aware adaptive modulation module FAAM brings significant improvements to the ultra-high-definition image restoration task.
Claims
1. A super-high-definition image restoration method based on a frequency-enhanced variational autoencoder, characterized in that It is carried out according to the following steps: Step 1: Obtain a degraded image under any degradation scenario and its corresponding clear image , where represents the length of the image, represents the width of the image, represents the number of channels of the image; Step 2: Construct a Frequency Enhancement Variational Autoencoder (FE-VAE), including an image input embedding layer, a frequency domain enhancement processing module, and an image output embedding layer, and perform processing on to obtain a reconstruction result , thereby constructing the total loss of the first stage , and perform the first-stage training on the FE-VAE to obtain the frequency enhancement variational autoencoder after the first-stage training; Step 3: Construct a frequency-enhanced VAE ultra-high-definition image restoration network FEVAE-UHD, including: an encoder frequency-domain replay module, a latent space image restoration network, and a decoder high-frequency information injection module; and perform processing on to obtain a restored image , thereby constructing the total loss of the second stage ; and freeze the parameters of the frequency-enhanced variational autoencoder after the first stage of training, and perform the second stage of training on the frequency-enhanced VAE ultra-high-definition image restoration network to obtain a frequency-enhanced variational autoencoder after the second stage of training, which is used to process the degraded image to obtain a restored image .
2. The super-high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 1, wherein The said step 2 includes: Step 2.1, the image input embedding layer performs feature extraction to obtain clear encoded features ; where represents the length of the encoded feature, represents the width of the encoded feature, represents the number of channels of the encoded feature; Step 2.2, Encoding Features After being processed by the frequency-domain enhancement processing module, the output features are obtained ; Step 2.3, Output Features After being processed by the output image embedding layer, a reconstruction result is obtained ; Step 2.
4. Use Equations (3)-(5) to construct the reconstruction loss function, distribution loss function, and frequency loss function of the Frequency Enhanced Variational Autoencoder (FE-VAE) respectively, so as to obtain the total loss of the first stage using Equation (6), and conduct the first-stage training on the FE-VAE to obtain the Frequency Enhanced Variational Autoencoder after the first-stage training: , the distribution loss function , the frequency loss function , so as to obtain the total loss of the first stage using Equation (6) , and conduct the first-stage training on the FE-VAE to obtain the Frequency Enhanced Variational Autoencoder after the first-stage training: (3) (4) (5) (6) In Formula (3) - Formula (6), represents the KL divergence, is the encoded latent variable of the given input image , is the approximate posterior distribution of the encoded latent variable , is the prior distribution of the encoded latent variable and and are three weighting coefficients.
3. The super high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 2, wherein, The frequency-domain enhancement processing module in the said step 2.2 consists of an encoder with N layers and a decoder with N layers. Each layer of the encoder and decoder includes: a spatial frequency adaptive decomposition module SFAD, a frequency-aware feature extraction module FAFE, and a spatial frequency interaction module SFIM; Step 2.2.1: The spatio-frequency adaptive decomposition module SFAD in the encoder of the first layer processes to obtain frequency-domain features and spatial-domain features ; Step 2.2.2: The Fourier-aware feature extraction module FAFE in the encoder of the first layer processes and to obtain the output feature of the frequency domain branch and the output feature of the spatial domain branch ; Step 2.2.3, the Spatial-Frequency Interaction Module (SFIM) in the encoder of the first layer processes and to obtain the fused feature ; Step 2.2.4: According to the process of Step 2.2.1 - Step 2.2.3, input it into the encoder of the second layer for processing until the fused feature output by the encoder of the (N - 1)-th layer is input into the encoder of the N-th layer for processing to obtain the encoded fused feature , map it to a distribution to obtain the encoded latent variable , and then, after sampling , obtain the decoded fused feature , input it into the decoder of the first layer, and process it according to the process of Step 2.2.1 - Step 2.2.3 until the final fused feature is output by the decoder of the N-th layer .
4. The super high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 3, characterized in that, Step 2.2.1 is carried out according to the following steps: Step 2.2.1.1, for After performing channel Fourier transform, the amplitude spectrum and the phase spectrum are obtained; Step 2.2.1.2, perform point-by-point convolution on the amplitude spectrum and then perform inverse Fourier transform together with to obtain the global perception feature ; Step 2.2.1.3, respectively perform two pointwise convolution operations with non-shared parameters to correspondingly obtain frequency domain features and spatial domain features .
5. A super high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 4, characterized in that Step 2.2.2 is carried out according to the following steps: Step 2.2.2.1, perform Fourier transform on in the spatial dimension to obtain the frequency-domain amplitude spectrum and the frequency-domain phase spectrum ; Step 2.2.2.2, only amplitude spectrum After performing point-by-point convolution, and then together with perform inverse Fourier transform to obtain the output feature of the frequency domain branch ; Step 2.2.2.3, use the residual network module composed of depthwise separable convolutions to process it to obtain the output feature of the spatial domain branch .
6. The super high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 5, wherein Step 2.2.3 is carried out according to the following steps: Step 2.2.3.1: According to formula (1) and formula (2), and Perform cross-gating to obtain frequency domain gating output characteristics And spatial domain gated output features : (1) (2) In formulas (1) and (2), represents the Sigmoid activation function, represents element-wise multiplication; Step 2.2.3.2: After concatenating and along the channel dimension, perform pointwise convolution fusion processing to obtain the fused feature . 7. A super-high definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 6, characterized in that, The encoder frequency-domain replay module in the said step 3 extracts the image input embedding layer and the N-layer encoder of the frequency-domain enhancement processing module in the pre-trained frequency-enhanced variational autoencoder, and respectively adds an adapter encoding insertion module after each layer of the encoder; The decoder high-frequency information injection module extracts the N-layer decoder of the frequency-domain enhancement processing module in the pre-trained frequency-enhanced variational autoencoder and the image output embedding layer, and respectively adds an adapter decoding insertion module after each layer of the decoder; Step 3.
1. The encoder frequency-domain replay module processes to obtain the Nth-layer fusion enhanced feature , the Nth-layer prior output feature and the Nth-layer high-frequency component feature ; Step 3.2: Construct a latent space restoration network, and encode into a degraded latent variable , then perform processing to obtain a clean latent variable , so that after sampling in , the input features of the decoder are obtained ; Step 3.
3. The decoder high-frequency information injection module processes , and to obtain the final restoration result ; Step 3.
4. Construct the restoration loss and the frequency loss respectively using Equations (10) and (11), so as to construct the total loss of the second stage using Equation (12): and the frequency loss respectively, so as to construct the total loss of the second stage using Equation (12): (10) (11) (12) In formula (12), and are two weighting coefficients.
8. A super high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 7, characterized in that Step 3.1 is carried out according to the following steps: Step 3.1.1: The image input embedding layer of the pre-trained frequency enhancement variational autoencoder in the encoder frequency domain replay module performs feature extraction to obtain the degraded encoded features ; Step 3.1.
2. When i = 1, the i-th layer encoder in the frequency domain enhancement processing module of the pre-trained frequency enhancement variational autoencoder in the encoder frequency domain replay module processes to obtain the i-th layer fusion feature ; Step 3.1.3, and input it into the i-th layer adapter encoding insertion module for processing to obtain the i-th layer fusion enhanced feature ; Step 3.1.
4. When \(i = 2, 3, \ldots, N\), use the fused enhanced feature of the \((i - 1)\)th layer as the input of the \(i\)th layer encoder, and use the \((i - 1)\)th prior output by the \((i - 1)\)th adapter encoding insertion module as the \(i\)th prior of the \(i\)th adapter encoding insertion module, and process and in accordance with the process of Steps 3.1.2 - 3.1.3 until the \((N - 1)\)th fused feature output by the \((N - 1)\)th layer encoder and the \((N - 1)\)th prior output by the \((N - 1)\)th adapter encoding insertion module are input into the \(N\)th adapter encoding insertion module for processing, and output the \(N\)th fused enhanced feature , the \(N\)th prior output feature and the \(N\)th high-frequency component feature .
9. A method for ultra-high definition image restoration based on a frequency-enhanced variational autoencoder according to claim 8, characterized in that, Step 3.1.3 is carried out according to the following steps: Step 3.1.3.1: When i = 1, insert as the prior of the (i - 1)-th layer into the i-th layer adapter coding insertion module, denoted as ; the i-th layer adapter coding insertion module performs wavelet transform on to obtain the i-th layer subbands , and includes: low-frequency component , high-frequency component in the horizontal direction , high-frequency component in the vertical direction , and high-frequency component in the diagonal direction ; Step 3.1.3.2: After is fused with the four frequency components in respectively, four fused features of the i-th layer are obtained accordingly. Then, through a convolutional layer and a sigmoid activation function, the four fused features of the i-th layer are respectively mapped to the four weights of the filter kernel. Thus, the spectrum of is modulated by the four weights of the filter kernel, and four modulated spectral features are obtained accordingly and a pointwise convolution operation is performed, thereby obtaining the frequency fusion feature of the i-th layer; Step 3.1.3.2: Obtain the modulated subbands of the i-th layer using Equation (7) , and after cascading and on the channel, use three pointwise convolutions to decompose the cascaded features to obtain the i-th layer fusion enhanced features , the i-th layer prior and the i-th layer high-frequency encoded features ; (7) In formula (7), represents average pooling, represents the Sigmoid activation function, represents element-wise multiplication; Step 3.1.3.3, after performing zero convolution processing on , add it to to obtain the fused and enhanced feature of the i-th layer .
10. A super high-definition image restoration method based on a frequency-enhanced variational autoencoder according to claim 9, characterized in that, Step 3.3 is carried out according to the following steps: Step 3.3.
1. When i = 1, the i-th layer decoder of the pre-trained frequency enhancement variational auto-encoder in the decoder high-frequency information injection module processes to obtain the i-th layer decoded original feature ; When i = 1, the i-th layer adapter decoding insertion module in the decoder high-frequency information injection module performs and inverse wavelet transform to obtain the i-th layer wavelet reconstruction feature , and then uses equations (8) and (9) to obtain the i-th layer enhanced decoding feature ; (8) (9) In Equations (8) and (9), represents the inverse wavelet transform, represents zero convolution; Step 3.3.
2. When i = 2, 3, …, N, the (i - 1)-th layer of enhanced decoded features are input into the i-th layer decoder of the pre-trained frequency-enhanced variational autoencoder in the decoder high-frequency information injection module for processing to obtain the i-th layer decoded original features ; The i-th layer adapter decoding insertion module in the decoder high-frequency information injection module decodes and inserts the high-frequency component features output by the layer adapter encoding insertion module and the wavelet reconstruction features of the (i - 1)-th layer to perform inverse wavelet transform to obtain the wavelet reconstruction features of the i-th layer , and then use equations (8) and (9) to obtain the enhanced decoding features of the i-th layer ; until the enhanced decoding features of the (N - 1)-th layer output by the (N - 1)-th layer decoder are input into the N-th layer decoder of the pre-trained frequency-enhanced variational autoencoder in the decoder high-frequency information injection module for processing, and the enhanced decoding features of the N-th layer are output ; and the N-th layer adapter decoding insertion module performs inverse wavelet transform on the and the wavelet reconstruction features of the (N - 1)-th layer to obtain the wavelet reconstruction features of the N-th layer ; Step 3.3.3: Inject the image output embedding layer of the input decoder high-frequency information injection module to obtain a restored image .
Citation Information
Cited By
Self-adaptive heterogeneous expert image restoration method based on frequency guidance
CN122222851A
A Frequency-Guided Adaptive Heterogeneous Expert Image Restoration Method
CN122222851B