Compressed aperture coding image super-resolution reconstruction method based on learnable wavelet enhancement network

Through the compressed aperture coded image super-resolution reconstruction method based on the learnable wavelet enhancement network, the alternating direction multiplier method and the U-shaped codec network are used to solve the problem of insufficient retention of interpretability and details of the existing methods, and the efficient image super-resolution reconstruction effect is achieved.

CN120339076APending Publication Date: 2025-07-18NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510508658.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing image super-resolution reconstruction methods have shortcomings in interpretability, image smoothing and detail retention, especially methods based on traditional physical models lead to image blur, while methods based on deep learning lack interpretability.

Method used

The super-resolution reconstruction method of compressed aperture coded image based on the learnable wavelet enhancement network is adopted. The constraint optimization problem is decomposed into two sub-problems of projection and denoising through the alternating direction multipliers method, and the U-shaped codec network is used for iterative solution. Combined with the learnable wavelet enhancement downsampling module and the spatial channel residual block, the interpretability and detail retention ability of the model are improved.

Benefits of technology

It has achieved a significantly improved interpretability and detail retention effect in image super-resolution reconstruction, and it has shown that it has superior performance in peak signal-to-noise ratio and structural similarity evaluation indicators through qualitative and quantitative analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339076A_ABST
    Figure CN120339076A_ABST
Patent Text Reader

Abstract

The invention discloses a compressed aperture coding image super-resolution reconstruction method based on a learnable wavelet enhancement network. The compressed aperture coding image super-resolution reconstruction method comprises the steps of determining a degradation model during imaging of a compressed aperture coding image; constructing a compressed aperture coding image super-resolution reconstruction problem model based on the degradation model, and rewriting the problem model into a constraint optimization problem; a constraint optimization problem is solved and decomposed into a projection sub-problem and a denoising sub-problem through an alternating direction multiplier method, the two sub-problems are solved alternately and iteratively, a super-resolution image is obtained, the projection sub-problem obtains a closed-form solution through derivation, and the denoising sub-problem is updated and solved through a U-shaped coding and decoding network. The image subjected to super-resolution reconstruction through the method is excellent in peak signal-to-noise ratio and structural similarity evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image super-resolution reconstruction, and specifically relates to a compressed aperture coding image super-resolution reconstruction method based on a learnable wavelet enhancement network. Background Art

[0002] Image super-resolution reconstruction technology aims to recover high-resolution images from low-resolution images. Its core goal is to improve the detail performance and quality of images, especially for image enhancement when the original images are limited. This technology has far-reaching significance and is widely applied in many important fields, including medical imaging, satellite remote sensing, security monitoring, video enhancement, virtual reality, etc. With the increasing demand for high-quality images in these fields, super-resolution reconstruction technology has played a key role in enhancing image details, improving visual experience, and enriching information content.

[0003] In the field of image super-resolution reconstruction, the images reconstructed by algorithms usually suffer from deficiencies in fine texture and edge information and the lack of high-frequency information. Existing super-resolution reconstruction methods can be divided into two categories: those based on traditional physical models and those based on deep learning. Among the methods based on traditional physical models, the Bicubic method performs weighted averaging using 16 neighboring pixels around the target pixel. This method is computationally simple; however, it introduces blurring when processing images, details are lost, and the image edges are relatively smooth and not sharp enough. Among the methods based on deep learning for super-resolution reconstruction, methods such as SRCNN, SRResNet, SwinIR, and RRDB-LINF-LP are based on black-box models and lack interpretability. Summary of the Invention

[0004] To overcome the above-mentioned deficiencies of the prior art, the present invention provides a compressed aperture coding image super-resolution reconstruction method based on a learnable wavelet enhancement network to improve the deficiencies of existing image super-resolution reconstruction algorithms in terms of interpretability, image smoothing, and detail preservation.

[0005] The technical solution for achieving the purpose of the present invention is as follows: A compressed aperture coding image super-resolution reconstruction method based on a learnable wavelet enhancement network, comprising the following steps:

[0006] Step 1: Determine the degradation model during the imaging of the compressed aperture coding image;

[0007] Step 2: Based on the degradation model, construct a super-resolution reconstruction problem model for the compressed aperture coding image and rewrite the problem model as a constrained optimization problem;

[0008] Step 3: Solve the constrained optimization problem by the alternating direction method of multipliers, decompose it into two sub-problems of projection and denoising, and alternately iterate to solve the two sub-problems to obtain the super-resolution image, where the projection sub-problem obtains a closed-form solution by taking the derivative, and the denoising sub-problem is updated and solved by a U-shaped encoding and decoding network.

[0009] Preferably, the specific method for determining the degradation model during the imaging of the compressive aperture encoded image is as follows:

[0010] Select high-resolution images from the dataset to generate compressive aperture encoded images. The process of generating the compressive aperture encoded image y from the high-resolution image x is represented as the degradation model, specifically:

[0011] y = Hx + z

[0012] where H represents the encoding matrix, x represents the high-resolution image, and y represents the compressive aperture encoded image.

[0013] Preferably, the specific method for constructing the super-resolution reconstruction model of the compressive aperture encoded image based on the degradation model is as follows:

[0014] The super-resolution reconstruction problem model of the compressive aperture encoded image is the inverse problem of the degradation model, and its goal is to recover x from the observed data y. Thus, a regularized inverse problem optimization model is established: The super-resolution reconstruction problem model of the compressive aperture encoded image is represented by the augmented Lagrangian method:

[0015]

[0016] where Φ(x), λ1, and γ1 represent the prior regularization, Lagrange multiplier, and penalty parameter, respectively.

[0017] Preferably, the specific method for rewriting the problem model as a constrained optimization problem is as follows:

[0018] Rewrite the problem model as:

[0019]

[0020] Introduce an auxiliary variable v to transform the problem model into a constrained optimization problem:

[0021]

[0022] Preferably, the specific method for decomposing the solution of the constrained optimization problem into two sub-problems of projection and denoising by the alternating direction method of multipliers is as follows:

[0023] By the alternating direction method of multipliers, write the super-resolution reconstruction problem model of the compressive aperture encoded image as:

[0024]

[0025] Among them, λ2 and γ2 respectively represent the Lagrange multiplier and the penalty parameter in the update process of variable v;

[0026] According to the solution method of the alternating direction method of multipliers, the model of the compressed aperture coded image super-resolution reconstruction problem is divided into two solution problems: updating x and updating v:

[0027]

[0028] Among them, i represents the i-th iteration process. In the optimization iteration, the U-shaped encoder-decoder network is first used to update v:

[0029]

[0030] There is an analytical solution in the process of updating x, and the analytical solution is obtained by taking the derivative of x and setting the derivative to zero:

[0031]

[0032] Compared with the prior art, the features of the present invention are as follows: (1) The ADMM algorithm of the present invention is used as the core architecture, and the denoising step in the iterative process is expanded into a network, realizing the interpretability of the network model. (2) The present invention designs a downsampling module based on learnable wavelet enhancement. This module sets the wavelet basis as a learnable parameter, and combines wavelet transform with parameter-free attention, which not only enhances the feature extraction ability of the model, but also improves the flexibility of the model. (3) The present invention designs an iterative connection module based on spatial-channel residual blocks. The introduction of this module enables the network to more effectively process the feature information from the previous iteration process, and while not occupying too much computing resources, it better retains the detail information of the image.

[0033] The present invention will be further described below with reference to the accompanying drawings of the specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic flowchart of a method for super-resolution reconstruction of a compressed aperture coded image based on a learnable wavelet enhancement network according to the present invention.

[0035] Figure 2 It is a network structure diagram of the present invention.

[0036] Figure 3The processing results of a set of test images on various algorithms, where the images are respectively (a) the ground truth image, (b) the magnified image of the region, (c) the test result image of the Bicubic algorithm, (d) the test result image of the SRCNN algorithm, (e) the test result image of the SRResNet algorithm, (f) the test result image of the SwinIR algorithm, (g) the test result image of the RRDB-LINF-LP algorithm, and (h) the test result image of the algorithm of the present invention.

[0037] Figure 4 The processing results of a set of test images on various algorithms, where the images are respectively ((a) the ground truth image, (b) the magnified image of the region, (c) the test result image of the Bicubic algorithm, (d) the test result image of the SRCNN algorithm, (e) the test result image of the SRResNet algorithm, (f) the test result image of the SwinIR algorithm, (g) the test result image of the RRDB-LINF-LP algorithm, and (h) the test result image of the algorithm of the present invention.

[0038] Figure 5 The processing results of a set of test images on various algorithms, where the images are respectively (a) the ground truth image, (b) the magnified image of the region, (c) the test result image of the Bicubic algorithm, (d) the test result image of the SRCNN algorithm, (e) the test result image of the SRResNet algorithm, (f) the test result image of the SwinIR algorithm, (g) the test result image of the RRDB-LINF-LP algorithm, and (h) the test result image of the algorithm of the present invention.

[0039] Figure 6 The processing results of a set of test images on various algorithms, where the images are respectively (a) the ground truth image, (b) the magnified image of the region, (c) the test result image of the Bicubic algorithm, (d) the test result image of the SRCNN algorithm, (e) the test result image of the SRResNet algorithm, (f) the test result image of the SwinIR algorithm, (g) the test result image of the RRDB-LINF-LP algorithm, and (h) the test result image of the algorithm of the present invention. Specific implementation

[0040] In the field of image super-resolution reconstruction, existing algorithms have problems in interpretability and detail restoration. To address this challenge, the present invention proposes a method for super-resolution reconstruction of compressive aperture coded images based on learnable wavelet enhancement. By designing the network structure, the present invention constructs an ADMM algorithm as the main solution architecture and unfolds the ADMM optimization iteration process into a U-shaped encoder-decoder network. The U-shaped encoder-decoder network consists of a learnable wavelet enhanced downsampling module, an upsampling module, and a spatial channel residual block. Finally, qualitative and quantitative analyses are performed on the super-resolution images obtained by the present invention and the denoised images obtained by other traditional algorithms and deep learning algorithms. The analysis results show that the present invention has significant performance advantages in image super-resolution reconstruction, can effectively perform super-resolution reconstruction on low-resolution images, and achieves the expected research goals.

[0041] A method for super-resolution reconstruction of compressive aperture coded images based on learnable wavelet enhancement, comprising the following steps:

[0042] Step 1, determine the degradation model during the imaging of compressive aperture coded images, and construct a super-resolution reconstruction problem model and an objective function for compressive aperture coded images based on the degradation model; the specific method is:

[0043] Select high-resolution images from the dataset to generate compressive aperture coded images. The process of generating compressive aperture coded images y from high-resolution images x is represented by the following degradation model:

[0044] y = Hx + z

[0045] where H represents the encoding matrix, x represents the high-resolution image, and y represents the compressive aperture coded image.

[0046] Step 2, determine the super-resolution reconstruction model for compressive aperture coded images based on the degradation model, and determine that the goal is to recover x from the observed data y. To this end, a regularized inverse problem optimization model is established to represent the super-resolution reconstruction problem model of compressive aperture coded images by the augmented Lagrangian method:

[0047]

[0048] where Φ(x), λ1, and γ1 represent prior regularization, Lagrange multiplier, and penalty parameter, respectively;

[0049] For more clarity, rewrite the problem model as:

[0050]

[0051] To further optimize the solution, introduce an auxiliary variable v to transform the above problem into a constrained optimization problem:

[0052]

[0053] Step 3: Use the alternating direction method of multipliers to solve the optimization problem of the compressed aperture coded image super-resolution reconstruction model, and decompose the optimization problem into two sub-optimization problems. The specific method is as follows:

[0054] Using the alternating direction method of multipliers, the compressed aperture coded image super-resolution reconstruction problem model is written as:

[0055]

[0056] where λ2 and γ2 respectively represent the Lagrange multiplier and the penalty parameter in the update process of the variable v. According to the solution method of the alternating direction method of multipliers, the compressed aperture coded image super-resolution reconstruction problem model is divided into two solution problems: updating x and updating v:

[0057]

[0058] where i represents the i-th iteration process. In the optimization iteration, first use the U-shaped encoding and decoding network to update v:

[0059]

[0060] There is an analytical solution in the process of updating x, and the analytical solution is obtained by taking the derivative of x and setting the derivative to zero:

[0061]

[0062] In the present invention, the process of using the alternating direction method of multipliers to solve the optimization problem of the compressed aperture coded image super-resolution reconstruction model is also expanded into a deep unfolding network for training.

[0063] The process of each training contains a set number of iteration processes, and the output after the iteration ends is used as the final result. After calculating the loss function, the learnable parameters λ1, γ1, λ2, γ2 in the deep unfolding network and the parameters in the U-shaped encoding and decoding network are updated.

[0064] In a further embodiment, the constructed U-shaped encoding and decoding network includes an input layer, an encoder, a decoder, an output layer, and a spatial channel residual block. The input layer is used to input the compressed coded image and includes two convolutional layers. The encoder part includes 2 learnable wavelet enhancement downsampling modules, the decoder part includes 2 upsampling modules, and the output layer includes two convolutional layers. The specific structure of the U-shaped encoding and decoding network is as follows:

[0065] Input layer: In the i-th iteration process, an image of H*W with an input channel number of C is input After passing through two convolutional layers, a feature map F1 of H*W with an output channel number of 128 is output i ;

[0066] Spatial Channel Residual Block 1: Using the feature map F1 output from the input layer in the (i - 1)-th iteration process i-1 as the input, after passing through a spatial channel residual block, output a feature map of H * W with 128 output channels;

[0067] Concatenation Layer 1: Concatenate the output of Spatial Channel Residual Block 1 and the output of the input layer;

[0068] Learnable Wavelet Enhancement Downsampling Module 1: Using the output of Concatenation Layer 1 as the input, output a feature map of (H / 2) * (W / 2) with 256 output channels

[0069] Spatial Channel Residual Block 2: Using the feature map output from Learnable Wavelet Enhancement Downsampling Module 1 in the (i - 1)-th iteration process as the input, after passing through a spatial channel residual block, output a feature map of (H / 2) * (W / 2) with 256 output channels;

[0070] Concatenation Layer 2: Perform a concatenation operation on the output of Spatial Channel Residual Block 2 and the output of Learnable Wavelet Enhancement Downsampling Module 1;

[0071] Learnable Wavelet Enhancement Downsampling Module 2: Using the output of Concatenation Layer 2 as the input, output a feature map of (H / 4) * (W / 4) with 512 output channels

[0072] Spatial Channel Residual Block 3: Using the feature map output from Learnable Wavelet Enhancement Downsampling Module 2 in the (i - 1)-th iteration process as the input, after passing through a spatial channel residual block, output a feature map of (H / 4) * (W / 4) with 512 output channels;

[0073] Concatenation Layer 3: Perform a concatenation operation on the output of Spatial Channel Residual Block 3 and the output of Learnable Wavelet Enhancement Downsampling Module 2;

[0074] Upsampling Module 1: Using the output of Concatenation Layer 3 as the input, obtain a feature map of (H / 2) * (W / 2) with 256 channels through upsampling

[0075] Fusion Layer 1: Add and fuse the output of Learnable Wavelet Enhancement Downsampling Module 2 and the output of Upsampling Module 1 for addition and fusion;

[0076] Spatial Channel Residual Block 4: Using the feature map output from Fusion Layer 1 in the (i - 1)-th iteration process as the input, after passing through a spatial channel residual block, output a feature map of (H / 2) * (W / 2) with 256 output channels;

[0077] Concatination layer 4: Perform a concatination operation on the output of the spatial channel residual block 4 and the output of the fusion layer 1;

[0078] Upsampling module 2: Use the feature map output by the concatination layer 3 as the input. Add the feature map with 128 channels and size H*W obtained by upsampling to the feature map output by the input layer, and output a feature map with 128 channels and size H*W

[0079] Fusion layer 2: The output of the upsampling module 2 and the output F1 of the input layer i are added and fused;

[0080] Spatial channel residual block 5: Use the output of the fusion layer 2 in the (i - 1)-th iteration process as the input. After passing through a spatial channel residual block, output a feature map with 128 channels and size H*W;

[0081] Concatination layer 5: Perform a concatination operation on the output of the spatial channel residual block 5 and the output of the fusion layer 2, and output a feature map with 128 channels and size H*W;

[0082] Output layer: Use the feature map output by the concatination layer 5 as the input. After passing through two convolutional layers, obtain an image with C channels and size H*W

[0083] Fusion layer 3: The output of the output layer and the output F1 of the input layer i are added and fused to obtain the desired v i 。

[0084] Step 4.1: Construct a wavelet enhanced downsampling module, including a learnable wavelet decomposition layer, a parameter-free attention enhancement layer, and a dynamic feature fusion layer; The specific structure of the multi-scale residual block is as follows:

[0085] Learnable wavelet decomposition layer: Use the obtained feature map with size M*N*C as the input. After decomposing the input into four subbands through a learnable wavelet basis convolution kernel, divide the 4 subbands into one low-frequency component and three high-frequency components, perform pixel-by-pixel addition and fusion on the high-frequency components, and output a low-frequency and high-frequency fusion feature map with 2C channels and size (M / 2)*(N / 2);

[0086] Parameter-free attention enhancement layer 1: Input the low-frequency feature map with 2C channels and size (M / 2)*(N / 2) from the learnable wavelet decomposition layer, calculate the importance weight of each position using the mean and variance of the eigenvalues, and output an enhanced low-frequency feature map with 2C channels and size (M / 2)*(N / 2);

[0087] Parameter-free attention enhancement layer 2: Input the high-frequency fusion feature map of (M / 2)*(N / 2) with 2C input channels from the learnable wavelet decomposition layer, calculate the importance weight of each position using the mean and variance of the eigenvalues, and output the enhanced high-frequency fusion feature map of (M / 2)*(N / 2) with 2C channels.

[0088] Dynamic feature fusion layer: Take the outputs of parameter-free attention enhancement layer 1 and parameter-free attention enhancement layer 2 as inputs, multiply the activation values of the low-frequency component and the high-frequency component element-wise to obtain the modulated low-frequency feature map, multiply the activation values of the high-frequency component and the low-frequency component element-wise to obtain the modulated high-frequency feature map, concatenate the obtained modulated low-frequency feature map and modulated high-frequency feature map, and compress them through a convolutional layer to output the enhanced fused feature map of (M / 2)*(N / 2) with 2C channels.

[0089] Step 4.2: Construct a spatial-channel residual block, including spatial feature processing and channel processing; the specific structure of the spatial-channel residual block is as follows:

[0090] Spatial feature processing: Take the obtained M*N*C feature map as input, and pass it through a normalization layer, a depthwise separable convolution, a pointwise convolution, and Dropout respectively to obtain an M*N*C feature map. Further, perform a residual connection with the original input and output an M*N*C feature map.

[0091] Channel processing: After passing the feature output by the spatial feature processing through batchnorm and MLP, perform a residual connection with the initial input feature and output an M*N*C feature map.

[0092] Step 4.3: Construct an upsampling module, including a basic feature extraction layer and interpolation; the specific structure of the upsampling module is as follows:

[0093] Basic feature extraction layer: Take the obtained M*N*C feature map as input, pass it through two convolutional layers, and output an M*N feature map with C / 2 output channels.

[0094] Interpolation: Take the feature map obtained by the basic feature extraction layer as input, and after bilinear interpolation, output a (2M)*(2N) feature map with C / 2 channels.

[0095] In a further embodiment, the designed network structure is trained using a compressed coding image dataset until the maximum number of iterations is reached to obtain a trained network model. The training process is as follows:

[0096] First, preprocess the training images. Each image patch undergoes data augmentation operations such as random rotation, flipping, Mixup, etc., and the training image pairs are loaded in batches of size 16.

[0097] Then, the hyperparameters in the training process are defined. The selection of hyperparameters depends on past experimental experience. Among them, the number of training epochs is set to 500, and the optimizer selected in the present invention is the Adam optimizer. The initial learning rate is set to 3×10−4, and the ReduceLROnPlateau learning rate scheduler is used to adjust the learning rate.

[0098] The training image pairs are input into the designed network for training. The network model is verified and the current model is saved every 10 training epochs until the maximum number of training epochs is reached, and the last model is saved as the final model.

[0099] The compressed coded image test set is input into the trained network model to obtain the denoised image, and the obtained denoised image is compared and evaluated with the ground truth image.

[0100] Example

[0101] In this experimental example, the method proposed in the present invention is compared and tested with a traditional method Bicubic and four deep learning algorithms SRCNN, SRReaNet, SwinIR, and RRDB-LINF-LP in the DIV2K dataset. Among them, the Bicubic algorithm is a weighted average method based on neighboring pixels, which realizes super-resolution reconstruction by using 16 neighboring pixels of the input image to estimate the value of the target pixel. As the first CNN-based super-resolution network, SRCNN uses a three-layer convolutional structure to learn end-to-end mapping. However, limited by the shallow structure, it is difficult to model complex textures. SRResNet introduces residual learning and sub-pixel convolution, and constructs a deep network through 16 residual blocks. However, there is a problem of excessive high-frequency sharpening. SwinIR introduces the SwinTransformer architecture, which uses the window multi-head self-attention mechanism to model long-range dependencies. In this way, SwinIR can capture richer features, especially suitable for processing complex image details. RRDB-LINF-LP introduces learned priors to help capture the correlation between image patches, thereby reducing the discontinuity between different patches in the generated image and effectively solving the grid artifact problem.

[0102] The test results of each method are as Figure 3As shown, which are in sequence (a) true value image, (b) magnified image of the true value area, (c) test result graph of the Bicubic algorithm, (d) test result graph of the SRCNN algorithm, (e) test result graph of the SRResNet algorithm, (f) test result graph of the SwinIR algorithm, (g) test result graph of the RRDB-LINF-LP algorithm, and (h) test result graph of the algorithm of the present invention. The Bicubic algorithm performs excellently in preserving the original image information, but its ability to restore details and process edges is limited, resulting in blurred images and smoothed edges in most test results. The SRCNN algorithm has achieved remarkable results in image clarity. However, with only a few convolutional layers, it cannot extract overly complex high-frequency details. The SRResNet algorithm introduces a deep network design, enhancing the network's learning ability. Nevertheless, it still faces problems such as limited high-frequency detail restoration, insufficient texture accuracy, unstable training, and smoothed edges. The SwinIR algorithm can capture global information well, thus restoring finer details. However, it has a large computational overhead and poor model interpretability. The RRDB-LINF-LP algorithm has reached a relatively high level in the effect of image super-resolution reconstruction, but there is still room for improvement in detail preservation. The method proposed by the present invention, by constructing a deep unfolding network, while having interpretability, optimizes the preservation of image detail information, thus achieving the best results in image quality.

[0103] In addition, the commonly used objective evaluation indicators for image super-resolution reconstruction test experiments are: peak signal-to-noise ratio (PSNR) and structural similarity (SSIM). The test results are shown in Table 1:

[0104] Table 1 Results of Evaluation Indicators

[0105] PSNR SSIM Bicubic 26.25 0.7363 SRCNN 27.74 0.7854 SRResNet 27.87 0.8013 SwinIR 28.18 0.7962 RRDB-LINF-LP 28.00 0.7889 The present invention 28.65 0.8953

[0106] Among them, PSNR can be used to evaluate the overall distortion degree of the image. The larger its value, the smaller the distortion and the higher the quality. SSIM comprehensively considers the three aspects of image brightness, contrast, and structure, and comprehensively measures the similarity of the image from an angle more in line with human visual judgment. The closer its value is to 1, the better the super-resolution reconstruction effect of the image to be evaluated. It can be seen from Table 1 that the model of the present invention performs best in both PSNR and SSIM indicators. In summary, the present invention has the best image super-resolution reconstruction effect under the experimental conditions of the present invention.

Claims

1. A super-resolution reconstruction method for compressive aperture encoded images based on a learnable wavelet enhancement network, characterized in that It includes the following steps: Step 1: Determine the degradation model during the imaging of the compressed aperture encoded image; Step 2: Based on the degradation model, construct a super-resolution reconstruction problem model for the compressed aperture encoded image, and rewrite the problem model as a constrained optimization problem; Step 3: Use the alternating direction method of multipliers to decompose the solution of the constrained optimization problem into two sub-problems: projection and denoising, and alternately iterate to solve the two sub-problems to obtain the super-resolution image. Among them, the projection sub-problem obtains a closed-form solution by taking the derivative, and the denoising sub-problem is updated and solved through a U-shaped encoder-decoder network.

2. The compressive aperture encoded image super-resolution reconstruction method based on the learnable wavelet enhancement network according to claim 1, wherein The specific method for determining the degradation model during the imaging of the compressed aperture encoded image is as follows: Select a high-resolution image from the dataset to generate a compressed aperture encoded image. The process of generating the compressed aperture encoded image y from the high-resolution image x is represented as the degradation model, specifically: y = Hx + z where, H represents the encoding matrix, x represents the high-resolution image, and y represents the compressed aperture encoded image.

3. The super-resolution reconstruction method of compressive aperture coded images based on a learnable wavelet enhancement network according to claim 2, wherein The specific method for constructing a super-resolution reconstruction model for the compressed aperture encoded image based on the degradation model is as follows: The super-resolution reconstruction problem model for the compressed aperture encoded image is the inverse problem of the degradation model. Its goal is to recover x from the observed data y. Thus, a regularized inverse problem optimization model is established: The super-resolution reconstruction problem model for the compressed aperture encoded image is represented by the augmented Lagrangian method: where, Φ(x), λ1, and γ1 represent the prior regularization, Lagrange multiplier, and penalty parameter respectively.

4. The super-resolution reconstruction method of compressive aperture coded images based on the learnable wavelet enhancement network according to claim 3, characterized in that The specific method for rewriting the problem model as a constrained optimization problem is as follows: Rewrite the problem model as: Introduce an auxiliary variable v to transform the problem model into a constrained optimization problem:

5. The super-resolution reconstruction method of compressive aperture coded images based on a learnable wavelet enhancement network according to claim 3, wherein The specific method for using the alternating direction method of multipliers to decompose the solution of the constrained optimization problem into two sub-problems: projection and denoising is as follows: Using the alternating direction method of multipliers, write the super-resolution reconstruction problem model for the compressed aperture encoded image as: where, λ2 and γ2 represent the Lagrange multiplier and penalty parameter during the update process of the variable v respectively; According to the solution method of the alternating direction method of multipliers, divide the super-resolution reconstruction problem model for the compressed aperture encoded image into two solution problems: updating x and updating v: where, i represents the i-th iteration process. In the optimization iteration, first use the U-shaped encoder-decoder network to update v: There is an analytical solution in the process of updating x, and the analytical solution is obtained by taking the derivative of x and setting the derivative to zero:

6. The compressive aperture encoded image super-resolution reconstruction method of the learnable wavelet enhancement network according to claim 1, characterized in that The U-shaped encoder-decoder network includes an input layer, an output layer, a spatial-channel residual block 1, a spatial-channel residual block 2, a spatial-channel residual block 3, a spatial-channel residual block 4, a spatial-channel residual block 5, a learnable wavelet enhanced downsampling module 1, a learnable wavelet enhanced downsampling module 2, an upsampling module 1, an upsampling module 2, a splicing layer 1, a splicing layer 2, a splicing layer 3, a splicing layer 4, a splicing layer 5, a fusion layer 1, a fusion layer 2, and a fusion layer 3; The specific process of updating v through the U-shaped encoder-decoder network is as follows: Input layer: In the $i$-th iteration, an image of size $H \times W$ with $C$ channels is input After passing through two convolutional layers, a feature map $F1$ of size $H \times W$ with 128 output channels is obtained i ; Spatial channel residual block 1: Take the feature map F1 output by the input layer in the (i - 1)-th iteration process i-1 as the input. After passing through a spatial channel residual block, output a feature map of H * W with 128 output channels; Splicing layer 1: Splice the output of the spatial-channel residual block 1 and the output of the input layer; Learnable Wavelet Enhancement Downsampling Module 1: Taking the output of Concatenation Layer 1 as input and outputting a feature map of size (H / 2)*(W / 2) with 256 output channels Spatial Channel Residual Block 2: Using the feature map output by the learnable wavelet enhancement downsampling module 1 in the (i - 1)-th iteration process as the input, after passing through a spatial channel residual block, it outputs a feature map of (H / 2)*(W / 2) with 256 output channels; Splicing layer 2: Perform a splicing operation on the output of the spatial-channel residual block 2 and the output of the learnable wavelet enhanced downsampling module 1; Learnable Wavelet Enhancement Downsampling Module 2: Taking the output of Concatenation Layer 2 as input and outputting a feature map of size (H / 4)*(W / 4) with 512 output channels Spatial channel residual block 3: Using the feature map output by the learnable wavelet enhancement downsampling module 2 in the (i - 1)-th iteration process as the input, after passing through a spatial channel residual block, a feature map of (H / 4)*(W / 4) with 512 output channels is output; Concatenation layer 3: Perform a concatenation operation on the output of the spatial-channel residual block 3 and the output of the learnable wavelet enhancement downsampling module 2; Upsampling Module 1: Taking the output of the concatenation layer 3 as input, and obtaining a feature map of (H / 2)*(W / 2) with 256 channels through upsampling Fusion layer 1: Additively fuse the output of the learnable wavelet enhancement downsampling module 2 and the output of the upsampling module 1 ; Spatial-channel residual block 4: Take the feature map of the output of the fusion layer 1 in the (i - 1)-th iteration process as the input. After passing through a spatial-channel residual block, output a feature map of (H / 2)*(W / 2) with 256 output channels; Concatenation layer 4: Perform a concatenation operation on the output of the spatial-channel residual block 4 and the output of the fusion layer 1; Upsampling module 2: Taking the feature map output by the splicing layer 3 as input, adding the feature map with 128 channels and size H*W obtained by upsampling to the feature map output by the input layer, and outputting a feature map with 128 channels and size H*W Fusion layer 2: Additively fuse the output of the upsampling module 2 and the output F1 of the input layer i together; Spatial-channel residual block 5: Take the output of the fusion layer 2 in the (i - 1)-th iteration process as the input. After passing through a spatial-channel residual block, output a feature map of H*W with 128 output channels; Concatenation layer 5: Perform a concatenation operation on the output of the spatial-channel residual block 5 and the output of the fusion layer 2, and output a feature map of H*W with 128 output channels; Output layer: Taking the feature map output by the splicing layer 5 as the input, after passing through two convolutional layers, an image of H*W with the number of channels C is obtained Fusion layer 3: Add the output of the output layer and the output F1 of the input layer i to perform additive fusion and obtain the desired v i .

7. The super-resolution reconstruction method of compressive aperture coded images based on the learnable wavelet enhancement network according to claim 6, wherein The learnable wavelet enhancement downsampling module includes a learnable wavelet decomposition layer, a parameter-free attention enhancement layer, and a dynamic feature fusion layer; the specific structure of the learnable wavelet enhancement downsampling module is: Learnable wavelet decomposition layer: Take the obtained M*N*C feature map as the input. After decomposing the input into four subbands through a learnable wavelet basis convolution kernel, divide the 4 subbands into a low-frequency component and three high-frequency components, perform pixel-wise addition fusion on the high-frequency components, and output a low-frequency and high-frequency fusion feature map of (M / 2)*(N / 2) with 2C output channels; Parameter-free attention enhancement layer 1: Input the low-frequency feature map of (M / 2)*(N / 2) with 2C input channels from the learnable wavelet decomposition layer, calculate the importance weight of each position using the mean and variance of the eigenvalues, and output an enhanced low-frequency feature map of (M / 2)*(N / 2) with 2C output channels; Parameter-free attention enhancement layer 2: Input the high-frequency fusion feature map of (M / 2)*(N / 2) with 2C input channels from the learnable wavelet decomposition layer, calculate the importance weight of each position using the mean and variance of the eigenvalues, and output an enhanced high-frequency fusion feature map of (M / 2)*(N / 2) with 2C output channels; Dynamic feature fusion layer: Take the outputs of the parameter-free attention enhancement layer 1 and the parameter-free attention enhancement layer 2 as the input, perform element-wise multiplication on the activation values of the low-frequency component and the high-frequency component to obtain a modulated low-frequency feature map, perform element-wise multiplication on the activation values of the high-frequency component and the low-frequency component to obtain a modulated high-frequency feature map, concatenate the obtained modulated low-frequency feature map and the modulated high-frequency feature map, and perform compression through a convolutional layer to output an enhanced and fused feature map of (M / 2)*(N / 2) with 2C output channels.

8. The super-resolution reconstruction method of compressive aperture coded images based on a learnable wavelet enhancement network according to claim 4, wherein The spatial-channel residual block includes spatial feature processing and channel processing; the specific structure of the spatial-channel residual block is: Spatial feature processing: Take the obtained M*N*C feature map as the input, and respectively pass through a normalization layer, a depthwise separable convolution, a pointwise convolution, and Dropout to obtain an M*N*C feature map. Further, perform a residual connection with the original input and output an M*N*C feature map. Channel processing: After processing the features output from spatial feature processing through batch normalization and MLP, a residual connection is made with the initial input features, and a feature map of M*N*C is output.

9. The compressive aperture coded image super-resolution reconstruction method based on a learnable wavelet enhancement network according to claim 4, wherein The upsampling module includes a basic feature extraction layer and interpolation; the specific structure of the upsampling module is as follows: Basic feature extraction layer: Using the obtained M*N*C feature map as input, through two convolutional layers, a feature map of M*N with an output channel number of C / 2 is output. Interpolation: Using the feature map obtained from the basic feature extraction layer as input, after bilinear interpolation, a feature map of (2M)*(2N) with an output channel number of C / 2 is output.