Two-stage lens-free image enhancement method based on physical prior and generation prior
By adopting a two-stage image enhancement method based on physical priors and generative priors, the problems of poor image edge consistency and loss of high-frequency details in lensless imaging systems are solved, and high-quality image reconstruction results are achieved. This method is applicable to scene reconstruction and quality optimization in lensless imaging systems.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TIANJIN UNIV
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-01
AI Technical Summary
Lensless imaging systems suffer from poor image edge consistency, reconstruction distortion, severe blurring, and loss of high-frequency details. Existing methods struggle to achieve a balance between physical consistency, detail richness, and computational efficiency.
A two-stage image enhancement method based on physical and generative priors is adopted. By unifying the size specifications of lensless measurement images and point spread functions, spatial adaptive deconvolution is performed to obtain low-frequency content, and a generative prior-driven adversarial high-frequency enhancement network is used to restore high-frequency details. The model is trained in stages to achieve a balance between structural consistency and detail realism.
It significantly improves the quality of lensless imaging, solves the problems of poor edge consistency and loss of high-frequency details caused by PSF spatial variations, and achieves high fidelity and high detail enhancement of images. It is applicable to mainstream lensless datasets and images of different resolutions.
Smart Images

Figure CN121961890A_ABST
Abstract
Description
A two-stage lensless image enhancement method based on physical and generative priors Technical Field
[0001] This invention relates to the fields of image processing and computer vision technology, and in particular to a lensless image enhancement method based on the collaboration of physical prior and generative prior, applicable to scene reconstruction and quality optimization of lensless imaging systems. Background Technology
[0002] As imaging technology advances towards lightweight, miniaturized, and low-cost designs, lensless imaging systems, with their significant advantages of small size, light weight, simple structure, and low manufacturing cost, have found widespread application in various fields such as medical diagnostics, security monitoring, portable electronic devices, and industrial inspection. Lensless imaging systems eliminate the optical lenses relied upon in traditional imaging. Instead, they encode scene information into complex diffraction or scattering patterns by placing modulation elements such as phase masks or diffusers in front of the sensor. The original scene image is then reconstructed through computational reconstruction algorithms, overcoming the size and cost limitations of traditional lens systems.
[0003] However, the computational reconstruction process of lensless imaging faces numerous technical challenges, making it difficult for its image quality to reach the level of traditional lens systems. In physics, the lensless imaging process can be modeled as the convolutional superposition of noise between the scene and the point spread function (PSF), that is, the observed image is a blurred measurement of the scene after being modulated by the spatially varying PSF. This modulation characteristic presents several technical challenges: First, the spatial variation of the PSF is significant. Traditional methods, which assume a single fixed PSF, cannot accurately model the imaging process across the entire field of view, leading to distortion in the reconstruction of image edges and peripheral regions, and poor structural consistency. Second, lensless measurement images generally suffer from severe blurring and loss of high-frequency details. While deconvolution methods based solely on physical models (such as Wiener filtering and Tikhonov regularization) can recover some low-frequency structures, they struggle to replenish lost high-frequency details, resulting in overly smooth reconstructions that lack visual realism. Third, while pure generative models (such as diffusion models) have emerged in recent years and can generate rich high-frequency details, they are prone to deviating from the physical constraints of the original measurement data, leading to issues such as content tampering and the injection of false details, thus compromising scene consistency. Furthermore, sensor noise and diffraction noise during the imaging process further exacerbate the reconstruction difficulty, making it difficult for existing methods to achieve a balance between physical consistency, detail richness, and computational efficiency.
[0004] These problems severely restrict the application expansion of lensless imaging systems, especially in scenarios requiring high image detail accuracy and scene realism, where existing enhancement methods cannot meet practical needs. Therefore, developing a lensless image enhancement technique that can accurately model the physical imaging process, effectively supplement high-frequency details, and simultaneously ensure scene consistency has become crucial for promoting the practical application of lensless imaging technology. To this end, this invention proposes a two-stage lensless image enhancement method based on physical priors and generative priors. Summary of the Invention
[0005] The purpose of this invention is to propose a two-stage lensless image enhancement method based on physical priors and generative priors to solve the core problems of poor edge consistency caused by PSF spatial variations in lensless imaging, loss of high-frequency details due to simple deconvolution, and easy content manipulation of pure generative models. This invention achieves a balance between physical structural consistency and the realism of generated details, effectively supplements the high-frequency details lost in lensless imaging, and has good enhancement effects on mainstream lensless datasets and images of different resolutions, significantly improving imaging quality and adaptability to subsequent applications.
[0006] To achieve the above objectives, this invention proposes the following technical solution: a two-stage lensless image enhancement method based on physical and generative priors. This method utilizes the two-stage collaborative enhancement of physical and generative priors to obtain high-fidelity and high-detail lensless enhanced images. Specifically, it includes the following steps: S1, unifying the size specifications of the lensless measurement image and the point spread function (PSF) to ensure compatibility with subsequent deconvolution operations; S2, using spatially adaptive deconvolution driven by physical priors to obtain low-frequency content; S3, recovering high-frequency details lost during lensless imaging through an adversarial high-frequency enhancement network driven by generative priors; S4, training the two-stage model in stages, and outputting the final enhanced image after inference post-processing.
[0007] Preferably, S1 specifically includes the following: S101, acquiring the original measurement image captured by a lensless camera. and the point spread function (PSF) data obtained from system calibration, including the original measurement image. It is a three-channel RGB image with a size of The point spread function (PSF) size is And satisfy S102, using the original measurement image Using the center as a reference, an image is cropped to the same size as the point spread function (PSF). To ensure the preservation of the core content region of the image and achieve frequency domain alignment with the point spread function (PSF), the calculation formula is as follows:
[0008] S103. The cropped image Mirror padding is performed to avoid periodic boundary effects during frequency domain convolution. The padding width is set to... , To ensure that the filled image I pad Adapted to the frequency domain of the point spread function (PSF), the calculation formula is as follows:
[0009] S104. Standardize the point spread function (PSF) data, normalizing its pixel values to the [0,1] interval while preserving the spatial distribution characteristics of the PSF, ensuring consistency with the filled image. The grayscale range adaptation lays the foundation for subsequent deconvolution operations.
[0010] Preferably, S2 specifically includes the following: S201, filling the image... The space is uniformly divided into K partitions, each partition corresponding to a control vertex. , The standardized point spread function (PSF) data is centered according to the partition location to obtain the point spread function (PSF) block corresponding to each partition. Ensure the point spread function (PSF) block is secure. Spatial correspondence with image partitions; S202, Point Spread Function (PSF) block for each partition Perform frequency domain transformation to obtain its frequency domain representation. The calculation formula is:
[0011] in, S203, a Wiener filter is constructed based on the frequency domain point spread function (PSF) to balance noise reduction and detail preservation. The calculation formula is as follows:
[0012] in, for The conjugate of complex numbers, The square of the modulus, S204 is the regularization parameter for the k-th partition; S204, for any pixel position in the image Calculate its distance to each control vertex The Euclidean distance is used to calculate the initial weights based on the reciprocal of the distance. The calculation formula is as follows:
[0013] in, The minimum value (set to 1e-6) is used to avoid the denominator being zero; S205, the initial weights are normalized to obtain the final weight matrix, calculated using the following formula:
[0014] Ensure the weights of each pixel position sum to 1 to achieve a smooth transition in the partitioning results; S206, process the filled image. Perform a Fourier transform to obtain the frequency domain image. ; and combine it with the Wiener filters of each partition. Element-wise multiplication, followed by inverse Fourier transform, yields the deconvolution result for each partition, as shown in the formula:
[0015] Where F represents the Fourier transform, Indicates the inverse Fourier transform.
[0016] S207. Deconvolution the results of each partition. The low-frequency content is obtained by multiplying each pixel and summing the results. The calculation formula is:
[0017] in, This indicates pixel-by-pixel multiplication. Preserving the basic structure and contours of the image lays the foundation for subsequent high-frequency enhancement.
[0018] Preferably, S3 specifically includes the following: S301, selecting a variational autoencoder with a pre-trained diffusion model to output low-frequency content. Encode into the latent space to obtain the initial latent representation. The calculation formula is:
[0019] in, This is a VAE encoder used to map images from pixel space to a low-dimensional latent space, preserving low-frequency structural information while reducing computational complexity. It's important to note that the weights of the VAE decoder are frozen throughout the VAE encoder training process. S302 uses the pre-trained diffusion model's Unet as the initial weights for the enhancement network, embedding low-rank adaptation modules in its key layers (cross-attention and residual blocks, etc.) to efficiently adapt to high-frequency enhancement tasks for lensless images. For any original weight matrix... Its effective update form is:
[0020] in, , For a trainable low-rank matrix, the rank Simultaneously, a discriminator based on the ConvNeXt architecture is introduced. This forms a conditional adversarial training framework with the augmentation network. During the training of the augmentation network, , All original UNet weights are frozen, except for the LoRA parameters and the discriminator. Learnable; S303, latent representation By directly inputting the finely tuned augmented network, the latent representation is predicted and reconstructed through a single forward propagation. The goal of augmenting the network is to improve the final output. In the discriminator It is indistinguishable from real high-definition images in the eye, thus injecting high-frequency details that conform to the statistical laws of natural images.
[0021] S304, Enhanced latent representation Input VAE decoder The reconstructed image, with added high-frequency details, is a complete enhanced image. The calculation formula is as follows:
[0022] in, Retained It maintains structural consistency while supplementing the texture and high-frequency edge details lost in lensless imaging.
[0023] Preferably, S4 specifically includes the following: S401, fixing all parameters of the prior generation module, and training the diffusion function of each partition point of the adaptive deconvolutional network. and regularization parameters The optimization objective is to ensure the consistency of low-frequency content reconstruction. The loss function is a combination of MSE loss and LPIPS perceptual loss, expressed as follows:
[0024] in, This is for the low-frequency content output of the spatially adaptive deconvolution network. Labels for low-frequency components of real-world scene images. For mean square error loss, For the perceptual loss based on the VGG network; S402, after the physical prior module is trained, its parameters and the augmentation network UNet and discriminator are fixed. decoder The weights are adjusted only by fine-tuning the VAE encoder. This aims to improve the potential representation quality of lensless degraded images. The optimization objective is to make the response of low-frequency images consistent with that of real high-definition images in the discriminator. The loss function is defined as:
[0025] in This is for the low-frequency content output of the spatially adaptive deconvolution network. Real-world scene images It is a fine-tuning encoder. It is a fixed decoder; S403, fixed first-stage physical prior module parameters and encoder-decoder parameters, train the LoRA adaptation layer and discriminator of the augmentation network, utilize the natural image modeling capability of the generated prior to optimize the realism of high-frequency details, and the loss function adopts a multi-objective combination form:
[0026] in, To generate the high-frequency enhanced image output by the prior module, Real-world scene images For mean square error loss, For the perceptual loss based on VGG networks, To address Wasserstein's adversarial loss, UNet enhances the network. A discriminator based on the ConvNeXt backbone; These are the weighting coefficients for each loss term.
[0027] The present invention further protects a computer device, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the instruction, program, code set or instruction set being loaded and executed by the processor to implement the above-described two-stage lensless image enhancement method based on physical prior and generative prior.
[0028] The present invention further protects a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the instruction, program, code set, or instruction set is loaded and executed by a processor to implement the above-described two-stage lensless image enhancement method based on physical prior and generative prior.
[0029] Compared with the prior art, the present invention provides a two-stage lensless image enhancement method based on physical prior and generative prior, which has the following beneficial effects: The present invention ensures frequency domain adaptation by unifying the size specifications of the lensless measurement image and PSF; then, spatial adaptive deconvolution is used to obtain low-frequency content that retains the basic structure; then, a pre-trained diffusion model is used to restore the high-frequency details lost in lensless imaging; finally, the two-stage modules are trained in stages, and the lensless enhanced image is output after inference post-processing, taking into account both structural consistency and detail realism. Specifically, the following beneficial effects are included: (1) It solves the core problem of inaccurate modeling of PSF spatial variation in lensless imaging. By generating a multi-kernel PSF set through a spatial adaptive deconvolution network, and combining partitioned frequency domain deconvolution and distance-weighted fusion, it accurately adapts to the PSF variation characteristics of the entire field of view. Compared with the traditional single convolution kernel method, it significantly improves the consistency and edge integrity of low-frequency structure reconstruction.
[0030] (2) A two-stage collaborative architecture of "physical prior + generative prior" is adopted. The first stage relies on the physical imaging law to ensure data fidelity, and the second stage uses a pre-trained diffusion model to supplement high-frequency details. This avoids the problem of missing details in simple deconvolution and overcomes the defect of easy content tampering in pure generative models, thus achieving a balance between structural consistency and detail realism. Attached Figure Description
[0031] Figure 1 is a flowchart illustrating a two-stage lensless image enhancement method based on physical prior and generative prior proposed in this invention; Figure 2 is the processed lensless measurement image in an example of this invention; Figure 3 is the point spread function corresponding to the processed lensless measurement image in an example of this invention; Figure 4 is the spatial weight map of the spatial adaptive deconvolution stage driven by physical prior in an example of this invention; Figure 5 is the low-frequency image obtained by spatial adaptive deconvolution driven by physical prior in an example of this invention; Figure 6 is the final result restored by the high-frequency enhancement network driven by generative prior in an example of this invention; Figure 7 is a visual comparison of this invention with other lensless image reconstruction algorithms on the PhlatCam test set. Detailed Implementation
[0032] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0033] Example 1: Referring to Figure 1, this invention proposes a two-stage lensless image enhancement method based on physical priors and generative priors, specifically including the following steps: S1. Selecting the PhlatCam dataset and performing preprocessing: To ensure that the lensless measurement image and the point spread function are strictly aligned in the spatial and frequency domains to adapt to subsequent deconvolution operations, this embodiment selects the PhlatCam dataset and performs a systematic preprocessing process, the specific steps of which are as follows: S101. Overview of the PhlatCam dataset: The PhlatCam dataset contains 10,000 images covering 1,000 categories. Each image is adjusted to 384×384 pixels before display, and then optical modulation and sensor capture are completed through the lensless PhlatCam system to generate a RAW format measurement image with a resolution of 3036×4024 pixels. In this embodiment, images of 990 categories are selected for model training, and images of the remaining 10 categories are used for performance testing; S102. Performing de-mosaic and resolution conversion: For the original RAW image De-mosaic processing is performed to obtain the RGB representation, and the final output is a multi-channel image with a resolution of 1518×2012. .
[0034] S103. Cropping based on calibration center pairs: Based on the system calibration results, determine the coordinates of the alignment center point between the image and the point spread function. , Centered on this point, the de-mosaiced image... Perform a center-cropping operation on the point spread function image, setting the cropping size to [value missing]. , Obtain the cropped image This operation ensures that the core imaging region is preserved while matching the image size to the input specifications of subsequent networks. The specific calculation formula is as follows:
[0035]
[0036] S103. Perform mirror padding to adapt to frequency domain convolution: To avoid artifacts introduced during frequency domain deconvolution due to the periodic boundary assumption, the cropped image is... A mirror fill operation is performed to match the filled size with the point spread function size. In this embodiment, due to the cropped image... and the clipped point spread function Since the dimensions are consistent, the fill amount is set to 0. The specific calculation formula is as follows:
[0037] The visualization results are shown in Figure 2; S104, PSF normalization processing: The PSF data obtained from system calibration is normalized to linearly map its pixel values to the [0,1] interval. The specific calculation formula is as follows:
[0038] This operation preserves the spatial energy distribution characteristics of the PSF while also adjusting its grayscale dynamic range to match... Consistency provides numerically stable input for the first-stage deconvolution operation. Standardized The visualization results are shown in Figure 3.
[0039] S2. Low-frequency content is obtained using spatially adaptive deconvolution driven by physical priors. This step aims to utilize the system-calibrated point spread function as a physical prior, construct multiple spatially varied Wiener filters, and combine them with a spatially adaptive weight fusion mechanism to recover structurally accurate low-frequency content from the preprocessed measurement image. Simultaneously, learnable parameters are optimized through end-to-end training to improve deconvolution accuracy and low-frequency content integrity. The specific process is as follows: S201, Construct multiple sets of learnable partition point spread functions: The standardized point spread functions are... Copy 9 times to get (k=1,2,…,9), the gradients of each part can be updated, thus allowing the model to learn the optimal diffusion function variants for different spatial regions.
[0040] S202, Construct a spatial Wiener filter for each partition point spread function kernel: For each Perform frequency domain transformation to obtain its frequency domain representation. The calculation formula is:
[0041] in, For frequency domain size, For spatial coordinates, For frequency domain coordinates; S203, construct a Wiener filter based on the frequency domain PSF, balancing noise reduction and detail preservation, the calculation formula is:
[0042] in, for The conjugate of complex numbers, for The square of the modulus, is the regularization parameter for the k-th partition.
[0043] S204. Perform parallel frequency domain convolution: convert the padded input image... Perform frequency domain transformation to obtain the frequency domain image. ; and combine it with the Wiener filters of each partition. Element-wise multiplication, followed by inverse Fourier transform, yields the deconvolution result for each partition, as shown in the formula:
[0044] Where F represents the Fourier transform, Indicates the inverse Fourier transform. S205, Generate Spatial Adaptive Fusion Weights: To smoothly fuse the K deconvolution results, K control vertices are uniformly sampled in the image plane. For any pixel position in the image Calculate its distance to each control vertex The Euclidean distance is used to calculate the initial weights based on the reciprocal of the distance. The calculation formula is as follows:
[0045] in, To minimize the value (set to 1e-6), avoid a denominator of zero. Normalize the initial weights to obtain the final weight matrix, calculated using the following formula:
[0046] To ensure a smooth transition of the partitioning results, the weights at each pixel position are summed to 1. Each control vertex weight matrix... The visualization is shown in Figure 4; S206, weighted fusion yields the final low-frequency image: the deconvolution results of each partition are... With the corresponding normalized weights The low-frequency content is obtained by multiplying each pixel and summing the results. The calculation formula is:
[0047] in, This indicates pixel-by-pixel multiplication. The basic structure and outline of the image are preserved, as shown in Figure 5, which lays the foundation for subsequent high-frequency enhancement.
[0048] S207, Network Training and Parameter Optimization: The learnable parameters in this step include the spread function of the learnable points in each partition. and regularization parameters An end-to-end training approach was used to recover low-frequency content. The difference from the true reference image is used as the loss function, and the optimization objective is to minimize this loss. The formula for the loss function is as follows:
[0049] in, This is for the low-frequency content output of the spatially adaptive deconvolution network. Labels for low-frequency components of real-world scene images. For mean square error loss, This describes the perceptual loss mechanism based on the VGG network. During training, the AdamW optimizer is used with a momentum coefficient of 0.9, a second-order moment coefficient of 0.999, an initial learning rate of 4e-10, and 100 iterations. Cosine annealing is used as the learning rate scheduling strategy, with an initial period of 1, a period doubling factor of 2, and a decay step size of 2. Gradient pruning is employed concurrently during training to prevent gradient explosion and ensure stable parameter convergence. After training convergence, this step can be directly used for low-frequency content extraction during the testing phase without additional parameter adjustments.
[0050] S3. Recover high-frequency details lost during lensless imaging by generating a prior-driven adversarial high-frequency enhancement network. This step aims to utilize the natural image generation prior inherent in a large-scale pre-trained diffusion model to recover high-frequency textures and details lost due to physical degradation during lensless imaging through a lightweight adversarial fine-tuning strategy. The specific process is as follows: S301. Encode low-frequency content into the latent space of the diffusion model: Select the variational autoencoder of the pre-trained diffusion model SD2.1 to encode the low-frequency content output in step S2. Encode into the latent space to obtain the initial latent representation. The calculation formula is:
[0051] in, This is a VAE encoder used to map images from pixel space to a low-dimensional latent space, preserving low-frequency structural information while reducing computational complexity. It's important to note that the weights of the VAE decoder are frozen throughout the VAE encoder training process. S302, Constructing an adversarial enhancement network and discriminator based on diffusion prior: The pre-trained diffusion model SD2.1's Unet is used as the initial weights for the enhancement network, and low-rank adaptation modules are embedded in its cross-attention, residual block, and space-channel transformation modules to efficiently adapt to high-frequency enhancement tasks for lensless images. For any original weight matrix... Its effective update form is:
[0052] in, , It is a trainable low-rank matrix with a rank of 256. A discriminator based on the ConvNeXt architecture is also introduced. This forms a conditional adversarial training framework with the augmentation network. During the training of the augmentation network, , All original UNet weights are frozen, except for the LoRA parameters and the discriminator. Learnable; S303, Latent Representation Augmentation: Augmenting the latent representation By directly inputting the finely tuned augmented network, the latent representation is predicted and reconstructed through a single forward propagation. The goal of augmenting the network is to improve the final output. In the discriminator It is indistinguishable from real high-definition images in the eye, thus injecting high-frequency details that conform to the statistical laws of natural images.
[0053] S304. Decoding and reconstruction to obtain high-frequency enhanced images: The enhanced latent representation... Input VAE decoder The reconstructed image, with added high-frequency details, is a complete enhanced image. The calculation formula is as follows:
[0054] in, Retained The structure is consistent, and the texture and high-frequency edge details lost in lensless imaging are supplemented, as shown in Figure 6.
[0055] S305. Network Training and Parameter Optimization: The training process is divided into two consecutive stages: the first stage is encoder fine-tuning, and the second stage is adversarial high-frequency augmentation network training. The overall approach is a phased end-to-end training method, with the optimization goal of maximizing the consistency between the final augmented image and the original high-resolution image at the pixel, perceptual, and distribution levels.
[0056] S3051, Encoder Fine-tuning: This stage only applies to VAE encoders. Fine-tuning is performed to enhance its potential representation of lensless degraded images. The optimization objective is to make the response of low-frequency images consistent with that of real high-definition images in the discriminator. The loss function is defined as:
[0057] in This is for the low-frequency content output of the spatially adaptive deconvolution network. Real-world scene images It is a fine-tuning encoder. It uses a fixed decoder. Training employs the AdamW optimizer with a momentum coefficient of 0.9, a second-order moment coefficient of 0.999, and a learning rate of 5e-5. The batch size is set to 32, and the number of training steps is 5000. Cosine annealing is used for learning rate scheduling, with a warm-up step of 500, an initial period of 1, a period multiplication factor of 2, and a decay step size of 1000 to ensure stable convergence.
[0058] S3052, Fine-tuning of the adversarial high-frequency enhancement network: With the parameters of the first-stage physical prior module and the encoder / decoder parameters fixed, the LoRA adaptation layer and discriminator of the enhancement network are trained. The natural image modeling capability generated from the prior is used to optimize the realism of high-frequency details. The loss function adopts a multi-objective combination form.
[0059] in, To generate the high-frequency enhanced image output by the prior module, Real-world scene images For mean square error loss, For the perceptual loss based on VGG networks, To address Wasserstein's adversarial loss, UNet enhances the network. A discriminator based on the ConvNeXt backbone; The values are 1, 0.2, and 0.1 respectively. Discriminator The training process alternates between the generator and the augmentation network; the augmentation network is updated once for every two updates to the discriminator. Training uses the AdamW optimizer with a momentum coefficient of 0.9 and a second moment coefficient of 0.999. The initial learning rate for both the generator and discriminator is 1e-4, the batch size is 16, and the training steps are 20,000. The learning rate remains constant for the first 10,000 steps, then linearly decays to 1e-6. Gradient clipping is used during training to prevent gradient explosion. After training, all module parameters are fixed in this step and can be applied to the high-frequency detail restoration task of the test image without additional parameter tuning.
[0060] S4. Output Results: Input the PhlatCam test set into the final model to obtain the corresponding reconstruction output results.
[0061] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A two-stage lensless image enhancement method based on physical priors and generative priors, characterized in that, A two-stage collaborative enhancement approach using physical and generative priors is employed to obtain high-fidelity, high-detail lensless enhanced images. The specific steps include: S1, unifying the size specifications of the lensless measurement image and the point spread function to ensure compatibility with subsequent deconvolution operations; S2, using spatially adaptive deconvolution driven by physical priors to obtain low-frequency content; S3, recovering high-frequency details lost during lensless imaging through an adversarial high-frequency enhancement network driven by generative priors; and S4, training the two-stage model in stages, followed by inference post-processing to output the final enhanced image.
2. The two-stage lensless image enhancement method based on physical prior and generative prior according to claim 1, characterized in that, S1 specifically includes the following: S101, acquiring the original measurement image captured by the lensless camera. and the point spread function data obtained from system calibration, including the original measurement image. It is a three-channel RGB image with a size of The point spread function size is And satisfy S102, using the original measurement image Using the center as a reference, an image with the same size as the point spread function is cropped. To ensure the preservation of the core content region of the image and achieve frequency domain alignment with the point spread function, the calculation formula is as follows: S103. The cropped image Mirror padding is performed to avoid periodic boundary effects during frequency domain convolution. The padding width is set to... 、 To ensure that the filled image I pad Adapted to the frequency domain of the point spread function, the calculation formula is as follows: S104. Standardize the point spread function data, normalizing its pixel values to the [0,1] interval while preserving the spatial distribution characteristics of the point spread function to ensure consistency with the filled image. The grayscale range adaptation lays the foundation for subsequent deconvolution operations.
3. The two-stage lensless image enhancement method based on physical prior and generative prior according to claim 2, characterized in that, S2 specifically includes the following: S201, filling the image... The space is uniformly divided into K partitions, each partition corresponding to a control vertex. , The standardized point spread function data is centered according to the partition position to obtain the point spread function block corresponding to each partition. Ensure the point spread function block Spatial correspondence with image partitions; S202, point spread function block for each partition Perform frequency domain transformation to obtain its frequency domain representation. The calculation formula is: in, For frequency domain size, For spatial coordinates, Frequency domain coordinates; S203, a Wiener filter is constructed based on the frequency domain point spread function to balance noise reduction and detail preservation. The calculation formula is as follows: in, for The conjugate of complex numbers, for The square of the modulus, S204 is the regularization parameter for the k-th partition; S204, for any pixel position in the image Calculate its distance to each control vertex The Euclidean distance is used to calculate the initial weights based on the reciprocal of the distance. The calculation formula is as follows: in, The minimum value is used to avoid the denominator being zero; S205, the initial weights are normalized to obtain the final weight matrix, calculated using the following formula: Ensure the weights of each pixel position sum to 1 to achieve a smooth transition in the partitioning results; S206, process the filled image. Perform a Fourier transform to obtain the frequency domain image. ; and combine it with the Wiener filters of each partition. Element-wise multiplication, followed by inverse Fourier transform, yields the deconvolution result for each partition, as shown in the formula: Where F represents the Fourier transform, Indicates the inverse Fourier transform. S207 represents element-wise multiplication; deconvolution results of each partition. With the corresponding normalized weights The low-frequency content is obtained by multiplying each pixel and summing the results. The calculation formula is: in, This indicates pixel-by-pixel multiplication. Preserving the basic structure and contours of the image lays the foundation for subsequent high-frequency enhancement.
4. The two-stage lensless image enhancement method based on physical prior and generative prior according to claim 3, characterized in that, S3 specifically includes the following: S301, using a variational autoencoder with a pre-trained diffusion model to output low-frequency content. Encode into the latent space to obtain the initial latent representation. The calculation formula is: in, The VAE encoder is used to map images from pixel space to a low-dimensional latent space, preserving low-frequency structural information while reducing computational complexity. During VAE encoder training, the weights of the VAE decoder are frozen throughout. S302: The pre-trained diffusion model's Unet is used as the initial weights for the enhancement network, and a low-rank adaptation module is embedded in its key layers to efficiently adapt to high-frequency enhancement tasks for lensless images. For any original weight matrix... Its effective update form is: in, 、 For a trainable low-rank matrix, the rank Simultaneously, a discriminator based on the ConvNeXt architecture is introduced. This forms a conditional adversarial training framework with the augmentation network; during the training of the augmentation network, 、 All original UNet weights are frozen, except for the LoRA parameters and the discriminator. Learnable; S303, latent representation By directly inputting the finely tuned augmented network, the latent representation is predicted and reconstructed through a single forward propagation. The goal of augmenting the network is to improve the final output. In the discriminator The image is indistinguishable from a real high-definition image in the eye, thus injecting high-frequency details that conform to the statistical laws of natural images; S304, the enhanced latent representation Input VAE decoder The reconstructed image, with added high-frequency details, is a complete enhanced image. The calculation formula is as follows: in, Retained It maintains structural consistency while supplementing the texture and high-frequency edge details lost in lensless imaging.
5. The two-stage lensless image enhancement method based on physical prior and generative prior according to claim 4, characterized in that, S4 specifically includes the following: S401, fixing all parameters of the prior generation module, and training the space-adaptive deconvolution network's partition point spread function. and regularization parameters The optimization objective is to ensure the consistency of low-frequency content reconstruction. The loss function is a combination of MSE loss and LPIPS perceptual loss, expressed as follows: in, This is for the low-frequency content output of the spatially adaptive deconvolution network. Labels for low-frequency components of real-world scene images. For mean square error loss, For the perceptual loss based on the VGG network; S402, after the physical prior module is trained, its parameters and the augmentation network UNet and discriminator are fixed. decoder The weights are adjusted only by fine-tuning the VAE encoder. To improve its potential representation quality for lensless degraded images; the optimization objective is to make the response of low-frequency images consistent with that of real high-definition images in the discriminator, and the loss function is defined as: in This is for the low-frequency content output of the spatially adaptive deconvolution network. Real-world scene images It is a fine-tuning encoder. It is a fixed decoder; S403, fixed first-stage physical prior module parameters and encoder-decoder parameters, train the LoRA adaptation layer and discriminator of the augmentation network, utilize the natural image modeling capability of the generated prior to optimize the realism of high-frequency details, and the loss function adopts a multi-objective combination form: in, To generate the high-frequency enhanced image output by the prior module, Real-world scene images For mean square error loss, For the perceptual loss based on VGG networks, To address Wasserstein's adversarial loss, UNet enhances the network. A discriminator based on the ConvNeXt backbone; These are the weighting coefficients for each loss term.
6. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement the two-stage lensless image enhancement method based on physical prior and generative prior as described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement the two-stage lensless image enhancement method based on physical priors and generative priors as described in any one of claims 1-5.