An image reconstruction method for physical degradation perception of planar diffractive lens array
By introducing a weight modulation mechanism guided by physical priors, the non-stationary degradation problem of planar diffractive lens arrays is solved, achieving high-quality image reconstruction and improving reconstruction accuracy and detail preservation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF SCI & TECH
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-23
AI Technical Summary
Existing deep learning fusion algorithms lack the ability to utilize the physical distribution of optical degradation when dealing with the non-stationary degradation of planar diffractive lens arrays. This results in decreased contrast, blurred edge details, and artifacts in the reconstructed image, leading to unsatisfactory reconstruction results.
By introducing a physical prior-guided weight modulation mechanism, adaptive dynamic weighting of features at different spatial locations and sub-apertures is achieved through edge enhancement feature alignment, pseudo-array feature fusion, and adaptive group upsampling, thereby improving reconstruction accuracy.
It effectively solves the problem of spatial non-stationary blurring, accurately fits the quantitative relationship between feature response and point spread function, significantly improves image reconstruction quality, suppresses diffraction blurring and preserves high-frequency details, and provides a reconstruction approach with greater physical consistency.
Smart Images

Figure CN122265433A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computational imaging and computer vision, specifically to an image reconstruction method for physical degradation perception of planar diffractive lens arrays. Background Technology
[0002] In recent years, planar diffractive lens array imaging technology, based on the principles of diffractive optics, has gradually become a research hotspot. The core advantage of this technology lies in achieving ultra-thin and lightweight imaging systems, enabling efficient image acquisition solely through micro / nano structures modulating the light field. These characteristics give planar diffractive lens arrays broad application potential in various fields such as consumer electronics, micro-drones, smart security, and portable medical imaging. Because multi-aperture arrays can compensate for the shortcomings of single-aperture arrays in signal-to-noise ratio and spatial resolution through feature fusion, they can significantly improve the detail representation of images while maintaining an extremely low system profile.
[0003] While array imaging improves system integration, it also introduces complex and non-uniform image degradation patterns. Specifically, planar diffractive lenses exhibit severe diffraction aberrations and spatially non-stationary blurring during imaging. This physical degradation manifests as a significant change in the point spread function with the field of view, resulting in distinctly different degradation characteristics across different sub-apertures and even within the same sub-aperture. In multi-frame image joint reconstruction, most existing deep learning fusion algorithms rely on purely data-driven attention mechanisms to assign sub-aperture weights, lacking direct utilization of the physical distribution patterns of optical degradation. This "black box" weight allocation mechanism easily introduces poor-quality features severely affected by diffraction into the fusion process, manifesting in the spatial domain as decreased contrast, blurred edge details, and artifacts in the reconstructed image. Therefore, current joint reconstruction algorithms are not ideal in handling the unique non-stationary degradation of planar diffractive lens arrays, and still lack the ability to perceive physical degradation and perform precise weight compensation. Summary of the Invention
[0004] This invention provides an image reconstruction method for physical degradation perception of planar diffractive lens arrays. The method aims to improve the reconstruction accuracy of monochrome imaging systems by introducing a physically prior-guided weight modulation mechanism to achieve adaptive dynamic weighting for different spatial locations and sub-aperture features.
[0005] The technical solution for implementing this invention is: an image reconstruction method for physical degradation perception of planar diffraction lens arrays, comprising the following steps:
[0006] Step 1: Obtain a monochromatic degraded image sequence with multiple apertures using a planar diffraction lens array, and perform preprocessing to obtain the network input tensor;
[0007] Step 2: Input the network input tensor obtained in Step 1 into the pre-trained physical degradation-aware image reconstruction network for feature alignment, dimensionality reduction fusion and adaptive upsampling processing, and finally output the reconstructed high-definition grayscale image.
[0008] The physical degradation-aware image reconstruction network includes, in sequence, an edge enhancement feature alignment module, a pseudo-array feature fusion module, an adaptive grouping upsampling module, and a channel recombination unit;
[0009] The specific reconstruction process is as follows:
[0010] Step 2.1: Use deformable convolution in the edge enhancement feature alignment module to perform implicit geometric alignment of the spatial offset of each sub-aperture image caused by parallax and diffraction aberration, and extract preliminary alignment features;
[0011] Step 2.2: Using the pseudo-array feature fusion module, the aligned multi-aperture features obtained in Step 2.1 are stitched together in the channel dimension, and dimensionality reduction and remapping are performed using convolution to construct pseudo-array features with spatial topological associations, and deep features are extracted.
[0012] Step 2.3: In the adaptive grouping upsampling module, the pseudo-array features obtained in Step 2.2 are processed using a weight modulation mechanism guided by physical priors, and weights are assigned according to the degradation degree of different sub-apertures at different spatial locations; then transposed convolutional upsampling is performed.
[0013] Step 2.4: The channel reconstruction unit performs channel reconstruction on the upsampled features obtained in step 2.3, and finally outputs the reconstructed high-definition grayscale image.
[0014] Compared with the prior art, the present invention has the following significant advantages: (1) The present invention models the optical degradation distribution by using a degradation sensing unit, thereby achieving adaptive feature fusion for planar diffraction lens arrays and field edge regions with different physical parameters, effectively solving the problem of spatial non-stationary blurring; (2) The present invention utilizes physical prior to guide weight modulation, accurately fitting the quantitative relationship between feature response and point spread function, suppressing diffraction blur while fully preserving high-frequency details, significantly improving the reconstruction quality of the image; (3) The present invention automatically shields the interference of inferior physical features through a weight allocation mechanism, greatly reducing the model convergence difficulty under complex aberrations, and providing a more physically consistent reconstruction approach for lightweight array imaging systems.
[0015] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the image reconstruction algorithm of the present invention.
[0017] Figure 2 This is a schematic diagram of the edge enhancement feature alignment module structure.
[0018] Figure 3 This is a schematic diagram of the adaptive grouping upsampling and weight modulation strategy.
[0019] Figure 4 This is a diagram illustrating the internal structure and mechanism of the degenerate sensing unit.
[0020] Figure 5 This is a diagram of the experimental setup.
[0021] Figure 6 This is a weighting diagram for a 4×4 array.
[0022] Figure 7 This is a diagram showing the result of image reconstruction. Detailed Implementation
[0023] like Figure 1 As shown, the present invention is conceived as follows: an image reconstruction method for physical degradation perception of planar diffraction lens arrays. This method utilizes a deep learning network integrating an optical degradation perception mechanism to optimize sub-aperture feature fusion and weight allocation strategies, aiming to improve the imaging contrast and detail recovery accuracy of planar diffraction lens arrays. The method first acquires a sequence of monochromatic sub-aperture images of a planar diffraction lens array carrying spatially non-stationary blur and diffraction degradation features, and preprocesses them as input to the network. Subsequently, the feature responses of each sub-aperture are implicitly geometrically aligned through an edge enhancement feature alignment module, and a pseudo-array feature representation is constructed. During adaptive grouping upsampling, a physically prior-guided weight modulation strategy is introduced. The distribution law of the spatial heterogeneous point diffusion function is learned through the degradation perception unit, and the feature contribution is spatially adaptively modulated. Finally, a high-quality grayscale image is output through channel reconstruction. During the network training phase, an end-to-end optimization is performed using a composite loss function for grayscale image detail recovery to ensure that the model can adaptively perceive and compensate for optical degradation, effectively suppress diffraction blur, and preserve high-frequency details of the image. The specific steps are as follows:
[0024] Step 1: Obtain a sequence of monochromatic degraded images with multiple apertures using a planar diffraction lens array, and preprocess it to obtain the network input tensor, specifically:
[0025] Step 1.1: Perform parallel observations of the target scene using a planar diffraction lens array to obtain data containing... A sequence of monochrome degraded images from different sub-aperture viewpoints;
[0026] Step 1.2: Crop and normalize the acquired original degraded image, and stack all sub-aperture images along the channel dimension to construct a dimension of Input tensor ,in Number of sub-apertures and These represent the height and width of the image, respectively.
[0027] Step 2: Input the network input tensor obtained in Step 1 into the physical degradation-aware image reconstruction network for feature alignment, dimensionality reduction fusion and adaptive upsampling processing, and finally output the reconstructed high-definition grayscale image.
[0028] The physically degraded image reconstruction network comprises, in sequence, an edge enhancement feature alignment module, a pseudo-array feature fusion module, an adaptive grouping upsampling module, and a channel reconstruction unit; the specific reconstruction process is as follows:
[0029] Step 2.1: As Figure 2 As shown, the input tensor is fed into the edge enhancement feature alignment module, which uses deformable convolution to perform implicit geometric alignment on the spatial offset of each sub-aperture image caused by parallax and diffraction aberration, and extracts preliminary alignment features, specifically:
[0030] Step 2.1.1: Convert the input tensor obtained in Step 1.2 into a single input tensor. The input tensor is fed into the feature processing submodule, where the convolutional network structure within the feature processing submodule is used to process the input tensor. Multi-level feature extraction is performed to obtain an initial feature map. Subsequently, residual global context attention blocks are used to perform spatial and channel-dimensional feature weighting on the initial feature map to enhance its ability to express structural information, ultimately obtaining the extracted multi-aperture features. ;
[0031] Step 2.1.2: Apply modulated deformable convolution to the extracted majority aperture features obtained in Step 2.1.1. The responses are implicitly aligned to mitigate spatial response inconsistencies caused by differences in sub-aperture optical degradation, resulting in aligned multi-aperture features. The feature alignment process is achieved through modulated deformable convolution, mathematically expressed as:
[0032]
[0033] in, For the aligned majority aperture features, The extracted sub-aperture features obtained in step 2.1 are the input features of the current layer. This is the current pixel position. Preset the offset for the convolution kernel. For convolution kernel weights, and These are the learnable bias (offset) and modulation amplitude factor obtained through network learning for the spatial non-stationary offset of a planar diffractive lens array. This represents the total number of sampling points for the convolution kernel.
[0034] Step 2.2: Using the pseudo-array feature fusion module, the aligned multi-aperture features obtained in Step 2 are concatenated along the channel dimension, and dimensionality reduction and remapping are performed using convolution to construct pseudo-array features with spatial topological correlation. Deep features are then extracted, specifically:
[0035] Step 2.2.1: The aligned sub-aperture features output from Step 2.1 are concatenated along the channel dimension to form a unified high-dimensional feature tensor;
[0036] Step 2.2.2: Perform dimensionality reduction and remapping on the concatenated high-dimensional feature tensor through 1×1 convolution operation to construct a pseudo-array feature sequence, so that each pseudo-array channel integrates complementary information from all sub-apertures;
[0037] Step 2.2.3: Input the pseudo-array feature sequence into the U-Net structure with shared weights for multi-scale feature extraction, enhance the ability to model structural details and contextual information, and output the fused deep features.
[0038] Step 2.3: As Figure 3 As shown, in the adaptive grouping upsampling module, the features are first processed using a weight modulation mechanism guided by physical priors, and weights are assigned according to the degree of degradation of different sub-apertures at different spatial locations; then transposed convolutional upsampling is performed.
[0039] Prior to this, the core computational steps of the physics-prior-guided weight modulation mechanism are as follows:
[0040] Step 2.3.1: Divide the fused deep features from Step 2.2.3 into several feature groups;
[0041] Step 2.3.2: The original attention branch generates an initial dense weight map based on data-driven processing. That is, it uses the grouping features obtained in step 2.3.1 as input to the attention generation network to predict the dense spatial attention map, and obtains the basic weight distribution after Softmax normalization. ;
[0042] Step 2.3.3: The degradation sensing unit branch implicitly characterizes the degree of degradation at different spatial locations using the high-frequency distribution information in the grouped features obtained in Step 2.3.1, and generates a spatial modulation map related to the diffusion function of the spatial heterogeneous points of the planar diffraction lens. ;
[0043] Step 2.3.4: Apply the basic weight distribution obtained in Step 2.3.2 The spatial modulation map obtained in step 2.3.3 Perform element-wise multiplication to obtain the modulated effective weights. The effective weight The detailed mathematical formula is as follows:
[0044]
[0045] The final modulation weights are determined based on the effective weights:
[0046]
[0047] in, Represents the spatial coordinates of the image. Indicates the sub-aperture channel index. This represents the Hadamard product (element-by-element multiplication). This indicates a normalization operation. This represents the final modulation weights used for subsequent feature fusion. This is a temperature adjustment factor used to control the smoothness of the weight distribution, thereby suppressing the severely degraded edge sub-aperture response during the fusion process;
[0048] Step 2.3.5: After weighted fusion of the final modulation weights obtained in Step 2.3.4 and the grouped features obtained in Step 2.3.1, the result is fed into a transposed convolutional layer to achieve adaptive upsampling. The calculation formula is as follows:
[0049]
[0050] in, This represents a weighted operation in the spatial dimension. This indicates the transpose convolution operation, which is used to specifically compensate for diffraction blur in a planar diffraction lens and improve spatial resolution.
[0051] Step 2.4: The upsampled features obtained in Step 2.3 are reconstructed into channels to finally output the reconstructed high-resolution grayscale image. During the network training phase, to address the I / O bottleneck caused by the multi-channel high-dimensional data of the FDL array, discrete sub-aperture images are stacked along the channel dimension and serialized into multi-dimensional tensors before training to significantly improve data throughput efficiency. Network parameter updates employ the AdamW optimizer combined with a cosine annealing learning rate scheduling strategy to ensure the numerical stability of feature learning and smoothly decay the learning rate, avoiding parameter oscillations in the later stages of training. The optimization objective uses a composite loss function instead of mean squared error, leveraging its robustness to outliers to induce a sparse error distribution, prompting the network to focus on the detail recovery of high-frequency edges and diffraction fringes, effectively suppressing excessive image smoothing. Furthermore, to adapt to limited GPU memory resources, the scheme decouples the training and validation processes. The main loop only performs parameter updates and periodically saves the model for offline evaluation, thus ensuring high GPU utilization and system stability during long-term training.
[0052] Preferably, the composite loss function used in the training phase of the method is expressed as:
[0053]
[0054] in, This is a pixel-level reconstruction loss used to constrain the consistency between the reconstructed image and the real image in the pixel space. Specifically, it is expressed as follows:
[0055]
[0056] in The total number of pixels in the image. Let i be the pixel value of the reconstructed image at the i-th pixel position. Let be the pixel value of the actual label image at the i-th pixel position.
[0057] This is a perceptual loss used to guide the network in recovering structural and texture information that conforms to visual perception. Specifically, it is expressed as follows:
[0058]
[0059] in For the j-th layer of the pre-trained classification network, The activation feature map output by the j-th layer. Let be the number of channels in the j-th layer feature map. Let be the height of the feature map of layer j. is the width of the feature map of the j-th layer.
[0060] Gradient loss is used to strengthen the reconstruction constraints of high-frequency edge information that disappears in planar diffraction lens imaging. Specifically, it is expressed as:
[0061]
[0062] in For the horizontal gradient operator, This is the gradient operator in the vertical direction. , , These are the weighting coefficients corresponding to each type of loss.
[0063] like Figure 4As shown, the degradation sensing unit of this invention extracts optical degradation-related information from input features through convolutional layers. Its core utilizes a deep learning network to learn and estimate the distribution law of spatial heterogeneous point spread functions caused by physical characteristics, and dynamically generates adaptive spatial and channel weights accordingly. Subsequently, the weights and extracted features are element-wise modulated to achieve targeted degradation compensation and feature suppression guided by physical priors. Finally, the degradation sensing features carrying physical degradation compensation information are output through multi-scale feature fusion, providing key physical feature support for the accurate reconstruction of subsequent high-quality grayscale images.
[0064] like Figure 5 The diagram shows the experimental setup, with the device on the left being a multi-aperture imaging system and the right side showing multiple sub-aperture images captured by the imaging system.
[0065] like Figure 6 The diagram shows the weight distribution of the 16 sub-apertures.
[0066] like Figure 7 The image shown is a reconstruction result. The first row shows the sub-aperture images captured by the multi-aperture imaging system, and the second row shows the corresponding reconstruction images. The window railings are reconstructed completely, the text is reconstructed clearly, the butterfly wing patterns and animal fur are reconstructed relatively clearly, and the facial features of the human face are reconstructed well.
Claims
1. An image reconstruction method for physical degradation perception of planar diffractive lens arrays, characterized in that, Includes the following steps: Step 1: Obtain a monochromatic degraded image sequence with multiple apertures using a planar diffraction lens array, and perform preprocessing to obtain the network input tensor; Step 2: Input the network input tensor obtained in Step 1 into the pre-trained physical degradation-aware image reconstruction network for feature alignment, dimensionality reduction fusion and adaptive upsampling processing, and finally output the reconstructed high-definition grayscale image. The physical degradation-aware image reconstruction network includes, in sequence, an edge enhancement feature alignment module, a pseudo-array feature fusion module, an adaptive grouping upsampling module, and a channel recombination unit; The specific reconstruction process is as follows: Step 2.1: Use deformable convolution in the edge enhancement feature alignment module to perform implicit geometric alignment of the spatial offset of each sub-aperture image caused by parallax and diffraction aberration, and extract preliminary alignment features; Step 2.2: Using the pseudo-array feature fusion module, the aligned multi-aperture features obtained in Step 2.1 are stitched together in the channel dimension, and dimensionality reduction and remapping are performed using convolution to construct pseudo-array features with spatial topological associations, and deep features are extracted. Step 2.3: In the adaptive grouping upsampling module, the pseudo-array features obtained in Step 2.2 are processed using a weight modulation mechanism guided by physical priors, and weights are assigned according to the degradation degree of different sub-apertures at different spatial locations; then transposed convolutional upsampling is performed. Step 2.4: The channel reconstruction unit performs channel reconstruction on the upsampled features obtained in step 2.3, and finally outputs the reconstructed high-definition grayscale image.
2. The image reconstruction method for physical degradation perception of planar diffractive lens arrays according to claim 1, characterized in that, The specific process of obtaining a monochromatic degraded image sequence with multiple apertures using a planar diffraction lens array is as follows: Step 1.1: Perform parallel observations of the target scene using a planar diffraction lens array to obtain data containing... A sequence of monochrome degraded images from different sub-aperture viewpoints; Step 1.2: Crop and normalize the acquired original degraded image, and stack all sub-aperture images along the channel dimension to construct a dimension of Input tensor ,in Number of sub-apertures and These represent the height and width of the image, respectively.
3. The image reconstruction method for physical degradation perception of planar diffractive lens arrays according to claim 1, characterized in that, The process of implicitly aligning the spatial offsets of each sub-aperture image caused by parallax and diffraction aberrations using deformable convolution in the edge enhancement feature alignment module, and extracting preliminary alignment features, is as follows: Step 2.1.1: The input tensor first enters the feature processing submodule of the edge enhancement feature alignment module, and the convolutional network structure within the feature processing submodule is used to process the input tensor. Multi-level feature extraction is performed to obtain an initial feature map; Subsequently, the initial feature map is subjected to spatial and channel-dimensional feature weighting using residual global context attention blocks to obtain the extracted multi-aperture features. ; Step 2.1.2: Apply modulated deformable convolution to the extracted majority aperture features obtained in Step 2.1.
1. The response is implicitly aligned to obtain the aligned majority aperture features.
4. The image reconstruction method for physical degradation perception of planar diffractive lens arrays according to claim 3, characterized in that, The aligned sub-aperture features are as follows: in, For the aligned output features, For input features, This is the current pixel position. Preset the offset for the convolution kernel. For convolution kernel weights, and These are the learnable bias and modulation amplitude factors obtained through network learning for the spatial non-stationary offset of a planar diffractive lens array. This represents the total number of sampling points for the convolution kernel.
5. The image reconstruction method for physical degradation perception of planar diffractive lens arrays according to claim 1, characterized in that, The specific process of constructing pseudo-array features with spatial topological correlations and extracting deep features through the pseudo-array feature fusion module is as follows: Step 2.2.1: The aligned sub-aperture features output from Step 2.1 are concatenated along the channel dimension to form a unified high-dimensional feature tensor; Step 2.2.2: Perform dimensionality reduction and remapping on the concatenated high-dimensional feature tensor through 1×1 convolution operation to construct a pseudo-array feature sequence, so that each pseudo-array channel integrates complementary information from all sub-apertures; Step 2.2.3: Input the pseudo-array feature sequence into the U-Net structure with shared weights for multi-scale feature extraction, enhance the ability to model structural details and contextual information, and output the fused deep features.
6. The image reconstruction method for physical degradation perception of planar diffractive lens arrays according to claim 1, characterized in that, The specific process of using a weight modulation mechanism guided by physical priors to process features, assigning weights according to the degradation degree of different sub-apertures at different spatial locations, and performing transposed convolutional upsampling is as follows: Step 2.3.1: Divide the fused deep features from Step 2.2.3 into several feature groups; Step 2.3.2: The original attention branch generates an initial dense weight map based on data-driven processing. That is, it uses the grouping features obtained in step 2.3.1 as input to the attention generation network to predict the dense spatial attention map, and obtains the basic weight distribution after Softmax normalization. ; Step 2.3.3: The degradation sensing unit branch implicitly characterizes the degree of degradation at different spatial locations using the high-frequency distribution information in the grouped features obtained in Step 2.3.1, and generates a spatial modulation map related to the diffusion function of the spatial heterogeneous points of the planar diffraction lens. ; Step 2.3.4: Apply the basic weight distribution obtained in Step 2.3.2 The spatial modulation map obtained in step 2.3.3 Perform element-wise multiplication to obtain the modulated effective weights. Effective weights The detailed mathematical formula is as follows: The final modulation weights are determined based on the effective weights: in, Represents the spatial coordinates of the image. Indicates the sub-aperture channel index. It represents the Hadamah accumulation. This indicates a normalization operation. This represents the final modulation weights used for subsequent feature fusion. This is a temperature adjustment factor used to control the smoothness of the weight distribution, thereby suppressing the severely degraded edge sub-aperture response during the fusion process; Step 2.3.5: After weighted fusion of the final modulation weights obtained in step 2.3.4 and the grouped features obtained in step 2.3.1, the weights are fed into the transposed convolutional layer to achieve adaptive upsampling.
7. The image reconstruction method for physical degradation perception of planar diffractive lens arrays according to claim 1, characterized in that, During the training phase of the physically degraded image reconstruction network, a composite loss function for grayscale image detail recovery is used for end-to-end optimization. The composite loss function used during training is expressed as follows: in, This is a pixel-level reconstruction loss used to constrain the consistency between the reconstructed image and the real image in the pixel space. This is a perceptual loss used to guide the network to recover structural and texture information that conforms to visual perception; Gradient loss is used to strengthen the reconstruction constraints on high-frequency edge information that disappears in planar diffraction lens imaging; , , These are the weighting coefficients corresponding to each type of loss.
8. The image reconstruction method for physical degradation perception of a planar diffractive lens array according to claim 7, characterized in that, Pixel-level reconstruction loss Specifically, it is expressed as follows: in, The total number of pixels in the image. Let i be the pixel value of the reconstructed image at the i-th pixel position. Let be the pixel value of the actual label image at the i-th pixel position.
9. The image reconstruction method for physical degradation perception of a planar diffractive lens array according to claim 7, characterized in that, Perceived loss Specifically, it is expressed as follows: in For the j-th layer of the pre-trained classification network, The activation feature map output by the j-th layer. Let be the number of channels in the j-th layer feature map. Let be the height of the feature map of layer j. is the width of the feature map of the j-th layer.
10. The image reconstruction method for physical degradation perception of a planar diffractive lens array according to claim 7, characterized in that, Gradient loss Specifically, it is expressed as follows: in For the horizontal gradient operator, For the vertical gradient operator, , , These are the weighting coefficients corresponding to each type of loss.