Lensless imaging reconstruction method based on advanced visual task guidance

By constructing a reconstruction-task cascade architecture and dynamic weighted joint loss, combined with the CBAM attention module and feature adaptation layer, the problem of synergy between lensless imaging reconstruction quality and advanced vision tasks is solved, achieving simultaneous optimization of high-quality reconstruction and high-precision tasks, and adapting to the hardware requirements of miniaturized devices.

CN121937504APending Publication Date: 2026-04-28TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN UNIV
Filing Date
2026-01-15
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In the process of lensless imaging reconstruction, the reconstruction quality is insufficient and the coordination with advanced vision tasks is poor, resulting in blurred reconstructed images, loss of texture, low accuracy in advanced vision tasks, and difficulty in supporting high-precision applications.

Method used

A reconstruction-task cascade architecture is constructed, which combines an improved reconstruction network with dynamic weight joint loss based on the CBAM attention module. Performance is improved through knowledge distillation and edge loss, and a feature adaptation layer is inserted to ensure cascade compatibility, thereby achieving simultaneous optimization of reconstruction quality and task accuracy.

Benefits of technology

It significantly improves the PSNR/SSIM, classification accuracy, and segmentation mIoU of reconstructed images, adapts to the needs of resource-constrained hardware scenarios, and ensures stable operation of the model in lensless imaging hardware systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937504A_ABST
    Figure CN121937504A_ABST
Patent Text Reader

Abstract

The invention discloses a lens-free imaging reconstruction method based on advanced visual task guidance, and belongs to the technical field of lens-free imaging and digital image signal processing. According to the method, a cascade depth architecture is designed according to the requirement of an advanced vision task for image quality, a lens-free input signal is preliminarily reconstructed, then a preliminary reconstruction result is input to an advanced vision task network, and by means of a joint loss function fused with advanced vision task performance feedback, reconstruction performance optimization is strengthened through task guidance. According to the method, multiple deep network structures are adopted, and the lens-free imaging reconstruction quality is efficiently improved under the guidance of advanced visual tasks such as image classification and semantic segmentation. Experiments show that the reconstruction quality of lens-free imaging is remarkably improved while the accuracy of advanced vision tasks is guaranteed, task-oriented multi-task collaborative robust optimization is achieved, and key technical support is provided for advanced vision application of lens-free imaging.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of digital image signal processing technology, and more specifically, relates to a lensless imaging reconstruction method based on advanced vision task orientation. Background Technology

[0002] With deep learning technology continuously empowering various fields, lensless imaging technology in the field of image signal processing has achieved rapid development due to its unique technical characteristics, gradually becoming one of the important solutions for scenarios such as portable imaging and embedded vision. The rise of this technology has not only driven innovation in areas such as miniaturized imaging devices, privacy-preserving monitoring, and on-site medical diagnosis, but also demonstrated irreplaceable value in resource-constrained application scenarios (such as wearable devices and vision modules for miniature robots).

[0003] Lensless imaging has gained widespread attention due to its series of superior characteristics: First, the system is extremely simple to assemble, requiring no complex lens group structure; imaging can be completed simply through direct coupling between the sensor and the target, significantly reducing hardware costs and size. Second, it is highly flexible, adapting to different imaging distances and scene types by adjusting algorithm parameters without replacing hardware components. Third, it offers outstanding privacy protection; the raw measurement data is a non-intuitive frequency and spatial domain mixed signal, requiring algorithmic reconstruction to obtain a visualized image, thus inherently possessing data privacy protection capabilities. Fourth, it is highly adaptable to miniaturized devices; its minimalist hardware structure allows it to be embedded in thin and light devices such as mobile phone camera modules and miniature medical detectors, aligning with the current trend of device miniaturization.

[0004] However, despite these significant advantages, lensless imaging still faces several core challenges in its practical application. Firstly, there is a performance bottleneck in the lensless imaging reconstruction process: the raw measurement data from lensless imaging is susceptible to photon noise, circuit interference, and other factors, resulting in a low signal-to-noise ratio; traditional reconstruction algorithms (such as classical Wiener filtering), while computationally efficient, are insufficient in recovering detailed information, often leading to blurry images and texture loss; even deep learning-based single reconstruction networks often prioritize visual quality as the sole optimization objective, lacking adaptation to subsequent task requirements, making it difficult for reconstruction results to support high-precision advanced visual tasks.

[0005] Secondly, there is insufficient synergy between lensless imaging and advanced vision tasks: In most current solutions, lensless reconstruction and advanced vision tasks such as image classification and semantic segmentation are independent processes—first, a visualization image is obtained through a reconstruction network, and then it is input into a separately trained task model. This "separate" process has obvious drawbacks: on the one hand, the reconstruction process does not take into account the semantic requirements of the task, and may lose feature information that is crucial for classification and segmentation; on the other hand, the independently trained task model is not adapted to the feature distribution of the lensless reconstructed image, which can easily lead to a significant drop in task accuracy. Even the few attempts at cascading solutions often fail to achieve ideal reconstruction quality and task performance simultaneously due to the lack of a reasonably designed joint loss function and a lack of feature alignment mechanisms.

[0006] In summary, while lensless imaging offers advantages such as miniaturization, low cost, and high privacy, the synergistic optimization of its reconstruction quality and performance for advanced vision tasks, as well as its adaptability to resource-constrained devices, remain key issues that urgently need to be addressed. How to integrate high-quality reconstruction with high-precision advanced vision tasks while ensuring the hardware advantages of lensless imaging has become the core breakthrough for promoting the large-scale deployment of this technology.

[0007] In view of this, the present invention proposes a lensless imaging reconstruction method based on advanced vision task orientation. Summary of the Invention

[0008] The purpose of this invention is to propose a lensless imaging reconstruction method based on advanced vision task orientation to address the problems mentioned in the background art. This invention constructs a "reconstruction-task" cascade architecture, combining an improved reconstruction network with a CBAM attention module and a dynamic weighted joint loss to simultaneously optimize reconstruction quality and task accuracy. This invention performs lightweight modifications to the advanced vision task model and introduces knowledge distillation and edge loss to improve performance. A feature adaptation layer is inserted to ensure cascade compatibility, ultimately achieving hardware adaptation and deployment, balancing the hardware advantages of lensless imaging with the efficiency and adaptability of the algorithm.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: A lensless imaging reconstruction method based on advanced vision task orientation specifically includes the following steps: S1. Acquisition and preprocessing of raw data for lensless imaging: Create a dataset that can be used simultaneously for lensless imaging reconstruction and advanced visual tasks such as image classification and semantic segmentation. S2. Reconstruction Network Pre-training: Pre-training can train the Wiener filter and the lensless reconstruction network to obtain a preliminary lensless imaging reconstruction network. S3. Pre-training of advanced visual task models: Pre-training advanced visual task models such as image classification networks and semantic segmentation networks to obtain preliminary advanced visual task models. S4. End-to-end training of cascaded networks: The reconstruction network is cascaded with the high-level vision task network, and the Wiener filter parameters and the high-level vision task model parameters are frozen for end-to-end training. During training, the reconstruction loss (such as perceptual loss, MSE loss, adversarial loss, etc.) and the high-level vision task loss (such as classification loss, segmentation loss) are calculated, and their weighted summation is used to obtain the joint loss, which guides the training optimization of the lensless reconstruction network. S5. Real-world deployment and testing of the model: Deploy the adapted model to a lensless imaging hardware system and test it in real-world scenarios such as microscopic imaging and object recognition: Collect new lensless measurement data, output reconstructed images, classification results, and segmentation masks through the model, verify the reconstruction quality and task accuracy, and meet the needs of practical applications.

[0010] Preferably, S1 specifically includes the following: A lensless imaging hardware system was built and key parameters were configured, including sensor resolution, object distance, gain, and exposure time. N×P-dimensional raw measurement data from multiple scenes were collected. Preprocessing operations were performed on the raw data: an adaptive threshold method was used to filter out abnormal pixels, and the data was normalized to the [0,1] interval to finally obtain clean lensless measurement input data.

[0011] Preferably, S2 specifically includes the following: S201. Construct a trainable Wiener filter module (parameter matrix W with dimension N×P), using lensless measurement data as input and high-resolution target image as supervision, define a joint loss function of "frequency domain MSE loss + spatial domain gradient loss", and train the filter to learn the frequency domain reconstruction mapping. S202. Embed CBAM attention modules after the output of each layer of the encoder (first weighting the feature channels with channel attention, then enhancing the detail regions with spatial attention); introduce residual connections in the decoder to alleviate the gradient vanishing problem. During the training phase, the intermediate image output by the Wiener filter is used as input and the high-resolution target image is used as the supervision signal. The network is trained using a joint loss function of "MSE pixel loss + perceptual loss + edge gradient loss", and a preliminary reconstructed network is obtained after iteration.

[0012] Preferably, S3 specifically includes the following: S301. Pre-training of the lightweight classification model: ResNet50 is selected as the base classification model, and its backbone network is replaced with MobileNetV2 to achieve lightweighting; the classification labels of the labeled dataset are used as supervision, and the model is trained by "cross-entropy loss + knowledge distillation loss (distilling the classification probability distribution of ResNet101)" and the pre-training of the classification model is completed after iteration.

[0013] S302. Model improvement and pre-training based on the DeepLab-LargeFOV framework: its backbone network is replaced with the lightweight ShuffleNetV2, and an edge detection branch is added before the segmentation head (with the real edges extracted by the Canny operator as supervision); the semantic segmentation mask is used as the core supervision, and the model is trained using the joint loss function of "edge consistency loss + mIoU loss", and the segmentation model is pre-trained after iteration.

[0014] Preferably, S4 specifically includes the following: S401. Analyze the distribution differences between the output features of the reconstruction network and the input features of the high-level vision task model, and insert a feature adaptation layer between the reconstruction network and the task model: unify the mean and variance of the features through adaptive layer normalization (AdaLN), and use a 1×1 convolutional layer to adjust the number of feature channels to the input dimension of the task model to ensure the compatibility of feature transfer after cascading.

[0015] S402. Construction of the dynamic weighted joint loss function: Define the joint loss for multiple tasks:

[0016] in, Indicates the weighting coefficient; Indicates the losses incurred during reconstruction; Indicates task loss; The reconstruction loss L Recon The expression is:

[0017] in, Indicates the weighting coefficient; L MSE This represents the mean squared error loss; L Percep Indicates perceived loss; L adv This represents the adversarial loss; it is output by the discriminator Disc. The task loss The expression is:

[0018] in, Indicates the weighting coefficient; L Ce Represents cross-entropy loss; L mIoU This represents the mean intersection and union ratio loss; S403, End-to-End Training and Multi-Metric Monitoring of Cascaded Networks: The Wiener filter, attention Unet, feature adaptation layer, and classification / segmentation model are cascaded into a complete network. The parameters of the Wiener filter and the backbone network of the task model are frozen, and only the parameters of Unet, feature adaptation layer, and task model output layer are opened. The AdamW optimizer (with the learning rate set) is used for training. A certain number of training epochs are set, and core metrics such as reconstruction PSNR / SSIM, classification accuracy, and segmentation mIoU are verified at fixed intervals. When the metrics do not show significant improvement for several consecutive epochs, training is terminated early.

[0019] Compared with existing technologies, this invention provides a lensless imaging reconstruction method based on advanced vision task guidance, which has the following beneficial effects: This invention proposes a lensless imaging reconstruction method based on advanced vision task orientation. By constructing a "reconstruction-task" cascade architecture and a dynamic weight joint loss function, it achieves synergistic optimization of lensless imaging reconstruction quality and advanced vision task accuracy. It effectively solves the problems of feature loss and poor task adaptability in the traditional "separate" process, and significantly improves the PSNR / SSIM, classification accuracy, and segmentation mIoU of the reconstructed image.

[0020] Secondly, strategies such as embedding the CBAM attention module into the reconstruction network and adopting lightweight backbone replacement and knowledge distillation for the task model enhance performance while controlling the number of model parameters and computational complexity, which meets the hardware requirements of miniaturized and resource-constrained lensless imaging.

[0021] Finally, the feature compatibility of the cascaded network is ensured by the feature adaptation layer, and the hardware adaptation deployment optimization ensures that the model can run stably on lensless imaging hardware systems. It demonstrates high reliability and wide applicability in practical scenarios such as microscopic imaging and object recognition, and provides an efficient solution for advanced vision applications of lensless imaging. Attached Figure Description

[0022] Figure 1 This is an overall flowchart of a lensless imaging reconstruction method based on advanced vision task orientation proposed in this invention; Figure 2 This refers to the assembly scheme of the lensless imaging system mentioned in Embodiment 1 of the present invention. Figure 3This is the training method flow of a lensless imaging reconstruction method based on advanced vision task orientation mentioned in Embodiment 1 of the present invention. Detailed Implementation

[0023] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0024] Example 1: Please see Figure 1 This invention proposes a lensless imaging reconstruction method based on advanced vision task guidance, which specifically includes the following steps: S1. Lensless Imaging System Setup. The lensless imaging system consists of four core components: a fixed support, a light-shielding box, an optical mask, and an image sensor. The fixed support provides mechanical support for the entire system, ensuring the stable relative positions of the components. The light-shielding box isolates stray light interference from the scene, preventing ambient light from contaminating the measurement data. The optical pupil is tightly connected to the optical mask, which is made of sandblasted colloidal material. Its placement is 400 mm from the imaging object (in this embodiment, a display is used as the imaging target) and only 1.9 mm from the image sensor. The image sensor uses the same model as the CM3-U3-13Y3C camera, with a resolution of 1280×1024 (i.e., N=1280, P=1024), a frame rate of 148FPS, a pixel size of 4.8μm, and contains 1.3 million pixel units. Figure 3 The diagram shown is a schematic of a lensless imaging system.

[0025] S2. Lens-free raw dataset acquisition. 50 object classes (including daily necessities, animals, plants, vehicles, etc., 200 images per class) from a public classification dataset (such as a subset of ImageNet) were selected, with each image size of 224×224 and a uniform black background. 2000 labeled images (including indoor scenes, street scenes, etc., with labeled categories including people, vehicles, buildings, vegetation, etc., totaling 20 classes) from a public semantic segmentation dataset (such as a subset of PASCAL VOC) were selected, with each image size of 512×512.

[0026] The control screen displays the above source images in sequence, and the corresponding raw measurement data is collected through the lensless system: the sensor gain is set to 10dB, the exposure time is 50ms, and all data dimensions are 1280×1024.

[0027] Please see Figure 3 The training method flow of the lensless imaging reconstruction method based on advanced vision task-oriented proposed in this invention is as follows: Figure 3 As shown, it specifically includes: S3. Pre-training of the trainable Wiener filter: Construct a 1280×1024 dimension parameter matrix W, using preprocessed lensless measurement data as input and the corresponding high-resolution target image as supervision. Train for 10 epochs using the Adam optimizer (learning rate 0.001), verifying PSNR every 10 epochs. Stop training when the loss converges, obtaining the initial frequency domain reconstruction mapping.

[0028] Training an improved Unet network: The Unet network structure is designed (4 layers for the encoder, 4 layers for the decoder, each layer with a 3×3 convolutional kernel and a stride of 2). A CBAM attention module (using a 2-layer MLP for channel attention and a 3×3 convolution for spatial attention) is embedded after the output of each encoder layer. Residual connections are added between each layer of the decoder and the corresponding layer of the encoder. Using the intermediate image output from the Wiener filter as input and the high-resolution target image as supervision, a joint loss function is defined:

[0029] The Adam optimizer (learning rate 0.0005) was used to train for 50 epochs. Eventually, the loss converged and no longer decreased, resulting in the preliminary reconstructed network.

[0030] S4, Advanced Vision Task Model Pre-training (Lightweight Classification + Segmentation Model) Lightweight classification model pre-training: Based on ResNet50, its backbone network is replaced with MobileNetV2 (depth multiplier 1.0), reducing the number of parameters to 40% of the original model. The classification labels of the dataset are used as supervision, and the loss function is defined as cross-entropy loss. The SGD optimizer (initial learning rate 0.01, cosine annealing decay) is used for 100 epochs of training, stopping when the classification accuracy on the validation set is ≥95%.

[0031] Pre-training of the edge-enhanced segmentation model: Based on the DeepLab-LargeFOV framework, the backbone network is replaced with ShuffleNetV2, and an edge detection branch is added before the segmentation head (using the real edges extracted by the Canny operator as supervision). The loss functions are defined as Intersection over Union (IoU) loss and edge consistency loss. The model is trained for 180 epochs using the Adam optimizer (learning rate 0.001), and pre-training is completed when the mIoU on the validation set is ≥85%.

[0032] S5, End-to-end training of cascaded networks (feature adaptation + dynamic joint loss) Feature adaptation layer design and network cascading: The feature distribution differences between the reconstruction network output (3-channel image) and the input of the high-level vision task model (256 channels for classification, 384 channels for segmentation) were analyzed, and a feature adaptation layer was inserted: First, an AdaLN layer (mean μ=0.5, variance σ=0.2) was used to unify the feature distribution, and then 1×1 convolutional layers were used to adjust the number of channels to 256 (classification branch) and 384 (segmentation branch), respectively. The "Wiener filter → improved Unet → feature adaptation layer → classification / segmentation model" was cascaded into a complete network. All parameters of the Wiener filter and the backbone network parameters of the classification / segmentation model were frozen, and only the parameters of Unet, the feature adaptation layer, and the task model output layer were exposed.

[0033] Construction of dynamic joint loss: Define total loss :

[0034] in, Indicates the weighting coefficient; Indicates the losses incurred during reconstruction; Indicates task loss; The reconstruction loss L Recon The expression is:

[0035] in, Indicates the weighting coefficient; L MSE This represents the mean squared error loss; L Percep Indicates perceived loss; L adv This represents the adversarial loss, which is output by the discriminator Disc. The task loss The expression is:

[0036] in, Indicates the weighting coefficient; L Ce Represents cross-entropy loss; L mIoU This represents the mean crossover ratio loss.

[0037] End-to-end training and monitoring: The AdamW optimizer (learning rate 0.0001, weight decay 1e-5) was used for 50 training rounds. In each round, the four core metrics of "reconstruction PSNR / SSIM, classification accuracy, and segmentation mIoU" were verified. When the metrics did not show significant improvement for 10 consecutive rounds (improvement < 0.1%), training was terminated early (convergence was achieved after 30 training rounds).

[0038] S6. Macroscopic Object Test Scenario Design: Select two typical macroscopic object scenarios, both involving everyday, non-microscopic objects (size 5-50 cm): Indoor desktop scenarios include plastic water cups, hard-sided laptops, wireless keyboards, and metal spoons; outdoor portable scenarios include folding umbrellas, bicycle helmets, and plastic storage boxes, totaling 9 types of objects.

[0039] Data Acquisition and Model Inference: The aforementioned macroscopic objects were placed sequentially at their original screen positions (400mm from the optical mask), and the lensless system was controlled to acquire raw measurement data (sensor gain 12 dB, exposure time 35 ms, 5 acquisitions per object to reduce random noise); the data was input into the deployed model, and three results were output: Reconstructed image: A 224×224 resolution visualization image generated by the model; Classification results: Label the object category and confidence level (e.g., "water cup, confidence level 0.96"); Segmentation mask: distinguishes the object region from the background (the background is the desktop / ground, represented by a grayscale value of 0, and the object region is represented by a grayscale value of 255); Reconstruction quality: The average PSNR and SSIM of reconstructed images for both indoor and outdoor scenes are superior to traditional task-free reconstruction methods (PSNR improvement of 0.5-1.0 dB).

[0040] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A lensless imaging reconstruction method based on advanced vision task guidance, characterized in that, Specifically, the following steps are included: S1. Acquisition and preprocessing of raw data for lensless imaging: Create a dataset that can be used for both lensless imaging reconstruction and advanced visual tasks such as image classification and semantic segmentation. S2. Reconstruction Network Pre-training: Pre-train the Wiener filter and the lensless reconstruction network to obtain a preliminary lensless imaging reconstruction network; S3. Pre-training of advanced visual task models: Pre-training image classification networks and semantic segmentation networks to obtain preliminary advanced visual task models; S4. End-to-end training of cascaded networks: The reconstruction network is cascaded with the high-level vision task network, and some parameters of the Wiener filter and the parameters of the high-level vision task model are frozen for end-to-end training. During the training process, the reconstruction loss and the high-level vision task loss are calculated and weighted to obtain the joint loss, which guides the training optimization of the lensless reconstruction network. S5. Real-world deployment and testing of the model: Deploy the adapted model to the lensless imaging hardware system and test it in a real-world scenario: Collect new lensless measurement data, output reconstructed images, classification results, and segmentation masks through the model, verify the reconstruction quality and task accuracy, and meet the needs of practical applications.

2. The lensless imaging reconstruction method based on advanced vision task guidance according to claim 1, characterized in that, S1 specifically includes the following: Build a lensless imaging hardware system and configure key parameters, including sensor resolution, object distance, gain and exposure time, and collect N×P dimension raw measurement data from multiple scenes; Preprocessing operations are performed on the raw data: an abnormal pixel is filtered out using an adaptive threshold method, and the data is normalized to the [0,1] interval to finally obtain clean lens-free measurement input data.

3. The lensless imaging reconstruction method based on advanced vision task guidance according to claim 1, characterized in that, S2 specifically includes the following: S201. Construct a Wiener filter module, using lensless measurement data as input and high-resolution target image as supervision, define a joint loss function of "frequency domain MSE loss + spatial domain gradient loss", and train the filter to learn the frequency domain reconstruction mapping; S202. Embed CBAM attention modules after the output of each layer of the encoder, and introduce residual connections in the decoder to alleviate the gradient vanishing problem. During the training phase, the intermediate image output by the Wiener filter is used as input and the high-resolution target image is used as the supervision signal. The network is trained using a joint loss function of "MSE pixel loss + perceptual loss + edge gradient loss". After iteration, a preliminary reconstruction network is obtained.

4. The lensless imaging reconstruction method based on advanced vision task guidance according to claim 1, characterized in that, S3 specifically includes the following: S301. Pre-training of the lightweight classification model: ResNet50 was selected as the base classification model, and its backbone network was replaced with MobileNetV2 to achieve lightweighting; the classification labels of the labeled dataset were used as supervision, and the model was trained using "cross-entropy loss + knowledge distillation loss"; the pre-training of the classification model was completed after iteration. S302. Model improvement and pre-training based on the DeepLab-LargeFOV framework: its backbone network is replaced with the lightweight ShuffleNetV2, and an edge detection branch is added before the segmentation head; the semantic segmentation mask is used as the core supervision, and the model is trained using the joint loss function of "edge consistency loss + mIoU loss", and the segmentation model is pre-trained after iteration.

5. The lensless imaging reconstruction method based on advanced vision task guidance according to claim 1, characterized in that, S4 specifically includes the following: S401. Analyze the distribution differences between the output features of the reconstruction network and the input features of the high-level visual task model, and insert a feature adaptation layer between the reconstruction network and the task model: The mean and variance of the unified features are normalized by an adaptive layer, and the number of feature channels is adjusted to the input dimension of the task model by a 1×1 convolutional layer to ensure the compatibility of feature transfer after cascading. S402. Construction of the dynamic weighted joint loss function: Define the joint loss for multiple tasks: in, Indicates the weighting coefficient; Indicates the losses incurred during reconstruction; Indicates task loss; The reconstruction loss L Recon The expression is: in, Indicates the weighting coefficient; L MSE This represents the mean squared error loss; L Percep Indicates perceived loss; L adv This represents the adversarial loss, which is output by the discriminator Disc. The task loss The expression is: in, Indicates the weighting coefficient; L Ce Represents cross-entropy loss; L mIoU This represents the mean intersection and union ratio loss; S403, end-to-end training and multi-metric monitoring of cascaded networks: The Wiener filter, attention Unet, feature adaptation layer, and classification / segmentation model are cascaded into a complete network. The parameters of the Wiener filter and the backbone network of the task model are frozen, and only the parameters of Unet, feature adaptation layer, and task model output layer are opened. The AdamW optimizer is used for training, and a certain number of training epochs are set. The core indicators of reconstruction PSNR / SSIM, classification accuracy, and segmentation mIoU are verified at fixed intervals. When the indicators do not show significant improvement for several consecutive epochs, the training is terminated early.

6. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, code set, or instruction set, and the instruction, program, code set, or instruction set is loaded and executed by the processor to implement a lensless imaging reconstruction method based on any one of claims 1-5.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, at least one program, code set, or instruction set, which is loaded and executed by a processor to implement a lensless imaging reconstruction method based on advanced vision task orientation as described in any one of claims 1-5.