Lensless imaging reconstruction method and device

By combining the Le-ADMM model, loss compensation network, and dual-feature cross-transformer network, the problems of low image quality and artifacts in lensless imaging reconstruction methods are solved, achieving high-quality image reconstruction and improving the applicability of lensless imaging technology.

CN119478092BActive Publication Date: 2025-12-12ZHOUKOU NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411551837.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-01
Publication Date
2025-12-12
Estimated Expiration
2044-11-01

AI Technical Summary

Technical Problem

Existing lensless imaging reconstruction methods suffer from low image perception quality and artifacts, which limit the development of lensless imaging technology and prevent the reproduction of high-quality images.

Method used

By combining the Le-ADMM model, loss compensation network, and dual-feature cross-transformer network, an image pair dataset is constructed by acquiring optically encoded images and lens images in pairs, and high-quality real-world scene images are reconstructed through model training.

Benefits of technology

The reconstructed image quality is improved, containing rich details and textures that conform to human visual perception, avoiding artifacts, providing interpretability, and enhancing global feature reasoning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478092B_ABST
    Figure CN119478092B_ABST
Patent Text Reader

Abstract

The application discloses a lensless imaging reconstruction method and device, belongs to the technical field of minimal optical system imaging, and is used for solving the technical problems that the existing optical coding image reconstruction method has low image perception quality and artifacts, limits the development of the lensless imaging technology, and cannot reproduce the high-quality images as the traditional lens camera. The method comprises the following steps: collecting optical coding images and lens images in pairs, and constructing an image pair dataset; constructing a Le-ADMM model and a corresponding loss compensation network; constructing a double-feature cross network, and taking the output feature maps of the Le-ADMM model and the loss compensation network as inputs of the double-feature cross network to obtain a lensless imaging reconstruction model; training the lensless imaging reconstruction model through the image pair dataset, inputting a to-be-reconstructed optical coding image into the trained lensless imaging reconstruction model, and obtaining a reconstructed optical coding image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of minimalist optical system imaging technology, and in particular to a lensless imaging reconstruction method and device. BACKGROUND

[0002] Emerging needs such as wearable / integrated sensors, Internet of Things, and augmented / virtual reality have greatly driven imaging devices towards miniaturization, multifunctionality, and high cost performance. At the same time, traditional imaging optical technology has the disadvantages of low design freedom, large volume, high cost, and limited multi-dimensional data acquisition capability due to the use of lens imaging method. Therefore, people are actively exploring new imaging modes to overcome the limitations of traditional imaging optical technology.

[0003] Lensless imaging technology is a new and hot field in the field of computational imaging. This technology encodes a real scene through a mask, records the optical encoding image after the mask through a sensor, and then directly reconstructs the scene from the optical encoding image using a lensless imaging reconstruction method. Lensless imaging technology eliminates the dependence on heavy and complex optical lenses, further improving the miniaturization and integration of the imaging system, and is very suitable for integration into portable or wearable devices.

[0004] Existing lensless imaging reconstruction methods can be divided into model-based methods and data-driven methods. The model-based method describes the lensless imaging system as a model-based inverse problem, and uses an iterative optimization algorithm to minimize the loss function, thereby reconstructing the real scene. The optimization function of the model consists of a data fidelity term and a regularization term for satisfying the image prior knowledge. This method utilizes the prior information in the lensless imaging system, which is called the Point Spread Function (PSF). However, due to the use of physical models that are often too idealized and single, it is difficult to obtain an ideal PSF in practice due to system errors and calibration errors. In addition, the sparsity prior used in the iterative optimization method is not applicable to all scenes. This type of method cannot provide convergence guarantees within a few iterations, and the reconstructed image has low perceptual quality due to model mismatch and calibration errors. The data-driven method does not incorporate any prior information in the lensless reconstruction process. Compared with the traditional model-based method, the data-driven method lacks interpretability, has no structured way to insert knowledge of the image system, and cannot mathematically prove the convergence of the results. Therefore, although the data-driven method based on deep learning can reduce the reasoning time while improving image quality, it often has various artifacts due to the lack of guidance from physical models.

[0005] It can be seen that the existing optical coding image reconstruction method has the problems of low image perception quality and artifacts, which limits the development of lensless imaging technology and cannot reproduce high-quality images like traditional lens cameras. It is urgent to develop and improve the existing lensless imaging reconstruction method to promote the development of lensless imaging technology. SUMMARY

[0006] The embodiment of the present application provides a lensless imaging reconstruction method and device, which solves the technical problem that the existing optical coding image reconstruction method has the problems of low image perception quality and artifacts, which limits the development of lensless imaging technology and cannot reproduce high-quality images like traditional lens cameras.

[0007] The embodiment of the present application adopts the following technical scheme:

[0008] On the one hand, the embodiment of the present application provides a lensless imaging reconstruction method, which comprises: acquiring optical coding images and lens images in pairs, and constructing an image pair dataset;

[0009] constructing a Le-ADMM model and a corresponding loss compensation network; the input of the loss compensation network is the original optical coding image and the reconstructed feature map output by the Le-ADMM model;

[0010] constructing a double-feature cross transformer network, and taking the output feature maps of the Le-ADMM model and the loss compensation network as inputs of the double-feature cross transformer network to obtain a lensless imaging reconstruction model;

[0011] training the lensless imaging reconstruction model through the image pair dataset, and inputting a to-be-reconstructed optical coding image into the trained lensless imaging reconstruction model to obtain a reconstructed optical coding image.

[0012] In a feasible implementation, the optical coding images and the lens images are acquired in pairs to construct the image pair dataset, which specifically comprises:

[0013] The optical coding images and the lens images are acquired in pairs for the same field of view target in the same scene, and an image pair is formed; wherein the optical coding image is acquired by a lensless camera, and the lens image is acquired by an RGB camera;

[0014] The scene or the field of view target is changed, different image pairs are acquired, and the acquired image pairs are summarized as an image pair dataset Wherein, P I is an optical coding image, P O is a corresponding lens image; k∈(1,2,……,K), indicating the serial number of the image pair, and K is the total number of image pairs in the image pair dataset.

[0015] In an implementable embodiment, the Le-ADMM model and its corresponding loss compensation network are constructed, specifically including:

[0016] The Le-ADMM model is constructed; wherein the Le-ADMM model comprises an image reconstruction module and a feature extraction module;

[0017] The image reconstruction module is calculated for N times of iteration to obtain reconstructed images I1, I2, …, In; N and input into the feature extraction module for feature extraction, and output the reconstructed feature map

[0018] Based on the reconstructed feature map, the loss compensation network of the Le-ADMM model is constructed to obtain the compensation feature map through the loss compensation network.

[0019] In an implementable embodiment, based on the reconstructed feature map, the loss compensation network of the Le-ADMM model is constructed to obtain the compensation feature map, specifically including:

[0020] According to obtain the compensation feature map wherein P I is the original optical encoding image;

[0021] Based on the model formula of the loss compensation network is constructed; wherein F 1 , F 2 , …, F N are auxiliary variables, is the compensation feature map; Conv represents the convolution feature extraction operation; Dense represents the dense connection between network layers;

[0022] Concate represents the aggregation of two features along the channel direction;

[0023] The compensation feature map is substituted into the model formula, and all the compensation feature maps are iteratively calculated.

[0024] In an implementable embodiment, a double-feature cross transformer network is constructed, and the output feature maps of the Le-ADMM model and the loss compensation network are both taken as inputs of the double-feature cross transformer network to obtain a lens-free imaging reconstruction model, specifically including:

[0025] Construct a first convolutional layer and a second convolutional layer; the first convolutional layer is used to extract the first shallow features of the reconstructed feature map, and the second convolutional layer is used to extract the second shallow features of the compensation feature map output by the loss compensation network;

[0026] A feature fusion module based on the dual-feature cross-transformer network DFCT is constructed; wherein, the feature fusion module is composed of several identical DFCT blocks;

[0027] The outputs of the first and second convolutional layers are used as inputs to the feature fusion module to obtain the lensless imaging reconstruction model.

[0028] In one feasible implementation, constructing a first convolutional layer and a second convolutional layer specifically includes:

[0029] Construct the first convolutional layer H with a kernel size of 3×3. FE1 (·), and according to Extracting and reconstructing feature maps First shallow features in, W and H represent the length and width of the feature, respectively, and C represents the number of channels of the feature;

[0030] Construct a second convolutional layer H with a kernel size of 3×3. FE2 (·), and according to Extracting compensated feature maps The second shallow feature F y Among them, F y ∈R W×H×C .

[0031] In one feasible implementation, a feature fusion module based on the dual-feature cross-transformer network DFCT is constructed, specifically including:

[0032] Based on a dual-feature cross-transformer network, i identical DFCT blocks are constructed; each DFCT block includes two cross-attention blocks, namely a first cross-attention block and a second cross-attention block.

[0033] Let the input of the i-th DFCT block be the output of the (i-1)-th DFCT block. and the second shallow feature F y Where i is a positive integer greater than 0; the input of the first DFCT block is the first shallow feature.

[0034] The input features are fused by the first and second cross-attention blocks in the i-th DFCT block to obtain the output features of the i-th DFCT block.

[0035] The i DFCT blocks are sequentially connected, finally connected with two convolutional layers, and a reconstruction image calculation formula is constructed: The construction of the feature fusion module is completed.

[0036] In a feasible implementation manner, the input features are fused by the first cross-attention block and the second cross-attention block in the i-th DFCT block to obtain the output features of the i-th DFCT block Specifically, it comprises:

[0037] The F and F y are respectively expanded into non-overlapping blocks and Wherein, M=HW / L 2 , represents the total number of expanded blocks, and L represents the size of the expanded block;

[0038] Through the first cross-attention block, the F is linearly mapped into a key vector and a value vector The F is linearly mapped into a query vector Wherein, m=(1,2,……,M); W K , W V , W Q are mapping matrices; D is the dimension of the vector;

[0039] According to the F , the first self-attention result F CA1 is obtained; wherein, softmax(·) is a normalization operation;

[0040] According to the F , the output result F1 of the first cross-attention block is obtained;

[0041] Through the second cross-attention block, the F is linearly mapped into a key vector and a value vector The F is linearly mapped into a query vector And the second self-attention result F CA2 is calculated correspondingly; wherein,

[0042] According to F2=MLP(LN(F y +F CA2 ))+F y +F CA2 , the output result F2 of the i-th DFCT block is obtained., to obtain an output result F2 of the second cross-attention block; wherein, MLP(·) represents a multi-layer perceptron operation, and LN(·) represents a layer normalization operation.

[0043] According to to obtain an initial output result of the i-th DFCT block;

[0044] According to optimizing the initial output result and densely connecting the output results of the first i-1 DFCT blocks to obtain an output feature of the i-th DFCT block; wherein, DFCTB(·) represents a calculation operation of the DFCT block.

[0045] In a feasible implementation, the lens-free imaging reconstruction model is trained by the image pair dataset, and specifically includes:

[0046] The image pair dataset is divided into a training set and a test set according to a preset ratio;

[0047] The neural network parameters in the Le-ADMM model, the loss compensation network and the double-feature cross-transformer network are trained by the training set until the network converges, and the training is completed;

[0048] The optical encoding image in the test set is input into the trained lens-free imaging reconstruction model to obtain a reconstructed test image wherein, is the k-th optical encoding image;

[0049] According to the lens image in the test set and the reconstructed test image, a loss function is subjected to gradient descent operation, and the neural network parameters in the Le-ADMM model, the loss compensation network and the double-feature cross-transformer network are optimized according to the operation result; wherein, represents the square root of the L2 norm.

[0050] On the other hand, the embodiment of the present application also provides a lens-free imaging reconstruction device, the device includes:

[0051] at least one processor; and

[0052] a memory in communication connection with the at least one processor; wherein

[0053] The memory stores instructions executable by the at least one processor, so that the at least one processor can execute the lens-free imaging reconstruction method according to any one of claims 1-9.

[0054] Compared with the prior art, the lens-free imaging reconstruction method and device provided by the embodiment of the application has the following beneficial effects:

[0055] (1) The application can reconstruct a higher-quality real scene image from a known optical encoding image by combining a physical reconstruction model and a neural network model. The image contains rich detailed textures and is more natural, conforming to human visual perception.

[0056] (2) Compared with the traditional physical model-based method, the application avoids the problem of low perceived quality of the reconstructed image caused by model mismatch and calibration error. Compared with the pure data-driven method, the application combines the physical model and the neural network model, adds prior information in the lens-free reconstruction process, and makes the whole more interpretable, avoiding some artifacts that usually exist.

[0057] (3) In addition, by inputting the physical model and its corresponding loss features into the double-feature cross transformer network, the network considers the multiplexing characteristics in the lens-free imaging process, does not have local inductive bias, encodes through the self-attention mechanism, directly considers the global context, enhances the global feature reasoning ability, can understand the global information in the optical encoding image, and makes the reconstructed image consistent with the real scene. BRIEF DESCRIPTION OF DRAWINGS

[0058] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the application, and those skilled in the art can obtain other drawings according to these drawings without creative labor. In the drawings:

[0059] Figure 1 A lens-free imaging reconstruction method flowchart is provided for the embodiment of the application;

[0060] Figure 2 A lens-free imaging reconstruction model architecture diagram is provided for the embodiment of the application;

[0061] Figure 3 A double-feature cross transformer network architecture diagram is provided for the embodiment of the application;

[0062] Figure 4 An internal structure diagram of a double-feature cross transformer network block is provided for the embodiment of the application;

[0063] Figure 5This is a comparison chart of the visual quality of images reconstructed by different methods, provided in an embodiment of the present invention.

[0064] Figure 6 This is a schematic diagram of the structure of a lensless imaging reconstruction device provided in an embodiment of the present invention. Detailed Implementation

[0065] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.

[0066] This invention provides a lensless imaging reconstruction method, such as... Figure 1 As shown, the method specifically includes steps S101-S104:

[0067] S101. Acquire optically encoded images and lens images in pairs to construct an image pair dataset.

[0068] Specifically, optically encoded images and lens images are first acquired for the same target in the same scene and then combined into an image pair. The optically encoded image is acquired by a lensless camera, and the lens image is acquired by an RGB camera.

[0069] Then, change the scene or field of view target and acquire different image pairs until a sufficient number have been acquired. The acquired image pairs are then compiled into an image pair dataset. Among them, P I For optically encoded images, P O The corresponding lens image; k∈(1,2,……,K), represents the index of the image pair, and K is the total number of image pairs in the image pair dataset.

[0070] As a feasible implementation, in the same scene, for the same target in the same field of view, an optically coded image is acquired using a lensless camera, and a lensed image is acquired using an RGB camera. The real-world scene content and size represented by these two images are kept consistent. The two images are then bound together as an image pair and stored in an image pair dataset. After acquiring a sufficient number of image pairs, the image pair dataset is stored for use during subsequent model training.

[0071] S102. Construct the Le-ADMM model and its corresponding loss compensation network; the input to the loss compensation network is the original optically encoded image and the reconstructed feature map output by the Le-ADMM model.

[0072] Specifically, first, a Le-ADMM model is constructed; wherein the Le-ADMM model comprises an image reconstruction module and a feature extraction module. Then, the image reconstruction module is calculated for N times of iterations to obtain reconstructed images I 1 ,I 2 ,……,I N , and input into the feature extraction module for feature extraction, and output reconstructed feature maps

[0073] As a feasible implementation manner, the Le-ADMM model is a lens-free imaging reconstruction method based on a physical model, which obtains a reconstructed image by solving a minimization optimization problem of the following formula: Wherein A represents a measurement matrix constructed by a point spread function (PSF) of a lens-free optical system, C represents a clipping operation, which is mainly used to limit the size of the output, n and represent a real scene photographed and an optical encoded image acquired by a sensor respectively, represents a sparse transformation, and represents a weight of a sparse term, which is used to balance a data fidelity term and a sparse term in an optimization process.

[0074] Then, by using an alternating direction method of multipliers (ADMM), the minimization optimization problem in the formula is solved by separating variables, and a clear image is reconstructed from an optical encoded image. The Le-ADMM is to set a hyperparameter in the algorithm as a learnable parameter.

[0075] Figure 2 A lens-free imaging reconstruction model architecture provided by the embodiment of the present application is shown in Fig. 1. Figure 2 As shown in Fig. 1, intermediate images I1, I2, …, I n are reconstructed by the Le-ADMM model in the nth iteration, the total number of iterations is N, and the reconstructed image obtained by the last iteration is I N .

[0076] Then, the obtained intermediate images I1, I2, …, I N-1 and the reconstructed image I N obtained by the last iteration are subjected to feature extraction by separate convolutional neural networks g(.), g(.) … g(.), g(.) represents a convolutional layer with a three-layer convolution kernel size of 3x3, and the size of the image after convolution remains unchanged, as shown in Fig. 2. Figure 2 After feature extraction, the reconstructed feature maps obtained by I1, I2, …, I N are

[0077] Further, based on the reconstructed feature map, a loss compensation network of the Le-ADMM model is constructed to obtain a compensation feature map through the loss compensation network, specifically comprising:

[0078] According to the compensation feature map is obtained ; wherein P I is the original optical code image.

[0079] Then, a model formula of the loss compensation network is constructed based on . Among them, F 1 , F 2 , …, F N are auxiliary variables, is the compensation feature map; represents a convolution feature extraction operation; Dense represents a dense connection between network layers;

[0080] Concate represents aggregating two features along the channel direction;

[0081] Then, the compensation feature map is substituted into the above model formula, and all compensation feature maps are iteratively calculated.

[0082] As a feasible implementation manner, as shown in Figure 2 , the compensation feature map and the reconstructed feature map obtained by the first iteration of the Le-ADMM model are calculated through the above formula, and the auxiliary variable F 1 is output, and the auxiliary variable F 1 is obtained after convolution operation , and the compensation feature map is obtained. The compensation feature map and the reconstructed feature map obtained by the second iteration of the Le-ADMM model are calculated through the above formula, and the auxiliary variable F 2 is output, and the auxiliary variable F 2 is obtained after convolution operation , and the compensation feature map is obtained. In this way, the auxiliary variable F N-1 is obtained after convolution operation , and the compensation feature map is obtained.

[0083] Then, the reconstructed feature map obtained by the Nth iteration of the Le-ADMM model and the compensation feature map Input into the Dual Feature Cross Transformer (DFCT) network.

[0084] S103. Construct a dual-feature cross-transformer network, and use the output feature maps of both the Le-ADMM model and the loss compensation network as inputs to the dual-feature cross-transformer network to obtain a lensless imaging reconstruction model.

[0085] Specifically, a first convolutional layer and a second convolutional layer are constructed; the first convolutional layer is used to extract the first shallow features of the reconstructed feature map, and the second convolutional layer is used to extract the second shallow features of the compensation feature map output by the loss compensation network.

[0086] As a feasible implementation method, a first convolutional layer H with a kernel size of 3×3 is constructed. FE1 (·), and according to Extracting and reconstructing feature maps First shallow features in, W and H represent the length and width of the feature, respectively, and C represents the number of channels in the feature. A second convolutional layer H with a kernel size of 3×3 is constructed. FE2 (·), and according to Extracting compensated feature maps The second shallow feature F y Among them, F y ∈R W×H×C .

[0087] Furthermore, a feature fusion module based on a dual-feature cross-transformer network is constructed; the feature fusion module consists of several identical DFCT blocks (DFCTB). Then, the outputs of the first and second convolutional layers are used as inputs to the feature fusion module to obtain a lensless imaging reconstruction model.

[0088] The feature fusion module based on the dual-feature cross-transformer network DFCT specifically includes:

[0089] Based on the dual-feature cross-transformer network DFCT, i identical DFCTBs are constructed; each DFCTB includes two cross-attention blocks, namely the first cross-attention block and the second cross-attention block.

[0090] Let the input of the i-th DFCT block be the output of the (i-1)-th DFCT block. and the second shallow feature F y Where i is a positive integer greater than 0; the input of the first DFCT block is the first shallow feature.

[0091] fusing the input features by the first cross-attention block and the second cross-attention block in the ith DFCT block to obtain output features of the ith DFCT block

[0092] connecting the i DFCT blocks in turn, finally connecting with two convolutional layers, and constructing a reconstruction image calculation formula: The construction of the feature fusion module is completed.

[0093] As a feasible implementation manner, Figure 3 A double-feature cross-transformer network architecture schematic diagram provided by the embodiment of the application is shown in Figure 1. Figure 3 As shown in the figure, the reconstructed feature map obtained by the Le-ADMM model in the Nth iteration and the compensation feature map output by the loss compensation network are input into the DFCT network, and first, each of them is subjected to shallow feature extraction to obtain a first shallow feature and a second shallow feature F y , and then is input into the first block DFCTB-1 of the DFCT network for calculation, and the output is input into the next block DFCTB-2 after being densely connected with , and the second shallow feature F y is also input into the next block. In this way, the output of each block is densely connected with the outputs of all the previous blocks, and is input into the next block together with F y , until the output feature of the ith block DFCTB-i is obtained. Then, according to the reconstruction image calculation formula , two convolutional operations are performed to obtain the final reconstructed optical code image.

[0094] In a further implementation manner, Figure 4 An internal structure schematic diagram of a double-feature cross-transformer network block provided by the embodiment of the application is shown in Figure 2. Figure 4 The output feature of the ith DFCT block is obtained by fusing the input features by the first cross-attention block and the second cross-attention block in the ith DFCT block Specifically, it includes:

[0095] F and F y are respectively expanded into non-overlapping blocks and , wherein M = HW / L 2 , indicating the total number of expanded blocks, and L represents the size of the expanded block.

[0096] The first cross-attention block linearly maps to key vectors and value vectors linearly maps to query vectors where m=(1, 2, …, M); W K , W V , W Q are mapping matrices. d is the dimension of the vector.

[0097] Then, according to a first self-attention result F CA1 is obtained; wherein softmax(·) is a normalization operation.

[0098] Then, according to an output result F1 of the first cross-attention block is obtained.

[0099] The second cross-attention block is opposite to the first cross-attention block, which linearly maps to key vectors and value vectors linearly maps to query vectors and according to a second self-attention result F CA2 is calculated; wherein

[0100] According to F2=MLP(LN(F y +F CA2 ))+F y +F CA2 , an output result F2 of the second cross-attention block is obtained; wherein MLP(·) represents a multi-layer perceptron operation, and LN(·) represents a layer normalization operation.

[0101] Then, according to an initial output result of the i-th DFCT block is obtained.

[0102] Since the features extracted by the deep neural network at different depths have different information, the application adopts a dense connection method to enhance feature propagation and promote feature reuse, and also reduces gradient disappearance. The method is: according to optimizing the initial output result, densely connecting the output results of the first i-1 DFCT blocks, and obtaining the output features of the i-th DFCT block, which is also embodied in Figure 3 ; wherein DFCTB(·) represents the calculation operation of the DFCT block.

[0103] S104, training the lensless imaging reconstruction model through the image pair dataset, and inputting the optical coded image to be reconstructed into the trained lensless imaging reconstruction model to obtain a reconstructed optical coded image.

[0104] Specifically, the image pair dataset is divided into a training set and a test set according to a preset ratio. The preset ratio is preferably training set: test set = 8:2.

[0105] Further, the training set is input, and various hyperparameters during training are included, including a learning rate, a batch size, a number of channels of the network, and a size of a convolution kernel. The neural network parameters in the Le-ADMM model, the loss compensation network, and the double-feature cross transformer network are trained through the training set until the network converges, and the training is completed.

[0106] Further, the optical coded image in the test set is input into the trained lensless imaging reconstruction model to obtain a reconstructed test image wherein, is the kth optical coded image.

[0107] Then, according to the lens image in the test set and the reconstructed test image, the loss function is subjected to gradient descent operation, and the neural network parameters in the Le-ADMM model, the loss compensation network, and the double-feature cross transformer network are optimized according to the operation result; wherein, represents the square root of the L2 norm. The above process is cycled to update the neural network parameters in the model until the loss function converges, and the trained lensless imaging reconstruction model is obtained.

[0108] Finally, the optical coded image to be reconstructed is input into the trained lensless imaging reconstruction model to obtain a reconstructed optical coded image.

[0109] In order to prove the effectiveness of the method, an open dataset from DiffuserCam of the University of California, Berkeley is used for experimental verification, and the dataset contains a total of 25000 pairs of optical coded images-lens images of the same scene. The original image is down-sampled by 4 times to obtain an image with a size of 480x270, so as to avoid the decline in the quality of the lens image due to the cloud stripe. The dataset is split, of which 20000 pairs are used for training and 5000 pairs are used for testing.

[0110] In the implementation process, the number of double-feature cross transformer blocks i is set to 4, the channel number C is set to 64, the size of the convolution kernel in all convolution layers is 3*3, the network is trained by using an Adam optimizer, the number of cycles of the data set in the training is set to 200, the size of the batch is 8, the initial learning rate is set to 0.001, and the learning rate is attenuated to one-tenth of the original value every 50 training cycles.

[0111] In the experiment, the mean square error (MSE), the peak signal-to-noise ratio (PSNR), and the learned perceptual image patch similarity (LPIPS) are used as evaluation indexes of the reconstructed images, wherein the lens image obtained by the traditional lens camera is used as a reference image. As shown in Table 1, the comparison results of the present application and other methods in the above indexes can be found that the present application has a lower MSE and a higher PSNR, indicating that the present application can more accurately reconstruct the same information as the lens image. In addition, the results on the LPIPS are also obviously better than those of other methods. As shown in Figure 5 the visual quality comparison chart of the images reconstructed by different methods, the results show that the present application not only performs well in the quantitative indexes, but also has more outstanding visual quality of the reconstructed images, and the reconstructed images are more clear in vision and can more accurately reflect the real scene and reconstruct the color and detail information.

[0112] Table 1

[0113]

[0114]

[0115] In addition, the present application also provides a lens-free imaging reconstruction device, as shown in Figure 6 The lens-free imaging reconstruction device specifically comprises:

[0116] at least one processor; and a memory connected with the at least one processor in communication; wherein

[0117] The memory stores instructions executable by the at least one processor, so that the at least one processor can execute:

[0118] pairing acquisition of optical coding images and lens images, and construction of an image pair data set;

[0119] construct a Le-ADMM model and a corresponding loss compensation network; the input of the loss compensation network is an original optical encoding image and a reconstructed feature map output by the Le-ADMM model;

[0120] construct a double-feature cross transformer network, and input the output feature maps of the Le-ADMM model and the loss compensation network into the double-feature cross transformer network to obtain a lens-free imaging reconstruction model;

[0121] train the lens-free imaging reconstruction model through the image pair dataset, and input an optical encoding image to be reconstructed into the trained lens-free imaging reconstruction model to obtain a reconstructed optical encoding image.

[0122] The lens-free imaging reconstruction method using Le-ADMM reconstruction and loss compensation feature cross attention can combine the advantages of the physical model method and the data-driven method, significantly improve the quality of the reconstructed image, and verify the robustness and superiority of the algorithm through experimental results on multiple datasets. The algorithm effectively improves the reconstruction quality of the lens-free imaging system and increases the applicability of the lens-free imaging technology. The clear real scene image can be reconstructed from the optical encoding image.

[0123] Each of the embodiments in the present application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment mainly describes the difference from other embodiments. Especially, the device, equipment and non-volatile computer storage medium embodiments are basically similar to the method embodiments, so the description is relatively simple, and the related parts can be referred to the part of the method embodiment.

[0124] The above describes specific embodiments of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or can be advantageous.

[0125] The above only describes the embodiments of the present application and is not used to limit the present application. The embodiments of the present application can be variously changed and modified by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.

Claims

1. A lensless imaging reconstruction method, characterized in that, The method comprises: Pairing acquisition of optical coding images and lens images to construct an image pair dataset; Constructing a Le-ADMM model and a corresponding loss compensation network, wherein the input of the loss compensation network is an original optical coding image and a reconstructed feature map output by the Le-ADMM model, and the loss compensation network specifically comprises: constructing a Le-ADMM model; wherein the Le-ADMM model comprises an image reconstruction module and a feature extraction module; performing N times of iterative calculation on the image reconstruction module to obtain reconstructed images I1, I2, …, I N , and inputting the reconstructed images into the feature extraction module for feature extraction to output a reconstructed feature map Based on the reconstructed feature map, a loss compensation network of the Le-ADMM model is constructed to obtain a compensation feature map through the loss compensation network, specifically comprising: According to obtaining a compensated feature map wherein P I is the original optical code image; Based on A model formula for constructing the loss compensation network; wherein F 1 ,F 2 ,……,F N are auxiliary variables, is a compensation feature map; represents a convolution feature extraction operation; Dense represents a dense connection between network layers; Concate represents aggregation of two features along the channel direction; Compensation feature map Substitute into the model formula, iterative calculation of all compensation feature map; A double-feature cross transformer network is constructed, and the output feature maps of the Le-ADMM model and the loss compensation network are both taken as inputs of the double-feature cross transformer network to obtain a lens-free imaging reconstruction model, specifically comprising: A first convolutional layer and a second convolutional layer are constructed, specifically comprising: A first convolutional layer H with a kernel size of 3x3 is constructed FE1 (·), and according to extracting a reconstructed feature map a first shallow feature of wherein, W and H represent the length and width of the feature respectively, and C represents the number of channels of the feature A second convolutional layer H with a kernel size of 3x3 is constructed FE2 (·), and according to The compensation feature map is extracted The second shallow feature F of the second convolutional layer H y ; wherein F y ∈R W×H×C ; the first convolutional layer is used to extract the first shallow feature of the reconstruction feature map, and the second convolutional layer is used to extract the second shallow feature of the compensation feature map output by the loss compensation network. A feature fusion module based on a double-feature cross transformer network (DFCT) is constructed; wherein the feature fusion module is composed of a plurality of identical DFCT blocks; each DFCT block includes two cross-attention blocks; specifically comprising: Based on a double-feature cross transformer network, i identical DFCT blocks are constructed; wherein each DFCT block includes two cross-attention blocks, namely a first cross-attention block and a second cross-attention block; Let the input of the i-th DFCT block be the output of the (i-1)-th DFCT block and the second shallow feature F y ; wherein i is a positive integer greater than 0; the input of the 1st DFCT block is the first shallow feature The first cross-attention block and the second cross-attention block in the ith DFCT block fuse the input features to obtain output features of the ith DFCT block Specifically comprises: will be described below. and F y are respectively unfolded into non-overlapping blocks and where M = HW / L 2 , M represents the total number of unfolded blocks, and L represents the size of the unfolded blocks. Through the first cross attention block, Linear mapping to key vector Sum value vector Will Linear mapping to query vector Where m = (1, 2, ..., M); W K W V W Q All are mapping matrices; d is the dimension of the vector; According to obtaining a first self-attention result F CA1 ; wherein softmax(·) is a normalization operation; According to obtaining an output result F1 of the first cross-attention block; Through the second cross attention block, Linear mapping to key vector Sum value vector Will Linear mapping to query vector And the corresponding second self-attention result F is calculated. CA2 ;in, According to F2=MLP(LN(F y +F CA2 ))+F y +F CA2 , an output result F2 of the second cross-attention block is obtained; wherein MLP(·) represents a multi-layer perception operation, and LN(·) represents a layer normalization operation. According to obtaining an initial output result of the ith DFCT block; According to The initial output result is optimized, and the output results of the first i-1 DFCT blocks are densely connected to obtain the output feature of the i-th DFCT block; wherein DFCTB(·) represents the calculation operation of the DFCT block. The i DFCT blocks are connected in sequence, finally connected with two convolutional layers, and a reconstruction image calculation formula is constructed: The construction of the feature fusion module is completed. The outputs of the first convolutional layer and the second convolutional layer are taken as inputs of the feature fusion module to obtain the lens-free imaging reconstruction model; The lens-free imaging reconstruction model is trained through the image pair dataset, and a to-be-reconstructed optical coding image is input into the trained lens-free imaging reconstruction model to obtain a reconstructed optical coding image.

2. The lensless imaging reconstruction method of claim 1, wherein, Pairing acquisition of optical coding images and lens images to construct an image pair dataset, specifically comprising: Optical coding images and lens images are respectively acquired for the same field of view target under the same scene, and an image pair is formed; wherein the optical coding image is acquired by a lens-free camera, and the lens image is acquired by an RGB camera; Transforming a scene or field of view target, collecting different image pairs, and collecting the image pairs into an image pair dataset wherein P I is an optical encoding image, P O is a corresponding lens image; k ∈ (1, 2, …, K) represents the serial number of the image pair, and K is the total number of image pairs in the image pair dataset.

3. The lensless imaging reconstruction method of claim 1, wherein, The lens-free imaging reconstruction model is trained through the image pair dataset, specifically comprising: The image pair dataset is divided into a training set and a test set according to a preset ratio; The neural network parameters in the Le-ADMM model, the loss compensation network and the double-feature cross transformer network are trained through the training set until the network converges, and the training is completed; inputting the optical coding images in the test set into the trained lensless imaging reconstruction model to obtain reconstructed test images wherein, is the kth optical coding image; According to the lens images in the test set and the reconstructed test images, gradient descent operation is performed on a loss function , and neural network parameters in the Le-ADMM model, the loss compensation network and the double-feature cross transformer network are optimized according to an operation result; wherein, denotes a square root of an L2 norm.

4. A lensless imaging reconstruction device, characterized in that, The device comprises: At least one processor; and The memory is in communication connection with the at least one processor; wherein The memory stores instructions executable by the at least one processor, so that the at least one processor can execute the lens-free imaging reconstruction method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Three-dimensional face reconstruction method based on double-branch feature fusion

    CN117853664A