Method for design and optimization of dynamic phase mask based on deep learning

By designing a dynamic phase mask using an end-to-end optimization method based on deep learning, the limitations of dynamic phase masks in terms of quantization accuracy and optical imaging performance are overcome, achieving more efficient depth estimation and depth-of-field extension, which is suitable for computational imaging systems.

CN119335805BActive Publication Date: 2025-11-04奈米科学仪器装备(杭州)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411291608.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-11-04
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

Existing technologies struggle to generate high-quality time-averaged dynamic phase masks due to limited phase accuracy during quantization, and traditional methods also have limitations in optical imaging performance.

Method used

A deep learning-based end-to-end approach is used to design and optimize a time-averaged dynamic phase mask. A set of PSFs is generated using Zemax optical simulation software. The parameters of the phase mask and the image reconstruction network are optimized using image processing and neural networks to achieve multi-frame intensity averaging and improve accuracy.

Benefits of technology

It improves the depth estimation and depth-of-field extension performance of computational imaging systems, enhances the expressiveness and imaging efficiency of phase masks, and is suitable for dynamic phase mask design in computational imaging systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119335805B_ABST
    Figure CN119335805B_ABST
Patent Text Reader

Abstract

The application provides a design and optimization method of a dynamic phase mask based on deep learning, and comprises the following steps: step 1, obtaining a PSF set; step 2, obtaining an encoded image; and step 3, synchronously optimizing the dynamic phase mask and an image reconstruction network. The application designs and optimizes a time-averaged dynamic phase mask through deep learning in an end-to-end mode, and the dynamic phase mask has stronger expressiveness than a traditional static mask in a computational imaging system, and can enhance the performance of depth estimation and extended depth of field imaging in the computational imaging.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computational imaging, and particularly relates to a design and optimization method of dynamic phase mask based on deep learning. BACKGROUND

[0002] Computational imaging is a discipline that enhances or extracts image information through computational methods. It combines techniques from optics, signal processing, computer vision, and other fields, enabling breakthroughs in traditional optical imaging limitations, achieving more efficient and flexible image processing and analysis. In the field of computational imaging, phase mask technology has attracted widespread attention due to its ability to expand the depth of field (DOF) and improve the accuracy of depth estimation. Compared with static phase masks used in traditional computational imaging engineering, dynamic phase masks are considered to have stronger expressiveness because the PSF set output by dynamic phase masks is strictly non-convex, so the design space of PSF generated by dynamic phase masks is fundamentally more suitable for computational imaging engineering than static masks. However, in the process of designing dynamic phase masks, due to the involvement of a large number of quantization processes, the phase precision of modulation is limited, thus creating a demand for high-quality time-averaged dynamic phase masks.

[0003] Spatial light modulator (SLM) is an optical device that can modulate the spatial distribution characteristics of light waves. SLM controls the light field by changing the phase, amplitude, or polarization of incident light, so it is widely used in holographic display, optical information processing, adaptive optics, optical communication, and other fields. In recent years, MEMS-based SLM has achieved rapid development, which is a kind of optical device that uses micro-mechanical structures to realize light wave modulation. MEMS SLM adjusts the phase, amplitude, or direction of light waves through micron-level mechanical movement to achieve precise control of the light field, with high response speed, low power consumption, and integration characteristics. Due to its very small inertia, it has faster response speed than traditional liquid crystal SLM, so it is very suitable for high-speed dynamic modulation, and for the design of dynamic phase masks, the effective precision can be improved by intensity averaging of multiple frames to overcome the impact of quantization.

[0004] Traditional optical depth estimation methods use sensors and optical devices to encode and recover depth information. This approach exploits the depth-dependent blurring caused by the aperture, which retains more high-frequency information about the scene through aperture encoding to obtain more depth information. Similar to depth estimation, static phase masks are widely used to produce invariable custom PSFs, achieving better depth-of-field imaging effects. However, these traditional optical-driven methods have gradually been replaced by deep learning neural networks that allow joint optimization of optical elements and image reconstruction networks. Deep learning techniques can be used to jointly train optical parameters and image reconstruction network parameters by encoding images to retain additional information about the scene, and then using a neural network to reconstruct the image. A differentiable model is used for light forward propagation in this process, and backpropagation is used to update system parameters simultaneously with the neural network. The effectiveness of this method has been demonstrated in extended depth of field, depth estimation, and holography. Although these previous methods have successfully improved imaging performance, they have generally focused on designing and optimizing simple optical structures. SUMMARY

[0005] The purpose of the present application is to provide a deep learning-based dynamic phase mask design and optimization method, which aims to design and optimize time-averaged dynamic phase masks through deep learning end-to-end, which have stronger expressiveness than traditional static masks in computational imaging systems and can enhance the performance of depth estimation and extended depth of field imaging in computational imaging. The technical solution adopted is:

[0006] A deep learning-based dynamic phase mask design and optimization method, comprising the following steps:

[0007] Step 1, set N initial phase masks to be optimized, generate their respective corresponding PSFs;

[0008] Specifically comprising the following steps:

[0009] Step 1A, call Zemax optical simulation software, input an initial phase mask and the parameters of the imaging lens structure, and obtain the PSF of the optical system;

[0010] Step 1B, switch different initial phase mask inputs, repeat step 1A, and output the PSF set of the optical system under the influence of the corresponding N phase masks for subsequent optimization;

[0011] Step 2, input the initial image, use the generated series of PSFs to convolve the image to obtain depth-dependent convolution results, and average the images produced by each phase mask to obtain the final encoded image, specifically comprising the following steps:

[0012] Step 2A, the image processing unit receives the initial image C of the target, the initial image is convolved with each PSF to generate blurred images corresponding to different phase masks;

[0013] Step 2B, the N blurred images are averaged to generate the encoded image B;

[0014] Step 3, the dynamic phase mask and the image reconstruction network are optimized synchronously, which specifically includes the following steps:

[0015] Step 3A, the image reconstruction network receives the encoded image and outputs the network image W;

[0016] Step 3B, the image processing unit compares the network image W with the initial image C in step 2A in terms of loss, and adjusts the training parameters in the loss function of the phase mask sequence and the image reconstruction network according to the comparison result.

[0017] Preferably, the specific steps of calculating the loss in step 3B include:

[0018] Based on the monocular depth estimation network, the predicted depth D of the image is calculated, and then the mean square error loss L Depth is calculated.

[0019]

[0020] Where N is the number of pixels.

[0021] -True depth.

[0022] Preferably, the specific steps of calculating the loss in step 3B include:

[0023] The Attention U-Net is used to reconstruct the full-focus image, and the reconstructed predicted image I' is output, and then the mean square error loss L AiF is calculated.

[0024]

[0025] Where N is the number of pixels; AiF is the full-focus.

[0026] -True full-focus image.

[0027] Preferably, the loss function in step 3B is:

[0028]

[0029] Where I is the initial image C.

[0030] -Encoded image B.

[0031] - network image W;

[0032] - image simulation consistency loss;

[0033] - original image loss;

[0034] - reconstruction loss;

[0035] d - depth of the image;

[0036] d' - depth of the other image;

[0037] ω1 - first hyperparameter; ω2 - second hyperparameter;

[0038] wherein the neural network weights are adjusted in step 4B to reduce and

[0039] Preferably, the specific steps of phase switching in step 1 during actual application are as follows:

[0040] In the actual imaging process, high-speed switching between the optimized N phase masks is realized by a MEMS spatial light modulator.

[0041] Preferably, step 2B specifically comprises the following steps:

[0042] The blurred image is cropped or resized, the intensity of each corresponding pixel is accumulated and the average value is calculated, the calculated average intensity value matrix is converted into an image format, and a final average encoding image is generated, which is used in subsequent end-to-end optimization.

[0043] Compared with the prior art, the application has the following advantages:

[0044] The application optimizes the time-averaged dynamic phase mask through end-to-end design.

[0045] The method uses high-speed phase modulation realized by an SLM to switch between multiple phase masks in a single exposure process. Since only a single frame is captured, the system uses less memory. Moreover, the light efficiency of single-exposure imaging is higher.

[0046] Single exposure allows capturing all the light in the scene within a fixed time interval. The deep learning-based end-to-end optimization system can simultaneously optimize the parameters of the dynamic phase mask and the image reconstruction system.

[0047] The dynamic phase mask has stronger expressiveness than the traditional static mask in the computational imaging system, and can enhance the performance of depth estimation and extended depth of field imaging in the computational imaging.

[0048] The application can effectively improve the performance of the optical imaging system containing the phase mask, and provides a new technical path for the depth estimation and extended depth of field imaging in the computational imaging task. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 It is a schematic diagram of the overall process of the method of the application.

[0050] Figure 2 It is a schematic diagram of the specific process using the method of the application.

[0051] Figure 3 It is a schematic diagram of the process of generating a dynamic phase mask coded image using an SLM.

[0052] Figure 4 It is a schematic diagram of the process of end-to-end optimization based on deep learning.

[0053] Figure 5 It is a schematic diagram of phase mask sequence optimization. DETAILED DESCRIPTION

[0054] The design and optimization method of the dynamic phase mask based on deep learning of the application will be described in more detail below in conjunction with the schematic diagram, wherein the preferred embodiments of the application are represented, and it should be understood that the application described herein can be modified by those skilled in the art, while still achieving the advantageous effects of the application. Therefore, the following description should be understood as extensive knowledge for those skilled in the art, and not as a limitation on the application.

[0055] As Figures 1-5 The design and optimization method of the dynamic phase mask based on deep learning can effectively enhance the performance of depth estimation and extended depth of field imaging in the computational imaging engineering. Compared with the existing phase mask design method, the phase precision of the dynamic phase mask designed and optimized by the application is higher, and the expressiveness is better, which opens up a new direction for the phase mask technology and the design of the computational imaging system.

[0056] Specifically comprising the following steps:

[0057] The MEMS spatial light modulator is a spatial light modulator (SLM) of micro-electro-mechanical system (MEMS) technology, which can realize high-response phase modulation, can construct a series of time-averaged (the time of each switching is consistent) dynamic phase mask sequence and obtain its coded image.

[0058] Step 1, set N initial phase masks to be optimized, and generate their respective PSFs.

[0059] Step 1A, call Zemax optical simulation software, input an initial phase mask and the parameters of the imaging lens structure, get the PSF of the optical system;

[0060] Step 1B, switch different initial phase mask input, repeat step 1A, output the PSF set of the optical system under the influence of corresponding N phase masks for subsequent optimization.

[0061] In the whole optimization process, the optimization system adjusts the phase mask sequence according to the neural network feedback to ensure that the PSF set of the optical system under the influence of the optimized dynamic phase mask can adapt to the computational imaging task (such as depth of field expansion). In the actual imaging process, the SLM based on MEMS is used to realize high response phase modulation, and switching between multiple phase masks is realized in a single exposure process, so as to obtain the encoded image output by multiple phase masks.

[0062] Because the PSF set output by multiple phase masks is strictly non-convex, the design space of the point spread function generated by the dynamic SLM phase mask is fundamentally more expressive than that of the static mask.

[0063] As shown in Figure 3 The optical model images the static scene through multiple phase masks.

[0064] First, a series of dynamic phase masks are designed, and the SLM is used for phase modulation. In a single exposure process, the SLM switches between the determined phase patterns at a high speed, so as to obtain the encoded image under the influence of the dynamic phase mask.

[0065] This process can model the total exposure time as 100 milliseconds, and the switching time of each exchange ranges from 1 to 16 milliseconds. As the switching time between phase masks increases, the performance of the joint optimization system will decrease.

[0066] When the time spent on switching is less than 25% of the total exposure, the effective precision can still be improved by the intensity average of multiple frames to overcome the influence of noise.

[0067] Step 2, obtain the encoded image.

[0068] Step 2A, the image processing unit receives the initial image C of the target, and the initial image is convolved with each PSF to generate a blurred image corresponding to the phase mask;

[0069] Step 2B, average the N blurred images to generate the encoded image B.

[0070] That is, input the initial static scene for imaging, and use the designed time-averaged phase mask sequence to generate a series of depth-dependent PSFs, which are convolved with the depth-independent image to obtain the depth-dependent convolution result.

[0071] The average intensity value matrix is converted to image format, and the final average encoded image is generated. The encoded image is input into the downstream depth network for image processing. The image output by the network is compared with the actual image, and the system adjusts the parameters of the dynamic phase mask and the image reconstruction system according to the comparison result, thereby realizing the cooperative optimization of the dynamic phase mask and the image reconstruction network, and finally obtaining a dynamic phase mask more suitable for computational imaging engineering.

[0072] Step 3, synchronously optimize the phase mask and the image reconstruction network.

[0073] This step includes synchronous optimization of the phase mask and the reconstruction algorithm (image reconstruction network).

[0074] In practical applications, the designed dynamic phase mask can be applied to the phase mask technology related computational imaging system, and better extended depth of field imaging and depth estimation performance can be obtained.

[0075] Step 3A, the image reconstruction network receives the encoded image and outputs the network image W;

[0076] Step 3B, the image processing unit compares the network image W with the initial image C in step 2A, and then adjusts the training parameters in the loss function of the phase mask sequence and the image reconstruction network according to the comparison result.

[0077] In the end-to-end optimization process, the system will continuously optimize the phase mask sequence through deep learning.

[0078] The mixed surface representing the phase mask plate can be represented by a function containing the radial distance of the optical axis, the curvature radius, the aspherical coefficient, the high-order polynomial coefficient, and the odd-order polynomial with phase mask characteristics. By continuously training and optimizing the related optical parameters, a dynamic phase mask more suitable for extended depth of field imaging and other computational imaging systems can be obtained.

[0079] The image reconstruction network is an important part of the optimization process.

[0080] NAFNet can be used as the image reconstruction network, which is a UNet-shaped network with excellent inter-block and intra-block networks, making it computationally efficient and easy to train.

[0081] In addition, NAFNet shows excellent performance in multiple image deblurring tasks, achieving a balance between performance and computational efficiency, making the network suitable for end-to-end optimization systems. The dynamic phase mask optimized by the system performs better than the traditional static mask in computational imaging tasks involving phase mask technology.

[0082] For the depth estimation task, the pre-trained weights of the MiDaS Small architecture can be used for monocular depth estimation, which is a well-known convolutional monocular depth estimation network designed to receive natural images and output relative depth maps.

[0083] The network is trained end-to-end using a phase mask, according to the depth reconstruction prediction D and the true depth The mean square error (MSE) loss term is defined as:

[0084]

[0085] where N is the number of pixels, this process allows simultaneous optimization of the phase template, as well as fine-tuning of the MiDaS implementation to average the final encoded image's reconstruction.

[0086] For the extended depth of field task, the Attention U-Net can be used to reconstruct the all-in-focus image, which is jointly optimized with the phase mask sequence, according to the reconstruction prediction I and the true all-in-focus image The MSE error loss term is defined as:

[0087]

[0088] where N is the number of pixels.

[0089] As Figure 4 In the forward image simulation process (black arrow), the optimization system tracks the gradient of each parameter,

[0090] and the error is backpropagated from the network output of the end-to-end optimization system (red arrow),

[0091] During this process, the dynamic phase mask and the image reconstruction network are jointly optimized.

[0092] During this optimization process, the system will continuously change the phase mask sequence through deep learning feedback, and the PSF of the optical system corresponding to the dynamic phase mask will also change. The goal of optimizing the phase mask sequence is to make the final encoded image more suitable for implementing computational imaging tasks such as extended depth of field or depth estimation.

[0093] NAFNet is used as the image reconstruction network, which is a UNet-shaped network with high computational efficiency and easy training, making it suitable for end-to-end optimization systems.

[0094] In this stage, in order to design a dynamic phase mask more suitable for the extended depth of field and other computational imaging systems, the purpose of optimizing the dynamic phase mask can be achieved by setting different depth training parameters and pursuing the consistency of the image simulation thereof.

[0095] The system loss function is designed as:

[0096]

[0097] In the formula, I, and respectively represent the real image, the encoded image and the reconstruction result.

[0098] The hyperparameters ω1 and ω2 balance different loss terms.

[0099] By optimizing the optical parameters of the dynamic phase mask and the parameters of the reconstruction network, the image simulation consistency of different depths is maximized, and the highest simulation quality is also pursued so as to realize the joint optimization of the dynamic phase mask and the image reconstruction network.

[0100] The setting of the depth of training can be a reasonable value around the initial setting of the focal length value of the different optical imaging systems.

[0101] As shown in Figure 5 , each circle in the figure represents a visualization of a phase mask. Taking the example that there are three phase masks in the phase mask sequence, a phase mask with a random initial starting point can be set for optimization (in the actual optimization process, a suitable optimization starting point can also be selected), in the optimization process, as the neural network feedback constantly changes the network parameters, the structure of the phase mask sequence is also adjusted, and both of them are optimized towards the common goal, and the optimized phase mask sequence finally constitutes the dynamic phase mask.

[0102] The above is only the preferred embodiment of the present application, and does not limit the present application in any way. Any person skilled in the art, without departing from the scope of the technical solutions of the present application, makes any form of equivalent replacement or modification of the technical solutions and technical content disclosed by the present application, etc. change, still belongs to the protection scope of the present application.

Claims

1. A method for designing and optimizing a dynamic phase mask based on deep learning, characterized in that, Includes the following steps: Step 1: Set N initial phase masks to be optimized and generate their respective PSFs; Specifically, the following steps are included: Step 1A: Call the Zemax optical simulation software, input the parameters of an initial phase mask and imaging lens structure, and obtain the PSF of this optical system; Step 1B: Switch to different initial phase mask inputs, repeat step 1A, and output the optical system PSF set under the influence of N phase masks for subsequent optimization; Step 2: Input the initial image, convolve it with the generated series of PSFs to obtain depth-dependent convolution results, and average the images generated by each phase mask to obtain the final encoded image. This includes the following steps: Step 2A: The image processing unit receives the initial image C of the target. The initial image is convolved with each PSF to generate blurred images corresponding to different phase masks. Step 2B: Average the N blurred images to generate coded image B; Step 3: Simultaneously optimize the dynamic phase mask and the image reconstruction network, specifically including the following steps: Step 3A: The image reconstruction network receives the encoded image and outputs the network image W; Step 3B: The image processing unit performs a loss comparison between the network image W and the initial image C in step 2A, and adjusts the training parameters in the phase mask sequence and the loss function of the image reconstruction network based on the comparison results.

2. The design and optimization method for a deep learning-based dynamic phase mask according to claim 1, characterized in that, The specific steps for calculating the loss in step 3B include: Based on a monocular depth estimation network, the predicted depth D of the image is calculated, followed by the mean squared error loss L. Depth ; Where N represents the number of pixels; -True depth.

3. The design and optimization method for a deep learning-based dynamic phase mask according to claim 1, characterized in that, The specific steps for calculating the loss in step 3B include: Attention U-Net is used to reconstruct the full-focus image, outputting the reconstructed predicted image I′, and then the mean squared error loss L is calculated. AiF ; Where N is the number of pixels; AiF is the focal point. -True full-focus image.

4. The design and optimization method for a deep learning-based dynamic phase mask according to claim 1, characterized in that, Loss function in step 3B for: Wherein, I - initial image C; -Encoded image B; -Network image W; - Image simulation consistency loss; -Loss of the original image; -Reconstruction losses; d - the depth of the image; d' - another image depth; ω1 - First hyperparameter; ω2 - Second hyperparameter; In step 3B, the neural network weights are adjusted to reduce... and 5. The design and optimization method for a deep learning-based dynamic phase mask according to claim 1, characterized in that, The specific steps of phase switching in step 1 during practical application are as follows: In the actual imaging process, high-speed switching between the optimized N phase masks is achieved through a MEMS spatial light modulator.

6. The design and optimization method for a deep learning-based dynamic phase mask according to claim 1, characterized in that, Step 2B specifically includes the following steps: The blurred image is cropped or resized, the intensity of each corresponding pixel is accumulated and the average value is calculated. The calculated average intensity value matrix is ​​converted into an image format to generate the final average encoded image, which is used in subsequent end-to-end optimization.

Citation Information

Patent Citations

  • Reconstruction of phase images using deep learning

    CN114730477A

  • KR20240055328A