Snapshot-style full Stokes imaging and depth estimation methods, devices and electronic equipment

By coordinating the design of cascaded metasurface modules and polarization cameras, it is possible to simultaneously acquire the full Stokes polarization information and depth information of a scene in a single exposure, solving the problem of low information acquisition efficiency in existing technologies and achieving efficient multi-dimensional information acquisition.

CN120876571BActive Publication Date: 2026-01-06ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511385817.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-01-06
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing technologies cannot achieve snapshot-style simultaneous acquisition of full Stokes polarization and depth information of a scene in a compact system, and single-function encoding limits the dimensionality and efficiency of information acquisition.

Method used

A cascaded metasurface module is used, including a first-level metasurface with depth encoding and a second-level metasurface with polarization encoding. In conjunction with a polarization camera, the original image with dual encoding is captured by a single exposure and decoded by a back-end reconstruction unit to output a depth map and a full Stokes image of the scene.

Benefits of technology

This technology enables the simultaneous acquisition of full Stokes polarization and depth information of a scene in a single exposure, solving the problem of information acquisition in existing technologies, improving information acquisition efficiency, and overcoming the shortcomings of traditional methods that require time-division scanning or combining multiple complex systems to obtain the same information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876571B_ABST
    Figure CN120876571B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of light field imaging, and provides a snapshot full-Stokes imaging and depth estimation method and device and electronic equipment. In the device, a cascaded super surface module comprises a first-stage super surface and a second-stage super surface; the first-stage super surface is a depth coding super surface, is used for carrying out depth-related phase modulation on incident scene light, and obtains first intermediate light; the second-stage super surface is a polarization coding super surface, and the second-stage super surface is used for jointly carrying out polarization-related intensity adjustment on the first intermediate light in cooperation with a polarization camera; the polarization camera captures an original image which is double coded in depth and polarization through single exposure; and a rear-end reconstruction unit receives the original image and decodes and outputs a depth map and a full-Stokes image of a scene through a preset reconstruction algorithm. The application provides a snapshot imaging system capable of simultaneously acquiring scene polarization information and depth information, and greatly improves information acquisition efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of light field imaging technology, and in particular to snapshot-type full Stokes imaging and depth estimation methods, devices and electronic equipment. Background Technology

[0002] In many cutting-edge fields such as autonomous driving, robot vision, biomedical imaging, industrial inspection, and materials science, simply acquiring two-dimensional intensity images of a scene is far from sufficient to meet application requirements. To obtain more dimensional information beyond light intensity (such as spectrum, polarization, and depth), the following techniques can be used: time-division multiplexing acquisition systems, spatial segmentation array systems (spatial multiplexing), and computational imaging systems based on optical coding elements.

[0003] Although the aforementioned existing technologies have achieved the acquisition of multi-dimensional information to a certain extent, they still have the following significant drawbacks and shortcomings:

[0004] (1) It cannot balance the snapshot-style acquisition of multi-dimensional information with system compactness:

[0005] Time-division acquisition systems contain mechanically moving parts (such as rotating waveplates / filters), resulting in large system size and complex structure. Most importantly, they cannot perform "snapshot" imaging, making it difficult to capture instantaneous information of dynamic scenes. While spatial segmentation array systems achieve snapshot imaging, this comes at the cost of sacrificing spatial resolution.

[0006] To acquire both polarization and depth information simultaneously, traditional methods require combining two or more of the aforementioned systems, which further increases the complexity, size, and cost of the system, making miniaturization and integration difficult.

[0007] (2) Single-function coding limits the dimensions and efficiency of information acquisition:

[0008] Existing computational imaging systems based on a single diffractive optical element (DOE) or metasurface are typically designed to encode information in a single dimension, such as depth or polarization only. To simultaneously acquire both depth and polarization dimensions, the design freedom of a single element is often insufficient for efficient, decoupled encoding, leading to severe crosstalk and difficulties in reconstruction.

[0009] Separating depth estimation and polarization imaging requires two independent systems and two separate images, which goes against the original intention of snapshot imaging and cannot guarantee the consistency of the scene between the two acquisitions.

[0010] In summary, existing technologies lack a method to simultaneously and efficiently acquire full Stokes polarization and depth information of a scene in a single exposure using a compact optical system. Summary of the Invention

[0011] This application provides a snapshot-type full Stokes imaging and depth estimation method, apparatus, and electronic device to solve the problem that existing technologies cannot achieve the coordinated acquisition of snapshot-type full Stokes polarization and depth information in a compact system.

[0012] This application provides a snapshot-type full Stokes imaging and depth estimation device, including a cascaded metasurface module, a polarization camera, and a back-end reconstruction unit. The cascaded metasurface module includes a first-level metasurface and a second-level metasurface. The first-level metasurface is a depth-coded metasurface used to perform depth-related phase modulation on incident scene light to obtain a first intermediate ray. The second-level metasurface is a polarization-coded metasurface, which works in conjunction with the polarization camera to jointly perform polarization-related intensity adjustment on the first intermediate ray. The polarization camera is also used to capture the original image encoded by both depth and polarization in a single exposure. The back-end reconstruction unit is used to receive the original image and decode and output the depth map and full Stokes image of the scene using a preset reconstruction algorithm.

[0013] According to the snapshot-type full Stokes imaging and depth estimation device provided in this application, the first-level metasurface is a polarization-insensitive pure phase element. By optimizing the phase distribution, a point spread function that varies with the target depth is generated, and the polarization state of the light remains unchanged. The phase distribution in the first-level metasurface is optimized based on Zernike polynomial fitting, and a circular nanopillar array is disposed on the first-level metasurface.

[0014] According to the snapshot-type full Stokes imaging and depth estimation device provided in this application, the second-level metasurface adopts a checkerboard-style regional design, and the second-level metasurface is aligned with the micro-polarizer array of the polarization camera; wherein, by optimizing the size and orientation of the rectangular nanopillars on the second-level metasurface, the second-level metasurface and the polarization camera together form a pixel-level varying polarization encoding mask; the polarization encoding mask is used to compress and encode the Stokes vector of the first intermediate ray.

[0015] According to the snapshot-type full Stokes imaging and depth estimation device provided in this application, the second-level metasurface includes a first region and a second region corresponding to different polarization channels of the polarization camera, and the first region and the second region are arranged alternately; wherein, a rectangular nanopillar array is disposed on the first region for working together with the polarization camera to compress and encode the first intermediate light; the second region is a transparent substrate for directly transmitting the first intermediate light to the polarization camera.

[0016] According to the snapshot-type full Stokes imaging and depth estimation device provided in this application, the back-end reconstruction unit is a pre-trained reconstruction neural network; the weight parameters of the pre-trained reconstruction neural network, the nanostructure parameters of the first-level metasurface, and the nanostructure parameters of the second-level metasurface are the result of joint optimization; wherein, the joint optimization is achieved in the following way: by using a preset physical model, the nanostructure parameters of the first-level metasurface, the nanostructure parameters of the second-level metasurface, and the weight parameters of the reconstruction neural network are associated, and the network is trained using sample data, and iteratively optimized using a backpropagation algorithm until the loss function between the reconstruction result and the true value is minimized.

[0017] This application also provides a snapshot-type full Stokes imaging and depth estimation method. Using the aforementioned snapshot-type full Stokes imaging and depth estimation device, the method includes: using a first-level metasurface to perform depth-related phase modulation on incident scene light to obtain a first intermediate ray; the first-level metasurface is a depth-encoded metasurface; using a second-level metasurface in conjunction with a polarization camera to jointly perform polarization-related intensity adjustment on the first intermediate ray; the second-level metasurface is a polarization-encoded metasurface; using a polarization camera to capture an original image that is dually encoded by depth and polarization through a single exposure; based on the original image, using a back-end reconstruction unit to decode and output the scene's depth map and full Stokes image through a preset reconstruction algorithm.

[0018] According to the snapshot-style full Stokes imaging and depth estimation method provided in this application, the back-end reconstruction unit includes a first restoration network and a second restoration network. Based on the original image, the back-end reconstruction unit decodes and outputs the depth map and full Stokes image of the scene using a preset reconstruction algorithm, including: obtaining first channel data, second channel data, and third channel data based on the original image; the first channel data and second channel data have not undergone polarization adjustment by the second-level metasurface; the third channel data has undergone polarization adjustment by the second-level metasurface; inputting the first channel data and second channel data into the first restoration network to obtain the depth map output by the first restoration network; the first restoration network is a Unet neural network based on dual-polarization image processing; inputting the third channel data into the second restoration network to obtain the full Stokes image output by the second restoration network; the second restoration network is a recurrent neural network based on the point spread function characteristics in optical imaging.

[0019] According to the snapshot-type full Stokes imaging and depth estimation method provided in this application, the first restoration network includes a first encoder, a second encoder, and an upsampling decoder; inputting first channel data and second channel data into the first restoration network to obtain a depth map output by the first restoration network includes: inputting the first channel data into the first encoder for downsampling processing to obtain a first output result; the first channel data is a first polarization image; inputting the second channel data into the second encoder for downsampling processing to obtain a second output result; the second channel data is a second polarization image; inputting the first output result and the second output result into the upsampling decoder for processing to obtain a depth map.

[0020] According to the snapshot-type full Stokes imaging and depth estimation method provided in this application, the second recovery network is used to perform the following steps: downsampling the third channel data to obtain a third output result; inputting the third output result into the loop processing module for a preset number of iterations until the preset number of iterations are completed, and then outputting a full Stokes image; wherein, each iteration includes: using at least one residual convolution module to calculate the input of the current iteration; using a Wiener filter layer for deblurring; using an upsampling layer to enlarge the size of the output result of the previous iteration and use it as the input of the next iteration.

[0021] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the snapshot-type full Stokes imaging and depth estimation methods described above.

[0022] This application provides a snapshot-type full Stokes imaging and depth estimation method, apparatus, and electronic device. The snapshot-type full Stokes imaging and depth estimation apparatus includes a cascaded metasurface module, a polarization camera, and a back-end reconstruction unit. The cascaded metasurface module includes a first-level metasurface and a second-level metasurface. The first-level metasurface is a depth-encoded metasurface used to perform depth-related phase modulation on incident scene light rays to obtain a first intermediate ray. The second-level metasurface is a polarization-encoded metasurface, which, in conjunction with the polarization camera, is used to jointly perform polarization-related intensity adjustment on the first intermediate ray. The polarization camera is also used to capture the original image encoded with both depth and polarization in a single exposure. The back-end reconstruction unit receives the original image and decodes and outputs the scene's depth map and full Stokes image using a preset reconstruction algorithm. Therefore, this application provides a snapshot-type imaging system capable of simultaneously acquiring scene polarization and depth information, achieving synchronous acquisition of snapshot-type full Stokes polarization and depth, greatly improving information acquisition efficiency. The core advantage of this application lies in its ability to capture both the full Stokes polarization information and depth information of a scene in a single exposure. This directly overcomes the fundamental deficiency of traditional methods that require time-division scanning or combining multiple complex systems to obtain the same information. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the structure of the snapshot-type full Stokes imaging and depth estimation device provided in the embodiments of this application.

[0025] Figure 2 This is a flowchart illustrating the snapshot-style full Stokes imaging and depth estimation method provided in the embodiments of this application.

[0026] Figure 3 This is a schematic diagram of phase optimization in the first-level metasurface provided in the embodiments of this application.

[0027] Figure 4 This is a schematic diagram comparing the simulation and actual measurement of the point spread function in the first-level metasurface provided in the embodiments of this application.

[0028] Figure 5 This is a schematic diagram of the structure of nanopillars in the first-level metasurface provided in the embodiments of this application.

[0029] Figure 6This is a schematic diagram of the final optimized phase distribution of the first-level metasurface provided in the embodiments of this application.

[0030] Figure 7 This is a schematic diagram of the structure of nanopillars in the second-level metasurface provided in the embodiments of this application.

[0031] Figure 8 This is a schematic diagram of the polarization coding mask formed by the second-level metasurface and the polarization camera provided in the embodiments of this application.

[0032] Figure 9 This is a scatter plot of the polarization coding mask provided in the embodiments of this application.

[0033] Figure 10 This is a schematic diagram of the structure of the first-level metasurface provided in the embodiments of this application.

[0034] Figure 11 This is a schematic diagram of the structure of the second-level metasurface provided in the embodiments of this application.

[0035] Figure 12 This is a schematic diagram of three channels of image for network input provided in an embodiment of this application.

[0036] Figure 13 This is a schematic diagram of the neural network structure in the back-end reconstruction unit provided in the embodiments of this application.

[0037] Figure 14 This is a schematic diagram of the original image, Stokes vector, and depth map provided in the embodiments of this application.

[0038] Figure 15 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0040] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the embodiments of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0041] In related technologies, the following techniques can be used to obtain more dimensional information besides light intensity (such as spectrum, polarization, depth, etc.):

[0042] 1. Time-division acquisition system (time multiplexing): This is the most traditional and direct method for acquiring multi-dimensional information. For example, in polarization imaging, by placing rotatable polarizers and waveplates in the optical path, images under different polarization states are acquired at different time points, and finally synthesized into a full-Stokes polarization image. Similarly, multispectral imaging can be achieved by switching filters of different wavelengths. This method sacrifices temporal resolution in exchange for an increase in the dimensionality of information.

[0043] 2. Spatial Segmentation Array System (Spatial Multiplexing): To achieve snapshot imaging with a single exposure, a miniature optical array is integrated on the focal plane of the detector. For example, a Focal Plane Array Polarization Camera covers each 2×2 pixel group of the detector with micro-polarizers at 0°, 45°, 90°, and 135° respectively, thereby capturing images of four linear polarization directions simultaneously in a single exposure.

[0044] 3. Computational Imaging Systems Based on Optical Coding Elements: Optical elements with specific coding functions are combined with backend reconstruction algorithms to acquire information in specific dimensions. For example, by designing diffractive optical elements (DOEs) or metasurfaces with depth-dependent point spread functions (PSFs), objects at different depths can produce different morphologies of blur on the detector. Then, algorithms analyze these blur features to reconstruct a depth map.

[0045] Although the aforementioned technologies have achieved the acquisition of multi-dimensional information to some extent, they still have the following significant drawbacks and shortcomings, such as the inability to balance snapshot-style acquisition of multi-dimensional information with system compactness, and the limitation of single-function coding on the dimensions and efficiency of information acquisition.

[0046] Furthermore, multi-task collaborative design is still in the exploratory stage. Although optical-algorithm co-design has become a trend, related technologies mainly focus on optimizing individual optical elements to complete a single imaging task. How to collaboratively design multiple cascaded optical elements (such as two or more metasurfaces), assigning them different encoding functions, and jointly optimize them with backend algorithms to simultaneously complete multiple imaging tasks (such as polarization and depth measurements) remains a pressing technical challenge. This complex multi-device, multi-task joint optimization places higher demands on both the design framework and the algorithms.

[0047] Based on this, this application provides a snapshot-type full Stokes imaging and depth estimation method, apparatus and electronic device, which can simultaneously and efficiently acquire the full Stokes polarization information and depth information of a scene in a single exposure.

[0048] Please see Figure 1 , Figure 1 This is a schematic diagram of the snapshot-type full Stokes imaging and depth estimation device provided in an embodiment of this application. In this embodiment, the snapshot-type full Stokes imaging and depth estimation device may include a cascaded metasurface module 110, a polarization camera 120, and a back-end reconstruction unit 130.

[0049] Specifically, the cascaded metasurface module 110 may include a first-stage metasurface 111 and a second-stage metasurface 112.

[0050] Among them, the first-level metasurface 111 is a depth-coded metasurface used to perform depth-related phase modulation on the incident scene light to obtain the first intermediate light.

[0051] The second-level metasurface 112 is a polarization-coded metasurface. The second-level metasurface 112 works in conjunction with the polarization camera 120 to jointly adjust the intensity of the first intermediate light beam based on polarization.

[0052] In addition, the polarization camera 120 is also used to capture raw images that are dually encoded in terms of depth and polarization in a single exposure.

[0053] The back-end reconstruction unit 130 is used to receive the original image and decode and output the depth map and full Stokes image of the scene through a preset reconstruction algorithm.

[0054] In this embodiment, the cascaded metasurface module 110, the polarization camera 120, and the back-end reconstruction unit 130 work collaboratively. The cascaded metasurface module 110 performs dual modulation of the light field using depth and polarization encoding, coupling these two types of physical information into the same original image. The polarization camera 120 completes data acquisition through a single exposure, achieving snapshot-style acquisition. The back-end reconstruction unit 130 can decouple the depth and polarization information in the original image, ultimately outputting a high-resolution depth map and a full Stokes image.

[0055] Specifically, the first-level metasurface 111 is designed as a depth-coded metasurface, whose function is to perform depth-related phase modulation on incident scene light. In addition, the two-dimensional planar properties of the metasurface make it possible to replace traditional bulky optical systems (such as lens groups and diffraction elements), enabling system miniaturization.

[0056] Optionally, the first-level metasurface 111 can apply a phase delay that is a function of depth to light rays at different depths (i.e., different object distances) through the phase modulation capability of subwavelength structures (such as nanopillars, nanopores, etc.).

[0057] The second-level metasurface 112 is designed as a polarization-coded metasurface, working in conjunction with the polarization camera 120 to perform polarization-dependent intensity adjustment on the first intermediate ray. The second-level metasurface 112 can impose different transmittances or phase delays on each polarization component (e.g., linear polarization of 0°, 45°, 90°, 135° or circular polarization of left-hand / right-hand rotation).

[0058] Furthermore, the second-level metasurface 112, combined with the polarization camera 120, can perform compressed sampling of the Stokes vector.

[0059] The polarization camera 120 may be equipped with a micro-polarizer array, such as a pixel-level polarization filter array. The polarization camera 120, in conjunction with the second-level metasurface 112, can further convert the polarization state into a measurable intensity signal.

[0060] Among them, the second-level metasurface 112 can dynamically adjust the polarization coding efficiency (e.g., by optimizing the geometric parameters of the metasurface structure) to improve the signal-to-noise ratio.

[0061] The sensor array of the polarization camera 120 can capture the dual-coded light field intensity distribution to generate the original image. The polarization camera 120 does not require time-division or aperture-division; it can acquire complete information in a single exposure, making it suitable for dynamic scenes. Furthermore, by compressing high-dimensional data containing depth and polarization information into a two-dimensional original image, the depth map and full Stokes image of the scene can be obtained through backend algorithms.

[0062] In the snapshot-type full Stokes imaging and depth estimation device of this embodiment, a cascaded metasurface cooperative coding architecture is adopted. By cascading a first-level metasurface as depth phase encoding and a second-level metasurface as polarization intensity encoding, functional reuse and system miniaturization are achieved. In addition, multi-dimensional information acquisition in a single exposure is realized. Through joint modulation of metasurface and polarization camera, depth and full Stokes polarization information are captured simultaneously in a single exposure.

[0063] For example, scene light first enters a first-level metasurface (depth-encoded metasurface), which modulates the light field with depth-related phase. The modulated light continues to propagate and enters a second-level metasurface (polarization-encoded metasurface), which, together with a subsequent polarization camera, modulates the light field with polarization-related intensity. The polarization camera captures a raw image encoded with both depth and polarization in a single exposure. This raw image is fed into a pre-trained reconstruction neural network (which can be provided by a back-end reconstruction unit), which ultimately decodes and outputs a full Stokes image (S0, S1, S2, S3) and a depth map of the scene.

[0064] The present application provides a snapshot imaging device capable of simultaneously acquiring scene polarization and depth information, enabling synchronous acquisition of full Stokes polarization and depth information in a snapshot format, thus greatly improving information acquisition efficiency. The core advantage of this embodiment lies in its ability to capture both full Stokes polarization and depth information of a scene in a single exposure, directly overcoming the fundamental deficiency of related methods that require time-division scanning or combining multiple complex systems to acquire the same information.

[0065] This application also provides a snapshot-type full-Stokes imaging and depth estimation method, using the above-described snapshot-type full-Stokes imaging and depth estimation device.

[0066] Please see Figure 2 , Figure 2 This is a flowchart illustrating the snapshot-type full Stokes imaging and depth estimation method provided in this application embodiment. In this embodiment, the snapshot-type full Stokes imaging and depth estimation method may include steps S210 to S240, and the specific steps are as follows:

[0067] S210: The first intermediate ray is obtained by performing depth-related phase modulation on the incident scene light using the first-level metasurface; the first-level metasurface is a depth-coded metasurface.

[0068] S220: The second-level metasurface works in conjunction with a polarization camera to adjust the intensity of the first intermediate light beam based on polarization; the second-level metasurface is a polarization-coded metasurface.

[0069] S230: Utilizes a polarization camera to capture raw images encoded with both depth and polarization in a single exposure.

[0070] S240: Based on the original image, the back-end reconstruction unit decodes and outputs the scene's depth map and full Stokes image using a preset reconstruction algorithm.

[0071] This embodiment uses the aforementioned snapshot-type full Stokes imaging and depth estimation device. It achieves depth encoding by using a first-stage metasurface in a cascaded metasurface module to perform depth-related phase modulation on the incident light; and by using a second-stage metasurface in conjunction with a polarization camera to perform polarization-related intensity adjustment on the light, completing polarization encoding. The polarization camera captures the dual-encoded original image in a single exposure. Finally, the back-end reconstruction unit receives the original image, decodes it using a preset algorithm, and outputs the scene's depth map and full Stokes image.

[0072] The above embodiment provides a snapshot imaging method that can simultaneously acquire scene polarization and depth information, enabling synchronous acquisition of full Stokes polarization and depth information in a snapshot, greatly improving information acquisition efficiency; it can capture full Stokes polarization and depth information of a scene in a single exposure.

[0073] In some embodiments, the first-level metasurface is a polarization-insensitive pure phase element that generates a point spread function that varies with the target depth by optimizing the phase distribution, while the polarization state of the light remains unchanged; wherein the phase distribution in the first-level metasurface is optimized based on Zernike polynomial fitting, and a circular nanopillar array is disposed on the first-level metasurface.

[0074] Please see Figures 3-6 , Figure 3 This is a schematic diagram of phase optimization in the first-level metasurface provided in the embodiments of this application. Figure 4 This is a schematic diagram comparing the simulation and measured results of the point spread function in the first-order metasurface provided in this application embodiment. Figure 5 This is a schematic diagram of the structure of nanopillars in the first-level metasurface provided in the embodiments of this application. Figure 6 This is a schematic diagram of the final optimized phase distribution of the first-level metasurface provided in the embodiments of this application.

[0075] like Figure 3 As shown, in depth encoding, the phase of the first-order metasurface can be expressed by a 40-term Zernike polynomial. and its optimization coefficients This is achieved by stacking. This design allows it to generate point spread functions (PSFs) with deep dependencies; such as... Figure 4As shown, point light sources at different depths will form light spots of different shapes after passing through the first-level metasurface; from left to right and from top to bottom, the point spread functions at different depths are shown, ranging from 0.16m to 1.27m. The first one on the top left is the point spread function at 0.16m, and the last one on the bottom right is the point spread function at 1.27m.

[0076] To achieve polarization insensitivity, the first-level metasurface in this embodiment can be an array of cylindrical nanopillars, such as... Figure 5 As shown.

[0077] The first-level metasurface comprises a substrate and nanopillars. The substrate can be made of silicon dioxide (SiO2), and the nanopillars can be made of silicon nitride (SiNx). Each nanopillar has a substrate side length of P, a nanopillar length of h, and a circular cross-section with radius r. The optimized phase distribution of the first-level metasurface can be found in [reference needed]. Figure 6 .

[0078] In some embodiments, the second-level metasurface adopts a checkerboard-style regional design, and the second-level metasurface is aligned with the micro-polarizer array of the polarization camera; wherein, by optimizing the size and orientation of the rectangular nanopillars on the second-level metasurface, the second-level metasurface and the polarization camera together form a pixel-level varying polarization coding mask; the polarization coding mask is used to compress and encode the Stokes vector of the first intermediate ray.

[0079] In some embodiments, the second-level metasurface includes a first region and a second region corresponding to different polarization channels of the polarization camera, the first region and the second region being arranged alternately; wherein, a rectangular nanopillar array is disposed on the first region for working together with the polarization camera to compress and encode the first intermediate light; the second region is a transparent substrate for directly transmitting the first intermediate light to the polarization camera.

[0080] Please see Figures 7-9 , Figure 7 This is a schematic diagram of the structure of nanopillars in the second-level metasurface provided in the embodiments of this application. Figure 8 This is a schematic diagram of the polarization-encoded mask formed by the second-level metasurface and the polarization camera provided in the embodiments of this application. Figure 9 This is a scatter plot of the polarization coding mask provided in the embodiments of this application.

[0081] The second-level metasurface adopts a checkerboard pattern design. It consists of alternating regions with rectangular nanopillars and transparent regions without nanopillars. The second-level metasurface is precisely aligned with the micro-polarizer array of the polarization camera.

[0082] like Figure 7As shown, the second-level metasurface includes a substrate and nanopillars. The substrate can be made of silicon dioxide (SiO2), and the nanopillars can be made of silicon nitride (SiNx). The side length of the substrate corresponding to each nanopillar is P1, the length of the nanopillar is h1, the cross-section of the nanopillar is rectangular, the width of the rectangle is w, and the length of the rectangle is l.

[0083] By optimizing the size and orientation of rectangular nanopillars on the second-level metasurface to form a pixel-level varying polarization-coded mask MS together with a polarization camera, the Stokes vector of the incident light is compressed and encoded.

[0084] like Figure 8 As shown, the four coded masks (M0, M1, M2 and M3 in the left figure) determined by the second-level metasurface and the polarization camera, when applied to the 4-channel Stokes vector, can produce a coded image (the apple basket in the right figure).

[0085] like Figure 9 The image shows a scatter plot of the encoded mask optimized using the proposed optical-algorithm co-optimization framework. This encoded mask enables compressed sampling and reconstruction of Stokes vectors.

[0086] Specifically, in the 3D point map, the coding mask is mainly divided into two lobes, which is the effect of the 0-degree and 90-degree polarization channels of the polarization camera (these two channels directly correspond to the metasurface). In addition, the M2 component is basically concentrated near 0, which is also predictable, because the Stokes component corresponding to M2 represents the polarization information of 45 degrees and 135 degrees, and these two types of information can be obtained directly from the region where no nanopillars are placed.

[0087] Please see Figure 10 and Figure 11 , Figure 10 This is a schematic diagram of the structure of the first-level metasurface provided in the embodiments of this application. Figure 11 This is a schematic diagram of the structure of the second-level metasurface provided in the embodiments of this application.

[0088] in, Figure 10 and Figure 11 These are scanning electron microscope (SEM) images of the first-order metasurface and the second-order metasurface, respectively.

[0089] In related technologies, a single coding element has limited degrees of freedom when simultaneously modulating multiple physical dimensions, which easily leads to crosstalk. This embodiment employs a cascaded double metasurface structure to decompose the complex dual coding task. Specifically, as follows:

[0090] The first-order metasurface, responsible for encoding depth information, is designed as a polarization-insensitive pure phase element. Its phase distribution is optimized (e.g., by fitting with Zernike polynomials) to produce a unique point spread function (PSF) that varies with the target depth. This means that objects at different distances will be imaged with different, recognizable blur patterns, while the polarization state of light remains unchanged throughout the process.

[0091] ;

[0092] in, S D This represents a deep-encoded Stokes vector. S in The Stokes vector representing the incident radiation. M z D For depth z Binarized occlusion mask at the location, It refers to the dotted extension function.

[0093] The second-level metasurface is specifically responsible for encoding polarization information and works in conjunction with the subsequent polarization camera to perform polarization modulation on the depth-encoded light field.

[0094] In this embodiment, a functionally decoupled cascaded metasurface architecture is employed to achieve efficient and independent encoding of multi-dimensional information. This "functionally decoupled" design avoids the complexity and crosstalk issues associated with depth and polarization hybrid encoding on a single element. Each metasurface level focuses on a single task, significantly improving encoding efficiency and subsequent decoding accuracy, successfully overcoming the bottleneck that a single encoding element cannot support multi-dimensional, high-quality encoding.

[0095] Furthermore, this embodiment employs a unique checkerboard pattern design in the second-level metasurface, which is a "compression + multiplexing" hybrid polarization encoding scheme that enables full Stokes measurement and improves depth estimation accuracy.

[0096] Compression Encoding: Optimized rectangular nanopillar structures are arranged in specific areas of the checkerboard pattern (e.g., corresponding to pixels in the 0° and 90° channels of a polarization camera). These structures work in conjunction with the polarization camera to compress and encode the Stokes vectors (S0, S1, S2, S3) of the incident light. This not only measures linear polarization but also captures circular polarization information (S3) that conventional polarization cameras cannot measure, thus achieving true full-Stokes imaging while avoiding the spatial resolution loss of traditional array cameras.

[0097] ;

[0098] in, SD This represents a deep-encoded Stokes vector. I S-D This represents the intensity image acquired after depth-Stokes encoding. M S This represents the encoding mask determined by the Mueller matrix of the metasurface and the polarization camera. N Represents noise.

[0099] Polarization multiplexing: No nanopillar structures (i.e., a transparent substrate) are used in the remaining areas of the checkerboard pattern (e.g., pixels corresponding to the 45° and 135° channels of the polarization camera). This allows the light from both polarization channels to be directly captured by the polarization camera, resulting in two clear, high-quality polarization images. These two images are specifically used as input for monocular depth estimation tasks.

[0100] In summary, this embodiment solves several problems at once. First, compressed encoding overcomes the limitations of traditional polarization cameras, such as their inability to measure circular polarization (S3) and low spatial resolution. Second, the innovative "polarization multiplexing" strategy provides the depth estimation algorithm with richer dual-channel input information than a single intensity map, thereby significantly improving the accuracy and robustness of monocular depth estimation.

[0101] In some embodiments, the back-end reconstruction unit is a pre-trained reconstruction neural network; the weight parameters of the pre-trained reconstruction neural network, the nanostructure parameters of the first-level metasurface, and the nanostructure parameters of the second-level metasurface are the result of joint optimization; wherein, the joint optimization is achieved in the following way: by using a preset physical model, the nanostructure parameters of the first-level metasurface, the nanostructure parameters of the second-level metasurface, and the weight parameters of the reconstruction neural network are associated, training is performed using sample data, and iterative optimization is performed using the backpropagation algorithm until the loss function between the reconstruction result and the true value is minimized.

[0102] In this embodiment, an end-to-end optical-algorithm co-optimization framework is constructed to ensure optimal overall system performance. Specifically, the physical parameters of the two-stage metasurface at the front end (e.g., the Zernike coefficient of the first-stage phase and the nanostructure size corresponding to the Mueller matrix of the second-stage metasurface) and the weight parameters of the reconstruction neural network in the back-end reconstruction unit are all incorporated into a unified differentiable model. By defining a loss function that measures the difference between the final reconstruction results (polarization image and depth map) and the ground truth, the backpropagation algorithm can be used to simultaneously optimize the design of the optical hardware and the software algorithm.

[0103] ;

[0104] in, p,m,θThese represent the phase of the first-order metasurface, the Mueller matrix of the second-order metasurface, and the weights of the reconstruction neural network in the back-end reconstruction unit, respectively. Represents the optimal value. SD in This represents the input Stokes-depth image. SD recover This represents the restored Stokes-depth image.

[0105] It should be noted that the Mueller matrix is ​​a 4×4 matrix that describes the polarization characteristics of a device. Therefore, in this embodiment, the Mueller matrix can be used to describe the response of the second-stage metasurface and the polarization camera to incident light with different polarization states. The Mueller matrices of the second-stage metasurface and the polarization camera together form an encoding mask for the incident light.

[0106] In related technologies, optical design and algorithm development are separate, making it impossible to guarantee optimal matching between the two. However, the end-to-end framework in this embodiment can automatically explore the complex collaborative relationship between hardware encoding and software decoding, finding a globally optimal solution for "hardware-software synergy." This ensures that the encoding pattern generated by the designed metasurface is best suited for subsequent neural network decoding, enabling the entire system to achieve its theoretically optimal imaging quality and efficiency, thereby solving the performance degradation problem caused by hardware-algorithm mismatch.

[0107] In some embodiments, the back-end reconstruction unit includes a first restoration network and a second restoration network; based on the original image, the back-end reconstruction unit decodes and outputs a depth map and a full Stokes image of the scene using a preset reconstruction algorithm, including: obtaining first channel data, second channel data, and third channel data based on the original image; neither the first channel data nor the second channel data has undergone polarization adjustment by the second-level metasurface; the third channel data has undergone polarization adjustment by the second-level metasurface; inputting the first channel data and the second channel data into the first restoration network to obtain a depth map output by the first restoration network; the first restoration network is a Unet neural network based on dual-polarization image processing; inputting the third channel data into the second restoration network to obtain a full Stokes image output by the second restoration network; the second restoration network is a recurrent neural network based on the point spread function characteristics in optical imaging.

[0108] Optionally, the first restoration network includes a first encoder, a second encoder, and an upsampling decoder; inputting the first channel data and the second channel data into the first restoration network to obtain the depth map output by the first restoration network includes: inputting the first channel data into the first encoder for downsampling processing to obtain a first output result; the first channel data is a first polarization image; inputting the second channel data into the second encoder for downsampling processing to obtain a second output result; the second channel data is a second polarization image; inputting the first output result and the second output result into the upsampling decoder for processing to obtain the depth map.

[0109] Optionally, the second recovery network performs the following steps: downsampling the third channel data to obtain a third output result; inputting the third output result into a loop processing module for a preset number of iterations until the preset number of iterations are completed, and then outputting a full Stokes image; wherein each iteration includes: using at least one residual convolution module to calculate the input of the current iteration; using a Wiener filter layer for deblurring; using an upsampling layer to enlarge the size of the output result of the previous iteration and use it as the input of the next iteration.

[0110] A schematic diagram of the acquired raw images is shown below. Figure 8 As shown on the right. According to the design of this application, the image can be decomposed into three channels (i.e., channel 1, channel 2 and channel 3) and fed into the back-end reconstruction unit to reconstruct the neural network.

[0111] Please see Figure 12 , Figure 12 This is a schematic diagram of three channels of image for network input provided in an embodiment of this application.

[0112] Channels 1 & 2 (for depth estimation): These images (i.e., the first polarization image and the second polarization image) are directly extracted from the raw data by the images captured by the 45° and 135° channels of the polarization camera. These images are not subjected to the complex modulation of the polarization-coded metasurface, resulting in lower noise and purer information, specifically designed to improve the accuracy of depth estimation.

[0113] Channel 3 (for polarization reconstruction): Pixel information captured by the 0° and 90° channels of the polarization camera and compressed using a polarization-coded metasurface.

[0114] The reconstruction neural network in the back-end reconstruction unit can be divided into two parts: a first recovery network and a second recovery network. Please refer to [link / reference]. Figure 13 , Figure 13 This is a schematic diagram of the neural network structure in the back-end reconstruction unit provided in the embodiments of this application.

[0115] The first recovery network can be a dual-polarization input Unet neural network, which inputs 45° and 135° polarization images into two encoders with different downsampling, and then uses an upsampling decoder to receive and recover the depth estimation map.

[0116] The dual-polarization input method used in this embodiment contains more input information than depth estimation from a single intensity map, and therefore has higher accuracy.

[0117] The second restoration network can employ a PSF-aware recurrent neural network, which receives the original image (Raw) and the PSF and restores the fully focused Stokes image. For the input image, the second restoration network first downsamples it to one-eighth of its original size, and then feeds it into a series of residual convolutional modules for computation. In addition, a Wiener filter layer is added to the second restoration network. The Wiener filter layer receives the image and the PSF, performs deblurring processing, and sets the signal-to-noise ratio as a learnable parameter.

[0118] It should be noted that because the image size output by the second recovery network in the first loop does not match the original image (for example, the output image size is not equal to the input image), the output image will be upsampled by a factor of two and then input into the network again for calculation. After four loops, the recovered fully focused Stokes image can be obtained.

[0119] This embodiment uses a multi-scale loop to realize the image restoration process from coarse to fine, and the addition of point spread function and Wiener filter further improves the network's non-blind deconvolution capability by utilizing physical principles.

[0120] The above embodiment provides a snapshot-style full Stokes imaging and depth estimation method and apparatus to achieve end-to-end framework optimization: through a differentiable physical model, the nanostructure parameters of the two-level metasurface at the front end are associated with the neural network weights of the reconstruction unit at the back end. The model is trained using a large amount of synthetic data and iteratively optimized through the backpropagation algorithm until the loss function between the reconstruction result and the true value is minimized.

[0121] The two-level metasurfaces can be prepared using standard semiconductor manufacturing processes (electron beam lithography, plasma etching, etc.). Figure 4 The changes in PSF with depth were compared between simulation and experimental measurements, and the two showed good consistency, demonstrating the effectiveness of the depth coding function.

[0122] Furthermore, to verify the performance of the designed cascaded metasurface system in all-Stokes imaging and depth estimation, this embodiment also constructed an experimental optical path for quantitative analysis. The results are shown in Table 1, demonstrating extremely low measurement error with a standard polarized light source. The average depth estimation error was only 5.1%, verifying the system's quantitative measurement capability.

[0123] Table 1. Quantitative Measurement Error

[0124]

[0125] In addition, to verify the imaging capabilities, this embodiment also performed imaging and reconstruction of various complex scenes. Please refer to [link / reference]. Figure 14 , Figure 14 This is a schematic diagram of the original image, Stokes vector, and depth map provided in the embodiments of this application. S0-S3 represent the Stokes vector.

[0126] Figure 14 (a) in the image shows a pair of 3D glasses whose two lenses can transmit circularly polarized light with opposite chirality. This is difficult to observe in S0-S2, but it can be seen from S3 that the two lenses have significantly opposite values. Due to the limited field of view of the imaging system, the accuracy of restoring the darker parts of the image edges is somewhat reduced.

[0127] Figure 14 (b) in the image shows an image of a Grey butterfly specimen. The microstructures on the butterfly's wings interact with light to produce structural colors, similar to the effects of multilayer interference or diffraction gratings. This selectively reflects light of specific wavelengths and polarization directions, which can be clearly seen in the Stokes vector.

[0128] Figure 14 Images (c) and (d) show the reconstruction results of an acrylic sheet under natural and compressed conditions, respectively. A comparison reveals that the Stokes image of the compressed acrylic sheet exhibits a distinct striped pattern, reflecting the stress distribution within the material, which is difficult to obtain directly from a single intensity map. In the depth map, although the acrylic sheet and the background can be distinguished, the differences are minimal, primarily due to the object's transparent material. The depth encoding only shows significant contrast changes at the object's edges.

[0129] In addition, a scene consisting of a linear polarizer, a toy car, and a model tower was photographed, and the reconstructed result is as follows. Figure 14 As shown in (e), the polarization information and depth distribution of objects in the scene are well reflected.

[0130] Compared with related technologies, the embodiments of this application have the following significant technical advantages:

[0131] 1. This application achieves simultaneous acquisition of full Stokes polarization and depth information in a snapshot-style manner, significantly improving information acquisition efficiency. The core advantage of this application lies in its ability to simultaneously capture full Stokes (S0-S3) polarization and depth information of a scene in a single exposure (“snapshot”). This directly overcomes the fundamental deficiency of related methods that require time-division scanning or combining multiple complex systems to acquire the same information. For dynamic scenes, it effectively avoids data mismatch and motion artifacts caused by target movement or changes in illumination between different dimensions of information (such as polarization and depth), ensuring the instantaneous consistency and accuracy of the acquired multi-dimensional data. Regarding system efficiency, no mechanical scanning components are required, and the imaging speed is only limited by the camera frame rate, greatly improving the throughput and efficiency of information acquisition.

[0132] 2. The system boasts a compact structure with no moving mechanical parts, facilitating miniaturization and integration. This application utilizes a metasurface as the core optical element, essentially a planar, lightweight "flat-panel optics" device. It exhibits high integration, embedding complex depth and polarization encoding functions onto two paper-thin metasurfaces, allowing for compact encapsulation with an image sensor to form a miniaturized imaging module. Furthermore, it demonstrates high stability, containing no macroscopic moving mechanical parts (such as rotating waveplates), resulting in a robust structure, strong resistance to vibration and environmental interference, and high reliability. This system has broad application prospects; its miniaturization, lightweight design, and high stability make it easily integrated into platforms with space, power consumption, and weight constraints, such as smartphones, drones, endoscopes, and sensing units in autonomous vehicles.

[0133] 3. Hardware and software co-optimization for optimal overall system performance. This embodiment adopts an end-to-end joint optimization design concept. The physical structural parameters of the front-end metasurface and the algorithm parameters of the back-end neural network are jointly optimized within the same framework, ensuring that the encoded pattern generated by the optical hardware is "tailor-made" for the back-end decoding algorithm. This deep coupling and co-evolution of hardware and software enables the finding of the globally optimal solution for the entire imaging chain. Compared to the fragmented process of "designing hardware first, then developing algorithms" in related technologies, this application maximizes the potential of the system, achieving significant improvements in key performance indicators such as signal-to-noise ratio and reconstruction accuracy.

[0134] On the other hand, this application also provides an electronic device, please refer to... Figure 15 , Figure 15 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of this application, such as... Figure 15As shown, the electronic device may include memory 1520, processor 1510, and a computer program stored in memory 1520 and executable on processor 1510. When processor 1510 executes the program, it can implement a snapshot-style full Stokes imaging and depth estimation method, which may include:

[0135] The first intermediate ray is obtained by performing depth-related phase modulation on the incident scene light using a first-level metasurface; the first-level metasurface is a depth-encoded metasurface. The second-level metasurface, in conjunction with a polarization camera, performs polarization-related intensity adjustment on the first intermediate ray; the second-level metasurface is a polarization-encoded metasurface. The polarization camera captures the original image, which is dually encoded by depth and polarization, through a single exposure. Based on the original image, the back-end reconstruction unit decodes and outputs the scene's depth map and full Stokes image using a preset reconstruction algorithm.

[0136] Optionally, the electronic device may further include a communication bus 1530 and a communication interface 1540, wherein the processor 1510, the communication interface 1540, and the memory 1520 communicate with each other through the communication bus 1530. The processor 1510 can call the computer program in the memory 1520 to execute the snapshot-style full Stokes imaging and depth estimation methods provided by the above methods.

[0137] Furthermore, the logical instructions in the aforementioned memory 1520 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0138] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0139] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A snapshot full-Stokes imaging and depth estimation apparatus, characterized by, The device comprises a cascaded metasurface module, a polarization camera and a backend reconstruction unit; the cascaded metasurface module comprises a first-stage metasurface and a second-stage metasurface; The first-stage metasurface is a depth-encoding metasurface, which is used for depth-dependent phase modulation of incident scene light to obtain first intermediate light; The second-stage metasurface is a polarization-encoding metasurface, which cooperates with the polarization camera to jointly perform polarization-dependent intensity adjustment on the first intermediate light; The polarization camera is further configured to capture a raw image that is doubly encoded by depth and polarization through a single exposure; The backend reconstruction unit is configured to receive the raw image and decode output a depth map and a full Stokes image of a scene through a preset reconstruction algorithm; The backend reconstruction unit is a pre-trained reconstruction neural network; the weight parameters of the pre-trained reconstruction neural network, the nanostructure parameters of the first-stage metasurface and the nanostructure parameters of the second-stage metasurface are the results of joint optimization; The joint optimization is realized based on the following manner: the nanostructure parameters of the first-stage metasurface, the nanostructure parameters of the second-stage metasurface and the weight parameters of the reconstruction neural network are associated through a preset physical model, sample data is used for training, and the loss function between the reconstruction result and the true value is minimized through iterative optimization of a back propagation algorithm.

2. The snapshot full-Stokes imaging and depth estimation apparatus of claim 1, wherein, The first-stage metasurface is a polarization-insensitive pure phase element, which produces a point spread function varying with target depth by optimizing the phase distribution, and the polarization state of the light remains unchanged. The phase distribution in the first-stage metasurface is optimized and designed based on Zernike polynomial fitting, and the first-stage metasurface is provided with a circular nanocolumn array.

3. The snapshot full-Stokes imaging and depth estimation apparatus of claim 1, wherein, The second-stage metasurface adopts a regionalized design in a chessboard format, and is aligned with a micro-polarization plate array of the polarization camera. The size and direction of the rectangular nanocolumns on the second-stage metasurface are optimized to make the second-stage metasurface and the polarization camera jointly form a pixel-level variable polarization encoding mask, which is used for compressively encoding the Stokes vector of the first intermediate light.

4. The snapshot full-Stokes imaging and depth estimation apparatus of claim 3, wherein, The second-stage metasurface comprises a first region and a second region corresponding to different polarization channels of the polarization camera, and the first region and the second region are staggered. The first region is provided with a rectangular nanocolumn array, which is used for jointly acting with the polarization camera to compressively encode the first intermediate light; and the second region is a transparent substrate, which is used for directly transmitting the first intermediate light to the polarization camera.

5. A method of snapshot full stokes imaging and depth estimation, the method comprising: The method comprises the following steps: a first-stage metasurface is used to perform depth-dependent phase modulation on incident scene light to obtain first intermediate light; the first-stage metasurface is a depth-encoding metasurface; a second-stage metasurface is used to jointly perform polarization-dependent intensity adjustment on the first intermediate light with a polarization camera; the second-stage metasurface is a polarization-encoding metasurface. capture, by a single exposure, a raw image doubly encoded with depth and polarization by the polarized camera; decode, based on the raw image, a depth map and a full Stokes image of the scene by a preset reconstruction algorithm using a backend reconstruction unit.

6. The snapshot full-Stokes imaging and depth estimation method of claim 5, wherein, The backend reconstruction unit comprises a first recovery network and a second recovery network; the decoding, based on the raw image, a depth map and a full Stokes image of the scene by a preset reconstruction algorithm using a backend reconstruction unit comprises: obtaining first channel data, second channel data and third channel data based on the raw image; the first channel data and the second channel data are not subjected to polarization adjustment by the second-order hyper-surface; the third channel data has been subjected to polarization adjustment by the second-order hyper-surface; inputting the first channel data and the second channel data into the first recovery network to obtain a depth map output by the first recovery network; the first recovery network is a Unet neural network based on dual-polarization image processing; inputting the third channel data into the second recovery network to obtain a full Stokes image output by the second recovery network; the second recovery network is a recurrent neural network based on the point spread function characteristics in optical imaging.

7. The snapshot full-Stokes imaging and depth estimation method of claim 6, wherein, The first recovery network comprises a first encoder, a second encoder and an up-sampling decoder; the inputting the first channel data and the second channel data into the first recovery network to obtain a depth map output by the first recovery network comprises: inputting the first channel data into the first encoder for down-sampling processing to obtain a first output result; the first channel data is a first polarization image; inputting the second channel data into the second encoder for down-sampling processing to obtain a second output result; the second channel data is a second polarization image; inputting the first output result and the second output result into the up-sampling decoder for processing to obtain the depth map.

8. The snapshot full-Stokes imaging and depth estimation method of claim 6, wherein, The second recovery network is configured to perform the following steps: performing down-sampling processing on the third channel data to obtain a third output result; inputting the third output result into a recurrent processing module for iterative calculation for a preset number of times, and outputting the full Stokes image after the preset number of iterations is completed; wherein each iteration comprises: calculating the input of the current iteration by using at least one residual convolution module; performing deblurring processing by using a Wiener filtering layer, enlarging the size of the output result of the previous iteration by using an up-sampling layer, and taking the output result as the input of the next iteration.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the snapshot full Stokes imaging and depth estimation method according to any one of claims 5 to 8. The processor executes the computer program to implement the snapshot full Stokes imaging and depth estimation method according to any one of claims 5 to 8.

Citation Information

Patent Citations

  • Large-depth metasurface polarization holographic 3D display method

    US20250164929A1

  • KR20230085797A