Design method and device of monocular depth imaging equipment based on vector light field control
Through the combination of polarization-dependent regulation devices and deep reconstruction neural networks, the degree of freedom of optical regulation is expanded, the problem of limited depth information in monocular depth imaging is solved, high-throughput coding is achieved, and imaging quality and accuracy are improved.
Patent Information
- Application Number
- CN202510046977.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Existing single-lens defocus blurs provide limited depth information and confusing, limiting the accuracy and robustness of monocular depth imaging in areas such as autonomous driving, augmented reality and robotic vision.
The polarization-dependent regulation device (liquid crystal space optical modulator) and linear polarization measurement are adopted, combined with deep reconstruction neural networks, and by building a differentiable imaging model and an end-to-end training framework, the degree of freedom of optical regulation is expanded to achieve high-throughput deep information coding.
It significantly improves the quality and accuracy of monocular depth imaging, improves the deep information encoding capabilities, and enhances its application potential in fields such as autonomous driving and augmented reality.
Smart Images

Figure CN119903750B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computational photography and computer vision, and in particular to a design method and device for a monocular depth imaging device based on vector light field control. Background Art
[0002] Monocular depth imaging based on defocused images has shown significant potential for applications in emerging industries such as autonomous driving, augmented reality, and robotic vision. However, the depth information provided by a single lens's defocus blur is extremely limited and subject to ambiguity. This limitation limits the accuracy and robustness of these methods, hindering their widespread adoption and promotion in practical applications.
[0003] To improve the above-mentioned defects, a series of methods for efficiently encoding depth information through coded apertures have emerged in the field of computational photography. These methods use amplitude or phase control devices to make the optical imaging system produce a point spread function that is highly correlated with depth, providing richer visual cues for subsequent depth recovery. In recent years, a significant development in this field has been the use of end-to-end training of optical systems and deep reconstruction neural networks to replace manually designed coding schemes, thereby obtaining designs suitable for application scenarios. However, the limited optical control freedom still limits the depth information flux contained in the point spread function, resulting in the depth information encoding capability of the coded aperture system being restricted. To solve this problem, it is urgent to explore new methods to expand the optical control freedom. Summary of the Invention
[0004] Taking the vector properties (polarization distribution) of the light field as the starting point, this paper proposes a design method for a monocular depth imaging device consisting of a polarization-dependent control device (liquid crystal spatial light modulator) and linear polarization measurement. This expands the degrees of freedom of optical control and realizes high-throughput depth information encoding, thereby significantly improving the quality of depth imaging.
[0005] The technical solution adopted in the present invention is as follows:
[0006] A method for designing a monocular depth imaging device based on vector light field control includes the following steps:
[0007] Step 1: obtain a dataset containing intensity maps and true depth maps of natural scenes and divide it into a training set and a test set;
[0008] Step 2: Use a light intensity detection device and a Twyman-Green interferometer to calibrate the mapping relationship between the loaded grayscale value and the output Jones matrix in the liquid crystal spatial light modulator in the monocular depth imaging device at a specific working wavelength, wherein the light intensity detection device is used to calibrate the mapping relationship between the loaded grayscale value and the amplitude coefficient of the output Jones matrix, and the Twyman-Green interferometer is used to calibrate the mapping relationship between the loaded grayscale value and the phase coefficient of the output Jones matrix;
[0009] Step 3: constructing a differentiable imaging model based on the Jones-Mueller matrix in tensor form, using the loaded grayscale value of the liquid crystal spatial light modulator as a learnable control parameter, calculating its vector control effect and the vector light field distribution of the imaging plane, and then calculating the two-dimensional linear polarization measurement value corresponding to the natural scene in the monocular depth imaging device based on the calculated vector control effect and the vector light field distribution of the imaging plane;
[0010] Step 4: constructing a depth reconstruction neural network to reconstruct a depth map corresponding to the natural scene from the two-dimensional linear polarization measurement values;
[0011] Step 5: Establish an end-to-end training framework based on the differentiable imaging model constructed in step 3 and the deep reconstruction neural network described in step 4, using the training set described in step 1 as the framework input and the depth map reconstructed by the deep reconstruction neural network as the framework output, and use the Cramer-Rao lower bound loss function and the deep reconstruction loss function to supervise the joint optimization of the control parameters and the neural network parameters;
[0012] Step 6: Input the test set described in step 1 into the jointly optimized differentiable imaging model and depth reconstruction neural network, and evaluate the monocular depth imaging performance of the monocular depth imaging device based on the accuracy error between the output reconstructed depth map and the corresponding true depth map.
[0013] The present disclosure also provides a monocular depth imaging device design apparatus based on vector light field control, characterized in that the apparatus includes: a component for acquiring a data set containing an intensity map and a true depth map of a natural scene and dividing the data set into a training set and a test set; a component for calibrating the mapping relationship between the loaded grayscale value and the output Jones matrix in the liquid crystal spatial light modulator in the monocular depth imaging device at a specific working wavelength using a light intensity detection device and a Twyman-Green interferometer, wherein the light intensity detection device is used to calibrate the mapping relationship between the loaded grayscale value and the amplitude coefficient of the output Jones matrix, and the Twyman-Green interferometer is used to calibrate the mapping relationship between the loaded grayscale value and the phase coefficient of the output Jones matrix; and a component for constructing a differentiable imaging model based on the Jones-Mueller matrix in tensor form, using the loaded grayscale value of the liquid crystal spatial light modulator as a learnable control parameter to calculate its vector control effect and the vector light field of the imaging plane. Distribution, and then calculating the two-dimensional linear polarization measurement value corresponding to the natural scene in the monocular depth imaging device based on the calculated vector control effect and the vector light field distribution of the imaging plane; a component for constructing a deep reconstruction neural network to reconstruct the depth map corresponding to the natural scene from the two-dimensional linear polarization measurement value; a component for establishing an end-to-end training framework based on the constructed differentiable imaging model and the deep reconstruction neural network, using the training set as the framework input, the depth map reconstructed by the deep reconstruction neural network as the framework output, and using the Cramer-Rao lower bound loss function and the depth reconstruction loss function to supervise the joint optimization of the control parameters and the neural network parameters; a component for inputting the test set into the jointly optimized differentiable imaging model and deep reconstruction neural network, and evaluating the monocular depth imaging performance of the monocular depth imaging device according to the accuracy error between the output reconstructed depth map and the corresponding true depth map.
[0014] The present disclosure also provides a monocular depth imaging device design device based on vector light field control, characterized in that the device includes: a processor, and a memory, the memory storing computer-executable instructions, and the computer-executable instructions, when executed by the processor, prompt the processor to execute any of the methods described above.
[0015] The present disclosure also provides a computer program product comprising computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor is prompted to perform any of the above methods.
[0016] By introducing the polarization-dependent Jones and Mueller matrix theories, this paper constructs a differentiable imaging model capable of accurately calculating the vector control effects and vector light field distribution characteristics of polarization-dependent control devices. By combining this differentiable imaging model with a deep reconstruction neural network into an end-to-end training framework, and using statistical information theory and image-level loss functions to jointly supervise the optimization of control parameters and neural network parameters, high-throughput deep information encoding is achieved, achieving excellent performance on the constructed dataset. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the process of the present invention;
[0018] Figure 2 The figure is an overall framework diagram for realizing the method of the present invention.
[0019] Figure 3 This is a comparison chart of the results of the present invention in simulation experiments.
[0020] Figure 4 is a block diagram illustrating a design apparatus 400 according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0021] The method of the present invention is described in detail below, with detailed implementation methods and specific operating procedures provided, but the protection scope of the present invention is not limited to the following implementations.
[0022] This embodiment provides a method for designing a monocular depth imaging device based on vector light field control. Figure 1 and Figure 2 The design method can be executed by a computer or other device with processing capabilities. The specific steps are as follows:
[0023] In step S110, an intensity map of a natural scene (such as a living room, an office, and an outdoor park) is obtained. and the true depth map The dataset is divided into training set and test set according to the pre-set appropriate ratio (such as 8:2). It is a visible light image that describes the light intensity distribution of the scene, while the real depth map It reflects the physical distance of each pixel in the scene to the imaging device. Both the intensity map and the true depth map reflect the information of the same scene.
[0024] In step S120, the operating wavelength is set 532 , using a light intensity detection device (including but not limited to a thermoelectric power meter or a photoelectric power meter with a millimeter-scale detection area) and a Twyman-Green interferometer, calibrate the mapping relationship between the loaded grayscale value (or called the loaded grayscale value, the input grayscale value, or the input grayscale value) and the output Jones matrix (or called the output Jones matrix) in the liquid crystal spatial light modulator at the wavelength. The output Jones matrix of the liquid crystal spatial light modulator can be expressed as:
[0025]
[0026] in, and Indicates the direction of electromagnetic wave vibration, represents a liquid crystal spatial light modulator, and denote the amplitude and phase coefficients respectively, and The distribution represents the Euler number and the imaginary unit. The light intensity detection device is used to calibrate the mapping relationship between the loaded gray value and the amplitude coefficient of the output Jones matrix. The Twyman-Green interferometer is used to calibrate the mapping relationship between the loaded gray value and the phase coefficient of the output Jones matrix. After calibration, at the set working wavelength and load grayscale values The coordinates on the liquid crystal spatial light modulator are The Jones matrix corresponding to the pixel can be expressed as .
[0027] In step S130, a differentiable imaging model is constructed based on the tensor-based Jones-Mueller matrix. Specifically, the grayscale value of each pixel on the liquid crystal spatial light modulator is first converted to Initialized to 0 Then, the Jones matrix corresponding to the liquid crystal spatial light modulator can be derived according to the mapping relationship obtained by calibration in step S120. Taking the loaded grayscale value of the liquid crystal spatial light modulator as the learnable control parameter, the vector control effect of the entire imaging device can be expressed as:
[0028]
[0029] in, Represents the imaging device as a whole, represents the primary lens that converges scene light, represents a liquid crystal spatial light modulator, represents the Hadamard convolution between two tensors, represents the control effect of the main lens. A focal length of The control effect of the main lens can be expressed as ,in and represent Euler number and imaginary unit respectively. It represents the amplitude transmittance corresponding to the aperture plane of the main lens, which is constant at 1 within the aperture and constant at 0 outside the aperture. represents the phase control item of the primary mirror, is the wave number. After obtaining the overall vector control effect of the imaging device, the Fresnel diffraction integral can be used to calculate the vector diffraction process from the aperture plane to the imaging plane, and the vector light field in the form of the Jones matrix of the imaging plane can be obtained:
[0030]
[0031] in represents the imaging plane, and represents the spatial frequency, is the distance between the aperture plane and the sensor plane, represents the two-dimensional Fourier transform, and is a depth of The transmission factor corresponding to the incident point light source. Since the vector light field in the form of the Jones matrix is only applicable to the fully polarized incident scene, in order to make the method applicable to any incident scene, the vector light field in the form of the Jones matrix is converted to the Mueller matrix form:
[0032]
[0033] in Indicates that the matrix is in the form of a Mueller matrix, represents the vector light field in Jones matrix form, represents the complex conjugate, Indicates The two-dimensional Kronecker product of the plane, Indicates Two-dimensional matrix multiplication performed in the plane. It is used to convert the coordinate system of the two-dimensional Kronecker product calculation result from the Pauli coordinate system to the standard Cartesian coordinate system, which can be expressed as:
[0034]
[0035] and express Assuming that the imaging scene applicable to this imaging method is a natural light scene, the vector light field in the Stokes vector form, also known as the Stokes vectorized point spread function, can be expressed as:
[0036]
[0037] in Indicates Matrix product in the direction, The Stokes vector corresponding to natural light.
[0038] Finally, the linear polarization sensor simulation configuration is used on the imaging plane to collect the projection of the vector light field in the linear polarization space, that is, in the four linear polarization channels The corresponding point spread function (PSF):
[0039]
[0040] in Indicates Dimensional matrix multiplication, It is the transfer matrix that projects the Stokes vector into the linear polarization space. It simulates the working principle of the linear polarization sensor at the simulation level and can be expressed as:
[0041]
[0042] At this time, the intensity map corresponding to the natural scene in the green channel is taken as the working wavelength Input intensity map under , and discretize its true depth map into 12 layers, and use the image matting method to The depth layer is assigned to each position of discrete depth layers Scene distribution .
[0043] Combined with classical imaging theory formulas (such as the thin lens formula (Lens's Equation), the simplified camera model (Simplified Camera Model), the incoherent imaging principle (Incoherent Imaging Principle), etc.), the measurement value of the imaging device can be expressed as the convolution operation result of the scene input intensity map and the point spread function. Therefore, the linear polarization channel The corresponding two-dimensional measurements can be expressed as:
[0044]
[0045] in represents the imaging noise, represented by the standard deviation The Gaussian distribution is generated. In step S140, the depth reconstruction neural network adopts a six-layer deep learning convolutional neural network (U-Net) and includes five consecutive downsampling and five upsampling stages. Correspondingly, the depth map reconstruction process can be expressed as:
[0046]
[0047] in represents the reconstructed depth map, represents the reasoning process of the deep reconstruction neural network, represents a two-dimensional linear polarization measurement. Specifically, The inference process consists of successive downsampling and upsampling stages. First, the downsampling stage gradually extracts multi-scale features from the two-dimensional linear polarization measurements, capturing richer global information while reducing computational complexity. The upsampling stage then gradually restores resolution while effectively preserving underlying local details using the unique skip connection structure of the U-Net, enabling more accurate depth reconstruction.
[0048] In step S150, since the mathematical operations involved in the differentiable imaging model constructed in step S130 and the deep reconstruction neural network in step S140 are all differentiable operations, they can be implemented in software using the scientific computing library Pytorch. On this basis, the differentiable imaging model and the deep reconstruction neural network are encapsulated into an end-to-end training framework (such as Figure 2 ). This framework uses the training set described in step S110 as input and the depth map reconstructed by the deep reconstruction neural network as output. The design of this framework enables shared propagation of reverse gradients, thereby achieving joint optimization of control parameters and neural network parameters. Furthermore, it should be noted that the differentiable imaging model can simulate the imaging device's measurement of a natural scene, namely the two-dimensional linear polarization measurement described in step S130, and the deep reconstruction neural network can generate the corresponding depth map, thereby enabling monocular depth imaging.
[0049] For optical imaging devices, the corresponding Cramer-Rao lower bound loss function can be derived through a variety of methods, including but not limited to information theory, statistical estimation theory, maximum likelihood estimation, or variational inference. As an example, in this design method, the derivation process of the Cramer-Rao lower bound loss function corresponding to the imaging device can be as follows:
[0050] First, set the observation variable to the collected point spread function , the parameter to be estimated is the three-dimensional spatial information, expressed as For linear polarization channels , depth plane The corresponding elements of the Fisher information matrix can be expressed as:
[0051]
[0052] in and Indicates the matrix row and column index corresponding to the element, It is assumed that the collected The background noise in the 0.1 of the total energy Secondly, the Fisher information matrix is inverted, and the diagonal elements of the inverse matrix constitute the linear polarization channel In the depth plane The corresponding Cramer-Rao lower bound vector in the Cramer-Rao lower bound vector can be expressed as:
[0053]
[0054] To guide the four linearly polarized channels ( ) The corresponding point spread function (PSF) encodes more depth information, and the Cramer-Rao lower bound loss function is defined as:
[0055]
[0056] Secondly, the existing predefined deep reconstruction loss function for directly evaluating the neural network reconstruction performance can be divided into two parts. The first part can be directly expressed using the L1 norm:
[0057]
[0058] The subscript depth indicates that the loss function is used to measure the numerical difference between the depth map reconstructed by the neural network and the real depth map. Indicates the number of pixels on the depth map. In order to make the reconstructed depth map have higher fidelity, the second part is the real depth map. and reconstruct the depth map The gradient of L1 norm is solved:
[0059]
[0060] The subscript grad indicates that the loss function is used to measure the gradient difference between the depth map reconstructed by the neural network and the real depth map. and Indicates that the depth map is and The gradient component in the direction.
[0061] In summary, the loss function for joint optimization of supervisory control parameters and neural network parameters can be expressed as:
[0062]
[0063] in and Empirically determined to be 10 and 0.004.
[0064] In step S160, the jointly optimized control parameters and neural network parameters are first loaded into the liquid crystal spatial light modulator and the depth reconstruction neural network in the differentiable imaging model respectively, and then the test set is input into the differentiable imaging model to obtain the regulated two-dimensional linear polarization measurement value. , and then the two-dimensional linear polarization measurement value is fed into the deep reconstruction neural network, and finally the reconstructed depth map is output .
[0065] See also Figure 3 , the proposed design method is compared with existing different coded aperture methods on real data sets in simulation experiments. The results show that the proposed design method has good performance in common indicators for measuring the accuracy and fidelity of reconstruction depth, such as mean absolute error (MAE), root mean square error (RMSE), logarithmic absolute error ( ) and relative error ( ) all achieved the best performance. Among them, the mean absolute error (MAE), root mean square error (RMSE), logarithmic absolute error ( ) should be as small as possible, and the relative error ( ) should be as large as possible.
[0066] In addition to providing the above-mentioned design method, the present disclosure also provides a design device. The above description of the design method is also applicable to the design device, unless otherwise explicitly stated.
[0067] The design device provided by the present disclosure includes: a component for acquiring a data set containing an intensity map and a real depth map of a natural scene and dividing it into a training set and a test set; a component for calibrating the mapping relationship between the loaded grayscale value and the output Jones matrix in the liquid crystal spatial light modulator in the monocular depth imaging device at a specific working wavelength using a light intensity detection device and a Twyman-Green interferometer, wherein the light intensity detection device is used to calibrate the mapping relationship between the loaded grayscale value and the amplitude coefficient of the output Jones matrix, and the Twyman-Green interferometer is used to calibrate the mapping relationship between the loaded grayscale value and the phase coefficient of the output Jones matrix; a component for constructing a differentiable imaging model based on the Jones-Mueller matrix in tensor form, taking the loaded grayscale value of the liquid crystal spatial light modulator as a learnable control parameter, calculating its vector control effect and the vector light field distribution of the imaging plane, and then based on the calculated vector control The invention relates to a component for calculating the two-dimensional linear polarization measurement value corresponding to the natural scene in the monocular depth imaging device based on the vector light field distribution of the effect and the imaging plane; a component for constructing a deep reconstruction neural network to reconstruct the depth map corresponding to the natural scene from the two-dimensional linear polarization measurement value; a component for establishing an end-to-end training framework based on the constructed differentiable imaging model and the deep reconstruction neural network, using the training set as the framework input, the depth map reconstructed by the deep reconstruction neural network as the framework output, and using the Cramer-Rao lower bound loss function and the depth reconstruction loss function to supervise the joint optimization of the control parameters and the neural network parameters; a component for inputting the test set into the jointly optimized differentiable imaging model and deep reconstruction neural network, and evaluating the monocular depth imaging performance of the monocular depth imaging device according to the accuracy error between the output reconstructed depth map and the corresponding true depth map.
[0068] The present disclosure also provides a design device for a monocular depth imaging device based on vector light field control. The above description of the design method also applies to the design device, unless otherwise explicitly stated.
[0069] Figure 4 is a block diagram illustrating a design apparatus 400 according to an embodiment of the present disclosure.
[0070] See also Figure 4 , the design device 400 may include a processor 401 and a memory 402. The processor 401 and the memory 402 may be connected via a bus 403.
[0071] Processor 401 can perform various actions and processes according to the program stored in memory 402. Specifically, processor 401 can be an integrated circuit chip with signal processing capabilities. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor, such as an X86 architecture or an ARM architecture.
[0072] Memory 402 stores computer instructions that, when executed by processor 401, implement the aforementioned design method. Memory 402 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM), which functions as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory used in the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0073] The present disclosure also provides a computer program product comprising computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor is prompted to perform the method described above.
[0074] In addition, the present disclosure also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, can implement the above-described method. Similarly, the computer-readable storage medium in the embodiments of the present disclosure can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memory. It should be noted that the computer-readable storage medium described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0075] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a portion of code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.
[0076] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuitry, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented, as non-limiting examples, in hardware, software, firmware, dedicated circuitry or logic, general-purpose hardware or a controller or other computing device, or some combination thereof.
[0077] The exemplary embodiments of the present invention described in detail above are merely illustrative and not restrictive. It will be appreciated by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present invention, and such modifications should fall within the scope of the present invention.
Claims
1. A design method for a monocular depth imaging device based on vector light field control, characterized in that: The steps include: Step 1: obtain a dataset containing intensity maps and true depth maps of natural scenes and divide it into a training set and a test set; Step 2: Use a light intensity detection device and a Twyman-Green interferometer to calibrate the mapping relationship between the loaded grayscale value and the output Jones matrix in the liquid crystal spatial light modulator in the monocular depth imaging device at a specific working wavelength, wherein the light intensity detection device is used to calibrate the mapping relationship between the loaded grayscale value and the amplitude coefficient of the output Jones matrix, and the Twyman-Green interferometer is used to calibrate the mapping relationship between the loaded grayscale value and the phase coefficient of the output Jones matrix; Step 3: constructing a differentiable imaging model based on the Jones-Mueller matrix in tensor form, using the loaded grayscale value of the liquid crystal spatial light modulator as a learnable control parameter, calculating its vector control effect and the vector light field distribution of the imaging plane, and then calculating the two-dimensional linear polarization measurement value corresponding to the natural scene in the monocular depth imaging device based on the calculated vector control effect and the vector light field distribution of the imaging plane; Step 4: constructing a depth reconstruction neural network to reconstruct a depth map corresponding to the natural scene from the two-dimensional linear polarization measurement values; Step 5: Establish an end-to-end training framework based on the differentiable imaging model constructed in step 3 and the deep reconstruction neural network described in step 4, using the training set described in step 1 as the framework input and the depth map reconstructed by the deep reconstruction neural network as the framework output, and use the Cramer-Rao lower bound loss function and the deep reconstruction loss function to supervise the joint optimization of the control parameters and the neural network parameters; Step 6: Input the test set described in step 1 into the jointly optimized differentiable imaging model and depth reconstruction neural network, and evaluate the monocular depth imaging performance of the monocular depth imaging device based on the accuracy error between the output reconstructed depth map and the corresponding true depth map.
2. The method for designing a monocular depth imaging device based on vector light field control according to claim 1, characterized in that: In step 1, the dataset includes indoor and outdoor natural scenes, and the true depth range is set between 1m and 5m.
3. The method for designing a monocular depth imaging device based on vector light field control according to claim 1, characterized in that: In step 2, the output Jones matrix of the liquid crystal spatial light modulator is expressed as: in, and Indicates the direction of electromagnetic wave vibration, represents a liquid crystal spatial light modulator, and denote the amplitude and phase coefficients respectively, and The distribution represents the Euler number and the imaginary unit; at a specific operating wavelength Under the condition of grayscale value, the liquid crystal spatial light modulator There is a variable electro-optic effect, and the coordinates on the liquid crystal spatial light modulator are The pixel corresponds to a Jones matrix that changes with the loaded gray value .
4. The method for designing a monocular depth imaging device based on vector light field control according to claim 1, characterized in that: In step 3, the differentiable imaging model is used to simulate the measurement of the natural scene by the monocular depth imaging device, and the corresponding calculation process is as follows: Initialize the loaded grayscale value of the liquid crystal spatial light modulator, and derive the corresponding Jones matrix based on the mapping relationship obtained by calibration in step 2 to describe the vector control effect of the liquid crystal spatial light modulator; Based on the vector control effect of the liquid crystal spatial light modulator, the vector light field distribution on the imaging plane is calculated by Fresnel diffraction integral, wherein the calculation result is expressed in the form of Jones matrix; The vector light field in the form of Jones matrix is further converted into a vector light field in the form of Mueller matrix; The linear polarization sensor simulation configuration is used on the imaging plane to collect the projection of the vector light field in the linear polarization space, which is recorded as the projection of the vector light field in the four linear polarization channels ( ) corresponding point spread function (PSF); The intensity map and the true depth map corresponding to the natural scene at the specific working wavelength are input, and the two-dimensional linear polarization measurement value corresponding to the natural scene in the monocular depth imaging device is calculated by combining the obtained point spread function.
5. The method for designing a monocular depth imaging device based on vector light field control according to claim 1, characterized in that: In step 4, the depth reconstruction neural network uses a six-layer deep learning convolutional neural network (U-Net) and includes five consecutive downsampling and five upsampling stages; the depth map reconstruction process is expressed as: in represents the reconstructed depth map, represents the reasoning process of the deep reconstruction neural network, Represents linear polarization measurements.
6. The method for designing a monocular depth imaging device based on vector light field control according to claim 1, characterized in that: In step 5, the Cramer-Rao lower bound loss function is used to measure the four linear polarization channels in the joint optimization ( ) contains the depth information flux contained in the point spread function (PSF) corresponding to the image, and the depth reconstruction loss function is used to directly evaluate the reconstruction performance of the neural network.
7. The method for designing a monocular depth imaging device based on vector light field control according to claim 1, characterized in that: In step 6, the jointly optimized control parameters and neural network parameters are first loaded into the liquid crystal spatial light modulator and the depth reconstruction neural network, respectively. The test set is then input into the differentiable imaging model to obtain the controlled two-dimensional linear polarization measurement values. The two-dimensional linear polarization measurement values are then fed into the depth reconstruction neural network, and finally a reconstructed depth map is output. The monocular depth imaging performance of the monocular depth imaging device is evaluated by comparing the accuracy error between the reconstructed depth map and the corresponding real depth map using mean absolute error, root mean square error, logarithmic absolute error, and relative error indicators.
8. A design device for a monocular depth imaging device based on vector light field control, characterized in that: The design device comprises: A component for obtaining a dataset containing intensity maps and true depth maps of natural scenes and dividing the dataset into a training set and a test set; A component for calibrating the mapping relationship between the loaded grayscale value and the output Jones matrix in the liquid crystal spatial light modulator in the monocular depth imaging device at a specific working wavelength using a light intensity detection device and a Twyman-Green interferometer, wherein the light intensity detection device is used to calibrate the mapping relationship between the loaded grayscale value and the amplitude coefficient of the output Jones matrix, and the Twyman-Green interferometer is used to calibrate the mapping relationship between the loaded grayscale value and the phase coefficient of the output Jones matrix; A component for constructing a differentiable imaging model based on a Jones-Mueller matrix in tensor form, calculating its vector control effect and the vector light field distribution of the imaging plane using the loaded grayscale value of the liquid crystal spatial light modulator as a learnable control parameter, and then calculating the corresponding two-dimensional linear polarization measurement value of the natural scene in the monocular depth imaging device based on the calculated vector control effect and the vector light field distribution of the imaging plane; A component for constructing a deep reconstruction neural network to reconstruct a depth map corresponding to a natural scene from the two-dimensional linear polarization measurements; A component for establishing an end-to-end training framework based on the constructed differentiable imaging model and the deep reconstruction neural network, using the training set as the framework input, using the depth map reconstructed by the deep reconstruction neural network as the framework output, and using the Cramer-Rao lower bound loss function and the deep reconstruction loss function to supervise the joint optimization of control parameters and neural network parameters; A component for inputting the test set into the jointly optimized differentiable imaging model and depth reconstruction neural network, and evaluating the monocular depth imaging performance of the monocular depth imaging device based on the accuracy error between the output reconstructed depth map and the corresponding true depth map.
9. The monocular depth imaging device design apparatus based on vector light field control according to claim 8, characterized in that: The dataset contains both indoor and outdoor natural scenes, and its true depth range is set between 1m and 5m.
10. The monocular depth imaging device design device based on vector light field control according to claim 8, characterized in that: The output Jones matrix of the liquid crystal spatial light modulator is expressed as: in, and Indicates the direction of electromagnetic wave vibration, represents a liquid crystal spatial light modulator, and denote the amplitude and phase coefficients respectively, and The distribution represents the Euler number and the imaginary unit; at a specific operating wavelength Under the condition of grayscale value, the liquid crystal spatial light modulator There is a variable electro-optic effect, and the coordinates on the liquid crystal spatial light modulator are The pixel corresponds to a Jones matrix that changes with the loaded gray value .
Citation Information
Patent Citations
Learable flight time imaging method and system guided by fee house information amount
CN115201852A
Light field depth estimation method based on light field sequence feature analysis
CN115272435A