A Design Method of a Single-Lens Computational Imaging System for Depth Estimation

By designing a single-lens computing imaging system for depth estimation and a depth-associated fuzzy feature extraction model, the problem of difficulty in designing the optical system and the algorithm is solved, and a high-accuracy monocular depth estimation system is achieved.

CN117830115BActive Publication Date: 2025-06-20HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410016949.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-06-20
Estimated Expiration
2044-01-05

AI Technical Summary

Technical Problem

In the prior art, in the introduction of optical prior monocular depth estimation method, it is difficult to design separately for the optical system and the algorithm part, making it difficult for the system to optimize and improve the accuracy of depth estimation.

Method used

By designing a single-lens computing imaging system for depth estimation, combined with a matching depth-associated fuzzy feature extraction model, the parameters of a single-lens computing imaging system are optimized to solve the problem that the optical system and algorithm parts are difficult to design separately.

Benefits of technology

The monocular depth estimation system with high accuracy is designed, which can more accurately distinguish aberration characteristics of different depths, and improve the robustness of the system and the accuracy of depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117830115B_ABST
    Figure CN117830115B_ABST
Patent Text Reader

Abstract

The present invention discloses a design method for a single-lens computational imaging system for depth estimation, including: reading RGB images and depth maps in a dataset, dividing the RGB images into different depth regions through the depth maps, and calculating the adaptive feature depths of each depth region; aiming at the amplified PSF feature difference evaluation coefficient, constructing an initial single-lens encoding model, encoding depth information into the RGB images through the point spread functions at the adaptive feature depths of each depth region to obtain simulated images with different aberration characteristics at different depth positions; constructing a depth-correlated fuzzy feature extraction model to obtain a predicted depth image and a restored clear image, and optimizing the parameters of the single-lens computational imaging system based on the initial single-lens encoding model and the depth-correlated fuzzy feature extraction model. The present invention solves the problem that it is difficult to separately design the optical system and the algorithm part, and obtains a monocular depth estimation imaging system with high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of monocular depth estimation, and particularly relates to a design method of a single-lens computational imaging system for depth estimation. Background Art

[0002] Monocular depth estimation is an important research content in computer vision and has a wide range of applications in many fields, such as autonomous driving, virtual reality, etc. At present, there are mainly two methods for monocular depth estimation. One is a depth estimation method that completely relies on data-driven: by learning features such as the size of objects in the image, relative spatial position relationships, and semantic information through a neural network to predict depth information. Another method introduces an optical imaging part into the model and introduces prior knowledge of the imaging process into the extraction of depth features. The depth estimation that introduces optical priors believes that during the imaging process of an object, the point spread functions at different depths have different characteristics, and there are differences in its spherical aberration, chromatic aberration, defocus, etc. This depth-related feature is encoded into the image by the optical system during the imaging process of the lens, and then through the extraction of depth-related features by the neural network, the accurate depth of the object can be obtained by decoding. The completely data-driven depth estimation method has problems such as relying on datasets and poor interpretability. In the method that introduces optical priors, the depth-related features encoded into the image and the decoding of such features require high coordination, which makes it difficult to separately design the optical encoding part and the algorithm decoding part. Summary of the Invention

[0003] To solve the above technical problems, the present invention proposes a design method of a single-lens computational imaging system for depth estimation. By designing a single-lens optical system and combining a phase-matched depth-correlated fuzzy feature extraction model, the problem that it is difficult to separately design the optical system and the algorithm part in monocular depth estimation that introduces optical priors is solved, and a monocular depth estimation system with high accuracy can be designed.

[0004] To achieve the above object, the present invention provides a design method of a single-lens computational imaging system for depth estimation, including the following steps:

[0005] Read the RGB images and depth maps in the dataset, divide the RGB images into different depth regions through the depth maps, and calculate the adaptive feature depths of each depth region;

[0006] With the amplification PSF feature difference evaluation coefficient as the target, construct an initial single-lens encoding model, and encode depth information into the RGB images through the point spread functions at the adaptive feature depths of each depth region to obtain simulation images with different aberration characteristics at different depth positions;

[0007] Construct a depth - correlated fuzzy feature extraction model to obtain a predicted depth image and a restored clear image, and optimize the parameters of the single - lens computational imaging system based on the initial single - lens encoding model and the depth - correlated fuzzy feature extraction model.

[0008] Optionally, the method for dividing the RGB image into different depth regions through the depth map and calculating the adaptive feature depth of each depth region includes:

[0009] Evenly divide a given depth estimation range into several depth layers;

[0010] Determine the depth layer corresponding to each pixel according to the depth map, and segment the corresponding RGB image into different depth regions;

[0011] Calculate the average depth of different pixels in different depth regions respectively as the adaptive feature depth of each depth region.

[0012] Optionally, the method for obtaining the PSF feature difference evaluation coefficient is:

[0013]

[0014] where represents the mean of the point - spread functions of different depth layers, represents the L2 norm of the difference between the nth point - spread function and the average point - spread function, and N represents the number of depth layers.

[0015] Optionally, the method for constructing a depth - correlated fuzzy feature extraction model to obtain a predicted depth image and a restored clear image, and optimizing the parameters of the single - lens computational imaging system based on the initial single - lens encoding model and the depth - correlated fuzzy feature extraction model includes:

[0016] Perform Canny edge detection on the simulation image to extract object edge information, obtain an edge feature image with three channels. After splicing the three - channel edge feature image and the simulation image, input them into the depth - correlated fuzzy feature extraction model. Extract the aberration features in the spliced three - channel edge feature image and the simulation image through the depth - correlated fuzzy feature extraction model. The depth - correlated fuzzy feature extraction model outputs a four - channel matrix. The first channel represents the predicted depth image of the input scene, and the last three channels represent the restored clear image after aberration elimination;

[0017] Optimize the parameters of the single - lens computational imaging system based on the initial single - lens encoding model and the depth - correlated fuzzy feature extraction model to obtain a single - lens encoding model that maximizes the PSF feature difference evaluation coefficient and a depth - correlated fuzzy feature extraction model that minimizes the error between the predicted depth image and the restored clear image, thus constituting a single - lens computational imaging system for depth estimation.

[0018] Technical effects of the present invention:

[0019] (1) The present invention designs an imaging simulation method for scenes containing objects at different depths. The pixels of the real depth images in the dataset are evenly divided into N depth regions, and the average depth of the pixels in each depth region is used as the adaptive feature depth of all pixels in this region, improving the accuracy of imaging simulation. At the same time, compared with the sampling method that examines a fixed depth, the computational imaging system trained using the adaptive depth feature is more robust.

[0020] (2) The present invention designs a PSF feature difference evaluation coefficient for depth estimation to evaluate the differences between point spread functions at different depth positions. By increasing the differences between some depth point spread function features in the optical system design, the depth information encoded into the simulation image is more distinguishable, enabling the system to more accurately distinguish the aberration features at different depths.

[0021] (3) The present invention proposes a design method for a single-lens computational imaging system for depth estimation. Through the design of a single-lens optical system and in combination with a phase-matched depth-correlated fuzzy feature extraction model, it solves the problem that it is difficult to separately design the optical encoding part and the algorithm decoding part in monocular depth estimation that introduces optical priors, and a monocular depth estimation system with high accuracy can be designed. Description of the drawings

[0022] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0023] Figure 1 is a schematic flow chart of a design method for a single-lens computational imaging system for depth estimation according to an embodiment of the present invention;

[0024] Figure 2 is a schematic structural diagram of a constructed image reconstruction and depth prediction model according to an embodiment of the present invention;

[0025] Figure 3 is a comparison diagram between the RGB image, depth image in the dataset, the simulated image obtained through the single-lens encoding model, the clear restored image obtained through the computational imaging system, and the predicted depth map according to an embodiment of the present invention. Detailed implementation manners

[0026] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.

[0027] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0028] As Figure 1 shown, in this embodiment, a design method for a single-lens computational imaging system for depth estimation is provided, which is used to design a set of computational imaging systems that can be used for depth estimation in indoor scenes, including the following steps:

[0029] Step S1: Obtain the RGB image and depth map of the area to be measured, divide the RGB image into different depth regions through the depth map, and calculate the adaptive feature depth of each depth region.

[0030] Step S2: Construct an initial single-lens encoding model with the amplification PSF feature difference evaluation coefficient as the target, encode the depth information into the RGB image through the point spread function at the adaptive feature depth of each depth region, and obtain a simulated image with different aberration characteristics at different depth positions.

[0031] Step S3: Construct a depth-correlated fuzzy feature extraction model, obtain a predicted depth image and a restored clear image, and optimize the parameters of the single-lens computational imaging system based on the single-lens encoding model and the fuzzy feature extraction model.

[0032] In step S1, obtain the RGB image and depth map of the area to be measured, use the scene data as the input of the imaging model. The RGB image is a two-dimensional RGB three-channel image, and the depth map is composed of a single-channel depth map, and the single-channel depth map is composed of the actual depth at each pixel of the image. The given depth estimation range is evenly divided into N depth layers, and then the depth layer to which each pixel belongs is determined according to the depth map, so as to divide the corresponding RGB image into different depth regions. The nth depth region is marked as I n , n = 0, 1, … N - 1. Next, calculate the average depth of different pixels in the N depth regions respectively as the adaptive feature depth of each depth layer. The adaptive feature depth of the nth depth region is denoted as d n .

[0033] Specifically, the NYU Depth dataset is divided into a training set and a test set. 1400 256×256 pixel two-dimensional RGB three-channel images plus a single-channel depth map are selected as the training set. The designed depth estimation range is 1m to 5m. The given depth range is divided into 15 uniform depth layers. The RGB image and depth map of the area to be tested are read in and the pixels are classified into depth layers according to the depth. The average depth of the pixel values ​​in the 15 depth layers is calculated to represent the depth of all pixels in the depth layer. The RGB images in the dataset are divided into depth regions I according to the depth layer. n , n=0,1,2...14, calculate the depth mean of the pixels in each depth layer as the feature depth of each depth layer, denoted as d n ,n=0,1,2...14.

[0034] In step S2, a PSF feature difference evaluation coefficient is used as a quantitative indicator of the degree of difference between point spread function features at different depths. For the point spread functions corresponding to N depth layers, the PSF feature difference evaluation coefficient is expressed as:

[0035]

[0036] in, represents the mean of the point spread function at different depth layers, represents the L2 norm of the difference between the nth point spread function and the mean point spread function, where N represents the number of depth layers.

[0037] The single lens surface parameters are designed with the goal of enlarging the PSF feature difference evaluation coefficient to form a single lens coding model. n Point spread function psf n Respectively and depth area I n Convolution and superposition are performed to encode the depth information in the point spread function into the RGB image to obtain simulated images with different aberration characteristics at different depth positions. The simulated image is obtained as follows:

[0038]

[0039] Here, η represents Gaussian noise.

[0040] Specifically, a single-lens imaging model is established based on geometric ray tracing. The optical system is a single-lens optical system with even-order aspheric surfaces on both sides. The parameters to be optimized are the 4th, 6th, 8th, 10th, 12th, 14th, and 16th-order even-order aspheric coefficients of the front and back surfaces, which are defined as The initial values are all 0. Calculate the coordinates of the light rays starting from the axial object point at the average depth of each layer above on the image plane after passing through the single lens, and calculate the point spread function obtained above from this. Encode the depth information in the point spread function into an RGB image to obtain a simulation image with different aberration characteristics at different depth positions.

[0041] As an implementation manner of an embodiment of the present invention, in step S3, construct a depth-associated fuzzy feature extraction model, perform Canny edge detection on the simulation images with different aberration characteristics at different depth positions obtained in step S2 to extract the object edge information, obtain a three-channel edge feature image, splice it with the simulation image as the input of the model, extract the aberration characteristics in the simulation image through the model, and the model outputs a four-channel matrix. The first channel represents the predicted depth image of the input scene, and the last three channels represent the restored clear image after aberration elimination.

[0042] Specifically, use a convolutional neural network to extract the aberration characteristics in the simulation image obtained in step S1, decode the included depth information, splice the simulation image and the edge image obtained by the optical system to form a six-channel matrix, and convert it into a 32-channel feature map through a 1×1 convolution. The network sets four consecutive downsamplings and upsamplings, and uses max pooling and bilinear interpolation to perform downsampling and upsampling. Learn features through convolution at each scale, and the number of channels of the features is set to 32, 64, 64, 128, 256 respectively. The network output is the predicted depth image and the restored image. The network structure used in the embodiment is as Figure 2 shown. The network outputs a four-channel matrix. The first channel represents the predicted depth image of the input scene, and the last three channels represent the restored clear image after aberration elimination.

[0043] In the training of the depth-associated fuzzy feature extraction model, a system-level loss function is composed of the L1 loss between the predicted depth map and the original depth map, the L1 loss between the restored map and the original RGB image, and the reciprocal of the PSF depth feature difference coefficient. Use the gradient descent method to optimize the single-lens parameters to increase the difference in aberration characteristics of imaging at different depths, and at the same time train the model parameters that can identify the aberration characteristics at different depths.

[0044] The system-level loss function can be expressed as:

[0045]

[0046] In the formula, respectively represent the L1 loss between the predicted depth map and the original depth map, the L1 loss between the original RGB image and the restored map, and the regularization weight of the reciprocal of the PSF feature difference coefficient. In the embodiment, take Take 1, 0.5, 1 respectively.

[0047] Steps S1 to S3 implement the design of a single-lens computational imaging system for depth estimation. Based on the design of the single-lens computational imaging system, the single-lens encoding model is used to image the depth scene to be measured, obtaining blurred images with different aberration characteristics at different depths. Then, through processing by the depth-correlated blurred feature extraction model that matches, a scene predicted depth map and a clear restored image are obtained.

[0048] Select 200 RGB images with a size of 576×416 pixels and real depth maps as the test set and input them into the single-lens computational imaging system. Figure 3 It is a comparison diagram among the RGB image, the depth image, the simulated image obtained by the single-lens encoding model, the clear restored image obtained by the computational imaging system, and the predicted depth map in the dataset.

[0049] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A single-lens computational imaging system design method for depth estimation, characterized in that: include: Read the RGB image and depth map in the data set, divide the RGB image into different depth regions according to the depth map, and calculate the adaptive feature depth of each depth region; Taking the magnified PSF feature difference evaluation coefficient as the target, an initial single-lens coding model is constructed, and the depth information is encoded into the RGB image through the point spread function at the adaptive feature depth of each depth region to obtain simulated images with different aberration characteristics at different depth positions; Constructing a depth-correlated fuzzy feature extraction model to obtain a predicted depth image and a restored clear image, and optimizing the parameters of a single-lens computational imaging system based on the initial single-lens encoding model and the depth-correlated fuzzy feature extraction model; A method for constructing a depth-correlated fuzzy feature extraction model to obtain a predicted depth image and a restored clear image, and optimizing the parameters of a single-lens computational imaging system based on the initial single-lens encoding model and the depth-correlated fuzzy feature extraction model comprises: Based on the simulated image, Canny edge detection is performed to extract object edge information to obtain a three-channel edge feature image, the three-channel edge feature image and the simulated image are spliced ​​and input into the depth-related fuzzy feature extraction model, and the aberration features in the spliced ​​three-channel edge feature image and the simulated image are extracted by the depth-related fuzzy feature extraction model. The depth-related fuzzy feature extraction model outputs a four-channel matrix, the first channel represents the predicted depth image of the input scene, and the last three channels represent the restored clear image after eliminating the aberration; Based on the initial single-lens coding model and the depth-correlated fuzzy feature extraction model, the parameters of the single-lens computational imaging system are optimized to obtain a single-lens coding model that maximizes the PSF feature difference evaluation coefficient and a depth-correlated fuzzy feature extraction model that minimizes the error between the predicted depth image and the restored clear image, thereby forming a single-lens computational imaging system for depth estimation.

2. The single-lens computational imaging system design method for depth estimation according to claim 1, characterized in that: The method of dividing the RGB image into different depth regions by using the depth map and calculating the adaptive feature depth of each of the depth regions includes: Divide the given depth estimation range evenly into several depth layers; Determine the depth layer corresponding to each pixel according to the depth map, and divide the corresponding RGB image into different depth regions; The depth averages of different pixels in different depth regions are calculated respectively as the adaptive feature depths of each depth region.

3. The single-lens computational imaging system design method for depth estimation according to claim 1, characterized in that: The method for obtaining the PSF feature difference evaluation coefficient is: in, represents the mean of the point spread function at different depth layers, represents the L2 norm of the difference between the nth point spread function and the mean point spread function, where N represents the number of depth layers.

Citation Information

Patent Citations

  • Image acquiring method for removing image blurring and image acquiring system

    CN102436639A

  • Method for designing a passive single-channel imager capable of estimating depth of field

    CN104813217A