A light field depth estimation method based on light field sequence feature analysis

By constructing the light field neural network model LFRNN, combined with local depth estimation and global optimization modules, the existing light field depth estimation methods are solved in terms of accuracy and robustness, achieving higher quality depth estimation, and supporting three-dimensional reconstruction and defect detection of complex scenarios.

CN115272435BActive Publication Date: 2025-05-06NANJING INST OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210721840.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-05-06
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

The existing light field depth estimation methods have bottlenecks in terms of accuracy and robustness, especially in occlusion and noise environments, which are difficult to effectively estimate depth.

Method used

Using a method based on light field sequence feature analysis, the optical field neural network model LFRNN is constructed, and the local depth estimation module and the global optimization module are combined to perform end-to-end depth estimation. The model includes a sliding window processing layer, a sequence feature extraction subnet and a feature map deformation layer, which uses a recurrent neural network to extract deep features, and performs deep optimization through a conditional random field model.

Benefits of technology

It significantly improves the accuracy and robustness of light field depth estimation, can handle complex scenarios such as occlusion and noise more effectively, provides higher-quality depth information, and supports applications such as three-dimensional reconstruction and defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115272435B_ABST
    Figure CN115272435B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for light field depth estimation based on light field sequence feature analysis, extracts a central sub-aperture image from 4D light field data represented by a dual plane, and calculates and generates an EPI synthetic image; designs an LFRNN network with central sub-aperture image and EPI synthetic image as input and disparity map as output, the network includes a local depth estimation module based on light field sequence analysis and a global depth optimization module based on a conditional random field model; trains and evaluates the LFRNN network in two stages: local depth estimation and global optimization; tests and uses the LFRNN network to evaluate network performance. The present invention takes a different approach to analyzing light fields from the perspective of sequence data, designs a deep feature extraction sub-network based on a recurrent neural network, and significantly improves the local depth estimation capability; models global depth information, and designs an end-to-end optimization network, which significantly improves the accuracy and robustness of depth estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and artificial intelligence, and in particular relates to a light field depth estimation method based on light field sequence feature analysis. Background Art

[0002] Microlens light field cameras have entered the field of consumer electronics and have great industrial applications and academic research value. Microlens light field cameras provide a new way to solve the problem of depth estimation. On the one hand, light field imaging can not only record the position of light, but also the direction of light, providing a geometric basis for depth estimation; on the other hand, microlens light field cameras have the ability to collect monocular multi-view images, which facilitates the deployment of visual systems and lays a physical foundation for expanding the application of depth estimation.

[0003] Depth estimation based on microlens light field cameras has been a hot research topic in the past decade, and can be roughly divided into two categories: traditional depth estimation and learning-based depth estimation. Traditional depth estimation methods are mainly based on extracting features based on structure operators and inverting depth information based on light field imaging geometry. For example, Tao et al. used variance and mean operators to measure and crop EPI, fused defocus and corresponding two depth cues, and then estimated the scene depth. Wanner et al. used 2D structure tensors to estimate the slope of the line on the EPI image to obtain depth information. Zhang et al. introduced parallelogram operators to locate the position of the line on the EPI image for depth estimation. The Chinese invention patent "A depth estimation based on light field (ZL201510040975.1)" also uses structure tensors as a method for initial depth estimation. Traditional depth estimation methods based on structure operators have strong interpretability, but the operator description ability is limited, and there is a bottleneck in improving the accuracy of depth estimation.

[0004] In recent years, with the rise of deep learning technology, learning-based light field depth estimation methods have gained favor. Heber et al. first proposed a light field depth estimation method based on deep learning, using a convolutional neural network to extract the features of the EPI block, and then regressing to obtain the corresponding depth value. Shin et al. performed convolution processing on the input stream of EPI images in multiple directions and fused them to obtain depth values. Han et al. proposed a method for generating EPI synthetic images, and then used multi-stream convolution and jump-layer fusion to estimate the scene depth. Tsai et al. introduced an attention mechanism to select a more effective sub-aperture image, and then used convolution to extract features and obtain the scene depth. The Chinese invention patent "A method for light field image depth estimation based on a hybrid convolutional neural network (ZL201711337965.X)" discloses a light field image depth estimation method, which uses a convolutional neural network to extract information from horizontal and vertical EPI blocks, and then regresses and fuses them to obtain a depth map. The Chinese invention patent application "A method for light field depth estimation based on multimodal information (publication number: CN 112767466 A)" uses convolution and dilated convolution to analyze and process the focal stack and center view to predict the scene depth.

[0005] Theoretical modeling of depth estimation, light field data extraction methods, neural network design, etc. affect the depth estimation effect. At present, learning-based methods have become the mainstream of light field depth estimation methods and have made great progress; however, the accuracy of depth estimation and its robustness in terms of occlusion, noise, etc. need to be improved, especially technical links such as light field data extraction and neural network data processing, which urgently need innovation and breakthroughs. To this end, the present invention discloses a light field image processing method based on vector sequence analysis, and designs an end-to-end depth estimation network that integrates local depth estimation and global optimization. The network is used to perform light field depth estimation, and the accuracy is significantly improved, which provides good support for applications such as three-dimensional reconstruction and three-dimensional defect detection. Summary of the invention

[0006] Purpose of the invention: The present invention provides a light field depth estimation method based on light field sequence feature analysis, which can estimate high-accuracy depth results and support applications such as light field three-dimensional reconstruction and defect detection.

[0007] Technical solution: The light field depth estimation method based on light field sequence feature analysis described in the present invention specifically includes the following steps:

[0008] (1) Extracting the central sub-aperture image from 4D light field data Where (i C ,j C ) represents the viewing angle coordinates of the central sub-aperture image;

[0009] (2) Generate EPI synthetic image I from 4D light field dataSEPI ;

[0010] (3) Construct the light field neural network model LFRNN, receive I SEPI , Input, output and central subaperture image A disparity map D of the same resolution; the light field neural network model LFRNN includes a local depth estimation module based on light field sequence analysis and a depth optimization module based on a conditional random field model;

[0011] (4) Training the light field neural network model LFRNN constructed in step (3) to obtain the optimal parameter set P of the network: the training is divided into two stages, and the mean absolute error is used as the loss function in both stages; in the first stage, only the local depth estimation module based on light field sequence analysis is trained to obtain the optimal parameter set P1 of the module; in the second stage, the optimal parameter set P1 of the local depth estimation module based on light field sequence analysis is frozen, and the entire network is trained to update the parameters of the depth optimization module based on the conditional random field model to obtain the optimal parameter set P of the LFRNN network.

[0012] Furthermore, the implementation process of step (1) is as follows:

[0013] 4D light field data is a decoded representation of the light field image collected by the light field camera, denoted as L:(i,j,k,l)→L(i,j,k,l), where (i,j) represents the pixel index coordinates or viewing angle coordinates of the microlens image, (k,l) represents the index coordinates of the microlens center, i,j,k,l are all integers, and L(i,j,k,l) ​​represents the radiation intensity of the light passing through the position (k,l) under the viewing angle (i,j); the central pixel of each microlens image is extracted and arranged according to the microlens position index to obtain a two-dimensional image, that is, Where (i C ,j C ) represents the viewing angle coordinates of the central sub-aperture image.

[0014] Furthermore, the implementation process of step (2) is as follows:

[0015] (21) Initialize I according to the dimension of the input 4D light field SEPI A matrix of all zeros:

[0016] In the 4D light field L:(i,j,k,l)→L(i,j,k,l), the angular resolution is N Ai ×N Aj , that is, i∈[0,N Ai ), j∈[0,N Aj ); the spatial resolution is N Sk ×N Sl , that is, k∈[0,N Sk), l∈[0,N Sl ); then I SEPI Yes (N Sk ×N Aj )×N Sl The two-dimensional matrix is ​​initialized to a matrix of all zeros;

[0017] (22) For each row of the third dimension k of the 4D light field, the row number is k * , calculate its corresponding EPI image and use Update I SEPI Some areas:

[0018] The third dimension k is calculated from the 4D light field data * The process of mapping the rows of the EPI image to each other can be considered as a mapping: That is, fix the first and third dimensions in the 4D light field, change the other two dimensions to obtain the two-dimensional slice image, let i = i C , k = k*;

[0019] Use the proceeds Update I SEPI part of the area, namely Here, I SEPI ((k*-1)×N Aj :k*×N Aj ,0:N Sl ) indicates I SEPI (k*-1)×N Aj Go to the k*×Nth Aj -1 row, column 0 to column N Sl - An area of ​​1 column;

[0020] (23) Perform step (22) on each row of the third dimension of the 4D light field to calculate and generate the EPI composite image I SEPI .

[0021] Furthermore, the local depth estimation module based on light field sequence analysis in step (3) actually includes a sliding window processing layer, a sequence feature extraction subnetwork, and a feature map deformation layer;

[0022] The sliding window processing layer is responsible for synthesizing the image I in EPI SEPI Swipe up to capture EPI block I EPI-p , input to the sequence feature extraction sub-network; the sliding window size is (N Aj ,16), the horizontal sliding step is 1, and the vertical sliding step is N Aj , Sliding Window Beyond I SEPI When , fill with 0;

[0023] The sequence feature extraction subnetwork is used to extract EPI block I EPI-p The proposed recurrent neural network is based on the sequence features of the EPI image, including serialization splitting processing, bidirectional GRU layer and fully connected network; the serialization splitting processing is based on the unique observation that the straight lines containing depth information on the EPI image are distributed in multiple columns of pixels. Aj ×16 EIP image block I EPI-p Each column of pixels is considered as a column vector Among them, x and y represent EPI image block I respectively. EPI-p The row and column coordinates of the pixel above, Represents EPI image block I EPI-p The gray value of the pixel at (x, y); an N Aj ×16 EPI image block I EPI-p Can be serialized into 16 column vectors G y , 0≤y≤15 and y is an integer; vector G y The bidirectional GRU layer is composed of GRU units in two directions, and the dimension of each GRU unit in each direction is 256. Each GRU unit is set to non-sequential working mode, receiving vector inputs at 16 moments and generating 1 output. The bidirectional GRU layer generates a total of 512 outputs. The fully connected network contains two fully connected layers. The first fully connected layer receives 512 outputs of the bidirectional GRU layer and generates 16 outputs. The fully connected layer is configured with the ReLU activation function. The second fully connected layer receives 16 outputs of the previous fully connected layer and outputs 1 disparity value. The fully connected layer is not configured with an activation function.

[0024] The feature map deformation layer will (N Sk ×N Sl ) disparity value sequences, transformed into N Sk ×N Sl The matrix is ​​called the feature map and is denoted by U.

[0025] Furthermore, the depth optimization module based on the conditional random field model in step (3) includes two parts: central sub-aperture image kernel parameter extraction and feature map iterative optimization; the central sub-aperture image kernel parameter extraction part calculates the filter kernel parameters according to the input central sub-aperture image; the feature map iterative optimization part is based on the conditional random field theory, and according to the filter kernel parameters obtained by the central sub-aperture image kernel parameter extraction part, the feature map is iteratively optimized to obtain the disparity map D;

[0026] The central sub-aperture image kernel parameter extraction part is based on the central sub-aperture image For input, calculate the spatial and color convolution kernel F1 and the spatial convolution kernel F2:

[0027]

[0028]

[0029] Among them, p i 、p j Represent the central sub-aperture images The position information of the i-th and j-th pixels on the image, c i 、c j Represent the central sub-aperture images The color information of the i-th and j-th pixels, θ α ,θ β ,θ γ It is a custom bandwidth radius;

[0030] The feature map iterative optimization part includes four modules: parallel filtering, unary term superposition, normalization factor calculation, and normalization. The parallel filtering module uses two channels to respectively input μ t-1 Filtering: The first channel uses the convolution kernel F1 to filter μ t-1 Filtering, that is Then, the filtering result Multiply by the weight parameter θ1, that is The second path uses the convolution kernel F2 to t-1 Filtering, that is Then, the filtering result Multiply by the weight parameter θ2, that is In the first iteration, μ t-1 Initialize to feature map U; θ1 and θ2 are randomly initialized and updated through network training; the results of the two paths Add element by element to get the output of the parallel filtering module Right now The unary addition module is the result of combining the feature map U with the parallel filtering module Superposition, we get Right now The normalization factor calculation module also performs parallel filtering and unary term addition operations to obtain the normalization factor γ; the data processing object is the all-1 matrix J, not μ t-1 and feature map U; the specific processing steps of the normalization factor calculation module are: The normalized module is the result of the module to which the unary term is added. Divide the normalization factor γ element by element to get the output μ of this iteration t ,Right now The output of the last iteration is the optimized disparity map D.

[0031] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: 1. Based on the geometry of light field imaging, the present invention analyzes the light field from the perspective of sequence data in an alternative way, and designs a deep feature extraction subnetwork based on recurrent neural networks, which replaces the feature extraction method of traditional convolutional neural networks and significantly improves the local depth estimation capability; 2. Based on the conditional random field theory, the present invention models the global depth information and designs an end-to-end optimization network, which significantly improves the accuracy and robustness of depth estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a flow chart of the present invention;

[0033] Figure 2 It is a dual-plane representation diagram of the 4D light field in the present invention;

[0034] Figure 3 A flow chart for generating an EPI composite image for calculation in the present invention;

[0035] Figure 4 It is an example of the central sub-aperture image and the EPI image in the present invention;

[0036] Figure 5 Schematic diagram of the EPI composite image in the present invention;

[0037] Figure 6 LFRNN network architecture diagram designed for the present invention;

[0038] Figure 7 The structural diagram of the sequence feature extraction subnetwork designed for the present invention;

[0039] Figure 8 A structural diagram of a deep optimization module based on a conditional random field model designed for the present invention;

[0040] Fig. 9 The figure is a schematic diagram for comparing the results of the present invention with those of the prior art method. DETAILED DESCRIPTION

[0041] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0042] like Figure 1 As shown, the present invention discloses a method for estimating light field depth based on light field sequence feature analysis, comprising the following steps:

[0043] Step 1: Extract the central sub-aperture image from the 4D light field data Where (i C ,j C ) represents the viewing angle coordinates of the central sub-aperture image.

[0044] 4D light field data is a decoded representation of the light field image captured by the light field camera, such as Figure 2 As shown, the dual plane method (2PP) is usually used to represent light field data. In the figure, Π and Ω are parallel planes, representing the viewing angle plane and the position plane respectively. A ray of light is represented by the intersection of the ray and the dual plane, and all rays of light that are not parallel to the dual plane form a light field. The light field is usually recorded as L:(i,j,k,l)→L(i,j,k,l), where (i,j) represents the pixel index coordinates or viewing angle coordinates of the microlens image, (k,l) represents the index coordinates of the microlens center, i,j,k,l are all integers, and L(i,j,k,l) ​​represents the radiation intensity of the light passing through the position (k,l) at the viewing angle (i,j). The method of extracting the central sub-aperture image is to extract the central pixel of each microlens image, and arrange them according to the microlens position index to obtain a two-dimensional image, that is, Where (i C ,j C ) represents the viewing angle coordinates of the central sub-aperture image.

[0045] Step 2: Generate EPI synthetic image I by calculating 4D light field data SEPI ;like Figure 3 As shown, the following steps are included:

[0046] (2.1) Initialize I according to the dimension of the input 4D light field SEPI is a matrix of all 0s.

[0047] In the 4D light field L:(i,j,k,l)→L(i,j,k,l), if we assume that the angular resolution is N Ai ×N Aj , that is, i∈[0,N Ai ), j∈[0,N Aj ); the spatial resolution is N Sk ×N Sl , that is, k∈[0,N Sk ), l∈[0,N Sl ). Then I SEPI Yes (N Sk ×N Aj )×N Sl The two-dimensional matrix is ​​initialized to a matrix of all zeros.

[0048] (2.2) For each row of the third dimension (k) of the 4D light field (row number: k * ), calculate the corresponding EPI image and use Update I SEPI parts of the area.

[0049] The third dimension k is calculated from the 4D light field data *The process of converting rows into EPI images can be viewed as a mapping: That is, fix the first and third dimensions in the 4D light field, change the other two dimensions to obtain the two-dimensional slice image, let i = i C , k=k*. Figure 4 As shown, the upper image is a central sub-aperture image generated by a scene light field data, and the lower image is an EPI image corresponding to the row where the solid line of the central sub-aperture image is located.

[0050] Then, use the resulting Update I SEPI part of the area, namely Here, I SEPI ((k*-1)×N Aj :k*×N Aj ,0:N Sl ) indicates I SEPI (k*-1)×N Aj Go to the k*×Nth Aj -1 row, column 0 to column N Sl -1 column area.

[0051] (2.3) Perform the operation of step (2.2) on each row of the third dimension of the 4D light field to calculate and generate the EPI composite image I SEPI To demonstrate the effect, Figure 5 Take an area in the EPI composite image as an example. This area is Figure 4 The EPI composite image corresponding to the 14 rows of pixels above and below the solid line on the central sub-aperture image.

[0052] Step 3: Construct the light field neural network model LFRNN and receive I SEPI , Input, output and central subaperture image The disparity map D of the same resolution. Figure 6 As shown, the light field neural network model LFRNN includes a local depth estimation module based on light field sequence analysis and a depth optimization module based on a conditional random field model.

[0053] The local depth estimation module based on light field sequence analysis includes a sliding window processing layer, a sequence feature extraction sub-network, and a feature map deformation layer. Among them, the sliding window processing layer is responsible for synthesizing the EPI image I SEPI Swipe up to capture EPI block I EPI-p , input to the sequence feature extraction sub-network. The sliding window size is (N Aj ,16), the horizontal sliding step is 1, and the vertical sliding step is N Aj , Sliding Window Beyond I SEPI When , fill with 0.

[0054] The sequence feature extraction subnetwork is used to extract EPI block I EPI-p The recurrent neural network specially designed for the sequence features of Figure 7 The serialization splitting process is based on the unique observation that the straight lines containing depth information in the EPI image are distributed in multiple columns of pixels, and the proposed EPI image block serialization mechanism is shown in Figure 1. Aj ×16 EIP image block I EPI-p Each column of pixels is considered as a column vector Among them, x and y represent EPI image block I respectively. EPI-p The row and column coordinates of the pixel above, Represents EPI image block I EPI-p The grayscale value of the pixel at (x, y) on the Aj ×16 EPI image block I EPI-p Can be serialized into 16 column vectors G y , 0≤y≤15 and y is an integer. These vectors will be used as inputs to the subsequent bidirectional GRU layer at each moment.

[0055] The bidirectional GRU layer is composed of GRU units in two directions. The dimension of each directional GRU unit is 256. Each GRU unit is set to non-sequential working mode, that is, it receives vector inputs at 16 moments and generates 1 output. The bidirectional GRU layer generates a total of 512 outputs.

[0056] The next fully connected network contains two fully connected layers. The first fully connected layer receives 512 outputs from the bidirectional GRU layer and produces 16 outputs; this layer is fully connected and configured with the ReLU activation function. The second fully connected layer receives 16 outputs from the previous fully connected layer and outputs 1 disparity value; this fully connected layer is not configured with an activation function.

[0057] The task of the feature map deformation layer is to transform (N Sk ×N Sl ) disparity value sequences, transformed into N Sk ×N Sl The matrix is ​​called a feature map, denoted as U. The previous sliding window processing layer synthesizes the image I in the EPI according to the set step size. SEPI Slide up to capture (N Sk ×N Sl ) EPI Block I EPI-p , each EPI block I EPI-p After the sequence feature extraction sub-network processes 1 disparity value, all EPI blocks generate a total of (N Sk ×N Sl) disparity values, the feature map deformation layer calls Reshape to deform it into N Sk ×N Sl Matrix, denoted as U.

[0058] The deep optimization module based on the conditional random field model includes two parts: central sub-aperture image kernel parameter extraction and feature map iterative optimization. Figure 8 The main function of the central sub-aperture image kernel parameter extraction part is to calculate the filter kernel parameters according to the input central sub-aperture image; the feature map iterative optimization part is based on the conditional random field theory, and according to the filter kernel parameters obtained by the central sub-aperture image kernel parameter extraction part, the feature map is iteratively optimized to obtain the disparity map D.

[0059] The central sub-aperture image kernel parameter extraction part is based on the central sub-aperture image As input, calculate the parameters of two globally connected convolution kernels: 1) Calculate the spatial / color convolution kernel F1, the calculation method is Among them, p i 、p j Represent the central sub-aperture images The position information of the i-th and j-th pixels on the image, c i 、c j Represent the central sub-aperture images The color information of the i-th and j-th pixels, θ α ,θ β is the custom bandwidth radius (here, it is set to 1). 2) Calculate the spatial convolution kernel F2, the calculation method is Likewise, p i 、p j Represent the central sub-aperture images The position information of the i-th and j-th pixels on the γ is the custom bandwidth radius (set here as ).

[0060] The iterative optimization part of the feature map includes four modules: parallel filtering, unary item addition, normalization factor calculation, and normalization.

[0061] The parallel filtering module uses two paths to filter the input μ t-1 Filtering: The first channel uses the convolution kernel F1 to filter μ t-1 Filtering, that is Then, the filtering result Multiply by the weight parameter θ1, that is Similarly, the second path uses the convolution kernel F2 to t-1 Filtering, that is Then, the filtering result Multiply by the weight parameter θ2, that is In the first iteration, μ t-1 Initialize to feature map U; θ1 and θ2 are randomly initialized and updated through network training. Results of two paths Add element by element to get the output of the parallel filtering module Right now

[0062] The unary addition module is the result of combining the feature map U with the parallel filtering module Superposition, we get Right now

[0063] The normalization factor calculation module also performs parallel filtering and unary term addition operations to obtain the normalization factor γ; the difference is that the data processing object is the all-1 matrix J instead of μ t-1 And feature map U. The specific processing steps of the normalization factor calculation module are:

[0064] The normalized module is the result of the module to which the unary term is added. Divide the normalization factor γ element by element to get the output μ of this iteration t ,Right now

[0065] The feature map iterative optimization part is an iterative process consisting of four modules. Usually, 6 iterations can achieve the ideal optimization effect. The output of the last iteration is the optimized disparity map D.

[0066] Step 4, training the LFRNN described in step 3 to obtain the optimal parameter set P of the network. The feature is that the training step is divided into two stages, and both stages use the mean absolute error as the loss function. In the first stage, only the local depth estimation module based on light field sequence analysis is trained to obtain the optimal parameter set P1 of the module; in the second stage, the optimal parameter set P1 of the local depth estimation module based on light field sequence analysis is frozen, and the entire network is trained to update the parameters of the depth optimization module based on the conditional random field model, and finally the optimal parameter set P of the LFRNN network is obtained.

[0067] Training the LFRNN network includes the following steps:

[0068] (4.1) Prepare a light field dataset and divide it into a training set, a validation set, and a test set. The light field dataset must contain scene light field data and scene parallax true values. Specifically, the currently available HCI light field dataset can be used, or the light field data can be synthesized through Blender simulation software. The light field data and depth true values ​​can also be collected through a light field camera and a rangefinder. The light field dataset is randomly divided into a training set, a validation set, and a test set in a ratio of 5:3:2.

[0069] (4.2) Prepare the input data and true value data required for network training. The input data includes the central subaperture image and the EPI synthetic image, which are calculated from the light field dataset according to steps 1 and 2 respectively; the true value data is directly provided by the light field dataset.

[0070] (4.3) The local depth estimation module based on light field sequence analysis is trained and verified as an independent network. First, the input is the EPI synthetic image, the output feature map is used as the estimated disparity value, and the true value data provided by the dataset is used as the true disparity value. The mean absolute error is calculated, and the network parameters are optimized by back propagation. After training, the optimal parameter set P1 of the module is obtained. Among them, the hyperparameter batch is set to 64, and the hyperparameter epoch is set to 10000; the learning rate for the first 2000 epochs is 0.1×10 -3 , the learning rate for the next 8000 epochs is 0.1×10 -4 . Secondly, the generalization ability of the network module is verified on the validation set.

[0071] (4.4) Train and verify LFRNN to obtain the optimal parameter set P. First, the local depth estimation module based on light field sequence analysis is used as a pre-trained network, its parameter set P1 is loaded, and the parameter update of the module is frozen; then, the EPI synthetic image and the central sub-aperture image are input, the estimated disparity value is output, the mean absolute error is calculated with reference to the true disparity value, and the parameters of the depth optimization module based on the conditional random field model in the LFRNN network are optimized by back propagation, and finally the optimal parameter set P of LFRNN is obtained. Among them, the hyperparameter batch is set to 64, the hyperparameter epoch is set to 3000, and the learning rate is set to 0.1×10 -4 Finally, the generalization ability of the entire network is tested on the validation set.

[0072] Testing and practical application of LFRNN network. For the test set described in step 4 or the 4D light field data collected by the light field camera, the central sub-aperture image can be obtained by processing according to step 1, and the EPI composite image can be obtained by processing according to step 2; then, the obtained central sub-aperture image and EPI composite image are input into the LFRNN network described in step 3; then, the optimal parameter set P described in step 4 is loaded, and forward calculation is performed to obtain the disparity map D.

[0073] Fig. 9 An example of performance comparison between the method in this paper and other neural network-based depth estimation methods is given. Taking 4 typical scenes as examples, mainstream light field depth estimation methods such as EPINet, FusionNet and VommaNet are compared. The first column is the central sub-aperture image of the scene, and the second to fifth columns are the processing results of the method disclosed in this invention, EPINet, FusionNet and VommaNet respectively; the processing results of the same scene are arranged in the same row; the comparative evaluation indicator is the mean square error (MSE), and the number above each processing result image represents the MSE value obtained by the corresponding method on the scene; a grayscale ruler is attached to each row, indicating the error distribution of the processing result at each pixel position. The lighter the color, the smaller the error, and the darker the color, the larger the error. Fig. 9 It can be seen that the LFRNN depth estimation method disclosed in the present invention achieved the best MSE index in the first two example scenes. In the last two example scenes, although the overall MSE index is not as good as the VammaNet method, the depth estimation results of most pixels are closer to the true value, and the visual effect is significantly better than the results of VammaNet.

Claims

1. A light field depth estimation method based on light field sequence feature analysis, characterized in that: The following steps are involved: (1) Extracting the central sub-aperture image from 4D light field data Where (i C ,j C ) represents the viewing angle coordinates of the central sub-aperture image; (2) Generate EPI synthetic image I from 4D light field data SEPI ; (3) Construct the light field neural network model LFRNN, receive I SEPI , Input, output and central subaperture image A disparity map D of the same resolution; the light field neural network model LFRNN includes a local depth estimation module based on light field sequence analysis and a depth optimization module based on a conditional random field model; (4) Training the light field neural network model LFRNN constructed in step (3) to obtain the optimal parameter set P of the network: the training is divided into two stages, and the mean absolute error is used as the loss function in both stages; in the first stage, only the local depth estimation module based on light field sequence analysis is trained to obtain the optimal parameter set P1 of the module; in the second stage, the optimal parameter set P1 of the local depth estimation module based on light field sequence analysis is frozen, and the entire network is trained to update the parameters of the depth optimization module based on the conditional random field model to obtain the optimal parameter set P of the LFRNN network; The implementation process of step (2) is as follows: (21) Initialize I according to the dimension of the input 4D light field SEPI A matrix of all zeros: In the 4D light field L:(i,j,k,l)→L(i,j,k,l), the angular resolution is N Ai ×N Aj , that is, i∈[0,N Ai ), j∈[0,N Aj ); the spatial resolution is N Sk ×N Sl , that is, k∈[0,N Sk ), l∈[0,N Sl ); then I SEPI Yes (N Sk ×N Aj )×N Sl The two-dimensional matrix is ​​initialized to a matrix of all zeros; (22) For each row of the third dimension k of the 4D light field, the row number is k * , calculate its corresponding EPI image and use Update I SEPI Some areas: The third dimension k is calculated from the 4D light field data * The process of mapping the rows of the EPI image to each other can be considered as a mapping: That is, fix the first and third dimensions in the 4D light field, change the other two dimensions to obtain the two-dimensional slice image, let i = i C , k = k*; Use the proceeds Update I SEPI part of the area, namely Here, I SEPI ((k*-1)×N Aj :k*×N Aj ,0:N Sl ) indicates I SEPI (k*-1)×N Aj Go to k*×N Aj -1 row, column 0 to column N Sl - An area of ​​1 column; (23) Perform step (22) on each row of the third dimension of the 4D light field to calculate and generate the EPI composite image I SEPI .

2. The method for light field depth estimation based on light field sequence feature analysis according to claim 1, characterized in that: The implementation process of step (1) is as follows: 4D light field data is a decoded representation of the light field image collected by the light field camera, denoted as L:(i,j,k,l)→L(i,j,k,l), where (i,j) represents the pixel index coordinates or viewing angle coordinates of the microlens image, (k,l) represents the index coordinates of the microlens center, i,j,k,l are all integers, and L(i,j,k,l) ​​represents the radiation intensity of the light passing through the position (k,l) under the viewing angle (i,j); the central pixel of each microlens image is extracted and arranged according to the microlens position index to obtain a two-dimensional image, that is, Where (i C ,j C ) represents the viewing angle coordinates of the central sub-aperture image.

3. The method for light field depth estimation based on light field sequence feature analysis according to claim 1, characterized in that: The local depth estimation module based on light field sequence analysis in step (3) actually includes a sliding window processing layer, a sequence feature extraction subnetwork, and a feature map deformation layer; The sliding window processing layer is responsible for synthesizing the image I in EPI SEPI Swipe up to capture EPI block I EPI-p , input to the sequence feature extraction sub-network; the sliding window size is (N Aj ,16), the horizontal sliding step is 1, and the vertical sliding step is N Aj , Sliding Window Beyond I SEPI When , fill with 0; The sequence feature extraction subnetwork is used to extract EPI block I EPI-p The proposed recurrent neural network is based on the sequence features of the EPI image, including serialization splitting processing, bidirectional GRU layer and fully connected network; the serialization splitting processing is based on the unique observation that the straight lines containing depth information on the EPI image are distributed in multiple columns of pixels. Aj ×16 EIP image block I EPI-p Each column of pixels is considered as a column vector Among them, x and y represent EPI image block I respectively. EPI-p The row and column coordinates of the pixel above, Represents EPI image block I EPI-p The gray value of the pixel at (x, y); an N Aj ×16 EPI image block I EPI-p Can be serialized into 16 column vectors G y , 0≤y≤15 and y is an integer; vector G y The bidirectional GRU layer is composed of GRU units in two directions, and the dimension of each GRU unit in each direction is 256. Each GRU unit is set to non-sequential working mode, receiving vector inputs at 16 moments and generating 1 output. The bidirectional GRU layer generates a total of 512 outputs. The fully connected network contains two fully connected layers. The first fully connected layer receives 512 outputs of the bidirectional GRU layer and generates 16 outputs. The fully connected layer is configured with the ReLU activation function. The second fully connected layer receives 16 outputs of the previous fully connected layer and outputs 1 disparity value. The fully connected layer is not configured with an activation function. The feature map deformation layer will (N Sk ×N Sl ) disparity value sequences, transformed into N Sk ×N Sl The matrix is ​​called the feature map and is denoted by U.

4. The method for light field depth estimation based on light field sequence feature analysis according to claim 1, characterized in that: The depth optimization module based on the conditional random field model in step (3) includes two parts: central sub-aperture image kernel parameter extraction and feature map iterative optimization; the central sub-aperture image kernel parameter extraction part calculates the filter kernel parameters based on the input central sub-aperture image; The iterative optimization part of the feature map is based on the conditional random field theory. According to the central sub-aperture image kernel parameter extraction part of the obtained filter kernel parameters, the feature map is iteratively optimized to obtain the disparity map D; The central sub-aperture image kernel parameter extraction part is based on the central sub-aperture image For input, calculate the spatial and color convolution kernel F1 and the spatial convolution kernel F2: Among them, p i 、p j Represent the central sub-aperture images The position information of the i-th and j-th pixels on the image, c i 、c j Represent the central sub-aperture images The color information of the i-th and j-th pixels, θ α ,θ β ,θ γ It is a custom bandwidth radius; The feature map iterative optimization part includes four modules: parallel filtering, unary term superposition, normalization factor calculation, and normalization. The parallel filtering module uses two channels to respectively input μ t-1 Filtering: The first channel uses the convolution kernel F1 to filter μ t-1 Filtering, that is Then, the filtering result Multiply by the weight parameter θ1, that is The second path uses the convolution kernel F2 to t-1 Filtering, that is Then, the filtering result Multiply by the weight parameter θ2, that is In the first iteration, μ t-1 Initialize to feature map U; θ1 and θ2 are randomly initialized and updated through network training; the results of the two paths Add element by element to get the output of the parallel filtering module Right now The unary addition module is the result of combining the feature map U with the parallel filtering module Superposition, we get Right now The normalization factor calculation module also performs parallel filtering and unary term addition operations to obtain the normalization factor γ; the data processing object is the all-1 matrix J, not μ t-1 and feature map U; the specific processing steps of the normalization factor calculation module are: The normalized module is the result of the module to which the unary term is added. Divide the normalization factor γ element by element to get the output μ of this iteration t ,Right now The output of the last iteration is the optimized disparity map D.

Citation Information

Patent Citations

  • Depth estimation method based on optical field

    CN104598744A

  • Light field image depth estimation method based on hybrid convolutional neural network

    CN107993260A

  • Multi-modal information-based light field depth estimation method

    CN112767466A

  • Light field image depth estimation method based on deep convolutional neural network

    CN112116646A

  • Light field depth self-supervised learning method based on occlusion region iterative optimization

    CN112288789A