A method for estimating depth in large baseline light field video
By generating wide-pixel EPI images and combining them with a macro-pixel block segmentation method based on semantic texture fusion, the problem of depth estimation with large parallax between viewpoints in array-type light field video acquisition equipment is solved, achieving high-precision depth estimation for large baseline light field data and supporting 3D scene reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2022-09-16
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to effectively address the challenges in scene depth estimation caused by large parallax between viewpoints in array-type light field video acquisition devices, especially lacking effective solutions for depth estimation of large baseline light field data.
A large baseline light field video depth estimation method is adopted. By generating wide pixel EPI images, macro-pixel block segmentation based on semantic texture fusion and multi-scale hybrid feature index based on standard squared difference and correlation coefficient are used to perform macro-pixel block matching and depth estimation.
It improves the accuracy and precision of depth estimation in large baseline light field videos, broadens the applicability of depth estimation technology in the field of light field images, and supports subsequent 3D scene reconstruction.
Smart Images

Figure CN115496790B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of visual image processing and relates to scene depth estimation in the field of light field, specifically a large baseline light field video depth estimation method. Background Technology
[0002] With the rise of virtual reality technologies such as VR and AR, online media audiences have increasingly higher requirements for image resolution and viewing angle. Unlike traditional cameras, which can only capture two-dimensional information from a certain direction of a scene, light field cameras, due to their unique imaging characteristics, can record the intensity and direction information of three-dimensional light in a single image, thus obtaining richer scene information.
[0003] Because light field data contains rich horizontal and vertical parallax information, depth estimation of light fields has a significant advantage over traditional single-camera scene capture. Accurate scene depth estimation results also have a crucial impact on applications such as scene 3D modeling and viewpoint interpolation rendering. Existing depth estimation methods in the field of light fields have achieved good results with synthetic light field data where parallax is relatively small. However, due to limitations in current production processes, the parallax between adjacent viewpoints in light field video acquisition equipment remains large. Currently, there is a lack of effective solutions for depth estimation of such large-baseline light field data. To enable continuous 3D reconstruction of real-time acquired light field scenes, accurate depth estimation of large-parallax array-type light field data is a problem with practical application significance and urgently needs to be solved.
[0004] Light field image arrays can contain multi-view information of a scene at the same time, which makes light field depth estimation a great advantage over traditional single-camera scene capture. In recent years, many light field depth recovery methods have been proposed and are being extended to more and more scenes. Depth estimation based on light field models is mainly divided into multi-view stereo matching and methods based on the epipolar slope of EPI images.
[0005] Multi-view stereo matching depth estimation methods evolved from monocular stereo matching. Different loss functions are constructed based on the difference between the central view image and adjacent view images to obtain the matching quantity. It is also possible to construct a cost quantity based on the pixel consistency of microlens images focused at different depths, and then estimate the depth. Heber et al. [1] applied principal component analysis (PCA) to depth estimation, and aligned the sub-aperture images by projecting the image onto the column vector of the shared matrix and using low-rank structure normalization. Chen et al. [2] introduced bilateral consistency measure into light field depth estimation. This method was used to deal with the depth estimation problem in occluded scenes. Jeon et al. [3] proposed a multi-view stereo matching method based on cost function to achieve sub-pixel accuracy depth estimation.
[0006] Methods for estimating depth based on epipolar slope in EPI images mainly include methods for directly extracting EPI slope information and deep learning methods. EPI images are epipolar plane maps obtained by adjusting the dimensional order of a light field array image set. The slope of the epipolar lines is related to the depth of spatial points, and many studies on light field depth estimation are based on this characteristic of EPI images. Zhang et al. [4] designed a rotating parallelogram operator to measure the slope of the EPI epipolar line, which improved the accuracy of EPI-based depth estimation in occluded and noisy scenes; Sheng et al. [5] designed an EPI patch that combines CAD and CAE cost functions in four directions: horizontal, vertical and diagonal, and obtained a depth estimation result with high accuracy; Shin et al. [6] designed a fully convolutional neural network that uses information from the horizontal, vertical and two diagonal directions of the EPI image as the network input to estimate the depth information of the scene; Wu et al. [7] designed a convolutional neural network to reconstruct the light field with high angular resolution, thereby obtaining a more accurate EPI epipolar plane map; Leistner et al. [8] introduced the concept of EPI-Shift and designed a fully neural network structure to solve the depth estimation problem of wide baseline light field; Li et al. [9] proposed an end-to-end fully convolutional network to estimate the depth value of the intersection point on the horizontal and vertical EPI images.
[0007] The aforementioned depth estimation algorithms in the field of light fields have demonstrated high performance on ideal light field datasets. However, due to limitations in current manufacturing processes, the parallax between adjacent viewpoints in light field video acquisition equipment is large, and the factory quality of cameras also carries certain errors. As a result, there is still a lack of effective solutions for depth estimation of such large baseline light field data.
[0008] [1]Heber, Stefan, R.Ranftl, and T.Pock. "Variational Shape from LightField." Energy Minimization Methods in Computer Vision and PatternRecognition. Springer Berlin Heidelberg, 2013: 66-79.
[0009] [2]Chen C, Lin H, Zhan Y, et al.Light Field Stereo Matching Using Bilateral Statistics of Surface Cameras[C] / / Computer Vision&PatternRecognition.IEEE,2014:1518-1525.
[0010] [3]Jeon H G,Park J,Choe G,et al.Accurate depth map estimation from alenslet light field camera.IEEE,2015:1547-1555.
[0011] [4]Zhang S,Hao S,Chao L,et al.Robust depth estimation for light fieldvia spinning parallelogram operator[J].Computer Vision&Image Understanding,2016,145(apr.):148-159.
[0012] [5]Sheng H,Zhao P,Zhang S,et al.Occlusion-Aware Depth Estimation forLight Field Using Multi-Orientation EPIs[J].Pattern Recognition,2017:587-599.
[0013] [6]Shin C,Jeon H G,Yoon Y,et al.EPINET:A Fully-Convolutional NeuralNetwork Using Epipolar Geometry for Depth from Light Field Images[J].IEEE,2018:4748-4757.
[0014] [7]Wu G,Zhao M,Wang L,et al.Light field reconstruction using deepconvolutional network on EPI[C].In:IEEE Conf.Comput.Vision and PatternRecognition,2017:1638-1646.
[0015] [8] Leistner T, Schilling H, Mackowiak R, et al. Learning to Think Outside the Box: Wide-Baseline Light Field Depth Estimation with EPI-Shift[J]. 2019.
[0016] [9]Li K, Zhang J, Sun R, et al.EPI-based Oriented Relation Networks for Light Field Depth Estimation[C] / / British Machine Vision Conference(BMVC).2020. Summary of the Invention
[0017] To address the challenge of scene depth estimation when there is significant parallax between viewpoints in array-type light field video acquisition devices, this invention provides a large baseline light field video depth estimation method to achieve continuous depth estimation in wide baseline light fields, thus broadening the applicability of depth estimation technology in the field of light field imaging.
[0018] The specific steps of the large baseline light field video depth estimation method are as follows:
[0019] Step 1: Use a light field camera to acquire images, read in the original single-viewpoint video file sequentially, split the frames, and store the video frame images from different viewpoints at the same time as a group to obtain a multi-viewpoint light field sub-aperture image set.
[0020] Specifically, each group of video frame images includes N sub-aperture images, and each sub-aperture image corresponds to a viewpoint.
[0021] Step 2: For each group of images, use the row pixels and column pixels in each frame of sub-aperture image to generate the horizontal EPI image and vertical EPI image of each frame respectively;
[0022] For each sub-aperture image, starting from the first row, pixels with the same height value are extracted sequentially from left to right to form the horizontal EPI (wide pixel outer polar plane) image of that sub-aperture image; similarly, starting from the first column, pixels with the same width value are extracted sequentially from top to bottom to obtain the vertical EPI image.
[0023] Step 3: Use the center row / center column of each frame of horizontal / vertical EPI image to perform gradient map template detection and graph-based image segmentation to obtain the macro-pixel block segmentation reference results corresponding to each frame of horizontal / vertical EPI image;
[0024] First, for the current horizontal EPI image, gradient map template detection is performed by selecting the horizontal width pixels of the center row to obtain the texture segmentation result;
[0025] Then, graph-based image segmentation is performed again on the central row to obtain its semantic segmentation result;
[0026] Logical OR operation is performed on the texture and semantic segmentation results to determine the original wide pixels of the horizontal EPI image that are related, and macro pixels are defined to obtain the macro pixel block segmentation reference result in the frame.
[0027] Similarly, for the current vertical EPI image, gradient map template detection is performed by selecting the vertical width pixels of the center column to obtain the texture segmentation result;
[0028] Then, graph-based image segmentation is performed again on the central column to obtain its semantic segmentation results;
[0029] The texture and semantic segmentation results are judged by logical OR operation. The original wide pixels with correlation in the vertical EPI image are defined as macro pixels to obtain the macro pixel block segmentation reference result in the frame.
[0030] Step 4: Using the macro-pixel block segmentation reference results of each frame of horizontal / vertical EPI images, find the target pixel position in each frame of horizontal / vertical EPI images as the preliminary depth estimation result.
[0031] Specifically:
[0032] First, select each frame of sub-aperture image one by one, take the macro pixel block segmentation reference result of the current frame horizontal EPI image as image a1, and copy the remaining pixel rows in the image except the center row to form a new copied image b1.
[0033] The length of the pixel row extension in image b1 is the maximum length of the macro-pixel block segmented in image a1.
[0034] Then, each row of the copied image b1 is cyclically compared with the pixels in image a1. The standard deviation of squares (Sq_diff) and the correlation coefficient (R) are calculated in multi-scale space, and their average values are taken.
[0035] The closer the standard squared deviation is to 0, the higher the matching degree between the two pixels, and vice versa. The closer the correlation coefficient R is to 1, the more it indicates that the target regions of the two pixels are completely matched, and vice versa, it indicates that there is no correlation between the two.
[0036]
[0037]
[0038] in:
[0039] T'(x',y')=T(x',y')-1 / (w·h)·∑ x”,y” T(x”,y”) (3)
[0040] I'(x+x',y+y')=I(x+x',y+y')-1 / (w·h)·∑ x”,y” I(x+x”,y+y”) (4)
[0041] (x, y) represents the Cartesian coordinates of the top-left corner of image b1 retrieved in the current loop, indicating the position of the loop;
[0042] (x',y') represents the offset of each point in image b1 relative to the top left corner;
[0043] T(x',y') obtains the coordinates of the point in image a1, which is the coordinates of the top left corner of image a1 plus the offset;
[0044] I(x+x',y+y') obtains the coordinates of the point in image b1, which is the coordinates of the top left corner of image b1 plus the offset;
[0045] w and h are the width and height of the corresponding image, respectively;
[0046] T(x”,y”) obtains the coordinates of a point in image a1 in the complete EPI;
[0047] I(x+x”,y+y”) obtains the coordinates of a point in image b1 within the complete EPI;
[0048] Finally, sort the average values in descending order and select the pixel value corresponding to the maximum average value as the target pixel;
[0049] The slope information of the line connecting each target pixel value is calculated in the horizontal EPI image and converted to the corresponding [0,255] gray space as a preliminary depth estimation result.
[0050] Similarly, the macro-pixel block segmentation reference result of the current frame vertical EPI image is used as image a2, and the remaining pixel rows in this image except for the center column are copied to form a new copied image b2.
[0051] Then, each column of the copied image b2 is cyclically compared with the pixels in image a2. The standard deviation of squares (Sq_diff) and the correlation coefficient (R) are calculated in multi-scale space, and their average values are taken.
[0052] Finally, the average values are sorted in descending order, and the pixel value corresponding to the maximum average value is selected as the target pixel. The slope information of the line connecting each target pixel value in the vertical EPI image is calculated using the coordinates of each target pixel value, and then converted to the corresponding [0,255] gray space as the preliminary depth estimation result.
[0053] Step 5: Using the preliminary estimation results, the triangular coordinate mapping method in computer vision is used to obtain the depth information of the target pixel in the three-dimensional matching result, that is, the depth information of the original pixels contained in the horizontal / vertical EPI image.
[0054] The depth values of the target pixel positions in all horizontal / vertical EPI images are weighted and averaged, and then mapped back to the original video frames according to the coordinates to obtain the corresponding depth map output results of the video frames.
[0055] The advantages of this invention are:
[0056] (1) A large baseline light field video depth estimation method proposes the concept of wide pixel EPI image, selects continuous multiple rows or columns of pixels of each sub-aperture image to generate EPI image, and better preserves the integrity of some image information in the light field spatial domain.
[0057] (2) A large baseline light field video depth estimation method is proposed. A macro-pixel block segmentation method based on semantic texture fusion is designed. The detection and segmentation method based on gradient graph template enables objects with different texture information to be segmented. Graph-based semantic segmentation can ensure that all pixels of the same macro-pixel block belong to the same scene objects as much as possible.
[0058] (3) A large baseline light field video depth estimation method introduces a macro-pixel block search and matching method based on standard squared difference Sq_diff and correlation coefficient R multi-scale hybrid features. By selecting the matching result with the highest probability in multiple search structures as the final output matching coordinates, the coordinate matching accuracy of macro-pixel blocks is improved to a certain extent, and the accuracy of depth estimation results is also improved. Attached Figure Description
[0059] Figure 1 This is a flowchart of a large baseline light field video depth estimation method according to the present invention;
[0060] Figure 2 This is the horizontal / vertical wide pixel EPI image of the present invention;
[0061] Figure 3 This is a schematic diagram illustrating how the slope of the line connecting each target pixel value coordinate is calculated in the EPI image according to the present invention.
[0062] Figure 4 This is a comparison diagram of the present invention and the prior art. Detailed Implementation
[0063] To facilitate understanding and implementation of the present invention by those skilled in the art, the present invention will be further described in detail and in depth below with reference to the accompanying drawings.
[0064] This invention addresses the issue of significant parallax between viewpoints in array-type light field video acquisition equipment due to manufacturing limitations. It proposes a large-baseline light field video depth estimation method. First, horizontal and vertical wide-pixel EPI images are constructed from the original light field video data. Then, macro-pixel block segmentation based on semantic texture fusion is performed on the center viewpoint row / column of each EPI image to obtain the macro-pixel block segmentation results. A multi-scale hybrid feature search strategy based on standard squared difference and correlation coefficient is used to locate the position coordinates of the target macro-pixel block in each row / column of the EPI image, thereby obtaining the corresponding macro-pixel block depth information. This search and matching process is performed once for each of the horizontal and vertical wide-pixel EPI images in a sampled light field frame, and the final depth map output is obtained after weighted averaging. This achieves rapid depth estimation of large-baseline light field video frame data, laying the foundation for subsequent 3D scene reconstruction and broadening the application of depth estimation technology in the field of light field acquisition.
[0065] The large baseline light field video depth estimation method, such as Figure 1 As shown, the specific steps are as follows:
[0066] Step 1: Use a light field camera to acquire images, read in the original single-viewpoint video file sequentially, split the frames, and store the video frame images from different viewpoints at the same time as a group to obtain a multi-viewpoint light field sub-aperture image set.
[0067] Specifically, all viewpoint video files are split into frames sequentially, and the frame images at different times are stored in different folders. Frame images at the same time from different viewpoints are stored as a group in the same folder. Each group of video frame images includes N sub-aperture images, and each sub-aperture image corresponds to a viewpoint.
[0068] Step 2: For each group of images, use the row pixels and column pixels in each frame of sub-aperture image to generate the horizontal EPI image and vertical EPI image of each frame respectively;
[0069] Traditional EPI images are generated by selecting only one row or column of pixels from each sub-aperture image, which often easily destroys the integrity of texture information in the light field spatial domain. Compared with small parallax light field EPI images, the texture information contained in large parallax light field EPI images is completely disordered and distorted. Pixels generated by the same point light source in the scene are no longer spatially adjacent in the EPI image, which poses a great challenge to depth estimation algorithms. Therefore, this invention proposes the concept of wide-pixel EPI images, that is, selecting multiple consecutive rows or columns of pixels from each sub-aperture image to generate the EPI image.
[0070] like Figure 2 As shown, for each sub-aperture image, starting from the first row, pixels with the same height value are extracted sequentially from left to right to form the horizontal EPI (wide pixel outer polar plane) image of that sub-aperture image; similarly, starting from the first column, pixels with the same width value are extracted sequentially from top to bottom to obtain the vertical EPI image.
[0071] Step 3: Use the center viewpoint row / column pixels of each frame of horizontal / vertical EPI image to perform gradient map template detection and graph-based image segmentation to obtain macro-pixel block segmentation reference results;
[0072] First, for the current sub-aperture image, starting from the center row viewpoint, pixels with the same height value are extracted sequentially from left to right to obtain a horizontal EPI image with only the center row. Gradient map template detection is performed by selecting the horizontal width pixels of the center row to obtain the texture segmentation result.
[0073] The gradient information of an image reflects the rate of change of image grayscale, and can well show the differences in texture information and scene edge information in different regions. Gradient information detection is performed on the center viewpoint row of each wide pixel EPI image to obtain its gradient map; template detection is performed on the pixels of the gradient map to obtain the texture segmentation result of the wide pixel EPI image.
[0074] Specifically, a template with the same height as the center viewpoint row is used for horizontal scanning detection, with a template length of len. The center viewpoint image is segmented into different macro-pixel blocks by detecting changes in the gradient mean within the template region. The gradient mean AvgGrad at the current position of the template is compared with the gradient mean AvgGrad at the previous position. last The difference between them exceeds the preset segmentation threshold T h At that time, it is determined whether there is a significant change in the gradient between two detections, and the current position of the template is set to Loc. pat Stored as a separator between adjacent macro-pixel blocks in sequence S out and AvgGrad last Updated to AvgGrad. After traversing the center viewpoint rows S outIt is possible to divide the central viewpoint map into different macro-pixel blocks;
[0075] Then, graph-based image segmentation is performed again on the central row to obtain its semantic segmentation result;
[0076] The method of extraction from the macroscopic perspective is adopted: First, graph-based image segmentation is performed on the complete view of the center viewpoint to obtain the overall semantic segmentation result; second, the semantics of the corresponding position of the wide pixel EPI image is found by the iterative method in step two, and it is extracted as the semantic segmentation result of the center viewpoint row of the wide pixel EPI (graph-based image segmentation is performed on the center viewpoint image of the current frame to obtain the overall semantic segmentation result, and then the row (column) corresponding to the horizontal (vertical) wide pixel EPI image is extracted to obtain the corresponding local semantic segmentation result).
[0077] Finally, a logical OR operation is performed on the texture and semantic segmentation results to determine the original wide pixels of the horizontal EPI image that are related to each other, and a macro pixel is defined as a macro pixel to obtain the macro pixel block segmentation reference result in the frame.
[0078] Similarly, for the current sub-aperture image, starting from the center column from top to bottom, the same width value of pixels is extracted sequentially to obtain a vertical EPI image with only the center column. Gradient map template detection is performed on the vertical width pixels of the center column to obtain the texture segmentation result.
[0079] Then, graph-based image segmentation is performed again on the central column to obtain its semantic segmentation results;
[0080] The texture and semantic segmentation results are judged by logical OR operation. The original wide pixels with correlation in the vertical EPI image are defined as macro pixels to obtain the macro pixel block segmentation reference result in the frame.
[0081] Step 4: Using the macro-pixel block segmentation reference results of each frame of horizontal / vertical EPI images, find the target pixel position in each frame of horizontal / vertical EPI images as the preliminary depth estimation result.
[0082] Boundary missing issues directly impact pixel block matching results, easily causing the depth of all boundary macro-pixel blocks to tend towards infinity. To address this, the pixel columns at the boundaries of the wide-pixel EPI image are copied and extended by a template macro-pixel block length region, providing sufficient space for detecting pixel block displacement between viewpoints in the wide-pixel EPI image; specifically:
[0083] First, select each frame of sub-aperture image one by one, take the macro pixel block segmentation reference result of the current frame horizontal EPI image as image a1, and copy the remaining pixel rows in the image except the center row to form a new copied image b1.
[0084] The length of the pixel row extension in image b1 is the maximum length of the macro-pixel block segmented in image a1.
[0085] Then, each row of the copied image b1 is cyclically compared with the pixels in image a1. The standard deviation of squares (Sq_diff) and the correlation coefficient (R) are calculated in multi-scale space, and their average values are taken.
[0086] The closer the standard squared deviation is to 0, the higher the matching degree between the two pixels, and vice versa. The closer the correlation coefficient R is to 1, the more it indicates that the target regions of the two pixels are completely matched, and vice versa, it indicates that there is no correlation between the two.
[0087] If we define the macro-pixel target template as follows, and the EPI viewpoint pixel behavior to be searched is defined as follows:
[0088]
[0089]
[0090] in:
[0091] T'(x',y')=T(x',y')-1 / (w·h)·∑ x”,y” T(x”,y”) (3)
[0092] I'(x+x',y+y')=I(x+x',y+y')-1 / (w·h)·∑ x”,y” I(x+x”,y+y”) (4)
[0093] (x, y) represents the Cartesian coordinates of the top-left corner of image b1 retrieved in the current loop, indicating the position of the loop;
[0094] (x',y') represents the offset of each point in image b1 relative to the top left corner;
[0095] T(x',y') obtains the coordinates of the point in image a1, which is the coordinates of the top left corner of image a1 plus the offset;
[0096] I(x+x',y+y') obtains the coordinates of the point in image b1, which is the coordinates of the top left corner of image b1 plus the offset;
[0097] w and h are the width and height of the corresponding image, respectively;
[0098] T(x”,y”) obtains the coordinates of a point in image a1 in the complete EPI;
[0099] I(x+x”,y+y”) obtains the coordinates of a point in image b1 within the complete EPI;
[0100] Finally, the average values are sorted in descending order, and the average value is calculated in multiple search structures. The pixel value corresponding to the maximum average value is selected as the target pixel.
[0101] The slope information of the line connecting each target pixel value is calculated in the horizontal EPI image and converted to the corresponding [0,255] gray space as a preliminary depth estimation result.
[0102] Similarly, the macro-pixel block segmentation reference result of the current frame vertical EPI image is used as image a2, and the remaining pixel rows in this image except for the center column are copied to form a new copied image b2.
[0103] Then, each column of the copied image b2 is cyclically compared with the pixels in image a2. The standard deviation of squares (Sq_diff) and the correlation coefficient (R) are calculated in multi-scale space, and their average values are taken.
[0104] Finally, the average values are sorted in descending order, and the pixel value corresponding to the maximum average value is selected as the target pixel. The slope information of the line connecting each target pixel value in the vertical EPI image is calculated using the coordinates of each target pixel value, and then converted to the corresponding [0,255] gray space as the preliminary depth estimation result.
[0105] like Figure 3 As shown, its slope is: Convert it to the [0,255] grayscale space: Where dmin is the minimum slope of all macro pixel blocks, and dmax is the maximum slope of all macro pixel blocks.
[0106] Step 5: Using the preliminary estimation results and the triangular coordinate mapping commonly used in computer vision, the depth information of the target pixel is determined by the 3D matching result, thus obtaining the depth information of the original pixels contained in the horizontal / vertical EPI image.
[0107] The depth values of the target pixel positions in all horizontal / vertical EPI images are weighted and averaged for each original pixel, and then mapped back to the original video frame according to the coordinates to obtain the corresponding depth map output result of the video frame.
[0108] The light field data captured by the array-type light field video acquisition device of this invention is a four-dimensional sample of the light field. Uniform horizontal and vertical parallax exists between viewpoints in the same row and column, providing a basis for depth estimation. Epipolar Plane Images (EPI) are a commonly used tool in light field signal analysis and processing. The slope of the diagonal lines in an EPI image reflects the depth of the scene. Objects with larger horizontal displacements have greater parallax corresponding to diagonal lines in the EPI image, meaning they have smaller depths, and vice versa. This key characteristic of light field EPI images is used for scene depth estimation. Since the texture information of EPI images with large parallax light fields is relatively messy and no longer smooth and continuous diagonal lines, ordinary pixels in such EPI images are converted into macro-pixel blocks. A multi-scale hybrid feature search scheme is used to match the offset of the macro-pixel blocks to estimate the scene depth.
[0109] Example:
[0110] The light field data used in the testing of this invention includes two sets of real-world light field data (office and wall) and four sets of composite light field images with large parallax rendered using 3D software Blender: Bed, Pencil Case, Chair, and Oil Drum. The specific process is as follows:
[0111] Step s101: Acquire light field video array data, sequentially read in single-viewpoint original video files, and use ffmpeg to perform frame splitting processing on multiple light field videos to obtain scene multi-viewpoint sub-aperture image data with rich parallax information.
[0112] Step s102: Select consecutive rows and columns of pixels from each sub-aperture image to generate horizontal and vertical wide pixel EPI images.
[0113] Step s103: Using a macro-pixel block segmentation method that fuses semantic and texture information, gradient map template detection is performed on the center viewpoint row of each wide-pixel EPI image to obtain texture segmentation results; graph-based image segmentation is performed on the center viewpoint row of each wide-pixel EPI image to obtain its semantic segmentation results; logical OR operation is performed on the texture and semantic segmentation results to determine the final macro-pixel block segmentation results.
[0114] Step s104: Match the position of the target macro pixel block in each row of the wide pixel EPI image, and use two feature indicators, standard squared error and correlation coefficient, to search for the optimal matching solution. This is used to search repeatedly in multi-scale space, and select the highest probability of occurrence in multiple search structures as the matching coordinates (x,y).
[0115] Based on the coordinates of the matching results of the macro pixel block in the center row of the wide pixel EPI image, the slope information of the target macro pixel block in the EPI image is estimated and converted to the corresponding [0,255] grayscale space as the preliminary depth estimation result of the pixel region.
[0116] Step s105: Determine the depth information of the macro pixel block by matching the position of the macro pixel block in each row of the wide pixel EPI image. Perform step s104 on all horizontal and vertical wide pixel EPI images generated by the same frame of light field data to obtain the preliminary depth estimation results of the horizontal and vertical EPI of the scene at the current time. Perform the depth estimation results of the horizontal and vertical wide pixel EPI images on each pixel in the scene, and obtain the final depth map output result after weighted averaging.
[0117] The results indicate that:
[0118] This invention selects three algorithms from recent work in the field of light field depth estimation: Zhang-SPO, Shin-EPINET, Jeon-CVPR15, and the stereo estimation algorithm SGBM as experimental control groups. Additionally, a scheme MPSM-baseline is added to the control group. This scheme masks the macro-pixel block segmentation algorithm with fused semantics in the proposed schemes, using only a single texture information segmentation algorithm for subsequent search and matching of the output results. The complete depth estimation algorithm proposed in this invention is referred to as MPSM-Fusion, and the resulting wide-pixel EPI image is shown below. Figure 4 As shown.
[0119] (1) Depth estimation results of different schemes on the actual light field datasets Office and Wall.
[0120] Because real-world light field data acquired using array-type light field acquisition devices lacks true depth maps, it's impossible to effectively compare the performance of different solutions using objective metrics. However, from... Figure 4 The subjective results of the first four schemes demonstrate that traditional light field depth estimation algorithms and binocular estimation algorithms are ineffective when processing large parallax light field data. However, the light field depth estimation algorithm based on semantic texture fusion and macro-pixel block matching proposed in this invention can effectively overcome the impact of parallax and achieve a more efficient scene depth estimation effect.
[0121] (2) Comparison of PSNR (dB) of different depth estimation algorithms on synthetic large parallax light field datasets.
[0122] The PSNR of an image is an objective standard for measuring the level of image distortion or noise. The larger the PSNR value between the depth map estimated by the algorithm and the true depth map, the more similar the two images are, that is, the more accurate the depth estimation result of the algorithm is. It can be seen that the depth estimation method proposed in this invention has good performance.
[0123]
[0124]
[0125] (3) Comparison of SSIM indexes of different depth estimation algorithms on synthetic large parallax light field datasets.
[0126] The SSIM of an image is also a metric for measuring the similarity between two images. The larger the SSIM value between the depth map estimated by the algorithm and the ground truth depth map, the more similar the two images are, meaning that the depth estimation result of the algorithm is more accurate. The depth estimation method proposed in this invention has good performance.
[0127]
Claims
1. A large baseline light field video depth estimation method, characterized in that, The specific steps are as follows: First, images are acquired using a light field camera. The original single-viewpoint video files are read in sequentially and split into frames. Video frame images from different viewpoints at the same time are stored as a group to obtain a multi-viewpoint light field sub-aperture image set. For each set of images, the horizontal EPI image and the vertical EPI image of each frame are generated by using the row pixels and column pixels in each frame sub-aperture image; Then, gradient map template detection and graph-based image segmentation are performed using the center row of each frame of horizontal EPI image to obtain macro-pixel block segmentation reference results for each frame of horizontal EPI image; Gradient map template detection and graph-based image segmentation are performed using the center column of each frame of vertical EPI image to obtain macro-pixel block segmentation reference results for each frame of vertical EPI image; Next, using the macro-pixel block segmentation reference results of each frame of horizontal EPI image, the target pixel position is found in each frame of horizontal EPI image as a preliminary depth estimation result; Using the macro-pixel block segmentation reference results of each frame of vertical EPI image, the target pixel position is found in each frame of vertical EPI image as a preliminary depth estimation result; Specifically: First, select each frame of sub-aperture image one by one, take the macro pixel block segmentation reference result of the current frame horizontal EPI image as image a1, and copy the remaining pixel rows in the image except the center row to form a new copied image b1. The length of the pixel row extension in image b1 is the maximum length of the macro-pixel block segmented in image a1; Then, each row of the copied image b1 is cyclically compared with the pixels in image a1. The standard deviation of squares (Sq_diff) and the correlation coefficient (R) are calculated in multi-scale space, and their average values are taken. The closer the standard squared deviation is to 0, the higher the matching degree between the two pixels, and vice versa; the closer the correlation coefficient R is to 1, the more it indicates that the target regions of the two pixels are completely matched, and vice versa, it indicates that there is no correlation between the two. in: (x,y) represents the Cartesian coordinates of the top-left corner of image b1 retrieved in the current loop, indicating the position of the loop; (x',y') represents the offset of each point in image b1 relative to the top left corner; T(x',y') obtains the coordinates of the point in image a1, which is the coordinates of the top left corner of image a1 plus the offset; I(x+x',y+y') obtains the coordinates of the point in image b1, which is the coordinates of the top left corner of image b1 plus the offset; w and h are the width and height of the corresponding image, respectively; T(x”,y”) obtains the coordinates of a point in image a1 in the complete EPI; I(x+x”,y+y”) obtains the coordinates of a point in image b1 within the complete EPI; Finally, sort the average values in descending order and select the pixel value corresponding to the maximum average value as the target pixel; The slope information of the line connecting each target pixel value is calculated in the horizontal EPI image and converted to the corresponding [0,255] gray space as a preliminary depth estimation result. Similarly, the macro-pixel block segmentation reference result of the current frame's vertical EPI image is used as image a2. The remaining pixel rows in this image, excluding the center column, are copied to form a new copied image b2. Then, each column of the copied image b2 is cyclically compared with the pixels in image a2. The standard squared error Sq_diff and the correlation coefficient R are calculated in a multi-scale space, and the average value is taken. The average values are sorted in descending order, and the pixel value corresponding to the maximum average value is selected as the target pixel. The slope information of the line connecting each target pixel value is calculated in the vertical EPI image using the coordinates of each target pixel value, and converted to the corresponding [0,255] grayscale space as the preliminary depth estimation result. Finally, using the triangular coordinate mapping method in computer vision, the depth information of the target pixel is determined by the preliminary estimation results, that is, the depth information of the original pixels contained in the horizontal or vertical EPI image.
2. The method of claim 1, wherein, Each group of video frame images includes N sub-aperture images, and each sub-aperture image corresponds to a viewpoint.
3. The method of claim 1, wherein, The generation of the horizontal and vertical EPI images for each frame is specifically as follows: For each sub-aperture image, starting from the first row, pixels with the same height value are extracted sequentially from left to right to form the horizontal EPI image of that sub-aperture image; similarly, starting from the first column, pixels with the same width value are extracted sequentially from top to bottom to obtain the vertical EPI image.
4. The large baseline light field video depth estimation method according to claim 1, characterized in that, The process of obtaining the macro-pixel block segmentation reference results corresponding to each frame of horizontal or vertical EPI image is as follows: First, for the current horizontal EPI image, gradient map template detection is performed by selecting the horizontal width pixels of the center row to obtain the texture segmentation result; Then, graph-based image segmentation is performed again on the central row to obtain its semantic segmentation result; Logical OR operation is performed on the texture and semantic segmentation results to determine the original wide pixels of the horizontal EPI image that are related, and macro pixels are defined to obtain the macro pixel block segmentation reference result in the frame. Similarly, for the current vertical EPI image, gradient map template detection is performed by selecting the vertical width pixels of the center column to obtain the texture segmentation result; Then, graph-based image segmentation is performed again on the central column to obtain its semantic segmentation results; A logical OR operation is performed on the texture and semantic segmentation results to determine the original wide pixels that are relevant to the vertical EPI image, which are defined as macro pixels, thus obtaining the macro pixel block segmentation reference result in the frame.
5. The large baseline light field video depth estimation method according to claim 1, characterized in that, The depth information of the target pixel is specifically obtained by: taking a weighted average of the depth values of the target pixel positions in all horizontal or vertical EPI images, mapping them back to the original video frame according to the coordinates, and obtaining the depth map output result of the corresponding video frame.
Citation Information
Patent Citations
Depth estimation method based on optical field
CN104598744A
Light field image depth estimation method
CN108596965A