Multi-view based annular heterogeneous panoramic depth perception device and three-dimensional reconstruction method
Patent Information
- Application Number
- CN202210891995.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-07-27
AI Technical Summary
其中,深度学习方法过于依赖数据,尤其是单目深度估计,泛化能力存在巨大挑战,导致基于深度学习的单目重建难以在实际场景中稳定运行;而而对称性双目结构虽然可以直接获取深度,但在大场景中标定非常困难
[0047]本发明结合多相机视觉和深度学习的优势,提出一种差异化异构成像的非结构化全景深度感知装置,并在此基础上,利用局部相机的优势,通过融合方法得到全景图像和全景深度感知结果;本发明所提出的全景深度感知装置和三维重建方法,对数据的依赖更少,泛化能力更强,并且适用于各种场景,在不同的大场景下均能非常容易地实现数据的标定。
Smart Images

Figure CN115272118B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computational imaging technology, and in particular to a multi-viewpoint ring-shaped heterogeneous panoramic depth sensing device and a three-dimensional reconstruction method. Background Technology
[0002] According to technical research, panoramic synthesis is widely used in security, consumer entertainment, and other fields, such as large-scene monitoring, panoramic real estate viewing, and Baidu Street View photography, demonstrating broad application areas and value. Panoramic synthesis algorithms solve the problem of synthesizing two-dimensional panoramic images. As artificial intelligence gradually emerges in the three-dimensional field, high-resolution panoramic spaces in three-dimensional visual space will have even greater application value.
[0003] To achieve panoramic high-resolution 3D reconstruction, existing methods fall into two categories: deep learning methods and traditional multi-view geometric methods. Deep learning methods are overly reliant on data, especially monocular depth estimation, which faces significant challenges in generalization, making it difficult for deep learning-based monocular reconstruction to operate stably in real-world scenarios. While symmetrical binocular structures can directly obtain depth, calibration in large scenes is extremely difficult. Traditional multi-view geometric methods, after years of development, have reached performance bottlenecks, making them unsuitable for scenarios with no texture or motion blur, and calibration remains challenging. Summary of the Invention
[0004] Based on this, and to address the aforementioned technical problems, a multi-viewpoint ring-shaped heterogeneous panoramic depth sensing device and a 3D reconstruction method are provided.
[0005] In a first aspect, a multi-viewpoint-based ring-shaped heterogeneous panoramic depth sensing device includes a ring-shaped fixed bracket. Multiple global reconstruction auxiliary cameras are fixed at equal intervals along the circumference of the ring-shaped fixed bracket, and a global synthesis camera is fixed between every two adjacent global reconstruction auxiliary cameras. At the position of each global synthesis camera, an array camera fixed bracket is also correspondingly fixed along the axial direction of the ring-shaped fixed bracket. The middle of each array camera fixed bracket is fixed to the ring-shaped fixed bracket, and each of the two ends of each array camera fixed bracket has a local camera fixed to it. The two local cameras on each array camera fixed bracket and the corresponding global synthesis camera on that array camera fixed bracket constitute a set of array cameras.
[0006] Multiple global reconstruction auxiliary cameras and multiple global synthesis cameras respectively cover multiple peripheral areas around the annular fixed support. Adjacent global reconstruction auxiliary cameras and global synthesis cameras have overlapping perception areas. The vertical field of view of each global synthesis camera is greater than or equal to twice the vertical field of view of the local camera.
[0007] Optionally, nine global reconstruction auxiliary cameras are fixed at equal intervals on the annular fixed bracket, and there are a total of nine global synthesis cameras.
[0008] Further optionally, the horizontal field of view of each global reconstruction auxiliary camera and the horizontal field of view of each global synthesis camera are both greater than or equal to 60°, and the perceptual image resolution of each global synthesis camera and local camera is greater than or equal to 14 million pixels.
[0009] Alternatively, the field-of-view optical axes of all global reconstruction auxiliary cameras and all global synthesis cameras are coplanar with the annular fixing bracket.
[0010] Alternatively, all global reconstruction auxiliary cameras and all global synthesis cameras are of the same model.
[0011] Secondly, a multi-viewpoint ring-shaped heterogeneous panoramic 3D reconstruction method includes:
[0012] Step 1: Construct the above-mentioned ring-shaped heterogeneous panoramic depth perception device, perform visual calibration on the global cameras in it, obtain the intrinsic parameters of 18 global cameras and the extrinsic parameter matrices of adjacent global cameras, and calculate the extrinsic parameter matrices between all global synthetic cameras based on the global reconstruction auxiliary camera.
[0013] Step 2: Based on the intrinsic parameters of the 18 global cameras and the extrinsic parameter matrices of the adjacent global cameras, calculate the dense depth map from the perspective of each global synthetic camera.
[0014] Step 3: Images are acquired using the constructed ring-shaped heterogeneous panoramic depth sensing device. Nine sets of array camera data are obtained at the same time stamp, including nine globally synthesized camera images and 18 locally synthesized camera images. The deformation matrix between each pair of adjacent images in the nine globally synthesized camera images is calculated, and the obtained deformation matrix is used to fuse the panoramic image and the dense depth map to obtain the panoramic image I. G And panoramic depth perception results D G .
[0015] Optionally, step one specifically includes:
[0016] The above-mentioned ring-shaped heterogeneous panoramic depth perception device was constructed, and the 18 global cameras were numbered sequentially from 1 to 18. Odd-numbered cameras were global reconstruction auxiliary cameras among the global cameras, and even-numbered cameras were global synthesis cameras among the global cameras.
[0017] Based on Zhang Zhengyou's calibration method, a checkerboard pattern was used to visually calibrate the global cameras and adjacent global cameras in the ring-shaped heterogeneous panoramic depth sensing device, obtaining the intrinsic parameters of 18 global cameras. in, The intrinsic parameters represent the global reconstruction auxiliary camera within the global camera. The intrinsic parameters of the global composite camera in the global camera array; and the extrinsic parameter matrix {T} of the adjacent global cameras. 1,2 T 2,3 , ..., T 17,18 T 18,1 Each extrinsic parameter matrix contains a rotation matrix R and a translation vector t;
[0018] m takes values of 2, 4, 6...16 in sequence. Perform the following operation for each m:
[0019] Based on the extrinsic parameter T between the m-th global synthetic camera and the m+1-th global reconstruction auxiliary camera m,m+1 {R m,m+1 , t m,m+1}, and the extrinsic parameter T between the m+1 global reconstruction auxiliary camera and the m+2 global synthesis camera. m+1,m+2 {R m+1,m+2 , t m+1,m+2 The extrinsic parameter T between the m-th global composite camera and the m+2-th global composite camera was calculated. m,m+2 {R m,m+2 , t m,m+2},in:
[0020] R m,m+2 =R m+1,m+2 *R m,m+1
[0021] t m,m+2 =R m+1,m+2 t m,m+1 +t m+1,m+2 ;
[0022] This allows us to calculate the extrinsic matrix between all global synthetic cameras.
[0023] Further, optionally, step two specifically includes:
[0024] Using the m-1 global reconstruction auxiliary camera and the m-1 global synthesis camera as a stereo camera pair, based on the intrinsic and extrinsic parameters of the m-1 global reconstruction auxiliary camera and the m-1 global synthesis camera... and T m-1,m The stereo image composed of the m-1 global reconstruction auxiliary camera and the m global synthesis camera is used to compare I. m-1 and I m A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the first dense depth map D from the viewpoint of the m-th global synthetic camera. m-1,m ;
[0025] Using the m-th global composite camera and the m+1-th global reconstruction auxiliary camera as a stereo camera pair, based on the intrinsic and extrinsic parameters of the m-th global composite camera and the m+1-th global reconstruction auxiliary camera... and T m,m+1 The stereo image composed of global synthetic camera m and global reconstruction auxiliary camera m+1 is used to compare I. m and I m+1 A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the second dense depth map D from the viewpoint of the m-th global synthetic camera. m,m+1 ;
[0026] Obtain the stereo vision pair consisting of the m-1 global reconstruction auxiliary camera and the m global synthesis camera, and obtain the mask M of the effective correction area during the stereo correction process. m-1,m And to obtain the stereo vision pair consisting of global synthesis camera m and global reconstruction auxiliary camera m+1, and the mask M of the effective correction area obtained during the stereo correction process. m,m+1 In the mask layer, a pixel index value of 1 represents valid, and 0 represents invalid.
[0027] According to the mask M m-1,m and mask M m,m+1 For the first dense depth map D m-1,m Second Dense Depth Map D m,m+1 The merged data is obtained as a merged dense depth map D. m The merging principle is as follows, for pixel indices (p, q):
[0028]
[0029] Bilateral filtering is applied to Dm to obtain the final dense depth map Dm from the perspective of the m-th global synthetic camera;
[0030] m is taken as 2, 4, 6..., 16, and the above operation is performed for each m.
[0031] With m set to 18, the 17th global reconstruction auxiliary camera and the 18th global synthesis camera are used as a stereo camera pair. Based on the intrinsic and extrinsic parameters of the 17th global reconstruction auxiliary camera and the 18th global synthesis camera... and T 17,18 The stereo image composed of the 17th global reconstruction auxiliary camera and the 18th global synthesis camera is used to compare I. 17 and I 18 A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the first dense depth map D from the perspective of the 18th global synthetic camera. 17,18 ;
[0032] Using global composite camera 18 and global reconstruction auxiliary camera 1 as a stereo camera pair, based on the intrinsic and extrinsic parameters of global composite camera 18 and global reconstruction auxiliary camera 1... and T 18,1 The stereo image composed of global synthetic camera No. 18 and global reconstruction auxiliary camera No. 1 is used to compare I. 18 Together with I1, a binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the second dense depth map D from the perspective of the global synthetic camera No. 18. 18,1 ;
[0033] Obtain the stereo vision pair consisting of global reconstruction auxiliary camera 17 and global synthesis camera 18, and obtain the mask M of the effective correction area during stereo correction. 17,18 And to obtain the stereo vision pair consisting of global synthesis camera No. 18 and global reconstruction auxiliary camera No. 1, and the mask M of the effective correction area obtained during the stereo correction process. 18,1 In the mask layer, a pixel index value of 1 represents valid, and 0 represents invalid.
[0034] According to the mask M 17,18 and mask M 18,1 For the first dense depth map D 17,18 Second Dense Depth Map D 18,1 The merged data is obtained as a merged dense depth map D. 18 The merging principle is as follows, for pixel indices (p, q):
[0035]
[0036] For D 18 Bilateral filtering was performed to obtain the final dense depth map D′ from the perspective of the 18th global synthetic camera. 18 .
[0037] Further, optionally, step three specifically includes:
[0038] Using a constructed ring-shaped heterogeneous panoramic depth sensing device, image data was acquired using external synchronization triggering or software timestamp synchronization methods. At the same timestamp, nine sets of array camera data were obtained, including nine globally synthesized camera images G2, G4, ..., G... taken by a global synthesized camera. 18 , and 18 partial camera images taken by a partial camera Among them G i , Images from a set of array cameras;
[0039] At the same timestamp, for a set of globally synthesized camera image sequences G2, G4, ..., G 18Feature detection algorithms are used to detect features in adjacent images;
[0040] For a pair of adjacent images in a sequence of globally synthesized camera images, image registration is performed based on the detected feature points to obtain the deformation matrix TG between each pair of adjacent images in the globally synthesized camera image sequence. 2,4 TG 4,6 ... TG 16,18 ;
[0041] Using the obtained deformation matrix, the nine globally synthesized images G2, G4, ..., G... 18 The images are stitched together to obtain a 360-degree panoramic stitched image I. g ;
[0042] Using the obtained deformation matrix, the dense depth maps D′2, D′4, ..., D′ from the global synthetic camera viewpoint are plotted. 18 The images are stitched together to obtain a 360-degree panoramic depth perception result D. g ;
[0043] Using a cross-scale image fusion network, 18 local camera images are fused into I g In the process, a panoramic image at the 100-megapixel level was obtained. G ;
[0044] Using super-resolution methods, for D g To improve the resolution, obtain the same as I G Panoramic depth perception results at the same resolution (100 million pixels) D G .
[0045] Further, alternatively, a random consistency sampling algorithm is used for image registration; the cross-scale image fusion network is, for example, the CrossNet network.
[0046] The present invention has at least the following beneficial effects:
[0047] This invention combines the advantages of multi-camera vision and deep learning to propose an unstructured panoramic depth perception device with differentiated heterogeneous imaging. Based on this, it utilizes the advantages of local cameras to obtain panoramic images and panoramic depth perception results through a fusion method. The panoramic depth perception device and 3D reconstruction method proposed in this invention have less dependence on data, stronger generalization ability, and are applicable to various scenarios. Data calibration can be easily achieved in different large-scale scenarios.
[0048] When generating the final panoramic image, this invention fuses local camera images into a 360-degree panoramic stitched image composed of global images to obtain a panoramic image with a resolution of hundreds of millions of pixels. It also improves the 360-degree panoramic depth perception result through super-resolution methods to obtain a panoramic depth perception result with a resolution of hundreds of millions of pixels. Therefore, this invention improves the resolution and accuracy of global depth and imaging perception. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the structure of a multi-viewpoint-based ring-shaped heterogeneous panoramic depth sensing device according to an embodiment of the present invention;
[0050] Figure 2 A detailed structural schematic diagram of a multi-viewpoint-based ring-shaped heterogeneous panoramic depth sensing device is provided in one embodiment of the present invention;
[0051] Figure 3 A flowchart illustrating a multi-viewpoint ring-shaped heterogeneous panoramic 3D reconstruction method provided in one embodiment of the present invention;
[0052] Figure 4 This is a schematic diagram showing the numbering of a ring-shaped heterogeneous panoramic depth sensing device in one embodiment of the present invention.
[0053] Explanation of reference numerals in the attached figures:
[0054] 101. Global reconstruction auxiliary camera; 102. Global synthesis camera; 103. Circular fixed support; 104. Array camera fixed support; 105. Local camera. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] In one embodiment, to achieve panoramic 360-degree high-resolution depth perception, such as Figure 1 As shown, a multi-viewpoint ring-shaped heterogeneous panoramic depth sensing device is provided, wherein the sensing hardware adopts a ring-shaped heterogeneous design in structure; the ring-shaped heterogeneous panoramic depth sensing device includes a ring-shaped fixed bracket 103, on which multiple global reconstruction auxiliary cameras 101 are fixed at equal intervals along the circumference. Figure 1 The white camera in the middle), and a global synthesis camera 102 is fixed between every two adjacent global reconstruction auxiliary cameras 101. Figure 1(The black camera in the image). In other words, global cameras (global reconstruction auxiliary camera 101 and global compositing camera 102 are both collectively referred to as global cameras) are evenly distributed on the ring-shaped fixed bracket 103. In terms of characteristics, all global reconstruction auxiliary cameras 101 and all global compositing cameras 102 are exactly the same model. Global cameras can be divided into two types according to their functions: white cameras (represented by 101) are global reconstruction auxiliary cameras, and black cameras (represented by 102) are global compositing cameras.
[0057] Furthermore, such as Figure 2 As shown, each global composite camera 102 is also correspondingly fixed with an array camera mounting bracket 104 along the axial direction of the annular mounting bracket 103. The middle part of each array camera mounting bracket 104 is fixed to the annular mounting bracket 103, and each end of each array camera mounting bracket 104 is fixed with a local camera 105. The two local cameras 105 on each array camera mounting bracket 104 and the corresponding global composite camera 102 of the array camera mounting bracket 104 constitute a set of array cameras. In other words, each global composite camera 102 also includes two additional local cameras 105 in a columnar structure (vertical direction).
[0058] Furthermore, multiple global reconstruction auxiliary cameras 101 and multiple global synthesis cameras 102 respectively cover multiple peripheral areas around the annular fixed bracket 103. Adjacent global reconstruction auxiliary cameras 101 and global synthesis cameras 102 have a perceptual overlap area, that is, there is a perceptual overlap area between each pair of global cameras. For a single columnar camera array, the vertical field of view of each global synthesis camera 102 should be at least twice the vertical field of view of the local camera 105.
[0059] As an optional solution, such as Figure 1 As shown, nine global reconstruction auxiliary cameras 101 are fixed at equal intervals on the annular fixed bracket 103, and there are also nine corresponding global synthesis cameras 102. This configuration of nine cameras satisfies both the field of view and resolution requirements and represents the minimum number of cameras required under these conditions. Theoretically, the horizontal field of view of the 18 global cameras should be no less than 40 degrees. To ensure a sufficiently large overlapping area for subsequent algorithm development, the horizontal viewing angle of a single global camera should be at least 60 degrees. That is, the horizontal viewing angle of each global reconstruction auxiliary camera 101 and each global synthesis camera 102 should be greater than or equal to 60°. Furthermore, the perceived image resolution of each global synthesis camera 102 and the local camera 105 should be greater than or equal to 14 megapixels. The annular heterogeneity in this panoramic depth sensing device mainly reflects the differentiated settings of the global cameras and the arbitrarily configurable angles of the local cameras 105 in the array camera.
[0060] In addition, the field-of-view optical axes of all global reconstruction auxiliary cameras 101 and all global synthesis cameras 102 are coplanar with the annular fixed bracket 103.
[0061] As a preferred embodiment, the reverse extension lines of the field of view optical axes of all global reconstruction auxiliary cameras 101 and all global synthesis cameras 102 pass through the center of the annular fixed bracket 103.
[0062] This invention combines the advantages of multi-camera vision and deep learning to propose an unstructured panoramic depth perception device with differentiated heterogeneous imaging. It has less dependence on data, stronger generalization ability, and is applicable to various scenarios. It can easily achieve data calibration in different large-scale scenarios.
[0063] In one embodiment, such as Figure 3 As shown, a multi-viewpoint ring-shaped heterogeneous panoramic 3D reconstruction method is provided, including the following steps:
[0064] Step S301, according to Figure 1 and Figure 2 The structure described in the above embodiment describes the construction of a ring-shaped heterogeneous panoramic depth sensing device provided in one of the optional solutions. The ring-shaped heterogeneous panoramic depth sensing device includes 9 global synthetic cameras and 9 global reconstruction auxiliary cameras. Visual calibration is performed on the global cameras to obtain the intrinsic parameters of 18 global cameras and the extrinsic parameter matrices of adjacent global cameras. Based on the global reconstruction auxiliary cameras, the extrinsic parameter matrices between all global synthetic cameras are calculated.
[0065] Step S301 specifically includes:
[0066] (1) Construct the ring-shaped heterogeneous panoramic depth sensing device provided in one of the optional solutions of the above embodiments. The ring-shaped heterogeneous panoramic depth sensing device includes 9 global synthetic cameras and 9 global reconstruction auxiliary cameras, and the 18 global cameras are numbered sequentially from 1 to 18, such as... Figure 4 As shown, odd-numbered cameras are global reconstruction auxiliary cameras in the global camera system, and even-numbered cameras are global synthesis cameras in the global camera system.
[0067] (2) Based on Zhang Zhengyou's calibration method, a checkerboard pattern was used to perform visual calibration on the global cameras and adjacent global cameras in the ring-shaped heterogeneous panoramic depth sensing device, and the intrinsic parameters of 18 global cameras were obtained. in, The intrinsic parameters represent the global reconstruction auxiliary camera within the global camera. The intrinsic parameters of the global composite camera in the global camera array; and the extrinsic parameter matrix {T} of the adjacent global cameras. 1,2 T 2,3 , ..., T 17,18T 18,1 Each extrinsic parameter matrix contains a rotation matrix R and a translation vector t;
[0068] (3) Calculate the extrinsic parameters between all global composite cameras according to the following steps:
[0069] m takes values of 2, 4, 6...16 in sequence. Perform the following operation for each m:
[0070] Based on the extrinsic parameter T between the m-th global synthetic camera and the m+1-th global reconstruction auxiliary camera m,m+1 {R m,m+1 , t m,m+1}, and the extrinsic parameter T between the m+1 global reconstruction auxiliary camera and the m+2 global synthesis camera. m+1,m+2 {R m+1,m+2 , t m+1,m+2 The extrinsic parameter T between the m-th global composite camera and the m+2-th global composite camera was calculated. m,m+2 {R m,m+2 , t m,m+2},in:
[0071] R m,m+2 =R m+1,m+2 *R m,m+1
[0072] t m,m+2 =R m+1,m+2 t m,m+1 +t m+1,m+2 ;
[0073] Therefore, the extrinsic parameter matrix between all global synthetic cameras can be calculated using the global reconstruction auxiliary camera.
[0074] For example, taking the extrinsic parameter determination between global composite cameras 2 and 4 as an example, this process requires the assistance of global auxiliary camera 3. The extrinsic parameter between global composite camera 2 and global reconstruction auxiliary camera 3 is T. 2,3 {R 2,3 , t 2,3 The extrinsic parameter between the No. 3 global reconstruction auxiliary camera and the No. 4 global synthesis camera is T. 3,4 {R 3,4 , t 3,4}, substitute the extrinsic parameter T 2,4 {R 2,4 , t 2,4 The calculation method is as follows:
[0075] R 2,4 =R 3,4 *R 2,3
[0076] t 2,4 =R 3,4t 2,3 +t 3,4 ;
[0077] Similarly, by using a global reconstruction auxiliary camera, the extrinsic parameter T between all global composite cameras is calculated. 2,4 {R 2,4 , t 2,4}、T 4,6 {R 4,6 , t 4,6}、T 6,8 {R 6,8 , t 6,8}、......
[0078] Step S302: Based on the intrinsic parameters of the 18 global cameras and the extrinsic parameter matrices of the adjacent global cameras, calculate the dense depth map from the perspective of each global synthetic camera.
[0079] Step S302 specifically includes:
[0080] m takes values of 2, 4, 6..., 16 in sequence, and performs the following operations for each m;
[0081] (1) Using the m-1 global reconstruction auxiliary camera and the m-1 global synthesis camera as a stereo camera pair, based on the intrinsic and extrinsic parameters of the m-1 global reconstruction auxiliary camera and the m-1 global synthesis camera... and T m-1,m The stereo image composed of the m-1 global reconstruction auxiliary camera and the m global synthesis camera is used to compare I. m-1 and I m A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the first dense depth map D from the viewpoint of the m-th global synthetic camera. m-1,m ;
[0082] (2) Using the m-th global composite camera and the m+1-th global reconstruction auxiliary camera as a stereo camera pair, based on the intrinsic and extrinsic parameters of the m-th global composite camera and the m+1-th global reconstruction auxiliary camera... and T m,m+1 The stereo image composed of global synthetic camera m and global reconstruction auxiliary camera m+1 is used to compare I. m and I m+1 A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the second dense depth map D from the viewpoint of the m-th global synthetic camera. m,m+1 ;
[0083] (3) Obtain the stereo vision pair consisting of the m-1 global reconstruction auxiliary camera and the m global synthesis camera, and obtain the mask M of the effective correction area during the stereo correction process.m-1,m And to obtain the stereo vision pair consisting of global synthesis camera m and global reconstruction auxiliary camera m+1, and the mask M of the effective correction area obtained during the stereo correction process. m,m+1 In the mask layer, a pixel index value of 1 represents valid, and 0 represents invalid.
[0084] (4) According to the mask M m-1,m and mask M m,m+1 For the first dense depth map D m-1,m Second Dense Depth Map D m,m+1 The merged data is obtained as a merged dense depth map D. m The merging principle is as follows, for pixel indices (p, q):
[0085]
[0086] (5) Regarding D m Bilateral filtering is performed to obtain the final dense depth map D′ from the viewpoint of the m-th global synthetic camera. m ;
[0087] When m is 18, perform the following operations on it:
[0088] (1) Using the 17th global reconstruction auxiliary camera and the 18th global synthesis camera as a stereo camera pair, based on the intrinsic and extrinsic parameters of the 17th global reconstruction auxiliary camera and the 18th global synthesis camera... and T 17,18 The stereo image composed of the 17th global reconstruction auxiliary camera and the 18th global synthesis camera is used to compare I. 17 and I 18 A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the first dense depth map D from the perspective of the 18th global synthetic camera. 17,18 ;
[0089] (2) Using the 18th global composite camera and the 1st global reconstruction auxiliary camera as a stereo camera pair, based on the intrinsic and extrinsic parameters of the 18th global composite camera and the 1st global reconstruction auxiliary camera... and T 18,1 The stereo image composed of global synthetic camera No. 18 and global reconstruction auxiliary camera No. 1 is used to compare I. 18 Together with I1, a binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the second dense depth map D from the perspective of the global synthetic camera No. 18. 18,1 ;
[0090] (3) Obtain the stereo vision pair consisting of the global reconstruction auxiliary camera No. 17 and the global synthesis camera No. 18, and obtain the mask M of the effective correction area during the stereo correction process. 17,18 And to obtain the stereo vision pair consisting of global synthesis camera No. 18 and global reconstruction auxiliary camera No. 1, and the mask M of the effective correction area obtained during the stereo correction process. 18,1 In the mask layer, a pixel index value of 1 represents valid, and 0 represents invalid.
[0091] (4) According to the mask M 17,18 and mask M 18,1 For the first dense depth map D 17,18 Second Dense Depth Map D 18,1 The merged data is obtained as a merged dense depth map D. 18 The merging principle is as follows, for pixel indices (p, q):
[0092]
[0093] (5) Regarding D 18 Bilateral filtering was performed to obtain the final dense depth map D′ from the perspective of the 18th global synthetic camera. 18 .
[0094] Taking global compositing camera #2 as an example, the other global compositing cameras are similar:
[0095] (1) Using cameras 1 and 2 as a stereo camera pair, based on the intrinsic and extrinsic parameters of cameras 1 and 2 ( T 1,2 The stereo image pairs (I1 and I2) composed of cameras 1 and 2 are used to perform binocular depth calculation, resulting in the dense depth map D from the perspective of camera 2. 1,2 ;
[0096] (2) Similarly, the stereo image pair composed of cameras 2 and 3 includes intrinsic and extrinsic parameters ( T 2,3 Stereo depth calculations were performed on the stereo image pairs (I3 and I2) consisting of images 2 and 3, to obtain the dense depth map D from the perspective of camera 2. 2,3 ;
[0097] (3) Obtain the mask M of the effective correction area obtained during the stereo correction process for the stereo vision pair composed of No. 1 and No. 2. 1,2 Similarly, the mask M, which forms the stereo vision pair composed of pairs 2 and 3, is obtained as the effective correction area during the stereo correction process. 2,3 In the mask layer, a pixel index value of 1 represents valid, and 0 represents invalid.
[0098] (4) According to the mask M1,2 and M 2,3 For D 1,2 and D 2,3 The images are then merged to obtain the final merged dense depth map D2. The merging principle is as follows: for pixel indices (p, q):
[0099]
[0100] (5) Perform bilateral filtering on D2 to obtain the final dense depth map D′2 from the perspective of the No. 2 global synthetic camera.
[0101] In addition, the stereo depth calculation will also use the extrinsic parameters between the global synthetic cameras calculated in step S301. The specific calculation process is common knowledge to those in the art and will not be elaborated here.
[0102] Step S303: Images are acquired using the constructed ring-shaped heterogeneous panoramic depth sensing device. Nine sets of array camera data are obtained at the same time stamp, including nine global composite camera images and 18 local camera images. The deformation matrix between each pair of adjacent images in the nine global composite camera images is calculated, and the obtained deformation matrix is used to fuse the panoramic image and the dense depth map to obtain the panoramic image I. G And panoramic depth perception results D G .
[0103] Step S303 specifically includes:
[0104] (1) Using the constructed ring-shaped heterogeneous panoramic depth sensing device, image data is acquired using external synchronization triggering or software timestamp synchronization methods. At the same timestamp, nine sets of array camera data can be obtained, including nine global composite camera images G2, G4, ..., G6 taken by the global composite camera. 18 , and 18 partial camera images taken by a partial camera Among them G i , Images from a set of array cameras;
[0105] (2) At the same timestamp, for a set of globally synthesized camera image sequences G2, G4, ..., G 18 Feature detection algorithms are used to detect features in adjacent images in a globally synthesized image sequence.
[0106] (3) For a pair of adjacent images in a set of globally synthesized camera image sequences, RANSAC (Random Consistency Sampling Algorithm) is used to register the detected feature points, resulting in the deformation matrix TG of the adjacent images. This deformation matrix can be used to stitch the adjacent images together. Based on this method, the deformation matrix TG between each pair of adjacent images in the globally synthesized camera image sequence can be obtained separately. 2,4 TG 4,6 ... TG 16,18 :
[0107] (4) Using the deformation matrix obtained above, the nine globally synthesized images G2, G4, ..., G 18 The images are stitched together to obtain a 360-degree panoramic stitched image I. g ;
[0108] (5) Using the deformation matrix obtained above, the dense depth maps D′2, D′4, ..., D′ from the perspectives of the nine global synthetic cameras are transformed. 18 The images are stitched together to obtain a 360-degree panoramic depth perception result D. g ;
[0109] (6) Use a cross-scale image fusion network (such as CrossNet, an end-to-end super-resolution fusion method) to fuse 18 local camera images into I g In the process, a panoramic image at the 100-megapixel level was obtained. G ;
[0110] (7) Using super-resolution methods, D g To improve the resolution, obtain the same as I G Panoramic depth perception results at the same resolution (100 million pixels) D G .
[0111] Leveraging the advantages of deep learning in complex scene learning, and considering its stringent data requirements, this invention proposes a multi-viewpoint unstructured panoramic high-resolution 3D reconstruction device and method, aiming to solve the following two problems: Problem 1: Monocular depth estimation requires a large amount of data, and the reconstruction results lack dimension. While symmetrical binocular structures can directly obtain depth, calibration is very difficult in large scenes. Therefore, how to utilize the advantages of deep learning and solve the camera calibration problem is a key consideration. Problem 2: How can panoramic camera arrays improve resolution and enhance the refinement of local perception in large scenes while solving panoramic depth perception?
[0112] Binocular depth estimation based on deep learning has a geometrical prior advantage and requires less training data. Therefore, this invention combines the advantages of multi-camera vision and deep learning to propose an unstructured panoramic depth perception device with differentiated heterogeneous imaging. Based on this, it utilizes the advantages of local cameras to obtain panoramic images and panoramic depth perception results through a fusion method. The panoramic depth perception device and 3D reconstruction method proposed in this invention have less dependence on data, stronger generalization ability, and are applicable to various scenarios. Data calibration can be easily achieved in different large-scale scenarios.
[0113] When generating the final panoramic image, this invention fuses local camera images into a 360-degree panoramic stitched image composed of global images to obtain a panoramic image with a resolution of hundreds of millions of pixels. It also improves the 360-degree panoramic depth perception result through super-resolution methods to obtain a panoramic depth perception result with a resolution of hundreds of millions of pixels. Therefore, this invention improves the resolution and accuracy of global depth and imaging perception.
[0114] It should be understood that, although Figure 3 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 3 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.
[0115] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program relating to all or part of the processes in the methods of the above embodiments.
[0116] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon relating to all or part of the processes in the methods of the above embodiments.
[0117] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0118] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0119] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A multi-viewpoint ring-shaped heterogeneous panoramic depth sensing device, characterized in that, The system includes a ring-shaped fixed bracket on which multiple global reconstruction auxiliary cameras are fixed at equal intervals along the circumference. A global synthesis camera is fixed between every two adjacent global reconstruction auxiliary cameras. At the position of each global synthesis camera, an array camera fixed bracket is also fixed along the axial direction of the ring-shaped fixed bracket. The middle part of each array camera fixed bracket is fixed to the ring-shaped fixed bracket, and a local camera is fixed at each end of each array camera fixed bracket. The two local cameras on each array camera fixed bracket and the corresponding global synthesis camera on that array camera fixed bracket constitute a set of array cameras. Multiple global reconstruction auxiliary cameras and multiple global synthesis cameras respectively cover multiple peripheral areas around the annular fixed support. Adjacent global reconstruction auxiliary cameras and global synthesis cameras have overlapping perception areas. The vertical field of view of each global synthesis camera is greater than or equal to twice the vertical field of view of the local camera. The horizontal field of view of each global reconstruction auxiliary camera and the horizontal field of view of each global synthesis camera are greater than or equal to 60°; The field-of-view optical axes of all global reconstruction auxiliary cameras and all global synthesis cameras are coplanar with the annular fixing bracket; the backward extensions of the field-of-view optical axes of all global reconstruction auxiliary cameras and all global synthesis cameras pass through the center of the annular fixing bracket.
2. The multi-viewpoint ring-shaped heterogeneous panoramic depth sensing device according to claim 1, characterized in that, Nine global reconstruction auxiliary cameras are fixed at equal intervals on the annular fixed bracket, and there are a total of nine global synthesis cameras.
3. The multi-viewpoint-based ring-shaped heterogeneous panoramic depth sensing device according to claim 2, characterized in that, The perceived image resolution of each global synthetic camera and local camera is greater than or equal to 14 megapixels.
4. The multi-viewpoint ring-shaped heterogeneous panoramic depth sensing device according to claim 1, characterized in that, All global reconstruction auxiliary cameras and all global compositing cameras have the same model number.
5. A multi-viewpoint ring-shaped heterogeneous panoramic 3D reconstruction method, characterized in that, include: Step 1: Construct the ring-shaped heterogeneous panoramic depth perception device as described in claim 3, perform visual calibration on the global cameras therein, obtain the intrinsic parameters of 18 global cameras and the extrinsic parameter matrices of adjacent global cameras, and calculate the extrinsic parameter matrices between all global synthetic cameras based on the global reconstruction auxiliary camera. Step 2: Based on the intrinsic parameters of the 18 global cameras and the extrinsic parameter matrices of the adjacent global cameras, calculate the dense depth map from the perspective of each global synthetic camera. Step 3: Image acquisition is performed using the constructed ring-shaped heterogeneous panoramic depth sensing device. Nine sets of array camera data are obtained at the same time stamp, including nine global composite camera images and 18 local camera images. The deformation matrix between each pair of adjacent images in the nine global composite camera images is calculated. The obtained deformation matrix is then used to fuse the panoramic image and dense depth map through a neural network to obtain the panoramic image. And panoramic depth perception results .
6. The multi-viewpoint ring-shaped heterogeneous panoramic 3D reconstruction method according to claim 5, characterized in that, Step one specifically includes: Construct the ring-shaped heterogeneous panoramic depth sensing device as described in claim 3, and number the 18 global cameras therein sequentially from 1 to 18, with odd-numbered cameras being global reconstruction auxiliary cameras among the global cameras, and even-numbered cameras being global synthesis cameras among the global cameras. Based on Zhang Zhengyou's calibration method, a checkerboard pattern was used to visually calibrate the global cameras and adjacent global cameras in a ring-shaped heterogeneous panoramic depth sensing device, obtaining the intrinsic parameters of 18 global cameras. , , , , , ,…, , },in, (n=1,3,5…,17) represents the intrinsic parameters of the global reconstruction auxiliary camera in the global camera. (n=2,4,6…,18) represents the intrinsic parameters of the global composite camera in the global camera; and the extrinsic parameter matrix of the adjacent global cameras is obtained. Each extrinsic parameter matrix contains a rotation matrix. and a translation vector ; Take 2, 4, 6...16 in sequence, for each Perform the following operations: according to Global Synthetic Camera No. 1 and External parameters between +1 global reconstruction auxiliary camera ,as well as +1 global reconstruction auxiliary camera and Extrinsic parameters between +2 global synthetic cameras Calculations yielded Global Synthetic Camera No. 1 and Extrinsic parameters between +2 global synthetic cameras ,in: ; This allows us to calculate the extrinsic matrix between all global synthetic cameras.
7. The multi-viewpoint ring-shaped heterogeneous panoramic 3D reconstruction method according to claim 6, characterized in that, Step two specifically includes: Will -1 Global Reconstruction Auxiliary Camera and The global synthetic camera is used as a stereo camera pair, based on -1 Global Reconstruction Auxiliary Camera and Intrinsic and extraterrestrial parameters of the global synthetic camera , and ,as well as -1 Global Reconstruction Auxiliary Camera and Stereo images composed of a global synthetic camera and Binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation to obtain... First dense depth map from the perspective of the global synthetic camera. ; Will Global Composite Camera and +1 global reconstruction auxiliary camera serves as a stereo camera pair, based on Global Composite Camera and +1 Global Reconstruction Auxiliary Camera's Intrinsic and Extrinsic Parameters , ,and ,as well as Global Composite Camera and The stereo images composed of the +1 global reconstruction auxiliary camera and Binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation to obtain... Second dense depth map from the perspective of the global synthetic camera. ; Get -1 Global Reconstruction Auxiliary Camera and A stereo vision pair composed of two global synthetic cameras, with a mask of the effective correction area obtained during stereo correction. and obtaining Global Composite Camera and The stereo vision pair consisting of the +1 global reconstruction auxiliary camera, and the masking of the effective correction area obtained during stereo correction. In the mask layer, a pixel index value of 1 represents a valid value, and 0 represents an invalid value. According to the mask and mask For the first dense depth map Second Dense Depth Map Merging is performed to obtain a merged dense depth map. The merging principle is as follows, for pixel index : ; right Perform bilateral filtering to obtain the final result. Dense Depth Map from the Perspective of the Global Synthetic Camera ; Take 2, 4, 6..., 16 in sequence, for each... Perform the above operations; Choosing 18, we treat the global reconstruction auxiliary camera (17) and the global synthesis camera (18) as a stereo camera pair, based on the intrinsic and extrinsic parameters of the global reconstruction auxiliary camera (17) and the global synthesis camera (18). , and The stereo image pairs formed by the 17th global reconstruction auxiliary camera and the 18th global synthesis camera. and A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain the first dense depth map from the perspective of the 18th global synthetic camera. ; Using global composite camera 18 and global reconstruction auxiliary camera 1 as a stereo camera pair, based on the intrinsic and extrinsic parameters of global composite camera 18 and global reconstruction auxiliary camera 1... , ,and The stereo image pairs composed of the No. 18 global synthetic camera and the No. 1 global reconstruction auxiliary camera. and A binocular stereo depth calculation is performed, which includes distortion correction, stereo correction, disparity solving, and depth calculation, to obtain a second dense depth map from the perspective of the 18th global synthetic camera. ; Obtain the stereo vision pair consisting of global reconstruction auxiliary camera 17 and global synthesis camera 18, and obtain the mask of the effective correction area during stereo correction. And to obtain the stereo vision pair consisting of global synthesis camera No. 18 and global reconstruction auxiliary camera No. 1, and the mask of the effective correction area obtained during the stereo correction process. In the mask layer, a pixel index value of 1 represents a valid value, and 0 represents an invalid value. According to the mask and mask For the first dense depth map Second Dense Depth Map Merging is performed to obtain a merged dense depth map. The merging principle is as follows, for pixel index : ; right Bilateral filtering was performed to obtain the final dense depth map from the perspective of global synthetic camera #18. .
8. The multi-viewpoint circular heterogeneous panoramic 3D reconstruction method according to claim 7, characterized in that, Step three specifically includes: Using a constructed ring-shaped heterogeneous panoramic depth sensing device, image data was acquired using either external synchronization triggering or software timestamp synchronization methods. At the same timestamp, nine sets of array camera data were obtained, including nine globally synthesized camera images captured by the global synthetic camera. , … , and 18 partial camera images taken by a partial camera , , , ... , ;in , , Images from a set of array cameras; At the same timestamp, for a set of globally synthesized camera image sequences , … Feature detection algorithms are used to detect features in adjacent images; For a pair of adjacent images in a sequence of globally synthesized camera images, image registration is performed based on the detected feature points to obtain the deformation matrix T between each pair of adjacent images in the globally synthesized camera image sequence. T ... T ; Using the obtained deformation matrix, the 9 globally synthesized images , … The images are stitched together to obtain a 360-degree panoramic image. ; Using the obtained deformation matrix, the dense depth map from the global synthetic camera viewpoint is... , ... The images are stitched together to obtain a 360-degree panoramic depth perception result. ; Using a cross-scale image fusion network, 18 local camera images are fused together. In the process, a panoramic image with a resolution of hundreds of millions of pixels was obtained. ; Using super-resolution methods, for To improve the resolution, obtain the same as Panoramic depth perception results at the same resolution with a resolution of hundreds of millions of pixels .
9. The multi-viewpoint circular heterogeneous panoramic 3D reconstruction method according to claim 8, characterized in that, Image registration is performed using a random consistency sampling algorithm; the cross-scale image fusion network is a network such as CrossNet.
Citation Information
Patent Citations
External parameter calibration method for multi-camera system of unmanned vehicle
CN109360245A
One billion-pixel virtual reality video collection device, system and method
CN111343367A