Depth estimation method and device based on multi-view fisheye camera, equipment and medium

Through the multi-view fisheye physic mechanism, the Cassini projection stereo pair is constructed to solve the problem of insufficient accuracy and efficiency of depth estimation in the prior art, and achieve a comprehensive depth estimation with high accuracy, high efficiency and strong generalization capabilities.

CN120182347AActive Publication Date: 2025-06-20GUANGXI UNIV

Patent Information

Application Number
CN202510243886.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

The existing multi-view fisheye camera depth estimation method based on convolutional neural networks is insufficient in terms of accuracy and computing efficiency, especially in complex environments and unknown scenarios, with poor generalization capabilities.

Method used

The 360-degree panoramic image was obtained through multiple fisheye cameras, and the stereo pairs of Cassini projection were constructed to match the stereoscopic dimensions, and the parallax map and confidence map were obtained, converted into depth maps and fused. Finally, the fusion depth map was refined to improve accuracy and efficiency.

Benefits of technology

It improves the accuracy and efficiency of depth estimation, enhances the generalization ability in complex environments and unknown scenarios, can more accurately capture environmental depth details and achieve real-time all-round depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182347A_ABST
    Figure CN120182347A_ABST
Patent Text Reader

Abstract

The invention discloses a depth estimation method and device based on a multi-view fisheye camera, equipment and a medium, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a fisheye image covering a 360-degree panorama through a plurality of fisheye cameras; constructing a three-dimensional pair of Karisi projection in a view overlapping area in fisheye images shot by two adjacent fisheye cameras; performing stereo matching on each stereo pair to obtain a disparity map and a confidence map of each stereo pair; converting each disparity map into a depth map; converting each confidence map and each depth map from Karisy projection to equidistant columnar projection, and fusing each depth map according to each confidence map to obtain a fused omnidirectional depth map; and carrying out refining processing on the fused omnidirectional depth map to eliminate a difference error generated by conversion from the Karzeni projection to the equidistant columnar projection and a detail error generated during fusion so as to obtain a target omnidirectional depth map. According to the method, the accuracy and the calculation efficiency are improved by constructing the three-dimensional pair of the Karisi projection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and particularly to a depth estimation method, device, equipment and medium based on a multi-view fisheye camera. Background Art

[0002] In many application scenarios of navigation, fast and reliable omnidirectional 3D perception is crucial. Currently, the main methods for obtaining omnidirectional 3D information are lidar solutions and stereo camera solutions. The lidar solution has problems such as high price and large volume, and the generated depth map is relatively sparse and lacks detailed information. The stereo camera solution has problems such as an increase in the weight, size, and cost of the system due to an increase in the number of cameras, and there are blind spots in the field of view between the lenses. In contrast, using a relatively small number of fisheye cameras for omnidirectional depth estimation can significantly reduce costs and weight, which is a better solution.

[0003] In the field of omnidirectional depth estimation based on multi-view fisheye cameras, many methods have been proposed. Existing methods based on convolutional neural networks have made certain progress in improving accuracy and robustness, but there are still many deficiencies. For example, some methods have poor accuracy when calculating omnidirectional depth, have problems such as uneven stratification, and the generalization ability of the model in unknown scenarios is poor, and it relies heavily on the minimum depth prior information. In terms of computational efficiency, most existing methods usually need to construct a 4D cost volume containing a large amount of redundant information or learn complex epipolar geometry, resulting in low efficiency. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a depth estimation method, device, equipment and medium based on a multi-view fisheye camera to improve the accuracy and efficiency of depth estimation.

[0005] To achieve the above object, on the one hand, an embodiment of the present application proposes a depth estimation method based on a multi-view fisheye camera, and the method includes the following steps:

[0006] Obtain fisheye images corresponding to respective viewpoints through multiple fisheye cameras; wherein, the field of view angles of the multiple fisheye cameras cover a 360-degree panorama;

[0007] Construct a stereo pair of Cassini projections in the overlapping field of view regions of the fisheye images captured by two adjacent ones of the fisheye cameras;

[0008] Perform stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map for each of the stereo pairs;

[0009] Convert each of the disparity maps into a depth map as a first depth map;

[0010] Convert each of the first confidence maps and each of the first depth maps from the Cassini projection to the equidistant cylindrical projection, and then fuse each of the first depth maps according to each of the first confidence maps to obtain a fused omnidirectional depth map;

[0011] Perform refinement processing on the fused omnidirectional depth map to eliminate the difference error caused by the conversion from the Cassini projection to the equidistant cylindrical projection and the detail error generated during fusion, and obtain a target omnidirectional depth map.

[0012] In some embodiments, constructing a stereoscopic pair of Cassini projections in the overlapping field of view regions of the fisheye images captured by two adjacent fisheye cameras includes the following steps:

[0013] Project each of the fisheye images onto the unit sphere in the Cartesian coordinate system using the fisheye camera model, then convert the projections of each of the fisheye images from the Cartesian coordinate system to the Cassini spherical coordinate system, and further map each of the fisheye images onto a Cassini projection image;

[0014] Construct the stereoscopic pair of Cassini projections on the Cassini projection image corresponding to the overlapping field of view region;

[0015] Wherein, define the Cassini spherical coordinate system as (ρ, φ, θ), where ρ represents the distance between the coordinate origin O and the point P, is the angle between the straight line OP and the plane yOz, and θ is the angle between the straight line OP' and the positive Z-axis, where the point P' is the projection of the point P on the plane yOz; the coordinate (x, y, z) of the Cartesian coordinate system and the coordinate of the Cassini spherical coordinate system have the following conversion relationship:

[0016]

[0017] In the Cassini projection, define the point P l and the point P r The angular parallax d = φ l - φ r , the baseline length is B, and the depth ρ l from the point P to the origin O l is calculated by the following formula:

[0018]

[0019] In some embodiments, performing stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map for each of the stereo pairs includes the following steps:

[0020] Construct a cost volume for each of the stereo pairs according to the ACVNet method;

[0021] Regularize the relevant stereo matching volume and the cost volume respectively using an hourglass network;

[0022] Upsample the regularized cost volume to obtain the disparity map and the first confidence map for each of the stereo pairs.

[0023] In some embodiments, the fusing each of the first depth maps according to each of the first confidence maps to obtain a fused omnidirectional depth map includes the following steps:

[0024] Align each of the first depth maps and each of the first confidence maps from different perspectives to the same direction and position through rotation and translation to obtain aligned second confidence maps and second depth maps;

[0025] Fuse each of the second depth maps into an initial omnidirectional depth map according to the confidence information of the second confidence maps;

[0026] Use a lightweight fusion network with an encoder-decoder architecture and skip connections to adjust the depth values of the initial omnidirectional depth map according to the adjacent depth information at each position to obtain the fused omnidirectional depth map.

[0027] In some embodiments, the refining the fused omnidirectional depth map to eliminate the difference error generated from the conversion from the Cassini projection to the equidistant cylindrical projection and the detail error generated during fusion to obtain a target omnidirectional depth map includes the following steps:

[0028] Sample each position of the fused omnidirectional depth map within a depth neighborhood to generate depth hypotheses;

[0029] Map the feature maps of each of the fisheye images to the equidistant cylindrical projection using the fisheye camera imaging model and the depth hypotheses, and then construct a 4D cost volume;

[0030] Perform cost regularization on the 4D cost volume through an hourglass network, and then use the softargmin function to obtain the target omnidirectional depth map according to the regularized 4D cost volume.

[0031] In some embodiments, the method further includes the following steps:

[0032] Generate navigation information according to the target omnidirectional depth map.

[0033] In some embodiments, the obtaining the fisheye images corresponding to the perspectives through multiple fisheye cameras includes the following steps:

[0034] Obtain the fisheye images corresponding to the perspectives through multiple of the fisheye cameras disposed on a vehicle, a robot, or a drone.

[0035] To achieve the above object, on the other hand, an embodiment of the present application provides a depth estimation device based on a multi-view fisheye camera, the device comprising:

[0036] An image acquisition unit, configured to acquire fisheye images corresponding to respective perspectives through multiple fisheye cameras; wherein, the field of view angles of the multiple fisheye cameras cover a 360-degree panorama;

[0037] A stereo pair construction unit, configured to construct a stereo pair of Cassini projections in an overlapping field of view region in the fisheye images captured by two adjacent ones of the fisheye cameras;

[0038] A stereo matching unit, configured to perform stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map for each of the stereo pairs;

[0039] A disparity conversion unit, configured to convert each of the disparity maps into a depth map as a first depth map;

[0040] A depth fusion unit, configured to convert each of the first confidence maps and each of the first depth maps from the Cassini projection to an equidistant cylindrical projection, and then fuse each of the first depth maps according to each of the first confidence maps to obtain a fused omnidirectional depth map;

[0041] An error processing unit, configured to perform refinement processing on the fused omnidirectional depth map to eliminate the difference error generated from the conversion from the Cassini projection to the equidistant cylindrical projection and the detail error generated during fusion, so as to obtain a target omnidirectional depth map.

[0042] To achieve the above object, on the other hand, an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned depth estimation method based on a multi-view fisheye camera when executing the computer program.

[0043] To achieve the above object, on the other hand, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, and the computer program implementing the above-mentioned depth estimation method based on a multi-view fisheye camera when executed by a processor.

[0044] The embodiments of the present application at least include the following beneficial effects:

[0045] This application can obtain fisheye images corresponding to different perspectives through multiple fisheye cameras; among them, the field of view angles of multiple fisheye cameras cover a 360-degree panorama; construct a stereoscopic pair of Cassini projections in the overlapping field of view area of the fisheye images captured by two adjacent fisheye cameras; perform stereo matching on each stereoscopic pair to obtain the disparity map and the first confidence map of each stereoscopic pair; convert each disparity map into a depth map as the first depth map; convert each first confidence map and each first depth map from Cassini projection to equidistant cylindrical projection, and then fuse each first depth map according to each first confidence map to obtain a fused omnidirectional depth map; perform refinement processing on the fused omnidirectional depth map to eliminate the difference error generated by converting from Cassini projection to equidistant cylindrical projection and the detail error generated during fusion, so as to obtain the target omnidirectional depth map. This application projects each fisheye image onto a specific plane by constructing a stereoscopic pair of Cassini projections, making the epipolar lines of adjacent views arranged horizontally after projection, thus simplifying the photometric matching process, reducing the computational difficulty of depth estimation, and enabling subsequent stereo matching to quickly and accurately obtain the depth map from the stereoscopic pair of Cassini projections, thereby improving the accuracy and computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 It is a schematic flowchart of the depth estimation method based on multi-perspective fisheye cameras provided by the embodiments of the present application;

[0048] Figure 2 It is a schematic diagram of arranging four fisheye cameras on the top of a vehicle provided by the embodiments of the present application;

[0049] Figure 3 It is an exemplary flowchart of the depth estimation method based on multi-perspective fisheye cameras provided by the embodiments of the present application;

[0050] Figure 4 It is an exemplary diagram of constructing a stereoscopic pair of Cassini projections from fisheye images provided by the embodiments of the present application;

[0051] Figure 5 It is a schematic diagram of the result of depth estimation according to fisheye images provided by the embodiments of the present application;

[0052] Figure 6 It is a schematic structural diagram of the depth estimation device based on multi-perspective fisheye cameras provided by the embodiments of the present application;

[0053] Figure 7 This is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0054] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the embodiments of the present application. They are only examples of devices and methods that are consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0055] It can be understood that the terms "first", "second", etc. used in the present application can be used in this document to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "while...", or "in response to determining".

[0056] The terms "at least one", "multiple", "each", "any one", etc. used in the present application, at least one includes one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any one refers to any one of the multiple.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0058] Before the detailed description of the embodiments of the present application, some terms and related technologies involved in the embodiments of the present application are described as follows:

[0059] 1. Omnidirectional depth estimation:

[0060] It refers to using multiple cameras to obtain the depth information of the surrounding 360°, realizing the three-dimensional geometric perception of the entire scene, and having important applications in fields such as autonomous driving, robot navigation, and drone navigation.

[0061] 2. Field of View (FOV):

[0062] 3. It refers to the angular measurement of the scene range that a camera can capture. In the multi-view fisheye camera system of this application, the field of view determines the size of the area that each camera can cover. Fisheye cameras usually have a large field of view and can capture more extensive scene information.

[0063] 4. Cassini projection:

[0064] A spherical projection method, which is used in this application to project fisheye images onto a specific plane, so that the epipolar lines of adjacent views are arranged horizontally after projection, thereby simplifying the photometric matching process and reducing the computational difficulty of depth estimation. It is one of the key solutions for efficient depth estimation in this application.

[0065] 5. Equirectangular Projection (ERP):

[0066] A projection method that maps the surface of a sphere onto a plane. In omnidirectional depth estimation, fisheye images or other panoramic images are often converted to equirectangular projection for processing and analysis. In this application, the depth map needs to be converted to equirectangular projection at a certain stage for multi-view fusion and other operations.

[0067] 6. Stereo matching:

[0068] In computer vision, it is the process of calculating depth information by analyzing the relationship between corresponding pixel points in images from different perspectives and calculating parallax information. In this application, stereo matching is one of the core links of depth estimation. By constructing an efficient cost volume and combining a specific network structure and algorithm, depth information can be quickly and accurately estimated from the Cassini projection stereo pair.

[0069] The existing technologies mainly have the following defects: First, the depth estimation accuracy is insufficient. Whether it is the method based on spherical scanning or epipolar line reconstruction, it is difficult to accurately capture the depth details of the environment in complex and unfamiliar scenes. The spherical hierarchical scanning method is limited by the number and uniformity of depth layers and is sensitive to prior information such as the minimum depth. The epipolar line reconstruction method has problems such as mapping errors and depth discontinuities. Second, the computational efficiency is low. Most methods have a high computational complexity when processing multi-view fisheye images and cannot meet the real-time requirements, seriously restricting the feasibility in practical applications. Third, the model generalization ability is weak. The performance drops significantly when facing unknown scenes and it is difficult to be reliably applied to different environments. This application aims to solve the above problems and achieve high-precision, high-efficiency and strong generalization ability of full-depth estimation through innovative technical solutions.

[0070] The embodiments of the present application provide a depth estimation method, device, equipment and medium based on a multi-view fisheye camera, which relates to the field of computer vision technology. The depth estimation method, device, equipment and medium based on a multi-view fisheye camera provided by the embodiments of the present application can be applied to a terminal, or can be applied to a server, or can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements the depth estimation method based on a multi-view fisheye camera, etc., but is not limited to the above forms.

[0071] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0072] Referring to Figure 1 , the embodiments of the present application provide a depth estimation method based on a multi-view fisheye camera. The method can include but is not limited to S100 to S150, as follows:

[0073] S100: Obtain fisheye images corresponding to the respective viewpoints through multiple fisheye cameras; wherein, the field of view angles of the multiple fisheye cameras cover a 360-degree panorama.

[0074] Exemplarily, Figure 2 is a schematic diagram showing the arrangement of four fisheye cameras on the top of a vehicle provided by the embodiments of the present application.

[0075] Further, S100 may include the following steps:

[0076] Obtain the fisheye images of corresponding viewpoints through multiple said fisheye cameras disposed on a vehicle, a robot, or a drone.

[0077] S110: Construct a stereo pair of Cassini projection in the field-of-view overlapping region of the fisheye images captured by two adjacent said fisheye cameras.

[0078] Further, S110 may include the following steps S111 to S112:

[0079] S111: Project each of the fisheye images onto the unit sphere in the Cartesian coordinate system using a fisheye camera model, then convert the projection of each fisheye image from the Cartesian coordinate system to the Cassini spherical coordinate system, and further map each fisheye image onto a Cassini projection image;

[0080] S112: Construct the stereo pair of the Cassini projection on the Cassini projection image corresponding to the field-of-view overlapping region;

[0081] Wherein, define the Cassini spherical coordinate system as (ρ, φ, θ), where ρ represents the distance between the coordinate origin O and the point P, is the angle between the straight line OP and the plane yOz, θ is the angle between the straight line OP' and the positive Z-axis, where the point P' is the projection of the point P on the plane yOz; the coordinates (x, y, z) of the Cartesian coordinate system and the coordinates of the Cassini spherical coordinate system have the following conversion relationship:

[0082]

[0083] In the Cassini projection, define the point P l and the point P r The angular parallax d = φ l - φ r , the baseline length is B, and the depth ρ l from the point P l to the origin O is calculated by the following formula:

[0084]

[0085] S120: Perform stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map for each of the stereo pairs.

[0086] Further, S120 may include the following steps S121 to S123:

[0087] S121: Construct a cost volume for each of the stereo pairs according to the ACVNet method;

[0088] S122: Regularize the relevant stereo matching volume and the cost volume respectively using an hourglass network;

[0089] S123: Upsample the regularized cost volume to obtain the disparity maps and the first confidence maps of each of the stereo pairs.

[0090] S130: Convert each of the disparity maps into a depth map as the first depth map.

[0091] S140: Convert each of the first confidence maps and each of the first depth maps from the Cassini projection to the equidistant cylindrical projection, and then fuse each of the first depth maps according to each of the first confidence maps to obtain a fused omnidirectional depth map.

[0092] Further, fusing each of the first depth maps according to each of the first confidence maps in S140 to obtain a fused omnidirectional depth map may include the following steps S141 - S143:

[0093] S141: Align each of the first depth maps and each of the first confidence maps from different perspectives to the same direction and position through rotation and translation to obtain aligned second confidence maps and second depth maps;

[0094] S142: Fuse each of the second depth maps into an initial omnidirectional depth map according to the confidence information of the second confidence maps;

[0095] S143: Use a lightweight fusion network with an encoder - decoder architecture and skip connections to adjust the depth values of the initial omnidirectional depth map according to the adjacent depth information at each position to obtain the fused omnidirectional depth map.

[0096] S150: Refine the fused omnidirectional depth map to eliminate the difference error generated from the conversion from the Cassini projection to the equidistant cylindrical projection and the detail error generated during fusion to obtain a target omnidirectional depth map.

[0097] Further, S150 may include the following steps S151 - S153:

[0098] S151: Sample at each position of the fused omnidirectional depth map within a depth neighborhood to generate depth hypotheses;

[0099] S152: Map the feature maps of each of the fisheye images to the equidistant cylindrical projection using the fisheye camera imaging model and the depth hypotheses, and then construct a 4D cost volume;

[0100] S153: Cost-regularize the 4D cost volume through an hourglass network, and then use the softargmin function to obtain the target omnidirectional depth map based on the regularized 4D cost volume.

[0101] As an alternative implementation, after S150, the embodiments of the present application may further include the following steps:

[0102] S160: Generate navigation information based on the target omnidirectional depth map.

[0103] It can be understood that, in this embodiment, accurate navigation information can be quickly generated based on the target omnidirectional depth map, and thus accurate navigation information can be provided for scenarios such as autonomous driving, robot navigation, or drone navigation.

[0104] Next, the solutions of the embodiments of the present application will be introduced and described in detail with specific application examples.

[0105] Specifically, referring to Figure 3 , this embodiment may include the following solutions:

[0106] 1. Construct a Cassini projection stereo pair.

[0107] Referring to Figure 4 , in this embodiment, the Cassini spherical projection is used to construct a stereo pair in the overlapping field of view of two adjacent fisheye cameras. The epipolar lines of the Cassini stereo pair are horizontally arranged after projection, which can simplify the calculation process of stereo matching and improve the efficiency and accuracy of depth estimation. It should be noted that the step of constructing the stereo pair of Cassini projection in this embodiment can be replaced by constructing the stereo pair of pinhole projection.

[0108] In specific implementation, first project the fisheye image onto the unit sphere in the Cartesian coordinate system using the fisheye camera model, then convert these Cartesian coordinates to Cassini spherical coordinates, and finally map them onto the Cassini projection image. Define the Cassini spherical coordinate system (ρ, φ, θ), where ρ represents the distance between the coordinate origin o and point P, is the angle between the straight line OP and the plane yOz, and θ is the angle between the straight line OP' and the positive Z-axis, where point P' is the projection of point P on the plane yOz. The conversion relationship between the Cartesian coordinates (x, y, z) and the Cassini spherical coordinates is as follows:

[0109]

[0110] In the Cassini projection, define the angular disparity d = φ l and point P r as φ l - φ r , the baseline length is B, and point Pl The depth ρ to the origin O l can be calculated by the following formula:

[0111]

[0112] In this embodiment, a Cassini stereo pair is constructed through precise coordinate transformation and projection operations, laying a foundation for subsequent stereo matching and depth estimation.

[0113] 2. Lightweight stereo matching.

[0114] In this embodiment, a lightweight stereo matching network is designed, which only contains 1.8M parameters. This network adopts a compact U-Net structure and can quickly extract image features. Different from the stereo matching method that directly constructs a cost volume with high computational cost and a large amount of redundant information, this embodiment constructs a compact and efficient cost volume according to the ACVNet method, effectively reducing the computational overhead. Among them, a small hourglass network is used to regularize the correlation volume and the cost volume respectively. Then, CUBG is applied to upsample the cost volume to obtain the disparity map and the confidence map C of each stereo pair. c Finally, the disparity map is converted into a depth map D c as the preliminary result of depth estimation.

[0115] 3. Multi-view depth fusion.

[0116] In this embodiment, a multi-view depth fusion module guided by confidence information is proposed to solve the depth discontinuity in the omnidirectional depth map. First, the depth maps D c and the confidence maps C c are converted from the Cassini projection to the equidistant cylindrical projection, and then the depth maps D c and the confidence maps C c from multiple views are aligned to the same direction and position through rotation and translation to obtain the aligned confidence maps C u and the depth maps D u . Guided by the confidence information, the depth maps from multiple views are fused into the initial omnidirectional depth map D I .

[0117] The initial omnidirectional depth map D I may have invalid pixels and noise due to alignment. In the second step of this module, a lightweight fusion network with an encoder-decoder architecture and skip connections is used to adjust the depth value according to the adjacent depth information at each position to obtain the fused omnidirectional depth map D F , further improving the accuracy of the omnidirectional depth map.

[0118] 4. Omnidirectional depth map refinement.

[0119] This embodiment proposes an omnidirectional depth refinement module to solve the interpolation error generated by correcting a fisheye image to Cassini projection and then to equidistant cylindrical projection, as well as the problem of detail blurring caused by the smoothing process of the fusion module.

[0120] This module samples each position of the fused omnidirectional depth map D F in the depth neighborhood to generate a depth hypothesis d h . Using the fisheye camera imaging model and the depth hypothesis d h , the feature map of the fisheye image is mapped to the equidistant cylindrical projection to construct a small 4D cost volume. Cost regularization is performed through a lightweight hourglass network, and the refined target omnidirectional depth map D R is obtained using the softargmin function. This module effectively restores the detail information of the omnidirectional depth map and improves the accuracy of the overall depth estimation. Exemplarily, Figure 5 is a schematic diagram of the result of depth estimation based on a fisheye image as an example.

[0121] In summary, this embodiment includes the following technical solutions:

[0122] 1. A lightweight stereo matching network that significantly reduces the number of parameters and computational overhead while ensuring a certain accuracy, and quickly generates a depth map and a confidence map.

[0123] 2. Multi-view depth fusion, which uses confidence information to fuse depth maps from multiple views, effectively solves the problem of depth discontinuity caused by directly stitching depth maps from different views, and handles the invalid points and noise problems of the omnidirectional depth map.

[0124] 3. Omnidirectional depth refinement, which enhances the detail information in the omnidirectional depth map through depth neighborhood sampling and reprojection, and improves the accuracy of the final omnidirectional depth estimation result.

[0125] Beneficial effects:

[0126] High depth estimation accuracy: Experimental results on multiple simulation datasets (such as Urban, OmniHouse, OmniThings) show that this embodiment is superior to other existing open-source methods in various evaluation metrics (including MAE, RMSE, SILog, δ1, δ2, δ3). This embodiment can capture environmental depth details more accurately and provide more reliable depth information for practical applications.

[0127] Fast depth estimation speed: This embodiment realizes real-time omnidirectional depth estimation. When outputting an omnidirectional depth map at a resolution of 640*320, on a high-performance device, the RTX TITAN GPU, a frame rate of 35 FPS can be achieved; on a performance-limited embedded device, the Jetson AGX Orin, a frame rate of 12.3 FPS can also be realized, which has an obvious speed advantage compared with other methods. This embodiment has a significant speed improvement on different datasets, while maintaining higher accuracy, effectively solving the problem of low computational efficiency of existing methods and being more suitable for real-time application scenarios.

[0128] Excellent generalization ability: When this embodiment is tested on the un-trained OmniTown dataset, it can still maintain high accuracy. Compared with other methods, the performance significantly drops in new scenarios. This embodiment can better adapt to different environments, reduce the dependence on specific scenarios, broaden its application scope, and has stronger practicality.

[0129] Next, more specific implementation manners will be described.

[0130] Specifically, in an actual application scenario, such as the environmental perception system of an autonomous vehicle, the method of this embodiment can be deployed to an in-vehicle computing platform. Four fisheye cameras are installed on the top of the vehicle, and the camera configuration can still refer to Figure 2 and are arranged in a specific layout to ensure that the surrounding environment can be covered omnidirectionally. In terms of hardware, this embodiment can select an embedded computing device such as Jetson AGX Orin and configure it according to different performance requirements and cost budgets. In terms of software implementation, this embodiment can build the network model of this embodiment based on the PyTorch framework. First, it is initially trained for 40 epochs on the OmniThing dataset, and then fine-tuned for 30 epochs on the mixed dataset of Urban and OmniHouse.

[0131] In actual applications, the fisheye cameras in four directions collect the surrounding environment images in real time and input them into the method model trained in this embodiment. This method model can quickly output a high-precision omnidirectional depth map and provide accurate environmental perception information.

[0132] Referring to Figure 6 , this application embodiment also provides a depth estimation device based on a multi-view fisheye camera, which can implement the above-mentioned depth estimation method based on a multi-view fisheye camera. The device includes:

[0133] An image acquisition unit, configured to acquire fisheye images corresponding to the perspectives through multiple fisheye cameras; wherein, the field of view angles of the multiple fisheye cameras cover a 360-degree panoramic view;

[0134] A stereo pair construction unit for constructing a stereo pair of Cassini projections in the overlapping field of view region in the fisheye images captured by two adjacent fisheye cameras;

[0135] A stereo matching unit for performing stereo matching on each stereo pair to obtain a disparity map and a first confidence map for each stereo pair;

[0136] A disparity conversion unit for converting each disparity map into a depth map as the first depth map;

[0137] A depth fusion unit for converting each first confidence map and each first depth map from the Cassini projection to the equidistant cylindrical projection, and then fusing each first depth map according to each first confidence map to obtain a fused omnidirectional depth map;

[0138] An error processing unit for refining the fused omnidirectional depth map to eliminate the difference error generated from the conversion from the Cassini projection to the equidistant cylindrical projection and the detail error generated during fusion, to obtain a target omnidirectional depth map.

[0139] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0140] An embodiment of the present application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the above-mentioned depth estimation method based on multi-view fisheye cameras. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0141] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented by the device embodiments of the present application are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0142] Please refer to Figure 7 , Figure 7 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0143] A processor 701, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, for executing relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0144] The memory 702 can be implemented in the form of a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 702 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 702 and are called by the processor 701 to execute the depth estimation method based on a multi-view fisheye camera according to the embodiments of this application;

[0145] The input / output interface 703 is used to implement information input and output;

[0146] The communication interface 704 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0147] The bus 705 transmits information between various components of the device (such as the processor 701, the memory 702, the input / output interface 703, and the communication interface 704);

[0148] Among them, the processor 701, the memory 702, the input / output interface 703, and the communication interface 704 achieve communication connections with each other inside the device through the bus 705.

[0149] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned depth estimation method based on a multi-view fisheye camera.

[0150] It can be understood that the content in the above method embodiments is applicable to the embodiments of this storage medium. The functions specifically implemented by the embodiments of this storage medium are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0151] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0152] The embodiments described in the embodiments of the present application are to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0153] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0155] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0156] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0157] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (items) or plural items (items). For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0158] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0159] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0160] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0161] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0162] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall fall within the scope of the rights of the embodiments of this application.

Claims

1. A depth estimation method based on a multi-view fisheye camera, characterized in that: The method comprises the following steps: Acquire fisheye images of corresponding viewing angles through multiple fisheye cameras; wherein the field of view of the multiple fisheye cameras covers a 360-degree panoramic view; Constructing a stereo pair of Cassini projections in the overlapping area of ​​the field of view in the fisheye images taken by two adjacent fisheye cameras; Performing stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map of each of the stereo pairs; Convert each of the disparity maps into a depth map as a first depth map; Convert each of the first confidence maps and each of the first depth maps from the Cassini projection to an equirectangular projection, and then fuse each of the first depth maps according to each of the first confidence maps to obtain a fused omnidirectional depth map; The fused omnidirectional depth map is refined to eliminate the difference error generated by converting from the Cassini projection to the equirectangular projection and the detail error generated during fusion, so as to obtain a target omnidirectional depth map.

2. The depth estimation method based on a multi-view fisheye camera according to claim 1, characterized in that: The method of constructing a stereo pair of Cassini projections in the overlapping area of ​​the field of view in the fisheye images taken by two adjacent fisheye cameras comprises the following steps: Using a fisheye camera model, projecting each of the fisheye images onto a unit sphere in a Cartesian coordinate system, then converting the projection of each of the fisheye images from the Cartesian coordinate system to a Cassini spherical coordinate system, and then mapping each of the fisheye images onto a Cassini projection image; constructing the stereo pair of the Cassini projections on the Cassini projection images corresponding to the overlapping areas of the visual field; The Cassini spherical coordinate system is defined as (ρ, φ, θ), where ρ represents the distance between the origin o and point P. is the angle between the straight line OP and the plane yOz, θ is the angle between the straight line OP′ and the positive Z axis, where point P′ is the projection of point P on the plane yOz; the coordinates (x, y, z) of the Cartesian coordinate system and the coordinates of the Cassini spherical coordinate system The conversion relationship is as follows: In the Cassini projection, point P is defined as l and point P r The angular parallax d = φ l -φ r , the baseline length is B, point P l Depth ρ from the origin O l Calculated by the following formula:

3. The depth estimation method based on a multi-view fisheye camera according to claim 1, characterized in that: The performing stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map of each of the stereo pairs comprises the following steps: Constructing a cost volume for each of the stereo pairs according to the ACVNet method; Regularizing the related stereo matching volume and the cost volume respectively using an hourglass network; The regularized cost volume is upsampled to obtain the disparity map and the first confidence map of each stereo pair.

4. The depth estimation method based on a multi-view fisheye camera according to claim 1, characterized in that: The step of fusing the first depth maps according to the first confidence maps to obtain a fused omnidirectional depth map comprises the following steps: Aligning each of the first depth maps and each of the first confidence maps at different viewing angles to the same direction and position by rotation and translation to obtain an aligned second confidence map and a second depth map; fusing each of the second depth maps into an initial omnidirectional depth map according to confidence information of the second confidence maps; A lightweight fusion network with an encoder-decoder architecture and skip connections is used to adjust the depth value of the initial omnidirectional depth map according to adjacent depth information at each position to obtain the fused omnidirectional depth map.

5. The depth estimation method based on a multi-view fisheye camera according to claim 1, characterized in that: The step of refining the fused omnidirectional depth map to eliminate the difference error generated by converting the Cassini projection to the equirectangular projection and the detail error generated during fusion to obtain a target omnidirectional depth map comprises the following steps: Sampling each position of the fused omnidirectional depth map within a depth neighborhood to generate a depth hypothesis; Mapping the feature maps of each of the fisheye images to an equirectangular projection using a fisheye camera imaging model and the depth hypothesis, thereby constructing a 4D cost volume; The 4D cost volume is cost regularized by an hourglass network, and then the target omnidirectional depth map is obtained according to the regularized 4D cost volume using a softargmin function.

6. The depth estimation method based on a multi-view fisheye camera according to claim 1, characterized in that: The method further comprises the following steps: Navigation information is generated according to the target omnidirectional depth map.

7. The depth estimation method based on a multi-view fisheye camera according to any one of claims 1 to 6, characterized in that: The method of obtaining fisheye images of corresponding viewing angles by using multiple fisheye cameras comprises the following steps: The fisheye images of corresponding viewing angles are acquired by using a plurality of the fisheye cameras arranged on a vehicle, a robot or a drone.

8. A depth estimation device based on a multi-view fisheye camera, characterized in that: The device comprises: An image acquisition unit, used to acquire fisheye images of corresponding viewing angles through multiple fisheye cameras; wherein the field of view of the multiple fisheye cameras covers a 360-degree panoramic view; A stereo pair construction unit, used for constructing a stereo pair of Cassini projection in an overlapping area of ​​the field of view in the fisheye images taken by two adjacent fisheye cameras; A stereo matching unit, configured to perform stereo matching on each of the stereo pairs to obtain a disparity map and a first confidence map of each of the stereo pairs; A disparity conversion unit, configured to convert each of the disparity maps into a depth map as a first depth map; A depth fusion unit, configured to convert each of the first confidence maps and each of the first depth maps from the Cassini projection to an equirectangular projection, and then fuse each of the first depth maps according to each of the first confidence maps to obtain a fused omnidirectional depth map; The error processing unit is used to refine the fused omnidirectional depth map to eliminate the difference error generated by converting from the Cassini projection to the equidistant cylindrical projection and the detail error generated during fusion, so as to obtain a target omnidirectional depth map.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the depth estimation method based on a multi-view fisheye camera as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the depth estimation method based on a multi-view fisheye camera according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and device for generating stereoscopic panoramic film

    CN108520232A

  • Panoramic perception method, device and equipment based on fisheye camera and medium

    CN116579962A

  • Target detection method and device based on fisheye camera, equipment and storage medium

    CN117746390A

  • Target detection method and system based on multi-scale fusion

    CN118628718A

  • Monocular depth estimation method and system

    CN119169066A

Cited By

  • Road obstacle detection method and device based on dual-camera combination and computer readable storage medium

    CN120783316A