A Feature Fusion and Mapping Method for Multi-View Self-Supervised Depth Estimation
Through the feature fusion and mapping method of multi-view self-supervised depth estimation, the depth estimation problem in complex environments such as autonomous driving is solved, and high-precision depth estimation and the effect of reducing training costs is achieved.
Patent Information
- Application Number
- CN202410607489.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-05-16
AI Technical Summary
The prior art is difficult to achieve high-precision depth estimation in complex environments such as autonomous driving, and traditional depth measurement methods are costly and are troubled by sparse data problems.
The feature fusion and mapping method of multi-view self-supervised depth estimation is adopted. By projecting and overlapping the features of multiple viewpoint images, a volume feature map and cylindrical feature map are established, and the self-attention mechanism and depth separation convolution are used for feature encoding and decoding, and a self-supervised loss function is constructed to optimize the model.
It realizes high-precision depth estimation in complex environments such as autonomous driving, reduces dependence on high-quality labeled data, reduces training costs, and improves the performance of depth estimation.
Smart Images

Figure CN118429770B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and specifically relates to a feature fusion and mapping method for multi-view self-supervised depth estimation. Background Art
[0002] Depth perception plays a fundamental role in fields such as autonomous driving, robotics, and augmented reality / virtual reality (AR / VR). Traditional depth measurement methods, such as lidar (Light Detection and Ranging) and time-of-flight (ToF) sensors, are costly and often plagued by sparse data problems. To address these challenges, researchers have begun to explore using deep learning methods to infer depth from camera images, which are not only less costly but also more accurate, and have received increasing attention. However, these methods require a large amount of training on high-quality, densely annotated ground truth depth maps, thus increasing the training cost.
[0003] To solve this problem, some research has started to use monocular temporal images for self-supervised depth estimation. Specifically, for a visual perception system equipped with multiple cameras, in addition to using temporal information, spatial information can also be used for self-supervised learning, and attempts are made to apply self-supervised monocular depth estimation methods by analyzing the characteristics of multi-camera systems. FSM establishes depth supervision using overlapping parts of views, VFDepth introduces a unified volumetric feature representation to fuse multi-view information, and SurroundDepth proposes a cross-view attention mechanism.
[0004] Meanwhile, in autonomous driving and other machine vision applications, obtaining accurate depth information is crucial. Traditionally, some methods have tried to use radar as a sensor for obtaining accurate depth information. However, radar systems often face some inherent challenges, such as data sparsity and high power consumption, which have been difficult to solve for a long time. To overcome these limitations, supervised depth estimation methods have been introduced to directly learn depth information from a large amount of annotated data, which has achieved remarkable success in indoor scenes. However, this method relies on a large number of accurately annotated depth maps, which is difficult to achieve in dynamic external environments such as autonomous driving because the complexity and variability of these scenes greatly increase the difficulty and cost of obtaining annotated data.
[0005] In view of this, self-supervised depth estimation technology has emerged. It does not rely on externally labeled depth maps, but gradually improves the accuracy of depth prediction through self-adjustment of algorithms. This method uses unlabeled images and optimizes the model through a designed loss function to enable it to predict coherent and accurate depth maps. This self-learning ability makes self-supervised depth estimation applicable not only to indoor scenes but also shows great potential in more complex application scenarios such as autonomous driving. Nevertheless, when implementing this technology, it is necessary to pay attention to keeping the computational overhead of the network within a reasonable range to ensure that the real-time processing efficiency will not be affected by an overly large model. Summary of the Invention
[0006] To solve the deficiencies of the prior art and achieve the purpose of maintaining prediction accuracy while controlling the consumption of computing resources, the present invention adopts the following technical solutions:
[0007] A feature fusion and mapping method for multi-view self-supervised depth estimation, comprising the following steps:
[0008] Step S1: Extract the features from the images of each viewpoint Use the intrinsic and extrinsic camera matrices of each view for projection and overlapping processing to establish a volume feature map V t , and use the volume feature map V t Fuse the image features from multiple views to establish a cylindrical feature map F surround_view , and obtain a bird's-eye view (BEV) feature map by flattening the volume feature map V t ;
[0009] Step S2: Use the self-attention mechanism to encode the cylindrical feature map F surround_view Apply natural cylindrical encoding to enhance the correlation between different views and within each view, and obtain the encoded cylindrical feature map F cylinder ;
[0010] Step S3: Depth and pose decoding; Decode the encoded cylindrical feature map F cylinder into the depth map of each view Decode the transformation relationship between different frames (Pose , Pose t-1 , Pose t+1 ) of the view from the bird's-eye view feature map.
[0011] Furthermore, in the step S1, the cylindrical features are divided into overlapping features and non-overlapping features Fuse the overlapping features with reduced dimensions and non-overlapping features to obtain a complete cylindrical feature map F surround_view .
[0012] Furthermore, in the step S1, between the volumetric feature map unit and the volumetric feature unit a mapping relationship is constructed:
[0013]
[0014]
[0015] z = z
[0016] where x, y, and z respectively represent three-dimensional coordinate values, Ox and Oy respectively represent the distances from the origin to the x and y coordinate values, and Backproject(·) represents the mapping function; in the obtained cylindrical map F surround_view contains three-dimensional features enhanced along the ray direction by the MLP network. In the subsequent encoding and depth decoding processes, the present invention can more conveniently utilize the features between the same viewpoint and different viewpoints, thereby avoiding negative optimization caused by different depth distributions between viewpoints and improving the depth estimation performance.
[0017] Furthermore, in the step S2, in order to make full use of the key points in the connected cylindrical feature map F surround_view , the present invention applies a cylindrical feature mapping before the depth decoder and uses a self-attention mechanism for encoding. In order to reduce the consumption of computing resources, the present invention uses depthwise separable convolution DS-Conv to reduce the scale of the cylindrical feature map F surround_view , performs cross-view information exchange by constructing a set (Z) of cross-view self-attention layers, and uses another depthwise separable convolution DS-Conv to enlarge the feature map and restore its original resolution to obtain the encoded cylindrical feature map F cylinder for subsequent feature decoding.
[0018] Furthermore, in the step S3, by performing depth decoding on the encoded cylindrical feature map F cylinder , an unfolded cylindrical depth map is obtained, and the camera parameters are used to project the cylindrical depth map to the corresponding viewpoints to obtain the decoded depth map
[0019] A continuous bird's-eye view feature map is obtained from the flattened volumetric feature maps (V t-1 , V t , V t+1 ) Then the map is compressed along the feature channels into a single feature map that fuses multi-view features, and the standard pose decoder of PoseNet is used to estimate the relative pose of the vision system over time.
[0020] Furthermore, in self-supervised depth estimation, due to the lack of real depth data, the camera viewpoint transformation matrix becomes crucial. The present invention can utilize the spatio-temporal invariance of key points to construct a self-supervised loss function. The depth decoding in step S3 constructs a self-supervised loss function based on features with spatio-temporal constants, maximizes the utilization of multi-view system attributes, and enriches the supervision signal. The formula is as follows:
[0021]
[0022] Among them, represents the reprojection loss, represents the adjacent view loss, represents the cylindrical self-supervised loss, λ r 、λ a 、λ cyl respectively represent the weights of the three losses;
[0023] The reprojection loss at time t i uses the transformation matrix and depth estimation to project the image onto the viewpoint image corresponding to time t j , and constructs the reprojection loss using the weighted sum of intensity difference and structural similarity
[0024] The adjacent view loss projects the overlapping part between viewpoints according to the known transformation matrix between viewpoints and establishes the adjacent view loss
[0025] The cylindrical self-supervised loss When the moving distance is short, the distance between point pairs in the corresponding cylindrical depth map approximately remains constant, and the cylindrical self-supervised loss is constructed based on this
[0026] Furthermore, in the reprojection loss , the inter-frame pose estimation network PoseEstimationNet is used to estimate the transformation matrix (Pose t-1 , Pose t+1 ) between time series frames. A transformation is established according to the known external and internal parameters [R|t] of the multi-vision system, and a transformation matrix for the next moment is created for each viewpoint
[0027]
[0028] Among them, represents the set composed of the camera internal parameters at time i, The set consisting of the camera internal parameters at time j, with subscript t i and t j represent different times.
[0029] Furthermore, the reprojection loss has the following formula:
[0030]
[0031] where represents the time t when the image j is projected, ||·||1 represents the L1 norm, SSIM represents the structural similarity function, and α represents the weight coefficient of the reprojection loss.
[0032] Furthermore, in the adjacent view loss , a transformation is established based on the known external and internal parameters [R|t] of the multi-vision system to create the transformation relationship between adjacent viewpoints for each viewpoint
[0033]
[0034] where K i represents the internal parameter matrix of the camera with serial number i, and K j represents the internal parameter matrix of the camera with serial number j, and [R|t] i represents the external parameter matrix of the camera with serial number i, and [R|t] j represents the external parameter matrix of the camera with serial number j.
[0035] Furthermore, in the cylindrical self-supervised loss , position embeddings are added to the cylindrical feature map F surround_view corresponding to the angle θ and height z indices (P θ , P z ) of the cylindrical coordinates, excluding the P direction because the P direction is the direction where the depth is located. After obtaining the cylindrical feature map F surround_view and decoding the depth to obtain the corresponding cylindrical depth map , based on the known multi-vision system parameters [R|t], it is possible to perform conversion surround_view between the cylindrical depth map D and the depth maps of each viewpoint;
[0036] There are many pairs of feature points in the cylindrical depth map These point pairs are connected by the viewpoint. Here, P represents one point in the point pair, and its connecting line passes through the viewpoint. Therefore, in cylindrical coordinates, the difference in the θ term of the point pair is π, r1 and r2 correspond to the depth of the point pair, and the distance dis between the feature point pairs remains constant between time t i and t j That is, a spatio-temporal invariant; according to θ in the cylindrical coordinate system, the cylindrical depth map is divided into n parts, and a mapping relationship is established. When the moving distance is short, the distance dis in the corresponding part of the cylindrical depth map is an approximately conservative quantity, and based on this, the cylindrical self-supervised loss is constructed:
[0037]
[0038] wherein, represents the average distance between the selected point pairs at time t1, represents the average distance between the selected point pairs at time t2.
[0039] The advantages and beneficial effects of the present invention are as follows:
[0040] The present invention makes full use of the inherent characteristics of different views in the multi-camera system, pays special attention to the relationship between depth and position and perspective, uses the learnable network ML to process the overlapping areas between different perspectives, and at the same time uses the constancy of the spatial distribution of static objects in the multi-camera system to propose the cylindrical self-supervised loss, enriching the content of self-supervised information. The present invention avoids the problems of high cost of traditional depth measurement and sparse data based on deep learning, reduces the dependence on high-quality and densely annotated ground-truth depth maps, thereby reducing the training cost, significantly improving the performance of depth estimation, and achieving an advanced performance level in the self-supervised depth estimation task. Brief Description of the Drawings
[0041] Figure 1 is the overall architecture diagram of the method in the embodiment of the present invention.
[0042] Figure 2a is the comparison diagram of the depth distribution analysis results of different views in the embodiment of the present invention.
[0043] Figure 2b is the comparison diagram of the depth estimation results of different views in the embodiment of the present invention.
[0044] Figure 3 is the flowchart of the method in the embodiment of the present invention.
[0045] Figure 4 is the example diagram of the point pairs for constructing the cylindrical self-supervised loss in the embodiment of the present invention.
[0046] Figure 5 is the visualization diagram of the depth and absolute relative results of the ablation study in the embodiment of the present invention. Detailed implementation manners
[0047] The following will describe in detail the detailed implementation manners of the present invention with reference to the accompanying drawings. It should be understood that the detailed implementation manners described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0048] Currently in the field of autonomous driving, there are several mainstream self-supervised depth estimation technology pipelines, which can be mainly classified into three categories. First is the non-fusion pipeline that treats multiple views as independent inputs. This method directly obtains data from multiple independent views without information fusion between views, and focuses on extracting depth information from each individual view. Second is the attention-based pipeline that uses a transformer for implicit information fusion. This method utilizes the powerful capabilities of the transformer to capture and fuse features from different views, and efficiently integrates multi-source information through the attention mechanism to improve the accuracy and robustness of depth estimation. Finally is the explicit pipeline that uses projection technology to establish a unified feature map. This strategy projects image data from different views into a common feature space, thereby achieving explicit fusion and optimization of features.
[0049] A feature fusion and mapping method for multi-view self-supervised depth estimation according to the present invention combines the advantages of these three technologies, and fuses the strategies of non-fusion, attention-based implicit fusion, and explicit fusion in an innovative way, as Figure 1 shown, the present invention first is the non-fusion pipeline that treats multiple views as independent inputs. This method directly obtains data from multiple independent views without information fusion between views, and focuses on extracting depth information from each individual view. Second is the attention-based pipeline that uses a transformer for implicit information fusion. This method utilizes the powerful capabilities of the transformer to capture and fuse features from different views, and efficiently integrates multi-source information through the attention mechanism to improve the accuracy and robustness of depth estimation. Finally is the explicit pipeline that uses projection technology to establish a unified feature map. This strategy projects image data from different views into a common feature space, thereby achieving explicit fusion and optimization of features. The model design of the present invention can process multi-view data simultaneously without prior fusion, and at the same time introduces an attention mechanism based on the transformer to effectively identify and integrate key visual cues, and finally ensures that the information in all views can be fully utilized within a unified feature framework through advanced explicit fusion technology. This comprehensive method not only improves the accuracy of depth estimation, but also enhances the adaptability and robustness of the model in the face of complex autonomous driving environments.
[0050] As shown Figure 2a , Figure 2b in the figure, the present invention calculates the average depth map of each view in the DDAD dataset and finds that there are significant differences in the depth distributions of different views, as shown by the white bright lines in Figure 2a . Inspired by these inherent characteristics, the present invention designs a CFD pipeline for multi-view self-supervised depth estimation, and Figure 2b the predicted depth map of the present invention is shown in
[0051] Figure 3 The present invention shows a novel feature fusion method, including cylindrical feature fusion, cylindrical feature mapping, and other modules for unsupervised depth estimation, such as a pose estimation network, a depth estimation network, and an unsupervised loss design. In a multi-view system, features are first extracted from the images of each viewpoint and then the overlaps between different views are processed to generate a cylindrical feature map F cylinder . To prevent negative optimization between different views, the present invention uses CVT to encode the cylindrical feature map and then decodes the depth map Meanwhile, the present invention proposes a novel method to construct a self-supervised loss, maximizing the utilization of multi-view system attributes and enriching the supervision signals.
[0052] Furthermore, using Resnet18, feature maps are extracted from the frames of N cameras at time t and considering the correlation between multi-view images, the present invention uses the intrinsic and extrinsic camera matrices [R|t] of each view for projection and overlap processing to establish a volume feature map V t ; subsequently, the present invention uses the volume feature map V t to establish a cylindrical feature map F surround_view and applies natural cylindrical encoding to enhance the correlation between different views and within each view, thereby generating a cylindrical feature map F for depth estimation cylinder ; then, the depth estimation module decodes F cylinder into the depth map of each view The transformation relationship between different frames (Pose t-1 , Pose t+1 ) is decoded from the BEV feature map which is obtained by flattening the volume feature map V t t+1 ; specifically, it includes the following steps:
[0053] 1) First, the present invention constructs cylindrical features. Using Resnet as the backend network, from the original images Extract image features, retain the last two levels to maintain the initial feature fusion within the same viewing point, and use bilinear interpolation to restore its shape. To fuse the features of multiple views, the present invention uses feature projection to establish a general volumetric feature map V t , specifically, receives the projected features from each viewing point through a multi-layer perceptron (MLP). In addition, the present invention processes two types of volumetric feature units through different MLPs to obtain a unified volumetric feature map V according to whether the viewing points overlap t .
[0054] After obtaining the volumetric feature map, it becomes necessary to perform 3D convolution operations or directly decode the depth value from the volumetric features. In addition, directly applying camera parameters for reprojection may result in loss of feature utilization between different viewing points. In addition, the distortion between image coordinates and real-world 3D coordinates may have a negative impact on the depth estimation task. To solve these problems, the present invention designs a Cylin-project Net to fuse image features from multiple perspectives using the volumetric feature map and establish a cylindrical feature map F surround_view . As shown in formula (1), the present invention establishes the mapping relationship between the volumetric feature map unit and the unit in the cylindrical coordinate system
[0055] In this context, the cylindrical features can be divided into overlapping features and non-overlapping features Making full use of the overlapping parts between different viewing points is the key to effectively using a multi-view system. The present invention designs the structure of the Cylin-project Net to process these two types of features. For overlapping features, the present invention uses an MLP to reduce its dimension and then fuses it with non-overlapping features to obtain a complete cylindrical feature map F surround_view for subsequent task processing
[0056]
[0057] At this point, the cylindrical map F obtained by the present invention surround_view contains three-dimensional features enhanced along the ray direction through the MLP network. In the subsequent encoding and depth decoding processes, the present invention can more conveniently utilize the features between the same viewing point and different viewing points, thereby avoiding negative optimization caused by different depth distributions between viewing points and improving the depth estimation performance
[0058] 2) Then, the present invention performs cylindrical feature mapping. To make full use of the key points in the connected cylindrical feature map F surround_view , the present invention applies cylindrical feature mapping before the depth decoder and uses a self-attention mechanism for encoding. The present invention performs operations on the cylindrical feature map F surround_viewAdd positional embeddings corresponding to the angular and height indices of the cylindrical coordinates (P θ , P z ). The P direction is not included because it is the direction where the depth lies.
[0059] Subsequently, to reduce the consumption of computing resources, the present invention uses a single layer of depthwise separable convolution (DS-Conv) to downscale F surround_view . Then, the present invention constructs Z cross-view self-attention layers to perform cross-view information exchange. After that, the present invention uses another DS-Conv to upscale the feature map and restore its original resolution. The feature map F cylinder obtained after this process is used for subsequent feature decoding.
[0060] 3) Finally, the present invention constructs a depth decoder and a pose decoder and proposes a novel loss function. Specifically, the fused and encoded cylindrical feature map F cylinder and the flattened bird's-eye view feature maps from three time points are obtained
[0061] For the depth decoder and the pose decoder
[0062] Due to the need for self-supervised depth estimation, the present invention needs to decode these two types of maps accordingly. For the cylindrical feature map, the present invention constructs a depth decoder to decode it.
[0063] The depth decoder consists of 4 convolutional layers. Through these layers, the present invention performs upsampling and feature channel normalization. Finally, the present invention obtains the unfolded cylindrical depth map. Using the camera parameters the present invention projects it onto the corresponding viewpoints to obtain the final result
[0064] The present invention adopts consecutive bird's-eye view feature maps These maps are obtained from the flattened volumetric feature maps (V t-1 , V t , V t+1 ). Then, these maps are compressed along the feature channels into a single feature map that fuses multi-view features. Next, the present invention uses the standard pose decoder of PoseNet to estimate the relative pose of the vision system over time. In self-supervised depth estimation, due to the lack of ground truth depth data, the camera viewpoint transformation matrices become crucial because they enable the present invention to utilize the spatio-temporal invariance of key points to construct a self-supervised loss function, which is the key to successful training.
[0065] Loss function
[0066] When establishing the loss function, the present invention selects three features with spatio-temporal constants, as shown in Equation (2):
[0067]
[0068] Among them, represents the reprojection loss, represents the adjacent view loss, represents the cylindrical self-supervised loss. The weights of these three loss functions are respectively represented by λ r , λ a , λ cyl .
[0069] Reprojection loss Given the external and internal parameters Rt of the known multi-camera system, a transformation can be established according to Equation (3). The present invention uses the inter-frame pose estimation network PoseEstimationNet to estimate the transformation matrix (Pose t-1 , Pose t+1 ) between the temporal frames. According to the known external and internal parameters [R|t] of the multi-camera system, a transformation as shown in Equation (3) can be established, which enables the creation of a transformation matrix for each viewpoint for the next moment and the transformation relationship between adjacent viewpoints
[0070]
[0071]
[0072] Among them, represents the set composed of the internal parameters of the camera at time i, represents the set composed of the internal parameters of the camera at time j, K i represents the internal parameter matrix of the camera with serial number i, K j represents the internal parameter matrix of the camera with serial number j, [R|t] i represents the external parameter matrix of the camera with serial number i, [R|t] j represents the external parameter matrix of the camera with serial number j.
[0073] The present invention projects the image i using the transformation matrix and depth estimation at time t to the viewpoint image j corresponding to time t . The present invention constructs the reprojection loss using the weighted sum of the intensity difference and the structural similarity
[0074]
[0075] Among them, ||·||1 represents the L1 norm, SSIM represents the structural similarity function, and α represents the weight coefficient of the reprojection loss. of the weight coefficient.
[0076] Adjacent view loss The known transformation matrix between given viewpoints The present invention adopts a calculation method similar to the reprojection loss to project the overlapping part between viewpoints and establish the adjacent view loss
[0077] Cylindrical self-supervised loss After obtaining the cylindrical feature map F surround_view and obtaining the corresponding cylindrical depth map through the depth decoder the present invention can easily convert the depth maps of different viewpoints in a multi-camera system where r corresponds to the image depth.
[0078] Given the known camera parameters [R|t] of the multi-camera system, the present invention can easily convert between the cylindrical depth map D surround_view and the depth maps of each viewpoint As shown, there are many point pairs in the cylindrical depth map Figure 4 These point pairs are connected by viewpoints, where P represents a point in the point pair, and its connection line passes through the viewpoint. Therefore, in the cylindrical coordinates, the difference in their θ terms is π, and r1 and r2 correspond to the depths of the point pair. The distance dis between these feature point pairs remains constant between time t and t i and t j This is a spatio-temporal invariant and can be utilized in a multi-camera system. The cylindrical feature fusion pipeline proposed by the present invention can easily utilize this temporal relationship.
[0079] The present invention divides the cylindrical depth map into n parts (n = 10) according to θ in the cylindrical coordinate system and establishes the Figure 4 shown mapping relationship. As shown in formula (5), when the moving distance is short, the distance dis in the corresponding part of the cylindrical depth map is an approximately conservative quantity. This can be used to establish the cylindrical self-supervised loss:
[0080]
[0081] Among them, represents the average distance between the selected points by the point pair at time t1, represents the average distance between the selected points by the point pair at time t2.
[0082] As Figure 4As shown, we analyze the characteristics of multiple sets of cylindrical point pairs obtained by the intersection of a straight line passing through the origin of the ego-vehicle plane coordinate system and the cylindrical depth map. It is found that the average distance change remains constant and can be used to construct a loss function.
[0083] As Figure 5 shown, the addition of cylindrical feature fusion, overlapping region processing, and cylindrical self-supervised loss improves the performance of depth estimation. Notably, cylindrical feature fusion contributes significantly to the enhancement of the left and right views.
[0084] In the embodiments of the present invention, training and evaluation are performed on commonly used and challenging multi-camera autonomous driving benchmark datasets, including DDAD [2020A] and nuScenes
[2020] . The DDAD [2020A] dataset is collected synchronously in multiple countries, while nuScenes
[2020] includes multiple urban driving scenarios with less overlap between camera views. In the training phase, the present invention obtained approximately 12,252 training samples and 3,850 validation samples from DDAD [2020A], and approximately 20,096 training samples and 5,417 validation samples from nuScenes
[2020] .
[0085] The present invention uses ResNet pre-trained on ImageNet as the backbone network of the present invention and constructs a volumetric feature map with a dimension of (100, 100, 30), and the resolution of each volume unit is (1m, 1m, 0.5m). In addition, the present invention also establishes a cylindrical feature map with a resolution of 480×48. During the self-attention mapping process, the present invention uses 8 attention heads with Z = 8. Due to the difference in the overlapping angle between camera views in the two datasets, the weights of the self-supervised loss L are different between the DDAD [2020A] and nuScenes
[2020] datasets. In the DDAD [2020A] dataset, the weights of the three losses (λ1, λ2, λ3) are (0.85, 0.1, 0.05) respectively. In the nuScenes
[2020] dataset, the weights of the three losses are (0.90, 0.09, 0.01) respectively.
[0086] In the ablation study, CFD_LC, CFD_LL, and CFD_LO represent the baseline network, the network without self-supervised loss, and the network without processing the overlapping part in cylindrical feature fusion respectively. The present invention uses the images at the previous moment t - 1 and the next moment t + 1 as the temporal context. All experiments are performed on 4 Nvidia 4090 GPUs.
[0087] The method of the present invention is compared with current advanced self-supervised depth estimation algorithms on the nuScenes
[2020] and DDAD [2020A] datasets, and is also compared with the baseline and the latest self-supervised multi-camera depth estimation methods, including Monodepth2 [2019a], PackNet-SfM [Guizilinietal., 2020b], FSM
[2022] , VFDepth [2022b], SurroundDepth
[2023] , and EGADepth [2023b]. The visualization results of the DDAD [2020A] and nuScenes
[2020] datasets are shown in the appendix.
[0088] The evaluation results of DDAD [2020A] and nuScenes
[2020] are shown in Table I. For fairness, the training and evaluation settings of all methods are the same.
[0089] Specifically, the resolution of the input images is adjusted to 640×384 and 640×352 on the DDAD [2020A] and nuScenes
[2020] datasets respectively by spline interpolation. The resolution of the feature maps extracted after passing through the backbone network is consistently 11×20×5. The ground truth of depth is generated by projecting the lidar point cloud onto each view, and the evaluation of all metrics is only carried out at the pixel coordinates where there are lidar ground truth points. On the DDAD [2020A] dataset, the method of the present invention has achieved continuous improvement in all metrics, with an enhancement of approximately 5% in (abs_rel, sq_rel, rmse). On the nuScenes
[2020] dataset, the method of the present invention shows continuous improvement, exceeding the baseline method in 4 out of 7 metrics. This indicates that the method of the present invention can continuously enhance the existing baseline.
[0090] Table I is an evaluation of the DDAD dataset and the nuScenes dataset, and the data provided in the table illustrates the comparison between the method of the present invention and the current state-of-the-art methods. In this table, the median accuracy of various evaluation metrics is selected by the present invention to evaluate the difference between the predicted depth map and the depth ground truth.
[0091] Table I
[0092]
[0093] To organize ablation experiments, the present invention established a baseline network architecture that omits the cylindrical feature fusion module and the cylindrical self-supervised loss Lcys. This architecture directly performs back-projection and depth decoding after volume feature map fusion. The present invention verified the effectiveness of the innovation of the present invention through three different scenarios: a network without the cylindrical self-supervised loss function CFD_LL, a network CFD_LO that does not process the overlapping regions in cylindrical feature fusion, and a baseline network CFD_LC that omits both of these modules.
[0094] Cylindrical feature fusion:
[0095] As shown in Table II, introducing cylindrical feature fusion significantly improved the performance of depth estimation, especially in the left and right views where the original performance was poor. This is because the cylindrical feature map Fcylinder minimized the adverse interference between depth estimates from different views.
[0096] Processing of overlapping regions:
[0097] The method of the present invention for handling the overlapping parts of different viewpoints generally enhanced the performance of depth estimation. The abs_rel of CFD and CFD_LO shows that when this method is missing, the error at the junction is relatively large.
[0098] Cylindrical self-supervised loss Lcyl
[0099] The performance improvement of multi-view self-supervised depth estimation stems from the integration of the cylindrical self-supervised loss of the present invention, which is also clearly shown in Table II. The integration of this module consistently improved the overall depth estimation performance. The fact that this loss function uses the invariant physical distance between two points passing through the line of sight in different frames to establish supervision is clearly reflected in the abs_rel graph.
[0100] Table II is an ablation study of the cylindrical feature fusion, overlapping region processing, and cylindrical self-supervised loss of the present invention. Among them, CFD_LO, CFD_LL, and CFD_LC represent not processing the overlapping regions, the cylindrical loss of not processing the overlapping regions, and all three modules of the network that does not process the overlapping regions, respectively. The cylindrical feature module introduced by the present invention significantly enhanced the performance of depth estimation, especially improving the depth estimation performance from viewpoints that were originally ineffective.
[0101] Table II
[0102]
[0103] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A feature fusion and mapping method for multi-view self-supervised depth estimation, characterized in that The steps include: Step S1: Project and overlap the features extracted from the image of each viewpoint to establish a volume feature map, use the volume feature map to fuse image features from multiple viewpoints to establish a cylindrical feature map, and obtain a bird's-eye view feature map by flattening the volume feature map; Step S2: Use the self-attention mechanism to encode the cylindrical feature map; Step S3: Depth and posture decoding; decode the encoded cylindrical feature map into a depth map of each perspective, and decode the transformation relationship between different frames of the view from the bird's-eye view feature map; the depth decoding constructs a self-supervised loss function based on features with spatiotemporal constants, and the formula is as follows: in, represents the reprojection loss, represents the adjacent view loss, represents the cylindrical self-supervised loss, λ r , a , cyl Represent the weights of the three losses respectively; The reprojection loss At time t i Using the transformation matrix and depth estimate, the image is projected onto the temporally corresponding viewpoint image, and the reprojection loss is constructed using a weighted sum of intensity differences and structural similarities. The adjacent view loss According to the known transformation matrix between viewpoints, the overlapping parts between viewpoints are projected and the adjacent view loss is established The cylindrical self-supervised loss When the moving distance is short, the distance between the corresponding cylindrical depth maps is approximately kept constant, thereby constructing the cylindrical self-supervised loss 2. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: In step S1, the cylindrical features are divided into overlapping features and non-overlapping features, and the overlapping features with reduced dimensions are fused with the non-overlapping features to obtain a complete cylindrical feature map F surround_view .
3. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: In step S1, in the volume feature map unit With volume feature unit V t xyz Build a mapping relationship between them: z=z Among them, xyz represents the three-dimensional coordinate values, Ox and Oy represent the distance from the point to the x and y coordinate values, respectively, and Backproject(·) represents the mapping function.
4. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: In step S2, a depthwise separable convolution is used to reduce the scale of the cylindrical feature map, a set of cross-view self-attention layers is constructed to perform cross-view information exchange, and another depthwise separable convolution is used to enlarge the feature map to obtain an encoded cylindrical feature map.
5. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: In the step S3, the encoded cylindrical feature map is depth decoded to obtain an expanded cylindrical depth map, and the cylindrical depth map is projected to a corresponding viewpoint using camera parameters to obtain a decoded depth map; A continuous bird's-eye view feature map is obtained from the flattened volume feature map, which is then compressed along the feature channel into a single feature map that fuses multi-view features. Pose decoding is used to estimate the relative pose of the visual system over time.
6. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: The reprojection loss In the inter-frame pose estimation network, the transformation matrix (Pose t-1 ,Pose t+1 ), establishes a transformation based on the known external and internal parameters [R|t] of the multi-vision system, and creates a transformation matrix for each viewpoint for the next moment in, represents the set of camera internal parameters at time i, Represents the set of camera internal parameters at time j, subscript t i and t j Indicates different times.
7. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: The reprojection loss The formula is as follows: in, Representing images Projection time t j The corresponding viewpoint image, ||·||1 represents the L1 norm, SSIM represents the structural similarity function, and α represents the reprojection loss The weight coefficient of .
8. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: The adjacent view loss In the example, the transformation is established based on the known external and internal parameters [R|t] of the multi-vision system, and the transformation relationship between adjacent viewpoints is created for each viewpoint. Among them, K i Represents the intrinsic parameter matrix of camera number i, K j Represents the intrinsic parameter matrix of camera number j, [R|t] i Represents the external parameter matrix of camera number i, [R|t] j Represents the extrinsic parameter matrix of camera number j.
9. The feature fusion and mapping method for multi-view self-supervised depth estimation according to claim 1, characterized in that: The cylindrical self-supervised loss In the process, position embedding is added to the cylindrical feature map, corresponding to the angle θ and height z index of the cylindrical coordinates. After obtaining the cylindrical feature map and obtaining the corresponding cylindrical depth map through depth decoding, based on the known multi-vision system parameters [R|t], it is possible to embed the cylindrical depth map D surround_view And the depth map of each viewpoint Convert between There are many feature point pairs in the cylindrical depth map These point pairs are connected by the viewpoint, where P represents a point in the point pair whose connecting line passes through the viewpoint. Therefore, in cylindrical coordinates, the difference in the θ term of the point pair is π, r1 and r2 correspond to the depth of the point pair, and the distance dis between the feature point pairs at time t i and t j Keep constant between them; according to θ in the cylindrical coordinate system, the cylindrical depth map is divided into n parts, and a mapping relationship is established. When the moving distance is short, the distance dis in the cylindrical depth map of the corresponding part is an approximately conservative amount, so as to construct the cylindrical self-supervised loss: in, represents the average distance between the selected point pairs at time t1, Represents the average distance between the selected point pairs at time t2.