Self-supervised monocular depth estimation method based on feature sharing

By integrating the attitude estimation unit in the monocular depth estimation network, using feature sharing and hybrid convolutional encoder methods, the existing monocular depth estimation methods are solved, and the detection accuracy, stability and real-time performance in unmanned driving is achieved, efficient and accurate monocular depth estimation is achieved, and it is suitable for low-power vehicle-consumer processors.

CN113034563BActive Publication Date: 2025-05-09SUZHOU YIHANG YUANZHI INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202110196301.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-22
Publication Date
2025-05-09
Estimated Expiration
2041-02-22

AI Technical Summary

Technical Problem

The existing monocular depth estimation methods are difficult to meet the requirements of detection accuracy, stability and real-time in unmanned driving applications, and it is difficult to obtain satisfactory comprehensive results on low-power vehicle-mounted processors.

Method used

Using a single network structure, the pose estimation unit is integrated into the depth estimation unit to realize the fusion of depth estimation and pose estimation to obtain a monocular single source depth estimation network based on feature sharing. The network includes a feature encoding unit, a depth estimation unit, an pose estimation unit and a supervised training unit. The computing efficiency and accuracy are improved through a hybrid convolutional encoder and a fusion module based on a spatial attention mechanism.

Benefits of technology

Real-time, efficient and accurate monocular depth estimation is achieved, reducing network computing volume and video memory usage, improving posture estimation accuracy and computing efficiency, and is suitable for running on low-power vehicle-mounted processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113034563B_ABST
    Figure CN113034563B_ABST
Patent Text Reader

Abstract

A self-supervised monocular depth estimation method based on feature sharing adopts a new single network structure, integrates the posture estimation module into the depth estimation module, realizes the fusion of the two functions of depth estimation and posture estimation, and obtains a monocular single-source depth estimation network based on feature sharing, the network includes: feature encoding unit, depth estimation unit, posture estimation unit and supervision training unit. The posture estimation unit realizes real-time posture output in video streaming based on feature matching, and improves the accuracy of posture estimation; the depth estimation unit is based on an efficient encoding and decoding module, which improves the calculation efficiency; combined with the output of the depth estimation unit and the posture estimation unit, the supervision information is extracted from the original picture, and the self-supervised network training process is completed; effectively solves the problem of obtaining high-precision monocular depth information in real time in unmanned driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of depth perception technology and computer vision technology in the unmanned driving industry, and specifically to a self-supervised monocular depth estimation method based on feature sharing in an unmanned driving scenario, and more particularly to a self-supervised monocular depth estimation method based on feature sharing implemented by a single network structure. Background Art

[0002] With the continuous development of computer vision technology, three-dimensional scene perception tasks have played a vital role in the unmanned driving industry. The three-dimensional perception task is different from the two-dimensional perception task, which is mainly reflected in the perception and detection of information such as pedestrians, vehicles and obstacles around the unmanned vehicle in the real three-dimensional space. The information obtained by three-dimensional perception is the key basis for the unmanned vehicle to make vehicle motion decisions, and the depth information is the basis of the three-dimensional scene perception task. Among them, due to the limitations of the installation layout of the actual unmanned vehicle sensor camera, and in view of the need to accurately detect information such as pedestrians, vehicles and obstacles around the unmanned vehicle in dynamic changing situations, whether it is for the monocular camera in the monocular acquisition system containing only one camera, or for any monocular camera in the multi-camera acquisition system containing multiple monocular cameras, the three-dimensional perception system needs to obtain depth information of the scene collected by the monocular camera to solve the problem of depth information acquisition in the application of the monocular camera or the effective scene information only comes from one monocular camera. In fact, the above situation often occurs during the dynamic operation of the unmanned vehicle. Therefore, using a monocular camera to obtain depth information is a key research content in the three-dimensional scene perception task.

[0003] Currently, for the monocular depth estimation task, existing research can be mainly divided into two directions: supervised and unsupervised. Supervised monocular depth estimation networks usually perform well on specific data sets, but there are problems such as many network model parameters and difficulty in obtaining labeled data; unsupervised monocular depth estimation networks can be flexibly applied to different data sets, but there are problems such as poor network model accuracy and unreasonable network training strategies.

[0004] In order to understand the development status of the prior art, the present disclosure searches, compares and analyzes the existing patents and papers:

[0005] Patent document CN 108961327 A "A monocular depth estimation method and its device, equipment and storage medium" proposes a method for training a monocular depth estimation network using binocular information. The method uses synthetic and real binocular sample data to train a monocular depth estimation network. The method uses a small amount of real binocular data and a large amount of synthetic binocular data, and the effect is more dependent on the accuracy of the synthetic data; at the same time, the method requires real binocular data to adjust the network, so the method can only be trained and learned in an offline manner, which increases the cost of data acquisition.

[0006] Patent document CN 111680554 A "Depth estimation method, device and autonomous vehicle for autonomous driving scenarios" proposes a cascade network approach to optimize and supplement the results of monocular depth estimation. First, the basic depth estimation information is generated through the depth estimation model, and then the deviation estimation information of the target area is output using the deviation estimation network, thereby solving the problem of insufficient depth estimation accuracy in the target area. In order to improve the accuracy of monocular depth estimation for the target area, this method introduces a target detection method and a cascade network, expands the overall scale of the network, has a complex network structure, high system cost, and high computational consumption, and the neural network method is difficult to run in real time.

[0007] Patent document CN 110599533 A "Fast monocular depth estimation method suitable for embedded platforms" proposes a lightweight monocular depth estimation method on an embedded platform. The method deploys a lightweight depth estimation network on the embedded platform and configures a model training framework on the edge server. The two interact through the network: the embedded platform provides data and labels to the edge server, and the edge server performs training after receiving the data and updates the server on the embedded platform. The method provides a method for deploying a depth estimation network on an embedded platform, but the method uses an RGB-D camera to collect monocular images and depth maps. Due to the limitations of the RGB-D camera itself, the depth map perception range is limited. Therefore, the method is limited to indoor scenes and is generally used in indoor robot sports occasions. It is not applicable to outdoor vehicle-side occasions.

[0008] It can be seen that in unmanned driving, the existing monocular depth estimation methods cannot meet the requirements of unmanned driving in terms of detection accuracy, stability and real-time performance, either in terms of network model accuracy or training strategy, and it is difficult to obtain satisfactory comprehensive results on low-power vehicle processors. Therefore, it is necessary to study new monocular depth estimation methods that can not only ensure the accuracy of monocular depth estimation, but also adapt to the needs of outdoor unmanned driving vehicles without adding additional computing overhead, and can be used in low-power vehicle processors without the need for complex and high-cost sensor system support. Summary of the invention

[0009] In order to adapt to unmanned driving applications, the present invention proposes a new self-supervised monocular depth estimation method for real-time, efficient, accurate and reliable depth estimation of video images acquired by a monocular camera to effectively obtain depth information in three-dimensional scene perception. The method adopts a single network structure, without the need for two independent networks, a depth estimation network and a posture estimation network. Instead, the posture estimation unit is integrated into the depth estimation unit by using feature sharing, and the two operations of depth estimation and posture estimation are integrated in a single network, thereby obtaining a new monocular single-source depth estimation network based on feature sharing, which simplifies the network structure, speeds up the network processing speed, and determines the real-time depth of the object detected by the monocular camera video stream in real time.

[0010] The monocular single-source depth estimation network includes: a feature encoding unit, a depth estimation unit, a posture estimation unit and a supervised training unit.

[0011] The posture estimation unit realizes real-time posture output in the form of video streaming based on feature matching, thereby improving the accuracy of posture estimation.

[0012] The depth estimation unit is based on an efficient encoding and decoding module, wherein an encoder based on hybrid convolution mixes depth separable convolution, SE module and residual convolution module, and combines with dilated convolution to improve computational efficiency and achieve high-precision depth output.

[0013] Combined with the output of the depth estimation unit and the posture estimation unit, supervision information is extracted from the original image, realizing a self-supervised network training process. The current frame features and image frames are obtained.

[0014] This results in a self-supervised monocular depth estimation method based on feature sharing.

[0015] The method reduces the network computational complexity and video memory occupancy, improves the output accuracy of the posture estimation network, reduces the network's demand for video memory and computing resources, and improves computing efficiency; an efficient feature encoding module and decoding module are designed to reduce the computational complexity and parameter quantity while improving the network's output accuracy, thereby ensuring the real-time performance of the depth estimation task.

[0016] Specifically, in order to solve the above technical problems, according to one aspect of the present invention, a self-supervised monocular depth estimation method based on feature sharing is provided, wherein:

[0017] A single network structure is adopted to integrate the posture estimation unit into the depth estimation unit, so as to realize the fusion of the two operations of depth estimation and posture estimation, and obtain a monocular single-source depth estimation network based on feature sharing; the monocular single-source depth estimation network comprises: a shared feature encoding unit, a depth estimation unit, a posture estimation unit and a supervised training unit;

[0018] The method comprises the following steps:

[0019] Step 1: Data collection: collect data from the video stream through a monocular camera and output image frames;

[0020] Step 2: shared feature encoding: after receiving the image frame, preprocessing the image frame and outputting a multi-scale shared feature group through an encoder;

[0021] Step 3: decoding, receiving the multi-scale shared feature group, processing the multi-scale features through the depth estimation unit, and outputting a depth map at the original resolution; performing feature matching and decoding on the features of the current frame and the features of the previous frame through the posture estimation unit, and outputting the posture transformation between the two frames;

[0022] Step 4: Loss calculation: combine the depth map output by the depth estimation unit with the pose transformation between the two frames output by the pose estimation unit to reconstruct the target frame, and then supervise the training of the network through the difference between the original target frame and the reconstructed target frame.

[0023] Preferably, the monocular camera is deployed on an unmanned vehicle.

[0024] Preferably, the monocular camera is deployed on the upper edge of the front windshield of the driverless vehicle.

[0025] Preferably, the target frame is reconstructed through projection and interpolation operations.

[0026] Preferably, the method further comprises the following steps:

[0027] Step 5: Storage: The original image of the frame and the features outputted from the feature encoding step are stored in a storage medium for use in the decoding step and the loss calculation step at the next moment.

[0028] Preferably, the monocular camera has a resolution of 720P or above;

[0029] The monocular camera is a monocular camera in a monocular acquisition system, or is any monocular camera in a multi-eye acquisition system including multiple monocular cameras.

[0030] Preferably, in step 1, in the video stream generated by the monocular camera, real-time sampling is performed at a certain frequency to generate image frames.

[0031] Preferably, when encoding the shared features, a hybrid convolution encoder is used to mix the depthwise separable convolution, the SE module (i.e., the compression and activation Squeeze-Excitation module) with the residual convolution module, and combine it with the dilated convolution.

[0032] Preferably, a deep neural network is used to process the image, and the features of each downsampling are stored to generate a multi-scale feature set.

[0033] Preferably, the deep neural network comprises a deep residual network (ResNet) series network.

[0034] Preferably, in the hybrid convolution process, 1×1 dimensionality-enhancing convolution is used to increase the dimensional space of the feature, and then 3×3 channel-by-channel convolution and 1×1 point-by-point convolution are combined into a depth-wise separable convolution to improve the computational efficiency of the feature and reduce the number of model parameters; the SE module composed of two full connections enhances the feature expression capability by re-evaluating the importance of the feature channel; and the receptive field in the feature extraction process is guaranteed by introducing a dilated convolution in the channel-by-channel convolution.

[0035] Preferably, the depthwise separable convolution decomposes the standard convolution into two steps:

[0036] The first step is to perform channel-by-channel convolution, where each convolution kernel is responsible for only one channel;

[0037] The second step is point-by-point convolution, where the convolution kernel is reduced to 1×1, the number of channels is the same as the number of input feature channels, and the channel information of the features is mixed to obtain enhanced features;

[0038] The computation amount Cal(DW) and parameter amount Parm(DW) of the depth-separable convolution are respectively:

[0039] Cal(DW)=K c ×K c ×C in ×C out ×W out ×H out (3)

[0040] Parm(DW)=K c ×K c ×C out ×C in (4)

[0041] Among them, C in is the number of channels of the input feature map, C out is the number of channels of the output feature map, K c is the size of the convolution kernel, W out and H out are the width and height of the output feature map respectively.

[0042] Preferably, the SE module is used to learn the correlation between channels in the feature map, and each channel is evaluated and scored, thereby selectively screening feature channels.

[0043] Preferably, the operation of the SE module includes:

[0044] The first step is to compress the feature map. Global average pooling is performed on the feature map with a dimension of C×W×H to obtain a feature map with a dimension of 1×1×C, which has a global receptive field.

[0045] The second step is feature excitation. Two full-connection operations are used to perform a global information interaction in the channel dimension on the 1×1×C feature map. Finally, the score of each channel is calculated through the Sigmoid activation function, and finally multiplied with the original feature to obtain the feature map after information channel weighting.

[0046] Preferably, dilated convolution is introduced in the fifth layer (layer5) and the sixth layer (layer6) of the network layer, and a lower dilation rate is selected to ensure that the backbone network extracts features with high receptive field and high resolution;

[0047] The maximum pooling layer is removed, and two layers of ordinary convolution are added after the sixth layer (layer6), namely the seventh layer (layer7) and the eighth layer (layer8), whose void rates are 2 and 1 respectively, and their skip connections are removed to obtain a smoother network output.

[0048] Preferably, in the multi-scale feature output, the features output by the second layer (layer2), the third layer (layer3), the fourth layer (layer4), the fifth layer (layer5) and the eighth layer (layer8) are respectively selected to form a multi-scale feature map set to represent the features at different scales for feature decoding in the decoding step.

[0049] Preferably, the depth estimation unit uses a fusion module based on a spatial attention mechanism to perform feature fusion operations, E is the feature of the encoder side, f D is the feature of the decoder, and the f E and f D After a 1×1 convolution, each of them is used to obtain the compact features with reduced dimension, which are recorded as f_h E and f_h D , then the feature f_h D With the feature f_h E After concatenation, a layer of 3×3 convolution is performed after activation to compress the features to 1 dimension, followed by a Sigmoid function to output a weight distribution map σ, which represents the screening of the encoder information after combining the decoder information; then, the weight distribution map σ is combined with the original encoder feature f E Multiply point by point and finally add the decoder information f D Splicing for deep decoding.

[0050] Preferably, the posture estimation unit uses correlation calculation to match features. The correlation calculation accepts feature maps f1 and f2 from two frames respectively, and performs correlation calculation on feature blocks of (2k+1)×(2k+1) centered on any feature x1, x2 in f1 and f2; wherein, for any feature block in f1, the similarity of all feature blocks in f2 is not calculated, but only the similarity of feature blocks in f2 corresponding to the corresponding position and moved up, down, left, and right by a length range of d is calculated; the calculation is shown in formula (5):

[0051]

[0052] Wherein, f1() and f2() represent the input feature map respectively, x1 and x2 represent the calculation center, k represents the calculation range, <·> represents the dot multiplication operation, o represents the moving step in the local area, and c(x1, x2) represents the result of the dot multiplication operation of the feature map centered on x1 and x2. The calculation direction of formula (5) is unidirectional and does not satisfy the commutative law, that is, c(x1, x2)≠c(x2, x1), thereby ensuring the unidirectionality of the posture transformation in the posture estimation process.

[0053] Preferably, dense convolution is used as a decoder to decode the correlation between features, and finally the output of the posture transformation matrix between two frames is realized through one layer of dense convolution and three layers of convolution.

[0054] Preferably, the loss calculation includes:

[0055] Reconstruction of target frame: Using the decoded output depth map and posture transformation matrix, calculate the correspondence between the coordinates in the source frame and the target frame, and then reconstruct the target frame;

[0056] Calculation of loss function: L1 loss, structural similarity loss and edge smoothing loss are used as loss functions.

[0057] Preferably, the source frame is reconstructed into a target frame using the depth map output by the depth estimation unit and the pose transformation relationship between the source frame and the target frame output by the pose estimation unit, the points in the two-dimensional space are projected into the three-dimensional space through the inverse projection operation, and then the points in the three-dimensional space are projected into the coordinate space of the adjacent frame using the coordinate system transformation and projection operation, and the target frame is reconstructed using the image sampling module; finally, the supervision information is extracted using the pixel relationship between the reconstructed target frame and the original target frame, and the network is trained and supervised.

[0058] Preferably, the depth map D of the target frame output by the network is used t And the pre-calibrated intrinsic parameter K, the pixel points of the target frame are projected into the camera coordinate system under the target frame to generate a sparse point cloud PC t, the calculation formula is as follows:

[0059] PC t (p t )=D t (p t )K -1 p t (6)

[0060] where p t Represents the coordinates of any pixel in the target frame;

[0061] The coordinate system conversion step uses the pose transformation matrix T between the source frame and the target frame output by the pose estimation unit t→s Sparse point cloud PC t Transform to the source frame coordinate system to get the point cloud PC s , the calculation formula is as follows:

[0062] PC s =T t→s PC t =R t→s PC t +t t→s (7)

[0063] Where R t→s and t t→s They are the rotation matrix and translation vector output by the attitude estimation unit;

[0064] The projection module receives the sparse point cloud PC in the source frame coordinate system s Then, the point cloud is reprojected to the pixel coordinate system of the source frame using the intrinsic parameter K to obtain the corresponding point coordinates p′ s , the calculation formula is as follows:

[0065] p′ s =KPC s (p t ) (8)

[0066] The corresponding relationship between the pixels of the two frames is as follows:

[0067] p′ s =KT t→s D t (p t )K -1 p t (9)

[0068] The camera internal parameter K is the pre-calibrated value, and the pose transformation matrix T t→s and the depth map D t They are output by the posture estimation unit and the depth estimation unit respectively.

[0069] Preferably, the loss function includes: L1 loss, structural similarity (Structural SIMilarity, SSIM) loss and edge smoothing loss.

[0070] Preferably, the L1 loss uses the absolute value of the difference in pixel level between the original target frame and the reconstructed target frame as the loss function, as shown in formula (10):

[0071]

[0072] Where S represents a continuous sequence of images during training, I t (p) and I′ t (p) represents the pixel value of a certain pixel point in the original target frame and the reconstructed target frame respectively; <I1,…,I N > represents the image sequence obtained by chronological order in a training process, corresponding to the 1st frame to the Nth frame respectively.

[0073] Preferably, the structural similarity loss function describes the similarity between the uncompressed original image and the compressed distorted image. The structural similarity loss function constrains the quality of the reconstructed image from three aspects: brightness, contrast and structure. The calculation is as follows:

[0074]

[0075]

[0076]

[0077] SSIM(x,y)=[l(x,y) α [c(x,y)] β ·[s(x,y)] γ (14)

[0078] Where l(x,y) represents brightness similarity, c(x,y) represents contrast similarity, and s(x,y) represents structural similarity; x and y are coordinate values, parameters α = β = γ = 1, and u x ,u y are the means of x and y, σ x ,σ y are the variances of x and y, σ xy is the covariance of x and y; parameter c1=(k1L) 2 , c2=(k2L) 2 , where L is the value range of pixel value; parameters k1 = 0.01, k2 = 0.03, c3 = 0.5c2;

[0079] Get the similarity metric loss L between the two frames ssim(x, y) is as shown in formula (15):

[0080]

[0081] Preferably, the edge smoothing loss L smooth As shown in formula (16):

[0082]

[0083] where d t is the depth value of the frame at time t, is d t The mean of That is, the inverse depth information after average normalization; I t is the pixel value; and Represents the differential operation in the x dimension and y dimension respectively.

[0084] Preferably, the modulus value of S is 3 or 5.

[0085] Preferably, the storage medium comprises a fixed memory or video memory, in which the space for storing features and pictures remains fixed.

[0086] To solve the above technical problem, according to another aspect of the present invention, a readable storage medium is provided, wherein:

[0087] The readable storage medium stores execution instructions, which, when executed by the processor, are used to implement the above-mentioned self-supervised monocular depth estimation method based on feature sharing.

[0088] To solve the above technical problems, according to another aspect of the present invention, a self-supervised monocular depth estimation system based on feature sharing is provided, comprising:

[0089] A memory storing a program for executing the self-supervised monocular depth estimation method based on feature sharing;

[0090] A processor; the processor executes the program.

[0091] In order to solve the above technical problems, according to another aspect of the present invention, an unmanned vehicle is provided, comprising:

[0092] An on-board processor, wherein the on-board processor executes the above-mentioned self-supervised monocular depth estimation method based on feature sharing.

[0093] Beneficial effects of the present disclosure:

[0094] 1. A hybrid convolutional module is proposed, which can encode features quickly, efficiently and accurately;

[0095] 2. A posture estimation decoder based on feature matching is proposed, which can achieve high-precision posture transformation output between two frames;

[0096] 3. A feature-sharing depth estimation method is proposed, which can save a lot of video memory in network training, speed up the network reasoning speed, and realize online learning on low-power devices.

[0097] 4. Realized real-time and accurate depth estimation of monocular camera;

[0098] 5. The simplified single network structure improves computing efficiency while ensuring the accuracy of the system; BRIEF DESCRIPTION OF THE DRAWINGS

[0099] The accompanying drawings illustrate exemplary embodiments of the present disclosure and are used together with the description thereof to explain the principles of the present disclosure, wherein these drawings are included to provide a further understanding of the present disclosure, and the drawings are included in and constitute a part of this specification. By describing the embodiments of the present disclosure in detail in conjunction with the accompanying drawings, the above and other purposes, features, and advantages of the present disclosure will become more apparent.

[0100] Figure 1 It is the overall flow chart;

[0101] Figure 2 It is a schematic diagram of the hybrid convolution module;

[0102] Figure 3 It is a feature fusion module based on the spatial attention mechanism;

[0103] Figure 4 is a schematic diagram of the posture estimation network;

[0104] Figure 5 It is a schematic diagram of target frame reconstruction. DETAILED DESCRIPTION

[0105] The present disclosure is further described in detail below in conjunction with the accompanying drawings and implementations. It is understood that the specific implementations described herein are only used to explain the relevant content, rather than to limit the present disclosure. It should also be noted that, for ease of description, only the parts related to the present disclosure are shown in the accompanying drawings.

[0106] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other. Unless otherwise specified, the exemplary embodiments / embodiments shown will be understood as exemplary features of various details of some ways in which the technical concept of the present disclosure can be implemented in practice. Therefore, unless otherwise specified, the features of various embodiments / embodiments can be further combined, separated, interchanged and / or rearranged without departing from the technical concept of the present disclosure.

[0107] The terms used herein are for the purpose of describing specific embodiments, and are not intended to be restrictive. As used herein, unless the context clearly indicates otherwise, the singular forms "one (kind, person)" and "said (the)" are also intended to include plural forms. In addition, when the terms "comprise" and / or "include" and their variations are used in this specification, it is explained that there are stated features, integral bodies, steps, operations, parts, assemblies and / or their groups, but it is not excluded that there are or add one or more other features, integral bodies, steps, operations, parts, assemblies and / or their groups. It should also be noted that, as used herein, the terms "substantially", "approximately" and other similar terms are used as approximate terms and not as degree terms, so that they are used to explain the inherent deviations of the measured values, calculated values ​​and / or the values ​​provided that will be recognized by those of ordinary skill in the art.

[0108] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0109] The purpose of this disclosure is to provide a new self-supervised monocular depth estimation method.

[0110] In order to reduce the cost of data collection and adapt to online training and learning methods, the present invention does not use binocular data, that is, no real binocular data or synthetic binocular data is required, and only relies on monocular camera collection data for monocular depth estimation. The monocular camera can be a monocular camera in a monocular collection system, or any monocular camera in a multi-camera collection system including multiple monocular cameras.

[0111] In order to adapt to the processing power of low-power on-board processors and be applicable in unmanned driving that cannot provide support for complex and high-cost sensor systems, the present invention redesigns the network structure, adjusts the training strategy, and adopts a new single network structure to achieve the fusion of depth estimation and posture estimation. The posture estimation is integrated into the depth estimation, which simplifies the network structure and reduces the computational overhead of the system from the algorithm principle.

[0112] The use of a single network structure overcomes the problem of complex network structures in the prior art that use depth estimation networks and posture estimation networks to obtain depth information and posture information respectively. It occupies less video memory and computing resources, has good system real-time performance, and is very suitable for use on low-power vehicle processors. The overall network speed and performance have also been improved.

[0113] By adopting the single network structure, a new monocular single-source depth estimation network based on feature sharing is obtained. The monocular single-source depth estimation network includes: a feature encoding unit, a depth estimation unit, a posture estimation unit and a supervised training unit.

[0114] In order to improve the accuracy of posture estimation and the real-time processing capability of the system, a monocular camera is used to collect data from a video stream, and a feature matching-based method is used to achieve real-time posture output in the form of video streaming, that is, the data processed by the network disclosed in the present invention is real-time video streaming data rather than static image data, thereby ensuring the real-time and accuracy of the output information and improving the accuracy of the system depth estimation. The single network structure processes the video streaming data in real time, thereby quickly and accurately determining the depth of the object in front of the unmanned vehicle in real time.

[0115] In order to achieve high-precision depth output, an efficient encoding and decoding module is designed. Among them, the encoder based on hybrid convolution mixes the depth separable convolution, SE module and residual convolution module, and combines it with the hole convolution, which greatly improves the computational efficiency. The posture estimation decoder based on feature matching can achieve high-precision posture transformation output between two frames. The depth estimation method based on feature sharing can save a lot of video memory in network training and speed up the network reasoning speed.

[0116] like Figure 1As shown, the present disclosure adopts a self-supervised monocular depth estimation method architecture based on feature sharing, which adopts a single network structure. There is no need for two independent depth estimation networks and posture estimation networks. Instead, feature sharing is used to integrate the posture estimation unit into the depth estimation unit, so as to achieve the fusion of the two operations of depth estimation and posture estimation in a single network, so that the network structure is simplified, the network processing speed is accelerated, and the real-time depth of the object detected by the monocular camera video stream is determined in real time. The single network structure includes a data acquisition module, a feature encoding module, a decoding module, a loss calculation module and a storage module.

[0117] The overall process of the self-supervised monocular depth estimation method based on feature sharing includes the following steps:

[0118] The first step is the data collection step, which uses a monocular camera deployed on an unmanned vehicle to collect data from the video stream and output image frames;

[0119] The second step is the shared feature encoding step. After receiving the image frame, the image is preprocessed and the encoder outputs a multi-scale feature group.

[0120] The third step is the decoding step. Based on the above-mentioned shared feature encoding, on the one hand, the depth map is decoded, and the multi-scale features are processed by the depth estimation unit to output the depth map at the original resolution. On the other hand, the posture is decoded. Through the posture estimation unit, the features of the current frame are matched and decoded with the features of the previous frame, and finally the posture change between the two frames is output;

[0121] The fourth step is the loss calculation step, which combines the depth map output by the depth estimation unit with the posture transformation between the two frames output by the posture estimation unit, reconstructs the target frame through operations such as projection and interpolation, and then supervises the training of the network through the difference between the original target frame and the reconstructed target frame;

[0122] The fifth step is the storage step, which stores the original image of this frame and the features output by the feature encoding step in a fixed memory or video memory for the decoding step and loss calculation step at the next moment.

[0123] 1. Data collection:

[0124] The data acquisition system disclosed in the present invention is composed of a monocular camera with a resolution of 720P or above. The monocular camera can be a monocular camera in a monocular acquisition system, or any monocular camera in a multi-camera acquisition system including multiple monocular cameras. Usually, in an unmanned driving system, the monocular camera is arranged at the upper edge of the front windshield of the unmanned vehicle to collect visual data in front of the vehicle in real time, thereby realizing video streaming data acquisition. The monocular camera needs to be calibrated for use in subsequent loss calculations. The calibration algorithm includes but is not limited to the Zhang Dingyou calibration method, and finally the distortion coefficient, internal parameters and external parameters of the monocular camera relative to the vehicle are obtained.

[0125] During the data acquisition process, according to the program, in the video stream generated by the monocular camera, real-time sampling is performed at a certain frequency to generate image frames, and the generated image frame data is transmitted in real time for use in subsequent steps or modules. It can be seen that the present disclosure does not need to use any real binocular data or any synthetic binocular data, so that the subsequent detection effect does not need to rely on the accuracy of the synthetic data. The accuracy of network training can be guaranteed only by the data collected in real time by the monocular camera, and the accuracy of depth estimation can be improved; the present disclosure does not need real binocular data to adjust the network, which overcomes the disadvantage that the traditional method can only be trained and learned in an offline manner, and greatly reduces the cost of system data acquisition.

[0126] 2. Shared feature encoding:

[0127] This paper proposes an encoder based on hybrid convolution, which mixes the depth-separable convolution, Squeeze-Excitation (SE) module and residual convolution module, and combines the dilated convolution, which improves the network computing efficiency while speeding up the network reasoning speed and the parameter scale of the encoder, so that the network can perform real-time reasoning on resource-limited embedded devices. The SE module realizes the weighting of feature channels.

[0128] For feature encoding, a deep neural network in the image classification task is used as an encoder to process the image, store the features of each downsampling, and generate a multi-scale feature set. The deep neural network includes a ResNet deep residual network series network. The present disclosure uses a hybrid convolution module as the basic network module of the feature encoder to help the encoder quickly and efficiently extract robust visual features. The schematic diagram of the hybrid convolution module is shown in Figure 2 shown.

[0129] On the basis of the residual network module used in the ResNet series of networks, the deep separable convolution, Squeeze-Excitation (SE) network module and dilated convolution are introduced to form a hybrid convolution module, in which the 1x1 dimensionality-enhancing convolution is used to increase the dimensional space of the feature, and then the 3x3 channel-by-channel convolution and 1x1 point-by-point convolution are combined into a deep separable convolution to improve the computational efficiency of the feature and reduce the number of model parameters; the SE module composed of two full connections is derived from the idea of ​​attention, and enhances the feature expression ability by re-evaluating the importance of the feature channel. In addition, by introducing dilated convolution in the channel-by-channel convolution, a sufficiently large receptive field is guaranteed during the feature extraction process, thereby ensuring the feature's ability to perceive detail information. Among them:

[0130] 1) Depthwise separable convolution comes from the MobileNet deep neural network, which is a network module designed specifically for embedded devices such as vehicle-mounted devices and mobile phones and is often used in lightweight networks.

[0131] Depthwise separable convolution decomposes the standard convolution into two steps:

[0132] The first step is to perform channel-by-channel convolution. Each convolution kernel is responsible for only one channel. Therefore, the number of convolution kernels is equal to the number of channels of the input feature and the number of channels of the feature remains unchanged. At this time, the information between channels is separated.

[0133] The second step is point-by-point convolution, where the convolution kernel is reduced to 1×1, but the number of channels is the same as the number of input feature channels, which is equivalent to mixing the channel information of the features to obtain enhanced features.

[0134] For conventional convolution operations, the number of channels of the output feature map is set to C in , the number of output channels is C out , the size of the convolution kernel is K c , and the width and height of the output feature map are W out and H out , then the calculation amount Cal(conv) and parameter amount Parm(conv) are respectively:

[0135] Cal(conv)=K c ×K c ×C in ×W out ×H out +C in ×C out ×W out ×H out (1)

[0136] Parm(conv)=K c ×K c×C in +C in ×C out (2)

[0137] Correspondingly, the computation amount Cal(DW) and parameter amount Parm(DW) of the depth-separable convolution are:

[0138] Cal(DW)=K c ×K c ×C in ×C out ×W out ×H out (3)

[0139] Parm(DW)=K c ×K c ×C out ×C in (4)

[0140] Then the number of parameters and computational complexity of depthwise separable convolution are respectively It can be seen that the use of depth-wise separable convolution can very effectively reduce the number of parameters and calculations of conventional convolution and reduce the system computational overhead.

[0141] 2) The SE module comes from SE-Net (Squeeze-Excitation), which uses the attention mechanism. The SE module is used in this disclosure to learn the correlation between channels in the feature map, evaluate and score each channel, and then selectively screen the feature channels. The SE module is divided into two steps: the first step is to compress the feature map, perform global average pooling on the feature map with a dimension of C×W×H, and obtain a feature map with a dimension of 1×1×C, which has a global receptive field; the second step is to excite the feature, use two full connection operations, perform a global information interaction of the channel dimension on the 1×1×C feature map, and finally calculate the score of each channel through the Sigmoid activation function, and finally multiply it with the original feature to obtain the feature map after the information channel weighting.

[0142] 3) In order to solve the problem of detail loss caused by downsampling, the encoder proposed in this disclosure introduces a dilated convolution to reduce the loss of information in the spatial dimension while ensuring a high receptive field.

[0143] The specific operation is as follows: introduce dilated convolution in the fifth layer (layer5) and the sixth layer (layer6) of the network layer, select a lower dilation rate (2 / 4) to ensure that the backbone network can extract features with high receptive field and high resolution; remove the maximum pooling layer to suppress the grid effect, and add two layers of ordinary convolution (seventh layer layer7 and eighth layer layer8) after the sixth layer (layer6), with dilation rates of 2 and 1 respectively, and remove their jump connections to obtain smoother network output. Improve the expressiveness of features.

[0144] In summary, the network structure of the hybrid convolutional encoder proposed in this disclosure is shown in Table 1:

[0145] Table 1 Encoder network

[0146]

[0147] Specifically, conv2d represents conventional two-dimensional convolution, and mix_conv represents the mixed convolution proposed in this method. In the multi-scale feature output, the features output by the second layer layer2, the third layer layer3, the fourth layer layer4, the fifth layer layer5, and the eighth layer layer8 are respectively selected to form a multi-scale feature map set, which is used to represent features at different scales. The auxiliary decoding step performs feature decoding.

[0148] It can be seen that the use of the hybrid convolution module can quickly, efficiently and accurately encode features and obtain a multi-scale shared feature group, which provides a basis for integrating the depth estimation unit into the posture estimation unit.

[0149] 3. Decoding

[0150] The decoding is divided into two parts, corresponding to the depth estimation unit that outputs the depth map and the posture estimation unit that outputs the posture transformation.

[0151] 1) The depth estimation unit proposed in this disclosure corresponds to a multi-scale feature map set. In the decoding process of gradually improving the feature resolution, it is fused with the encoder features at the corresponding scale to restore the detail information at the corresponding scale. This disclosure uses a fusion module based on a spatial attention mechanism to perform feature fusion operations. Its network structure is as follows: Figure 3 shown.

[0152] Assume the feature of the encoder side is f E , the decoder feature is f D , and each undergoes a 1×1 convolution to obtain a compact feature representation with reduced dimension, denoted as f_h E and f_h D , and then the feature f_h D With the feature f_hE After concatenation, a layer of 3×3 convolution is performed after activation to compress the features to 1 dimension, followed by a sigmoid function to output a weight distribution map σ. The weight distribution map σ represents the screening of the encoder information after combining the decoder information. Finally, the weight distribution map σ is combined with the original encoder feature f E Multiply point by point and finally add the decoder information f D Splicing is used for deep decoding.

[0153] 2) The posture estimation unit proposed in the present disclosure uses correlation calculation to match features. The specific process is as follows: the correlation calculation accepts the feature maps f1 and f2 from the two frames respectively, and through a convolution-like operation, performs correlation calculation on the (2k+1)×(2k+1) feature blocks centered on any feature x1, x2 in f1 and f2. In order to reduce the amount of calculation, for any feature block in f1, the similarity of all feature blocks in f2 is not calculated, but only the similarity of the feature blocks in f2 corresponding to the corresponding position and moved up, down, left, and right within a length range of d is calculated. The calculation formula is as follows:

[0154]

[0155] Among them, f1(), f2() represent the input feature map, x1, x2 represent the calculation center, k represents the calculation range, <·> represents the dot multiplication operation, o represents the moving step in the local area, and c(x1, x2) represents the result of the dot multiplication operation of the feature map centered on x1 and x2. It should be noted that the calculation direction of formula (5) is unidirectional and does not satisfy the commutative law, that is, c(x1, x2) ≠ c(x2, x1), thereby ensuring the unidirectionality of the posture transformation during the posture estimation process.

[0156] In addition, for the correlation calculation, the present disclosure also introduces dense convolution as a decoder to decode the correlation between features. Finally, the output of the posture transformation matrix between two frames is realized through one layer of dense convolution and three layers of convolution.

[0157] In summary, for the encoding and decoding of posture estimation, the network diagram is as follows Figure 4 shown.

[0158] Among them, f1 and f2 represent the features output by the encoding module after encoding two adjacent frames, f3 represents the correlation graph, and finally the posture transformation matrix R, T between the two frames is output through posture decoding.

[0159] For posture estimation decoding, it is necessary not only to accept the feature output after encoding the current frame, but also to perform correlation calculation on the features after encoding the previous frame.

[0160] 4. Loss Calculation

[0161] The loss calculation module adopted in the present invention is divided into two parts: the first part is the reconstruction of the target frame, which uses the depth map and posture transformation matrix output by decoding to calculate the correspondence between the coordinates in the source frame and the target frame, and then reconstructs the target frame; the second part is the calculation of the loss function, and the present invention adopts L1 loss, structural similarity loss and edge smoothing loss.

[0162] 1) The present invention reconstructs the target frame by using the depth map output by the depth estimation unit and the position transformation relationship between the source frame and the target frame output by the posture estimation unit. With the help of the image reconstruction algorithm, the points in the two-dimensional space are projected into the three-dimensional space through the inverse projection operation. Subsequently, the coordinate system transformation and projection operation are used to project the points in the three-dimensional space into the coordinate space of the adjacent frame. The target frame is reconstructed using the image sampling module. Finally, the supervision information is extracted using the pixel relationship between the reconstructed target frame and the original target frame to complete the training supervision of the network algorithm. The flowchart is as follows: Figure 5 shown.

[0163] Figure 5 The reverse projection part refers to the depth map D of the target frame output by the network. t And the pre-calibrated intrinsic parameter K, the pixel points of the target frame are projected into the camera coordinate system under the target frame to generate a sparse point cloud PC t , the calculation formula is as follows:

[0164] PC t (p t )=D t (p t )K -1 p t (6)

[0165] where p t Represents the coordinates of any pixel in the target frame.

[0166] The coordinate system conversion step uses the pose transformation matrix T between the source frame and the target frame output by the pose estimation unit t→s Sparse point cloud PC t Transform to the source frame coordinate system to get the point cloud PC s , the calculation formula is as follows:

[0167] PC s =T t→s PC t =R t→s PC t +t t→s (7)

[0168] Where R t→s and t t→sThese are the rotation matrix and translation vector output by the attitude estimation unit.

[0169] The projection module receives the sparse point cloud PC in the source frame coordinate system s After that, the point cloud is reprojected to the pixel coordinate system of the source frame using the internal parameter K. The corresponding point coordinates p′ are obtained s , the calculation formula is as follows:

[0170] p′ s =KPC s (p t )(8)

[0171] In summary, there is a corresponding relationship between pixels in two frames:

[0172] p′ s =KT t→s D t (p t )K -1 p t (9)

[0173] The camera internal parameter K is the pre-calibrated value, and the pose transformation matrix T t→s and the depth map D t They are output by the posture estimation unit and the depth estimation unit respectively.

[0174] Since the generated new coordinate values ​​are continuous values, the sampling process uses a fully differentiable bilinear interpolation algorithm to perform bilinear interpolation using the pixel values ​​of the four coordinate points adjacent to the sub-coordinate point to obtain the correspondence between the source frame and the target frame pixels, and generate the reconstructed target frame I′ t .

[0175] 2) The loss function used in the present disclosure includes three parts: L1 loss, structural similarity (SSIM) loss and edge smoothing loss. L1 loss refers to the L1 distance metric loss commonly used in the field of machine learning, using the absolute value of the pixel-level difference between the original target frame and the reconstructed target frame as the loss function, as shown in formula (10):

[0176]

[0177] Where S represents a continuous sequence of images during training. Generally, the modulus of S is 3 or 5. t (p) and I′ t (p) represents the pixel value of a certain pixel point in the original target frame and the reconstructed target frame respectively.

[0178] <I1,…,I N> represents the image sequence obtained by chronological order in a training process, corresponding to the 1st frame to the Nth frame respectively.

[0179] The structural similarity loss function is used to describe the similarity between the uncompressed original image and the distorted image after compression, and is used to measure the performance of the compression algorithm. In this disclosure, the SSIM loss function constrains the quality of the reconstructed image from three aspects: brightness, contrast, and structure. The calculation formulas are as follows:

[0180]

[0181]

[0182]

[0183] SSIM(x,y)=[l(x,y)] α [c(x,y)] β ·[s(x,y)] γ (14)

[0184] Where l(x,y) represents brightness similarity, c(x,y) represents contrast similarity, and s(x,y) represents structure similarity. x and y are coordinate values. Generally, the parameters α = β = γ = 1, u x ,u y are the means of x and y, σ x ,σ y are the variances of x and y, σ xy is the covariance of x and y; parameter c1=(k1L) 2 ,c2=(k2L) 2 , where L is the range of pixel values; parameters k1 = 0.01, k2 = 0.03, c3 = 0.5c2. In summary, the similarity measurement loss between the two frames can be L ssim (x,y), as shown in formula (15):

[0185]

[0186] Edge smoothing loss: Usually, the depth map output by the network will have a sense of discontinuity, that is, the depth map of the surface of the object at the same depth is not smooth. In order to make the depth value of the surface of the object in the depth map smoother and the depth value between objects more layered, the edge smoothing loss L is introduced. smoot h , as shown in formula (16):

[0187]

[0188] where d t is the depth value of the frame at time t, is d t The mean of That is, the inverse depth information after average normalization; I t is the pixel value. and Represents the differential operation in the x dimension and y dimension respectively.

[0189] 5. Storage

[0190] The present disclosure proposes that the encoder-shared depth estimation method not only needs to accept the feature output after encoding the current frame, but also needs to perform correlation calculation on the features after encoding the previous frame when performing posture estimation decoding. Therefore, after completing the training of this frame, the features encoded in this frame and the original picture need to be stored in a storage medium for use in the next frame training. The storage medium includes but is not limited to memory and video memory. To improve efficiency, the present disclosure proposes that the space used to store features and pictures in the storage medium should remain fixed.

[0191] In summary, since the present invention redesigns the posture estimation network from the perspective of feature matching, introduces a correlation calculation module and a dense convolution module, improves the performance of the posture estimation network, and realizes a high-precision posture transformation output between two frames based on the posture estimation decoder based on feature matching; due to the adoption of a feature sharing-based design method, the posture estimation unit and the depth estimation unit share the same feature encoder, realize a single network structure, reduce the network complexity, save a lot of video memory in network training, and speed up the network reasoning speed, thereby reducing the overall network during training and forward reasoning. The occupation of video memory and computing resources, speeding up the overall speed and performance of the network, can realize online learning on low-power devices; due to the encoder based on hybrid convolution, the depth separable convolution, SE module and residual convolution module are mixed, and the void convolution is combined, therefore, while improving the network computing efficiency, the network reasoning speed is accelerated, the parameter scale of the encoder is reduced, and the feature encoding can be performed quickly, efficiently and accurately, so that the network can perform real-time reasoning on resource-limited embedded devices.

[0192] Therefore, the method disclosed in the present invention reduces the amount of network calculation and video memory occupancy, improves the output accuracy of the posture estimation network, reduces the network's demand for video memory and computing resources, and improves computing efficiency; in an unmanned vehicle, it can ensure that the depth estimation task of the monocular camera is completed in real time and accurately, and can process the video streaming data detected by the monocular camera in real time in real time, and quickly and accurately determine the depth of the object in front of the unmanned vehicle in real time.

[0193] It can be seen that the new self-supervised monocular depth estimation method and system based on feature sharing proposed in the present invention are very suitable for the outdoor use background environment of unmanned vehicles, and the system has low computational overhead and can be used for low-power vehicle processors. At the same time, it does not require high-cost sensor system support. The depth estimation has good real-time performance and high accuracy, and has broad application prospects.

[0194] So far, the technical solutions of the present disclosure have been described in conjunction with the preferred implementation methods shown in the accompanying drawings. However, those skilled in the art should understand that the above implementations are only for the purpose of clearly illustrating the present disclosure, and are not intended to limit the scope of the present disclosure. The protection scope of the present disclosure is obviously not limited to these specific implementations. Without departing from the principles of the present disclosure, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present disclosure.

Claims

1. A self-supervised monocular depth estimation method based on feature sharing, characterized in that: A single network structure is adopted to integrate the posture estimation unit into the depth estimation unit, realizing the fusion of the depth estimation and posture estimation operations, and obtaining a monocular single-source depth estimation network based on feature sharing. The monocular single-source depth estimation network comprises: a shared feature encoding unit, a depth estimation unit, a posture estimation unit and a supervised training unit; The method comprises the following steps: Step 1: Data collection: collect data from the video stream through a monocular camera and output image frames; Step 2: shared feature encoding: after receiving the image frame, preprocessing the image frame and outputting a multi-scale shared feature group through an encoder; Step 3: decoding, receiving the multi-scale shared feature group, processing the multi-scale features through the depth estimation unit, and outputting a depth map at the original resolution; performing feature matching and decoding on the features of the current frame and the features of the previous frame through the posture estimation unit, and outputting the posture transformation between the two frames; Step 4: Loss calculation: combine the depth map output by the depth estimation unit with the pose transformation between the two frames output by the pose estimation unit to reconstruct the target frame, and then supervise the training of the network through the difference between the original target frame and the reconstructed target frame.

2. The self-supervised monocular depth estimation method based on feature sharing according to claim 1, characterized in that: The monocular camera is deployed on an unmanned vehicle.

3. The self-supervised monocular depth estimation method based on feature sharing according to claim 1 or 2, characterized in that: The monocular camera is deployed on the upper edge of the front windshield of the driverless vehicle.

4. The self-supervised monocular depth estimation method based on feature sharing according to claim 1 or 2, characterized in that: The target frame is reconstructed through projection and interpolation operations.

5. The self-supervised monocular depth estimation method based on feature sharing according to claim 1 or 2, characterized in that: The following steps are also included: Step 5: Storage: The original image of the frame and the features outputted from the feature encoding step are stored in a storage medium for use in the decoding step and the loss calculation step at the next moment.

6. The method for self-supervised monocular depth estimation based on feature sharing according to claim 1 or 2, characterized in that: The monocular camera has a resolution of 720P or above; The monocular camera is a monocular camera in a monocular acquisition system, or is any monocular camera in a multi-eye acquisition system including multiple monocular cameras.

7. The self-supervised monocular depth estimation method based on feature sharing according to claim 1 or 2, characterized in that: In the step 1, in the video stream generated by the monocular camera, real-time sampling is performed at a certain frequency to generate image frames.

8. The self-supervised monocular depth estimation method based on feature sharing according to claim 1 or 2, characterized in that: When encoding shared features, a hybrid convolution encoder is used to mix depth-wise separable convolution, SE module (compression and activation Squeeze-Excitation module) with residual convolution module and combine it with dilated convolution.

9. The method for self-supervised monocular depth estimation based on feature sharing according to claim 8, characterized in that: A deep neural network is used to process the image, and the features of each downsampling are stored to generate a multi-scale feature set.

10. The method for self-supervised monocular depth estimation based on feature sharing according to claim 9, characterized in that: The deep neural network includes a deep residual network (ResNet) series network.

11. The method for self-supervised monocular depth estimation based on feature sharing according to claim 10, characterized in that: In the hybrid convolution process, 1×1 dimensionality-enhancing convolution is used to increase the dimensional space of the feature, and then 3×3 channel-by-channel convolution and 1×1 point-by-point convolution are combined into a depth-wise separable convolution to improve the computational efficiency of the features and reduce the number of model parameters. The SE module, which consists of two full connections, enhances the feature expression capability by re-evaluating the importance of the feature channels. The receptive field in the feature extraction process is guaranteed by introducing dilated convolution in the channel-by-channel convolution.

12. The method for self-supervised monocular depth estimation based on feature sharing according to claim 11, characterized in that: The depthwise separable convolution decomposes the standard convolution into two steps: The first step is to perform channel-by-channel convolution, where each convolution kernel is responsible for only one channel; The second step is point-by-point convolution, where the convolution kernel is reduced to 1×1, the number of channels is the same as the number of input feature channels, and the channel information of the features is mixed to obtain enhanced features; The computation amount Cal(DW) and parameter amount Parm(DW) of the depth-separable convolution are respectively: Cal(DW)=K c ×K c ×C in ×C out ×W out ×H out (3) Parm(DW)=K c ×K c ×C out ×C in (4) Among them, C in is the number of channels of the input feature map, C out is the number of channels of the output feature map, K c is the size of the convolution kernel, W out and H out are the width and height of the output feature map respectively.

13. The method for self-supervised monocular depth estimation based on feature sharing according to claim 11, characterized in that: The SE module is used to learn the correlation between channels in the feature map, and each channel is evaluated and scored, thereby selectively screening feature channels.

14. The method for self-supervised monocular depth estimation based on feature sharing according to claim 13, characterized in that: The operations of the SE module include: The first step is to compress the feature map. Global average pooling is performed on the feature map with a dimension of C×W×H to obtain a feature map with a dimension of 1×1×C, which has a global receptive field. The second step is feature excitation. Two full-connection operations are used to perform a global information interaction in the channel dimension on the 1×1×C feature map. Finally, the score of each channel is calculated through the Sigmoid activation function, and finally multiplied with the original feature to obtain the feature map after information channel weighting.

15. The method for self-supervised monocular depth estimation based on feature sharing according to claim 1 or 2, characterized in that: The depth estimation unit uses a fusion module based on a spatial attention mechanism to perform feature fusion operations, f E is the feature of the encoder side, f D is the feature of the decoder, and the f E and f D After a 1×1 convolution, each of them is used to obtain the compact features with reduced dimension, which are recorded as f_h E and f_h D , then the feature f_h D With the feature f_h E After concatenation, a layer of 3×3 convolution is performed after activation to compress the features to 1 dimension, followed by a Sigmoid function to output a weight distribution map σ, which represents the screening of the encoder information after combining the decoder information; then, the weight distribution map σ is combined with the original encoder feature f E Multiply point by point and finally add the decoder information f D Splicing for deep decoding.

16. The method for self-supervised monocular depth estimation based on feature sharing according to claim 1 or 2, characterized in that: The posture estimation unit uses correlation calculation to match features. The correlation calculation accepts feature maps f1 and f2 from two frames respectively, and performs correlation calculation on feature blocks of (2k+1)×(2k+1) centered on any feature x1, x2 in f1 and f2; wherein, for any feature block in f1, the similarity of all feature blocks in f2 is not calculated, but only the similarity of feature blocks in f2 corresponding to the corresponding position and moved up, down, left, and right by a length range of d is calculated; the calculation is shown in formula (5): Wherein, f1(), f2() represent the input feature map, x1, x2 represent the calculation center, k represents the calculation range, <·> represents the dot multiplication operation, o represents the moving step in the local area, and c(x1, x2) represents the result of the dot multiplication operation of the feature map centered on x1 and x2. The calculation direction of formula (5) is unidirectional and does not satisfy the commutative law, that is, c(x1, x2) ≠ c(x2, x1), thereby ensuring the unidirectionality of the posture transformation in the posture estimation process.

17. The method for self-supervised monocular depth estimation based on feature sharing according to claim 16, characterized in that: Dense convolution is used as a decoder to decode the correlation between features, and finally the output of the posture transformation matrix between two frames is realized through one layer of dense convolution and three layers of convolution.

18. The method for self-supervised monocular depth estimation based on feature sharing according to claim 1 or 2, characterized in that: The loss calculation includes: Reconstruction of target frame: Using the decoded output depth map and posture transformation matrix, calculate the correspondence between the coordinates in the source frame and the target frame, and then reconstruct the target frame; Calculation of loss function: L1 loss, structural similarity loss and edge smoothing loss are used as loss functions.

19. The method for self-supervised monocular depth estimation based on feature sharing according to claim 18, characterized in that: The source frame is reconstructed into a target frame using the depth map output by the depth estimation unit and the pose transformation relationship between the source frame and the target frame output by the pose estimation unit. The points in the two-dimensional space are projected into the three-dimensional space through the inverse projection operation. Then, the points in the three-dimensional space are projected into the coordinate space of the adjacent frame using the coordinate system transformation and projection operations. The target frame is reconstructed using the image sampling module. Finally, the supervision information is extracted using the pixel relationship between the reconstructed target frame and the original target frame to supervise the network training.

20. The method for self-supervised monocular depth estimation based on feature sharing according to claim 19, characterized in that: Using the depth map D of the target frame output by the network t And the pre-calibrated intrinsic parameter K, the pixel points of the target frame are projected into the camera coordinate system under the target frame to generate a sparse point cloud PC t , the calculation formula is as follows: PC t (p t )=D t (p t )K -1 p t (6) where p t Represents the coordinates of any pixel in the target frame; The coordinate system conversion step uses the pose transformation matrix T between the source frame and the target frame output by the pose estimation unit t→s Sparse point cloud PC t Transform to the source frame coordinate system to get the point cloud PC s , the calculation formula is as follows: PC s =T t→s PC t =R t→s PC t +t t→s (7) Where R t→s and t t→s They are the rotation matrix and translation vector output by the attitude estimation unit; The projection module receives the sparse point cloud PC in the source frame coordinate system s Then, the point cloud is reprojected onto the pixel coordinate system of the source frame using the intrinsic parameter K to obtain the corresponding point coordinates p s ′ , the calculation formula is as follows: p s ′ =KPC s (p t ) (8) The corresponding relationship between the pixels of the two frames is as follows: p s ′ =KT t→s D t (p t )K -1 p t (9) The camera internal parameter K is the pre-calibrated value, and the pose transformation matrix T t→s and the depth map D t They are output by the posture estimation unit and the depth estimation unit respectively.

21. The method for self-supervised monocular depth estimation based on feature sharing according to claim 20, characterized in that: The L1 loss uses the absolute value of the pixel-level difference between the original target frame and the reconstructed target frame as the loss function, as shown in formula (10): Where S represents a continuous sequence of images during training, I t (p) and I t ′ (p) represents the pixel value of a certain pixel point in the original target frame and the reconstructed target frame respectively; <I1,…,I N > represents the image sequence obtained by chronological order in a training process, corresponding to the 1st frame to the Nth frame respectively.

22. The method for self-supervised monocular depth estimation based on feature sharing according to claim 20, characterized in that: The structural similarity loss function describes the similarity between the uncompressed original image and the compressed distorted image. The structural similarity loss function constrains the quality of the reconstructed image from three aspects: brightness, contrast, and structure. The calculation is as follows: SSIM(x,y)=[l(x,y)] α ·[c(x,y)] β ·[s(x,y)] γ (14) Where l(x,y) represents brightness similarity, c(x,y) represents contrast similarity, and s(x,y) represents structural similarity; x and y are coordinate values, parameters α = β = γ = 1, and u x ,u y are the means of x and y, σ x ,σ y are the variances of x and y, σ xy is the covariance of x and y; parameter c1=(k1L) 2 ,c2=(k2L) 2 , where L is the value range of pixel value; parameters k1 = 0.01, k2 = 0.03, c3 = 0.5c2; Get the similarity metric loss L between the two frames ssim (x,y) is as shown in formula (15):

23. The method for self-supervised monocular depth estimation based on feature sharing according to claim 20, characterized in that: The edge smoothing loss L smooth As shown in formula (16): where d t is the depth value of the frame at time t, is d t The mean of That is, the inverse depth information after average normalization; I t is the pixel value; and Represents the differential operation in the x dimension and y dimension respectively.

24. The method for self-supervised monocular depth estimation based on feature sharing according to claim 21, characterized in that: The modulus value of S is 3 or 5.

25. The method for self-supervised monocular depth estimation based on feature sharing according to claim 5, characterized in that: The storage medium includes a fixed memory or video memory, in which a space for storing features and pictures remains fixed.

26. A readable storage medium, characterized in that: The readable storage medium stores execution instructions, which, when executed by a processor, are used to implement the self-supervised monocular depth estimation method based on feature sharing as described in any one of claims 1 to 25.

27. A self-supervised monocular depth estimation system based on feature sharing, characterized in that: include: A memory storing a program for executing the self-supervised monocular depth estimation method based on feature sharing according to any one of claims 1 to 25; A processor; the processor executes the program.

28. An unmanned vehicle, characterized in that: include: An on-board processor, wherein the on-board processor executes the feature sharing-based self-supervised monocular depth estimation method as described in any one of claims 1-25.

Citation Information

Patent Citations

  • Monocular depth estimation method, device and equipment and storage medium

    CN108961327A

  • Rapid monocular depth estimation method suitable for embedded platform

    CN110599533A

  • Depth estimation method and device for automatic driving scene and autonomous vehicle

    CN111680554A

  • Pose estimation method based on self-supervised learning

    CN111325797A

  • Monocular video structure and motion prediction self-supervision method based on super resolution

    CN112270692A