Monocular depth estimation method and system based on convolutional neural network
By integrating local and global feature extraction using convolutional neural networks and visual models, the method enhances the accuracy of single-view depth estimation, overcoming the limitations of Transformer models in capturing long-range dependencies and local details.
Patent Information
- Application Number
- CN202510384233.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing monocular depth estimation method based on convolutional neural networks has problems such as information loss and limited receptive field when processing local and global information, resulting in insufficient estimation accuracy.
A monocular depth estimation method based on convolutional neural network is adopted to extract local and global features by combining convolutional neural networks and visual models, and feature fusion module is used to generate accurate monocular depth prediction maps.
It improves the accuracy of monocular depth estimation, overcomes the problems of information loss and limited receptive fields in traditional methods, and enhances the capture ability of local texture features and global dependence.
Smart Images

Figure CN120318289A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of image processing and deep learning, and particularly relates to a monocular depth estimation method and system based on a convolutional neural network. Background Art
[0002] In the fields of transportation and autonomous driving, a four-dimensional perspective (three-dimensional space plus time) is crucial for understanding driving behavior. External factors such as vehicles, pedestrians, road features, obstacles, and environmental conditions all affect driving decisions. Through various sensor technologies such as cameras, radars, lidar (LiDAR), ultrasonic sensors, and vehicle-to-everything (V2X), these information can be effectively obtained. These devices provide the necessary visual recognition, distance measurement, and precise positioning for advanced driver assistance systems (ADAS), supporting functions such as adaptive cruise control, collision prevention, and lane keeping. Depth information plays a core role in autonomous driving, not only helping the vehicle accurately measure and understand the surrounding environment, but also supporting real-time path planning and speed control. Depth perception technologies are divided into two categories: active (such as LiDAR and ToF cameras) and passive (such as stereo vision). Although these technologies provide accurate depth data, they also face problems such as high cost, high complexity, sensitivity to ambient light, and high maintenance requirements.
[0003] In the field of computer vision, due to the popularity of GPUs, the progress of deep learning technology, and the availability of big data, visual systems have developed rapidly and are widely used in traffic monitoring and autonomous driving. The problem of monocular image depth estimation is studied to explore how to predict a three-dimensional depth map from a two-dimensional image to improve the depth perception ability and adaptability of the model. Despite challenges such as insufficient diversity of training samples and difficulty in obtaining depth annotations, the research on monocular depth estimation still has important theoretical and application values. Monocular depth estimation based on deep learning has become the mainstream trend, and deep learning predicts scene depth information from a single image in an end-to-end manner. Deep learning methods usually adopt continuous regression to minimize the error between the actual depth and the predicted depth, thereby restoring the monocular depth map. By fusing features related to the appearance, geometry, semantics, and spatial relationships of objects, deep learning-based methods can effectively solve the ill-posed problem in monocular depth estimation.
[0004] Based on the basic network structure of convolutional neural networks (CNNs), numerous methods have designed encoder-decoder based architectures to solve the monocular depth estimation problem. However, the continuous downsampling in depth estimation may lead to the loss of some important feature information, which cannot be recovered in the decoder stage. To address this issue, many methods focus on improving the decoder to obtain a denser feature map and prevent information loss during the continuous upsampling process. Although these methods have improved in performance, CNN-based methods are still affected by limited receptive fields and less global representation. Therefore, obtaining dense long-range correlations in the encoder stage is crucial for achieving accurate depth estimation. Additionally, although Transformer has been successful in dealing with global dependencies in other tasks, due to the lack of spatial inductive bias, pure Transformer models lack the ability to model local information. Summary of the Invention
[0005] In view of the above technical problems, the present application provides a monocular depth estimation method and system based on a convolutional neural network, which overcomes the limitations of the Transformer model in solving the monocular depth estimation problem and improves the accuracy of the convolutional neural network for monocular depth estimation.
[0006] In a first aspect, an embodiment of the present application provides a monocular depth estimation method based on a convolutional neural network, including:
[0007] Obtain a two-dimensional image to be predicted;
[0008] Input the two-dimensional image into a preset depth estimation model, so that the depth estimation model uses a convolutional neural network model and a vision model to extract the local features and global features of the two-dimensional image respectively, and after feature fusion of the local features and global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image through a regression head;
[0009] Wherein, the depth estimation model is obtained by training an initial depth estimation model with a number of historical two-dimensional images and corresponding historical depth images, and the initial depth estimation model is jointly constructed by an initial convolutional neural network model, an initial vision model, an initial feature fusion module, and an initial regression head.
[0010] An embodiment of the present application provides a monocular depth estimation method based on a convolutional neural network. The global features and local features of a two-dimensional image are extracted through a preset depth estimation model, and a visual model is used to capture the global dependence and long-range correlation of the image, providing global context information for depth estimation. At the same time, a convolutional neural network model is used to extract the local texture features and details of the image, enhancing the model's perception ability for small-region features. Then, the obtained global features and local features are fused, and an accurate monocular depth prediction map is generated through a regression head. By effectively combining the ability of the convolutional neural network model to obtain local texture features of the image and the ability of the visual model to capture the long-range correlation between image pixels, the local and global information is fully integrated and utilized, so as to more accurately predict and restore the three-dimensional depth information in a single two-dimensional image, overcoming the limitations existing in the traditional Transformer model in solving the monocular depth estimation problem and improving the accuracy of the convolutional neural network for monocular depth estimation.
[0011] In a possible implementation manner, the depth estimation model uses a convolutional neural network model and a visual model to extract the local features and global features of the two-dimensional image respectively. After the local features and global features are fused through a feature fusion module, a monocular depth prediction map corresponding to the two-dimensional image is generated through a regression head, including:
[0012] Input the two-dimensional image into the convolutional neural network model, so that the convolutional neural network model extracts the local features of the two-dimensional image through a plurality of convolutional neural networks with different parameters, and generates a plurality of local feature maps with different dimensions;
[0013] Input the two-dimensional image into the visual model, so that the visual model extracts the global features of the two-dimensional image through a plurality of sub-visual models with different parameters, and generates a plurality of global feature maps with different dimensions, where each of the global feature maps and each of the local feature maps has a one-to-one correspondence based on the dimension;
[0014] Perform first feature fusion on each of the global feature maps and the corresponding local feature maps according to the one-to-one correspondence, and generate a plurality of first fusion feature maps with different dimensions;
[0015] Input each of the first fusion feature maps into the feature fusion module, so that the feature fusion module performs second feature fusion on each of the first fusion feature maps through a multi-scale feature fusion module with different parameters to obtain a second fusion feature map;
[0016] Input the second fusion feature map into the regression head, so that the regression head generates a monocular depth prediction map corresponding to the two-dimensional image.
[0017] An embodiment of the present application provides a method for depth prediction using a depth estimation model. In the depth estimation model, multiple convolutional networks with different parameters are deployed to extract local features of a two-dimensional image, and correspondingly, multiple sub-visual models with different parameters are deployed to extract global features of the two-dimensional image. The setting of multiple models with different parameters enables the constructed depth estimation model to extract features of the two-dimensional image in different dimensions, which is conducive to fully mining the useful information of the two-dimensional image and improving the prediction accuracy of the model. Correspondingly, in the feature fusion stage, the first fusion feature maps in each dimension are unified in dimension through several multi-scale feature fusion modules with different parameters, effectively integrating local and global information in different dimensions. By simultaneously improving the encoder stage (feature extraction stage) and the decoder stage (feature fusion stage), the embodiment of the present application avoids the problem of loss of some important feature information caused by continuous downsampling in traditional depth estimation and improves the prediction accuracy of the model.
[0018] Further, the convolutional neural networks are arranged and connected in sequence based on the output dimension size. The convolutional neural network model extracts local features of the two-dimensional image through several convolutional neural networks with different parameters, generating several local feature maps with different dimensions, including:
[0019] Each of the convolutional neural networks sequentially generates the local feature maps according to the arranged connection order, where the input data of any convolutional neural network is the local feature map output by the previous convolutional neural network, and the input data of the convolutional neural network ranked first is the two-dimensional image.
[0020] The embodiment of the present application further defines the connection structure and input-output process between each convolutional neural network. During the feature extraction process, local feature maps are sequentially generated according to the arranged connection order, and each generated local feature map participates in the feature extraction process of the next neural network as the input data of the next neural network. Such a setting can gradually improve the fineness of feature extraction, enabling each neural network model to fully mine the local feature information of the two-dimensional image and improving the accuracy of subsequent predictions.
[0021] Further, the sub-visual models are arranged and connected in sequence based on the output dimension size. The visual model extracts global features of the two-dimensional image through several sub-visual models with different parameters, generating several global feature maps with different dimensions, including:
[0022] Each of the sub-visual models sequentially generates the global feature maps according to the arranged connection order, where the input data of any sub-visual model is the global feature map output by the previous sub-visual model, and the input data of the sub-visual model ranked first is the two-dimensional image.
[0023] Corresponding to local feature extraction, in terms of global feature extraction, embodiments of the present application also define the connection structure and input / output process between each sub-visual model. During the feature extraction process, global feature maps are sequentially generated according to the order of arranged connections, and each generated global feature map participates in the feature extraction process of the next sub-visual model as the input data of the next sub-visual model. Such a setting can gradually improve the fineness of feature extraction, enabling each sub-visual model to fully exploit the global feature information of the two-dimensional image and improving the accuracy of subsequent predictions.
[0024] In a possible implementation manner, the first feature fusion of each of the global feature maps with the corresponding local feature map according to the one-to-one correspondence to generate a plurality of first fusion feature maps with different dimensions includes:
[0025] Reducing the feature channels of each of the global feature maps and each of the local feature maps to their respective corresponding preset channel numbers to obtain each corresponding channel-reduced global feature map and each corresponding channel-reduced local feature map;
[0026] Performing first feature fusion of each of the channel-reduced global feature maps with the corresponding channel-reduced local feature map to obtain a plurality of first fusion feature maps with different dimensions.
[0027] In embodiments of the present application, before performing feature fusion on the global feature map and the local feature map, the feature channels of each feature map are first reduced to a specified number to ensure the channel consistency between the global feature map and the local feature map during the fusion process, improving the effectiveness of feature fusion and the accuracy of subsequent predictions.
[0028] Further, during the first feature fusion process, each of the channel-reduced global feature maps and the channel-reduced local feature map are subjected to first feature fusion by a plurality of selective feature fusion modules with different parameters to obtain a plurality of first fusion feature maps with different dimensions. Among them, any selective feature fusion module performs first feature fusion on the channel-reduced global feature map and the channel-reduced local feature map, specifically as follows:
[0029] Performing preliminary feature fusion on the channel-reduced global feature map and the channel-reduced local feature map to obtain a preliminary fusion feature map;
[0030] Performing successive convolution operations and normalization processing operations on the preliminary fusion feature map in sequence to generate a normalized feature map;
[0031] Based on the selective attention mechanism and the channel values in the normalized feature map, perform an element-wise multiplication fusion operation on the local attention features and global attention features in the normalized feature map respectively, and add the obtained results to obtain the first fusion feature map.
[0032] The embodiments of the present application introduce selective feature fusion modules with multiple different parameters to perform feature fusion in different dimensions, which is adapted to the multi-dimensional feature extraction of the embodiments of the present application. At the same time, in the specific feature fusion process, the embodiments of the present application effectively distinguish the local attention features and global attention features in the normalized feature map based on the selective attention mechanism and channel values, and perform an element-wise multiplication fusion operation on the two types of features respectively. Finally, the calculation results are added to achieve the feature fusion of local features and global features. Different from the prior art that focuses on improving the decoder method, the embodiments of the present application overcome the problems of information loss, less global or local information in monocular depth estimation by improving the encoder structure, and improve the accuracy of convolutional neural networks for monocular depth estimation.
[0033] In a possible implementation manner, the multi-scale feature fusion modules are arranged and connected in sequence based on the output dimension size. The feature fusion module performs second feature fusion on each of the first fusion feature maps through a plurality of multi-scale feature fusion modules with different parameters to obtain a second fusion feature map, including:
[0034] Each of the multi-scale feature fusion modules performs second feature fusion in sequence according to the arranged connection order. Among them, the input data of any multi-scale feature fusion module is the output data of the previous multi-scale feature fusion module and the corresponding first fusion feature map. The input data of the multi-scale feature fusion module ranked first is the corresponding first fusion feature map, and the output data of the multi-scale feature fusion module ranked last is the second fusion feature map.
[0035] The embodiments of the present application further define the connection structure and input-output process between each multi-scale feature fusion module. Since several first fusion feature maps with different dimensions are obtained after feature fusion, several corresponding multi-scale feature fusion modules are required to perform secondary fusion on each first fusion feature map. In the embodiments of the present application, each multi-scale feature fusion module performs second feature fusion in sequence according to the arranged connection order, gradually superimposing the useful information between different feature maps and restoring the dimension. By gradually doubling the size of the feature map, the accuracy and details of the depth information are effectively restored, and the accuracy of depth prediction is improved.
[0036] Further, in the process of the second feature fusion, the output data end of each multi-scale feature fusion module is connected to a corresponding CBAM hybrid domain attention module, and the CBAM hybrid domain attention module is used to process the output data of the corresponding multi-scale feature fusion module, combine the attention of the spatial domain and the channel domain in the output data, and input the processed output data into the next multi-scale feature fusion module.
[0037] In the embodiment of the present application, a corresponding CBAM hybrid domain attention module is connected to the output data end of each multi-scale feature fusion module. By combining the attention of the spatial domain and the channel domain in the output data, the depth estimation performance of the model is further improved.
[0038] In a possible implementation manner, training the initial depth estimation model with a plurality of historical two-dimensional images and corresponding historical depth images to obtain the depth estimation model includes:
[0039] Taking each of the historical depth images as the label of the corresponding historical two-dimensional image, and then constructing a training data set based on each of the historical two-dimensional images;
[0040] Inputting the training data set into the initial depth estimation model, so that the initial depth estimation model generates a plurality of corresponding predicted depth images;
[0041] Calculating the loss according to each of the predicted depth images and the corresponding labels in the training data set, and then optimizing the parameters by backpropagation inside the initial depth estimation model according to the result of the loss calculation to obtain the depth estimation model.
[0042] In a second aspect, correspondingly, the embodiment of the present application provides a monocular depth estimation system based on a convolutional neural network, including an acquisition module and a depth estimation module;
[0043] The acquisition module is used to acquire a two-dimensional image to be predicted;
[0044] The depth estimation module is used to input the two-dimensional image into a preset depth estimation model, so that the depth estimation model uses a convolutional neural network model and a vision model to extract local features and global features of the two-dimensional image respectively, and after fusing the local features and the global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image through a regression head;
[0045] The depth estimation model is obtained by training an initial depth estimation model with a number of historical two-dimensional images and corresponding historical depth images. The initial depth estimation model is jointly constructed by an initial convolutional neural network model, an initial vision model, an initial feature fusion module, and an initial regression head. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 : A schematic flow chart of a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0047] Figure 2 : A schematic structural diagram of a convolutional neural network model in a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0048] Figure 3 : A schematic structural diagram of a vision model in a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0049] Figure 4 : A schematic structural diagram of a selective feature fusion module in a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0050] Figure 5 : A schematic structural diagram of a multi-scale feature fusion module in a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0051] Figure 6 : A schematic diagram of the depth prediction effect of a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0052] Figure 7 : A schematic diagram of the overall model architecture of a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0053] Figure 8 : A schematic structural diagram of the decoder of the depth estimation model in a monocular depth estimation method based on a convolutional neural network provided by an embodiment of the present application.
[0054] Figure 9 : A schematic structural diagram of a monocular depth estimation system based on a convolutional neural network provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present application.
[0056] It should be noted that the step numbers in the text are only for the convenience of explaining specific embodiments and do not serve as a limitation on the execution order of the steps. In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features.
[0057] Embodiment 1:
[0058] As Figure 1 shown, Embodiment 1 provides a monocular depth estimation method based on a convolutional neural network, including steps S1 - S2:
[0059] Step S1, obtain a two-dimensional image to be predicted;
[0060] Step S2, input the two-dimensional image into a preset depth estimation model, so that the depth estimation model uses a convolutional neural network model and a vision model to extract the local features and global features of the two-dimensional image respectively, and after fusing the local features and global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image through a regression head;
[0061] Among them, the depth estimation model is obtained by training an initial depth estimation model with a number of historical two-dimensional images and corresponding historical depth images, and the initial depth estimation model is jointly constructed by an initial convolutional neural network model, an initial vision model, an initial feature fusion module, and an initial regression head.
[0062] The embodiment of the present application provides a monocular depth estimation method based on a convolutional neural network. The global features and local features of a two-dimensional image are extracted through a preset depth estimation model, and a vision model is used to capture the global dependence and long-range correlation of the image, providing global context information for depth estimation. At the same time, a convolutional neural network model is used to extract the local texture features and details of the image, enhancing the model's perception ability of small-region features. Then, the extracted global features and local features are subjected to feature fusion, and an accurate monocular depth prediction map is generated through a regression head. The embodiment of the present application effectively combines the ability of the convolutional neural network model to obtain local texture features of the image and the ability of the vision model to capture the long-range correlation between image pixels, fully integrating and utilizing local and global information, thereby more accurately predicting and restoring the three-dimensional depth information in a single two-dimensional image, overcoming the limitations existing in the traditional Transformer model in solving the problem of monocular depth estimation, and improving the accuracy of the convolutional neural network for monocular depth estimation.
[0063] In a preferred embodiment, the convolutional neural network model is a ConvNeXt model, and the vision model is a Swin Transformer model.
[0064] In a possible implementation manner, in step S2, the depth estimation model uses a convolutional neural network model and a vision model to extract the local features and global features of the two-dimensional image respectively, and after performing feature fusion on the local features and global features through a feature fusion module, a monocular depth prediction map corresponding to the two-dimensional image is generated through a regression head, including steps S201 - S205:
[0065] Step S201: Input the two-dimensional image into the convolutional neural network model, so that the convolutional neural network model extracts the local features of the two-dimensional image through a plurality of convolutional neural networks with different parameters, and generates a plurality of local feature maps with different dimensions;
[0066] Step S202: Input the two-dimensional image into the vision model, so that the vision model extracts the global features of the two-dimensional image through a plurality of sub-vision models with different parameters, and generates a plurality of global feature maps with different dimensions, where there is a one-to-one correspondence based on dimensions between each global feature map and each local feature map;
[0067] Step S203: Perform first feature fusion on each global feature map and its corresponding local feature map according to the one-to-one correspondence, and generate a plurality of first fusion feature maps with different dimensions;
[0068] Step S204: input each of the first fused feature maps into the feature fusion module, so that the feature fusion module performs second feature fusion on each of the first fused feature maps through a plurality of multi-scale feature fusion modules with different parameters to obtain a second fused feature map;
[0069] Step S205: input the second fused feature map to the regression head, so that the regression head generates a monocular depth prediction map corresponding to the two-dimensional image.
[0070] The embodiment of the present application provides a method for depth prediction using a depth estimation model, in which a plurality of convolutional networks with different parameters are deployed in the depth estimation model to extract local features of a two-dimensional image, and a plurality of sub-visual models with different parameters are correspondingly deployed to extract global features of the two-dimensional image. The setting of a plurality of models with different parameters enables the constructed depth estimation model to extract features of the two-dimensional image in different dimensions, which is conducive to fully mining the useful information of the two-dimensional image and improving the prediction accuracy of the model. Accordingly, in the feature fusion stage, the first fused feature maps of each dimension are unified in dimension through a multi-scale feature fusion module with several different parameters, effectively integrating local and global information of different dimensions. The embodiment of the present application improves the encoder stage (feature extraction stage) and the decoder stage (feature fusion stage) at the same time, thereby avoiding the problem of loss of some important feature information caused by continuous downsampling in traditional depth estimation and improving the prediction accuracy of the model.
[0071] Further, in step S201, each of the convolutional neural networks is sequentially connected based on the output dimension size, and the convolutional neural network model extracts local features of the two-dimensional image through convolutional neural networks with several different parameters to generate several local feature maps of different dimensions, including:
[0072] Each of the convolutional neural networks generates the local feature map in sequence according to the order of arrangement and connection, wherein the input data of any convolutional neural network is the local feature map output by the previous convolutional neural network, and the input data of the convolutional neural network ranked first is the two-dimensional image.
[0073] The embodiment of the present application further defines the connection structure and input-output process between each convolutional neural network. In the feature extraction process, local feature maps are generated in sequence according to the order of arrangement and connection, and each local feature map generated is used as the input data of the next neural network to participate in the feature extraction process of the next neural network. Such a setting can gradually improve the precision of feature extraction, so that each neural network model can fully mine the local feature information of the two-dimensional image and improve the accuracy of subsequent predictions.
[0074] In a preferred embodiment, in step S201, the structural schematic diagram of the convolutional neural network model is as follows Figure 2 shown, including four ConvNeXt models with different parameters, which extract multiple local feature maps of different dimensions from the two-dimensional image by performing downsampling, convolution, and regression operations on the input data.
[0075] Furthermore, in step S202, each of the sub-visual models is arranged and connected in sequence based on the output dimension size. The visual model extracts the global features of the two-dimensional image through several sub-visual models with different parameters, generating several global feature maps of different dimensions, including:
[0076] Each of the sub-visual models generates the global feature map in sequence according to the arranged connection order. Among them, the input data of any sub-visual model is the global feature map output by the previous sub-visual model, and the input data of the sub-visual model ranked first is the two-dimensional image.
[0077] Corresponding to local feature extraction, in terms of global feature extraction, the embodiments of the present application also define the connection structure and input-output process between each sub-visual model. During the feature extraction process, global feature maps are generated in sequence according to the arranged connection order, and each generated global feature map participates in the feature extraction process of the next sub-visual model as the input data of the next sub-visual model. Such a setting can gradually improve the fineness of feature extraction, enabling each sub-visual model to fully mine the global feature information of the two-dimensional image and improving the accuracy of subsequent predictions.
[0078] In a preferred embodiment, in step S202, the model structure of the visual model is as follows Figure 3 shown, including four Swin Transformer models with different parameters, which perform operations such as layer normalization and multi-layer perception on the input data through multiple cyclic iterations based on the attention mechanism, and extract multiple global feature maps of different dimensions from the two-dimensional image.
[0079] In a possible implementation manner, in step S203, the first feature fusion of each of the global feature maps with the corresponding local feature map is performed according to the one-to-one correspondence, generating several first fusion feature maps of different dimensions, including:
[0080] Reduce the feature channels of each of the global feature maps and each of the local feature maps to their respective corresponding preset channel numbers, obtaining each corresponding channel-reduced global feature map and each corresponding channel-reduced local feature map;
[0081] Perform the first feature fusion of each of the channel-reduced global feature maps with the corresponding channel-reduced local feature map, obtaining several first fusion feature maps of different dimensions.
[0082] In the embodiment of the present application, before fusing the global feature map and the local feature map, each feature map is first subjected to channel reduction to reduce the feature channels to a specified number, ensuring the channel consistency between the global feature map and the local feature map during the fusion process, and improving the effectiveness of feature fusion and the accuracy of subsequent prediction.
[0083] In a preferred embodiment, the convolution kernel size of the channel reduction module is 1, the stride is 1, the padding is 0, and the output channel is a convolution of C. The formula is as follows:
[0084] c i = Conv 1×1 (s i (i = 1, 2, 3, 4)
[0085] Furthermore, during the first feature fusion process, the corresponding channel-reduced global feature map and the channel-reduced local feature map are respectively subjected to first feature fusion by a plurality of selective feature fusion modules with different parameters to obtain a plurality of first fusion feature maps with different dimensions. Among them, any selective feature fusion module performs first feature fusion on the channel-reduced global feature map and the channel-reduced local feature map, specifically as follows:
[0086] Perform preliminary feature fusion on the channel-reduced global feature map and the channel-reduced local feature map to obtain a preliminary fusion feature map;
[0087] Perform consecutive convolution operations and normalization processing operations on the preliminary fusion feature map in sequence to generate a normalized feature map;
[0088] Based on the selective attention mechanism and the channel values in the normalized feature map, perform element-wise multiplication fusion operations on the local attention feature and the global attention feature in the normalized feature map respectively, and add the obtained results to obtain the first fusion feature map.
[0089] The embodiment of the present application introduces a plurality of selective feature fusion modules with different parameters to perform feature fusion of different dimensions, which is adapted to the multi-dimensional feature extraction of the embodiment of the present application. At the same time, in the specific feature fusion process, the embodiment of the present application effectively distinguishes the local attention feature and the global attention feature in the normalized feature map based on the selective attention mechanism and the channel values, and performs element-wise multiplication fusion operations on the two features respectively. Finally, the calculation results are added to achieve the feature fusion of local features and global features. Different from the prior art that focuses on improving the decoder method, the embodiment of the present application overcomes problems such as information loss, less global or local information in monocular depth estimation by improving the encoder structure, and improves the accuracy of convolutional neural networks for monocular depth estimation.
[0090] In a preferred embodiment, in step S203, the structural schematic diagram of the selective feature fusion module is as follows Figure 4 shown, including operations such as preliminary feature fusion, convolution and batch normalization processing, Sigmoid function activation, attention fusion, output features, etc., specifically as follows:
[0091] 1. Feature fusion:
[0092] H0 = concat(G, L)
[0093] Where: G is the feature map corresponding to the output of different stages of the ConvNeXt model branch processed by the channel reduction module; L is the feature map corresponding to the output of different stages of the Swin transformer model branch processed by the channel reduction module.
[0094] 2. Convolution and batch normalization processing:
[0095] H1 = ReLu(BatchNorm(Conv 3×3 (H0)))
[0096] H2 = ReLu(BatchNorm(Conv 3×3 (H1)))
[0097] Where: H1 and H2 are the outputs of two consecutive convolutional blocks, ReLU() is the activation function; BatchNorm() is batch normalization; Conv3×3() is the convolution operation.
[0098] 3. Sigmoid activation function:
[0099] A = σ(Conv(H2))
[0100] Where: σ is the Sigmoid activation function.
[0101] 4. Selective attention mechanism:
[0102] A is divided into two parts according to the channel values, corresponding to the global attention A global and the local attention A local .
[0103]
[0104] Where: is the element-wise multiplication fusion operation.
[0105] 5. Output features:
[0106] H b = H global + H local
[0107] Wherein: H b is the weighted feature vector.
[0108] In a possible implementation manner, in step S204, each of the multi-scale feature fusion modules is arranged and connected in sequence based on the output dimension size, and the feature fusion module performs second feature fusion on each of the first fusion feature maps through a plurality of multi-scale feature fusion modules with different parameters to obtain a second fusion feature map, including:
[0109] Each of the multi-scale feature fusion modules performs second feature fusion in sequence according to the arranged connection order. Among them, the input data of any multi-scale feature fusion module is the output data of the previous multi-scale feature fusion module and the corresponding first fusion feature map. The input data of the multi-scale feature fusion module ranked first is the corresponding first fusion feature map, and the output data of the multi-scale feature fusion module ranked last is the second fusion feature map.
[0110] The embodiments of the present application further define the connection structure and input-output process between each multi-scale feature fusion module. Since a plurality of first fusion feature maps with different dimensions are obtained after feature fusion, a plurality of multi-scale feature fusion modules corresponding to the dimensions are required to perform secondary fusion on each first fusion feature map. In the embodiments of the present application, each multi-scale feature fusion module performs second feature fusion in sequence according to the arranged connection order, gradually superimposing the useful information between different feature maps and restoring the dimension. By gradually doubling the size of the feature map, the accuracy and details of the depth information are effectively restored, and the accuracy of depth prediction is improved.
[0111] Further, during the second feature fusion process, the output data end of each multi-scale feature fusion module is connected to a corresponding CBAM hybrid domain attention module. The CBAM hybrid domain attention module is used to process the output data of the corresponding multi-scale feature fusion module, combine the attention in the spatial domain and the channel domain of the output data, and input the processed output data into the next multi-scale feature fusion module.
[0112] The embodiments of the present application connect a corresponding CBAM hybrid domain attention module to the output data end of each multi-scale feature fusion module. By combining the attention in the spatial domain and the channel domain of the output data, the depth estimation performance of the model is further improved.
[0113] In a preferred embodiment, in step S204, the schematic diagram of the model structure of the multi-scale feature fusion module is as Figure 5As shown, it includes operations such as residual convolution, feature fusion, convolution, and upsampling. Taking the second multi-scale feature fusion module in the embodiment of the present application as an example, the input data of the second multi-scale feature fusion module is the first fusion feature map output by the third selective feature fusion (SFF) module in this embodiment and the processed feature output by the first CBAM hybrid domain attention module. The output data of the second multi-scale feature fusion module is then input to the second CBAM hybrid domain attention module. The specific formulas for each operation are as follows:
[0114] 1. Residual convolution unit 1:
[0115] R = ReLu(Conv 3×3 (ReLu(Conv 3×3 (X3)))) × 2
[0116] Where: X3 is the first fusion feature map processed by the third selective feature fusion (SFF) module in this embodiment.
[0117] 2. Feature fusion:
[0118] X c = R + M1
[0119] Where: M1 is the processed feature output by the first CBAM hybrid domain attention module in this embodiment.
[0120] 3. Residual convolution unit 2:
[0121] R1 = ReLu(Conv 3×3 (ReLu(Conv 3×3 (X c )))) × 2
[0122] 4. Convolution:
[0123] Y = Conv 3×3 (R1)
[0124] 5. Upsampling:
[0125] m2 = Upsample(Y)
[0126] Where: m2 is the output data of the second multi-scale feature fusion module in this embodiment.
[0127] In a preferred embodiment, in step S205, the specific working process of the regression head is as follows:
[0128] 1. Upsampling:
[0129] X u = Upsample(M4)
[0130] Where: M4 is the processed feature output by the fourth (i.e., the last) CBAM hybrid domain attention module in this embodiment, and Upsample is an operation in the pytorch library. Specifically, the function is that the number of channels remains unchanged after the input feature passes through this operation, and the resolution is doubled, that is, if the input resolution is H and W, the output resolution is 2H and 2W.
[0131] 2. Dropout:
[0132] X d = Dropout(X u )
[0133] 3. First convolution and activation:
[0134] X c1 = ReLU(Conv 3×3 (X d ))
[0135] 4. Feature fusion:
[0136] X f = X d + X c1
[0137] 5. Second Dropout:
[0138] X d2 = Dropout(X f )
[0139] 6. Second convolution:
[0140] X c2 = Conv 3×3 (X d2 )
[0141] 7. Sigmoid activation function:
[0142] Y = σ(X c2 )
[0143] Where: σ is the Sigmoid activation function.
[0144] In the actual application process, the schematic diagram of the depth prediction effect of using the depth estimation model provided by the embodiment of the present application is as Figure 6 shown, where the upper picture is the original data picture and the lower picture is the predicted depth picture.
[0145] In a possible implementation manner, the training of the initial depth estimation model through a plurality of historical two-dimensional images and corresponding historical depth images to obtain the depth estimation model includes:
[0146] Use each of the historical depth images as the label for the corresponding historical two-dimensional images, and then construct a training dataset based on each of the historical two-dimensional images;
[0147] Input the training dataset into the initial depth estimation model so that the initial depth estimation model generates a number of corresponding predicted depth images;
[0148] Calculate the loss based on each of the predicted depth images and the corresponding labels in the training dataset, and then optimize the parameters by backpropagation inside the initial depth estimation model according to the result of the loss calculation to obtain the depth estimation model.
[0149] In a preferred embodiment, the overall model architecture diagrams of the initial depth estimation model and the depth estimation model are as shown in Figure 7 shown, including an encoder, a decoder, and a regression head. Among them, the encoder includes four ConvNeXt models with different parameters, four Swin Transformer models with different parameters, eight corresponding channel reduction modules, and four corresponding selective feature fusion modules (SFF). The decoder includes four multi-scale feature fusion modules with different parameters and four corresponding CBAM hybrid domain attention modules. The structural diagram of the decoder is as shown in Figure 8 shown. The specific steps for training the initial depth estimation model are as follows:
[0150] 1. Build a diverse dataset, which can be sourced from public datasets or captured in actual scenarios. Use data augmentation techniques such as rotation, flipping, color adjustment, and noise injection for data augmentation to ensure the robustness of the model. Subsequently, adjust the data dimensions to the standard input size of the model through methods such as random cropping, padding, or compression to ensure that the data fits the model requirements. After this step of processing, the data dimensions should be (H, W, 3), and the corresponding label dimensions should be (H, W, 1).
[0151] 2. Input the image data into the model encoder. At Stage1, after being independently processed by the Swin Transformer module and the ConvNeXt module respectively, the output dimension is (H / 4, W / 4, C). The outputs here enter Stage2 and the channel reduction module along the model respectively. At Stage2, after being independently processed by the Swin Transformer module and the Con-vNeXt module respectively, the output dimension is (H / 8, W / 8, 2C). The outputs here enter Stage3 and the channel reduction module along the model respectively. At Stage3, after being independently processed by the Swin Transformer module and the ConvNeXt module respectively, the output dimension is (H / 16, W / 16, 4C). The outputs here enter Stage4 and the channel reduction module along the model respectively. At Stage4, after being independently processed by the Swin Transformer module and the ConvNeXt module respectively, the output dimension is (H / 32, W / 32, 8C). The output here enters the channel reduction module along the model. Among them, H and W are the height and width data of the input RGB image data, C is the number of channels of the output features of the Swin Transformer and ConvNeXt model branches at Stage1, and the number of channels of the output features at Stage2 - Stage4 are 2C, 4C, and 8C respectively, which are determined by the model and are preset values.
[0152] 3. After the output data of the Swin Transformer model and the ConvNeXt model in the four stages of Stage1 - Stage4 are respectively processed by the channel reduction module, the output dimensions are adjusted to (H / 4, W / 4, C), (H / 8, W / 8, C), (H / 16, W / 16, C), and (H / 32, W / 32, C).
[0153] 4. According to the four stages of the encoder at Stage1 - Stage4, the output data of the eight channel reduction modules are respectively input into the four Selective Feature Fusion (SFF) modules. The features corresponding to the ConvNeXt model branch enter the local feature entry of the Selective Feature Fusion (SFF) module, and the features corresponding to the Swin Transformer model branch enter the global feature entry of the Selective Feature Fusion (SFF) module. After being processed by the Selective Feature Fusion (SFF) module, the output dimensions corresponding to the four stages of Stage1 - Stage4 are (H / 4, W / 4, C), (H / 8, W / 8, C), (H / 16, W / 16, C), and (H / 32, W / 32, C), that is, the first fusion feature maps of four different dimensions.
[0154] 5. According to the dimension of the output data, the features corresponding to the encoder Stage4 with the dimension of (H / 32, W / 32, C) in the output of each Selective Feature Fusion (SFF) module enter the multi-scale feature fusion module in the decoder Stage1. After processing, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (H / 16, W / 16, C). Then, after being processed by the CBAM hybrid domain attention module, the dimension remains unchanged. After that, it enters the multi-scale feature fusion module in the decoder Stage2 and is fused with the features corresponding to the encoder Stage3 with the dimension of (H / 16, W / 16, C) in the output of the Selective Feature Fusion (SFF) module. After processing, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (H / 8, W / 8, C). Then, after being processed by the CBAM hybrid domain attention module, the dimension remains unchanged. After that, it enters the multi-scale feature fusion module in the decoder Stage3 and is fused with the features corresponding to the encoder Stage2 with the dimension of (H / 8, W / 8, C) in the output of the Selective Feature Fusion (SFF) module. After processing, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (H / 4, W / 4, C). Then, after being processed by the CBAM hybrid domain attention module, the dimension remains unchanged. After that, it enters the multi-scale feature fusion module in the decoder Stage1 and is fused with the features corresponding to the encoder Stage4 with the dimension of (H / 4, W / 4, C) in the output of the Selective Feature Fusion (SFF) module. After processing, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (H / 2, W / 2, C). Then, after being processed by the CBAM hybrid domain attention module, the final second fusion feature map is obtained.
[0155] 6. Input the second fusion feature map into the regression head. After being processed by the regression head, a depth prediction map with the dimension of (H, W, 1) is output.
[0156] 7. Calculate the loss between the output depth prediction map with the dimension of (H, W, 1) and the label with the dimension of (H, W, 1). Optimize the parameters according to the loss value through backpropagation, and then input new two-dimensional image data to repeat the above process. The final trained model is obtained through continuous iteration.
[0157] Among them, during the loss calculation process, one of the loss functions that can be selected is the scale-invariant loss proposed by Eigen, which is defined as follows:
[0158]
[0159] Where: n is the number of valid pixels in the two-dimensional picture; y is the true depth value; y i is the predicted depth value.
[0160] In a preferred embodiment, based on the above model training method and model structure, four different scales of depth estimation models can be constructed to handle image data of different sizes. Specifically, for the Swin Transformer model, the number of each module in the Swin Transformer model branches Stage1 - Stage4 can be selected as Swin-T: [2, 2, 6, 2], Swin-S: [2, 2, 18, 2], Swin-B: [2, 2, 18, 2], Swin-L: [2, 2, 18, 2]; the number of heads in the multi-head attention mechanism of each module in Stage1 - Stage4 can be correspondingly selected as Swin-T: [3, 6, 12, 24], Swin-S: [3, 6, 12, 24], Swin-B: [4, 8, 16, 32], Swin-L: [6, 12, 24, 48]; the number of channels C of each module in Stage1 - Stage4 can be correspondingly selected as Swin-T: [96, 192, 384, 768], Swin-S: [96, 192, 384, 768], Swin-B: [128, 256, 512, 1024], Swin-L: [192, 384, 768, 1536]; for the ConvNeXt model, the number of each module in the ConvNeXt model branches Stage1 - Stage4 can be selected as ConvNeXt-T: [3, 3, 9, 3], ConvNeXt-S: [3, 3, 27, 3], ConvNeXt-B: [3, 3, 27, 3], ConvNeXt-L: [3, 3, 27, 3] to match the scale of the number of each module in the Swin Transformer model branches Stage1 - Stage4; the number of channels C of each module in Stage1 - Stage4 can be selected as ConvNeXt-T: [96, 192, 384, 768], ConvNeXt-S: [96, 192, 384, 768], ConvNeXt-B: [128, 256, 512, 1024], ConvNeXt-L: [192, 384, 768, 1536] to match the scale of the number of channels of each module in the Swin Transformer model branches Stage1 - Stage4; according to the above parameter selections of the Swin Transformer model branches and ConvNeXt model branches, there can be four different scale models, namely SC-T, SC-S, SC-B, and SC-L.
[0161] Taking the SC-T model as an example, that is, the Swin Transformer and ConvNeXt model branches respectively select the Swin-T and ConvNeXt-T model backbones; the data input dimension is selected as (448, 1344, 3), and the label dimension is (448, 1344, 1); the dataset is selected as the KITTI dataset, then the training process of the SC-T model is specifically as follows:
[0162] 1. Use data augmentation techniques such as rotation, flipping, color adjustment, and noise injection to perform data augmentation on the KITTI dataset to ensure the robustness of the model. Subsequently, adjust the data dimension to the standard input size of the model through methods such as random cropping, padding, or compression to ensure that the data adapts to the model requirements. After this step of processing, the data dimension should be (448, 1344, 3), and the corresponding label dimension should be (448, 1344, 1).
[0163] 2. Input the image data into the model encoder. At Stage1, after being independently processed by the Swin Transformer module and the ConvNeXt module respectively, the output dimension is (112, 336, 96). The outputs here respectively enter Stage2 and the channel reduction module along the model. At Stage2, after being independently processed by the Swin Transformer module and the ConvNeXt module respectively, the output dimension is (56, 168, 192). The outputs here respectively enter Stage3 and the channel reduction module along the model. At Stage3, after being independently processed by the Swin Transformer module and the Con-vNeXt module respectively, the output dimension is (28, 84, 384). The outputs here respectively enter Stage4 and the channel reduction module along the model. At Stage4, after being independently processed by the SwinTransformer module and the ConvNe-Xt module respectively, the output dimension is (14, 42, 768). The output here enters the channel reduction module along the model.
[0164] 3. After the output data of the Swin Transformer model and the ConvNeXt model in the four stages of Stage1 - Stage4 are respectively processed by the channel reduction module, the output dimensions are respectively (112, 336, 96), (56, 168, 192), (28, 84, 384), (14, 42, 768).
[0165] 4. The output data processed by each channel reduction module enters the Selective Feature Fusion (SFF) module respectively. The features corresponding to the ConvNeXt model branch enter the local feature entrance of the Selective Feature Fusion (SFF) module, and the features corresponding to the SwinTransformer model branch enter the global feature entrance of the Selective Feature Fusion (SFF) module. After being processed by the Selective Feature Fusion (SFF) module, the output dimensions corresponding to the four stages of the encoder, Stage1 - Stage4, are (112, 336, 96), (56, 168, 192), (28, 84, 384), and (14, 42, 768) respectively, that is, the first fusion feature maps of four different dimensions.
[0166] 5. The features with a dimension size of (14, 42, 96) corresponding to the encoder Stage4 in the output of each Selective Feature Fusion (SFF) module enter the multi-scale feature fusion module in the decoder Stage1. After being processed, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (28, 84, 96). Then, after being processed by the CBAM mixed-domain attention module, the dimension remains unchanged. After that, it enters the multi-scale feature fusion module in the decoder Stage2 and is fused with the features with a dimension size of (28, 84, 96) corresponding to the encoder Stage3 in the output of the Selective Feature Fusion (SFF) module. After being processed, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (56, 168, 96). Then, after being processed by the CBAM mixed-domain attention module, the dimension remains unchanged. After that, it enters the multi-scale feature fusion module in the decoder Stage3 and is fused with the features with a dimension size of (56, 168, 96) corresponding to the encoder Stage2 in the output of the Selective Feature Fusion (SFF) module. After being processed, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (112, 336, 96). Then, after being processed by the CBAM mixed-domain attention module, the dimension remains unchanged. After that, it enters the multi-scale feature fusion module in the decoder Stage1 and is fused with the features with a dimension size of (112, 336, 96) corresponding to the encoder Stage4 in the output of the Selective Feature Fusion (SFF) module. After being processed, the resolution doubles and the number of channels remains unchanged, and the output feature dimension is (224, 672, 96). Then, after being processed by the CBAM mixed-domain attention module, the final second fusion feature map is obtained.
[0167] 6. The second fusion feature map is input into the regression head. After being processed, a depth prediction map with an output dimension of (448, 1344, 1) is obtained.
[0168] 7. Calculate the loss between the depth prediction map with a dimension of (448, 1344, 1) output by the regression head and the label with a dimension of (448, 1344, 1), and optimize the parameters through backpropagation according to the loss value. Continuously iterate to obtain the finally trained model.
[0169] In addition, during the actual application process, according to the requirements for model performance and computational efficiency, the Selective Feature Fusion (SFF) module and the CBAM hybrid domain attention module in the embodiments of the present application can be selectively used. When the Selective Feature Fusion (SFF) module is not enabled, directly add the features output by the corresponding Swin Transformer and ConvNeXt model branches in the same stage processed by the channel reduction module; when the CBAM hybrid domain attention is not enabled, simply skip it.
[0170] When tested on the KITTI dataset, when the SFF module is disabled and the CBAM module is enabled, compared with enabling both modules, the performance of the model in terms of Abs Rel, Sq Rel, RMSE, RMSE log, and δ2 decreases, by 2.6%, 1%, 0.5%, 0.8%, and 0.1% respectively.
[0171] When the SFF module is enabled and the CBAM module is disabled, compared with enabling both modules, the performance of the model in terms of Abs Rel, SqRel, RMSE, RMSE log, and δ1 decreases, by 9%, 10%, 1.8%, 3%, and 0.2% respectively.
[0172] When the SFF module and the CBAM module are disabled, compared with enabling both modules, the performance of the model in terms of Abs Rel, SqRel, RMSE, RMSE log, δ1, δ2, and δ3 decreases, by 12.8%, 20%, 3.9%, 6.3%, 1.3%, 0.5%, and 0.2% respectively.
[0173] Embodiment 2:
[0174] As Figure 9 shown, Embodiment 2 provides a monocular depth estimation system based on a convolutional neural network, including an acquisition module 10 and a depth estimation module 20;
[0175] Among them, the acquisition module 10 is used to acquire the two-dimensional image to be predicted;
[0176] The depth estimation module 20 is configured to input the two-dimensional image into a preset depth estimation model, so that the depth estimation model uses a convolutional neural network model and a vision model to extract local features and global features of the two-dimensional image respectively, and after fusing the local features and global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image through a regression head;
[0177] The depth estimation model is obtained by training an initial depth estimation model with a plurality of historical two-dimensional images and corresponding historical depth images. The initial depth estimation model is jointly constructed by an initial convolutional neural network model, an initial vision model, an initial feature fusion module, and an initial regression head.
[0178] In a possible implementation manner, the depth estimation model uses a convolutional neural network model and a vision model to extract local features and global features of the two-dimensional image respectively, and after fusing the local features and global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image, including:
[0179] Input the two-dimensional image into the convolutional neural network model, so that the convolutional neural network model extracts local features of the two-dimensional image through a plurality of convolutional neural networks with different parameters, and generates a plurality of local feature maps with different dimensions;
[0180] Input the two-dimensional image into the vision model, so that the vision model extracts global features of the two-dimensional image through a plurality of sub-vision models with different parameters, and generates a plurality of global feature maps with different dimensions, where each of the global feature maps and each of the local feature maps have a one-to-one correspondence based on dimensions;
[0181] Perform first feature fusion on each of the global feature maps and the corresponding local feature maps according to the one-to-one correspondence to generate a plurality of first fusion feature maps with different dimensions;
[0182] Input each of the first fusion feature maps into the feature fusion module, so that the feature fusion module performs second feature fusion on each of the first fusion feature maps through a plurality of multi-scale feature fusion modules with different parameters to obtain a second fusion feature map;
[0183] Input the second fusion feature map into the regression head, so that the regression head generates a monocular depth prediction map corresponding to the two-dimensional image.
[0184] Further, each of the convolutional neural networks is arranged and connected in sequence based on the output dimension size. The convolutional neural network model extracts local features of the two-dimensional image through a number of convolutional neural networks with different parameters, generating a number of local feature maps with different dimensions, including:
[0185] Each of the convolutional neural networks sequentially generates the local feature maps according to the arranged and connected order. Among them, the input data of any convolutional neural network is the local feature map output by the previous convolutional neural network, and the input data of the convolutional neural network ranked first is the two-dimensional image.
[0186] Further, each of the sub-visual models is arranged and connected in sequence based on the output dimension size. The visual model extracts global features of the two-dimensional image through a number of sub-visual models with different parameters, generating a number of global feature maps with different dimensions, including:
[0187] Each of the sub-visual models sequentially generates the global feature maps according to the arranged and connected order. Among them, the input data of any sub-visual model is the global feature map output by the previous sub-visual model, and the input data of the sub-visual model ranked first is the two-dimensional image.
[0188] In a possible implementation manner, the first feature fusion is performed on each of the global feature maps and the corresponding local feature maps according to the one-to-one correspondence, generating a number of first fusion feature maps with different dimensions, including:
[0189] Reduce the feature channels of each of the global feature maps and each of the local feature maps to their respective corresponding preset channel numbers, obtaining each corresponding channel-reduced global feature map and each corresponding channel-reduced local feature map;
[0190] Perform the first feature fusion on each of the channel-reduced global feature maps and the corresponding channel-reduced local feature maps, obtaining a number of first fusion feature maps with different dimensions.
[0191] Further, in the process of the first feature fusion, each of the channel-reduced global feature maps and the channel-reduced local feature maps are respectively subjected to the first feature fusion through a number of selective feature fusion modules with different parameters, obtaining a number of first fusion feature maps with different dimensions. Among them, any selective feature fusion module performs the first feature fusion on the channel-reduced global feature map and the channel-reduced local feature map, specifically:
[0192] Perform preliminary feature fusion on the channel-reduced global feature map and the channel-reduced local feature map, obtaining a preliminary fusion feature map;
[0193] Perform consecutive convolution operations and normalization processing operations on the preliminary fusion feature map in sequence to generate a normalized feature map;
[0194] Based on the selective attention mechanism and the channel values in the normalized feature map, perform element-wise multiplication fusion operations on the local attention features and global attention features in the normalized feature map respectively, and add the obtained results to obtain the first fusion feature map.
[0195] In a possible implementation manner, each of the multi-scale feature fusion modules is arranged and connected in sequence based on the output dimension size. The feature fusion module performs second feature fusion on each of the first fusion feature maps through a plurality of multi-scale feature fusion modules with different parameters to obtain a second fusion feature map, including:
[0196] Each of the multi-scale feature fusion modules performs second feature fusion in sequence according to the arranged connection order. Among them, the input data of any multi-scale feature fusion module is the output data of the previous multi-scale feature fusion module and the corresponding first fusion feature map. The input data of the multi-scale feature fusion module ranked first is the corresponding first fusion feature map, and the output data of the multi-scale feature fusion module ranked last is the second fusion feature map.
[0197] Further, during the second feature fusion process, a corresponding CBAM hybrid domain attention module is connected to the output data end of each of the multi-scale feature fusion modules. The CBAM hybrid domain attention module is used to process the output data of the corresponding multi-scale feature fusion module, combine the attention in the spatial domain and channel domain of the output data, and input the processed output data into the next multi-scale feature fusion module.
[0198] In a possible implementation manner, training the initial depth estimation model with a plurality of historical two-dimensional images and corresponding plurality of historical depth images to obtain the depth estimation model includes:
[0199] Take each of the historical depth images as the label of the corresponding historical two-dimensional image, and then construct a training data set based on each of the historical two-dimensional images;
[0200] Input the training data set into the initial depth estimation model so that the initial depth estimation model generates a plurality of corresponding predicted depth images;
[0201] Calculate the loss according to each of the predicted depth images and the corresponding labels in the training data set, and then optimize the parameters by backpropagation inside the initial depth estimation model according to the result of the loss calculation to obtain the depth estimation model.
[0202] The embodiment of the present application provides a monocular depth estimation system based on a convolutional neural network. The global features and local features of a two-dimensional image are extracted through a preset depth estimation model, and a vision model is used to capture the global dependence and long-range correlation of the image, providing global context information for depth estimation. At the same time, a convolutional neural network model is used to extract the local texture features and details of the image, enhancing the model's perception ability for small-region features. Then, the obtained global features and local features are fused, and an accurate monocular depth prediction map is generated through a regression head. The embodiment of the present application effectively combines the ability of the convolutional neural network model to obtain local texture features of the image and the ability of the vision model to capture the long-range correlation between image pixels, fully integrating and utilizing local and global information, so as to more accurately predict and restore the three-dimensional depth information in a single two-dimensional image, overcoming the limitations existing in the process of solving the monocular depth estimation problem by the traditional Transformer model and improving the accuracy of the convolutional neural network for monocular depth estimation.
[0203] The more detailed working principle and step flow of this embodiment can, but are not limited to, referring to the relevant records of Embodiment 1.
[0204] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above description is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. In particular, for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A monocular depth estimation method based on a convolutional neural network, characterized in that, Including: Obtain a two-dimensional image to be predicted; Input the two-dimensional image into a preset depth estimation model, so that the depth estimation model uses a convolutional neural network model and a vision model to extract the local features and global features of the two-dimensional image respectively, and after feature fusion of the local features and global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image through a regression head; Among them, the depth estimation model is obtained by training an initial depth estimation model with a number of historical two-dimensional images and corresponding historical depth images, and the initial depth estimation model is jointly constructed by an initial convolutional neural network model, an initial vision model, an initial feature fusion module, and an initial regression head.
2. The monocular depth estimation method based on a convolutional neural network according to claim 1, characterized in that The depth estimation model uses a convolutional neural network model and a vision model to extract the local features and global features of the two-dimensional image respectively, and after feature fusion of the local features and global features through a feature fusion module, generate a monocular depth prediction map corresponding to the two-dimensional image, including: Input the two-dimensional image into the convolutional neural network model, so that the convolutional neural network model extracts the local features of the two-dimensional image through a number of convolutional neural networks with different parameters, and generates a number of local feature maps with different dimensions; Input the two-dimensional image into the vision model, so that the vision model extracts the global features of the two-dimensional image through a number of sub-vision models with different parameters, and generates a number of global feature maps with different dimensions, where each of the global feature maps and each of the local feature maps has a one-to-one correspondence based on dimensions; Perform first feature fusion on each of the global feature maps with the corresponding local feature maps according to the one-to-one correspondence, and generate a number of first fusion feature maps with different dimensions; Input each of the first fusion feature maps into the feature fusion module, so that the feature fusion module performs second feature fusion on each of the first fusion feature maps through a number of multi-scale feature fusion modules with different parameters, and obtains a second fusion feature map; Input the second fusion feature map into the regression head, so that the regression head generates a monocular depth prediction map corresponding to the two-dimensional image.
3. The monocular depth estimation method based on convolutional neural network according to claim 2, characterized in that, Each of the convolutional neural networks is arranged and connected in sequence based on the output dimension size. The convolutional neural network model extracts the local features of the two-dimensional image through a number of convolutional neural networks with different parameters, and generates a number of local feature maps with different dimensions, including: Each of the convolutional neural networks generates the local feature maps in sequence according to the arranged and connected order. Among them, the input data of any convolutional neural network is the local feature map output by the previous convolutional neural network, and the input data of the convolutional neural network ranked first is the two-dimensional image.
4. The monocular depth estimation method based on a convolutional neural network according to claim 2, characterized in that, Each of the sub-vision models is arranged and connected in sequence based on the output dimension size. The vision model extracts the global features of the two-dimensional image through a number of sub-vision models with different parameters, and generates a number of global feature maps with different dimensions, including: Each of the sub-visual models generates the global feature map in sequence according to the order of arrangement and connection. Among them, the input data of any sub-visual model is the global feature map output by the previous sub-visual model, and the input data of the first sub-visual model in the order is the two-dimensional image.
5. The monocular depth estimation method based on a convolutional neural network according to claim 2, characterized in that Performing first feature fusion on each of the global feature maps and the corresponding local feature maps according to the one-to-one correspondence, to generate a number of first fusion feature maps with different dimensions, including: Reducing the feature channels of each of the global feature maps and each of the local feature maps to their respective corresponding preset channel numbers, to obtain each corresponding channel-reduced global feature map and each corresponding channel-reduced local feature map; Performing first feature fusion on each of the channel-reduced global feature maps and the corresponding channel-reduced local feature maps, to obtain a number of first fusion feature maps with different dimensions.
6. The monocular depth estimation method based on a convolutional neural network according to claim 5, characterized in that, During the first feature fusion process, each of the channel-reduced global feature maps and the channel-reduced local feature maps are respectively subjected to first feature fusion by a number of selective feature fusion modules with different parameters, to obtain a number of first fusion feature maps with different dimensions. Among them, for any selective feature fusion module to perform first feature fusion on the channel-reduced global feature map and the channel-reduced local feature map, specifically: Performing preliminary feature fusion on the channel-reduced global feature map and the channel-reduced local feature map, to obtain a preliminary fusion feature map; Performing consecutive convolution operations and normalization processing operations on the preliminary fusion feature map in sequence, to generate a normalized feature map; Based on the selective attention mechanism and the channel values in the normalized feature map, performing element-wise multiplication fusion operations on the local attention feature and the global attention feature in the normalized feature map respectively, and adding the obtained results, to obtain the first fusion feature map.
7. The monocular depth estimation method based on a convolutional neural network according to claim 2, wherein Each of the multi-scale feature fusion modules is arranged and connected in sequence based on the output dimension size. The feature fusion module performs second feature fusion on each of the first fusion feature maps through a number of multi-scale feature fusion modules with different parameters, to obtain a second fusion feature map, including: Each of the multi-scale feature fusion modules performs second feature fusion in sequence according to the order of arrangement and connection. Among them, the input data of any multi-scale feature fusion module is the output data of the previous multi-scale feature fusion module and the corresponding first fusion feature map, the input data of the first multi-scale feature fusion module in the order is the corresponding first fusion feature map, and the output data of the last multi-scale feature fusion module in the order is the second fusion feature map.
8. The monocular depth estimation method based on a convolutional neural network according to claim 7, characterized in that During the second feature fusion process, a corresponding CBAM hybrid domain attention module is connected to the output data end of each multi-scale feature fusion module. The CBAM hybrid domain attention module is used to process the output data of the corresponding multi-scale feature fusion module, combine the attention in the spatial domain and the channel domain of the output data, and input the processed output data into the next multi-scale feature fusion module.
9. A monocular depth estimation method based on a convolutional neural network according to claim 1, characterized in that, Training the initial depth estimation model with a number of historical two-dimensional images and corresponding historical depth images to obtain the depth estimation model includes: Using each of the historical depth images as the label for the corresponding historical two-dimensional image, and then constructing a training data set based on each of the historical two-dimensional images; Inputting the training data set into the initial depth estimation model so that the initial depth estimation model generates a number of corresponding predicted depth images; Calculating the loss based on each of the predicted depth images and the corresponding labels in the training data set, and then optimizing the parameters by backpropagation inside the initial depth estimation model according to the result of the loss calculation to obtain the depth estimation model.
10. A monocular depth estimation system based on a convolutional neural network, characterized in that, It includes an acquisition module and a depth estimation module; Among them, the acquisition module is used to acquire the two-dimensional image to be predicted; The depth estimation module is used to input the two-dimensional image into a preset depth estimation model, so that the depth estimation model uses a convolutional neural network model and a vision model to extract the local features and global features of the two-dimensional image respectively, and through a feature fusion module, fuse the local features and global features, and then generate a monocular depth prediction map corresponding to the two-dimensional image through a regression head; The depth estimation model is obtained by training the initial depth estimation model with a number of historical two-dimensional images and corresponding historical depth images. The initial depth estimation model is jointly constructed by an initial convolutional neural network model, an initial vision model, an initial feature fusion module, and an initial regression head.
Citation Information
Cited By
Depth estimation method based on double-branch deep network and multi-attention fusion
CN121280499A