A lightweight depth estimation method and system
By combining the lightweight depth estimation method of two-dimensional depth separable convolution and three-dimensional point-by-point convolution, the problems of difficult deployment, low memory utilization and waste of hardware resources in practical application scenarios are solved, and performance improvement is achieved.
Patent Information
- Application Number
- CN202510261430.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-06
AI Technical Summary
The existing binocular depth estimation network is difficult to deploy in practical application scenarios, with low memory utilization, wasted hardware resources, and performance losses.
A lightweight depth estimation method is adopted, combining two-dimensional depth separation convolution and three-dimensional point-by-point convolution, reducing the number of layers extracted by feature and the number of parameters of the depth estimation model.
It effectively solves the problems of difficult deployment, low memory utilization and waste of hardware resources in actual application scenarios, and improves performance.
Smart Images

Figure CN119762567B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of binocular system depth estimation, and specifically relates to a lightweight depth estimation method and system. Background Art
[0002] With the development of artificial intelligence, application fields such as autonomous driving, drones, and robots have received increasing attention, and environmental perception is the core technology among them. Depth estimation obtains depth information by reconstructing the three-dimensional information of the environment, thereby realizing environmental perception, which is an important research direction in the field of computer vision. Traditional methods for obtaining depth information come from lidar, monocular camera imaging, etc., and have problems such as insufficient richness of collected image features and only being applicable to short-distance scenarios. In recent years, convolutional neural networks, as an important part of deep learning, have made major breakthroughs in image fields such as object detection and object recognition, and can extract rich feature information from images, overcoming the deficiencies of traditional algorithms.
[0003] In the past few years, researchers have proposed some lightweight models by improving the convolutional layers of neural networks. In 2016, a team proposed the MobileNet model (Mobile Network Model), which selected depthwise separable convolutions to replace traditional convolutions, decomposed the convolution calculation into two steps of depth convolution and pointwise convolution, and introduced a width multiplier as a hyperparameter to select a model of appropriate size according to the actual application scenario. Compared with VGG16 (Visual Geometry Group Network with 16 layers), this lightweight network with the new convolutional method reduces the multiply-accumulate computational amount by 27 times and the number of parameters by 33 times on the premise of a small loss of accuracy, thus effectively compressing the model size. However, the network model structure is simple and does not reuse image features for feature fusion, and then more efficient network models were proposed. Chang et al. proposed the PSMNet (Pyramid Stereo Matching Network) network. In the feature extraction part, a deep residual neural network (ResNet) was used as the backbone network, and then the SPP (Spatial Pyramid Pooling) spatial pyramid structure was used to obtain feature information under different scale receptive fields, and the global information and local information were combined to form a matching cost volume.
[0004] However, as the requirements for accuracy in actual application scenarios are getting higher and higher, the large-scale parameters and high-complexity calculations of neural networks pose challenges to the storage, power consumption, and computing capabilities of devices. In resource-constrained mobile applications such as smart wearable devices, drones, and intelligent robots, neural networks will quickly exhaust their storage, memory, battery, and computing units. Therefore, reducing the parameter scale and computing complexity of neural networks is of great significance for deploying neural networks in mobile applications. Summary of the Invention
[0005] Object of the Invention: This application develops a lightweight depth estimation method and system, aiming to solve the technical problems of difficult deployment, low memory utilization, and waste of hardware resources of binocular depth estimation networks in actual application scenarios.
[0006] Technical Solution: In a first aspect, an embodiment of this application provides a lightweight depth estimation method, including:
[0007] Obtain a depth estimation model, input a binocular depth image into the depth estimation model, and obtain a depth estimation result. Among them, the step of obtaining the depth estimation model includes:
[0008] Determine input parameters, where the input parameters include a binocular depth image;
[0009] Extract features from the input parameters based on two-dimensional depthwise separable convolution to obtain feature parameters;
[0010] Obtain a cost volume through the feature parameters;
[0011] Perform stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information;
[0012] Determine output parameters, where the output parameters include multi-depth feature information, and the multi-depth feature information is used to represent the depth estimation result.
[0013] In some embodiments, the step of extracting features from the input parameters based on two-dimensional depthwise separable convolution to obtain feature parameters includes:
[0014] Perform multi-layer two-dimensional depthwise separable convolution operations on the binocular depth image, and perform non-linear processing using a rectified linear unit activation function after each layer of two-dimensional depthwise separable convolution operation to obtain a first feature map;
[0015] Process the first feature map through a multi-layer residual network to obtain a second feature map;
[0016] Process the second feature map through a fast spatial pyramid pooling layer to obtain a multi-channel feature map that fuses features of multiple scales, which is the third feature map;
[0017] The information of the third feature map is fused through a per-pixel convolutional layer to obtain feature parameters.
[0018] In some embodiments, the binocular depth image includes a left view and a right view, and the step of obtaining a cost volume from the feature parameters includes:
[0019] Obtain the difference data of the feature parameters of the left view and the right view corresponding to each pixel, perform matching by moving a window, respectively crop and match the left view and the right view to another view, and obtain the cost volume.
[0020] In some embodiments, the step of performing stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information includes:
[0021] Stack multiple hourglass networks, in which:
[0022] Perform a lightweight three-dimensional pointwise convolution operation on the cost volume to reduce the resolution of the cost volume and obtain a preliminary prediction result;
[0023] Perform a three-dimensional transposed convolution operation on the preliminary prediction result for upsampling to obtain the depth feature information.
[0024] In some embodiments, the depth estimation method further includes:
[0025] Train the depth estimation model based on the depth estimation result, and the model training includes a classification task and a regression task;
[0026] Obtain a depth map label, and obtain a classification label and a regression label based on the depth map label;
[0027] Obtain a classification loss based on the classification label and the classification task;
[0028] Obtain a regression loss based on the regression label and the regression task;
[0029] Determine a total loss function based on the classification loss and the regression loss;
[0030] Optimize the depth estimation model based on the total loss function.
[0031] In some embodiments, the step of obtaining a classification label based on the depth map label includes:
[0032] Perform clustering processing on the depth map label to obtain a clustering center;
[0033] Obtain a class mapping based on the clustering center and disparity classification;
[0034] Obtain the classification label based on the category mapping.
[0035] In some embodiments, the steps of the classification task include:
[0036] Process the depth estimation result through multi-layer pyramid three-dimensional lightweight convolution and multi-layer two-dimensional convolution to obtain a disparity estimation map;
[0037] The steps of the regression task include:
[0038] Perform channel fitting on the disparity estimation map through multi-layer pyramid three-dimensional lightweight convolution and multi-layer two-dimensional convolution, and obtain a continuous disparity map based on disparity regression.
[0039] In some embodiments, the stacked hourglass processing is performed through a multi-layer series-connected stacked hourglass network, and the step of determining the total loss function based on the classification loss and the regression loss includes:
[0040] In the multi-layer hourglass network,
[0041] For the output of each hourglass network, use the smooth absolute error loss function to determine the regression loss and assign a weight coefficient; among them, the larger the depth of the hourglass network, the larger the weight coefficient assigned to its output;
[0042] For the output of each hourglass network, use the cross-entropy loss function to determine the classification loss and assign a weight coefficient; among them, the larger the depth of the hourglass network, the larger the weight coefficient assigned to its output;
[0043] For the weight coefficient assignment between the classification loss and the regression loss, use dynamic coefficient assignment. When the number of optimization iterations is less than the iteration threshold, only use the classification loss as the total loss function; when the number of iterations is greater than or equal to the iteration threshold, then determine the total loss function based on the function after assigning weight coefficients to the classification loss and the regression loss.
[0044] In some embodiments, the lightweight depth estimation method further includes: performing linear quantization, weight pruning optimization, and activation pruning optimization on the depth estimation model to optimize the model storage requirements, and the steps of the linear quantization include:
[0045] Linearly map single-precision floating-point data to data of a specified bit, and its representation formula includes:
[0046]
[0047] s=(x max -x min ) / 2 n
[0048] z=-round((xmax +x min ) / (2×s));
[0049] Among them, x is the floating-point data to be quantized; s is the quantization scale; z is the quantization zero; n is the bit width of the data with specified bits; round is the rounding operation; clip is the truncation operation; x max is the maximum clipping threshold; x min is the minimum clipping threshold; q is the integer after quantizing the initial target parameter;
[0050] The steps of the weight clipping optimization include:
[0051] Determine the target convolution kernel in the depth estimation model;
[0052] Based on the particle swarm algorithm, initialize particles in the depth estimation model and determine the position range parameters of the particles;
[0053] Determine the positions and velocities of the particles based on the target convolution kernel and the position range parameters;
[0054] Determine the quantization deviation of the position and determine the position corresponding to the smallest quantization deviation as the optimal position;
[0055] In response to the quantization deviation being within the deviation threshold, determine the optimal position corresponding to the quantization deviation as the optimal weight clipping threshold;
[0056] In response to the quantization deviation being outside the deviation threshold, update the velocity based on the optimal position;
[0057] The steps of the activation clipping optimization include:
[0058] Obtain the activation distribution histogram of each channel in the same interval, which is the channel attention;
[0059] Compress the original feature map matrix based on principal component analysis to obtain the single-layer spatial attention;
[0060] Sum and normalize the single-layer spatial attention to obtain the global spatial attention;
[0061] Obtain the activation clipping attention based on the single-layer spatial attention, the global spatial attention and the channel attention;
[0062] The steps of compressing the original activation based on principal component analysis to obtain the single-layer spatial attention include:
[0063] Obtain the de-centered matrix of the original feature map matrix;
[0064] Perform singular value decomposition on the decentralized matrix, and compress the decomposed data by column vectors to obtain the projection of the original feature map matrix on the first column vector of the velocity feature;
[0065] Apply a linear rectification unit and a normalization operation to the projection to obtain the activated spatial attention, which is the single-layer spatial attention.
[0066] In a second aspect, an embodiment of the present application provides a lightweight depth estimation system, including:
[0067] A model construction module, which is used to obtain a depth estimation model, input a binocular depth image into the depth estimation model, and obtain a depth estimation result. Among them, the model construction module includes:
[0068] An input sub-module, which is used to determine input parameters, and the input parameters include a binocular depth image;
[0069] A feature extraction sub-module, which is used to extract features from the input parameters based on two-dimensional depthwise separable convolution to obtain feature parameters;
[0070] A cost volume acquisition sub-module, which is used to obtain a cost volume through the feature parameters;
[0071] A depth feature acquisition module, which is used to perform stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information;
[0072] An output sub-module, which is used to determine output parameters, and the output parameters include multi-depth feature information, and the multi-depth feature information is used to represent the depth estimation result.
[0073] Beneficial effects: Compared with the prior art, a lightweight depth estimation method provided by an embodiment of the present application obtains a depth estimation model, inputs a binocular depth image into the depth estimation model, and obtains a depth estimation result. When obtaining the depth estimation model, it determines the binocular depth image as the input parameter, extracts features from the input parameter based on two-dimensional depthwise separable convolution, obtains a cost volume based on the extracted feature parameters, performs stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution, obtains multi-depth feature information, and determines the multi-depth feature information as the depth estimation result, that is, the output parameter. The present application integrates lightweight feature extraction methods such as two-dimensional depthwise separable convolution and three-dimensional pointwise convolution, which can reduce the number of feature extraction layers, reduce the number of parameters of the depth estimation model, and solve the technical problems of difficult deployment, low memory utilization, waste of hardware resources, and performance loss of the binocular depth estimation network in actual application scenarios. Description of the Drawings
[0074] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0075] Figure 1 It is a flowchart of the steps for obtaining a depth estimation model in the lightweight depth estimation method provided by the embodiments of the present application;
[0076] Figure 2 It is a flowchart of the steps for obtaining feature parameters in the lightweight depth estimation method provided by the embodiments of the present application;
[0077] Figure 3 It is a flowchart of the steps for optimizing the depth estimation model in the lightweight depth estimation method provided by the embodiments of the present application;
[0078] Figure 4 It is a flowchart of the steps for obtaining classification labels in the lightweight depth estimation method provided by the embodiments of the present application;
[0079] Figure 5 It is a module connection diagram of the lightweight depth estimation system provided by the embodiments of the present application;
[0080] Figure 6 It is a module flowchart of the depth estimation method;
[0081] Figure 7 It is a module flowchart for obtaining cluster centers using the Kmeans clustering algorithm;
[0082] Figure 8 It is a module flowchart for extracting deep features from the input left and right views;
[0083] Figure 9 It is a schematic diagram of cost volume construction;
[0084] Figure 10 It is a module flowchart of the hourglass network and the loss function calculation;
[0085] Figure 11 It is a schematic diagram of constructing channel attention;
[0086] Reference numerals: 10, model construction module; 11, input sub-module; 12, feature extraction sub-module; 13, cost volume acquisition sub-module; 14, deep feature acquisition module; 15, output sub-module. Detailed implementation manners
[0087] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0088] With the development of artificial intelligence, application fields such as autonomous driving, drones, and robots have received increasing attention, and environmental perception is the core technology among them. Depth estimation obtains depth information by reconstructing three-dimensional information of the environment, thereby realizing environmental perception, which is an important research direction in the field of computer vision. Traditional methods for obtaining depth information come from lidar, monocular camera imaging, etc., and have problems such as insufficient richness of collected image features and only being applicable to short-distance scenarios. In recent years, convolutional neural networks, as an important part of deep learning, have made significant breakthroughs in image fields such as object detection and object recognition, and can extract rich feature information from images, overcoming the deficiencies of traditional algorithms.
[0089] In the past few years, researchers have proposed some lightweight models by improving the convolutional layers of neural networks. In 2016, a team proposed the MobileNet model (mobile network model), which selected depthwise separable convolutions to replace traditional convolutions, decomposed the convolution calculation into two steps: depth convolution and pointwise convolution, and introduced a width multiplier as a hyperparameter to select a model of appropriate size according to the actual application scenario. Compared with VGG16 (Visual Geometry Group Network with 16 layers), this new type of convolutional lightweight network reduces the multiply-accumulate computational amount by 27 times and the number of parameters by 33 times on the premise of a small loss of accuracy, thus effectively compressing the model size. However, the network model structure is simple and does not reuse image features for feature fusion, and then more efficient network models were proposed. Chang et al. proposed the PSMNet (Pyramid Stereo Matching Network) network. In the feature extraction part, a deep residual neural network (ResNet) was used as the backbone network, and then the SPP (Spatial Pyramid Pooling) spatial pyramid structure was used to obtain feature information under different scale receptive fields, and the global information and local information were combined to form a matching cost volume.
[0090] However, as the requirements for accuracy in actual application scenarios become increasingly high, the large-scale parameters and high-complexity calculations of neural networks pose challenges to the storage, power consumption, and computing capabilities of devices. In resource-constrained mobile applications such as smart wearable devices, drones, and intelligent robots, neural networks will quickly exhaust their storage, memory, batteries, and computing units. Therefore, reducing the parameter scale and computing complexity of neural networks is of great significance for deploying neural networks in mobile applications.
[0091] Researchers have proposed a series of methods for neural network compression and acceleration, such as distillation, pruning, and quantization. Among these techniques, model quantization exhibits significant advantages. Current quantization methods include quantization-based training and post-training quantization. Quantization-based training generally inserts quantization calculation nodes into the original network or directly designs low-bit computing units, and then uses backpropagation to learn the network weights and related quantization parameters. It can significantly improve the quantization performance, but the time cost is huge. Post-training quantization has been more widely used in mobile model deployment due to its convenience and efficiency. However, it causes greater performance loss to the network, especially for lightweight networks. It has been proven that pruning optimization can effectively improve the performance of post-training quantization, such as TensorRt, EasyQuant, MSE, ACIQ, etc., whose purpose is to select the optimal pruning threshold to minimize the quantization deviation. However, these methods ignore the fact that the network's attention to each weight or activation is different. Therefore, a new pruning optimization method that can find the optimal search step and search interval is needed.
[0092] In view of this, an embodiment of the present application provides a lightweight depth estimation method. Obtain a depth estimation model, input a binocular depth image into the depth estimation model to obtain a depth estimation result. When obtaining the depth estimation model, determine the binocular depth image as the input parameter, perform feature extraction on the input parameter based on two-dimensional depthwise separable convolution, and obtain a cost volume based on the extracted feature parameters. Perform stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information, and determine the multi-depth feature information as the depth estimation result, that is, the output parameter. The present application integrates lightweight feature extraction methods such as two-dimensional depthwise separable convolution and three-dimensional pointwise convolution, which can reduce the number of layers of feature extraction, reduce the number of parameters of the depth estimation model, and solve the technical problems of difficult deployment, low memory utilization, and waste of hardware resources of the binocular depth estimation network in actual application scenarios.
[0093] In some embodiments, refer to Figure 1 , Figure 1The flowchart of the steps for obtaining a depth estimation model in the lightweight depth estimation method provided by an embodiment of the present application. After obtaining the depth estimation model, a binocular depth image is input into the depth estimation model to obtain a depth estimation result. Specifically, the method for obtaining the depth estimation model is implemented through steps 100 to 500:
[0094] Step 100: Determine input parameters, where the input parameters include a binocular depth image.
[0095] Step 200: Perform feature extraction on the input parameters based on two-dimensional depthwise separable convolution to obtain feature parameters.
[0096] In some embodiments, please refer to Figure 2 , Figure 2 The flowchart of the steps for obtaining feature parameters in the lightweight depth estimation method provided by an embodiment of the present application. Specifically, the method for performing feature extraction on the input parameters based on two-dimensional depthwise separable convolution to obtain feature parameters is implemented through steps 210 to 240:
[0097] Step 210: Perform multi-layer two-dimensional depthwise separable convolution operations on the binocular depth image, and perform non-linear processing using a rectified linear unit activation function after each two-dimensional depthwise separable convolution operation to obtain a first feature map.
[0098] Step 220: Process the first feature map through a multi-layer residual network to obtain a second feature map.
[0099] Step 230: Process the second feature map through a fast spatial pyramid pooling layer to obtain a multi-channel feature map that fuses features of multiple scales, which is the third feature map.
[0100] Step 240: Fuse the information of the third feature map through a pixel-wise convolution layer to obtain feature parameters.
[0101] Step 300: Obtain a cost volume through the feature parameters.
[0102] In some embodiments, the binocular depth image includes a left view and a right view. When constructing the cost volume, the difference data of the feature parameters of the left view and the right view corresponding to each pixel is obtained, and matching is performed by means of a moving window. The left view and the right view are respectively cropped and matched with the other view to obtain the cost volume.
[0103] Step 400: Perform stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information.
[0104] In some embodiments, the stacked hourglass process is implemented by stacking multiple hourglass networks. In each hourglass network, a lightweight 3D pointwise convolution operation is performed on the cost volume to reduce the resolution of the cost volume and obtain a preliminary prediction result. A 3D transposed convolution operation is performed on the preliminary prediction result for upsampling to obtain depth feature information.
[0105] Step 500: Determine output parameters, where the output parameters include multi-depth feature information, and the multi-depth feature information is used to represent the depth estimation result.
[0106] In some embodiments, after the depth estimation model is constructed, it is necessary to train and optimize the depth estimation model. Please refer to Figure 3 , Figure 3 which is the flowchart of the steps for optimizing the depth estimation model in the lightweight depth estimation method provided by the embodiments of the present application. The method for optimizing the depth estimation model is specifically implemented through steps 610 to 660:
[0107] Step 610: Perform model training on the depth estimation model based on the depth estimation result. The model training includes a classification task and a regression task.
[0108] Step 620: Obtain a depth map label, and based on the depth map label, obtain a classification label and a regression label.
[0109] In some embodiments, please refer to Figure 4 , Figure 4 which is the flowchart of the steps for obtaining the classification label in the lightweight depth estimation method provided by the embodiments of the present application. The method for obtaining the classification label is specifically implemented through steps 621 to 623:
[0110] Step 621: Perform clustering processing on the depth map label to obtain a clustering center.
[0111] Step 622: Based on the clustering center and disparity classification, obtain a class mapping.
[0112] Step 623: Based on the class mapping, obtain the classification label.
[0113] Step 630: Based on the classification label and the classification task, obtain a classification loss.
[0114] In some embodiments, the classification task processes the depth estimation result through multi-layer pyramid 3D lightweight convolution and multi-layer 2D convolution to obtain a disparity estimation map.
[0115] Step 640: Based on the regression label and the regression task, obtain a regression loss.
[0116] In some embodiments, for the regression task, the disparity estimation map is subjected to channel fitting through multi-layer pyramid 3D lightweight convolution and multi-layer 2D convolution, and a continuous disparity map is obtained based on disparity regression.
[0117] Step 650: Determine the total loss function based on the classification loss and the regression loss.
[0118] In some embodiments, stacked hourglass processing is performed through a multi-layer series-connected stacked hourglass network. For the output of each hourglass network, the smooth absolute error loss function is used to determine the regression loss, and a weight coefficient is assigned; among them, the larger the depth of the hourglass network, the larger the weight coefficient assigned to the output of the hourglass network;
[0119] For the output of each hourglass network, the cross-entropy loss function is used to determine the classification loss, and a weight coefficient is assigned; among them, the larger the depth of the hourglass network, the larger the weight coefficient assigned to the output of the hourglass network;
[0120] For the weight coefficient assignment between the classification loss and the regression loss, dynamic coefficient assignment is used. When the number of optimization iterations is less than the iteration threshold, only the classification loss is used as the total loss function; when the number of iterations is greater than or equal to the iteration threshold, the total loss function is determined based on the function after assigning weight coefficients to the classification loss and the regression loss.
[0121] Step 660: Optimize the depth estimation model based on the total loss function.
[0122] In some embodiments, after the depth estimation model is constructed or after the optimization of the depth estimation model is completed, the present application further optimizes and clips the depth estimation model through linear quantization, weight pruning optimization, and / or activation pruning optimization to further reduce the memory consumption of the mobile device.
[0123] In some embodiments, the steps of linear quantization include:
[0124] Linearly map FP32 (Floating Point 32, 32-bit floating point number, also known as single-precision floating point number) data to data of a specified number of bits, and its representation formula includes:
[0125]
[0126] s = (x max - x min ) / 2 n
[0127] z = -round((x max + x min ) / (2 × s));
[0128] Wherein, x is the floating-point data to be quantized; s is the quantization scale; z is the quantization zero; n is the bit width of the data with specified bits; round is the rounding operation; clip is the truncation operation; x max is the maximum clipping threshold; x min is the minimum clipping threshold; q is the integer after quantizing the initial target parameter;
[0129] In some embodiments, the steps of weight clipping optimization include:
[0130] Determine the target convolution kernel in the depth estimation model.
[0131] Based on the particle swarm algorithm, initialize particles in the depth estimation model and determine the position range parameters of the particles.
[0132] Determine the positions and velocities of the particles based on the target convolution kernel and the position range parameters.
[0133] Determine the quantization deviation of the position, and determine the position corresponding to the minimum quantization deviation as the optimal position.
[0134] In response to the quantization deviation being within the deviation threshold, determine the optimal position corresponding to the quantization deviation as the optimal weight clipping threshold.
[0135] In response to the quantization deviation being outside the deviation threshold, update the velocity based on the optimal position.
[0136] In some embodiments, the steps of activation clipping optimization include:
[0137] Obtain the activation distribution histogram of each channel in the same interval, which is the channel attention.
[0138] Compress the original feature map matrix based on principal component analysis to obtain the single-layer spatial attention.
[0139] Sum and normalize multiple single-layer spatial attentions to obtain the global spatial attention.
[0140] Obtain the activation clipping attention based on the single-layer spatial attention, the global spatial attention, and the channel attention.
[0141] In some embodiments, the steps of compressing the original feature map matrix based on principal component analysis to obtain the single-layer spatial attention include:
[0142] Obtain the decentralized matrix of the original feature map matrix.
[0143] Perform singular value decomposition on the decentralized matrix, and compress the decomposed data by column vectors to obtain the projection of the original feature map matrix on the first column vector of the velocity feature.
[0144] Apply a linear correction unit and a normalization operation to the projection to obtain activated spatial attention, which is single-layer spatial attention.
[0145] Understandably, the embodiments of the present application provide a lightweight depth estimation method. Obtain a depth estimation model, input a binocular depth image into the depth estimation model to obtain a depth estimation result. When obtaining the depth estimation model, determine the binocular depth image as the input parameter, perform feature extraction on the input parameter based on two-dimensional depthwise separable convolution, and obtain a cost volume based on the extracted feature parameters. Perform a stacked hourglass process on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information, and determine the multi-depth feature information as the depth estimation result, that is, the output parameter. The present application integrates lightweight feature extraction methods such as two-dimensional depthwise separable convolution and three-dimensional pointwise convolution, which can reduce the number of layers of feature extraction, reduce the number of parameters of the depth estimation model, and solve the technical problems of difficult deployment, low memory utilization, and waste of hardware resources in the actual application scenario of the binocular depth estimation network.
[0146] Exemplarily, the embodiments of the present application also provide a specific implementation scheme of a lightweight depth estimation method, including:
[0147] (1) Please refer to Figure 6 , Figure 6 which is the module flowchart of the depth estimation method. When the steps of depth estimation are specifically implemented, it includes:
[0148] 1) Data clustering:
[0149] The purpose of clustering the depth map labels is to divide different depth values into 20 preset classes and find a "center" value for each class.
[0150] First, initialize the class centers. Divide the range of depth values [0, 192] into 20 classes, and use a uniform distribution method to initialize the center value of each class: centers = [9, 19, 29, 39,..., 189].
[0151] Then, match each pixel with its corresponding class center. Calculate the distance between the depth of each pixel and each class center. Then, assign each pixel to the class center with the smallest distance.
[0152] Please refer to Figure 7 , Figure 7It is the module flowchart for obtaining the clustering centers using the Kmeans clustering algorithm. When processing the dataset with the Kmeans clustering algorithm, first read the binocular depth image, use the flatten() method to flatten the two-dimensional depth image into a one-dimensional array, create a Kmeans clustering model, divide the depth values into 20 clusters, return the clustering centers of each category, and calculate the distance between each disparity value and the clustering center:
[0153]
[0154] Among them, d is the distance between the disparity value and the clustering center; x is the disparity value; μ is the clustering center; n is the number of categories; j is the serial number of the category. For each disparity value, assign it to the clustering center with the closest distance and classify the pixel into category j.
[0155] Finally, calculate the clustering centers. For each category, calculate the average depth value of all pixels in that category, which can ensure that the central value of each category represents the "typical" depth value of that category:
[0156]
[0157] Among them, N j is the number of all pixels in the j-th category; C j is the set of all pixels in the j-th category; x i is the depth value of the i-th pixel belonging to this category.
[0158] 2) Extract deep features from the input left and right views:
[0159] Please refer to Figure 8 , Figure 8 It is the module flowchart for extracting deep features from the input left and right views. First, use three-layer two-dimensional depthwise separable convolution (DSConv) to extract the features of the image, reducing the number of parameters while maintaining the feature extraction effect. After each layer of convolution, use the ReLU activation function to increase the nonlinearity of the network.
[0160] Then, more advanced features are gradually extracted through residual blocks. In layer1, the input features are processed by 3 BasicBlocks to capture finer-grained features. Layer2 increases the depth of the feature map while reducing the spatial size through downsampling. The main role of this stage is to gradually improve the network's representation ability and keep the information flowing smoothly through skip connections. In layer3, the network continues to extract more advanced features from the features output by layer2 while keeping the spatial size unchanged. Layer4 introduces dilated convolutions (dilation = 2) to keep the feature map resolution unchanged while increasing the receptive field, enabling the network to capture a larger range of context information. Through residual connections, the network can effectively utilize lower-level features while improving the flow of information and gradient transmission.
[0161] Subsequently, the Fast Spatial Pyramid Pooling layer (SPPF) is used to process different scales simultaneously, effectively enhancing the network's multi-scale perception ability. Multiple branches are created for different pooling scales, and each branch performs pooling and convolution operations on the input feature map at different scales. In this application, four pooling windows with scales of 8*8, 16*16, 32*32, and 64*64 are set to evenly capture local and global information. After pooling, the feature map is upsampled in the spatial dimension to the original image size and concatenated (concat) with the original feature map in the channel dimension to obtain a multi-channel feature map that fuses features of different scales.
[0162] Finally, a 1*1 convolutional layer (per-pixel convolutional layer) is used to fuse information of different scales to extract the features of the left and right views.
[0163] 3) Cost volume construction:
[0164] Please refer to Figure 9 , Figure 9 , which is the schematic diagram of cost volume construction (where H is the height of the image, i.e., the number of pixels in the vertical direction of the image; W is the width of the image, i.e., the number of pixels in the horizontal direction of the image; C is the number of channels of the image). First, a cost volume tensor of all zeros is initialized to store the cost volume of the left and right views for subsequent matching calculations. Then, the differences between the features of the left and right views are calculated pixel by pixel, and matching is performed by moving the window. The left and right views are respectively cropped to match the other view to obtain the cost volume. The matching cost between the left and right views is quantified by their features, and the cost volume represents the possibility of different disparities.
[0165] 4) Stacked hourglass to process the cost volume:
[0166] Process the image data using three - layer convolutional and transposed convolutional operations, and combine the top - down and bottom - up convolutional operations to capture feature information at different depths.
[0167] In the first hourglass network, the input cost volume undergoes two steps of lightweight P3D convolutional (3D pointwise convolutional) operations, which reduce the image resolution to 1 / 16H×1 / 16W×1 / 16D, and at the same time output the first preliminary prediction result. When extracting features, it gradually captures global information while reducing the spatial size. Then, use 3D transposed convolutional (transpose convolutional) operations for upsampling to restore the spatial size of the image. This stage helps the network refine the depth information in the image.
[0168] Stack three hourglass networks to gradually optimize the feature representation. Each unit outputs a preliminary prediction result for calculating the loss. This recursive processing makes the output of each stage gradually more accurate. By stacking the outputs of multiple hourglass networks, the network can obtain high - quality depth estimation results.
[0169] 5) Classification and regression mixed task:
[0170] For the classification task, the input is the three results output by the three - layer hourglass network. Use two - layer P3D convolutional (3D pointwise convolutional) and two - layer 2D depth - separable convolutional (DSConv) operations to output a 20 - class disparity estimation map, corresponding to 20 channels. For the regression task, use two - layer P3D convolutional (3D pointwise convolutional) and two - layer 2D depth - separable convolutional (DSConv) operations to fit the disparity in one channel, and use disparity regression to estimate the continuous disparity map. The probability of each disparity d is obtained from the predicted cost c via the softmax operation σ(·). d The predicted disparity is calculated as the sum of each disparity d weighted by its probability, and its characterization formula includes:
[0171]
[0172] where, is the predicted disparity value; c d represents the matching cost when the disparity is d, which is used to characterize the predicted cost;
[0173] 6) Loss function calculation:
[0174] The loss function of this model is divided into two parts, namely the classification loss and the regression loss.
[0175] Please refer to Figure 10 , Figure 10It is a flowchart of the hourglass network and the loss function calculation module. For the regression loss, the present invention uses the smooth absolute error loss (smooth L1 loss) to characterize the regression loss. The smooth L1 loss is widely used in the bounding box regression of object detection due to its robustness and low sensitivity to outliers. Its calculation formula is as follows:
[0176]
[0177] Where:
[0178]
[0179] Among them, L is the smooth L1 loss, which is used to characterize the regression loss; d is the distance between the disparity value and the clustering center, which is used to characterize the actual disparity; i is the sorting subscript; N is the number of labeled pixels; is the predicted disparity; x is the difference between the predicted value and the true value.
[0180] For the three outputs of the three hourglass networks in the regression task, the smooth absolute error loss function is used to determine the regression loss, and weight coefficients of 0.5, 0.7, and 1 are assigned to the three outputs. The deeper the depth, the finer the information obtained, and the larger the assigned weight coefficient.
[0181] For the classification task loss, the cross-entropy loss function is used to calculate the error between the predicted class and the true class. First, the dataset is constructed based on the regression task and belongs to a continuous depth map. Therefore, it needs to be processed into the label of the classification task. For the label disp_L (depth map) in the regression task, first, the disparity value is mapped to the nearest class center according to the predefined class center values (center array). Then, according to the class center mapping, the disparity value (disp_L) of each pixel is converted into the corresponding class label. In this way, the regression label disp_L becomes the label of the classification task. After obtaining the label of the classification task, the loss function of the classification task is calculated. The present invention uses the cross-entropy loss function to calculate the classification loss. For a multi-classification task with C classes, the cross-entropy loss is defined as follows:
[0182]
[0183] Where, Loss is the cross-entropy loss; C is the number of classes; y i is the one-hot encoding of the true class i. If the true class is the k-th class, then y k = 1, and for the remaining classes y i = 0; is the predicted probability of the model for class i; the logarithmic operation It is used to measure the matching degree between the predicted probability and the true probability. Similarly, for the three hourglass outputs of the classification task, weight coefficients of 0.5, 0.7, and 1 are also assigned.
[0184] For the assignment of weight coefficients between the classification task and the regression task, dynamic coefficient assignment is used. When the number of optimization iterations is less than 5, only the classification loss is used. When the number of optimization iterations is greater than or equal to 5, weight coefficients of 1 and 0.7 are assigned to the classification loss and the regression loss respectively.
[0185] (2) Pruning optimization method:
[0186] The quantization method after ordinary training performs well on the mobile deployment side but will cause performance loss. Pruning optimization is an advanced technology to improve quantization after training. However, the existing methods balance the quantization deviation, which limits the further improvement of quantization accuracy. Therefore, this application uses an algorithm that can give different attentions to each parameter and preferentially reduce the quantization deviation of important parameters. The specific steps are as follows:
[0187] 1) First, perform linear quantization to linearly map FP32 (Floating Point 32, 32-bit floating-point number, also known as single-precision floating-point number) data to low-bit data, which is defined as:
[0188]
[0189] s = (x max - x min ) / 2 n
[0190] z = -round((x max + x min ) / (2 × s)) (0.7)
[0191] where x is the floating-point data to be quantized; s is the quantization scale; z is the quantization zero; n is the bit width of the data with the specified number of bits; round is the rounding operation; clip is the truncation operation; x max is the maximum truncation threshold; x min is the minimum truncation threshold; q is the integer after quantizing the initial target parameter. Through the clip truncation operation, the quantized data is restricted to [-2 n-1 , 2 n-1 - 1], so that x min is mapped to -2 n-1 , and at the same time x max is mapped to 2 n-1 - 1.
[0192] 2) Weight pruning optimization:
[0193] The particle swarm optimization algorithm is added to the weight clipping to find the optimal weight clipping threshold. For the target convolution kernel w, the positions and velocities of several particles are initialized. The position P k of the k-th particle is (x1, x2), where x1 and x2 represent the minimum clipping threshold and the maximum clipping threshold respectively. Its moving velocity is V k =(v1, v2), where v1 and v2 represent the search direction and the step size respectively. The initial position and the initial velocity of the particle are defined by Equation (1.8), where Δ is a parameter that controls the range of the initial position of the particle.
[0194] x1 ∼ U(min(w) - Δ, min(w) + Δ)
[0195] x2 ∼ U(max(w) - Δ, max(w) + Δ)
[0196]
[0197] L = max(w) - min(w)(0.8)
[0198] where U is the uniform distribution; L is the difference between the maximum value and the minimum value of w; specifically, the quantization deviation is calculated at the initial position (p 1 , p 2 ,..., p k ) of each particle, and the minimum quantization deviation is recorded, which is defined as the optimal position G best . In addition, the optimal position of each particle is recorded The velocity of the particle changes based on G best and .
[0199]
[0200] where c1 and c2 are learning factors. When c1 is larger, the particle tends to learn its own experience and approach its best position. When c2 is larger, the particle tends to learn the experience of the entire particle swarm and move towards the global optimal position;
[0201] γ is the inertia factor of the original velocity, indicating that the particle will retain the initial motion trend to a certain extent. γ is defined as:
[0202]
[0203] P k = P k + V k (0.11)
[0204] Among them, T and t respectively represent the total number of iterations and the current iteration period of the particle swarm. After the particle swarm moves to a new position with a new velocity that conforms to (1.10), the quantization deviation of each particle is recalculated to obtain a new and G best , and the velocity is re-determined. Finally, most particles will converge to the global optimal position. The optimal position of the particle swarm is the optimal weight pruning threshold G best →[w min ,w max .
[0205] 3) Activation pruning optimization:
[0206] Specifically, please refer to Figure 11 , Figure 11 which is the schematic diagram for constructing channel attention. First, calculate the activation distribution histogram of each channel in the same interval [min(A), max(A)]. The information entropy I i of the i-th channel is calculated by Equation (1.12), where p j represents the frequency of the j-th group of data appearing in the distribution histogram.
[0207] I i =-∑ j p j log(p j ) (0.12)
[0208] Then, based on principal component analysis, the original feature map matrix is compressed to remove redundant information and obtain its principal components. Specifically, the eigenvectors in the original feature map matrix X are first subtracted by their own mean to obtain the decentralized matrix X′, and then X′ is subjected to singular value decomposition, where X is a matrix composed of h′×w′ groups of eigenvectors, h′ is the height of the activation, and w′ is the width of the activation.
[0209] X′=U∑V T (0.13)
[0210] U∈R h′w′×h′w′ is an orthogonal matrix, and its column vectors are left singular vectors. ∑∈R h′w′×n is a diagonal matrix with n singular values on its diagonal, and U∈R n×n is also an orthogonal matrix, and its column vectors are called right singular vectors. By compressing X in the column direction, the projection of X on the first column vector of V is obtained. Finally, ReLu and normalization are applied to the projection to obtain the activated spatial attention σ∈R h′w′ , that is, single-layer spatial attention:
[0211] σ=norm(ReLu(XV 1 )) (0.14)
[0212] According to the change of spatial attention from the shallow layer to the deep layer in MobileNet (Mobile Network V2) V2, in the shallow layer, the attention is concentrated on the boundary of the image, indicating that the shallow network layer focuses on extracting the local details of the input. As the network deepens, the receptive field gradually becomes larger, and the spatial attention is concentrated on the core part of the foreground object.
[0213] Since the quantization bias of the shallow activation will be transmitted to the deep layer, it is necessary to consider the spatial attention of the target layer and other network layers simultaneously during activation quantization. Therefore, the spatial attention of each layer of activation is summed and normalized to construct the global spatial attention σ g ∈R h′w′ 。
[0214] Based on the above single-layer spatial attention, global spatial attention and channel attention, the activated attention can be obtained:
[0215] M(l,m,n)=I l +α1·σ(m,n)+α2·σ g (m,n) (0.15)
[0216] Among them, M(l,m,n) is the second weight of the initial feature map parameter; α1 and α2 are the balance factors of single-layer spatial attention and global attention; I is the activation distribution histogram of the channel, used to characterize channel attention; σ is the activated spatial attention, used to characterize single-layer spatial attention; σ g is the global spatial attention.
[0217] To sum up, the optimization objective of activation clipping is as follows:
[0218]
[0219] Among them, s l represents the standard deviation of the l-th channel; C is the number of categories; h’ is the height of the feature map; w′ is the width of the feature map; A q is the quantization feature map parameter; A is the initial feature map parameter; A min ,A max are the minimum and maximum values of the activation clipping threshold respectively. It suppresses the adverse effects of numerical fluctuations in different channels. Since the activation amount is much larger than the weight, the consumption time of the particle swarm algorithm is too long, and because the activation data comes from a small number of samples, it may lead to overfitting problems. Therefore, grid search is used to solve the activation clipping threshold.
[0220] Correspondingly, the embodiment of the present application also provides a lightweight depth estimation system. Please refer to Figure 5 , Figure 5It is a module connection diagram of the lightweight depth estimation system provided by the embodiments of the present application. The depth estimation system provided by the embodiments of the present application includes:
[0221] A model construction module 10, which is used to obtain a depth estimation model, input a binocular depth image into the depth estimation model, and obtain a depth estimation result. Among them, the model construction module 10 includes:
[0222] An input sub-module 11, which is used to determine input parameters, and the input parameters include a binocular depth image;
[0223] A feature extraction sub-module 12, which is used to perform feature extraction on the input parameters based on two-dimensional depthwise separable convolution to obtain feature parameters;
[0224] A cost volume acquisition sub-module 13, which is used to obtain a cost volume through the feature parameters;
[0225] A depth feature acquisition module 14, which is used to perform stacked hourglass processing on the cost volume based on three-dimensional pointwise convolution to obtain multi-depth feature information;
[0226] An output sub-module 15, which is used to determine output parameters, and the output parameters include multi-depth feature information, and the multi-depth feature information is used to represent the depth estimation result.
[0227] In some embodiments, the feature extraction sub-module 12 is specifically used for:
[0228] Perform multi-layer two-dimensional depthwise separable convolution operations on the depth image, and use the rectified linear unit activation function for non-linear processing after each layer of two-dimensional depthwise separable convolution operation to obtain a first feature map;
[0229] Process the first feature map through a multi-layer residual network to obtain a second feature map;
[0230] Process the second feature map through a fast spatial pyramid pooling layer to obtain a multi-channel feature map that fuses features of multiple scales, which is the third feature map;
[0231] Fuse the information of the third feature map through a per-pixel convolution layer to obtain feature parameters.
[0232] In some embodiments, the cost volume acquisition sub-module 13 is specifically used for:
[0233] Obtain the feature difference data between the left view and the right view corresponding to each pixel, and perform matching by moving the window, and respectively crop and match the other view for the left view and the right view to obtain a cost volume.
[0234] In some embodiments, the depth feature acquisition module 14 is specifically configured to:
[0235] Stack multiple hourglass networks, in which:
[0236] Perform a lightweight 3D pointwise convolution operation on the cost volume to reduce the resolution of the cost volume and obtain a preliminary prediction result;
[0237] Perform a 3D transposed convolution operation on the preliminary prediction result for upsampling to obtain depth feature information.
[0238] The above has introduced in detail a lightweight depth estimation method and system provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A lightweight depth estimation method, characterized in that: include: Acquire a depth estimation model, input the binocular depth image into the depth estimation model, and obtain a depth estimation result, wherein the step of acquiring the depth estimation model includes: Determining input parameters, wherein the input parameters include a binocular depth image; Performing feature extraction on the input parameter based on two-dimensional depth-separable convolution to obtain feature parameters; Obtaining a cost volume through the characteristic parameters; Performing stacked hourglass processing on the cost volume based on three-dimensional point-by-point convolution to obtain multi-depth feature information; Determining an output parameter, wherein the output parameter includes multiple depth feature information, and the multiple depth feature information is used to characterize a depth estimation result; Performing model training on the depth estimation model based on the depth estimation result, wherein the model training includes a classification task and a regression task; Obtaining a depth map label, and obtaining a classification label and a regression label based on the depth map label: performing clustering processing on the depth map label to obtain a cluster center; Based on the cluster centers and the disparity classification, obtaining a category map; Acquire the classification label based on the category mapping; Based on the classification label and the classification task, obtaining a classification loss; Obtaining a regression loss based on the regression label and the regression task; The total loss function is determined based on the classification loss and the regression loss: the stacked hourglass processing is performed by a multi-layer serially stacked hourglass network, in which the multi-layer hourglass network, The output of each hourglass network is determined using a smoothed absolute error loss function, and a weight coefficient is assigned; wherein the output of the hourglass network with a greater depth is assigned a greater weight coefficient; The output of each hourglass network is subjected to a cross entropy loss function to determine the classification loss and to assign a weight coefficient; wherein the output of the hourglass network with a greater depth is assigned a greater weight coefficient; For the weight coefficient allocation between classification loss and regression loss, dynamic coefficient allocation is used. When the number of optimization iterations is less than the iteration threshold, only the classification loss is used as the total loss function; when the number of iterations is greater than or equal to the iteration threshold, the total loss function is determined based on the function after the weight coefficients are allocated to the classification loss and regression loss. The depth estimation model is optimized based on the total loss function.
2. The lightweight depth estimation method according to claim 1, characterized in that: The step of extracting features from the input parameters based on two-dimensional depth-separable convolution to obtain feature parameters includes: Performing a multi-layer two-dimensional depth-separable convolution operation on the binocular depth image, and performing nonlinear processing using a rectified linear unit activation function after each layer of the two-dimensional depth-separable convolution operation to obtain a first feature map; Processing the first feature map through a multi-layer residual network to obtain a second feature map; Processing the second feature map through a fast spatial pyramid pooling layer to obtain a multi-channel feature map that integrates multiple scale features as a third feature map; The information of the third feature map is fused through a pixel-by-pixel convolution layer to obtain feature parameters.
3. The lightweight depth estimation method according to claim 1, characterized in that: The binocular depth image includes a left view and a right view, and the step of obtaining a cost volume through the feature parameter includes: The difference data of the feature parameters of the left view and the right view corresponding to each pixel are obtained, and matching is performed by moving the window, and the left view and the right view are respectively cropped and matched with another view to obtain the cost volume.
4. The lightweight depth estimation method according to claim 1, characterized in that: The step of performing stacked hourglass processing on the cost volume based on three-dimensional point-by-point convolution to obtain multi-depth feature information includes: Multiple hourglass networks are stacked, in which: Performing a lightweight three-dimensional point-by-point convolution operation on the cost volume to reduce the resolution of the cost volume and obtain a preliminary prediction result; A three-dimensional deconvolution operation is performed on the preliminary prediction result to perform upsampling to obtain the depth feature information.
5. The lightweight depth estimation method according to claim 1, characterized in that: The steps of the classification task include: processing the depth estimation result through multi-layer pyramid three-dimensional lightweight convolution and multi-layer two-dimensional convolution to obtain a disparity estimation map; the steps of the regression task include: performing channel fitting on the disparity estimation map through multi-layer pyramid three-dimensional lightweight convolution and multi-layer two-dimensional convolution, and obtaining a continuous disparity map based on disparity regression.
6. The lightweight depth estimation method according to claim 1, characterized in that: The lightweight depth estimation method further includes: performing linear quantization, weight clipping optimization and / or activation clipping optimization on the depth estimation model, and the linear quantization step includes: The single-precision floating-point data is linearly mapped to the data of the specified bit. The characterization formula includes: Among them, x is the floating point data to be quantized; s is the quantization scale; z is the quantization zero; n is the bit width of the data of the specified bit; round is the rounding operation; clip is the truncation operation; x max is the maximum clipping threshold; x min is the minimum clipping threshold; q is the integer after quantizing the initial target parameter; The steps of weight clipping optimization include: Determining a target convolution kernel in the depth estimation model; Initializing particles in the depth estimation model based on a particle swarm algorithm and determining position range parameters of the particles; Determining the position and velocity of the particle based on the target convolution kernel and the position range parameter; Determine the quantization deviation of the position, and determine that the position corresponding to the minimum quantization deviation is the optimal position; In response to the quantization deviation being within the deviation threshold, determining the optimal position corresponding to the quantization deviation as the optimal weight clipping threshold; In response to the quantized deviation being outside a deviation threshold, updating the velocity based on the optimal position; The step of activating the cropping optimization comprises: Get the activation distribution histogram of each channel in the same interval, which is the channel attention; Compress the original feature map matrix based on principal component analysis to obtain single-layer spatial attention; Sum and normalize the single-layer spatial attention to obtain global spatial attention; Based on the single-layer spatial attention, the global spatial attention and the channel attention, the attention of activation clipping is obtained; and based on principal component analysis, the original activation is compressed, and the step of obtaining the single-layer spatial attention includes: Get the decentralized matrix of the original feature map matrix; Performing singular decomposition on the decentralized matrix, and performing column vector compression on the decomposed data to obtain the projection of the original feature map matrix on the first column vector of the velocity feature; A linear rectification unit and a normalization operation are applied to the projection to obtain the activated spatial attention, which is the single-layer spatial attention.
7. A lightweight depth estimation system, characterized in that: include: A model building module (10), the model building module (10) is used to obtain a depth estimation model, input the binocular depth image into the depth estimation model, and obtain a depth estimation result, wherein the model building module (10) includes: An input submodule (11), the input submodule (11) is used to determine an input parameter, the input parameter comprising a binocular depth image; a feature extraction submodule (12), the feature extraction submodule (12) is used to extract features from the input parameter based on a two-dimensional depth separable convolution to obtain a feature parameter; A cost body acquisition submodule (13), wherein the cost body acquisition submodule (13) is used to acquire a cost body through the feature parameter; A depth feature acquisition module (14), the depth feature acquisition module (14) is used to perform stacked hourglass processing on the cost volume based on three-dimensional point-by-point convolution to obtain multiple depth feature information; An output submodule (15), the output submodule (15) is used to determine an output parameter, the output parameter includes multiple depth feature information, the multiple depth feature information is used to characterize the depth estimation result; based on the depth estimation result, the depth estimation model is trained, the model training includes a classification task and a regression task; obtain a depth map label, and obtain a classification label and a regression label based on the depth map label: Performing clustering processing on the depth map labels to obtain cluster centers; Based on the cluster centers and the disparity classification, obtaining a category map; Acquire the classification label based on the category mapping; Based on the classification label and the classification task, obtaining a classification loss; Obtaining a regression loss based on the regression label and the regression task; The total loss function is determined based on the classification loss and the regression loss: the stacked hourglass processing is performed by a multi-layer serially stacked hourglass network, in which the multi-layer hourglass network, The output of each hourglass network is determined using a smoothed absolute error loss function, and a weight coefficient is assigned; wherein the output of the hourglass network with a greater depth is assigned a greater weight coefficient; The output of each hourglass network is subjected to a cross entropy loss function to determine the classification loss and to assign a weight coefficient; wherein the output of the hourglass network with a greater depth is assigned a greater weight coefficient; For the weight coefficient allocation between classification loss and regression loss, dynamic coefficient allocation is used. When the number of optimization iterations is less than the iteration threshold, only the classification loss is used as the total loss function; when the number of iterations is greater than or equal to the iteration threshold, the total loss function is determined based on the function after the weight coefficients are allocated based on the classification loss and regression loss. The depth estimation model is optimized based on the total loss function.
Citation Information
Patent Citations
Target detection method based on global convolution and local deep convolution fusion
CN111428765A
Binocular stereo matching method based on convolutional neural network
CN117808862A
Cited By
Monocular self-supervision depth estimation method fusing multi-resolution features and global context
CN120976282A
Power transmission line external damage prevention monitoring system and method based on improved DepthPro depth estimation
CN122023862A