A method for depth estimation of the scene in front of a train based on binocular vision
Through binocular vision technology, using multi-scale spatial attention network and error perception enhancement module, the accurate estimation of the scene depth ahead of the train is achieved, solving the problem of low detection accuracy in traditional methods and improving the safety of train driving.
Patent Information
- Application Number
- CN202510378666.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Traditional train depth of field detection methods lack specific details of the scene ahead, resulting in low detection accuracy.
Using a binocular vision-based train scene depth estimation method, the left and right views are obtained through a binocular camera, and then data enhancement is performed to extract multi-scale features using a lightweight feature extractor, and the cost body is constructed and parallax estimation is performed through a multi-scale spatial attention network and error perception enhancement module, and finally the accurate estimation result of the depth of field ahead of the train is obtained.
It improves the recognition accuracy of the depth of field ahead of the train, enhances the generalization ability of the model in different scenarios and weather, and improves the safety of the train driving process.
Smart Images

Figure CN119887851B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for depth estimation of the scene in front of a train based on binocular vision. Background Art
[0002] During the running of a train, objects in the scene in front may all interfere with the normal running of the train. Therefore, depth estimation of the scene in front of the train during the running of the train is crucial for the safe driving of the train.
[0003] Traditional train depth detection methods use radar as the main perception sensor, and its advantages lie in good robustness and strong anti-interference ability.
[0004] However, in practical applications, the traditional method of using radar for depth detection has its own defects, that is, it lacks the perception of specific details of the scene in front, that is, the detection accuracy is not high. Summary of the Invention
[0005] The present invention provides a method for depth estimation of the scene in front of a train based on binocular vision, which is used to solve the defect of low accuracy in depth estimation in front of a train in related technologies. In the solution of this application, through computer vision technology, depth estimation is performed based on the image data of the scene in front of the train, and it has a higher recognition accuracy compared with traditional technologies.
[0006] The present invention provides a method for depth estimation of the scene in front of a train based on binocular vision, including:
[0007] Obtain the left view and the right view of the scene in front of the train through a binocular camera;
[0008] After data augmentation of the left view and the right view, use a lightweight feature extractor to extract features from the left view and the right view to obtain multi-scale features;
[0009] Use the method of grouped correlation to construct a cost volume using features with a resolution of 1 / 4;
[0010] Multiply the geometric features contained in the cost volume by the extended image context features, and then input them into a pre-constructed multi-scale spatial attention network to obtain the aggregation result of the cost volume;
[0011] Perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling;
[0012] Apply the error-aware enhancement module to connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image, and input them into the attention enhancement module. Then, use the double hourglass model to optimize and obtain the depth residual map. Add the initial disparity map and the depth residual map to get the final estimated result of the depth of field in front of the train.
[0013] According to a method for estimating the depth of the scene in front of a train based on binocular vision provided by the present invention, the data enhancement includes: random image enhancement, geometric asymmetry enhancement, and random occlusion enhancement.
[0014] According to a method for estimating the depth of the scene in front of a train based on binocular vision provided by the present invention, the multi-scale spatial attention network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer;
[0015] The first convolutional layer is used to preprocess the fused features;
[0016] The second convolutional layer is used to obtain the multi-scale spatial information of the input data and splice the obtained spatial information in the channel dimension;
[0017] The third convolutional layer is used to convert the spliced spatial information into weights and perform weighted summation on the spatial information based on the weights;
[0018] The fourth convolutional layer is used to output the aggregation result of the cost volume after weighted summation.
[0019] According to a method for estimating the depth of the scene in front of a train based on binocular vision provided by the present invention, the disparity estimation of the aggregation result of the cost volume through disparity regression and obtaining the initial disparity map after upsampling include:
[0020] For the aggregated cost volume, calculate the expected disparity for the first two values at each pixel through the normalized exponential function, and upsample the expected disparity to the original resolution to obtain the initial disparity map.
[0021] According to a method for estimating the depth of the scene in front of a train based on binocular vision provided by the present invention, the construction process of the depth residual map includes:
[0022] Perform a warping operation using the right image and the initial disparity map, and calculate the reconstruction error through the warped right image and the original left image;
[0023] Connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image along the channel dimension, input them into the attention enhancement module, and use the double hourglass model to optimize and obtain the depth residual map.
[0024] A method for estimating the depth of the scene in front of a train based on binocular vision provided by the present invention. The application error perception enhancement module connects the initial disparity map, the left image, the reconstruction error, and the global features of the left image, and inputs them into the attention enhancement module. Then, the depth residual map is optimized using a double hourglass model. The initial disparity map and the depth residual map are added together to obtain the final estimated result of the depth of field in front of the train, including:
[0025] The initial disparity map, the left image, the reconstruction error, and the global features of the left image are connected through the following formula:
[0026] ;
[0027] where, is the initial disparity map, is the left image, is the reconstruction error, the global features of the left image, is the fused feature after connection in the channel dimension;
[0028] The final estimated result of the depth of field in front of the train is calculated through the following formula:
[0029] ;
[0030] ;
[0031] ;
[0032] where, is the attention enhancement module, is the fused feature after attention enhancement, is the double hourglass model, is the depth residual map, is the estimated result of the depth of field in front of the train.
[0033] In the method for estimating the depth of the scene in front of a train based on binocular vision provided by the present invention, the depth of field estimation result is obtained through feature extraction, cost volume construction, cost volume aggregation, and based on the aggregated cost volume. In this process, the cost volume aggregation is realized through a pre-constructed aggregation network and a multi-scale spatial attention network. Using multi-scale spatial attention, the geometric features of the cost volume and the extended context features are more effectively combined together, which can significantly improve the aggregation effect. Further, an error perception enhancement model is proposed when calculating the depth of field estimation result, introducing the global feature information of the left side, and combining with the attention mechanism to realize the accuracy optimization of the initial disparity map. At the same time, the network has good generalization in different scenarios and weather conditions. In summary, the method provided by this application can accurately identify the depth of field in front of the train and improve the safety of the train during driving. Brief Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0035] Figure 1 is one of the schematic flowcharts of a method for depth estimation of the scene in front of a train based on binocular vision provided by an embodiment of the present invention;
[0036] Figure 2 is the second schematic flowchart of a method for depth estimation of the scene in front of a train based on binocular vision provided by an embodiment of the present invention;
[0037] Figure 3 is the schematic structural diagram of a multi-scale spatial attention network provided by an embodiment of the present invention;
[0038] Figure 4 is the schematic structural diagram of an error perception enhancement module provided by an embodiment of the present invention;
[0039] Figure 5 is the schematic structural diagram of an attention enhancement module provided by an embodiment of the present invention;
[0040] Figure 6 is the schematic structural diagram of a system for depth estimation of the scene in front of a train based on binocular vision provided by an embodiment of the present invention;
[0041] Figure 7 is the schematic physical structure diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0042] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0043] Figure 1 is one of the schematic flowcharts of a method for depth estimation of the scene in front of a train based on binocular vision provided by an embodiment of the present invention.
[0044] Figure 2 is the second schematic flowchart of a method for depth estimation of the scene in front of a train based on binocular vision provided by an embodiment of the present invention.
[0045] As Figure 1 and Figure 2 shown, this embodiment provides a method for depth estimation of the scene in front of a train based on binocular vision, which can be executed by a processing device of the train, such as a processor of a driving system or a cloud processor. The method includes:
[0046] Step 101, obtain the left view and the right view of the scene in front of the train through a binocular camera.
[0047] The depth of field refers to the scene depth in this embodiment. Further, in this solution, by detecting the scene depth of the scene in front of the train, the relative distance information between each object in the scene in front of the train and the train can be determined. This information is crucial for train driving and can help the train driver or the train's automatic driving system identify potential dangers in time and respond in time.
[0048] Step 102, after performing data augmentation on the left view and the right view, use a lightweight feature extractor to extract features from the left view and the right view to obtain multi-scale features.
[0049] In this embodiment, random image augmentation can be achieved by randomly adjusting the brightness, gamma value, contrast, and saturation of the left and right images; geometric asymmetry augmentation can be achieved by random rotation, displacement, and random cropping; a random area in the image can be occluded and filled with the average value of its pixels to simulate the situation where some areas in the actual scene are invisible, thereby achieving random occlusion augmentation.
[0050] The purpose of data augmentation is to increase the diversity of data and improve the generalization ability of the model.
[0051] In implementation, a U-Net style encoder-decoder structure can be used to extract local features from the image pair. Specifically, the encoder structure can use the pre-trained MobileNetV2 as a lightweight backbone to obtain features at four scales; in the decoder stage, the spatial resolution of the feature map is gradually restored through three upsampling modules, and the feature map of the corresponding layer in the encoding stage is fused with the feature map of the current layer through skip connections, finally obtaining four scales of feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original resolution. Then, the feature map with a resolution of 1 / 4 is used to construct the cost volume, and the features at the other 1 / 8, 1 / 16, and 1 / 32 scales can be used as guidance for subsequent cost aggregation.
[0052] In practical applications, the size of the image data can be H×W×3, where H represents the height of the image, W represents the width of the image, and 3 represents the three channels of the image.
[0053] Step 103: Use the grouping-related method to construct a cost volume using features at 1 / 4 resolution.
[0054] The idea of grouping-related is to group features, and each group of features only calculates the correlation with another group of features, thereby reducing the computational amount. In this embodiment, the feature vectors can be grouped according to the channel dimension, that is, all channels are evenly divided into several groups, and each feature group contains the same number of feature channels.
[0055] The mathematical expression of grouping-related is:
[0056] ;
[0057] where represents the inner product of vectors, represents the channel dimension of features, represents the number of groups, is the feature map of the th group of the left feature map, is the feature map of the th group of the right feature map. The group correlation cost volume can be regarded as a set of cost volumes, and each cost volume is calculated by the corresponding feature group.
[0058] Step 104: During the cost aggregation process, multiply the geometric features contained in the cost volume by the extended image context features, and then input them into the pre-constructed multi-scale spatial attention network to optimize the aggregation result of the cost volume.
[0059] The structure of the aggregation network of the model includes three downsampling layers and three upsampling layers. Each downsampling layer includes a 3D convolution with a kernel of 3×3×3 and a stride of 2, and a 3D convolution with a kernel of 3×3×3 and a stride of 1. Each upsampling layer includes a 3D transposed convolution with a kernel of 4×4×4 and a stride of 2, and a 3D convolution with a kernel of 3×3×3 and a stride of 1.
[0060] The context features in this embodiment correspond to the image features at 1 / 8, 1 / 16, and 1 / 32 scales generated in step 102 above. The geometric features can be obtained through the three downsampling layers of the aggregation network. If the extended context features are directly added to the geometric features, the differences in the disparity dimension will be ignored, resulting in ineffective fusion. During the cost aggregation process, we alternately use the multi-scale attention fusion module and the upsampling module. In this embodiment, the multi-scale spatial attention network can more effectively combine the context and geometric information.
[0061] Figure 3It is a schematic diagram of the structure of the multi-scale spatial attention network provided by an embodiment of the present invention.
[0062] As Figure 3 shown, in an exemplary embodiment, the multi-scale spatial attention network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer;
[0063] Among them, the first convolutional layer multiplies the extended context feature by the geometric feature and then performs preprocessing through a 3D convolution operation. After that, the preprocessed feature is concatenated with the original feature in the channel dimension;
[0064] One branch of the second convolutional layer includes 2 3D convolutions with a kernel of 3×3×3 for obtaining low-level spatial information; the other branch includes 1 3D convolution with a kernel of 3×3×3, 1 3D convolution with a kernel of 5×5×5, 1 3D convolution with a kernel of 7×7×7, and 1 3D convolution with a kernel of 9×9×9, which are used to capture features at different scales, gradually increase the receptive field, and obtain richer multi-level spatial information. And the obtained results are concatenated in the channel dimension;
[0065] The third convolutional layer includes 1 3D convolution with a kernel of 1×1×1, which is used to adjust the channel dimension and convert the concatenated multi-scale spatial attention vector into weights;
[0066] The fourth convolutional layer is used to output the aggregation result of the weighted sum of the cost volume.
[0067] In practical applications, the above first convolutional layer satisfies the following formula:
[0068] ;
[0069] ;
[0070] ;
[0071] Among them, is the extended context feature, is the geometric feature, is the Hadamard product, is the fused feature, is the preprocessed feature, is the concatenated feature.
[0072] The above second convolutional layer satisfies the following formula:
[0073] ;
[0074] Among them, is the concatenated spatial information, For the splicing operation, is the first branch, is the second branch, is the input data.
[0075] The third convolutional layer satisfies the following formula:
[0076] ;
[0077] wherein, is the weight, is function.
[0078] The fourth convolutional layer satisfies the following formula:
[0079] ;
[0080] wherein, is the aggregation result of the output cost volume.
[0081] Step 105, perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling.
[0082] In practical applications, for the aggregated cost volume, the first two values at each pixel can be selected, and these values can be used to perform activation function to calculate the expected disparity, with a size of B×1×H / 4×W / 4. Then, the disparity map can be upsampled to the original resolution B×1×H×W using the weights around each pixel to obtain the initial disparity map.
[0083] Step 106, apply the error-aware enhancement module, connect the initial disparity map, the left image, the reconstruction error, and the left image global feature, input them into the attention enhancement module, and then optimize using the double hourglass model to obtain the depth residual map. Add the initial disparity map and the depth residual map to obtain the final estimation result of the depth of field in front of the train.
[0084] Figure 4 is the structural schematic diagram of the error-aware enhancement module provided by the embodiment of the present invention.
[0085] In implementation, the initial disparity map generated in the above embodiment can be used for estimating the depth of field in front of the train, but there are still certain limitations when applied to complex and fine regions. Therefore, in the solution of this embodiment, error-aware enhancement of the initial disparity map is performed through operations such as calculating the reconstruction error and attention enhancement. Specifically, it is to restore fine image details by integrating information from the left image and the right image.
[0086] Such as Figure 4As shown, in an exemplary embodiment, a distorted right image is obtained based on the right image and the initial disparity map, and a reconstruction error is determined based on the distorted right image and the original left image, including:
[0087] The distorted right image is obtained through the following formula:
[0088] ;
[0089] where is the distorted right image, is the data distortion operation, is the right image, is the initial disparity map;
[0090] The data distortion operation therein is a processing method for data, and the image data can be processed through methods such as Euclidean transformation, similarity transformation, affine transformation, and projective transformation.
[0091] The reconstruction error is determined through the following formula:
[0092] ;
[0093] where is the reconstruction error, is the left image.
[0094] In practical applications, during the process of disparity aggregation, continuous encoding-decoding steps are used, which will cause the feature map to lose some of its original semantic information and make the initial disparity map relatively blurred. To enhance the disparity prediction ability of the initial disparity map in pathological regions, more original feature information is introduced in this embodiment. The left image features with 1 / 4 resolution containing rich semantic information are upsampled to the original resolution through the context network to obtain global feature information, which contains more semantic information and edge high-frequency information. The initial disparity map, the left image, the reconstruction error, and the global features of the left image are connected along the channel dimension, and the formula is as follows:
[0095] ;
[0096] where is the initial disparity map, is the left image, is the reconstruction error, is the global feature of the left image, is the fused feature after connection in the channel dimension;
[0097] In traditional residual learning, there is a lack of evidence indicating the occurrence of errors. To alleviate this problem, in this embodiment, the spliced features are input into the attention enhancement module, so that the inaccurate regions are more concerned during the residual learning process. The attention enhancement module is as Figure 5As shown, it includes a channel attention layer and a spatial attention layer. First, it enters the channel attention layer to weight the importance of each channel, and then inputs the spatial attention layer to obtain the final weighted result. The formula is as follows:
[0098] ;
[0099] Among them, is the attention enhancement module, is the fused feature after attention enhancement.
[0100] It is further optimized using the double hourglass model to obtain the depth residual map, and then the initial disparity map and the residual map are added together to obtain the estimated result of the final depth of field. The formula is as follows:
[0101] ;
[0102] ;
[0103] Among them, is the double hourglass model, is the depth residual map, is the estimated result of the depth of field in front of the train.
[0104] In practical applications, the double hourglass model is constructed on the cascaded U-Net structure. Each U-Net structure contains 4 upsampling layers, 4 downsampling layers and skip connections, and each layer contains two 3×3 two-dimensional convolutions.
[0105] In an exemplary embodiment, the aggregation network and the multi-scale spatial attention network can both be trained in an end-to-end manner. The training process can use the Smooth L1 function for training, and the total loss function of the training process conforms to the following formula:
[0106] ;
[0107] Among them, and are both loss weights, is the initial disparity map with 1 / 4 resolution described in the above embodiment, is the final disparity map with full resolution, represents the true ground disparity map.
[0108] In the solution of this application, depth estimation of the scene in front of the train is based on binocular vision. While generating a disparity map with higher precision, a high speed is maintained. By using the method of multi-scale spatial attention fusion and error perception enhancement, the receptive field is expanded, spatial information of different scales is fused, and the ability of error perception refinement is enhanced. The matching accuracy of the model in shadow, edge, weak texture or textureless areas can be significantly improved, thereby reducing the occurrence of false matching and facilitating the subsequent work. The solution of this application has strong generalization ability. Data augmentation processing is performed on the left and right images to increase data diversity, enabling the model to achieve good results when predicting different scenes.
[0109] Next, a depth estimation system for the scene in front of the train based on binocular vision provided by the present invention will be described. The depth estimation system for the scene in front of the train based on binocular vision described below can be correspondingly referred to the depth estimation method for the scene in front of the train based on binocular vision described above.
[0110] Figure 6 It is a schematic structural diagram of a depth estimation system for the scene in front of the train based on binocular vision provided by an embodiment of the present invention.
[0111] As Figure 6 shown, a depth estimation system for the scene in front of the train based on binocular vision provided in this embodiment includes:
[0112] A data acquisition module 601, configured to obtain a left view and a right view of the scene in front of the train through a binocular camera
[0113] A feature extraction module 602, configured to perform data augmentation on the left view and the right view, and then use a lightweight feature extractor to extract features from the left view and the right view to obtain multi-scale features;
[0114] A cost volume construction module 603, configured to use the method of grouped correlation to construct a cost volume by using features with a 1 / 4 resolution;
[0115] A cost volume aggregation module 604, configured to multiply the geometric features contained in the cost volume by the extended image context features, and then input them into a pre-constructed multi-scale spatial attention network to obtain an aggregation result of the cost volume;
[0116] A disparity map generation module 605, configured to perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling;
[0117] A depth of field estimation module 606 is configured to apply an error-aware enhancement module to connect an initial disparity map, a left image, a reconstruction error, and global features of the left image, input them into an attention enhancement module, and then use a double hourglass model to optimize and obtain a depth residual map. Add the initial disparity map and the depth residual map to obtain a final estimation result of the depth of field in front of the train.
[0118] For the specific implementation method of a depth estimation system for the scene in front of a train based on binocular vision provided in this embodiment, reference may be made to the above embodiment for implementation, and details are not described herein again.
[0119] Figure 7 An entity structure diagram of an electronic device is exemplified, as Figure 7 shown. The electronic device may include: a processor 710, a communication interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 may call logical instructions in the memory 730 to execute a method for depth estimation of the scene in front of a train based on binocular vision. The method includes:
[0120] Obtain a left view and a right view of the scene in front of the train through a binocular camera;
[0121] After performing data enhancement on the left view and the right view, use a lightweight feature extractor to perform feature extraction on the left view and the right view to obtain multi-scale features;
[0122] Use a grouped correlation method to construct a cost volume using features at 1 / 4 resolution;
[0123] Multiply the geometric features contained in the cost volume by the extended image context features, and then input them into a pre-constructed multi-scale spatial attention network to obtain an aggregation result of the cost volume;
[0124] Perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling;
[0125] Apply an error-aware enhancement module to connect an initial disparity map, a left image, a reconstruction error, and global features of the left image, input them into an attention enhancement module, and then use a double hourglass model to optimize and obtain a depth residual map. Add the initial disparity map and the depth residual map to obtain a final estimation result of the depth of field in front of the train.
[0126] In addition, when the logical instructions in the aforementioned memory 730 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0127] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for depth estimation of the scene in front of a train based on binocular vision provided by the above-mentioned various methods. The method includes:
[0128] Obtain the left view and the right view of the scene in front of the train through a binocular camera;
[0129] After performing data augmentation on the left view and the right view, use a lightweight feature extractor to perform feature extraction on the left view and the right view to obtain multi-scale features;
[0130] Use the method of grouped correlation to construct a cost volume using features with a resolution of 1 / 4;
[0131] Multiply the geometric features contained in the cost volume by the extended image context features, and then input them into a pre-constructed multi-scale spatial attention network to obtain the aggregation result of the cost volume;
[0132] Perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling;
[0133] Apply an error-aware enhancement module, connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image, input them into the attention enhancement module, and then use a double hourglass model to optimize and obtain a depth residual map. Add the initial disparity map and the depth residual map to obtain the final estimation result of the depth of field in front of the train.
[0134] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for depth estimation of the scene in front of a train based on binocular vision provided by the above-mentioned various methods. The method includes:
[0135] Obtain the left view and the right view of the scene in front of the train through a binocular camera;
[0136] After performing data augmentation on the left view and the right view, use a lightweight feature extractor to perform feature extraction on the left view and the right view to obtain multi-scale features;
[0137] Use the method of grouped correlation to construct a cost volume using features with a resolution of 1 / 4;
[0138] Multiply the geometric features contained in the cost volume by the extended image context features, and then input them into a pre-constructed multi-scale spatial attention network to obtain the aggregation result of the cost volume;
[0139] Perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling;
[0140] Apply an error-aware enhancement module, connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image, input them into the attention enhancement module, and then use a double hourglass model to optimize and obtain a depth residual map. Add the initial disparity map and the depth residual map to obtain the final estimation result of the depth of field in front of the train.
[0141] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.
[0142] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for depth estimation of the scene in front of a train based on binocular vision, characterized in that, Including: Obtain the left view and the right view of the scene in front of the train through a binocular camera; After performing data augmentation on the left view and the right view, use a lightweight feature extractor to extract features from the left view and the right view to obtain multi-scale features; Use the method of grouped correlation to construct a cost volume using features at 1 / 4 resolution; Multiply the geometric features contained in the cost volume by the extended image context features, and then input them into a pre-constructed multi-scale spatial attention network to obtain the aggregation result of the cost volume; Perform disparity estimation on the aggregation result of the cost volume through disparity regression, and obtain an initial disparity map after upsampling; Apply an error-aware enhancement module, connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image, input them into the attention enhancement module, and then use a double hourglass model to optimize and obtain a depth residual map. Add the initial disparity map and the depth residual map to obtain the final estimation result of the depth of field in front of the train; The multi-scale spatial attention network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, and a fourth convolutional layer; Among them, the first convolutional layer multiplies the extended context features by the geometric features and performs preprocessing through 3D convolution operations, and then splices the preprocessed features and the original features in the channel dimension; One branch of the second convolutional layer includes 2 3D convolutions with kernels of 3×3×3 to obtain low-level spatial information; the other branch includes 1 3D convolution with a kernel of 3×3×3, 1 3D convolution with a kernel of 5×5×5, 1 3D convolution with a kernel of 7×7×7, and 1 3D convolution with a kernel of 9×9×9 to capture features at different scales, gradually increase the receptive field, obtain richer multi-level spatial information, and splice the obtained results in the channel dimension; The third convolutional layer includes 1 3D convolution with a kernel of 1×1×1 to adjust the channel dimension and convert the spliced multi-scale spatial attention vector into weights; The fourth convolutional layer is used to output the aggregation result of the weighted sum of the cost volume; The first convolutional layer satisfies the following formula: ; ; ; Among them, is the extended context feature, is the geometric feature, is the Hadamard product, is the fused feature, is the preprocessed feature, is the concatenated feature; The second convolutional layer satisfies the following formula: ; Among them, is the spliced spatial information, is the splicing operation, is the first branch, is the second branch, is the input data; The third convolutional layer satisfies the following formula: ; Among them, is the weight, is the sigmoid function; The fourth convolutional layer satisfies the following formula: ; Among them, is the aggregation result of the output cost volume.
2. The method for depth estimation of the scene in front of a train based on binocular vision according to claim 1, wherein, The data augmentation includes: random image augmentation, geometric asymmetry augmentation, and random occlusion augmentation.
3. A method for depth estimation of the scene in front of a train based on binocular vision according to claim 1, characterized in that, The performing disparity estimation on the aggregation result of the cost volume through disparity regression and obtaining an initial disparity map after upsampling includes: For the aggregated cost volume, calculate the expected disparity for the first two values at each pixel through a normalized exponential function, and upsample the expected disparity to the original resolution to obtain an initial disparity map.
4. A method for depth estimation of the scene in front of a train based on binocular vision according to claim 1, characterized in that, The construction process of the depth residual map includes: Perform a warping operation using the right image and the initial disparity map, and calculate the reconstruction error through the warped right image and the original left image; Connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image along the channel dimension, input them into the attention enhancement module, and use a double hourglass model to optimize and obtain a depth residual map.
5. A method for depth estimation of the scene in front of a train based on binocular vision according to claim 1, characterized in that, The application error perception enhancement module connects the initial disparity map, the left image, the reconstruction error, and the global features of the left image, inputs them into the attention enhancement module, and then uses the double hourglass model to optimize and obtain the depth residual map. Adding the initial disparity map and the depth residual map together gives the final estimated result of the depth of field in front of the train, including: Connect the initial disparity map, the left image, the reconstruction error, and the global features of the left image in the channel dimension through the following formula: ; Among them, is the initial disparity map, is the left image, is the reconstruction error, is the global feature of the left image, is the fused feature after concatenation in the channel dimension; Calculate the final estimated result of the depth of field in front of the train through the following formula: ; ; ; Among them, is the attention enhancement module, is the fused feature after attention enhancement, is the double hourglass model, is the deep residual map, is the estimated result of the depth of field in front of the train.
Citation Information
Patent Citations
Binocular depth estimation method and system based on attention mechanism and multilevel cost body
CN116258758A
Multi-scale feature fusion binocular parallax estimation algorithm based on cavity convolution
CN116612078A
Attention-guided multi-view depth estimation method
CN117911480A
Stereo matching method and system based on context geometric cube and distortion parallax optimization
CN119599967A