BEV elevation estimation method and system based on binocular data
Through the BEV elevation estimation method based on binocular data, the perception model is trained using lidar and IMU data, and the problem of low efficiency and low accuracy in complex terrain is solved, and efficient and accurate elevation estimation is achieved, which is suitable for complex terrain and resource-constrained equipment.
Patent Information
- Application Number
- CN202510480319.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
The existing elevation estimation technology is inefficient, has low accuracy and poor reliability in complex terrain. It is difficult to accurately capture potholes and irregular surface features in traditional camera perspective angles, and the depth estimation deviation is large and is susceptible to changes in perspective angles.
BEV elevation estimation method based on binocular data is adopted, and the perception model is trained using lidar point cloud and IMU data. The feature extraction network and elevation classification network are combined with the LiteMBConv module and the Hourglass module to perform multi-scale feature extraction and elevation classification, and the model is optimized using consistent voxel features and cross entropy loss function.
It improves the efficiency and accuracy of elevation estimation, enhances the stability and adaptability of the model in complex scenarios, is suitable for resource-constrained edge devices and real-time computing scenarios, reduces costs and improves the accuracy and robustness of elevation estimation.
Smart Images

Figure CN120388230A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and intelligent perception, and more specifically, relates to a BEV (Bird's Eye View) elevation estimation method and system based on binocular data, which is particularly applicable to terrain elevation estimation tasks in complex scenarios such as mine bulldozing sites. Background Art
[0002] In the existing elevation estimation technologies, the depth estimation method based on the traditional perspective (camera perspective) has significant defects: due to the limited camera perspective, it is difficult to accurately and clearly capture the potholes and irregular undulations on the road surface; at the same time, the depth estimation accuracy based on a single perspective is susceptible to factors such as angle changes and ground inclination, resulting in insufficient accuracy and stability of elevation estimation.
[0003] Specifically, the depth estimation from the traditional camera perspective has the following problems:
[0004] · Sparse and unclear features: It is difficult to effectively extract potholes and irregular surface features. The number of layers stacked by the used feature extraction network MobileNet-V3 reaches 15 layers, resulting in slow feature extraction speed and low efficiency;
[0005] · The depth direction is inconsistent with the elevation direction, resulting in a large depth estimation deviation and affecting the accuracy of elevation estimation;
[0006] · The accuracy is susceptible to perspective changes and has poor reliability when used in complex terrain conditions.
[0007] Therefore, the existing elevation estimation technologies have technical problems such as low efficiency, low accuracy, and poor reliability when used in complex terrains. Summary of the Invention
[0008] In view of the above defects or improvement requirements of the existing technology, the present invention provides a BEV elevation estimation method and system based on binocular data, thereby solving the technical problems of low efficiency, low accuracy, and poor reliability of the existing elevation estimation technologies when used in complex terrains.
[0009] To achieve the above object, according to one aspect of the present invention, there is provided a BEV elevation estimation method based on binocular data, including:
[0010] Collect binocular images and IMU data of the scene to be estimated, input them into a trained perception model, and output the elevation classification result of the scene;
[0011] The perception model includes a feature extraction network and an elevation classification network, and is trained in the following manner:
[0012] Collect the lidar point cloud data, binocular images, and IMU data of the acquisition scene, and use the lidar point cloud data to obtain the true value of the elevation classification of the scene;
[0013] Input the binocular images into the feature extraction network to extract visual features. Use the IMU data to map the three-dimensional voxel space in the BEV view to the binocular images, and fuse the voxel grids in the three-dimensional voxel space with the visual features of the binocular images to obtain consistent voxel features. Input the consistent voxel features into the elevation classification network, and use the error between the output elevation classification prediction value and the true value of the elevation classification as the loss function, and backpropagate to update the parameters of the perception model, and train until convergence to obtain a trained perception model.
[0014] Preferably, before using the lidar point cloud data to obtain the true value of the elevation classification of the scene,
[0015] Use the rotation matrix and displacement vector provided by the IMU data to correct the pose of the binocular images. Through the external parameter matrix from the lidar to the camera, convert the lidar point cloud data in the lidar coordinate system to the camera coordinate system, so that the lidar point cloud data and the binocular images are pixel-level spatially aligned.
[0016] Preferably, the feature extraction network includes LiteMBConv modules and convolution modules,
[0017] The LiteMBConv module has an MBConv structure, and the SE channel attention mechanism in the MBConv structure is replaced by a low-rank grouped attention mechanism. The number of LiteMBConv modules is multiple. Multiple LiteMBConv modules perform multi-scale feature extraction on the binocular images. The features extracted by the subsequent LiteMBConv modules except the first LiteMBConv module are interpolated and upsampled and then concatenated with the features extracted by the first LiteMBConv module to obtain a fused feature map. The convolution module performs convolution on the fused feature map to extract visual features.
[0018] Preferably, the elevation classification network includes: Hourglass modules and convolution modules,
[0019] The Hourglass module is used to extract global context features through downsampling and upsampling. The consistent voxel features are linearly interpolated after passing through multiple cross-arranged Hourglass modules and convolution modules to obtain the probability distribution of elevation classification. Perform Softmax normalization on the probability distribution of elevation classification to obtain the elevation classification prediction value.
[0020] Preferably, a hybrid attention mechanism is embedded in the Hourglass module. The hybrid attention mechanism fuses local multi-head self-attention and global multi-head self-attention, which can adaptively strengthen the features of key regions and enhance the adaptability of the model under rough terrains. By using the learnable weight parameter λ to fuse local features and global features, the calculation formula for the fusion process is:
[0021] Attention = λ·LMSA(X) + (1 - λ)·GMSA(X)
[0022] where Attention represents the fused attention features, λ represents the learnable weight parameter used to adaptively adjust the fusion ratio of local features and global features, LMSA(X) represents the local features extracted by local multi-head self-attention from X, GMSA(X) represents the global features extracted by global multi-head self-attention from X, and X represents the input consistent voxel features.
[0023] Preferably, the consistent voxel features are obtained in the following way:
[0024] A three-dimensional voxel space in the BEV view is established in the ground coordinate system. The three-dimensional voxel space is transformed from the ground coordinate system to the camera coordinate system through the rotation matrix and camera position provided by the IMU data. The three-dimensional voxel space is mapped to the image pixel coordinate system using the camera intrinsics, forming a projection index relationship between the voxel grid in the three-dimensional voxel space and the visual features in the binocular images. The visual features corresponding to a voxel grid in the binocular images are extracted using the projection index relationship and multiplied element by element to obtain the consistent voxel features.
[0025] Preferably, the loss function is:
[0026]
[0027] where L height represents the total loss during the training process, M(vg) represents whether the voxel grid vg where the consistent voxel features are located is valid. If the true elevation classification value of the voxel grid vg where the consistent voxel features are located is within the elevation interval, the voxel grid is valid, M(vg) = 1; otherwise, the voxel grid is invalid, M(vg) = 0. E(c, vg) represents the true elevation classification value of the voxel grid vg where the consistent voxel features are located, c is the class number representing the number of the elevation interval, ∑ represents the summation over all valid voxel grids vg and all classes c, and N c represents the total number of classes in the elevation classification task, and elepred(·, vg) is the predicted elevation classification value of the voxel grid vg where the consistent voxel features are located.
[0028] Preferably, the true elevation classification value is calculated in the following way:
[0029] Iteratively register and fuse consecutive frames of lidar point cloud data through the Iterative Closest Point (ICP) algorithm, and map the fused lidar point cloud data into a 3D voxel grid through the extrinsic parameter matrix from lidar to camera to generate the ground truth of elevation classification.
[0030] According to another aspect of the present invention, there is provided a BEV elevation estimation system based on binocular data, including:
[0031] A preprocessing module, configured to collect lidar point cloud data, binocular images, and IMU data of a scene, and obtain the ground truth of elevation classification of the scene by using the lidar point cloud data;
[0032] A training module, configured to input the binocular images into a feature extraction network to extract visual features, use the IMU data to map the 3D voxel space in the BEV view to the binocular images, fuse the visual features of the voxel grids in the 3D voxel space mapped to the binocular images to obtain consistent voxel features, input the consistent voxel features into an elevation classification network, use the error between the output elevation classification prediction value and the ground truth of elevation classification as a loss function, and backpropagate to update the parameters of the perception model until convergence to obtain a trained perception model;
[0033] An elevation estimation module, configured to collect binocular images and IMU data of a scene to be estimated, input them into the trained perception model, and output the elevation classification result of the scene.
[0034] According to another aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, a BEV elevation estimation method based on binocular data is implemented.
[0035] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0036] (1) During training, this application uses lidar point cloud data, binocular images, and IMU data. During real-time estimation, lidar point cloud data is not required, which can reduce costs. The present invention uses IMU data to map the three-dimensional voxel space in the BEV view to binocular images to achieve BEV elevation estimation. The visual feature fusion that maps the voxel grids in the three-dimensional voxel space to binocular images can effectively suppress the negative impact of one-sided image mis-matched features on elevation estimation, further improving the stability and reliability of the model in complex scenarios. The present invention uses a trained perception model to perform elevation estimation on the scene to be estimated, with high efficiency. At the same time, it directly uses elevation estimation without performing depth estimation, overcoming the defect of large deviation in traditional depth estimation and improving the elevation estimation accuracy. This shows that the present invention provides a BEV elevation estimation method that is highly efficient, accurate, stable, and reliable and applicable to complex scenarios.
[0037] (2) The present invention uses the rotation matrix and displacement vector provided by IMU data to correct the attitude of binocular images to avoid image distortion caused by camera vibration or tilt. Then, through the extrinsic parameter matrix from lidar to camera, the three-dimensional point cloud data obtained in the lidar coordinate system is converted to a position consistent with the camera coordinate system, enabling the lidar point cloud data and binocular images to be pixel-level spatially aligned. This effectively overcomes problems such as large angular deviation, sparse features, and inability to adapt to complex terrains in traditional image-view elevation estimation, significantly improving the accuracy and stability of elevation estimation, and can also effectively reduce errors and instabilities in the subsequent feature extraction process.
[0038] (3) In the design of the LiteMBConv module, considering the disadvantage of the large computational overhead of the SE module in lightweight networks, the SE channel attention mechanism in the traditional MBConv structure is replaced with the low-rank grouped attention mechanism ULSAM. This structural optimization significantly reduces the network parameter quantity and computational complexity, improving the network operation efficiency. The feature extraction network of the present invention is a lightweight multi-scale feature extraction network based on LiteMBConv. The features extracted by the subsequent LiteMBConv modules except the first LiteMBConv module are interpolated and upsampled and then concatenated with the features extracted by the first LiteMBConv module. This structure can fully fuse spatial hierarchical information of different scales while maintaining high computational performance, optimizing the ability to capture different terrain details, and improving the feature stability and expression ability of binocular images under disparity changes and texture differences. It is particularly suitable for resource-constrained edge devices and real-time computing scenarios, significantly improving the model deployment flexibility and real-time processing ability while ensuring elevation estimation accuracy.
[0039] (4) In the present invention, in the elevation classification network, Hourglass realizes the integration of features at different spatial scales through the downsampling-upsampling process. By adopting the Hourglass structure and the global-local self-attention mechanism, the network can capture both local terrain details and global scene structure features under lightweight conditions, adaptively enhance the features of key regions, enhance the adaptability of the model in rugged terrains, effectively improve the classification performance in complex terrains such as steep slopes and obstacle areas, and significantly improve the elevation classification performance and generalization ability of the model. To achieve the best local-global information fusion, the present invention uses a learnable parameter λ to dynamically adjust the ratio of local and global attention. This design effectively enhances the model's ability to understand local terrain changes and global structure features, and greatly improves the generalization performance and elevation estimation accuracy of the model.
[0040] (5) After completing the two-dimensional feature extraction, to accurately map the two-dimensional visual features to three-dimensional spatial positions, the present invention utilizes the projection index relationship between the voxel grid in the three-dimensional voxel space and the visual features in the binocular images, extracts the visual features corresponding to the voxel grid in the binocular images, and fuses the two-dimensional features of the left and right binocular images by element-wise multiplication to obtain consistent voxel features. This fusion method can effectively filter out the mismatched features brought by the single-side view, highlight the stable features jointly confirmed by the left and right views, and improve the accuracy and robustness of elevation estimation.
[0041] (6) To achieve efficient classification for the elevation estimation task, the present invention transforms the continuous elevation prediction problem into a classification problem of discrete categories, and uses the cross-entropy loss function to supervise the training of the network. Specifically, the model is optimized through the cross-entropy loss between the elevation classification prediction value output by the network and the true value, so as to achieve model convergence quickly and efficiently. This classification training method can significantly improve the training speed and stability of the network model while ensuring sufficient elevation accuracy, and is particularly suitable for real-time application scenarios.
[0042] (7) Due to the sparsity problem of single-frame lidar data, the present invention adopts continuous multi-frame point cloud data, and uses the ICP (Iterative Closest Point) algorithm to register and fuse the multi-frame point clouds to generate higher-density and more accurate elevation ground truth data, providing reliable supervision information for the subsequent training of the network model and improving the accuracy of the elevation estimation results. Brief Description of the Drawings
[0043] Figure 1 is a flowchart of a BEV elevation estimation method based on binocular data provided by an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of camera coordinate system conversion provided by an embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of voxel grid division provided by an embodiment of the present invention;
[0046] Figure 4(a) is a structural diagram of the LiteMBConv module using ULSAM provided by an embodiment of the present invention;
[0047] Figure 4(b) is a structural diagram of the feature extraction network provided by an embodiment of the present invention;
[0048] Figure 5(a) is a schematic diagram of the global-local self-attention mechanism fusion provided by an embodiment of the present invention;
[0049] Figure 5(b) is a schematic diagram of the structure of the Hourglass module provided by an embodiment of the present invention;
[0050] Figure 5(c) is a schematic diagram of the structure of the elevation classification network provided by an embodiment of the present invention;
[0051] Figure 6(a) is a binocular image in the horizontal scenario provided by an embodiment of the present invention;
[0052] Figure 6(b) is an elevation estimation map in the horizontal scenario provided by an embodiment of the present invention;
[0053] Figure 6(c) is an elevation error map in the horizontal scenario provided by an embodiment of the present invention;
[0054] Figure 7(a) is a binocular image in the high-low drop scenario provided by an embodiment of the present invention;
[0055] Figure 7(b) is an elevation estimation map in the high-low drop scenario provided by an embodiment of the present invention;
[0056] Figure 7(c) is an elevation error map in the high-low drop scenario provided by an embodiment of the present invention;
[0057] Figure 8(a) is a binocular image in the pothole scenario provided by an embodiment of the present invention;
[0058] Figure 8(b) is an elevation estimation map in the pothole scenario provided by an embodiment of the present invention;
[0059] Figure 8(c) is an elevation error map in the pothole scenario provided by an embodiment of the present invention. Detailed implementation manners
[0060] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0061] As shown Figure 1 in the figure, a BEV elevation estimation method based on binocular data includes:
[0062] Collect binocular images and IMU data of the scene to be estimated, input them into a trained perception model, and output the elevation classification result of the scene;
[0063] The perception model includes a feature extraction network and an elevation classification network, and is trained in the following way:
[0064] Collect lidar point cloud data, binocular images and IMU data of the scene, and use the lidar point cloud data to obtain the true value of the elevation classification of the scene;
[0065] Input the binocular image into the feature extraction network to extract visual features. Use the IMU data to map the three-dimensional voxel space in the BEV view to the binocular image, and fuse the voxel grid in the three-dimensional voxel space to the visual features of the binocular image to obtain consistent voxel features. Input the consistent voxel features into the elevation classification network, and use the error between the output elevation classification prediction value and the true value of the elevation classification as the loss function, and backpropagate to update the parameters of the perception model, and train until convergence to obtain a trained perception model.
[0066] Before using the trained perception model for elevation estimation, use the rotation matrix and displacement vector provided by the IMU data to correct the pose of the binocular image of the scene to be estimated, so as to avoid image distortion caused by camera vibration or tilt. When using the trained perception model for elevation estimation, the feature extraction network extracts visual features from the binocular image of the scene to be estimated, uses the IMU data to map the three-dimensional voxel space in the BEV view to the binocular image, and fuses the voxel grid in the three-dimensional voxel space to the visual features of the binocular image to obtain consistent voxel features. Input the consistent voxel features into the elevation classification network, and output the elevation classification result of the scene.
[0067] BEV elevation estimation refers to the technology of estimating the road surface elevation from a bird's-eye view. This technology is mainly used in autonomous driving and intelligent transportation systems. By estimating the road surface elevation information, it can help vehicles better understand the road surface conditions, thereby improving driving safety and comfort.
[0068] BEV elevation estimation has broad application prospects in practical applications. For example, in autonomous vehicles, by obtaining real-time road surface elevation information, the suspension control system can be optimized to reduce the bumpiness of the vehicle during driving and improve the riding comfort. In addition, this technology can also be used in intelligent traffic management to help traffic management departments better plan and maintain roads and reduce the occurrence of traffic accidents.
[0069] Example 1
[0070] In the present invention, through a front binocular camera, a lidar, and an IMU sensor installed on an unmanned bulldozing mechanical device, terrain data in a complex scene of a mine bulldozing yard is collected in real time. The binocular camera collects binocular images, the lidar collects lidar point cloud data, and the IMU sensor collects IMU data.
[0071] Input the binocular images and IMU data of the mine bulldozing scene to be estimated into the trained perception model, and output the elevation classification result of the scene;
[0072] The perception model includes a feature extraction network and an elevation classification network, and is trained in the following manner:
[0073] Collect the lidar point cloud data, binocular images, and IMU data of the scene, and use the lidar point cloud data to obtain the true value of the elevation classification of the scene;
[0074] Input the binocular images into the feature extraction network to extract visual features. Use the IMU data to map the three-dimensional voxel space in the BEV view to the binocular images, and fuse the voxel grids in the three-dimensional voxel space to the visual features of the binocular images to obtain consistent voxel features. Input the consistent voxel features into the elevation classification network, and use the error between the output elevation classification prediction value and the true value of the elevation classification as the loss function, and backpropagate to update the parameters of the perception model, and train until convergence to obtain the trained perception model.
[0075] Specifically, after collecting multi-source data, to ensure the spatial consistency of the data, data fusion and pixel-level alignment of multiple sensors are required. Specifically, as Figure 2 shown, the present invention uses the rotation matrix and displacement vector obtained by the IMU sensor to correct the pose of the image data captured by the binocular camera to avoid image distortion caused by equipment vibration or tilt. In the figure, the green arrow corresponds to X c , Y c and Z c is the camera coordinate system before pose correction, and the black arrow corresponds to X c ’, Y c ’ and Z c ’ is the camera coordinate system after pose correction. Then, through the pre-calibrated external parameter matrix from the lidar to the camera, the lidar coordinate system X r , Y r and Z rThe three-dimensional point cloud data obtained below is transformed to a position consistent with the camera coordinate system, thus completing the precise alignment of the point cloud data and the binocular images. This processing process can effectively reduce the errors and instabilities in the subsequent feature extraction process. The generation method of the elevation ground truth is specifically to perform ICP (Iterative Closest Point) algorithm registration and fusion on consecutive multi-frame lidar data to generate precise high-density point cloud data, and map the fused point cloud data into a three-dimensional voxel grid through the extrinsic parameter matrix of the camera and the lidar to generate high-precision elevation ground truth for network training.
[0076] As Figure 3 shown, based on data fusion, a three-dimensional voxel space in the BEV view is established in the ground coordinate system according to the actual size of the unmanned bulldozing machinery and equipment, and the three-dimensional voxel grid is divided to achieve a unified representation of the space. Considering the balance between the sensitivity to terrain undulations and the computational complexity in the elevation estimation task, the present invention selects 5 cm as the spatial resolution of the voxel grid. In addition, due to the sparse problem of single-frame lidar data, the present invention uses consecutive multi-frame point cloud data and uses the ICP (Iterative Closest Point) algorithm to register and fuse the multi-frame point clouds to generate higher-density and more accurate elevation ground truth data, providing reliable supervision information for the subsequent training of the network model.
[0077] Next, in order to meet the limitations of the network computing amount of the edge computing device, the present invention constructs a lightweight multi-scale feature extraction network based on the LiteMBConv module for visual feature extraction of the left and right binocular images. In Fig. 4(a), Conv represents the convolutional layer, BN (Batch Normalization) represents batch normalization, Swish is an activation function, Depthwise is used to efficiently extract features and reduce the computational complexity, Project is used for channel compression or alignment, and Stochastic Depth is used to randomly skip some layers (DropPath) during training to improve the generalization ability of the model. DW (Depthwise Convolution) is depthwise separable convolution, Max Pool is max pooling, Expand is to expand the dimension to make the attention weight dimension consistent with the input feature Figure 1 s, and PW is pointwise convolution (PointwiseConvolution). The LiteMBConv module includes a standard convolution and normalization module (Conv&BN), a depth convolutional layer (Depthwise), a lightweight channel attention module (UBSAM), a projection layer (Project), and a residual connection module (Stochastic Depth). In the design of the LiteMBConv module of the present invention,
[0078] Considering the disadvantage that the SE module has a large computational overhead in lightweight networks, the SE channel attention mechanism in the traditional MBConv structure is replaced with a low-rank grouped attention mechanism ULSAM. During the specific implementation, the input feature dimension is first reduced through a max-pooling operation, then channel compression is performed using stationary convolution, and the Softmax activation function is used to calculate the grouped attention weights. This structural optimization significantly reduces the network parameter quantity and computational complexity, improves the network operation efficiency, and is more suitable for real-time applications on edge devices. The input and output of this module are consistent.
[0079] As shown in Figure 4(b), the present invention designs a lightweight multi-scale feature extraction network for binocular images to extract intermediate feature maps with rich structures and stable semantics from the input binocular images. In Figure 4(b), ReLU represents the ReLU activation function, and Feature Map represents the feature map. Taking the left image as an example, its input size is (B, 3, H, W). First, it passes through an initial convolution module containing convolution, batch normalization (BN), and the Swish activation function to increase the number of channels from 3 to 32, and the spatial size is reduced to (H / 2, W / 2). Then, the image features sequentially pass through four backbone stages composed of LiteMBConv modules, namely: the LiteMBConv1 stage, with an output size of (B, 16, H / 2, W / 2); the LiteMBConv2 stage, containing 2 modules, with an output size of (B, 24, H / 4, W / 4); the LiteMBConv3 stage, containing 2 modules, with an output of (B, 40, H / 8, W / 8); the LiteMBConv4 stage, containing 3 modules, with an output of (B, 80, H / 16, W / 16). To effectively fuse multi-scale features, the present invention upsamples the output features of LiteMBConv2 to LiteMBConv4 to (H / 2, W / 2) through bicubic interpolation (BiCubic) and concatenates them with the output features of LiteMBConv1 in the channel dimension to obtain a fused feature map with 160 channels. Subsequently, this fused feature map sequentially passes through two convolution modules containing convolution, BN, and the ReLU activation function for further feature integration. The channel dimension is gradually compressed from 160 to 128 and then to 64, and finally, a unified feature map with a size of (B, 64, H / 2, W / 2) is output through a 1×1 convolution. While maintaining high computational performance, this structure can fully fuse spatial hierarchical information at different scales, improve the feature stability and expression ability of binocular images under disparity changes and texture differences, and is suitable for subsequent 3D understanding and elevation estimation tasks.
[0080] After completing the two-dimensional feature extraction, to accurately map the two-dimensional visual features to three-dimensional spatial positions, a three-dimensional voxel space in the BEV view is established in the ground coordinate system. The three-dimensional voxel space is transformed from the ground coordinate system to the camera coordinate system through the rotation matrix and camera position provided by the IMU data, and then mapped to the image pixel coordinate system using the camera intrinsic parameters, forming a projection index relationship between the voxel grid in the three-dimensional voxel space and the visual features in the binocular image. The visual features corresponding to a voxel grid in the binocular image are extracted using the projection index relationship and multiplied element by element to obtain consistent voxel features. The present invention fuses to obtain consistent voxel features (CVF) through the Hadamard product (element-by-element multiplication). This fusion method can effectively filter out the mismatched features caused by the unilateral view, highlight the stable features jointly confirmed by the left and right views, and improve the accuracy and robustness of elevation estimation.
[0081] After obtaining the consistent voxel features, dimensionality reduction processing is performed on the consistent voxel features, and then the Hourglass module is used to extract and aggregate multi-scale features; the spatial details of the voxel features at different scales are captured through the combination of upsampling and downsampling in the Hourglass structure. In the Hourglass module, a global-local self-attention mechanism is introduced, including local multi-head self-attention (LMSA) and global multi-head self-attention (GMSA). The local features and global features are fused through the learnable weight parameter λ, and the calculation formula for the fusion process is:
[0082] Attention = λ·LMSA(X) + (1 - λ)·GMSA(X)
[0083] where Attention represents the fused attention features, λ represents the learnable weight parameter for adaptively adjusting the fusion ratio of local features and global features, LMSA(X) represents the local features extracted by the local multi-head self-attention from X, GMSA(X) represents the global features extracted by the global multi-head self-attention from X, and X represents the input consistent voxel features.
[0084] To achieve the best local and global information fusion, the present invention uses the learnable parameter λ to dynamically adjust the ratio of local and global attention. This design effectively enhances the model's ability to understand the local changes and global structural features of the terrain, and greatly improves the generalization performance and elevation estimation accuracy of the model.
[0085] In Fig. 5(a), PW (Pointwise Conv) represents point convolution, Local Self-attention represents local self-attention, Global Self-attention represents global self-attention, Weighted Fusion represents weighted fusion, Concatenated Projection represents concatenated projection, Local head = 8 means the number of local attention heads = 8, and Global head = 4 means the number of global attention heads = 4. In Fig. 5(b), Attention block is the attention block, and ConvTrans&BN represents deconvolution + normalization. In Fig. 5(c), Hourglass is the upsampling and downsampling structure, Conv3d represents 3D convolution, and Trilinear Interpolate represents trilinear interpolation.
[0086] To improve the model's comprehensive understanding ability of terrain details and overall morphology, the present invention embeds a hybrid attention mechanism that fuses local multi-head self-attention (LMSA) and global multi-head self-attention (GMSA) in the Hourglass module (as shown in Fig. 5(a)). Specifically, the input voxel feature map (B, 64, 64, 164, 80) first performs LMSA calculation through local window partitioning (the number of heads is set to 8), and at the same time performs GMSA calculation within the entire feature map range (the number of heads is set to 4), and the results of the two are weighted and fused through the learnable weight parameter λ. The fused features remain dimension (B, 64, 64, 164, 80) after concatenated projection. Among them, B represents the batch size, 64 is the number of feature channels, and 64, 164, and 80 respectively represent the dimensions of the voxel in the elevation direction, forward direction, and left-right direction. The present invention uses multiple stacked Hourglass structures as the feature extraction backbone, and each Hourglass realizes the integration of features at different spatial scales through the downsampling-upsampling process, as shown in Fig. 5(b). The introduced attention enhancement module in the middle is used to capture long-range spatial dependence features and improve the context awareness ability without changing the spatial dimension. The input and output dimensions of each Hourglass module are consistent, both being (B, 64, 64, 164, 80), for subsequent elevation classification prediction. As shown in Fig. 5(c), after continuously stacking 3 to 4 Hourglass modules to extract global context features, the output features are compressed in the number of channels through a 3×3×3 3D convolution to generate a preliminary voxel elevation probability volume (B, 1, 64, 164, 80). Subsequently, it is upsampled to the preset 80 elevation classification categories in the elevation dimension (Z dimension) by trilinear interpolation, and the output size is (B, 80, 164, 64), that is, a set of elevation probability distributions is output at each forward × lateral BEV plane position for realizing the refined elevation estimation classification task.
[0087] To achieve efficient classification for the elevation estimation task, the present invention transforms the continuous elevation prediction problem into a classification problem of discrete categories, and uses the cross-entropy loss function to supervise the training of the network. Specifically, the model is optimized through the cross-entropy loss between the voxel classification probability output by the network and the true elevation category label, so as to achieve model convergence quickly and efficiently. Experimental results show that this classification training method not only ensures sufficient elevation accuracy, but also significantly improves the training speed and stability of the network model, and is particularly suitable for real-time application scenarios.
[0088] The loss function is as follows:
[0089]
[0090] where L height represents the total loss during the training process, M(vg) indicates whether the voxel grid vg where the consistent voxel feature is located is valid. If the true elevation classification value of the voxel grid vg where the consistent voxel feature is located is within the elevation range, the voxel grid is valid, M(vg)=1; otherwise, the voxel grid is invalid, M(vg)=0. E(c, vg) represents the true elevation classification value of the voxel grid vg where the consistent voxel feature is located, c is the category number, representing the number of the elevation range (for example, for the elevation range 0cm - 10cm, the category number is 1; for the elevation range 10cm - 20cm, the category number is 2; for the elevation range 20cm - 30cm, the category number is 3), ∑ represents the summation over all valid voxel grids vg and all categories c, N c represents the total number of categories in the elevation classification task, and elepred(·, vg) is the elevation classification prediction value of the voxel grid vg where the consistent voxel feature is located.
[0091] The present invention demonstrates the elevation estimation results under three different scenarios. In the horizontal scenario shown in Fig. 6(a), the original images of the left and right views show that the ground is uniform and has no obvious undulations. The corresponding elevation estimation map shown in Fig. 6(b) is smooth and continuous, and there are almost no obvious deviations in the error map shown in Fig. 6(c), indicating that the method has high accuracy and stability under regular terrains. In the complex undulating terrain shown in Fig. 7(a), obvious height differences can be seen in the original images. The elevation estimation map shown in Fig. 7(b) can capture the terrain changes well, but compared with the flat terrain, the estimation error increases slightly. There are certain fluctuations in the local areas of the error map shown in Fig. 7(c), but the overall error is still controllable, indicating that the method still has strong adaptability under complex terrains. For the terrain with pothole features shown in Fig. 8(a), the ground surface in the original images presents an irregular pothole structure. The elevation estimation map shown in Fig. 8(b) can still roughly reflect the terrain undulations, but the error map shown in Fig. 8(c) shows that the error is relatively large in the areas with deeper potholes, indicating that under extremely irregular terrains, the accuracy of elevation estimation will be affected to a certain extent, but the overall elevation information can still be maintained stably and has certain reference value. In summary, the method of the present invention can provide reliable elevation estimation results under different terrain conditions, and maintain a reasonable error range in complex environments, showing good robustness and practicality.
[0092] It is easy for those skilled in the art to understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A BEV elevation estimation method based on binocular data, characterized in that, Including: Collect binocular images and IMU data of the scene to be estimated, input them into the trained perception model, and output the elevation classification result of the scene; The perception model includes a feature extraction network and an elevation classification network, and is trained in the following way: Collect lidar point cloud data, binocular images and IMU data of the scene, and use the lidar point cloud data to obtain the ground truth of the elevation classification of the scene; Input the binocular image into the feature extraction network to extract visual features. Use the IMU data to map the three-dimensional voxel space in the BEV view to the binocular image, and fuse the voxel grids in the three-dimensional voxel space to the visual features of the binocular image to obtain consistent voxel features. Input the consistent voxel features into the elevation classification network, and use the error between the output elevation classification prediction value and the ground truth of the elevation classification as the loss function, and backpropagate to update the parameters of the perception model until convergence to obtain the trained perception model.
2. The BEV elevation estimation method based on binocular data according to claim 1, wherein, Before using the lidar point cloud data to obtain the ground truth of the elevation classification of the scene, Use the rotation matrix and displacement vector provided by the IMU data to correct the pose of the binocular image. Through the extrinsic parameter matrix from the lidar to the camera, convert the lidar point cloud data in the lidar coordinate system to the camera coordinate system, so that the lidar point cloud data and the binocular image are aligned at the pixel level space.
3. A BEV elevation estimation method based on binocular data according to claim 1 or 2, characterized in that The feature extraction network includes LiteMBConv modules and convolution modules, The LiteMBConv module is of the MBConv structure, and the SE channel attention mechanism in the MBConv structure is replaced by a low-rank grouped attention mechanism. The number of LiteMBConv modules is multiple. Multiple LiteMBConv modules perform multi-scale feature extraction on the binocular image. The features extracted by the subsequent LiteMBConv modules except the first LiteMBConv module are interpolated and upsampled and then concatenated with the features extracted by the first LiteMBConv module to obtain a fused feature map. The convolution module performs convolution on the fused feature map to extract visual features.
4. A BEV elevation estimation method based on binocular data according to claim 1 or 2, characterized in that, The elevation classification network includes: Hourglass modules and convolution modules, The Hourglass module is used to extract global context features through downsampling and upsampling. The consistent voxel features are linearly interpolated after passing through multiple cross-arranged Hourglass modules and convolution modules to obtain the probability distribution of elevation classification. Perform Softmax normalization on the probability distribution of elevation classification to obtain the elevation classification prediction value.
5. The BEV elevation estimation method based on binocular data according to claim 4, wherein A hybrid attention mechanism is embedded in the Hourglass module. The hybrid attention mechanism fuses local multi-head self-attention and global multi-head self-attention, and fuses local features and global features through the learnable weight parameter λ. The calculation formula for the fusion process is: Attention = λ·LMSA(X) + (1 - λ)·GMSA(X) Among them, Attention represents the fused attention feature, λ represents the learnable weight parameter used to adaptively adjust the fusion ratio of local features and global features, LMSA(X) represents the local features extracted from X by local multi-head self-attention, GMSA(X) represents the global features extracted from X by global multi-head self-attention, and X represents the input consistent voxel feature.
6. The BEV elevation estimation method based on binocular data according to claim 1 or 2, characterized in that, The consistent voxel feature is obtained in the following way: A three-dimensional voxel space in the BEV view is established in the ground coordinate system. The three-dimensional voxel space is transformed from the ground coordinate system to the camera coordinate system through the rotation matrix and camera position provided by the IMU data. The three-dimensional voxel space is mapped to the image pixel coordinate system using the camera intrinsic parameters, forming a projection index relationship between the voxel grid in the three-dimensional voxel space and the visual features in the binocular image. The visual features corresponding to a voxel grid in the binocular image are extracted using the projection index relationship and multiplied element by element to obtain the consistent voxel feature.
7. A BEV elevation estimation method based on binocular data according to claim 1 or 2, characterized in that The loss function is: Among them, L height represents the total loss during the training process. M(vg) indicates whether the voxel grid vg where the consistent voxel feature is located is valid. If the ground truth of the elevation classification of the voxel grid vg where the consistent voxel feature is located is within the elevation range, the voxel grid is valid, M(vg)=1; otherwise, the voxel grid is invalid, M(vg)=0. E(c, vg) represents the ground truth of the elevation classification of the voxel grid vg where the consistent voxel feature is located, c is the class number, representing the number of the elevation range, and ∑ represents the summation over all valid voxel grids vg and all classes c. N c represents the total number of classes in the elevation classification task, and elepred(·, vg) is the predicted value of the elevation classification of the voxel grid vg where the consistent voxel feature is located.
8. A method for BEV elevation estimation based on binocular data according to claim 1 or 2, characterized in that The ground truth of elevation classification is calculated in the following way: The iterative closest point is used to register and fuse consecutive frames of lidar point cloud data. The fused lidar point cloud data is mapped into a three-dimensional voxel grid through the extrinsic matrix from lidar to camera to generate the ground truth of elevation classification.
9. A BEV elevation estimation system based on binocular data, characterized in that, It includes: A preprocessing module for collecting lidar point cloud data, binocular images, and IMU data of the scene, and obtaining the ground truth of elevation classification of the scene using the lidar point cloud data; A training module for inputting the binocular image into a feature extraction network to extract visual features, using the IMU data to map the three-dimensional voxel space in the BEV view to the binocular image, fusing the visual features of the voxel grid in the three-dimensional voxel space mapped to the binocular image to obtain the consistent voxel feature, inputting the consistent voxel feature into an elevation classification network, taking the error between the output elevation classification prediction value and the ground truth of elevation classification as the loss function, and backpropagating to update the parameters of the perception model until convergence to obtain a trained perception model; An elevation estimation module for collecting binocular images and IMU data of the scene to be estimated, inputting them into the trained perception model, and outputting the elevation classification result of the scene.
10. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for BEV elevation estimation based on binocular data according to any one of claims 1 to 8.
Citation Information
Cited By
Unmanned aerial vehicle control system based on monocular vision
CN121297815A
Monocular vision-based unmanned aerial vehicle control system
CN121297815B
Vehicle control method and device, electronic equipment and storage medium
CN121469550A
Vehicle control method and device, electronic equipment and storage medium
CN121469550B