Multi-mode drivable occupancy prediction method for unstructured road environment

By combining multimodal fusion method of lidar and image data in an unstructured road environment, three-dimensional semantic scene completion and driving cost prediction are solved, and the problem of difficult to perceive and estimate terrain feasibility in an unstructured environment is achieved, achieving higher precision driving prediction and more robust path planning.

CN120071285APending Publication Date: 2025-05-30SHENZHEN AUTOMOTIVE RES INST BEIJING INST OF TECH (SHENZHEN RES INST OF NAT ENG LAB FOR ELECTRIC VEHICLES) +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510151350.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In unstructured road environments, it is difficult for the prior art to effectively perform precise perception and maneuverability estimation of the unstructured terrain ahead, especially in the presence of complex obstacles such as overhanging obstacles and negative obstacles.

Method used

By combining the geometric information of the lidar point cloud and the semantic characteristics of the monocular image, a multimodal fusion method is used to complete three-dimensional semantic scenes to generate fine-grained driving cost-occupation predictions.

Benefits of technology

It significantly improves the accuracy of driving prediction under complex terrain, can effectively characterize complex obstacles in unstructured environments, and improves the robustness of vehicle path planning and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071285A_ABST
    Figure CN120071285A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode drivable occupancy prediction method for an unstructured road environment, and belongs to the technical field of automatic driving. Comprising the following steps: step 1, processing occupation prediction model data for the cross-country environment; constructing a data set with a three-dimensional passable occupation mark based on any data set with a semantic laser radar point cloud mark so as to train a multi-modal drivable occupation prediction model; step 2, a multi-mode drivable area occupation prediction model ORDform is constructed; a multi-mode drivable area occupation prediction model designed for an unstructured road environment is adopted; a dense semantic occupancy prediction is generated from a forward perspective using LiDAR point cloud and monocular images. According to the method, complex obstacles in an unstructured environment can be represented, the drivable cost in a fine-grained environment can be predicted, the path planning and decision-making robustness of the vehicle is remarkably improved, and the vehicle is prevented from entering a high-risk area blindly.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention provides a multi-modal drivable occupancy prediction method for unstructured road environments, belonging to the technical field of autonomous driving. Background Art

[0002] With the development of autonomous driving technology, the demand for vehicles to autonomously drive in unstructured environments (such as mountainous areas, grasslands, hilly areas, forest rescue scenarios, mining transportation sections) has become increasingly urgent. However, different from the clear lane lines, regular road structures, and rich environmental markings in urban roads, unstructured scenarios, especially off-road scenarios, generally lack stable road features, have large terrain undulations, uneven surfaces, and complex environmental factors such as vegetation occlusion, rock accumulation, potholes, and slope changes. These factors make it difficult to directly apply autonomous driving perception and planning technologies based on traditional road markings or regular environments in off-road scenarios.

[0003] Currently, common autonomous driving perception algorithms are mostly based on structured scenarios (such as urban roads), using lidar point clouds and camera images for obstacle detection, target tracking, and map construction. However, these algorithms focus more on identifying clearly separated targets (such as vehicles, pedestrians, traffic signs) and flat road surfaces, and have limited analysis of the drivability of the terrain itself. At the same time, existing public datasets (such as the KITTI dataset, nuScenes dataset, etc.) mainly focus on urban and suburban environments, and there is a lack of three-dimensional drivable annotations for unstructured scenarios.

[0004] In order to enable unmanned vehicles to achieve safe and efficient autonomous driving in off-road environments, there is an urgent need for a method that can accurately perceive and estimate the drivability of the unstructured terrain ahead, that is, to judge whether each spatial position (voxel) can be safely driven through by the vehicle and give fine-grained difficulty or cost annotations (such as fatally non-drivable, high cost, low cost, and completely unobstructed). In addition, compared with relying solely on the geometric information of lidar or the semantic information of images, the multi-modal fusion of the two to improve the understanding accuracy of complex terrain features is also the focus and difficulty of the current research direction.

[0005] The BEVNet model proposed by Amirreza Shaban et al. (Shaban A, Meng X, Lee J H, et al. Semantic terrain classification for off-road autonomous driving[C] / / Conference on Robot Learning. PMLR, 2022: 619-629.) addresses the problem of feasible region estimation in unstructured environments. It uses single-frame lidar point clouds as input, converts it into a semantic scene completion (SSC) problem, and generates a 2D cost map that reflects terrain feasibility. The input point cloud is first voxelized to obtain a 3D grid input tensor. Then, a series of 3D sparse convolutional layers are used to extract features, and skip connections are introduced in the encoder-decoder structure to fuse fine-grained and high-level semantic features. The network finally outputs a 3D tensor with semantic class information.

[0006] This technology has the following disadvantages:

[0007] 1. The feasible region estimation result of this scheme is a 2D passable grid, and this representation form cannot represent common obstacles in unstructured environments such as overhanging obstacles and negative obstacles;

[0008] 2. This scheme only uses lidar point clouds to learn the passability of the scene, without using the semantic information of cameras. The classification of the passable area of the scene is rough, and the generalization ability is poor, and it cannot solve the problem of feasible region estimation in complex scenes.

[0009] Zhou Mengru et al. proposed a road passability analysis method in the article (Zhou Mengru, Chen Huiyan, Xiong Guangming, et al. Road passability analysis of unmanned tracked platforms in off-road environments[J]. Acta Armamentarii, 2022, 43(10): 2485-2496.). Based on the image semantic segmentation method, the road surface attributes are obtained and the passable area is roughly divided, and the 3D point cloud features are extracted to describe the ground geometric information. On the basis of combining the geometric passability constraints of the unmanned tracked platform, Gaussian mixture clustering is used to classify the passability of the road, and a 3D passable grid map is used to describe the passable degree of the environment.

[0010] This technology has the following disadvantages:

[0011] 1. This scheme requires multiple independent steps to process and integrate the data before finally outputting the feasible region prediction result.

[0012] 2. The semantic passability and geometric passability analysis processes are independent of each other. The image and point cloud data are only fused at the result layer, lacking deep fusion at the feature level. This decoupled fusion method limits the accuracy and robustness of the overall feasible region estimation. Summary of the Invention

[0013] The present invention provides a multi-modal drivable occupancy prediction method for unstructured road environments. The technical problems to be solved are as follows:

[0014] 1. Based on vehicle dynamics characteristics, terrain parameters (such as step height, slope, unevenness), and environmental semantic features, assign corresponding drivable cost levels (such as fatal obstacles, high-cost passage, low-cost passage, unobstructed passage) to each voxel, thereby generating training data for training the drivable area occupancy prediction model.

[0015] 2. Achieve three-dimensional representation of complex obstacles such as overhanging obstacles and negative obstacles in unstructured environments, especially off-road scenarios. At the same time, perform three-dimensional semantic scene completion on the front environment to achieve drivable cost prediction for the front road environment.

[0016] 3. The model effectively fuses the high-precision geometric information of lidar point clouds with the rich semantic features of monocular images to improve the drivable prediction accuracy in complex terrains (including scenarios such as vegetation occlusion, potholes, slope changes, rock stacks, etc.).

[0017] The complete technical solution provided by the present invention:

[0018] A multi-modal drivable occupancy prediction method for unstructured road environments, including two steps:

[0019] Step 1. Data processing of the occupancy prediction model for off-road environments:

[0020] Based on any dataset with semantic lidar point cloud annotations, construct a dataset with three-dimensional passable occupancy annotations, thereby training the multi-modal drivable occupancy prediction model.

[0021] Step 1.1: Dense semantic occupancy annotation;

[0022] Networks trained with sparse LiDAR point clouds are difficult to predict sufficiently dense occupancy information. Use the point clouds with semantic annotations in the RELLIS-3D dataset to generate dense annotations within the forward field of view. Convert multiple frames of LiDAR point data to a unified coordinate system and perform voxelization, including dynamic objects, so as to retain the dense point distribution in the grid. At the same time, the scanning of surface points by LiDAR enables this process to capture surface occupancy information.

[0023] Generate the occupancy ground truth within the range of [38.4m, 51.2m, 8m], and set the prediction range to [0m, 38.4m] along the X-axis, [-25.6m, 25.6m] along the Y-axis, and [-2m, 6m] along the Z-axis. The ground truth data is voxelized into a 192×256×40 grid with a voxel size of 0.2m, and the data is filtered to retain the voxels within the camera's field of view. Subsequently, the original categories are remapped to evaluate their impact on traversability.

[0024] Step 1.2: Calculate geometric feature parameters;

[0025] Construct a 3D occupancy voxel by fusing the registered dense point cloud. Each voxel G j ={x j , y j , z j} is centered at (x j , y j ) in the ground coordinate system with a height of z j . The voxel neighborhood Ω is defined as the area centered at G j , and the distribution of points within this area is approximately circular. Assign the points in the registered dense point cloud to each voxel, and all the points within the voxel form a matrix g:

[0026]

[0027] Step height h: Calculate the maximum elevation difference between the central voxel G and other voxels Gj within its neighborhood Ω to evaluate the terrain height change around the central voxel:

[0028]

[0029] where G z and are the elevations of the central voxel G and the neighboring voxel G j respectively.

[0030] Slope s: The slope is calculated from the elevation values within the neighborhood. For each voxel G, fit the center points of all voxels within its neighborhood Ω into a plane. The angle between the normal vector n of the plane and the unit vector Z=(0, 0, 1) of the Z-axis in the coordinate system is defined as the slope:

[0031]

[0032] where n is obtained by calculating the covariance matrix through principal component analysis (PCA) of the neighborhood points.

[0033] Roughness u: The roughness is estimated by calculating the logarithm of the mean square error MSE between the actual elevation of each voxel within the neighborhood and the fitted plane:

[0034]

[0035] Among them, α 0 and a 1 are the slope coefficients of the plane in the x and y directions respectively, c is the intercept, and z actual is the actual elevation of the current voxel.

[0036] Step 1.3: Analysis of vehicle obstacle crossing conditions;

[0037] Analyze the failure conditions for the vehicle to cross obstacles according to the principles of vehicle ground mechanics, and finally form StepMask.

[0038] Step Mask: The failure conditions for the vehicle to cross obstacles, analyzing four situations: the vehicle crossing vertical obstacles, ditches, overhanging obstacles, and longitudinal slopes;

[0039] Vehicle vertical obstacle crossing condition: For a 4×4 wheeled vehicle, evaluating its performance in crossing a stepped vertical protrusion obstacle is based on the maximum vertical obstacle height it can overcome. The dimensionless expression for the maximum crossable obstacle height is:

[0040]

[0041] Among them, h is the obstacle height, r is the wheel radius, μ is the friction coefficient, l is the vehicle wheelbase, and a is the horizontal distance from the front axle to the vehicle center of gravity.

[0042] Vehicle ditch crossing condition: The ditch width that the vehicle can cross is represented by the ratio of the ditch width l d to the wheel diameter D . The ratio of the crossable vertical obstacle height to the wheel diameter is converted to Given a specific obstacle height h, the corresponding ditch width ratio can be determined

[0043]

[0044] Vehicle overhanging obstacle crossing condition: For the vehicle to cross an overhanging obstacle, the height h obj of the lowest point of the obstacle from the ground must be greater than the sum of the height h pc of the obstacle point cloud relative to the LiDAR sensor and the height h Lid of the LiDAR from the ground. When the vehicle encounters an overhanging obstacle, it is necessary to satisfy:

[0045] h obj >h pc +h Lid

[0046] Vehicle longitudinal slope crossing condition: For longitudinal slope crossing, the platform shall satisfy:

[0047] θ max >α

[0048] where θ max is the maximum climbing angle that the vehicle can handle, and α is the current terrain slope angle.

[0049] Step 1.4: Generate passable cost occupancy annotations;

[0050] Perform an initial mapping of voxel semantic categories and passable costs according to the rules, and adjust based on the Step Mask obtained in Step 1.3. The Step Mask is obtained from the vehicle's geometric passability analysis, and voxels that do not meet the geometric passability are assigned a fatal non-passable cost. After the above steps, four passable cost labels are generated: fatal non-passable, high cost, medium cost, and directly passable, representing the passability from difficult to easy in sequence, and thus can be used to train the multi-modal drivable occupancy prediction model.

[0051] Step 2. Construct the multi-modal drivable area occupancy prediction model ORDformer:

[0052] Adopt a multi-modal drivable area occupancy prediction model designed for unstructured road environments. Use LiDAR point cloud and monocular image to generate dense semantic occupancy predictions from the forward view.

[0053] First, extract features from the two modalities: LiDAR data provides accurate geometric information, while image data provides rich semantic details. These features are then fused through a deformable attention mechanism to effectively integrate spatial and semantic information. Next, the fused features are processed through a deformable self-attention layer to enhance the 3D occupancy representation. Finally, the model generates a dense occupancy map to highlight the passable areas.

[0054] Step 2.1: Feature extraction;

[0055] Extract environmental features using multi-modal information. The model uses a ResNet-50 backbone network to extract image features from the monocular image It where l×b represents the spatial resolution of the feature map and d represents the feature dimension. At the same time, the sparse point cloud L from LiDAR t is voxelized to generate a sparse voxel occupancy supervision S spa ∈{0, 1} H×W×Z . Subsequently, use the query upsampling network to obtain the occupancy supervision S with a lower spatial resolution den ∈{0, 1} H / 2×W / 2×Z / 2, thus improving the robustness of the query. Among them, the Mask Token module is used to represent the empty voxel regions in the environment. When combined with voxel occupancy supervision, it forms a more comprehensive 3D voxel feature, which helps to improve the occupancy prediction performance.

[0056] Step 2.2: Feature fusion;

[0057] In the model, the deformable attention (DA) mechanism is adopted to achieve the interaction of local regions of interest in the 3D and 2D feature spaces. Deformable attention dynamically samples N s points around the reference point, adaptively learns the attention weights according to the query vector, and simultaneously predicts the spatial offsets of the sampled points relative to the reference point, thereby achieving flexible and efficient feature aggregation. Based on these sampled points, the attention result is calculated by the following formula:

[0058]

[0059] where, q is the query vector, p is the reference point corresponding to q, F is the input feature map, is the learnable weight matrix for each sampled point, used to generate the value vector value, A s ∈[0, 1] is the attention weight for each sampled point, adaptively learned according to the query, is the offset relative to the reference point p, F(p + δp s ) represents the feature vector at the position p + δp s extracted from the input feature by bilinear interpolation.

[0060] The deformable cross-attention (DCA) is used to enhance the interaction between the occupancy supervision query S den and the image features . Multilayer DCA is adopted to organically fuse multimodal information. First, the 3D grid is projected into the 2D image feature space The projected 2D points are used as the reference points of S den , and the surrounding features are sampled from them. The weighted sum of the sampled features is obtained as the output of the DCA layer, thereby obtaining a refined query that fuses the environmental semantics and geometric features

[0061]

[0062] For the query s den at the position p = (x, y, z), the query is mapped to the corresponding reference point on the image I through the camera projection function t .

[0063] The deformable self-attention (DSA) refines the query Concatenate with the mask token to form the initial 3D voxel feature F 3D ∈ H / 2×W / 2×Z / 2×d, and then obtain the enhanced voxel feature through multiple layers of DSA The DSA layer enables the model to learn the spatial 3D occupancy feature more efficiently and accurately. In the formula, the query q p is the Mask Token or the refined query at position p = (x, y, z):

[0064] DSA(I 3D , F 3D ) = DA(q p , p, F 3D )

[0065] Step 2.3: Output the 3D traversable cost occupancy prediction;

[0066] The enhanced voxel feature passes through the fully connected layer FC and through the softmax function to output the dense occupancy prediction Y t ∈ {c 0 , c 1 , …, c M}, where each voxel is either empty c H×W×Z , or a specific traversability category {c 0 , c 1 , …, c 2 , …, c M}. After mapping these predictions back to the 3D grid, a detailed traversable occupancy map is obtained, providing crucial annotation information for autonomous navigation in unstructured environments.

[0067] Step 2.4: Training loss;

[0068] The training of ORDformer is carried out in an end-to-end manner and minimizes three main loss components: semantic scale loss geometric scale loss and weighted cross-entropy loss The total training loss is defined as the sum of these losses:

[0069]

[0070] The weighted cross-entropy loss is calculated as follows:

[0071]

[0072] where Ω represents the set of valid voxels, M is the number of classes, tar i represents the true class of voxel i, is the weight of the true class where voxel i is located, yi,tar is the logit output of the model for the true category tar of voxel i i , and y i,c is the logit output of the model predicting voxel i as category c.

[0073] In the context of semantic scene completion, the precision P c , recall R c and specificity S c are used to evaluate the model performance and measure the ability of the model to correctly classify occupied and unoccupied voxels. The metrics are defined as follows:

[0074]

[0075] where p i,c represents the probability that the predicted voxel i belongs to category c, is the Iverson bracket.

[0076] To improve generality, the scale loss maximizes the above per-class metrics:

[0077]

[0078] According to the above formula, the semantic loss and geometric loss are calculated respectively, where tar sem , tar geo represent the true labels of semantics and geometry respectively, and p sem and p geo are the corresponding model prediction values.

[0079] Beneficial effects brought by the technical solution of the present invention:

[0080] The present invention fully fuses lidar and camera information at the feature level, greatly improving the three-dimensional semantic completion and occupancy prediction accuracy in complex scenes. The present invention can represent complex obstacles in an unstructured environment, such as overhanging obstacles and negative obstacles. And it has been verified by experiments that the intersection over union (IoU) index of scene completion has increased by more than 20%.

[0081] By analyzing terrain parameters and vehicle dynamics characteristics, the present invention can predict the traversable cost of a fine-grained environment, significantly improving the path planning and decision-making robustness of the vehicle and avoiding blindly entering high-risk areas.

[0082] The dataset generation method of the present invention provides a data basis for relevant research for annotating traversable costs, so as to be used for training a feasible region recognition algorithm based on semantic scene completion. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 Flow chart for generating densely occupied truth values of the present invention;

[0084] Figure 2 Scene diagram of a vehicle crossing an obstacle of the present invention;

[0085] Figure 3 Schematic diagram of the ORDformer model structure of the present invention. Detailed implementation manners

[0086] The specific technical solutions of the present invention are described in conjunction with embodiments.

[0087] A multi-modal drivable occupancy prediction method for unstructured road environments provided by the present invention includes two steps:

[0088] Step 1: Data processing of the occupancy prediction model for off-road environments;

[0089] Step 2: Construct a multi-modal drivable area occupancy prediction model (ORDformer).

[0090] The goal of the present invention is to use monocular images and LiDAR point clouds as inputs to predict a dense semantic scene within the field of view of the camera in front of the platform, so as to achieve accurate and reliable passability estimation in off-road environments. Specifically, the image I t and the LiDAR point cloud L t are used as inputs, and the output is a voxel grid Y t ∈{c 0 , c 1 , …, c M} H×W×Z , where each voxel is either empty c 0 , or is classified into one of the specific passability categories {c 1 , c 2 , …, c M}. In this context, M represents the total number of passability categories, and H, W, and Z represent the length, width, and height of the voxel grid, respectively.

[0091] Specifically:

[0092] Step 1. Data processing of the occupancy prediction model for off-road environments:

[0093] This method can construct a dataset with three-dimensional drivable occupancy annotations based on any dataset with semantic LiDAR point cloud annotations, so as to train a multi-modal drivable occupancy prediction model.

[0094] Step 1.1: Dense semantic occupancy annotation;

[0095] Networks trained with sparse LiDAR point clouds struggle to predict sufficiently dense occupancy information, highlighting the necessity of dense occupancy annotation. To this end, the present invention uses the point clouds with semantic annotations in the RELLIS-3D dataset to generate dense annotations within the forward field of view, such as Figure 1 as shown. This method converts multi-frame LiDAR point data to a unified coordinate system and voxelizes it, including dynamic objects, thus retaining the dense point distribution in the grid. At the same time, the scanning of surface points by LiDAR enables this process to capture the ground occupancy information.

[0096] Due to the complexity of obstacles and occlusions in unstructured environments, the present invention generates occupancy ground truth within the range of [38.4m, 51.2m, 8m], while the prediction range is set to [0m, 38.4m] along the X-axis, [-25.6m, 25.6m] along the Y-axis, and [-2m, 6m] along the Z-axis. The ground truth data is voxelized into a grid of 192×256×40 with a voxel size of 0.2m, and the data is filtered to retain the voxels within the camera's field of view. Subsequently, the original categories are remapped to evaluate their impact on traversability.

[0097] Step 1.2: Calculate geometric feature parameters;

[0098] Construct a three-dimensional occupancy voxel by fusing the registered dense point clouds. Each voxel G j ={x j , y j , z j} is centered at (x j , y j ) in the ground coordinate system with a height of z j . The voxel neighborhood Ω is defined as the area centered at G j , and the point distribution within this area is approximately circular. The points in the registered dense point cloud are assigned to each voxel, and all the points within the voxel form a matrix g:

[0099]

[0100] Step height h: Calculate the maximum elevation difference between the central voxel G and other voxels G j within its neighborhood Ω to evaluate the terrain height change around the central voxel:

[0101]

[0102] where G z and are the elevations of the central voxel G and the neighboring voxel G j respectively.

[0103] Slope s: The slope is calculated from the elevation values within the neighborhood. For each voxel G, the center points of all voxels within its neighborhood Ω are fitted to a plane. The angle between the normal vector n of the plane and the unit vector Z=(0, 0, 1) of the Z-axis in the coordinate system is defined as the slope:

[0104]

[0105] where n is obtained by calculating the covariance matrix through principal component analysis (PCA) of the neighborhood points.

[0106] Roughness u: The roughness is estimated by calculating the logarithmic form of the mean square error (MSE) between the actual elevation of each voxel within the neighborhood and the fitted plane:

[0107]

[0108] where a 0 and a 1 are the slope coefficients of the plane in the x and y directions respectively, c is the intercept, and z actual is the actual elevation of the current voxel.

[0109] Step 1.3: Analysis of vehicle obstacle crossing conditions;

[0110] In an unstructured road environment, the passable area of a vehicle is closely related to its geometric passability. The difficulty of terrain passage is related to various road geometric features. The present invention analyzes the failure conditions for a vehicle to cross an obstacle based on vehicle ground mechanics principles and finally forms a Step Mask.

[0111] Step Mask: The failure conditions for a vehicle to cross an obstacle include vehicle suspension, front-end collision, vertical obstacle, ditch, and rollover, etc. Since the extreme difficult scenarios are not involved in the dataset, here we focus on analyzing four situations: a vehicle crossing a vertical obstacle, a ditch, a suspended obstacle, and a longitudinal slope, as Figure 2 .

[0112] Vehicle vertical obstacle crossing condition. For a 4×4 wheeled vehicle, evaluating its performance in crossing a stepped vertical protrusion obstacle is mainly based on the maximum vertical obstacle height it can overcome. The dimensionless expression for the maximum passable obstacle height is:

[0113]

[0114] where h is the obstacle height, r is the wheel radius, μ is the friction coefficient, l is the vehicle wheelbase, and a is the horizontal distance from the front axle to the vehicle center of gravity.

[0115] Vehicle ditch crossing condition. The ditch width that a vehicle can cross is based on the ratio of the ditch width l d to the wheel diameter D Representation. The ratio of the height of the vertically traversable obstacle to the wheel diameter can be converted to as shown in the formula. Given a specific obstacle height h, the corresponding ditch width ratio can be determined

[0116]

[0117] Vehicle overhanging obstacle condition. For a vehicle to cross an overhanging obstacle, the height h of the lowest point of the obstacle from the ground obj must be greater than the height h of the obstacle point cloud relative to the LiDAR sensor pc and the height h of the LiDAR from the ground Lid sum. When the vehicle encounters an overhanging obstacle, it needs to satisfy:

[0118] h obj >h pc +h Lid

[0119] Vehicle longitudinal slope crossing condition. For longitudinal slope crossing, the platform needs to satisfy:

[0120] θ max >α

[0121] where θ max is the maximum climbing angle that the vehicle can handle, and α is the current terrain slope angle.

[0122] Step 1.4: Generate the traversable cost occupancy annotation;

[0123] In the present invention, the initial mapping of the voxel semantic category and the traversable cost is carried out according to the rules, and adjusted based on the Step Mask obtained in Step 1.3. The Step Mask is obtained from the vehicle geometric traversability analysis, and the voxels that do not meet the geometric traversability are assigned a fatal non-traversable cost. After the above steps, four traversable cost labels can be generated: fatal non-traversable, higher cost, medium cost, and directly traversable, representing the traversability from difficult to easy in sequence, and thus can be used to train the multi-modal drivable occupancy prediction model.

[0124] Step 2. Construct a multi-modal drivable area occupancy prediction model (ORDformer):

[0125] In an unstructured road environment, the lack of structured road and traffic control information makes the passability of the terrain closely related to the geometric constraints of the platform, and detailed analysis of ground geometric passability is required. Therefore, higher requirements are imposed on the perception algorithm, which needs to have advanced three-dimensional occupancy prediction capabilities. To address these challenges, the present invention proposes a multi-modal drivable area occupancy prediction model designed for unstructured road environments. This method uses LiDAR point clouds and monocular images to generate dense semantic occupancy predictions from the forward perspective. The overall process of ORDformer is as Figure 3 shown.

[0126] The model first extracts features from two modalities: LiDAR data provides precise geometric information, while image data provides rich semantic details. These features are then fused through a deformable attention mechanism to effectively integrate spatial and semantic information. Next, the fused features are processed through a deformable self-attention layer to enhance the three-dimensional occupancy representation. Finally, the model generates a dense occupancy map, thereby highlighting the passable areas.

[0127] Step 2.1: Feature extraction;

[0128] To achieve more accurate occupancy estimation, the present invention utilizes multi-modal information to extract environmental features. The model employs a ResNet-50 backbone network to extract image features from a monocular image I t from which l×b represents the spatial resolution of the feature map and d represents the feature dimension. At the same time, the sparse point cloud L from the LiDAR is voxelized to generate a sparse voxel occupancy supervision S t ∈ {0, 1} spa H×W×Z . Subsequently, a query upscale network is used to obtain a lower spatial resolution occupancy supervision S den ∈ {0, 1} H / 2×W / 2×Z / 2 to improve the robustness of the query. Among them, the Mask Token module is used to represent the empty voxel regions in the environment. When combined with the voxel occupancy supervision, this method forms a more comprehensive three-dimensional voxel feature, which helps to improve the occupancy prediction performance.

[0129] Step 2.2: Feature fusion;

[0130] The present invention adopts a Deformable Attention (DA) mechanism in the model to achieve the interaction of local regions of interest in the three-dimensional and two-dimensional feature spaces. Deformable attention dynamically samples N around the reference point s ​points, adaptively learn attention weights according to the query vector, and at the same time predict the spatial offset of the sampled points relative to the reference point, so as to achieve flexible and efficient feature aggregation. Based on these sampled points, the attention result is calculated by the following formula:

[0131]

[0132] where q is the query vector, p is the reference point corresponding to q, F is the input feature map, is the learnable weight matrix for each sampled point, used to generate the value vector, A s ∈[0, 1] is the attention weight for each sampled point, adaptively learned according to the query, is the offset relative to the reference point p, F(p + δp s ) represents the feature vector at the position p + δP s extracted from the input feature by bilinear interpolation.

[0133] Deformable Cross Attention (DCA). To enhance the interaction between the occupancy supervision query S den and the image features of the present invention adopts multi-layer DCA to organically fuse multi-modal information. First, project the 3D grid into the 2D image feature space The projected 2D points are used as the reference points of S den , and sample the surrounding features therefrom. The weighted sum of the sampled features is obtained as the output of the DCA layer, so as to obtain a refined query that fuses environmental semantics and geometric features

[0134]

[0135] For the query s den at the position p = (x, y, z), map the query to the corresponding reference point on the image I through the camera projection function t .

[0136] Deformable Self-Attention (DSA). Concatenate the refined query with the mask token to form the initial 3D voxel feature F 3D ∈H / 2×W / 2×Z / 2×d, and then obtain the enhanced voxel feature through multi-layer DSA The DSA layer enables the model to more efficiently and accurately learn the spatial 3D occupancy features. In the formula, the query q p can be a Mask Token or a refined query at the position p = (x, y, z):

[0137] DSA(I 3D , F3D ) = DA(q p , p, F 3D )

[0138] Step 2.3: Output the 3D traversable cost occupancy prediction;

[0139] The enhanced voxel features Pass through the fully connected layer (FC) and output the dense occupancy prediction Y t ∈ {c 0 , c 1 , …, c M} H×W×Z , where each voxel is either empty c 0 , or a specific traversability class {c 1 , c 2 , …, c M}. After mapping these predictions back to the 3D grid, a detailed traversable occupancy map can be obtained, providing crucial annotation information for autonomous navigation in unstructured environments.

[0140] Step 2.4: Training loss;

[0141] The training of ORDformer is carried out in an end-to-end manner and minimizes three main loss components: the semantic scale loss the geometric scale loss and the weighted cross-entropy loss The total training loss is defined as the sum of these losses:

[0142]

[0143] The weighted cross-entropy loss is calculated as follows:

[0144]

[0145] where Ω represents the set of valid voxels, M is the number of classes, tar i represents the true class of voxel i, is the weight of the true class where voxel i is located, y i,tar is the logit output of the model for the true class tar i of voxel i, and y i,c is the logit output of the model predicting voxel i as class c.

[0146] In the context of semantic scene completion, the precision P c , recall R c and specificity S cThe metric evaluates the model performance and is used to measure the ability of the model to correctly classify occupied and unoccupied voxels. The metric is defined as follows:

[0147]

[0148] where p i,c represents the probability that the predicted voxel i belongs to class c, is the Iverson bracket.

[0149] For generality, the scale loss maximizes the above per-class metric:

[0150]

[0151] According to the above formula, the semantic loss and the geometric loss can be calculated respectively. Where tar sem and tar geo represent the semantic and geometric ground truths respectively, and p sem and p geo are the corresponding model predictions.

Claims

1. A multimodal drivable occupancy prediction method for unstructured road environments, characterized in that: There are two steps: Step 1: Data processing of occupancy prediction model for off-road environment; Based on any dataset with semantic lidar point cloud annotations, a dataset with 3D drivable occupancy annotations is constructed to train a multimodal drivable occupancy prediction model. Step 2: Construct a multimodal drivable area occupancy prediction model ORDformer; A multimodal drivable area occupancy prediction model designed for unstructured road environments is used. Dense semantic occupancy predictions are generated from the forward perspective using LiDAR point clouds and monocular images. Features are first extracted from both modalities: LiDAR data provides precise geometric information, while image data provides rich semantic details. These features are then fused through a deformable attention mechanism to effectively integrate spatial and semantic information; the fused features are then processed through a deformable self-attention layer to enhance the 3D occupancy representation; finally, the model generates a dense occupancy map that highlights traversable areas.

2. The multimodal drivable occupancy prediction method for unstructured road environments according to claim 1, characterized in that: Step 1 specifically includes the following sub-steps: Step 1.1: Dense semantic occupancy annotation; The network trained with sparse LiDAR point clouds has difficulty predicting sufficiently dense occupancy information. The semantically annotated point cloud in the RELLIS-3D dataset is used to generate dense annotations in the forward field of view. Multi-frame LiDAR point data is converted to a unified coordinate system and voxelized, including dynamic objects, so that the dense point distribution is retained in the grid. At the same time, LiDAR scanning of surface points allows this process to capture surface occupancy information. The true occupancy value is generated within the range of [38.4m, 51.2m, 8m], and the prediction range is set to [0m, 38.4m] along the X axis, [-25.6m, 25.6m] along the Y axis, and [-2m, 6m] along the Z axis; The ground truth data was voxelized into a 192×256×40 grid with a voxel size of 0.2m, and the data was filtered to retain voxels within the camera field of view; subsequently, the original categories were remapped to assess their impact on traversability; Step 1.2: Calculate geometric feature parameters; The 3D occupied voxels are constructed by fusing the registered dense point clouds; each voxel G j ={x j ,y j , z j }In the ground coordinate system, (x j ,y j ) is the center and the height is z j ; The voxel neighborhood Ω is defined as G j The area centered in this area is approximately circular. The points in the registered dense point cloud are assigned to each voxel, and all the points in the voxel form a matrix g: Step height h: Calculate the central voxel G and other voxels G in its neighborhood Ω j The maximum elevation difference is used to evaluate the change in terrain height around the central voxel: Among them, G z and are the central voxel G and the neighboring voxel G respectively. j 's elevation; Slope s: The slope is calculated by the elevation value in the neighborhood. For each voxel G, all the voxel centers in its neighborhood Ω are fitted into a plane. The angle between the plane's normal vector n and the coordinate system's Z-axis unit vector Z = (0, 0, 1) is defined as the slope: Where n is obtained by calculating the covariance matrix through principal component analysis (PCA) of the neighborhood points; Roughness u: Roughness is estimated by calculating the logarithmic form of the mean square error (MSE) between the actual elevation of each voxel in the neighborhood and the fitted plane: Among them, a0 and a1 are the slope coefficients of the plane in the x and y directions respectively, c is the intercept, and z actual is the actual elevation of the current voxel; Step 1.3: Analysis of vehicle obstacle crossing conditions; According to the vehicle ground mechanics principle, the failure conditions of the vehicle crossing the obstacle are analyzed and finally the Step Mask is formed; Step Mask: Failure conditions for vehicles to cross obstacles. Four conditions are analyzed: vehicles crossing vertical obstacles, ditches, suspended obstacles, and longitudinal slopes. Conditions for vehicles to cross vertical obstacles: For 4×4 wheeled vehicles, the performance of crossing stepped vertical protrusions is evaluated based on the maximum vertical obstacle height that can be overcome; the dimensionless expression of the maximum obstacle height that can be crossed is: Where h is the obstacle height, r is the wheel radius, μ is the friction coefficient, l is the vehicle wheelbase, and a is the horizontal distance from the front axle to the center of gravity of the vehicle; Vehicle crossing trench conditions: The width of the trench that a vehicle can cross is the width of the trench l d Ratio to wheel diameter D Indicates the ratio of the vertical obstacle height that can be crossed to the wheel diameter Convert to Given a specific obstacle height h, the corresponding ditch width ratio can be determined Conditions for vehicles to cross overhanging obstacles: For vehicles to cross overhanging obstacles, the lowest point of the obstacle is at a height of h from the ground. obj Must be greater than the height h of the obstacle point cloud relative to the LiDAR sensor pc The height of LiDAR from the ground h Lid When the vehicle encounters an overhead obstacle, it must meet the following conditions: h obj >h pc +h Lid Vehicle crossing longitudinal slope conditions: For longitudinal slope crossing, the platform must meet the following requirements: i max >a Among them, θ max is the maximum climbing angle that the vehicle can handle, and α is the current terrain slope angle; Step 1.4: Generate passable cost occupation mark; The voxel semantic categories and passable costs are initially mapped according to the rules and adjusted based on the StepMask obtained in step 1.3; the Step Mask is obtained by analyzing the vehicle's geometric passability, and voxels that do not meet the geometric passability are assigned fatal impassable costs; after the above steps, four types of passable cost labels are generated: fatal impassable, high cost, medium cost, and directly passable, which represent the passability from difficult to easy, respectively, and can be used to train a multimodal drivable occupancy prediction model.

3. The multimodal drivable occupancy prediction method for unstructured road environments according to claim 1, characterized in that: Step 2 specifically includes the following sub-steps: Step 2.1: Feature extraction; Extract environmental features using multimodal information; The model uses the ResNet-50 backbone network to extract the image from the monocular image I t Extract image features Where l×b represents the spatial resolution of the feature map, and d represents the feature dimension; At the same time, the sparse point cloud L from LiDAR t Voxelization to generate sparse voxel occupancy supervision S spa ∈{0, 1} H×W×Z ; Then, the query upsampling network is used to obtain the occupancy supervision S at a lower spatial resolution. den ∈{0, 1} H / 2×W / 2×Z / 2 , thereby improving the robustness of the query; The Mask Token module is used to represent empty voxel areas in the environment. When combined with voxel occupancy supervision, it forms a more comprehensive 3D voxel feature, which helps improve occupancy prediction performance. Step 2.2: Feature fusion; The deformable attention DA mechanism is used in the model to realize the interaction of local interest regions in 3D and 2D feature spaces; the deformable attention dynamically samples N regions around the reference point. s points, and adaptively learns the attention weights according to the query vector, while predicting the spatial offset of the sampling point relative to the reference point, thereby achieving flexible and efficient feature aggregation; Based on these sampling points, the attention result is calculated as follows: Among them, q is the query vector, p is the reference point corresponding to q, F is the input feature map, is the learnable weight matrix for each sampling point, used to generate the value vector value, A s ∈[0, 1] is the attention weight of each sampling point, which is learned adaptively according to the query. is the offset relative to the reference point p, F(p+δp s ) represents the position p+δp extracted from the input features by bilinear interpolation s The eigenvector of Deformable Cross Attention DCA, for Enhanced Occupancy Supervision Query S den With image features The interaction between them is realized by adopting multi-layer DCA to organically integrate multimodal information. First, the three-dimensional grid is projected into the two-dimensional image feature space. The projected two-dimensional point is S den The reference point is used to sample surrounding features; the weighted sum of the sampled features is used to obtain the output of the DCA layer, thereby obtaining a refined query that integrates environmental semantics and geometric features. For a query s at position p = (x, y, z) den , through the camera projection function Mapping the query to image I t The corresponding reference point on Deformable self-attention DSA, the refined query Combined with the mask token to form the initial 3D voxel feature F 3D ∈H / 2×W / 2×Z / 2×d, and then multi-layer DSA is used to obtain enhanced voxel features The DSA layer enables the model to learn spatial 3D occupancy features more efficiently and accurately; in the formula, the query q p is a Mask Token or a refined query at position p = (x, y, z): DSA(F 3D ,F 3D )=DA(q p ,p,F 3D ) Step 2.3: Output three-dimensional traversable cost occupancy prediction; Enhanced voxel features Through the fully connected layer FC and the softmax function, the dense occupancy prediction Y is output t ∈{c0,c1,…,c M } H×W×Z , where each voxel is either empty c0 or has a specific accessibility class {c1, c2, …, c M }; Mapping these predictions back to a 3D grid yields a detailed traversable occupancy map and provides annotation information that is critical in autonomous navigation in unstructured environments; Step 2.4: Training loss; ORDformer is trained in an end-to-end manner and minimizes three main loss components: semantic scale loss Geometric scale loss and weighted cross entropy loss Total training loss It is defined as the sum of these losses: The weighted cross entropy loss is calculated as follows: Among them, Ω represents the set of valid voxels, M is the number of categories, and tar i represents the true category of voxel i, is the weight of the true category of voxel i, y i,tar is the model's true category tar for voxel i i The logit output, y i,c The model predicts that voxel i is the logit output of category c; In the context of semantic scene completion, the precision P c , recall rate R c and specificity S c The metric evaluates model performance and measures the model's ability to correctly classify occupied and unoccupied voxels; the metric is defined as follows: Among them, p i,c represents the probability of predicting that voxel i belongs to category c, For Iverson symbol; To improve versatility, scale loss Maximize the above class-by-class metric: According to the above formula, the semantic loss is calculated respectively and geometric loss where tar sem , tar geo Represent the true labels of semantics and geometry respectively, p sem and p geo is the corresponding model prediction value.

Citation Information

Cited By

  • Unstructured environment motion planning method, device and equipment

    CN120593793A

  • Semantic map occupation prediction method and device, storage medium and electronic device

    CN120997431A