Cross-country feasible region detection method based on Transform and multi-modal feature fusion
By introducing the Transformer framework and the Off-road network that integrates multimodal features into the autonomous driving system, the problem of insufficient segmentation accuracy and feature extraction of driving area detection in off-road environments is solved, and the detection effect and robustness are improved.
Patent Information
- Application Number
- CN202510420759.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-06
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has reduced segmentation accuracy and insufficient feature extraction for driving area detection of autonomous driving in off-road environments, especially poor performance under complex and inclement weather conditions. The traditional CNN model has limited receptive fields, making it difficult to effectively deal with complex intertwined road boundaries.
Using the Off-road network based on Transformer, combining RGB images and surface normal information of lidar point clouds, feature extraction and fusion are performed through the CM-FRM module and the FFX module, and a layered upsampling decoder is designed to improve the intensity of feature information and reduce the computational complexity.
It improves the segmentation accuracy of driving area detection in off-road environments and the performance of long-distance road scenarios, enhances the perception ability of autonomous vehicles and the robustness of decision-making systems, and adapts to complex and unknown scenarios.
Smart Images

Figure CN120298844A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving. Specifically, it realizes the deep integration of artificial intelligence and human-centered autonomous driving under Industry 5.0, and proposes a model training method from the basic model to the detection of drivable areas on off-road environment roads in cross-domain learning in the system structure. Background Art
[0002] As an integral part of Industry 5.0, autonomous driving plays an important role in accelerating human-machine cooperation and promoting a people-centered industrial world. In the long-term social life, humans learn how to make various decisions from social, learning, working and other scenarios. Therefore, humans have advantages in standard selection, pattern recognition, decision-making and group perception. However, humans are easily disturbed by emotions, which further affects reasonable decision-making during manual driving. For autonomous vehicles, their driving environment is not limited to urban roads, and sometimes involves unstructured roads such as mines and off-roads. In urban environments, traversable space is usually defined as free space, that is, paved road open spaces without obstacles. At the same time, in unstructured roads, the concept of passable areas is relatively vague, often full of obstacles such as weeds and stumps, and affected by unknown weather conditions, such as rain, snow, fog or low visibility, which poses great challenges to autonomous driving. In addition, when the vehicle's driving environment is in unstructured roads such as mines and off-roads, combined with the interference of complex high-dynamic scenarios, the vehicle's perception system cannot guarantee robustness, and the decision-making system is also easily affected. However, current research mainly focuses on urban structured road environments, and there is relatively little research on other types of road environments, such as rural roads and mountain roads. Considering the high complexity and diversity of off-road environments, compared with structured roads, directly adopting existing structured road segmentation models may lead to poor prediction results and even problems such as a decline in performance indicators. Although existing research has been dedicated to improving the perception ability of off-road environments, there is currently no optimal solution for off-road environments and challenges of adverse scenarios. Most previous work has utilized CNN. Although CNN-based networks perform well in structured urban navigation with clear roads and structures, due to the limited receptive field of CNN and in areas with unclear boundaries, overlaps and complex intersections in off-road environments, their performance will decline. In addition, in the task of detecting passable areas for autonomous vehicles, perception is an indispensable module, and using multi-modal fusion is currently recognized as the most effective method to enhance the vehicle's perception ability. LiDAR point cloud data contains spatial geometric information, but lacks semantic information, while monocular RGB images contain higher-level environmental semantic information. By fusing LiDAR and cameras, it can produce excellent perception performance and make it more suitable for coping with the challenges of off-road environments and adverse scenarios. Cross-attention is widely used for feature extraction and fusion, but in off-road environments, due to the challenges of bad weather and complex scenarios, the performance of cross-attention is not ideal, especially in the edge areas of the predicted map and the acquisition of long views. Finally, in the task of free space detection, since the surface normals are consistent on the same road plane, surface normal information is easier to identify than depth information. To improve the segmentation accuracy, there has already been work inferring surface normal information from dense depth images and fusing it with image information to enhance free space detection performance. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide an off-road feasible area detection method based on Transformer and multi-modal feature fusion, which is used to solve the problems of decreased segmentation accuracy in off-road area detection and insufficient feature extraction in autonomous driving.
[0004] The off-road feasible area detection method adopts an Off-road network, which includes an encoding module and a decoding module;
[0005] Among them, the encoding module further includes a path embedding module, a first encoder module, a CM-FRM module, an FFX module, a second encoder module, a third encoder module, and a fourth encoder module.
[0006] In the feature extraction process of the encoding module, first, an RGB image with a resolution of H×W×3 and a surface normal image derived from lidar point cloud are input into the path embedding module. The input image is segmented into image blocks of a set size. Subsequently, the first encoder processes the feature representation output by the above path embedding module and outputs the encoded RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C , where C represents the number of channels, and H and W respectively represent the height and width of the RGB image; then they are input into the CM-FR module again;
[0007] In the CM-FRM module, the modal features of the parallel streams are refined at each stage of feature extraction, which is divided into two parts: channel dimension feature correction and spatial dimension feature correction. Specifically:
[0008] In the channel dimension feature correction part, the features are processed in two dimensions: for the input 2 modal features, the RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C , global max pooling is performed on each modal along the channel dimension respectively, and then global average pooling is performed to obtain 2 feature vectors. A total of 4 feature vectors are generated for the two modalities, and they are concatenated to obtain the feature Y∈R 4C ; then, the weight W of the feature is calculated successively using the MLP and Sigmoid functions C ∈R 2C , and it is separated according to the modality it belongs to as and
[0009] The features are corrected according to the weights, and finally the corrected features of the two modalities are calculated and
[0010]
[0011] In spatial dimension feature correction, the input RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C are concatenated into the feature Y' ∈ R 2S , where S = C; the weight W S ∈R 2S of the feature Y' is calculated using the MLP and Sigmoid functions successively, and it is separated into weights and
[0012] according to the modality it belongs to. Calculate the cross-modal spatial correction feature:
[0013]
[0014] where * represents multiplication in the spatial domain; combining the channel and spatial correction features, the final corrected feature is obtained:
[0015]
[0016] where λ C and λ S are two hyperparameters to be trained;
[0017] In the feature fusion stage of the encoding module, the FFX module is adopted for feature fusion work, and the processing process is as follows:
[0018] First, the input RGB out and X out are processed through the channel embedding module for channel embedding to convert their dimensions into a form suitable for subsequent processing; then, these embedded features are normalized respectively; next, the normalized features of the two modalities are concatenated together along the channel dimension to form a fusion feature containing RGB and surface normal information, with the dimension of H / 4 × W / 4 × 2C;
[0019] Next, the concatenated feature map is processed by a 1x1 convolution for feature fusion, and a new feature map with the dimension of H / 4 × W / 4 × C is output; subsequently, the new feature map extracts and enhances features further through depth convolution, activation function, and convolution layer to obtain an enhanced feature map; the enhanced feature map performs residual connection and is normalized again, and then fused to obtain the feature map X with the feature dimension of H / 4 × W / 4 × C;
[0020] In the efficient attention module, first, the input feature map X, with dimensions H / 4×W / 4×C, is flattened into a matrix with dimensions n×C as the initial input of the efficient attention module, where n = H / 4×W / 4; after being processed by the efficient attention module, the resulting features are added to the flattened matrix features corresponding to the originally input feature map X to form a residual structure output X1, with dimensions H / 4×W / 4×C1;
[0021] The second encoder receives the input feature map X1 for encoding processing and outputs a feature map X2 with size H / 8×W / 8×C2 to the MLP layer;
[0022] The third encoder receives the input feature map X2 for encoding processing and outputs a feature map X3 with size H / 16×W / 16×C3 to the MLP layer;
[0023] The fourth encoder receives the input feature map X3 for encoding processing and outputs a feature map X4 with size H / 16×W / 16×C3 to the MLP layer;
[0024] Where C1, C2, C3, and C4 are set parameters;
[0025] The MLP layer receives the multi-level feature maps X1, X2, X3, and X4, aggregates the channel information of the features, and outputs the processed feature maps; the output feature maps of each layer will be upsampled to restore to the size of H / 4×W / 4×C; all upsampled feature maps are subjected to convolutional processing; finally, all convolved feature maps will enter the Fusion Module for fusion operation. In this module, first, the tensors corresponding to these feature maps are stacked into a new tensor in the index order of the original tensor list, and then a summation operation is performed along the first dimension of the new tensor to generate a new feature map; finally, the final prediction result is output through another MLP layer;
[0026] Finally, the Off-road network is trained: In the network training stage, the input data is the lidar point cloud and image pair data, including three categories, passable area, impassable area, and unreachable area;
[0027] Finally, the Off-road network is applied for recognition: The lidar point cloud and image pair data to be recognized are input into the trained Off-road network, and the feasible area is output.
[0028] Preferably, when training the Off-road network, SGDM is used as the optimizer, the number of epochs is set to 30, the initial learning rate is set to 0.00095, and the batch size is set to 2.
[0029] Preferably, when training the Off-road network, the image size for training and testing is set to 1280×704.
[0030] Preferably, when evaluating the model performance of the Off-road network, five metrics, namely Accuracy, Precision, Recall, F-score, and IOU, are adopted.
[0031] Preferably, the normalization method includes the Scaling function and the Softmax function.
[0032] Preferably, in the efficient attention module, the flattened matrix X is separated into Q, K, and V, which are respectively subjected to normalization calculations to obtain results with a dimension of n×d v ; then, this result is reshaped into H / 4×W / 4×d v ; if d v ≠C, a 1x1 convolution is applied to it to restore the dimension to C; finally, X1 is output.
[0033] Preferably, after the Softmax in the normalization method, it is calculated through E(Q,K,V) = ρ q (Q)(ρ k (K) T V) to obtain results with a dimension of n×d v ; where ρ q and ρ k are respectively the normalization functions for the query and key features.
[0034] The present invention has the following beneficial effects:
[0035] (1) The present invention proposes an Off-road model framework for a multi-modal perception and decision-making system based on off-road roads, which is used for detecting traversable areas in off-road environments. The aim is to solve problems such as insufficient extraction of modal features during the fusion of RGB images and depth images, unclear segmentation in autonomous driving, and poor performance in long-distance road scenarios.
[0036] (2) The present invention uses the CM-FRM module to replace the cross-attention mechanism in the benchmark model to meet the more complex feature extraction requirements in off-road scenarios. In the feature fusion part, the present invention proposes the FFX module to aggregate information from cameras and LiDAR sensors, thereby improving the effect of road detection tasks in off-road environments. Finally, the present invention redesigns the decoder part and adopts a hierarchical upsampling method to further enhance the intensity of feature information while reducing the number of parameters and computational complexity. Description of the Drawings
[0037] Figure 1 It is the algorithm architecture of the encoder.
[0038] Figure 2 It is the CM-FRM algorithm architecture.
[0039] Figure 3 It is the FFX algorithm architecture.
[0040] Figure 4 It is the algorithm structure of the decoder.
[0041] Figure 5 It is the visualization result of the segmentation result of the model proposed by the present invention. Detailed implementation manners
[0042] The present invention will be described in detail below in conjunction with the accompanying drawings and by way of examples.
[0043] The present invention proposes an Off-road network for combining camera and LiDAR information (i.e., surface normal information calculated from LiDAR point cloud). The Off-road network of the present invention includes an encoding module and a decoding module; wherein, the encoding module further includes a Path Embedding module, a first encoder module, a CM-FRM module, an FFX module, a second encoder module, a third encoder module, and a fourth encoder module.
[0044] In the free space detection task, since the surface normals are consistent on the same road plane. Therefore, the surface normal information is easier to identify than the depth information, which is beneficial for feature extraction and fusion with RGB images. The present invention selects the surface normal information as the input of the network. Since the passable area detection task requires the network to have a large receptive field, and the receptive field of CNN is limited, using the original CNN will lead to a decline in the image processing effect. While the Transformer framework performs very well in capturing local and global information, so the present invention also selects to use the Transformer framework for model design. According to the specific problems that occur, the present invention introduces the CM-FRM module to extract features, and proposes the feature fusion module FFX. At the same time, the decoder part adopts a hierarchical upsampling method to optimize the enhancement of feature information, and reduce the number of parameters and computational complexity. This model performs excellently in the experimental results. After the large model is migrated in the off-road environment, the vehicle's system will be more generalizable and can make correct perception and decisions in the face of complex and unknown scenarios, and can achieve safer and more intelligent decisions in applications such as route planning and cooperative driving.
[0045] In the encoder stage, the present invention uses the Transformer module and the CM-FRM module for feature extraction, and introduces the innovative FFX module to dynamically fuse multi-modal features. The process is as Figure 1As shown below. In addition, the present invention redesigned the encoder by means of a convolutional hierarchical output method, and the specific method is as follows.
[0046] (1) During the feature extraction process of the encoding module, first, an RGB image with a resolution of H×W×3 and a surface normal image derived from lidar point cloud are input into the Path Embedding module, and the input image is segmented into image blocks of a set size. Subsequently, the first encoder Transformer Encoder 1 processes the feature representation output by the above path embedding module and outputs the encoded RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C , where C = 64 represents the number of channels, and H and W represent the height and width of the RGB image respectively; then it is input into the CM-FR module again.
[0047] (2) In the CM-FRM module, the modal features of the parallel streams are refined at each stage of feature extraction, which is divided into two parts: channel dimension feature correction and spatial dimension feature correction, as Figure 2 shown.
[0048] In the channel dimension feature correction part, the CM-FRM module processes the features in two dimensions for the noise and uncertainty in different modes. Specifically, the two input modal features are the RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C . Global max pooling is performed on each modality along the channel dimension respectively, and then global average pooling is performed to obtain 2 feature vectors. A total of 4 feature vectors are generated for the two modalities, and they are concatenated to obtain the feature Y∈R 4C . Then, the weight W C ∈R 2C of the feature Y is calculated successively using the MLP and Sigmoid functions, and it is separated into and
[0049]
[0050] according to the modality to which it belongs. The features are corrected according to the weights, and finally, the corrected features of the two modalities and
[0051]
[0052] In the spatial dimension feature correction, the input RGB feature RGB in ∈RH / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C are concatenated into the feature Y' ∈ R 2S , where S = C. Similarly, the weights W of the feature Y' are calculated using the MLP and Sigmoid functions successively S ∈R 2S , and they are separated into weights according to their respective modalities and
[0053]
[0054] Calculate the cross-modal spatially corrected feature:
[0055]
[0056] where * represents spatial multiplication. Combining the channel and spatially corrected features, the final corrected feature is obtained:
[0057]
[0058] where λ C and λ S are two hyperparameters. They are both set to 0.5 as the default value. RGB out and X out are the rectified features after comprehensive calibration, and then they are sent to the feature fusion stage
[0059] (3) As Figure 3 shown, in the feature fusion stage of the encoding module, the present invention proposes the FFX module to perform efficient feature fusion work. The processing process of each module of the FFX (Feature Fusion eXtractor) network for data is as follows:
[0060] 1) First, the input RGB out and X out are processed through the Channel Embedding module for channel embedding, and their dimensions are converted into a form suitable for subsequent processing. Then, these embedded features are normalized respectively through the Add&Norm module to ensure the uniformity and stability of the data. Next, the normalized features of the two modalities are concatenated together along the channel dimension by the Concatenate and merge modules to form a fusion feature containing RGB and surface normal information, with the dimension of H / 4 × W / 4 × 2C
[0061] Next, the concatenated feature map is processed by a 1x1 convolution for feature fusion, and a new feature map with dimensions H / 4×W / 4×C is output. Subsequently, the new feature map is successively passed through a depthwise convolution (DWConv 3x3), an activation function (ReLU), and a convolutional layer (Conv 1) to further extract and enhance features, resulting in an enhanced feature map. After passing through this module, the enhanced feature map is passed to the second Add&Norm module for residual connection and another normalization process, and then fused through the Fmerged module to obtain the feature X with dimensions H / 4×W / 4×C.
[0062] 2) Efficient attention module. As Figure 4 shown, in the efficient attention module, first, the input feature map X with dimensions H / 4×W / 4×C is flattened into a matrix with dimensions n×C as the initial input of the efficient attention module, where n = H / 4×W / 4. Then, this matrix is separated into Q, K, and V, which are respectively normalized by ρ q and ρ k (the normalization methods include Scaling and Softmax, and Softmax is adopted in the present invention), and then calculated through E(Q,K,V)=ρ q (Q)(ρ k (K) T V) to obtain a result with dimensions n×d v . Then, this result is reshaped into H / 4×W / 4×d v . If d v ≠C, a 1x1 convolution is applied to restore the dimensions to C. Finally, the obtained feature is added to the flattened matrix feature corresponding to the initially input feature map X to form a residual structure and output X1 with dimensions H / 4×W / 4×C1.
[0063] Among them, ρ q and ρ k are respectively the normalization functions for the query and key features. The same two normalization methods in dot attention are adopted:
[0064]
[0065] Among them, σ row ,σ col respectively represent applying the softmax function along each row or each column of the matrix y.
[0066] (4) To obtain multi-level features, we designed 3 Transformer encoders. The input and output of each Transformer module are as follows:
[0067] ① Transformer Encoder 2 receives the input feature map X1 for encoding and obtains a feature map X2 of size H / 8×W / 8×C2 and outputs it to the MLP layer.
[0068] ②Transformer Encoder 3 receives the input feature map X2 for encoding and obtains a feature map X3 of size H / 16×W / 16×C3 and outputs it to the MLP layer.
[0069] ③Transformer Encoder 4 receives the input feature map X3 for encoding and obtains a feature map X4 of size H / 16×W / 16×C3 and outputs it to the MLP layer.
[0070] Among them, C1=64, C2=128, C3=320, C4=512.
[0071] In the decoding module, the present invention redesigns the decoder, such as Figure 4 As shown, local and global information can be effectively integrated.
[0072] First, the MLP layer included in the decoder module accepts the multi-level feature maps X1, X2, X3, and X4 from the encoder. The MLP layer aggregates the channel information of the features and outputs the processed feature maps. The output feature map of each layer will be upsampled through the Upsample layer and restored to the size of H / 4×W / 4×C. All upsampled feature maps are passed to the Convolutional Module for convolution processing. Finally, all convolved feature maps will enter the Fusion Module for fusion operation. In this module, the tensors corresponding to these feature maps will first be stacked into a new tensor according to the index order of the original tensor list, and then the summation operation will be performed along the first dimension of the new tensor to generate a new feature map. Finally, the final prediction result is output through the MLP layer.
[0073] (4) Finally, the Off-road network is trained; in the network training stage, the present invention uses the ORFD dataset, which has a total of 12198 frames of lidar point cloud and image pair data, including three categories, passable areas, impassable areas and inaccessible areas (such as the sky), and uses SGDM as the optimizer. The number of epochs is set to 30, the initial learning rate is set to 0.00095, and the batch size is set to 2. The image size used for training and testing is set to 1280×704. When evaluating the performance of the model, this embodiment uses the five indicators of Accuracy, Precision, Recall, F-score, and IOU. Among them, the hyperparameter λ also needs to be learned C and λ S .
[0074] (5) Finally, the trained Off-road network is applied for recognition: the input data is an image of 1280×704, and the output is the feasible region.
[0075] The results of this experiment are as Figure 5 shown, including RGB images, surface normals, segmentation images, and label images. The present invention quotes the research results in recent years and conducts a detailed analysis of the model results, and the results are shown in Table 2. Compared with the FuseNet and FtFoot models with depth and RGB image information as inputs, and the SNE-RoadSeg and FSN_Swin models and Off_Net models with RGB images and surface normals as inputs, better evaluation indicators are obtained.
[0076] Table II
[0077] Experimental results
[0078]
[0079] In summary, the above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An off-road feasible area detection method, which is characterized in that an Off-road network is adopted, which includes an encoding module and a decoding module; Among them, The encoding module further includes a path embedding module, a first encoder module, a CM-FRM module, an FFX module, a second encoder module, a third encoder module, and a fourth encoder module. During the feature extraction process of the encoding module, an RGB image with a resolution of H×W×3 and a surface normal image derived from lidar point cloud are first input into the path embedding module. The input image is segmented into image patches of a set size. Subsequently, the first encoder processes the feature representation output by the above path embedding module and outputs the encoded RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C , where C represents the number of channels, and H and W represent the height and width of the RGB image respectively; then it is input into the CM-FR module again; In the CM-FRM module, at each stage of feature extraction, the modal features of the parallel streams are refined, which are divided into two parts: channel dimension feature correction and spatial dimension feature correction. Specifically: In the channel dimension feature correction part, it is further divided into processing features in two dimensions: for the two input modal features, the RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C , global max pooling is performed along the channel dimension for each modality respectively, and then global average pooling is performed to obtain two feature vectors. A total of four feature vectors are generated for the two modalities, and they are concatenated to obtain the feature Y ∈ R 4C ; then, the weight W C ∈R 2C of the feature Y is calculated successively using the MLP and Sigmoid functions, and it is separated into and The features are corrected according to the weights, and finally the corrected features of the two modalities are calculated and In spatial dimension feature correction, the input RGB feature RGB in ∈R H / 4×W / 4×C and the point cloud feature X in ∈R H / 4×W / 4×C are concatenated into the feature Y' ∈ R 2S , where S = C; the weights W S ∈R 2S of the feature Y' are calculated using an MLP and a Sigmoid function successively, and are separated into weights and Calculate the cross-modal spatial correction feature: where * represents multiplication in the spatial domain; synthesize the channel and spatial correction features to obtain the final correction feature: Among them, λ C and λ S are two hyperparameters to be trained; In the feature fusion stage of the encoding module, the FFX module is used to perform feature fusion. The processing process is as follows: First, the input RGB out and X out are processed by the channel embedding module for channel embedding to convert their dimensions into a form suitable for subsequent processing. Then, these embedded features are respectively normalized. Next, the normalized features of the two modalities are concatenated together along the channel dimension to form a fused feature containing RGB and surface normal information, with a dimension of H / 4×W / 4×2C; Next, the concatenated feature map is processed by a 1x1 convolution for feature fusion, and a new feature map with dimensions H / 4 × W / 4 × C is output; subsequently, the new feature map passes through a depth convolution, an activation function, and a convolutional layer to further extract and enhance features, obtaining an enhanced feature map; the enhanced feature map is subjected to a residual connection and another normalization process, and then fused to obtain a feature map X with dimensions H / 4 × W / 4 × C; In the efficient attention module, first, the input feature map X with dimensions H / 4 × W / 4 × C is flattened into a matrix with dimensions n × C as the initial input of the efficient attention module, where n = H / 4 × W / 4; after being processed by the efficient attention module, the obtained feature is added to the flattened matrix feature corresponding to the originally input feature map X to form a residual structure output X1 with dimensions H / 4 × W / 4 × C1; The second encoder receives the input feature map X1 for encoding processing and outputs a feature map X2 with dimensions H / 8 × W / 8 × C2 to the MLP layer; The third encoder receives the input feature map X2 for encoding processing and outputs a feature map X3 with dimensions H / 16 × W / 16 × C3 to the MLP layer; The fourth encoder receives the input feature map X3 for encoding processing and outputs a feature map X4 with dimensions H / 16 × W / 16 × C3 to the MLP layer; where C1, C2, C3, and C4 are set parameters; The MLP layer receives the multi-level feature maps X1, X2, X3, and X4, aggregates the channel information of the features, and outputs the processed feature maps; the output feature maps of each layer will be upsampled to restore to the size of H / 4 × W / 4 × C; all upsampled feature maps are subjected to convolutional processing; finally, all the convolved feature maps will enter the Fusion Module for fusion operation. In this module, first, the tensors corresponding to these feature maps are stacked into a new tensor in the index order of the original tensor list, and then a summation operation is performed along the first dimension of the new tensor to generate a new feature map; finally, the final prediction result is output through another MLP layer; Finally, train the Off-road network: In the network training stage, the input data is the lidar point cloud and image pair data, including three categories: passable area, impassable area, and unreachable area; Finally, apply the Off-road network for recognition: Input the lidar point cloud and image pair data to be recognized into the trained Off-road network, and output the feasible area.
2. The off-road feasible area detection method according to claim 1, wherein when training the Off-road network, SGDM is used as the optimizer, the number of epochs is set to 30, the initial learning rate is set to 0.00095, and the batch size is set to 2.
3. The off-road feasible area detection method according to claim 2, wherein when training the Off-road network, the image size for training and testing is set to 1280×704.
4. The off-road feasible area detection method according to claim 2, wherein when evaluating the model performance of the Off-road network, five metrics are adopted: Accuracy, Precision, Recall, F-score, and IOU.
5. The off-road feasible area detection method according to claim 1, wherein the normalization method includes the Scaling function and the Softmax function.
6. The off-road feasible area detection method according to claim 1, wherein in the efficient attention module, the flattened matrix X is separated into Q, K, and V, and after normalization calculations respectively, results with dimensions of n×d are obtained. v Then, this result is reshaped into H / 4×W / 4×d. v If d v ≠C, a 1x1 convolution is applied to it to restore the dimension to C; finally, X1 is output.
7. The off-road feasible area detection method according to claim 6, wherein the normalization method includes, after Softmax, through calculation, to obtain a result with a dimension of n×d v ; where ρ q and ρ k are the normalization functions for the query and the key feature, respectively.
Citation Information
Cited By
Cross-country road identification method based on Transform network
CN122090415A
Off-road road recognition method based on a transformer network
CN122090415B