A Road Segmentation Method Based on Adaptive Progressive Fusion of Multimodal Features

By adopting the adaptive progressive fusion method of multimodal features in the road segmentation model, the problem of poor fusion effect of RGB images and depth maps in the prior art is solved, and high-precision and robust road segmentation in complex scenarios are achieved.

CN119399478BActive Publication Date: 2025-06-24ANHUI ANZHIXIN TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411979125.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-06-24
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The existing fusion method of RGB images and lidar depth maps is difficult to achieve ideal results in complex scenarios, and ignores the contextual relationship between different mode data, resulting in insufficient segmentation accuracy and robustness.

Method used

Adaptive progressive fusion method based on multimodal features is adopted, and by building a road segmentation model including RGB information encoding module, depth information encoding module, feature conversion module, feature progressive adaptive fusion module and multi-scale feature decoding module, the processing strategies of different regions are dynamically adjusted, and the information of RGB images and depth maps is fully utilized through feature conversion and progressive adaptive fusion mechanisms.

Benefits of technology

The accuracy and robustness of road segmentation are significantly improved in complex scenarios, especially in low-light, occlusion and dynamic environments, and the efficiency of multi-modal data fusion is improved through multi-scale feature decoding modules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399478B_ABST
    Figure CN119399478B_ABST
Patent Text Reader

Abstract

The present invention discloses a road segmentation method based on adaptive progressive fusion of multi-modal features, comprising the following steps: constructing a road image segmentation data set, which includes a number of road RGB images obtained by photographing and road depth images obtained by lidar; constructing a road segmentation model, which consists of an RGB information encoding module, a depth information encoding module, a feature conversion module, a feature progressive adaptive fusion module, and a multi-scale feature decoding module; inputting the road RGB images and road depth images into the road segmentation model for fusion segmentation to obtain the final road segmentation result; through the uniquely designed feature conversion module and feature progressive adaptive fusion module, the present invention enables the depth information to dynamically supplement the RGB information, thereby improving the accuracy and robustness of the road segmentation model in complex scenarios, especially performing more excellently in low-light, occlusion, and dynamic environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and specifically provides a road segmentation method based on adaptive progressive fusion of multi-modal features. Background Art

[0002] In the road segmentation task, traditional RGB image-based image segmentation methods have achieved remarkable results. However, relying solely on the features of RGB images usually has certain limitations. Especially in the case of lighting changes, weather changes, or complex scenes, the performance of RGB images may not be stable and reliable enough. Therefore, how to effectively fuse other types of perceptual data to improve segmentation accuracy and robustness has become a key research direction in recent years.

[0003] The depth map obtained by Light Detection and Ranging (LiDAR) can provide more accurate depth information than traditional RGB images due to its high-precision spatial measurement ability, greatly improving the accuracy of scene understanding. In the road segmentation task, depth information can effectively help the model understand the geometric shape of objects and distinguish the spatial relationship between the road and other objects, thereby reducing the occurrence of segmentation errors in complex backgrounds. However, there are still some defects and deficiencies in the existing technologies during the fusion process of RGB images and LiDAR depth maps. First, depth maps usually have a low resolution and large noise, which may introduce unnecessary errors when directly fusing them with high-resolution RGB images. Although existing fusion methods adopt mechanisms such as channel attention and spatial attention to strengthen the synergistic effect between different modalities, due to the huge differences in information expression between RGB images and depth maps, simply relying on traditional fusion methods is still difficult to achieve ideal results. Second, current fusion methods mostly ignore the context relationship between different modality data, such as the joint utilization of texture features with semantic information in RGB images and geometric features with structural information in depth maps. Existing methods often fail to perform effective multi-level and multi-scale fusion on the basis of fully considering the complementary nature of these different types of information.

[0004] Therefore, there is still much room for improvement in the accuracy, robustness, and efficiency of existing RGB-depth map fusion technologies. There is an urgent need to develop a new fusion mechanism that can better fuse these two modalities of data, make full use of the rich semantic information of RGB images and the spatial geometric characteristics of LiDAR depth maps, and further improve the performance of road segmentation. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a road segmentation method based on adaptive progressive fusion of multi-modal features, aiming to solve the problems mentioned in the background art.

[0006] To achieve the above object, the present invention provides the following technical solutions: A road segmentation method based on multi-modal feature adaptive progressive fusion, comprising the following steps:

[0007] Step S1: Construct a road image segmentation dataset, which includes a number of road RGB images obtained by photographing and road depth images obtained by lidar;

[0008] Step S2: Construct a road segmentation model, which consists of an RGB information encoding module, a depth information encoding module, a feature conversion module, a feature progressive adaptive fusion module, and a multi-scale feature decoding module; The RGB information encoding module is composed of a P2T network combined with an Agent attention mechanism, abbreviated as Agent P2T network; The Agent P2T network consists of 4 layers with the same structure; The depth information encoding module uses a ResNet network, and the ResNet network consists of 5 layers connected in sequence, and the 5-layer structure of the ResNet network is the same; The multi-scale feature decoding module consists of 4 layers with the same structure;

[0009] Step S3: First extract the surface normal of the road depth image to obtain the surface normal feature, and then input the surface normal feature into the depth information encoding module layer by layer for extraction to obtain the road depth feature information of each layer;

[0010] Step S4: Input the road RGB image into the first layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the first layer;

[0011] Step S5: Input the road depth feature information extracted from the second layer of the depth information encoding module and the road RGB feature information extracted from the first layer of the RGB information encoding module into the feature conversion module correspondingly to obtain the first improved depth feature information, and input the first improved depth feature information and the road RGB feature information extracted from the first layer of the RGB information encoding module into the feature progressive adaptive fusion module to obtain the first fusion feature information, and use the first fusion feature information as the input of the second layer of the RGB information encoding module, and so on, to obtain the second fusion feature information, the third fusion feature information, and the fourth fusion feature information;

[0012] Step S6: Input the fourth fusion feature information and the third feature fusion information into the first layer of the multi-scale feature decoding module to obtain the first-level upsampled feature. Input the first-level upsampled feature and the second feature fusion information into the second layer of the multi-scale feature decoding module to obtain the second-level upsampled feature. Input the second-level upsampled feature and the first feature fusion information into the third layer of the multi-scale feature decoding module to obtain the third-level upsampled feature. Input the third-level upsampled feature and the road depth feature information D1 extracted from the first layer of the depth information encoding module into the fourth layer of the multi-scale feature decoding module to obtain the fourth-level upsampled feature. After passing the fourth-level upsampled feature through the activation function, obtain the final output of the road segmentation model, that is, the final road segmentation result.

[0013] Further, the specific process of step S5 is as follows: Correspondingly input the road depth feature information extracted from the second layer of the depth information encoding module and the road RGB feature information extracted from the first layer of the RGB information encoding module into the feature conversion module to obtain the first improved depth feature information. Input the first improved depth feature information and the road RGB feature information of the first layer into the feature progressive adaptive fusion module to obtain the first fusion feature information. Input the first fusion feature information into the second layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the second layer. Input the road RGB feature information of the second layer and the road depth feature information extracted from the third layer of the depth information encoding module into the feature conversion module to obtain the second improved depth feature information. Input the second improved depth feature information and the road RGB feature information of the second layer into the feature progressive adaptive fusion module to obtain the second fusion feature information. Input the second fusion feature signal into the third layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the third layer. Input the road RGB feature information of the third layer and the road depth feature information extracted from the fourth layer of the depth information encoding module into the feature conversion module to obtain the third improved depth feature information. Input the third improved depth feature information and the road RGB feature information of the third layer into the feature progressive adaptive fusion module to obtain the third fusion feature information. Input the third fusion feature signal into the fourth layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the fourth layer. Input the road RGB feature information of the fourth layer and the road depth feature information extracted from the fifth layer of the depth information encoding module into the feature conversion module to obtain the fourth improved depth feature information. Input the fourth improved depth feature information and the road RGB feature information of the fourth layer into the feature progressive adaptive fusion module to obtain the fourth fusion feature information.

[0014] Further, in step S3, first perform surface normal extraction on the road depth image to obtain the surface normal feature, which is expressed as:

[0015] ;

[0016] ;

[0017] ;

[0018] In the formula, represents the obtained normal direction; represents normal normalization; represents the norm of the normal; represents the finally obtained surface normal feature; represents the road depth image; , respectively represent the pixel coordinates in the horizontal and vertical directions in the road depth image; represents the gradient in the horizontal direction of the road depth image; represents the gradient in the vertical direction of the road depth image; represents the vertical component of the normal.

[0019] Furthermore, the process of inputting the surface normal feature into the depth information encoding module for layer-by-layer extraction to obtain the road depth feature information of each layer is as follows:

[0020] Input the surface normal feature obtained after surface normal extraction into the first layer of the depth information encoding module, that is, the first layer of the ResNet network, to obtain the road depth feature information extracted in the first layer , which is expressed as:

[0021] ;

[0022] In the formula, is a 7×7 convolutional layer; is the convolutional kernel weight; is the batch normalization layer; is the activation function; is the max pooling layer of 3×3 convolution;

[0023] Input D1 into the second layer of the ResNet network to obtain the road depth feature information extracted in the second layer , and input into the third layer of the ResNet network to obtain the road depth feature information extracted in the third layer , and input into the fourth layer of the ResNet network to obtain the road depth feature information extracted in the fourth layer , and input into the fifth layer of the ResNet network to obtain the road depth feature information extracted in the fifth layer .

[0024] Furthermore, each layer of the Agent P2T network includes an image segmentation and vectorization layer, a Transformer encoding layer, a fusion layer, and a linear transformation layer. The Transformer encoding layer consists of an Agent attention mechanism and a feed-forward network FFN;

[0025] The specific process of step S4 is as follows: Input the road RGB image into the first layer of the RGB information encoding module, that is, the first layer of the Agent P2T network, for feature extraction. First, the road RGB image is embedded through the image segmentation and vectorization layer, and the output after the embedding process enters the Transformer encoding layer, and multi-scale features are extracted through the Agent attention mechanism and then fused with the input in the fusion layer to obtain multi-scale fusion features . The multi-scale fusion features are enhanced through the feed-forward network FFN, and the enhanced output and the multi-scale fusion features are input into the fusion layer for fusion to obtain multi-scale enhanced features . Finally, the multi-scale enhanced features are mapped to the target dimension through the linear transformation layer to obtain the road RGB feature information extracted by the first layer of the RGB information encoding module :

[0026] ;

[0027] ;

[0028] ;

[0029] ;

[0030] In the formula, is the input road RGB image; is a two-dimensional convolutional layer; represents the output after passing through the image segmentation and vectorization layer; represents the Agent attention mechanism; is a layer normalization operation; represents the feed-forward network FFN; is a linear transformation layer; is a bias term.

[0031] Furthermore, the road depth feature information extracted by the second layer of the depth information encoding module and the road RGB feature information extracted by the first layer of the RGB information encoding module The specific process of inputting into the feature transformation module to obtain the first improved depth feature information is as follows: First, the feature transformation module and are subjected to channel stacking, and the output after channel stacking goes through a convolution operation, and then passes through the ReLU activation function to obtain the first transformed feature map and the second transformed feature map , which is expressed as:

[0032] ;

[0033] ;

[0034] In the formula, represents convolution layer; represents and the output obtained after channel stacking;

[0035] First, goes through a convolution operation, and then performs a linear transformation on the features of each pixel point in . The result after the linear transformation passes through a ReLu activation function to obtain the preliminary transformed depth feature. Finally, through and , the preliminary transformed depth feature is scaled and biased to obtain the first improved depth feature information , which is expressed as:

[0036] .

[0037] Furthermore, the process of inputting the first improved depth feature information and the road RGB feature information of the first layer into the feature progressive adaptive fusion module to obtain the first fusion feature information is as follows: First, the feature progressive adaptive fusion module performs a splicing operation on and by adding features, and inputs the output after the splicing operation into the gated attention mechanism AG module. The gated attention mechanism AG module consists of channel attention and spatial attention, and generates an attention map from the output after the fusion operation through channel attention and spatial attention, which is expressed as:

[0038] ;

[0039] ;

[0040] ;

[0041] ;

[0042] Wherein, is a splicing operation; is and the output after the splicing operation; and are respectively the outputs after channel attention and spatial attention; is a 3×3 convolution; is an activation function; is an element-wise multiplication operation;

[0043] The generated attention map pairs and are weighted, and the weighted and are added to obtain the first fused feature information , expressed as:

[0044] ;

[0045] ;

[0046] ;

[0047] Wherein, represents the output weighted by the attention map; represents the output weighted by the attention map.

[0048] Furthermore, the fourth fused feature information and the third fused feature information are input into the first layer of the multi-scale feature decoding module to obtain the first-level upsampled feature . The specific process is as follows: First, the multi-scale feature decoding module uses bilinear interpolation to upsample the fourth fused feature information by two times, and then the upsampled is concatenated with the third fused feature information through a skip connection to obtain a preliminary concatenated feature map , expressed as:

[0049] ;

[0050] ;

[0051] Wherein, represents the two-times upsampling operation on ;

[0052] Subsequently, after passing through a depthwise separable convolution, the output after the depthwise separable convolution is compressed in the number of channels through a convolution, and then the output after compressing the number of channels passes through the ReLU activation function to increase non-linearity and is accelerated in training through batch normalization; the output after batch normalization passes through the first dilated convolutional layer, and the output of the first dilated convolutional layer passes through the second dilated convolutional layer to obtain the segmentation feature map , which is expressed as:

[0053] ;

[0054] ;

[0055] ;

[0056] In the formula, represents the depthwise separable convolution; and respectively represent the first dilated convolutional layer and the second dilated convolutional layer, and the kernel sizes of the first dilated convolutional layer and the second dilated convolutional layer are respectively and ;

[0057] Input into the SE attention module to obtain the weighted feature map , and multiply the weighted feature map with the segmentation feature map to obtain the segmentation weighted feature map , which is expressed as:

[0058] ;

[0059] ;

[0060] In the formula, represents the global average pooling operation; and respectively represent the first fully connected layer and the second fully connected layer;

[0061] Input , and are concatenated by channel to form the fused segmentation weighted feature map , and then is reduced in the number of channels through a convolution, and the after reducing the number of channels is regularized to obtain the first-level upsampling feature , expressed as:

[0062] ;

[0063] ;

[0064] ;

[0065] In the formula, represents the regularization operation.

[0066] Compared with the existing technology, the present invention has the following beneficial effects:

[0067] (1) By combining the multi-modal method of road RGB images and road depth images, the present invention can simultaneously utilize color features and spatial depth information, thereby improving the accuracy and robustness of road segmentation in complex scenarios, especially performing more excellently in low-light, occlusion, and dynamic environments.

[0068] (2) By using the P2T network combined with the Agent attention mechanism, the present invention can dynamically adjust the processing strategy of the road segmentation model in different regions. Compared with the traditional CNN network, it has stronger adaptability and flexibility, can more accurately process complex backgrounds, dynamic changes, and occlusion situations, and improves the accuracy and robustness of road segmentation.

[0069] (3) Through the uniquely designed feature transformation module and feature progressive adaptive fusion module, the present invention enables depth information to dynamically supplement RGB information, thereby improving the performance of the road segmentation model; the depth information is effectively transformed through the feature transformation module to better adapt to the fusion with RGB information; while the feature progressive adaptive fusion module dynamically adjusts the information weights of the two modalities according to the task requirements to ensure that the complementary advantages of depth information and RGB information can be fully exerted in different scenarios.

[0070] (4) By adopting a multi-scale feature decoding module composed of depthwise separable convolution, dilated convolution, and skip connection, the present invention improves the efficiency of multi-modal data fusion. Especially when processing RGB information and depth information, it can better fuse them. The adaptive selection of key channels is further strengthened through the SE attention mechanism, effectively enhancing the accuracy and generalization ability in the road segmentation task, especially performing more outstandingly in complex and dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is a schematic structural diagram of the road segmentation model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0072] The present invention provides a technical solution: a road segmentation method based on adaptive progressive fusion of multi-modal features, comprising the following steps:

[0073] Step S1: Construct a road image segmentation dataset, which includes a number of road RGB images obtained by photographing and road depth images obtained by lidar.

[0074] The road image segmentation dataset uses the KITTI dataset, which contains road RGB images collected by photographing in scenarios such as urban areas, rural areas, and highways, and road depth images obtained by lidar.

[0075] Among them, the road RGB images and road depth images contain various weather conditions and traffic conditions for better application in real life.

[0076] The roads in the KITTI dataset mainly have the following three types: UU (urban unmarked) urban unmarked lanes, UM (urban marked) urban marked lanes, and UMM (urban multiply marked) urban multi-marked lanes.

[0077] Step S2: Construct a road segmentation model, as Figure 1 shown, the road segmentation model is composed of an RGB information encoding module, a depth information encoding module, a feature transformation module, a feature progressive adaptive fusion module, and a multi-scale feature decoding module; among them, the RGB information encoding module and the depth information encoding module are in a parallel structure, and the feature transformation module and the feature progressive adaptive fusion module are interspersed between the layers of the RGB information encoding module and the depth information encoding module, and the multi-scale feature decoding module is connected in series after the RGB information encoding module; Figure 1 in represents the feature progressive adaptive fusion module.

[0078] The RGB information encoding module is composed of a P2T (Pyramid Pooling Transformer) network combined with an Agent attention mechanism, abbreviated as the Agent P2T network; the Agent P2T network is composed of 4 layers with the same structure.

[0079] The depth information encoding module uses a ResNet network, which is composed of 5 layers connected in sequence, and the 5-layer structure of the ResNet network is the same.

[0080] The multi-scale feature decoding module is composed of 4 layers with the same structure.

[0081] Step S3: First, perform surface normal extraction on the road depth image to obtain surface normal features, and then input the surface normal features into the depth information encoding module for extraction layer by layer to obtain the road depth feature information of each layer.

[0082] Among them, each layer of the depth information encoding module includes a convolutional layer, a batch normalization layer, a Relu activation function, and a max pooling layer.

[0083] Step S3.1: First, perform surface normal extraction on the road depth image. By calculating the normal of each point, local geometric information of the object surface is captured. Since the points in the road area are almost on a ground plane, the points on this ground plane have similar surface normals. After the surface normal extraction layer extracts features from the depth map, effective geometric information can be provided, expressed as:

[0084] (1);

[0085] (2);

[0086] (3);

[0087] In the formula, represents the obtained normal direction; represents normal normalization; represents the norm of the normal; represents the finally obtained surface normal features; represents the road depth image; , respectively represent the pixel coordinates in the horizontal and vertical directions of the road depth image; represents the gradient in the horizontal direction of the road depth image; represents the gradient in the vertical direction of the road depth image; represents the vertical component of the normal.

[0088] Step S3.2: Input the surface normal features obtained after surface normal extraction into the first layer of the depth information encoding module, that is, the first layer of the ResNet network. First, the surface normal features pass through a convolutional layer with a convolutional kernel size of 7×7 and a stride of 2, then pass through a batch normalization layer to accelerate training and improve generalization ability, and then pass through a Relu activation function to help the ResNet network introduce non-linear characteristics, avoid the possible problem of gradient disappearance, and at the same time improve computational efficiency. Finally, through a max pooling layer for convolutional pooling operation to further reduce the spatial dimension of the features, thereby effectively reducing the number of parameters and computational amount in the network; this not only helps to improve computational efficiency, but also enhances the robustness of the ResNet network to rotation and displacement changes, and finally obtains the road depth feature information extracted by the first layer for subsequent processing , expressed as:

[0089] (4);

[0090] In the formula, is a 7×7 convolutional layer; is the convolutional kernel weight; is the batch normalization layer; is the activation function; is the max pooling layer of 3×3 convolution.

[0091] Step S3.2: Input D1 into the second layer of the ResNet network to obtain the road depth feature information extracted by the second layer , and input into the third layer of the ResNet network to obtain the road depth feature information extracted by the third layer , and input into the fourth layer of the ResNet network to obtain the road depth feature information extracted by the fourth layer , and input into the fifth layer of the ResNet network to obtain the road depth feature information extracted by the fifth layer .

[0092] Step S4: Input the road RGB image into the first layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the first layer.

[0093] Among them, each layer of the Agent P2T network includes a patch embedding layer, a Transformer encoding layer (including the Agent attention mechanism and the feed-forward network FFN), a fusion layer, and a linear transformation layer.

[0094] The specific process of Step S4 is as follows: Input the road RGB image into the first layer of the RGB information encoding module, that is, the first layer of the Agent P2T network for feature extraction; First, the road RGB image is embedded through the patch embedding layer, and the output after the embedding process enters the Transformer encoding layer, and the multi-scale features of are extracted through the Agent attention mechanism and then fused with input into the fusion layer to obtain the multi-scale fusion feature . The multi-scale fusion feature is enhanced through the feed-forward network FFN, and the enhanced output and the multi-scale fusion feature are input into the fusion layer for fusion to obtain the multi-scale enhanced feature , finally, the multi-scale enhanced features are mapped to the target dimension through a linear transformation layer to obtain the road RGB feature information extracted by the first layer of the RGB information encoding module : where

[0095] (5)

[0096] (6);

[0097] (7);

[0098] (8);

[0099] In the formula, is the input road RGB image; is a two-dimensional convolutional layer; represents the output after passing through the Patch Embedding layer; represents the Agent attention mechanism; is a layer normalization operation; represents the feed-forward network FFN, which is used for feature projection and feature enhancement; is a linear transformation layer; is a bias term.

[0100] Step S5: The road depth feature information extracted by the second layer of the depth information encoding module and the road RGB feature information extracted by the first layer of the RGB information encoding module are correspondingly input into the feature conversion module to obtain the first improved depth feature information. The first improved depth feature information and the road RGB feature information extracted by the first layer of the RGB information encoding module are input into the feature progressive adaptive fusion module to obtain the first fusion feature information. The first fusion feature information is used as the input of the second layer of the RGB information encoding module, and so on, to obtain the second fusion feature information, the third fusion feature information, and the fourth fusion feature information.

[0101] The specific process of step S5 is as follows: The road depth feature information extracted by the second layer of the depth information encoding module and the road RGB feature information extracted by the first layer of the RGB information encoding module are correspondingly input into the feature conversion module. The feature conversion module learns conversion parameters through multiple convolution operations and transforms the road depth feature information to better supplement and improve the road RGB feature information, obtaining the first improved depth feature information. The first improved depth feature information and the road RGB feature information of the first layer are input into the feature progressive adaptive fusion module to obtain the first fusion feature information. The first fusion feature information is input into the second layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the second layer. The road RGB feature information of the second layer and the road depth feature information extracted by the third layer of the depth information encoding module are input into the feature conversion module to obtain the second improved depth feature information. The second improved depth feature information and the road RGB feature information of the second layer are input into the feature progressive adaptive fusion module to obtain the second fusion feature information. The second fusion feature signal is input into the third layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the third layer. The road RGB feature information of the third layer and the road depth feature information extracted by the fourth layer of the depth information encoding module are input into the feature conversion module to obtain the third improved depth feature information. The third improved depth feature information and the road RGB feature information of the third layer are input into the feature progressive adaptive fusion module to obtain the third fusion feature information. The third fusion feature signal is input into the fourth layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the fourth layer. The road RGB feature information of the fourth layer and the road depth feature information extracted by the fifth layer of the depth information encoding module are input into the feature conversion module to obtain the fourth improved depth feature information. The fourth improved depth feature information and the road RGB feature information of the fourth layer are input into the feature progressive adaptive fusion module to obtain the fourth fusion feature information.

[0102] Among them, the specific process of correspondingly inputting the road depth feature information extracted by the second layer of the depth information encoding module and the road RGB feature information extracted by the first layer of the RGB information encoding module into the feature conversion module to obtain the first improved depth feature information is as follows: The feature conversion module first performs and channel stacking on the outputs after channel stacking, performs a convolution operation with a convolution kernel size of 1×1, which maps the features of each pixel in the channel (without changing the spatial dimensions), and then passes through the ReLU activation function to obtain the first conversion feature map and the second conversion feature map , expressed as:

[0103] (9);

[0104] (10);

[0105] Wherein, represents the convolutional layer; represents and the output obtained after channel stacking.

[0106] First, is subjected to a convolution operation, and then for the features of each pixel point are linearly transformed in the channel, and the result after the linear transformation is passed through a ReLu activation function to obtain a preliminary transformed depth feature. Finally, through and the preliminary transformed depth feature is scaled and biased to obtain the first improved depth feature information , expressed as:

[0107] (11).

[0108] Similarly, the specific processes of obtaining the second improved depth feature information, the third improved depth feature information, and the fourth improved depth feature information are the same as the specific process of obtaining the first improved depth feature information, and will not be elaborated here.

[0109] Among them, the process of inputting the first improved depth feature information and the road RGB feature information of the first layer into the feature progressive adaptive fusion module to obtain the first fusion feature information is as follows: The feature progressive adaptive fusion module first performs a splicing operation on and by adding features, and inputs the output after the splicing operation into the gated attention mechanism AG module. The gated attention mechanism AG module consists of channel attention and spatial attention, and generates an attention map from the output after the fusion operation through channel attention and spatial attention , expressed as:

[0110] (12);

[0111] (13);

[0112] (14);

[0113] (15);

[0114] Wherein, is the splicing operation; is and the output after the splicing operation; and are respectively the outputs after channel attention and spatial attention; is a 3×3 convolution with a padding of 1; is the activation function; is the element-wise multiplication operation.

[0115] The attention map characterizes the importance of each spatial position in and and determines which parts should be emphasized and which parts should be suppressed in subsequent steps, making the fusion process an adaptive state rather than specifying the fusion weights of the two modalities.

[0116] Weight the generated attention map for and , add the weighted and to obtain the first fused feature information , expressed as:

[0117] (16);

[0118] (17);

[0119] (18);

[0120] In the formula, represents the output weighted by the attention map; represents the output weighted by the attention map.

[0121] Similarly, the specific processes of obtaining the second, third, and fourth fused feature information are the same as the specific process of obtaining the first fused feature information, which will not be elaborated here.

[0122] Step S6: Input the fourth fused feature information and the third feature fusion information into the first layer of the multi-scale feature decoding module to obtain the first-level upsampled feature, and the first-level upsampled feature The second feature fusion information and the first feature fusion information are input into the second layer of the multi-scale feature decoding module to obtain the secondary upsampled features. The secondary upsampled features and the first feature fusion information are input into the third layer of the multi-scale feature decoding module to obtain the tertiary upsampled features. The tertiary upsampled features and the road depth feature information D1 extracted from the first layer of the depth information encoding module are input into the fourth layer of the multi-scale feature decoding module to obtain the quaternary upsampled features. The quaternary upsampled features are passed through an activation function to obtain the final output of the road segmentation model, that is, the final road segmentation result.

[0123] Among them, the fourth fusion feature information and the third fusion feature information are input into the first layer of the multi-scale feature decoding module to obtain the primary upsampled features The specific process is as follows:

[0124] First, the multi-scale feature decoding module uses the method of bilinear interpolation to complete upsampling by a factor of two, and then the upsampled is concatenated with the third fusion feature information through a skip connection to obtain a preliminary concatenated feature map , which is expressed as:

[0125] (19);

[0126] (20);

[0127] In the formula, represents the operation of upsampling by a factor of two.

[0128] Subsequently, is passed through a depthwise separable convolution. By performing independent convolution on each input channel of , the computational complexity is reduced. The output after the depthwise separable convolution is passed through a 3×3 convolution to compress the number of channels. Then, the output after compressing the number of channels is passed through a ReLU activation function to increase non-linearity and is accelerated by batch normalization during training. The output after batch normalization is passed through the first dilated convolution layer, and the output of the first dilated convolution layer is passed through the second dilated convolution layer to obtain the segmentation feature map ; Among them, a batch normalization operation and a ReLU function are connected after the first dilated convolutional layer and the second dilated convolutional layer. The convolutional kernel sizes and dilation factors of the first dilated convolutional layer and the second dilated convolutional layer are (3, 2) and (2, 2) respectively. The role of using dilated convolution is to increase the receptive field of the convolutional kernel without increasing the computational amount and improve the segmentation accuracy; expressed as:

[0129] (21);

[0130] (22);

[0131] (23);

[0132] In the formula, represents depthwise separable convolution; and respectively represent the first dilated convolutional layer and the second dilated convolutional layer. The convolutional kernel sizes of the first dilated convolutional layer and the second dilated convolutional layer are and .

[0133] Input into the SE attention module. Through the SE attention module, adjust the weights of each channel, highlight the important channels and suppress the unimportant channels to obtain the weighted feature map . Multiply the weighted feature map with the segmentation feature map to obtain the segmentation weighted feature map , expressed as:

[0134] (24);

[0135] (25);

[0136] In the formula, represents the global average pooling operation; and respectively represent the first fully connected layer and the second fully connected layer.

[0137] Concatenate , and channel by channel to form the fused segmentation weighted feature map . Subsequently, pass through a 1 1 convolution to reduce the number of channels. Regularize the with the reduced number of channels to obtain the first-level upsampling feature , expressed as:

[0138] (26);

[0139] (27);

[0140] (28);

[0141] Wherein, represents a regularization operation.

[0142] Similarly, the specific processes for obtaining the secondary upsampling feature, the tertiary upsampling feature, and the quaternary upsampling feature are the same as the specific process for the primary upsampling feature, and will not be elaborated here.

[0143] Most road segmentation tasks use the cross-entropy loss function for optimization. However, this method may encounter difficulties when dealing with the situation of unbalanced samples. Especially in road segmentation tasks, the number of pixels in the road area is usually much less than that in the non-road area, resulting in the non-road area may have an excessive impact on the model during the model training process. This imbalance may cause the model to tend to identify as the non-road area when performing pixel classification, thus reducing the segmentation accuracy. To solve this problem, the present invention combines the boundary loss and the cross-entropy loss as the total loss of the road segmentation model , expressed as:

[0144] (29);

[0145] Wherein, and are the cross-entropy loss function and the boundary loss function respectively.

[0146] This loss function design enables the model to not only perform well in global pixel classification but also optimize in local details, and is particularly suitable for tasks that require fine boundary division.

[0147] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A road segmentation method based on adaptive progressive fusion of multimodal features, characterized in that: The steps include: Step S1: constructing a road image segmentation dataset, wherein the road image segmentation dataset includes a number of road RGB images acquired by photographing and road depth images acquired by laser radar; Step S2: constructing a road segmentation model, the road segmentation model consists of an RGB information encoding module, a depth information encoding module, a feature conversion module, a feature progressive adaptive fusion module and a multi-scale feature decoding module; the RGB information encoding module is composed of a P2T network combined with an Agent attention mechanism, referred to as an Agent P2T network; the Agent P2T network consists of four layers with the same structure; the depth information encoding module adopts a ResNet network, the ResNet network consists of five layers connected in sequence, and the five layers of the ResNet network have the same structure; the multi-scale feature decoding module consists of four layers with the same structure; Step S3: first extracting the surface normal of the road depth image to obtain the surface normal features, and then inputting the surface normal features into the depth information encoding module for extraction layer by layer to obtain the road depth feature information of each layer; Step S4: inputting the road RGB image into the first layer of the RGB information encoding module for feature extraction to obtain the first layer of road RGB feature information; Step S5: inputting the road depth feature information extracted by the second layer of the depth information encoding module and the road RGB feature information extracted by the first layer of the RGB information encoding module into the feature conversion module to obtain first improved depth feature information, inputting the first improved depth feature information and the road RGB feature information extracted by the first layer of the RGB information encoding module into the feature progressive adaptive fusion module to obtain first fused feature information, and using the first fused feature information as the input of the second layer of the RGB information encoding module, and so on, to obtain second fused feature information, third fused feature information and fourth fused feature information; Step S6: input the fourth fusion feature information and the third feature fusion information into the first layer of the multi-scale feature decoding module to obtain a first-level up-sampling feature, input the first-level up-sampling feature and the second feature fusion information into the second layer of the multi-scale feature decoding module to obtain a second-level up-sampling feature, input the second-level up-sampling feature and the first feature fusion information into the third layer of the multi-scale feature decoding module to obtain a third-level up-sampling feature, input the third-level up-sampling feature and the road depth feature information D1 extracted by the first layer of the depth information encoding module into the fourth layer of the multi-scale feature decoding module to obtain a fourth-level up-sampling feature, and obtain the final output of the road segmentation model, i.e., the final road segmentation result, after the fourth-level up-sampling feature is subjected to an activation function; First improve the depth feature information and the RGB feature information of the road at layer 1 The specific process of inputting the first fusion feature information into the feature progressive adaptive fusion module is as follows: the feature progressive adaptive fusion module first adds the features to the first fusion feature information. and Perform a splicing operation and input the output after the splicing operation into the gated attention mechanism AG module. The gated attention mechanism AG module consists of channel attention and spatial attention. The output after the fusion operation is generated into an attention map through channel attention and spatial attention. , expressed as: ; ; ; ; In the formula, for and Output after splicing operation; For splicing operation; and They are Output after channel attention and spatial attention; is the batch normalization layer; It is a 3×3 convolution; for Activation function; It is an element-by-element multiplication operation; for Activation function; The generated attention map is and Weighted, the weighted and Add and get the first fusion feature information , expressed as: ; ; ; In the formula, express The output weighted by the attention map; express Output weighted by the attention map.

2. The road segmentation method based on multi-modal feature adaptive progressive fusion according to claim 1, characterized in that: The specific process of step S5 is: inputting the road depth feature information extracted by the second layer of the depth information encoding module and the road RGB feature information extracted by the first layer of the RGB information encoding module into the feature conversion module to obtain the first improved depth feature information; inputting the first improved depth feature information and the road RGB feature information of the first layer into the feature progressive adaptive fusion module to obtain the first fused feature information; inputting the first fused feature information into the second layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the second layer; inputting the road RGB feature information of the second layer and the road depth feature information extracted by the third layer of the depth information encoding module into the feature conversion module to obtain the second improved depth feature information; inputting the second improved depth feature information and the road RGB feature information of the second layer into the feature progressive adaptive fusion module to obtain the second fused feature information; and inputting the second fused feature signal into the feature progressive adaptive fusion module to obtain the second fused feature information. The road RGB feature information of the third layer is input into the third layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the third layer; the road RGB feature information of the third layer and the road depth feature information extracted by the fourth layer of the depth information encoding module are input into the feature conversion module to obtain the third improved depth feature information; the third improved depth feature information and the road RGB feature information of the third layer are input into the feature progressive adaptive fusion module to obtain the third fused feature information; the third fused feature signal is input into the fourth layer of the RGB information encoding module for feature extraction to obtain the road RGB feature information of the fourth layer; the road RGB feature information of the fourth layer and the road depth feature information extracted by the fifth layer of the depth information encoding module are input into the feature conversion module to obtain the fourth improved depth feature information; the fourth improved depth feature information and the road RGB feature information of the fourth layer are input into the feature progressive adaptive fusion module to obtain the fourth fused feature information.

3. The road segmentation method based on multi-modal feature adaptive progressive fusion according to claim 2 is characterized by: In step S3, the road depth image is first subjected to surface normal extraction to obtain surface normal features, which are expressed as: ; ; ; In the formula, represents the obtained normal direction; Indicates normal normalization; represents the modulus of the normal line; Represents the final surface normal feature; represents a road depth image; , Respectively represent the pixel coordinates in the horizontal and vertical directions in the road depth image; Represents the horizontal gradient of the road depth image; Represents the vertical gradient of the road depth image; Represents the vertical component of the normal.

4. The road segmentation method based on multi-modal feature adaptive progressive fusion according to claim 3 is characterized by: The surface normal features are input into the depth information encoding module for extraction layer by layer. The specific process of obtaining the road depth feature information of each layer is as follows: The surface normal features obtained after surface normal extraction are input into the first layer of the depth information encoding module, i.e., the first layer of the ResNet network, to obtain the road depth feature information extracted by the first layer. , expressed as: ; In the formula, It is a 7×7 convolutional layer; is the convolution kernel weight; It is the maximum pooling layer of 3×3 convolution; Input D1 into the second layer of the ResNet network to obtain the road depth feature information extracted by the second layer ,Will Input to the third layer of the ResNet network to obtain the road depth feature information extracted by the third layer ,Will Input to the 4th layer of the ResNet network to obtain the road depth feature information extracted by the 4th layer ,Will Input to the 5th layer of the ResNet network to obtain the road depth feature information extracted by the 5th layer .

5. The road segmentation method based on multi-modal feature adaptive progressive fusion according to claim 4 is characterized by: Each layer of the Agent P2T network includes a picture segmentation vectorization layer, a Transformer encoding layer, a fusion layer, and a linear transformation layer. The Transformer encoding layer is composed of an Agent attention mechanism and a feedforward network FFN. The specific process of step S4 is: convert the road RGB image The input is sent to the first layer of the RGB information encoding module, that is, the first layer of the Agent P2T network for feature extraction. First, the road RGB image is embedded through the image segmentation and vectorization layer. The output after embedding is Enter the Transformer encoding layer and extract it through the Agent attention mechanism After the multi-scale features Input the fusion layer for fusion to obtain multi-scale fusion features , the multi-scale fusion features are enhanced through the feed-forward network FFN, and the enhanced output is fused with the multi-scale fusion feature input fusion layer to obtain the multi-scale enhanced features Finally, the multi-scale enhanced features are transformed through the linear transformation layer Mapped to the target dimension, the road RGB feature information extracted by the first layer of the RGB information encoding module is obtained : ; ; ; ; In the formula, is the input road RGB image; It is a two-dimensional convolutional layer; Represents the output of the image segmentation vectorization layer; Represents the Agent attention mechanism; It is the layer normalization operation; Represents the feed-forward network FFN; is the linear transformation layer; is the bias term.

6. The road segmentation method based on multi-modal feature adaptive progressive fusion according to claim 5, characterized in that: The road depth feature information extracted by the second layer of the depth information encoding module The road RGB feature information extracted by the first layer of the RGB information encoding module The specific process of inputting the corresponding information into the feature conversion module to obtain the first improved depth feature information is as follows: the feature conversion module first converts and Perform channel stacking and output the channel stacking After a convolution operation, the first conversion feature map is obtained by the ReLU activation function. and the second conversion feature map , expressed as: ; ; In the formula, express Convolutional layers; express and The output after channel stacking; Will After a convolution operation, The features of each pixel are linearly transformed on the channel, and the result of the linear transformation is then passed through a ReLu activation function to obtain the preliminary transformed depth features. and Scaling and biasing the initial transformed deep features to obtain the first improved deep feature information , expressed as: 。 7. The road segmentation method based on multi-modal feature adaptive progressive fusion according to claim 6, characterized in that: The fourth fusion feature information and the third feature fusion information are input into the first layer of the multi-scale feature decoding module to obtain the first-level up-sampled feature The specific process is as follows: First, the multi-scale feature decoding module converts the fourth fusion feature information Use bilinear interpolation to perform double upsampling, and then Fusion feature information with the third Splicing is performed through jump connections to obtain a preliminary splicing feature map , expressed as: ; ; In the formula, Express Perform a two-fold upsampling operation; Then will After a depth-wise separable convolution, the output of the depth-wise separable convolution is A convolution is used to compress the number of channels, and then the output after the compression of the number of channels is The ReLU activation function is used to increase nonlinearity, and batch normalization is used to accelerate training; the output after batch normalization is After the first dilated convolutional layer, the output of the first dilated convolutional layer After the second dilated convolutional layer, the segmentation feature map is obtained , expressed as: ; ; ; In the formula, represents depthwise separable convolution; and They represent the first dilated convolutional layer and the second dilated convolutional layer, respectively. The convolution kernel sizes of the first dilated convolutional layer and the second dilated convolutional layer are and ; Will Input SE attention module to get weight feature map , the weight feature map and segmentation feature map Multiply them together to get the segmentation weight feature map , expressed as: ; ; In the formula, represents the global average pooling operation; and Represent the first fully connected layer and the second fully connected layer respectively; Will , and Splice by channel to form a fused segmentation weight feature map , and then By reducing the number of channels through a convolution, the number of channels after reducing Perform regularization operation to obtain the first-level upsampling feature , expressed as: ; ; ; In the formula, represents the regularization operation.

Citation Information

Patent Citations

  • RGB-D saliency target detection method based on boundary deformable convolution guidance

    CN115830420A

  • RGB-D semantic segmentation method and system based on adaptive context sensing network

    CN116580192A