A Road Region Image Recognition Method Based on an Image and Point Cloud Fusion Network

By adopting an image-based and point cloud fusion network method in road area recognition, the point cloud input network is directly processed, which solves the problems of lighting conditions and point cloud conversion, and realizes high-precision road area detection.

CN113887349BActive Publication Date: 2025-05-27ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111098880.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-18
Publication Date
2025-05-27
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

The prior art is difficult to deal with the problem of changing outdoor lighting conditions in road area identification, especially when using RGB images for recognition, training is often difficult to achieve results in rainy days or nights under sunny days. At the same time, directly converting the point cloud into a pseudo-image form for fusion will lose the original structure of the point cloud and increase the computational complexity.

Method used

The road area recognition method based on the image and point cloud fusion network is adopted, and the features of images and point clouds are extracted by building a fusion backbone network and fusion is performed. The feature resolution is restored using the decoding network, and the pixel classification of the road area is finally performed through point-by-point convolution. This method does not require converting the point cloud into a pseudo-image, but directly input the point cloud into the network for processing.

Benefits of technology

It realizes high-precision detection of road areas in complex outdoor environments, avoids the impact of changes in lighting conditions on recognition, and reduces computing complexity and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887349B_ABST
    Figure CN113887349B_ABST
Patent Text Reader

Abstract

The present invention discloses a road area image recognition method based on image and point cloud fusion. A fusion backbone network is constructed to extract features from the original image and the original point cloud, and the two features are fused to obtain a fused feature map; a decoding layer is constructed using Upsampling, a 2D convolution layer, and a ReLU activation function layer, and a decoding network is constructed therefrom, and the fused feature map is input into the decoding network for processing to obtain a decoding feature result; a point-by-point convolution operation is used for the decoding feature result to obtain whether it is a road area classification category. The present invention solves the problem of direct fusion of images and point clouds, and directly inputs the original point cloud into the road area network without any pre-processing operation on the point cloud, so that the computational complexity of the entire method is low; it can stably and accurately detect road areas in complex environments with high precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and relates to a road image recognition method based on an image and point cloud fusion network for identifying road area images. Background Art

[0002] Autonomous vehicles need to identify road areas in the traffic environment to further plan their own driving trajectories. In diverse and complex traffic environments, accurately identifying road areas is very difficult due to factors such as the diversity of traffic scenes, traffic participants, and lighting conditions.

[0003] With the development of deep convolutional neural network technology, this technology has been successfully applied to various tasks, including road area recognition tasks. Such methods (typical representative: G.L. Oliveira, W. Burgard and T. Brox, "Efficient deep models for monocular road segmentation," 2016 IEEE / RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Korea (South), 2016, pp. 4885-4891) generally use RGB images captured by a monocular camera as input, and use a deep convolutional neural network as a feature extractor and classifier to classify each pixel in the image into two categories: "road" or "non-road". By connecting the pixels classified as the "road" category, a connected area is formed to obtain the finally recognized road area in the image. However, such methods face the challenge of being difficult to cope with the changing outdoor lighting conditions relying only on RGB images. For example, neural networks trained under sunny daytime conditions often have difficulty working on rainy days or at night.

[0004] To solve this problem, another type of method takes advantage of both the RGB images captured by a monocular camera and the point cloud scanned by a lidar as input, and improves the accuracy of road region recognition by designing a neural network that fuses image and point cloud information. This type of method (typical representative: Z. Chen, J. Zhang and D. Tao, "Progressive LiDAR adaptation for road detection," in IEEE / CAA Journal of Automatica Sinica, vol. 6, no. 3, pp. 693-702, May 2019) projects the point cloud information onto a 2D plane first, then rasterizes it, and represents the point cloud information in the form of a pseudo-image by constructing artificial features for each grid. Then, 2D convolution operations are used to extract feature point clouds and fuse them with the features extracted from the RGB images. However, all such methods require converting the point cloud into the form of a pseudo-image, losing the original structure of the point cloud during this conversion and increasing the operations, which affects both the accuracy and efficiency of the road recognition algorithm. Summary of the Invention

[0005] To break through the limitation that previous image and point cloud fusion technologies need to convert the point cloud into a pseudo-image, for complex outdoor scenes, the present invention proposes a road region image recognition method based on an image and point cloud fusion network.

[0006] As Figure 1 shown, the technical solution adopted by the present invention is as follows:

[0007] 1) Construct a fusion backbone network to extract features from the original image and the original point cloud, and fuse these two features to obtain a fused feature map;

[0008] 2) Then use Upsampling, 2D convolutional layers, and ReLU activation function layers to construct a decoding layer, and build a densely connected decoding network based on this. The decoding network is used to restore the resolution of the features, and the fused feature map is input into the decoding network for processing to obtain a decoded feature result;

[0009] The present invention uses the decoding network to improve the image information resolution for road region recognition. Specifically, it decodes the image features to restore the feature size to the size of the input image.

[0010] 3) Finally, perform pointwise convolution operations on the decoded feature result to obtain the classification category of each pixel in the original image as "road" or "non-road". Use pointwise convolution and feature detection to detect the pixels belonging to the road in the image.

[0011] The specific content of step 1) is as follows:

[0012] The fusion backbone network uses the image processing branch of ResNet-101 and the point cloud processing branch of PointNet++ to extract image appearance features and geometric feature point clouds from the original image and the original point cloud respectively, and uses a fusion module to fuse the image appearance features and the geometric feature point clouds to obtain a fused feature map.

[0013] The fusion of the image appearance features and the geometric feature point clouds is specifically to fuse the geometric feature point clouds onto the corresponding image appearance features.

[0014] The fusion of the image appearance features and the geometric feature point clouds is specifically divided into two steps: the alignment step of the image and the point cloud and the step of fusing the feature point cloud into the image:

[0015] In the described image and point cloud alignment step, by pre-calibrating the external parameter matrix of the lidar and the camera and the internal parameter matrix of the camera, first calculate the coordinates of the point cloud projected into the image coordinate system;

[0016] In the step of fusing the feature point cloud into the image, using the coordinates of the point cloud projected into the image coordinate system, select the corresponding points in the point cloud for each pixel in the image feature, and average the features of all corresponding points to obtain the feature obtained by the pixel from the point cloud as the final fused feature map.

[0017] The original point cloud and the original image of the present invention are obtained by detecting with a camera and a lidar installed at the front of the vehicle. The original point cloud is the front road data obtained simultaneously and synchronously with the original image.

[0018] The described image appearance features refer to the image features obtained by using the ResNet network as the feature extraction network and processing and outputting with an RGB image as the input.

[0019] The described geometric feature point cloud uses the PointNet++ network as the feature extraction network and obtains the feature point cloud by processing and outputting with a point cloud containing the three-dimensional coordinate information and the reflection future information of each point as the input.

[0020] As Figure 2 shown, the described fusion backbone network includes an image processing branch, a point cloud processing branch and a fusion module,

[0021] The described image processing branch includes five feature extraction blocks connected in cascade in sequence. The original image is input into the first feature extraction block and outputs their respective image features after being processed by the five feature extraction blocks in sequence; the feature extraction block is the structure in the ResNet-101 network,

[0022] The described point cloud processing branch includes four sequentially connected SA layers. The original point cloud is input into the first SA layer and, after being sequentially processed by five feature extraction blocks, outputs their respective feature point clouds. The SA layer is a structure in the PointNet++ network.

[0023] The results output by each feature extraction block, the results output by each SA layer, and the original point cloud are subjected to fusion transfer processing through multiple fusion modules and fed back into the feature extraction blocks. Specifically, the result output by the current feature extraction block and the feature point cloud / original point cloud output by the corresponding SA layer are subjected to fusion transfer processing through a fusion module and fed back into the next feature extraction block. That is, the image features output by the first feature extraction block and the original point cloud are subjected to fusion transfer processing through a fusion module and fed back into the second feature extraction block. The result output by the second feature extraction block and the feature point cloud output by the first SA layer are subjected to fusion transfer processing through a fusion module and fed back into the third feature extraction block. The result output by the third feature extraction block and the feature point cloud output by the second SA layer are subjected to fusion transfer processing through a fusion module and fed back into the fourth feature extraction block. The result output by the fourth feature extraction block and the feature point cloud output by the third SA layer are subjected to fusion transfer processing through a fusion module and fed back into the fifth feature extraction block. The result output by the fifth feature extraction block and the feature point cloud output by the fourth SA layer are subjected to fusion transfer processing through a fusion module and directly output.

[0024] Given an original image I 0 and an original point cloud P 0 , it is expressed as the following operation:

[0025]

[0026] F i = I i + Fusion(P j , I i ), j = i - 1, i ∈ {1, 2, 3, 4, 5}, j ∈ {0, 1, 2, 3, 4}

[0027]

[0028]

[0029] Among them, is the operation of the first feature extraction block, I i represents the image features output by the i-th feature extraction block, I 0 represents the original image, I 1 represents the image features output by the first feature extraction block, F irepresents the fused feature map output by the i-th fusion module, Fusion(·) is the operation of the fusion module, P j represents the feature point cloud output by the j-th SA layer, P 0 is the original point cloud, is the operation of the (j + 1)-th SA layer;

[0030] By looping the above operations, the output results of each fusion module are obtained, forming a set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5}.

[0031] The specific operation steps of the fusion module are as follows:

[0032] S1. Use the pre-calibrated extrinsic matrix of the lidar and the camera (this matrix is a 4x4 square matrix) and the intrinsic matrix K of the camera to find the pixel position in the image coordinate system of the image feature I i output by the i-th feature extraction block for each point in the feature point cloud P j :

[0033]

[0034] c i = 2 i

[0035] where P′ j is the homogeneous coordinate of P j , Q ij is the homogeneous coordinate of the feature point cloud P j in the image coordinate system of the image feature map I i , c i is the scaling scale constant corresponding to the image feature map I i , represents the floor operation;

[0036] S2. In this way, multiple points in the feature point cloud P j will be projected to the same pixel position in the image feature I i . Therefore, for each pixel of the image feature I i , select the points in the feature point cloud P j whose homogeneous coordinate is this pixel position to form a set, and take the average value of the feature values of all points in this set to obtain the feature of this pixel of the image feature I i acquired from the feature point cloud P j ;

[0037] S3. For the image feature Ii Each pixel in i is subjected to the above operations to form a complete image as the fused feature map F

[0038] As Figure 3 shown, the decoding network includes five decoding layers, denoted respectively as Each decoding layer is constructed by cascading Upsampling + 2D convolution + BN + ReLU + 2D convolution + BN + ReLU in sequence. Among them, Upsampling is implemented using bilinear interpolation, 2D convolution uses a convolution operation with a convolution kernel size of 3x3 and a padding size of 1, BN is a batch normalization layer, and ReLU is an activation function;

[0039] The five decoding layers are respectively processed in one-to-one correspondence with the five fused features in the set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5}. The current fused feature map in the set of fused feature maps {F Figure 1 , F 1 , F 2 , F 3 , F 4 , F 5} is input into its corresponding decoding layer for processing to obtain the current decoded feature, and the current decoded feature and the current fused feature Figure 1 are fed back to the next decoding layer for processing, specifically expressed as:

[0040]

[0041] Among them, is the call operation of the (i + 1)-th decoding layer, and U i represents the i-th decoded feature;

[0042] The (i + 1)-th decoding layer The specific steps are as follows: perform an Upsampling operation on the (i + 1)-th decoded feature U i+1 , then add the result obtained from the Upsampling operation to the (i + 1)-th fused feature map F 5-i , and then perform the operations of 2D convolution + BN + ReLU + 2D convolution + BN + ReLU on the added result in sequence;

[0043] The 5th fused feature map F 5 is used as the initial decoded feature U 0 ; for the 5th decoding layer the input is only the 4th decoded feature U4 , directly perform operations on the 4th decoded feature U 4 in sequence, including 2D convolution + BN + ReLU + 2D convolution + BN + ReLU operations to obtain the output of the 5th decoded feature U 5 .

[0044] Specifically, the pointwise convolution is to perform classification processing on the decoded feature results output by the decoding network through threshold judgment after convolution operation and Sigmoid operation in sequence.

[0045] The beneficial effects of the present invention are as follows:

[0046] 1) It solves the problem of direct fusion of images and point clouds. The original point cloud can be directly input into the road area network without any preprocessing operations on the point cloud, resulting in a relatively low computational complexity of the entire method;

[0047] 2) By fusing the information of images and point clouds, it can detect road areas in complex environments with high precision. For example Figure 4 as shown, in various environments, this method can stably and accurately detect road areas. Description of the Drawings

[0048] Figure 1 is the network flow chart of the present invention.

[0049] Figure 2 is the fusion backbone network diagram of the present invention.

[0050] Figure 3 is the densely connected decoding network in the present invention.

[0051] Figure 4 is the experimental result diagram for typical scenarios in the embodiments of the present invention. Each row in the figure represents an example scenario. The left figure in each row represents the schematic scenario, where the detection result is represented by a relatively light-colored area. To clearly show the detection result, see the right figure in each row, where the white part represents the detected road area. Specific Embodiments

[0052] The present invention will be further described below in conjunction with the drawings and specific embodiments.

[0053] The specific embodiment process of the present invention is as follows:

[0054] 1. Construct a fusion backbone network, extract features from images and point clouds, and fuse these two types of features. The specific steps are as follows:

[0055] 1.1. Use ResNet-101 to construct an image processing branch, which contains five feature extraction blocks, denoted as The operations of each feature extraction block are denoted as follows:

[0056]

[0057] Among them, is the operation of the i-th feature extraction block, I in is an input image feature or the original image, I out represents an image feature output after the operation of the feature extraction block, and its length and width dimensions are reduced to I in 1 / 2 of the length and width dimensions.

[0058] 1.2. Use PointNet++ to construct a point cloud processing branch, which contains four SA layers, denoted as The parameters required for constructing each SA layer are given in the following table:

[0059]

[0060] The operation of each SA layer is denoted as follows:

[0061]

[0062] Among them, is the operation of the i-th SA layer, P in is the input point cloud, P out is the output point cloud.

[0063] The set {P 0 formed by the input original point cloud P 1 , P 2 , P 3 , P 4 , P 5} is called the feature point cloud set, and each element in it is called a feature point cloud.

[0064] 1.3. Given an original image I 0 and the original point cloud P 0 , according to the results output by each current feature extraction block and the feature point cloud / original point cloud output by the corresponding SA layer, fusion transfer processing is performed through the current fusion module and fed back to the next feature extraction block. Such feedback transfer is represented by the following operation:

[0065]

[0066] F i = I i + Fusion(P j , I i ), j = i - 1, i ∈ {1, 2, 3, 4, 5}, j ∈ {0, 1, 2, 3, 4}

[0067]

[0068]

[0069] Among them, is the operation of the first feature extraction block, I i represents the feature point cloud output by the i-th feature extraction block, I 0 represents the original image, I 1 represents the image features output by the first feature extraction block, F i represents the fused feature map output by the i-th fusion module, Fusion(·) is the operation of the fusion module, P j represents the feature point cloud output by the j-th SA layer, P 0 is the original point cloud, is the operation of the (j + 1)-th SA layer;

[0070] By looping the above operations, the output results of each fusion module are obtained, forming a set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5}.

[0071] The specific operation steps of the fusion module in the specific implementation are as follows:

[0072] S1. Use the pre-calibrated extrinsic parameter matrix of the lidar and the camera (this matrix is a 4x4 square matrix) and the intrinsic parameter matrix K of the camera to find the pixel position in the image coordinate system of the image features I j output by the i-th feature extraction block for each point in the feature point cloud P i output by the j-th SA layer:

[0073]

[0074] c i = 2 i

[0075] where P′ j is the homogeneous coordinate of P j , Q ij is the homogeneous coordinate of the feature point cloud P j in the image coordinate system of the image feature map I i , c i is the scaling scale constant corresponding to the image feature map I i , represents the floor operation on the operation result;

[0076] S2. In this way, there will be a feature point cloud Pj Multiple points in i are projected onto the same pixel position in the image feature I, so for each pixel of the image feature I i a set of feature points P with homogeneous coordinates at this pixel position is selected j from the points in it, and the eigenvalues of all points in this set are averaged to obtain the feature of this pixel of the image feature I i from the feature point cloud P j ;

[0077] S3. Perform the above operations on each pixel in the image feature I i to form a complete image as the fused feature map F i .

[0078] 2. Use a decoding network and pointwise convolution to restore the feature size to the size of the input image and classify the pixels in the input picture as "road" and "non-road".

[0079] 2.1. Construct a densely connected decoding network

[0080] 2.1.1. Use Upsampling + 2D convolution + BN + ReLU + 2D convolution + BN + ReLU to construct a decoding layer.

[0081] Among them, Upsampling is implemented using bilinear interpolation;

[0082] The 2D convolution uses a convolution operation with a convolution kernel size of 3x3 and a padding size of 1; BN is a batch normalization layer, and ReLU is an activation function. The decoding layer is constructed in the above way.

[0083] 2.1.2. By constructing 5 decoding layers, denoted as construct a decoding network.

[0084] The input of the decoding network is the set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5}, and the specific representation of the decoding network is:

[0085]

[0086] Among them, is the call operation of the (i + 1)-th decoding layer, and U i represents the i-th decoded feature;

[0087] The (i + 1)-th decoding layer The specific steps are as follows: For the (i + 1)-th decoded feature U i+1Perform the upsampling operation, and then add the result obtained from the upsampling operation to the (i + 1)-th fused feature map F 5-i and then perform the operations of 2D convolution + BN + ReLU + 2D convolution + BN + ReLU on the added result in sequence;

[0088] The 5-th fused feature map F 5 is used as the initial decoded feature U 0 ; For the 5-th decoding layer the input is only the 4-th decoded feature U 4 and directly perform the operations of 2D convolution + BN + ReLU + 2D convolution + BN + ReLU on the 4-th decoded feature U 4 to obtain the output 5-th decoded feature U 5 .

[0089] 2.2 Pointwise Convolution

[0090] For the 5-th decoded feature U output by the decoding network 5 , use the convolution operation with a convolution kernel size of 1x1 and a number of channels of 1 as the pointwise convolution operation, and denote the result as S. S has the same property as the size of the input image.

[0091] Perform the Sigmoid operation on S to normalize the value of each pixel in S to the range (0, 1), and then make a judgment: when the value of a certain pixel in S is greater than or equal to 0.5, classify the pixel into the "road" category; when the value of a certain pixel in S is less than 0.5, classify the pixel into the "non-road" category.

[0092] 3. Training Process of the Neural Network. As can be seen from the previous description, the entire road area detection network used in the method is composed of a fused backbone network, a decoding network, and a pointwise convolution. The fused backbone network is further divided into an image processing branch and a point cloud processing branch.

[0093] 3.1. As can be seen from step 1.2, the point cloud processing branch is constructed by the PointNet++ network and trained on the Semantic-KITTI dataset. Only pre-train the point cloud processing branch of the fused backbone network to obtain its network parameter weights.

[0094] 3.2. Add the pre-trained network parameters of the point cloud processing branch of the fusion backbone network and freeze them. Then, train the entire network, including the fusion backbone network, the decoding network, and the pointwise convolution, on the Road task of the KITTI dataset. Use the negative log-likelihood loss, the SGD optimizer, and set the learning rate to 0.001 for mini-batch training, with the mini-batch set to 4. Through 1000 iterations of training, save the network parameter weights with the minimum loss during the training process.

[0095] 3.3. Take an image and the corresponding point cloud as input and feed them into the trained network to obtain the labels for each pixel in the image. The labels can only be "road" and "non-road". The area formed by all pixels belonging to "road" is the finally recognized road area.

[0096] A series of typical road scenarios were experimentally verified according to the embodiments of the present invention. The results are as Figure 4 shown. In the road detection task of the KITTI dataset, select the training set as the training data, construct the network and training method according to the description in the above invention specification, conduct training, and save the weight parameters with the minimum loss. Use the test set in the road detection task of the KITTI dataset for verification, and the Figure 4 results shown can be obtained. It can be seen from the results that the recognized road area has high accuracy in the original image.

Claims

1. A method for road area image recognition based on the fusion of images and point clouds, characterized in that: 1) Construct a fusion backbone network to extract features from the original image and the original point cloud, and fuse these two features to obtain a fused feature map; 2) Then use Upsampling, 2D convolutional layer and ReLU activation function layer to construct a decoding layer, and build a decoding network based on this. Input the fused feature map into the decoding network for processing to obtain a decoded feature result; 3) Finally, perform pointwise convolution operation on the decoded feature result to obtain the classification category of each pixel in the original image as "road" or "non-road"; The fusion backbone network includes an image processing branch, a point cloud processing branch and a fusion module. The image processing branch includes five feature extraction blocks connected in cascade in sequence. The original image is input into the first feature extraction block, and after being processed by the five feature extraction blocks in sequence, their respective image features are output; The point cloud processing branch includes four sequentially connected SA layers. The original point cloud is input into the first SA layer, and after being processed by the five feature extraction blocks in sequence, their respective feature point clouds are output; The results output by each feature extraction block, the results output by each SA layer and the original point cloud are fused and transmitted through multiple fusion modules and fed back into the feature extraction block; Expressed as the following operations: F i = I i + Fusion(P j , I i ), j = i - 1, i ∈ {1, 2, 3, 4, 5}, j ∈ {0, 1, 2, 3, 4} Among them, is the operation of the first feature extraction block, I i represents the image features output by the i-th feature extraction block, I 0 represents the original image, I 1 represents the image features output by the first feature extraction block, F i represents the fused feature map output by the i-th fusion module. Fusion(·) is the operation of the fusion module, P j represents the feature point cloud output by the j-th SA layer, P 0 is the original point cloud, is the operation of the (j + 1)-th SA layer; By looping the above operations, the output results of each fusion module are obtained, forming the set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5}; The specific operation steps of the fusion module are as follows: S1. Use the extrinsic matrix of the 4x4 square array of the pre-calibrated lidar and camera and the intrinsic matrix K of the camera to find the pixel position in the image coordinate system of each point in the feature point cloud P j output by the j-th SA layer in the image feature I i output by the i-th feature extraction block: c i =2 i Among them, P′ j is the homogeneous coordinate of P j , Q ij is the homogeneous coordinate of the feature point cloud P j in the image coordinate system of the image feature map I i , c i is the scaling scale constant corresponding to the image feature map I i , represents the floor operation; S2. For each pixel of the image feature I i select the feature point cloud P with homogeneous coordinates at the position of the pixel to form a set, and take the average value of the feature values of all points in the set to obtain the feature of the pixel of the image feature I j from the feature point cloud P i for this pixel of the image feature I j obtained from the feature point cloud P; S3. Perform the above operations on each pixel in the image feature I i to form a complete image as the fused feature map F i .

2. A method for road area image recognition based on the fusion of images and point clouds according to claim 1, characterized in that: The specific content of step 1) is: The fusion backbone network uses the image processing branch and the point cloud processing branch to extract image appearance features and geometric feature point clouds from the original image and the original point cloud respectively, and uses the fusion module to fuse the image appearance features and geometric feature point clouds to obtain a fused feature map.

3. A method for road area image recognition based on the fusion of images and point clouds according to claim 2, characterized in that: The fusion of the image appearance features and geometric feature point clouds is specifically to fuse the geometric feature point clouds onto the corresponding image appearance features.

4. A method for road area image recognition based on the fusion of images and point clouds according to claim 2 or 3, characterized in that: The fusion of the image appearance features and geometric feature point clouds is specifically divided into two steps: the alignment step of the image and the point cloud and the step of fusing the feature point clouds into the image: In the alignment step of the image and the point cloud, by pre-calibrating the external parameter matrix of the lidar and the camera and the internal parameter matrix of the camera, first calculate the coordinates of the point cloud projected into the image coordinate system; In the step of fusing the feature point clouds into the image, using the coordinates of the point cloud projected into the image coordinate system, select the corresponding points in the point cloud for each pixel in the image features, and average the features of all corresponding points to obtain the features obtained by the pixel from the point cloud as the final fused feature map.

5. A method for road area image recognition based on the fusion of images and point clouds according to claim 1, characterized in that: The described decoding network includes five decoding layers, and each decoding layer is constructed by cascading Upsampling + 2D convolution + BN + ReLU + 2D convolution + BN + ReLU in sequence. Among them, Upsampling is implemented using bilinear interpolation, the 2D convolution uses a convolution operation with a convolution kernel size of 3x3 and a padding size of 1, BN is a batch normalization layer, and ReLU is an activation function; The five decoding layers respectively correspond to and process the five fused feature maps in the set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5}, and each current fused feature map in the set of fused feature maps {F 1 , F 2 , F 3 , F 4 , F 5} is input into its corresponding decoding layer for processing to obtain the current decoded feature, and the current decoded feature and the current fused feature map are fed back to the next decoding layer for processing, which is specifically expressed as: Among them, is the call operation of the (i + 1)-th decoding layer, and U i represents the i-th decoded feature; The (i + 1)-th decoding layer Specifically, perform an upsampling operation on the (i + 1)-th decoded feature U i+1 Then add the result obtained from the upsampling operation to the (i + 1)-th fused feature map F 5-i And then perform operations of 2D convolution + BN + ReLU + 2D convolution + BN + ReLU on the added result in sequence; The 5th fused feature map F 5 is used as the initial decoded feature U 0 ; for the 5th decoding layer the input is only the 4th decoded feature U 4 , and directly perform operations of 2D convolution + BN + ReLU + 2D convolution + BN + ReLU on the 4th decoded feature U 4 in sequence to obtain the output of the 5th decoded feature U 5 .

6. A method for road area image recognition based on image and point cloud fusion according to claim 1, characterized in that: The pointwise convolution specifically classifies the decoded feature results output by the decoding network through threshold judgment after convolution operation and Sigmoid operation in sequence.