A method and system for extracting and classifying road markings based on MLS point cloud
Patent Information
- Application Number
- CN202410887551.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-03
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-07-03
AI Technical Summary
[0005]为了解决上述问题,本发明的目的是提供一种基于MLS点云的道路标线提取分类技术,意在解决现有技术侧重于从移动激光扫描(MLS)点云中提取道路标线,其性能受到路面背景复杂、道路标线尺度变化以及边界模糊等限制问题,提升提取分类道路标线的能力
本发明解决了现有技术中由于侧重于从移动激光扫描(MLS)点云中提取道路标线,而造成的性能受到路面背景复杂、道路标线尺度变化以及边界模糊等限制问题,提升了道路标线的提取分类能力,通过在道路标线的准确提取在智能交通系统、高精地图构建、自动驾驶等领域具有重要意义。
Smart Images

Figure CN118674994B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent traffic mapping and recognition technology, and more specifically, to a method and system for extracting and classifying road markings based on MLS point clouds. Background Technology
[0002] Road markings are various traffic signs and text painted on the road surface, including lane lines, stop lines, zebra crossings, arrows, and text. They play a vital role in road traffic safety, providing necessary guidance and protection for road users, thereby reducing the number of accidents and alleviating traffic congestion. Accurate identification and extraction of road markings are fundamental to technologies such as traffic monitoring, autonomous driving, and lane-level high-precision mapping.
[0003] Existing research largely relies on video and images acquired by vehicle-mounted digital cameras to extract and classify road markings. However, this method is easily affected by weather conditions and lighting. Unlike optical images, moving laser scanning (MLS) point clouds are not sensitive to lighting conditions and possess geometric information and reflectance intensity. Therefore, a series of studies have been conducted on extracting road markings using relevant features of MLS point clouds. For example, some methods use the geometric information of the point cloud to extract markings using Hough transform, but these methods are only applicable to specific types of markings and are easily affected by changes in road scenes. Other researchers have used the reflectance intensity information of the point cloud to set thresholds and extracted road markings using methods such as MSTV, global threshold segmentation, and multi-threshold segmentation. However, these methods are easily affected by changes in point cloud intensity and have poor adaptability across different scenes.
[0004] Driven by the continuous advancements in deep learning technology, scholars have proposed various instance segmentation network models, which have been widely applied in natural scenes. However, road markings differ from natural images; multiple instances may appear in a single image, and these objects occupy only a small portion of the entire image, making them highly susceptible to background noise during feature fusion. Summary of the Invention
[0005] To address the aforementioned issues, the present invention aims to provide a road marking extraction and classification technology based on MLS point clouds. This technology addresses the limitations of existing technologies, which focus on extracting road markings from moving laser scanning (MLS) point clouds. These limitations are caused by factors such as complex road surface backgrounds, varying road marking scales, and blurred boundaries. The invention also aims to improve the ability to extract and classify road markings.
[0006] To achieve the above technical objectives, this application provides a road marking extraction and classification method based on MLS point clouds, comprising the following steps: Based on Mask R-CNN, a two-dimensional intensity feature image with road markings is input into the feature extraction network ResNet-50 to extract feature maps. After multi-scale feature fusion, the feature maps with multi-scale information are then processed by the Region Proposal Network (RPN) to generate region candidate boxes. The region candidate boxes are used as input to the RoIAlign layer. The features corresponding to each region candidate box in the feature map are extracted and input into the detection branch and the mask branch to classify and recognize road markings.
[0007] Preferably, in the process of acquiring the intensity feature image, the point cloud data of the road is processed by the cloth simulation algorithm CSF to extract the point cloud data containing only the road surface background and road markings; and the point cloud data is converted into a two-dimensional intensity feature image using the intensity information of the road surface point cloud.
[0008] Preferably, in the process of generating a two-dimensional intensity feature image, the inverse distance weighted interpolation algorithm (IDW) is used to generate the two-dimensional intensity feature image based on point cloud data.
[0009] Preferably, during the multi-scale feature fusion process, the 1×1 convolution at the horizontal connection of PAFPN is replaced with the SA shuffle attention module. By refining the channels and spatial dimensions, a shuffle attention feature pyramid is constructed to perform multi-scale feature fusion on the four feature maps of different scales extracted by the feature extraction network ResNet-50.
[0010] Preferably, in the process of classifying and recognizing road markings, in the detection branch, the road markings are classified by correcting the bounding box coordinates of each instance object in the feature map and predicting its classification score.
[0011] Preferably, in the process of classifying and recognizing road markings, in the mask branch, a point rendering module is introduced to upsample the coarse segmentation result output by the CNN network by 2 times. In the upsampled result, N points with a prediction probability close to 0.5 are selected, and the point-by-point features of these N points are calculated. A multilayer perceptron is used to predict the label of the point-by-point features to recognize the road markings. The point-by-point features include coarse-grained features and fine-grained features on the selected N points.
[0012] Preferably, in the process of acquiring fine-grained features, fine-grained features are extracted by shuffling the attention feature pyramid.
[0013] Preferably, in the process of acquiring coarse-grained features, a multi-scale context-aware module is constructed to capture multi-scale feature information as coarse-grained features. Each branch consists of two parallel paths, CNN and Transformer. Each CNN path has a 3×3 dilated convolution. The structure of each branch is exactly the same, except for the dilation rate of the dilated convolution.
[0014] Preferably, in the process of acquiring coarse-grained features, CNN and Transformer are combined in parallel to capture global contextual relationships and local spatial details. In the CNN path, features are encoded from local to global by gradually increasing the receptive field. In the Transformer path, global contextual information is acquired first, local details are gradually restored, and finally the features with the same resolution extracted from the two paths are fused.
[0015] This invention discloses a road marking extraction and classification system based on MLS point clouds, used to implement a road marking extraction and classification method, including: The data processing module is used to input a two-dimensional intensity feature image with road markings into the feature extraction network ResNet-50 to extract feature maps based on Mask R-CNN, then perform multi-scale feature fusion, and then use the fused feature map with multi-scale information through the Region Proposal Network RPN to generate region candidate boxes. The classification and recognition module is used to take the region candidate boxes as input to the RoIAlign layer, extract the features corresponding to each region candidate box in the feature map, and input them into the detection branch and mask branch to classify and recognize road markings.
[0016] The present invention discloses the following technical effects: This invention addresses the limitations of existing technologies that focus on extracting road markings from mobile laser scanning (MLS) point clouds, which are constrained by complex road backgrounds, varying road marking scales, and blurred boundaries. It improves the ability to extract and classify road markings, and the accurate extraction of road markings is of great significance in fields such as intelligent transportation systems, high-precision map construction, and autonomous driving. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1This is the intensity image generated using IDW interpolation as described in this invention; Figure 2 This is an overall flowchart of the model described in this invention; Figure 3 This is a diagram of the mixed-wash attention feature pyramid structure described in this invention; Figure 4 This is a structural diagram of the SA mixed washing attention module described in this invention; Figure 5 This is a structural diagram of the multi-scale context-aware module described in this invention; Figure 6 This is a structural diagram of the point rendering module described in this invention; Figure 7 This is a schematic diagram of the method described in this invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0020] like Figure 1-7 As shown, this invention proposes a road marking extraction and classification technique based on MLS point clouds, which specifically includes the following technical processes: This invention designs a hybrid attention feature pyramid, which enhances the model's ability to perceive road marking features by focusing on the instance object region, thereby suppressing background noise. The specific details are as follows: Since the primary objective of this invention is road markings, and the original MLS point cloud also includes non-ground point clouds such as trees, vehicles, streetlights, and buildings, it is necessary to separate the road markings from these non-ground point clouds to improve computational efficiency in subsequent studies. This invention uses the Cloth Simulation (CSF) algorithm to process the point cloud data, extracting point cloud data containing only the road surface background and road markings. Subsequently, the intensity information of the road surface point cloud is used to convert it into a 2D intensity feature image.
[0021] Existing research indicates that the elevation information of road point clouds has a limited impact on road marking extraction. Therefore, although projecting the road point cloud onto a 2Dxy plane causes some accuracy loss, it preserves the intensity information required for road marking extraction. This invention employs an inverse distance weighted interpolation (IDW) algorithm to generate a two-dimensional intensity feature image. First, a blank raster image is created based on the maximum and minimum values (Xmax, Ymax, Xmin, Ymin) of the point cloud coordinates in the horizontal and vertical directions. The width of the raster image is W = (Xmax - Xmin) / R, and the height is H = (Ymax - Ymin) / R, where R is the image resolution.
[0022] To improve the processing efficiency of point cloud projection and the quality of the generated feature images, this invention sets the resolution of the blank image to 0.05m. Then, IDW interpolation is used to calculate the intensity value of each grid cell, and the calculated intensity values are normalized to a grayscale space of 0-255, thereby generating an intensity feature image of the entire road point cloud. Finally, the exposure of the generated feature image is adjusted to increase the contrast between the road markings and the background. Figure 1 Example of a generated intensity feature image.
[0023] like Figure 2 As shown, the overall architecture of this invention is built on top of Mask R-CNN. First, the image is input into the ResNet-50 feature extraction network to extract feature maps at four different scales. Then, the extracted feature maps are input into the shuffled attention feature pyramid for multi-scale feature fusion. After obtaining the enhanced features, they are input into the Region Proposal Network (RPN) to generate candidate bounding boxes for regions that may contain targets. Since road markings differ from natural images, the default anchor box ratio is not suitable for extracting road markings, so the anchor box ratio needs to be adjusted in the RPN. Next, the candidate bounding boxes are used as input to the RoIAlign layer, which can extract the features corresponding to each candidate bounding box in the feature map. Finally, the extracted features are input into the detection branch and the mask branch. In the detection branch, the model corrects the bounding box coordinates of each instance object in the feature map and predicts its classification score. In the mask branch, the point rendering module takes the coarse-grained information obtained by the multi-scale context-aware module and the fine-grained information extracted by the shuffled attention feature pyramid as input, and performs adaptive iterative optimization to generate a more accurate instance segmentation mask.
[0024] In FPN, semantic information propagates along a top-down path. From lower-level feature maps to higher-level feature maps, information needs to pass through multiple network layers, increasing the difficulty of obtaining initial image information. PAFPN constructs a feature fusion path from the bottom to the top, shortening the distance that lower-level feature information travels through the network and effectively improving the localization accuracy of features at each layer. For example... Figure 3 As shown, the red dashed line represents the top-down path of traditional FPN, where low-level feature maps need to pass through many network layers to reach the top layer, resulting in severe attenuation of detailed information. PAFPN introduces the bottom-up path shown by the green dashed line, where low-level feature maps can be directly fused into lower-level feature maps and then passed layer by layer to the top layer, effectively preserving the detailed information in the low-level features.
[0025] Although PAFPN reduces information loss in the lower-level feature maps, the distribution and quantity of road markings and road surface background pixels in the image samples input to the feature extraction network are not uniform, especially in some sub-images where there are very few road marking pixels. Therefore, the features extracted by the backbone network contain a large amount of background noise, which cannot be effectively filtered out by simple 1×1 convolutions. PAFPN still faces the challenge of extracting effective discriminative information from noise, and some semantic information is also lost during the fusion of features at adjacent scales. Therefore, in the feature fusion stage, the network's ability to suppress background and enhance key target features is crucial. The SA shuffling attention module can reconstruct the input feature map from both channel and spatial dimensions, effectively promoting the fusion of feature information. This allows the network model to focus on global features relevant to the foreground target and suppress irrelevant features. Furthermore, the SA shuffling attention module has lower computational complexity and fewer parameters, enabling the network to achieve higher performance with lower computational cost.
[0026] Therefore, considering the importance of spatial features and feature channels that are closely related to road markings, this invention replaces the 1×1 convolution at the lateral connection of PAFPN with the SA shuffling attention module. By refining the channels and spatial dimensions, the loss of semantic information in the feature fusion process is reduced, and the model's ability to discriminate input features is enhanced.
[0027] like Figure 4 As shown, firstly, feature maps of C2 to C5 are obtained using a bottom-up backbone network. Where Ci, H, W, and Si represent the number of channels, spatial height and width, and stride, respectively. , Subsequently, SA first analyzed the feature maps of each layer. The features are divided into G groups along the channel dimension, i.e. , Next, along the channel dimension, the sub-feature Bk is divided into two branches, namely... In the Bk1 branch, channel statistics are generated using Global Average Pooling (GAP). Then use Generate corresponding channel attention : (1) In the formula, W1 and b1 are shapes of... The parameters, using This enhances feature channels relevant to road marking information while suppressing the activation of irrelevant channels. For example, within a specific region of the feature map, some channels represent road markings, while others represent the road surface background. In this case, the former will be assigned a higher weight, while the latter will be assigned a lower weight. In the Bk2 branch, Group Norm (GN) is used to obtain spatial statistics. Generate corresponding spatial attention : (2) In the formula, W2 and B2 are shapes of... Parameters, spatial attention It can accurately locate important information in the spatial position of the feature map, effectively amplify the feature differences between road markings and the road surface background, assign higher weights to road marking pixels and lower weights to road surface background pixels. Then, the "channel shuffling" operator is used to overlay and shuffle the feature maps after channel and spatial attention are connected, so as to realize the flow of feature information between different groups and enhance feature extraction and expression.
[0028] After Ci is compressed and refined by the SA shuffling attention module, it is fused with feature map Pi+1 to obtain feature map Pi. Both feature maps have 256 channels, the difference being their resolution. Next, the detail information in Ni and the semantic information enhanced by the SA shuffling attention module in Pi+1 are fused to generate feature map Ni+1 with multi-scale information.
[0029] This invention constructs a multi-scale context-aware module. Each branch of the multi-scale context-aware module uses a parallel combination of Transformer and CNN to capture global context information and multi-scale spatial features. The specific details are as follows: Road markings exhibit more significant scale variations in images than in natural scenes. Relying solely on single-scale feature extraction is insufficient to fully capture the feature information of road markings. Therefore, this invention constructs a parallel multi-branch structure to capture multi-scale feature information. For example... Figure 5As shown in (a), each branch consists of two parallel paths: a CNN and a Transformer. Each CNN path contains a 3×3 dilated convolution. The three branches have identical structures; the only difference lies in the dilation rate of the dilated convolution. Each branch is trained by sampling only its corresponding receptive field to alleviate the undersegmentation problem of single-scale networks. Besides lacking multi-scale information, road markings may also be affected by the background or other nearby road markings, leading to semantic ambiguity. This requires the network to not only possess multi-scale local detail information but also the ability to capture long-range features. Figure 5 As shown in (b), traditional convolutional neural networks typically capture features in an image through local receptive fields. In contrast, the Transformer, through its multi-head self-attention module (MSA), models global relationships across the entire image, assigning different attention weights to each element based on its relationship with other elements. This allows the model to dynamically focus on relevant information at all locations in the image, thus achieving global context modeling. By combining CNN and Transformer in parallel, the model can simultaneously capture global contextual relationships and local spatial details. In the CNN path, features are encoded from local to global by progressively increasing the receptive field; in the Transformer path, global contextual information is first acquired, followed by progressive recovery of local details. Finally, features with the same resolution extracted from both paths are fused.
[0030] This invention introduces an adaptive iteration of a point rendering module to optimize the boundary mask of road markings, as detailed below: Existing road marking extraction methods focus on restoring the overall outline and connectivity of road markings, neglecting the boundaries. The similarity between pixels at road marking boundaries leads to inconsistent model predictions and coarser boundary segmentation, especially for markings with complex shapes. Previous research has shown that in semantic segmentation, the pixels most prone to misclassification are typically at object edges. For example, Mask R-CNN usually predicts masks in a low-resolution 28×28 regular grid, repeatedly using convolution and pooling operations to increase feature density, and then upsampling to restore the image size to its original size to obtain semantic information. However, during upsampling, the model oversamples the internal regions of road markings and undersamples the boundary regions, resulting in significant boundary segmentation errors and limiting the accuracy of semantic segmentation. Achieving high-resolution instance segmentation requires calculating each pixel individually, inevitably leading to high computational costs. Therefore, a trade-off between computational power and high-resolution instance segmentation is necessary. To improve the resolution of the output mask while using lower computational power, this invention introduces a point rendering module into the mask branch of Mask R-CNN. The point rendering module combines image rendering concepts with upsampling in the semantic segmentation process. It adaptively samples a small number of points with ambiguous categories in the boundary region and then iteratively predicts their true categories. This allows for the output of high-resolution road marking masks using only a small amount of data, avoiding computation on all pixels and achieving high-resolution prediction with limited computing resources.
[0031] The structure of the point rendering module is as follows: Figure 6 As shown, the coarse segmentation result output by the CNN network is first upsampled by a factor of 2. Then, N points with prediction probabilities close to 0.5 are selected from the upsampled result, and the point-by-point features of these N points are calculated. The point-by-point features are constructed by combining two types of features: coarse-grained features and fine-grained features on the selected N points. Coarse-grained features are obtained from the original mask features, which mainly focus on the local features of the road marking contours. Based on the details of the texture information, the approximate boundary positions can be located. Fine-grained features are obtained from the FPN, which focuses on extracting global context information and can capture similar road marking shape information in other areas of the image. Similar shape features help distinguish road marking boundaries. Finally, a multilayer perceptron (MLP) is used to predict the labels of the above point-by-point features to obtain more accurate prediction results. The above process is iteratively executed, preserving boundary information as much as possible, until the desired spatial resolution is obtained through upsampling.
[0032] The main task of the point rendering module is to adaptively select a small number of points located near the edges of road markings. Different point selection strategies and the number of points are used during inference and training. During inference, an iterative subdivision strategy is chosen, with an input image resolution of R0=7 and an output image resolution of R=224. Each iteration selects N=282 points, and the final predicted number of points is... The point rendering module only needs to predict 282 × 4.25 points, about 15 times fewer than the 2242 points calculated directly. During training, a random sampling strategy is chosen, which achieves a better balance between model training efficiency and prediction accuracy compared to iterative subdivision. First, random sampling is performed from uniformly distributed points. (k>1) This is the oversampling ratio. Next, we will... The probability of selecting a road marking as a result of a rough interpolation prediction based on a sample of points is approximately 0.5. One point, This represents the importance sampling ratio. Then, continue selecting (1-) points from a uniformly distributed point set. N points. By appropriately setting the oversampling ratio And important sampling ratio To optimize the sampling strategy, the sampling points are concentrated in the uncertain area, while also taking into account the overall uniform distribution.
[0033] For the improved model, the multi-task loss L = Lcls + Lbox + Lmask + Lpoint, where Lcls, Lbox, Lmask, and Lpoint are the classification loss, bounding box loss, mask loss, and loss for N uncertain points, respectively. During backpropagation, Cross Entropy Loss is used to calculate Lcls, Lmask, and Lpoint, and L1 Loss is used to calculate Lbox. In the extraction of road markings, foreground pixels are much smaller than background pixels in road marking instances. When calculating Lmask, Cross Entropy Loss often evaluates each category in the image equally, leading the network model to tend towards segmenting the road surface and background regions. Therefore, Cross Entropy Loss cannot accurately describe the model error, resulting in poor model training performance. To address the same problem in medical image segmentation, Dice Loss is typically used to increase the loss weights of foreground target samples and ignore a large amount of background information to solve the problem of imbalanced positive and negative samples. The definition of Dice Loss is as follows: (3) X represents the pixel label of the real image, and Y represents the pixel category of the predicted image. It is approximated as the dot product of the pixels in the predicted image and the pixels in the true label image, and the dot products are added together. and Each is approximated by adding the pixels in their respective corresponding images.
[0034] While Dice Loss focuses more on foreground region extraction, using it alone can lead to training instability and difficulty in minimizing the loss. Therefore, this invention uses Cross Entropy Loss as the base loss function and introduces Dice Loss to construct a novel mask loss function, HLF (Hybrid loss functions), which combines the stability of Cross Entropy Loss with the advantage of Dice Loss being unaffected by sample imbalance. The definition of HLF is as follows: (4) (5) In the formula, N is the number of pixels, C is the number of pixels in each class (excluding the background), and yi,j is the true class of pixel i. If the true class of sample i is j, then yi,j equals 1, otherwise it equals 0. This represents the probability that the network predicts sample i belongs to class j. α and β are the weighting factors for LDice and LCE, respectively, and the optimal weight ratio will be determined through subsequent experiments.
[0035] This invention designs a shuffled attention feature pyramid. This module uses PAFPN as the baseline model and introduces an SA shuffled attention module at the lateral connections of PAFPN. This module can highlight features related to road markings during feature fusion, thereby suppressing the influence of background noise. Because there are many categories of road markings, and different categories have scale differences, the single-scale mask branch segmentation capability in the model is insufficient. Therefore, this invention constructs a multi-scale context-aware module, which consists of three parallel bidirectional semantic association modules. Each module utilizes the local receptive field of CNN and the global attention mechanism of Transformer to effectively model the semantic correlation and discriminative features between pixels to handle the under-segmentation problem caused by multi-scale road marking instances. Furthermore, due to the semantic ambiguity of road marking boundary pixels, the model struggles to segment accurately. To address this issue, this invention introduces a point rendering module to optimize road marking boundaries. The point rendering module selects only the most uncertain points in the predicted mask for individual prediction, while other pixels are directly interpolated, thereby improving the problems of boundary ambiguity and lack of detail.
[0036] In summary, this invention proposes a road marking extraction and classification technique based on MLS point clouds for instance segmentation of road markings from MLS point clouds. First, this invention proposes a shuffled attention feature pyramid, which enhances the feature representation capability of road markings by learning feature dependencies in channels and spatial dimensions. Second, a multi-scale context-aware module is constructed, which improves the model's reasoning ability regarding contextual information by aggregating semantic information between different scales and distant regions. Finally, a point rendering module is introduced to adaptively select key points for prediction, thereby improving the segmentation effect of road marking boundaries. Experiments were conducted on the constructed point cloud road marking dataset for validation. The results show that compared with existing instance segmentation models, this invention outperforms existing technologies in all evaluation metrics, demonstrating significant performance advantages.
[0037] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0038] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0039] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A road marking extraction and classification method based on MLS point clouds, characterized in that, Includes the following steps: Based on Mask R-CNN, a two-dimensional intensity feature image with road markings is input into the feature extraction network ResNet-50 to extract feature maps. Then, multi-scale feature fusion is performed, and the fused feature map with multi-scale information is passed through the Region Proposal Network (RPN) to generate region candidate boxes. The candidate bounding boxes of the regions are used as input to the RoIAlign layer. The features corresponding to each candidate bounding box of the region in the feature map are extracted and input into the detection branch and the mask branch to classify and identify the road markings. In the process of multi-scale feature fusion, the 1×1 convolution at the horizontal connection of PAFPN is replaced with the SA shuffle attention module. By refining the channel and spatial dimensions, a shuffle attention feature pyramid is constructed to perform multi-scale feature fusion on the four feature maps of different scales extracted by the feature extraction network ResNet-50. In the process of classifying and recognizing road markings, in the mask branch, a point rendering module is introduced to upsample the coarse segmentation result output by the CNN network by 2 times. In the upsampled result, N points with a prediction probability close to 0.5 are selected, and the point-by-point features of these N points are calculated. The label of the point-by-point features is predicted using a multilayer perceptron to recognize the road markings. The point-by-point features include coarse-grained features and fine-grained features on the selected N points. In the process of acquiring fine-grained features, the fine-grained features are extracted through the shuffled attention feature pyramid; In the process of acquiring coarse-grained features, a multi-scale context-aware module is constructed to capture multi-scale feature information as the coarse-grained features. Each branch of the multi-scale context-aware module consists of two parallel paths: CNN and Transformer. Each CNN path has a 3×3 dilated convolution. The structure of each branch is exactly the same, except for the dilation rate of the dilated convolution.
2. The road marking extraction and classification method based on MLS point clouds according to claim 1, characterized in that: In the process of acquiring intensity feature images, the point cloud data of the road is processed by the cloth simulation algorithm CSF to extract point cloud data containing only the road surface background and road markings; and the point cloud data is converted into a two-dimensional intensity feature image using the intensity information of the road surface point cloud.
3. The road marking extraction and classification method based on MLS point clouds according to claim 2, characterized in that: In the process of generating a two-dimensional intensity feature image, the inverse distance weighted interpolation algorithm (IDW) is used to generate the two-dimensional intensity feature image based on the point cloud data.
4. The road marking extraction and classification method based on MLS point clouds according to claim 1, characterized in that: In the process of classifying and recognizing road markings, in the detection branch, the road markings are classified by correcting the bounding box coordinates of each instance object in the feature map and predicting its classification score.
5. The road marking extraction and classification method based on MLS point clouds according to claim 1, characterized in that: In the process of acquiring coarse-grained features, CNN and Transformer are combined in parallel to capture both global contextual relationships and local spatial details. In the CNN path, features are encoded from local to global by gradually increasing the receptive field. In the Transformer path, global contextual information is first acquired, followed by gradual recovery of local details, and finally the features with the same resolution extracted from the two paths are fused.
6. A road marking extraction and classification system based on MLS point clouds, used to execute the road marking extraction and classification method based on MLS point clouds as described in claim 1, characterized in that, include: The data processing module is used to input a two-dimensional intensity feature image with road markings into the feature extraction network ResNet-50 based on Mask R-CNN to extract feature maps, and after multi-scale feature fusion, the feature maps with multi-scale information are generated through the Region Proposal Network (RPN) to generate region candidate boxes. The classification and recognition module is used to take the candidate region boxes as input to the RoIAlign layer, extract the features corresponding to each candidate region box in the feature map, and input them into the detection branch and mask branch to classify and recognize the road markings.