Road point cloud data semantic segmentation method based on sparse perception three-dimensional fusion network
Through the preprocessing and feature extraction methods of sparse-perceptible three-dimensional fusion network, the waste of computing resources and insufficient accuracy caused by sparseness and irregularity of point cloud data is solved, and efficient point cloud semantic segmentation is achieved, which is suitable for scenarios such as autonomous driving and urban modeling.
Patent Information
- Application Number
- CN202510590714.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing point cloud data processing methods have problems of waste of computing resources and insufficient precision in terms of sparsity and irregularity, making it difficult to efficiently perform semantic segmentation.
Using a method based on sparse perception three-dimensional fusion network, point cloud data is preprocessed, sparse convolutional networks with encoder and decoder structures are built, and combined with the improved Focal Loss loss function for training, multi-scale features are extracted and context information is enhanced.
It significantly improves the segmentation accuracy and processing efficiency of point cloud data, especially in complex backgrounds and multi-scale target scenarios, and is suitable for applications such as autonomous driving and urban modeling.
Smart Images

Figure CN120451559A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of point cloud data processing and computer vision, and specifically relates to a semantic segmentation solution for road point cloud data based on a sparse perception three-dimensional fusion network. Background Art
[0002] With the rapid development of 3D scanning technologies such as LiDAR and depth cameras, point cloud data has been widely used in fields such as autonomous driving, urban modeling, environmental monitoring, and construction engineering. Point cloud data can accurately describe the 3D structure of objects and provide rich spatial information. However, due to its sparsity, irregularity, and large scale, the processing and analysis of point cloud data still faces many challenges.
[0003] In point cloud data processing, the goal of semantic segmentation is to assign each point in the point cloud to a corresponding category, such as ground, building, vegetation, or road. Semantic segmentation is a core issue in 3D computer vision, particularly when processing sparse and irregular point cloud data. Efficiently extracting semantically meaningful features and accurately classifying them is a key research focus. Traditional regular grid-based processing methods, such as voxelization and gridding, often suffer from wasted computational resources and insufficient accuracy due to the sparsity of point cloud data.
[0004] In recent years, deep learning techniques have become a mainstream approach to solving semantic segmentation problems with point cloud data. Three-dimensional network architectures based on convolutional neural networks (CNNs), in particular, have achieved remarkable results. However, the sparsity and irregularity of point cloud data remain two major challenges for deep learning methods. Traditional convolutional neural networks cannot directly process this data. Furthermore, preprocessing and feature conversion of point cloud data typically require extensive computation, further reducing computational efficiency. Therefore, improving the efficiency of point cloud data processing while maintaining segmentation accuracy has become a pressing issue. Summary of the Invention
[0005] To achieve the above objectives, the present invention provides a method and system for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network, which performs semantic segmentation on sparse point cloud data through an efficient three-dimensional convolutional neural network to improve the accuracy of segmentation and processing efficiency.
[0006] In order to solve the above technical problems, the technical solution of the present invention provides a road point cloud data semantic segmentation method based on a sparse perception three-dimensional fusion network, which includes the following steps: Pre-processing of raw point cloud data includes filtering out background points far from the road, segmenting point cloud data by road direction and performing slice enhancement, and simulating occlusion enhancement of pole-shaped assets. Constructing a sparse-aware 3D fusion network model, the sparse-aware 3D fusion network model comprising an encoder and a decoder, the encoder extracting multi-scale features through multiple sets of sparse convolutions, the decoder restoring spatial resolution through sparse voxel upsampling and performing skip connections on encoder features; Perform model training, perform semantic classification on point cloud data through the trained sparse perception 3D fusion network model, and output category prediction results.
[0007] Furthermore, the point cloud data is segmented according to the road direction, and is implemented by segmenting the filtered point cloud data according to the road direction to form point cloud slices along the road direction.
[0008] Moreover, when enhancing the segmented slices, conventional enhancement and hybrid enhancement operations are performed. The hybrid enhancement operation includes randomly mixing slices containing foreground objects to generate composite samples, and injecting foreground point clouds into slices without foreground objects after downsampling to balance the data ratio.
[0009] Furthermore, the pole-shaped asset occlusion enhancement simulation is achieved by constructing a cone-shaped mask to simulate the occlusion scene and adjust the color and intensity.
[0010] Moreover, the feature matrix of the input sparse perception 3D fusion network model is generated based on the point cloud data, including spatial coordinates, reflection intensity and color information.
[0011] Moreover, the improved Focal Loss loss function is used for model training; The improved Focal Loss loss function adopts the following formula:
[0012] Among them, n is the number of points in the point cloud, N is the total number of points in the point cloud, i is the number of categories, and ns is the total number of categories. is the focus factor, is the weight factor, is the predicted probability, is the smoothed label.
[0013] Moreover, the weight factor of the improved Focal Loss loss function is determined according to the ratio of the median of the category frequency to the frequency of a specific category.
[0014] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network as described above is implemented.
[0015] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network.
[0016] On the other hand, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned road point cloud data semantic segmentation method based on the sparse perception three-dimensional fusion network.
[0017] The present invention has the following positive effects: (1) Strong multi-scale feature extraction capability, significantly improving the segmentation accuracy of complex point cloud targets; This paper introduces multi-scale feature extraction modules at multiple encoder stages. Using a parallel multi-path residual architecture, this method performs multi-scale perceptual processing on point cloud features, effectively capturing geometric and semantic information at different spatial scales. Compared to traditional single-scale convolutional architectures, this model exhibits enhanced detail sensitivity and spatial structure representation, achieving higher segmentation accuracy in point cloud scenes with complex backgrounds and multi-scale objects.
[0018] (2) The context information enhancement mechanism improves the model’s ability to segment boundary areas and small objects; This paper designs a specialized feature extraction module based on the above context, leveraging multi-scale convolution and integrating global context information to enhance the model's ability to perceive class boundaries and small-sized objects in point cloud data. This approach is particularly useful in scenarios where point cloud data contains blurred boundaries and fragmented distributions, significantly improving the segmentation accuracy of small objects and those with complex boundaries.
[0019] (3) The structural design is flexible and adaptable, making it easy to migrate to various point cloud segmentation tasks; This method builds a backbone structure based on a U-shaped network, integrating high-resolution and high-semantic features through a skip connection mechanism. The overall architecture is clear and modular, making it easy to scale and migrate across different resolutions and task scenarios. This method is widely applicable to point cloud segmentation tasks such as autonomous driving, urban modeling, and environmental monitoring.
[0020] (4) Maintain the original point cloud information, avoid information loss, and improve segmentation accuracy; During the point cloud feature extraction and fusion process, the present invention retains multi-scale and multi-level information flows, avoiding the loss of point cloud information due to forced downsampling. Compared with traditional models, it can more comprehensively preserve the geometric, structural and semantic features of the original point cloud, thereby effectively improving the final point cloud segmentation accuracy.
[0021] In summary, this invention provides a high-precision, robust, and highly generalizable point cloud segmentation method, particularly suitable for the automatic segmentation of multi-category, multi-scale objects in complex scenes. It also provides a solid foundation for subsequent point cloud data analysis and application decision-making. This invention is applicable to a variety of application scenarios, including autonomous driving, urban modeling, and environmental monitoring, and effectively addresses the issues of insufficient sparsity handling and high computational complexity inherent in existing point cloud segmentation technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To further clarify the technical solution of the present invention, the accompanying drawings are briefly described below. It should be noted that these drawings represent only a few embodiments of the present invention. Those skilled in the relevant art can readily deduce other possible drawings based on these drawings without requiring additional creative effort.
[0023] Figure 1 This is a flowchart of semantic segmentation of road point cloud data based on a sparse perception 3D fusion network according to an embodiment of the present invention.
[0024] Figure 2 This is a network structure diagram based on the improved U-Net model in an embodiment of the present invention.
[0025] Figure 3 This is a visualization effect diagram of the point cloud segmentation result of an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.
[0027] The present invention mainly proposes a new three-dimensional fusion network architecture based on sparse convolution. This architecture avoids the problem of computational redundancy in traditional networks by introducing sparse convolution operations, and can efficiently process sparse point cloud data. Through the design of the encoder-decoder structure, the network can effectively extract and restore spatial features in point cloud data, thereby realizing efficient point cloud semantic segmentation. This design greatly reduces the amount of calculation while maintaining high precision, and is particularly suitable for processing large-scale sparse point cloud data. Although existing point cloud segmentation methods have achieved results in some application scenarios, there are still certain challenges in complex environments, such as occlusion, noise and differences in the morphology of different objects in point cloud data. The three-dimensional fusion network based on sparse convolution provides an effective technical solution to these problems by optimizing the network architecture and feature extraction method, and opens up a new path for improving the accuracy and efficiency of point cloud data in practical applications.
[0028] Example 1: Highway Scene Point Cloud Segmentation See also Figure 1 In an embodiment of the present invention, a method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network is proposed, comprising the following steps: Step 1: Point cloud preprocessing In practice, raw road scene point cloud data can be collected through 3D sensing equipment such as vehicle-mounted laser radar (LiDAR) and depth cameras. This provides the fundamental input for the entire system. Point cloud data contains rich information such as spatial geometry, distance, and intensity, which is used to describe the 3D structure of the road and its associated objects.
[0029] For the 3D point cloud data of road assets in highway scenarios, we use a variety of processing methods, including conventional and special enhancements, to improve the model's robustness and generalization capabilities. The details are as follows: Step 1.1: Filtering background points away from the road Color and intensity filters are used to remove background points from the original point cloud data. First, color features are used to remove background points in non-road areas. Intensity information is then used to further filter out noise points, ensuring that the retained point cloud is primarily focused on roads and their associated assets.
[0030] Step 1.2: Segment point cloud data by road direction The filtered point cloud data is segmented according to the road direction to form point cloud slices along the road. This strategy not only controls the volume of each segment of data, facilitating model processing, but also preserves the structural and semantic continuity of the road.
[0031] Step 1.3: General enhancement processing It is recommended to perform routine enhancement operations on each slice, including: Randomly rotate 0–360 degrees to improve the model’s rotation invariance; Randomly scale by 0.95–1.05 times to enhance scale adaptability; Normalize the RGB channels to the [0,1] interval to unify the color distribution; Apply sparse quantization technology to compress redundant points, retain only key structural points, and reduce data noise.
[0032] Step 1.4: Hybrid Enhancement Strategy In practice, for slices containing foreground objects (such as poles and signboards), objects from other slices can be randomly mixed to combine foreground objects with different backgrounds, enhancing data diversity. For slices without foreground objects, downsampling is used to reduce the number of points, and point clouds containing foreground objects are introduced to supplement the background to maintain a balanced foreground-background ratio.
[0033] Step 1.5: Pole-Shaped Asset Occlusion Enhancement Simulation For target pole-shaped assets (such as camera poles and streetlight poles), the top of the pole or a location near it is selected as the vertex of the cone model. The cone's axis is primarily oriented vertically upward, but can be adjusted to accommodate the growth orientation of actual obstructing objects (such as trees). Regarding the cone mask parameter settings, the half-apex angle (θ) is typically set between 5° and 10° to reflect the scale characteristics of pole-shaped assets on highways, simulating the diffusion trend of real-world obstructions. The cone length (L) is generally controlled between 1 and 5 meters to cover the typical obstruction distance around the pole-shaped asset. Within the spatial range defined by the cone mask, the foreground points (i.e., the pole point cloud) are density-downsampled to simulate local point cloud sparseness caused by occlusion. Furthermore, color and reflection intensity perturbations are superimposed to reflect brightness, material variations, and sensor noise effects under natural occlusion. After processing, the occlusion-enhanced point cloud and the original data are integrated back into the training set to enhance the model's ability to learn and recognize local occlusion features, effectively improving the segmentation accuracy of occluded objects in complex environments.
[0034] This method takes into account both point cloud integrity protection and occlusion authenticity. Through cone mask modeling, dynamic screening mechanism and composite enhancement strategy, it effectively alleviates the problem of feature loss caused by over-processing in traditional enhancement methods, and has achieved significant performance improvement in experimental results.
[0035] Step 2: Set up a three-dimensional deep learning network model for point cloud asset detection based on U-shaped structure sparse convolution. The three-dimensional deep learning network model adopts an encoder-decoder structure and is a sparse perception three-dimensional fusion network.
[0036] See also Figure 2 The three-dimensional deep learning network model for point cloud asset detection provided by the embodiment is implemented as follows: 1) This method aims to detect different types of road assets by encoding the input point cloud data into seven-dimensional features, including spatial coordinates (x, y, z), reflection intensity, and color information (R, G, B). This can be implemented to form an N × 7 input feature matrix, where N is the number of point clouds input in a batch. 2) On the encoder side, the feature matrix is first input into the basic module for preliminary feature extraction, and a 32-channel feature map is output to maintain spatial sparsity. The basic module consists of two sets of sparse convolutions arranged in sequence, each of which consists of a sparse convolution and a batch normalization layer convolution. Then, the features output by the basic module are fed into a four-layer feature extraction module in sequence. Each layer of the feature extraction module consists of four sets of sparse convolutions set in sequence. Each set of sparse convolution operations increases the feature dimension to 64, 128, and 256 dimensions respectively, and a downsampling operation with a stride of 2 is performed in the last set of convolutions to reduce the spatial resolution. In the feature extraction module, batch normalization is performed after each convolution operation, and sparse convolution is used to maintain feature sparsity while enhancing feature fusion capabilities; Among them, the specific working process of each layer feature extraction module is as follows: The first set of sparse convolutions consists of a sparse convolution and a batch normalization layer convolution, and the output of the first set of sparse convolutions is connected to the second set of sparse convolutions; The second group of sparse convolutions consists of the first sparse convolution and the first batch normalization layer convolution, the second sparse convolution and the second batch normalization layer convolution, which are arranged in sequence. The output of the first group of sparse convolutions and the output of the second group of sparse convolutions are added and input into the third group of sparse convolutions; The third group of sparse convolutions consists of the first sparse convolution and the first batch normalization layer convolution, the second sparse convolution and the second batch normalization layer convolution, and the output of the third group of sparse convolutions is connected to the fourth group of sparse convolutions. The fourth group of sparse convolutions consists of a sparse convolution and a batch normalization layer convolution. The output of the third group of sparse convolutions and the output of the fourth group of sparse convolutions are added together as the output of the feature extraction module of this layer.
[0037] 3) On the decoder side, the high-dimensional features output by the encoder are first fed into a four-layer upsampling module. Each upsampling module includes two layers of sparse voxel upsampling units, and the image is processed through the sparse voxel upsampling units eight times in sequence to gradually restore the spatial resolution. A jump connection is made between each upsampling module and the features output by the feature extraction module of the corresponding layer on the encoder side, achieving the fusion of deep features and shallow detailed features. Each layer of sparse voxel upsampling unit consists of four groups of sparse voxel convolution upsampling arranged in sequence. The specific working process is as follows: The first set of sparse voxel convolution upsampling consists of a sparse voxel convolution and a batch normalization layer convolution. The output of the first set of sparse voxel convolution upsampling is connected to the second set of sparse voxel convolution upsampling; The second group of sparse voxel convolution upsampling consists of the first sparse voxel convolution and the first batch normalization layer convolution, the second sparse voxel convolution and the second batch normalization layer convolution, and the output of the first group of sparse voxel convolution upsampling and the output of the second group of sparse voxel convolution upsampling are added and input into the third group of sparse voxel convolution upsampling; The third group of sparse voxel convolution upsampling consists of the first sparse voxel convolution and the first batch normalization layer convolution, the second sparse voxel convolution and the second batch normalization layer convolution, and the output of the third group of sparse voxel convolution upsampling is connected to the fourth group of sparse voxel convolution upsampling; The fourth group of sparse voxel convolution upsampling consists of a sparse voxel convolution and a batch normalization layer convolution. The output of the third group of sparse voxel convolution upsampling and the output of the fourth group of sparse voxel convolution upsampling are added together as the output of the upsampling module of this layer.
[0038] Then, after the upsampling module completes feature recovery, the category prediction results of each point are output through the last set of convolutional layers to obtain the final point cloud highway asset detection results.
[0039] Step 3: Set a category-weighted loss function for model training to solve the problem of uneven distribution of object categories in remote sensing images and further enhance the learning ability of the sparse perception 3D fusion network model for minority class targets.
[0040] The embodiment improves on the basis of the cross entropy loss function Focal Loss and introduces Adjustment factor, the improved Focal Loss formula is as follows:
[0041] Among them, n is the number of points in the point cloud, N is the total number of points in the point cloud, i is the number of categories, and ns is the total number of categories. is the focus factor, is the predicted probability, is the smoothed label.
[0042] Here It is a category The weight factor is adjusted inversely according to the frequency of the class in the dataset to alleviate class imbalance. The value of is calculated by the median frequency balance method, that is, the median of all category frequencies is divided by the specific category The frequency of each category is used to determine the weight of each category. This enhances the model's ability to learn minority class objects. The deep learning model is trained based on the training sample set to obtain a trained point cloud classification model.
[0043] Step 4: Use the network model trained in step 3 to evaluate the segmentation accuracy.
[0044] In order to facilitate understanding of the technical effects of the present invention, the following is a comparison of the point cloud segmentation accuracy indicators of different asset categories after adding data enhancement technology and introducing sparse voxel convolution and comparing with the baseline model as shown in Table 1. The visualization effect of the point cloud segmentation results can be seen in Figure 3 ,This figure shows the result of superimposing the extracted road asset point cloud on the ,original mobile laser scanning point cloud, which qualitatively shows the ,global perspective of the experimental results and demonstrates the ,accuracy of the extraction of this method.
[0045] Table 1
[0046] 2. Semantic Segmentation System for Road Point Cloud Data Based on Sparse Perception 3D Fusion Network The present invention provides a road point cloud data semantic segmentation system based on a sparse perception 3D fusion network, which is used to implement the above-mentioned point cloud segmentation method. The system includes the following modules: Module 1: Point cloud data acquisition and preprocessing module This module is used to collect and preprocess raw point cloud data, improve point cloud data quality, and ensure the effectiveness and robustness of subsequent model input. The module includes the following parts: Data Acquisition: This part uses 3D sensing equipment such as vehicle-mounted laser radar (LiDAR) and depth cameras to collect raw road scene point cloud data. This part provides the basic input for the entire system. The point cloud data contains rich information such as spatial geometry, distance, and intensity, which is used to describe the 3D structure of the road and its associated objects.
[0047] Point cloud cleaning and segmentation: The raw point cloud data is first filtered for background, using a color filter to remove background points in non-target areas. Intensity thresholding is then used to further remove noise points, ensuring the point cloud data is focused on the road and asset areas. The point cloud is then sliced and segmented along the road's direction to create structurally continuous and appropriately sized processing units.
[0048] General enhancement part: Perform general enhancement operations on the point cloud based on the slice, including: Point cloud slices are randomly rotated in the range of 0–360 degrees; Randomly scale by 0.95–1.05; Normalize the RGB values to the [0,1] interval; Implement sparse quantization operations to reduce the density of redundant points and increase the proportion of effective information.
[0049] Hybrid enhancement: For slices containing foreground assets (such as poles and traffic signs), probabilistic mixing is performed to generate composite samples with foreground diversity and occlusion changes. For slices without foreground, foreground points are injected after downsampling to balance the foreground-background data ratio.
[0050] Pole-Shaped Asset Occlusion Simulation: This section addresses the problem of slender objects like camera poles and emergency lights often being obscured by trees in real-world environments. This section simulates an occlusion environment by setting a cone-shaped mask, downsampling the target point cloud within its range, and perturbing its color and intensity information to simulate lighting conditions and enhance the model's robustness against occluded objects.
[0051] Module 2: U-shaped fully convolutional neural network model design module based on multi-scale feature extraction and context information enhancement Module 3: Category Weighted Loss Function Design Module This module is used to design a category-weighted loss function to solve the problem of uneven distribution of object categories in remote sensing images and further enhance the model's ability to learn minority class targets.
[0052] Module 4: Automatic classification and accuracy assessment module This module is used to apply the image classification model trained in the third module to the test image, complete the automatic object recognition and classification task of the image, and evaluate the classification accuracy.
[0053] Through the present invention, the system can realize automatic classification and recognition of vehicle-mounted laser radar three-dimensional point cloud data, significantly improve the efficiency and accuracy of point cloud data processing, and has good robustness and adaptability.
[0054] In specific implementations, those skilled in the art may use software technology to automate the above process. Accordingly, providing a semantic segmentation solution for road point cloud data based on a sparse-perception 3D fusion network, including a computer or server, and executing the above process on the computer or server to perform semantic segmentation of road point cloud data based on a sparse-perception 3D fusion network, should also fall within the scope of protection of the present invention.
[0055] In another embodiment, an electronic device is provided, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned road point cloud data semantic segmentation method based on a sparse perception three-dimensional fusion network.
[0056] In another embodiment, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned road point cloud data semantic segmentation method based on sparse perception three-dimensional fusion network.
[0057] In another embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network is implemented.
[0058] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of these modules may be selected to achieve the objectives of the present embodiments based on practical needs. Those skilled in the art will be able to understand and implement these embodiments without inventive effort. Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a required general-purpose hardware platform, or, of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or portions thereof. The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A semantic segmentation method for road point cloud data based on sparse perception 3D fusion network, characterized by: Including the following process, Pre-processing of raw point cloud data includes filtering out background points far from the road, segmenting point cloud data by road direction and performing slice enhancement, and simulating occlusion enhancement of pole-shaped assets. Constructing a sparse-aware 3D fusion network model, the sparse-aware 3D fusion network model comprising an encoder and a decoder, the encoder extracting multi-scale features through multiple sets of sparse convolutions, the decoder restoring spatial resolution through sparse voxel upsampling and performing skip connections on encoder features; Perform model training, perform semantic classification on point cloud data through the trained sparse perception 3D fusion network model, and output category prediction results.
2. The method according to claim 1, wherein: The point cloud data is segmented according to the road direction, and is implemented by segmenting the filtered point cloud data according to the road direction to form point cloud slices along the road direction.
3. The method according to claim 1, wherein: When enhancing the segmented slices, conventional enhancement and hybrid enhancement operations are performed. The hybrid enhancement operation includes randomly mixing slices containing foreground objects to generate composite samples, and injecting foreground point clouds into slices without foreground objects after downsampling to balance the data ratio.
4. The method according to claim 1, wherein: The enhanced simulation of rod-shaped asset occlusion is achieved by constructing a cone-shaped mask, simulating the occlusion scene, and adjusting the color and intensity.
5. The method according to claim 1, wherein: The feature matrix of the input sparse perception 3D fusion network model is generated based on the point cloud data, including spatial coordinates, reflection intensity and color information.
6. The method according to claim 1, wherein: Use the improved Focal Loss loss function for model training; The improved Focal Loss loss function adopts the following formula: Among them, n is the number of points in the point cloud, N is the total number of points in the point cloud, i is the number of categories, and ns is the total number of categories. is the focus factor, is the weight factor, is the predicted probability, is the smoothed label.
7. The method according to claim 6, wherein: The weight factor of the improved Focal Loss loss function is determined according to the ratio of the median of the category frequency to the frequency of a specific category.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network as described in any one of claims 1 to 7 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network as described in any one of claims 1 to 7 is implemented.
10. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, the method for semantic segmentation of road point cloud data based on a sparse perception three-dimensional fusion network as described in any one of claims 1 to 7 is implemented.