Point cloud real-time semantic segmentation method based on distance view

By reconstructing the point cloud data from geometric and reflection intensity, combining residual attention mechanism and multi-scale feature fusion, the problem of low accuracy of the point cloud semantic segmentation method based on distance view is solved, and higher segmentation accuracy and robustness are achieved.

CN120388177APending Publication Date: 2025-07-29CHENGDU UNIV OF INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510516631.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing point cloud semantic segmentation method based on distance view is problematic with low accuracy, mainly due to the distortion, overlapping and hollow pixel phenomena of original distance image features, unstable reflection intensity, and insufficient utilization of multi-stage point cloud features, resulting in unclear boundaries and difficult to balance high and low frequency features.

Method used

Point cloud data is preprocessed through geometric vision reconstruction and reflection intensity reconstruction modules, integrating spatial features and reflection intensity features, and fusion of residual attention mechanism and adaptive multi-scale feature fusion, extracting and updating two-dimensional distance image features, and semantic label prediction.

Benefits of technology

It improves the accuracy and robustness of point cloud semantic segmentation, enhances the ability of network to learn subtle features, balances multi-scale feature fusion, solves the problem of intensity instability in edge areas, and improves network performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388177A_ABST
    Figure CN120388177A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud real-time semantic segmentation method based on a distance view, and relates to the technical field of image analysis, and the method comprises the steps: carrying out the feature fusion of the spatial features and reflection intensity features of discrete point clouds in a plurality of projection spaces after reflection intensity reconstruction, carrying out the global feature extraction of the point cloud data after feature fusion, and carrying out the segmentation of the point clouds. Performing feature dimension compression on the global features to obtain distance view features; and projecting the distance view features to a two-dimensional distance image, performing two-dimensional distance image feature extraction on the two-dimensional distance image, updating the point cloud features and the projection space according to the extracted two-dimensional distance image features, and obtaining a new two-dimensional distance image by using the updated point cloud features and the projection space. Performing view feature updating on the new two-dimensional distance image through a residual attention mechanism; and predicting a semantic tag by using adaptive multi-scale feature fusion features. According to the method, the semantic segmentation accuracy and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image analysis, and particularly to a method for real-time semantic segmentation of point clouds based on a distance view. Background Art

[0002] Currently, lidar point cloud semantic segmentation mainly includes methods based on points, voxels, multi-source data fusion, and distance views. Compared with methods based on points, voxels, and multi-source data fusion, the method based on a distance view shows more significant advantages in terms of accuracy and computational efficiency, and has become a more ideal choice in the current situation.

[0003] However, although the method based on a distance view has real-time superiority and has been widely applied, its accuracy is significantly lower than that of the winners of other methods. The main reasons are as follows: 1. There are phenomena such as feature distortion, overlap, and void pixels in the original distance image; these phenomena are mainly manifested as the stretching or compression of feature information in some regions of the image, resulting in the destruction of the authenticity of the geometric structure; the overlap phenomenon is that too many point clouds are projected onto the same region, causing the boundaries of different objects to be blurred and making it difficult to accurately distinguish; the appearance of void pixels is due to data offset and even loss caused by precision loss. 2. The reflection intensity information is unstable and is affected by the distance between the sensor and the target object; the laser beam of the radar will attenuate during propagation, and various factors such as the material properties and reflectivity of the target surface will interfere with the reflection intensity, especially for the edge point clouds at a relatively long distance. 3. The current methods based on a distance view do not make full use of multi-stage point cloud features; the boundary of the feature map obtained based on a distance view is not clear, and traditional multi-scale feature fusion cannot effectively balance high and low frequency features.

[0004] Therefore, the applicant of the present invention has developed a method for real-time semantic segmentation of point clouds based on a distance view to solve the above problems. Summary of the Invention

[0005] The present invention proposes a method for real-time semantic segmentation of point clouds based on a distance view to solve the problem that the existing methods based on a distance view cannot effectively utilize point cloud feature information, resulting in low accuracy of semantic segmentation.

[0006] The present invention achieves the above object through the following technical solutions:

[0007] A method for real-time semantic segmentation of point clouds based on a distance view according to the present invention includes:

[0008] Obtain the initial point cloud data of the scene, where the initial point cloud data includes the initial three-dimensional spatial coordinates and reflection intensity of the point cloud;

[0009] Perform geometric field reconstruction on the initial point cloud data to obtain several divided projection spaces. Perform reflection intensity reconstruction on the point cloud data in the projection spaces, fuse the spatial features and reflection intensity features of the discrete point clouds in the several projection spaces after reflection intensity reconstruction, extract global features from the point cloud data after feature fusion, and compress the feature dimensions of the global features to obtain distance view features;

[0010] Project the distance view features onto a two-dimensional distance image, then extract two-dimensional distance image features from the two-dimensional distance image, update the point cloud features and projection spaces according to the extracted two-dimensional distance image features, then use the updated point cloud features and projection spaces to obtain a new two-dimensional distance image, and update the view features of the new two-dimensional distance image through a residual attention mechanism;

[0011] Perform adaptive multi-scale feature fusion on the two-dimensional distance image after view feature update and output the predicted semantic labels.

[0012] Furthermore, the initial point cloud data is , where I is the reflection intensity, and x, y, z are the initial three-dimensional spatial coordinates of the point cloud. The calculation formulas for x, y, and z are as follows:

[0013] ;

[0014] ;

[0015] ;

[0016] ;

[0017] is the time difference between the laser beam emission and reception, c is the speed of light, then d represents the distance between the acquisition device and the object, is the pitch angle of the laser beam relative to the plane, is the azimuth angle of the laser beam relative to the due front.

[0018] Furthermore, performing geometric field reconstruction on the initial point cloud data includes:

[0019] ;

[0020] ;

[0021] Among them, the is geometric field reconstruction, and the (u n , v n)(W, H) represent the abscissa and ordinate of the point cloud mapped to the distance image respectively, and (W, H) represent the width and height of the distance image respectively. represents the depression angle of the lidar acquisition device converted to radians, f represents the field of view angle of the lidar acquisition device converted to radians, and ∆ represents the offset in the vertical direction during the point cloud projection process to address the accuracy reduction problem caused by the laser beams not having the same starting point. represents the field of view angle in the vertical direction of the acquisition device, Envs represents the height of the device itself, as well as information such as the relative position with the ground. and represent the depression angle and elevation angle of the corresponding laser beam acquisition device respectively. is the K sub-regions divided from the initial point cloud data, where , R represents the real number space, n is the number of point clouds in each projection space, and c is the feature dimension of each point cloud.

[0022] Further, perform reflection intensity reconstruction on the point cloud data in the projection space, including performing mean pooling processing on the emission intensity of the point cloud data in the projection space.

[0023] Further, perform feature fusion on the spatial features and reflection intensity features of the discrete point clouds in several projection spaces after reflection intensity reconstruction, including:

[0024] Consider all points in each projection space as a set, and each point cloud has 4 basic attributes (x, y, z, I);

[0025] Encode the 4 basic attributes of each point cloud into 10 mutually related basic feature attributes P k n*4 -> P k n*10 ;

[0026] Among them, the basic feature attributes include the initial three-dimensional space coordinates and reflection intensity of the point cloud (x, y, z, I), the Manhattan distance between each point in the projection space and the virtual center of the set, the three-dimensional space coordinate vector difference, the reflection intensity difference, and the depth information of the point cloud.

[0027] Further, use a multi-layer perceptron to extract global features from the point cloud data after feature fusion.

[0028] Further, perform feature dimension compression on the global features to obtain distance view features, including:

[0029] Use max pooling to obtain the features of each projection space;

[0030] Perform feature dimension compression to obtain projection space features;

[0031] The point clouds within each projection space share the projection space features;

[0032] Then project the projection space features into the distance view features.

[0033] Furthermore, perform two-dimensional distance image feature extraction on the two-dimensional distance image, including:

[0034] ;

[0035] is a backbone network composed of a residual convolution structure, and the distance view features are denoted as the network input of the i-th stage, and the distance view features are denoted as the network output of the i-th stage.

[0036] Furthermore, update the point cloud features and the projection space according to the extracted two-dimensional distance image features, and then obtain a new two-dimensional distance image by using the updated point cloud features and projection space, including

[0037] Inverse map the extracted two-dimensional distance image features into the three-dimensional space and splice them with the point cloud feature vectors;

[0038] Then perform non-linear transformation and feature fusion on the splicing result, and update the point cloud features according to the feature fusion result. Then use the updated point cloud features to update the two-dimensional distance image features again;

[0039] The previous paragraph uses the two-dimensional distance image to inverse map into the three-dimensional space (become point cloud), and fuses with the point cloud features of the previous stage to update the point cloud features, and then projects it to two dimensions, that is, updates the two-dimensional distance image features. Fuse the updated two-dimensional distance image with the initial two-dimensional image (that is, update the two-dimensional image features). Thus, the consistency between two-dimensional features and three-dimensional features is achieved, and the subtle differences in image features are prevented from being amplified when projected into the three-dimensional space.

[0040] Then project the updated two-dimensional distance image features into the projection space corresponding to the updated point cloud features to obtain a new two-dimensional distance image.

[0041] Furthermore, during the two-dimensional distance image feature extraction process, use high-pass filtering to extract high-frequency information from high-resolution features and merge it with residual connections, and use low-pass filtering to extract low-frequency information from low-resolution features and align it for conventional upsampling.

[0042] The beneficial effects of the present invention are as follows:

[0043] The present invention proposes a method for real-time semantic segmentation of point clouds based on distance views, which fully integrates the three-dimensional spatial coordinate features and reflection intensity features of point clouds to improve the network's ability to learn fine features. Secondly, it balances multi-scale feature fusion. The high-resolution features of an image contain more global features, while the low-resolution features contain more semantic features. The present invention uses adaptive multi-scale feature fusion for distance images to make up for the deficiencies of linear multi-scale interpolation fusion, and can effectively identify abnormal point cloud intensity points. Especially in the edge area, the point cloud intensity is unstable, and there are a large number of intensity disappearance points. The fusion of high-frequency features and low-frequency features can effectively solve the problem of inconsistent intra-class segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a flowchart of a method for real-time semantic segmentation of point clouds based on distance views in this application.

[0045] Figure 2 It is a schematic diagram of reflection intensity reconstruction in this application.

[0046] Figure 3 It is a schematic diagram of the architecture of the residual feature extraction structure in this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0048] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0049] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0050] The following will describe the specific embodiments of the present invention in detail with reference to the accompanying drawings.

[0051] The present invention provides a method for real-time semantic segmentation of point clouds based on distance views, focusing on the feature extraction and utilization of unordered point clouds to achieve accurate and robust real-time semantic segmentation. A geometric field-of-view reconstruction module is designed to effectively address the feature misalignment problem caused by the projection from three-dimensional space to two-dimensional space. A reflection intensity reconstruction module is designed to update the zero-reflection points in the state where the reflection intensity disappears, compensating for the accuracy loss caused by the reflection disappearance problem. Adaptive multi-scale feature fusion is used to improve the network performance and efficiency.

[0052] As Figure 1 shown, the present invention relates to a method for real-time semantic segmentation of point clouds based on distance views, which specifically includes the following contents:

[0053] 1. Data acquisition

[0054] When the lidar is working, it emits laser beams at a fixed angle to scan the scene in a certain range of the environment, and obtains the spatial position and reflection intensity of the target by measuring the time, intensity, etc. of the emitted signal sent and returned. Since there are a variety of existing lidar devices with different parameters, the quality, reflection intensity, and other information of the collected point cloud data are also different. Using a single data processing method can no longer meet the actual application requirements. To solve the problem of inconsistent point cloud data, a geometric field-of-view reconstruction module and a reflection intensity reconstruction module are introduced. Among them, the initial point cloud data is , where I is the reflection intensity, and x, y, z are the initial three-dimensional coordinates of the point cloud. The calculation formulas for x, y, z are as follows:

[0055] ;

[0056] ;

[0057] ;

[0058] ;

[0059] is the time difference between the laser beam emission and reception, c is the speed of light, then d represents the distance between the acquisition device and the object, is the pitch angle of the laser beam relative to the plane, is the azimuth angle of the laser beam relative to the front.

[0060] 2.1 Feature encoding

[0061] The geometric field-of-view reconstruction module divides the point cloud scene into K sub-regions , where , where \(R\) represents the real number space, \(n\) is the number of point clouds in each projection space, and \(c\) is the feature dimension of each point cloud; the reflection intensity reconstruction module performs mean pooling on the intensity of the point cloud data in the projection space, effectively removing the influence of pseudo-noise points, solving the problem of abnormal reflection intensity values, and enhancing the quality and continuity of the data. This stage can be formulated as:

[0062] ;

[0063]

[0064] ;

[0065] GFVR represents the geometric field of view reconstruction module, aiming to improve the quality of the distance view. \((u n , v n ) represent the abscissa and ordinate of the point cloud mapped onto the distance image respectively, \((W, H)\) represent the width and height of the distance image respectively, represents the conversion of the depression angle of the lidar acquisition device to radians, \(f\) represents the conversion of the field of view angle of the lidar acquisition device to radians, \(\Delta\) represents the offset in the vertical direction during the point cloud projection process to address the accuracy reduction problem caused by the non-identical starting points of the laser beams, represents the field of view angle in the vertical direction of the acquisition device, Envs represents the height of the device itself and information such as its relative position to the ground, and represent the depression angle and elevation angle of the corresponding laser beam acquisition device respectively. RIR represents the reflection intensity reconstruction module, aiming to improve the model's ability to learn features, especially to learn the reflection intensity vanishing points. Specifically, in the case where most of the point cloud reflection intensities in the projection space are normal and only contain a small number of points with a reflection intensity of 0, the mean reflection intensity of the points with non-zero reflection intensity in the projection space is assigned to the points with a reflection intensity of 0, that is, excluding the influence of the reflection intensity vanishing points on the whole; when most of the reflection intensities in the projection space are 0, it is considered that the disappearance of the reflection intensity in this area is caused by the material of the object itself rather than the introduced noise, as shown in Figure 2 .

[0066] This application seamlessly integrates the spatial features and reflection intensity features of the discrete point clouds in the projection space, accurately captures the potential feature relationships of the point cloud structure and reflection characteristics, realizes the deep fusion of three-dimensional space coordinates and reflection intensity, and enhances the discrimination ability of the model in complex scenes. Specifically, all the points in each projection space are regarded as a set, and each point cloud has 4 basic attributes \((x, y, z, I)\), which are encoded into 10 mutually related basic feature attributes: \(P k n*4 -> P kn*10 , including the initial three-dimensional coordinates and reflection intensity (x, y, z, I) of the point cloud, the Manhattan distance, vector difference, and reflection intensity difference between each point in the projection space and the virtual center of the set, as well as the depth information of the point cloud.

[0067] Subsequently, the global features of the point cloud scene are extracted through a multi-layer perceptron , and the features of each projection space are obtained using max pooling, and feature dimension compression is performed to obtain the projection space features , and the point cloud within each projection space shares features , and then it is projected into distance view features :

[0068] ;

[0069] ;

[0070] ;

[0071] Among them, MLP is a multi-layer perceptron for extracting high-level features in the point cloud; MAX represents max pooling, refers to mapping the point cloud features to distance image features, C represents the number of channels after feature dimension compression, and H and W respectively represent the resolution of the distance image, namely height and width.

[0072] 2.2 Feature Extraction

[0073] Next, multi-scale features are extracted from the two-dimensional distance image stage by stage, and the features of the point cloud and the projection space are updated. The feature extraction at this stage is implemented through a residual network, and the distance view features are denoted as the network input of the i-th stage, are denoted as the output of the i-th stage. This process is expressed as:

[0074] ;

[0075] is the backbone network composed of residual convolutional structures, and the network structure is as Figure 3 shown, passing through a 3X3 convolutional layer, a SyncBN normalization layer, an Hswish activation layer, a 3X3 convolutional layer, a SyncBN normalization layer, feature fusion, and an Hswish activation layer in sequence from input to output.

[0076] After feature extraction, the output Inverse map it to the three-dimensional space and splice it with the point cloud feature vector, and then perform non-linear transformation and feature fusion through a multi-layer perceptron. Since the distance view is a projection of the real scene and cannot be directly used to predict the semantic information of the three-dimensional space, in order to maintain the consistency of the features between the point cloud and the projection space, feature fusion is required and the point cloud features are updated. This process can be expressed as:

[0077] ;

[0078] Among them, means inverse mapping the image features into point cloud features, is the updated global point cloud feature.

[0079] Subsequently, use the updated point cloud feature to update the two-dimensional distance image feature again, and obtain a new according to the projection space divided during the feature encoding process, and project it onto the distance image. Then update the view feature through the PFuse and residual attention weighting mechanism. This step can be formulated as:

[0080] ;

[0081] ;

[0082] Among them, PFuse is composed of a convolutional layer with a convolutional kernel size of and a stride of 1, a fully connected layer and an activation function layer. The convolutional layer extracts local features, the fully connected layer maps the features to the specified dimension, and the activation function further enhances the feature expression ability. The residual attention weighting mechanism can pay more attention to the key and important features, reduce the computational burden, and improve the efficiency and accuracy.

[0083] 2.3 Adaptive multi-scale feature fusion

[0084] During the extraction process of the two-dimensional view features, downsampling operations are used to reduce the resolution in order to extract high-dimensional features, but this also leads to the problem of boundary information loss. Since the distance view is only a projection of the three-dimensional space, the boundaries adjacent in the image may be far apart in the real scene, which will obviously affect the segmentation accuracy. Therefore, it is necessary to fuse the high-level and low-level features. Specifically, the present invention uses high-pass filtering to extract high-frequency information from the high-resolution features and merges it with residual connections to further enhance feature fusion; uses low-pass filtering to extract low-frequency information from the low-resolution features and aligns it for conventional upsampling to achieve multi-scale feature fusion with high quality. By reducing the problem of intra-class feature inconsistency during the upsampling process, enhancing the high-frequency detailed boundary information lost during the downsampling process, and improving the network performance.

[0085] This step is formulated as:

[0086] ;

[0087] and represent the high - resolution features and low - resolution features of each stage respectively; represent low - pass filtering and high - pass filtering respectively, which are used to extract the high - frequency features of high - resolution features and the low - frequency features of low - resolution features; is a conventional upsampling operation. And the multi - scale features of different stages are fused to obtain the feature .

[0088] 2.4 Predicting semantic labels

[0089] Use the multi - scale fusion features to predict the point cloud semantic labels, and use the cross - entropy loss to train the model parameters.

[0090] ;

[0091] Among them, C represents the total number of point cloud categories; represents the prediction result, 1 for correct and 0 otherwise; is the probability distribution predicted by the model.

[0092] The advantages of the present invention compared with the prior art are as follows:

[0093] The present invention inputs the point cloud captured by the lidar into the feature encoding part of the segmentation network, and uses the geometric field - of - view reconstruction module to divide the point cloud scene into different sub - regions; performs intensity reconstruction on the point cloud in the projection space through the reflection intensity reconstruction module to update the reflection intensity value of the pseudo - noise; proposes to seamlessly fuse the three - dimensional space features and the reflection intensity features, enhancing the network's ability to learn detailed features and distinguish pseudo - noise. By extracting high - frequency information from high - resolution features and using the way of residual connection for fusion, the effect of feature fusion is further enhanced. At the same time, extract low - frequency information from low - resolution features and achieve the fusion of multi - scale features through upsampling alignment. This method effectively reduces the problem of inconsistent intra - class features in the upsampling process and makes up for the loss of high - frequency details and boundary information in the downsampling process, thus improving the overall performance of the network. The present invention improves the feature expression ability of the distance image through geometric field - of - view reconstruction, enhances the network's ability to learn features, and reduces the impact of noise on the model accuracy. The present invention smooths the reflection intensity of the point cloud in the projection space through reflection intensity reconstruction and processes the pseudo - noise point cloud with linear time complexity. The present invention achieves the best balance between accuracy and real - time performance, and realizes a competitive network model with limited input and lightweight network structure design.

[0094] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principles of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A real-time semantic segmentation method for point clouds based on distance views, characterized in that, Including: Obtain the initial point cloud data of the scene, where the initial point cloud data includes the initial three-dimensional spatial coordinates and reflection intensity of the point cloud; Perform geometric field-of-view reconstruction on the initial point cloud data to obtain several divided projection spaces, perform reflection intensity reconstruction on the point cloud data in the projection spaces, fuse the spatial features and reflection intensity features of the discrete point clouds in the several projection spaces after reflection intensity reconstruction, extract global features from the point cloud data after feature fusion, and perform feature dimension compression on the global features to obtain distance view features; Project the distance view features onto a two-dimensional distance image, then extract two-dimensional distance image features from the two-dimensional distance image, update the point cloud features and projection spaces according to the extracted two-dimensional distance image features, then obtain a new two-dimensional distance image using the updated point cloud features and projection spaces, and update the view features of the new two-dimensional distance image through a residual attention mechanism; Perform adaptive multi-scale feature fusion on the two-dimensional distance image after view feature update and output the predicted semantic label.

2. The real-time semantic segmentation method of point cloud based on distance view according to claim 1, characterized in that The initial point cloud data is , where I is the reflection intensity, and x, y, z are the initial three-dimensional spatial coordinates of the point cloud. The calculation formulas for x, y, and z are as follows: ; ; ; ; where \(t\) is the time difference between the laser beam emission and reception, \(c\) is the speed of light, and \(d\) represents the distance between the acquisition device and the object. where \(\theta\) is the pitch angle of the laser beam relative to the plane. where \(\varphi\) is the azimuth angle of the laser beam relative to the front, and \(n\) is the number of point clouds in each projection space.

3. The real-time semantic segmentation method of point cloud based on distance view according to claim 2, characterized in that Performing geometric field-of-view reconstruction on the initial point cloud data includes: , , Among them, the is geometric field of view reconstruction. The (u n , v n ) respectively represent the abscissa and ordinate of the point cloud mapped onto the distance image. (W, H) respectively represent the width and height of the distance image. represents the conversion of the depression angle of the lidar acquisition device to radians. f represents the conversion of the field of view angle of the lidar acquisition device to radians. ∆ represents the offset in the vertical direction during the point cloud projection process to address the accuracy reduction problem caused by the non - same starting point of the laser beams. represents the field of view angle in the vertical direction of the acquisition device. Envs represents the height of the device itself and information such as its relative position to the ground. and respectively represent the depression angle and elevation angle of the corresponding laser beam acquisition device. are the K sub - regions divided from the initial point cloud data, where , represents the real number space, and c is the feature dimension of each point cloud.

4. A method for real-time semantic segmentation of point clouds based on a distance view according to claim 1 or 3, characterized in that Performing reflection intensity reconstruction on the point cloud data in the projection spaces includes performing mean pooling processing on the emission intensity of the point cloud data in the projection spaces.

5. A method for real-time semantic segmentation of point clouds based on a distance view according to claim 3, characterized in that Fusing the spatial features and reflection intensity features of the discrete point clouds in the several projection spaces after reflection intensity reconstruction includes: Regarding all points in each projection space as a set, and each point cloud has 4 basic attributes (x, y, z, I); Encode the four basic attributes of each point cloud into ten mutually related basic feature attributes P k n*4 ->P k n*10 ; Where the basic feature attributes include the initial three-dimensional spatial coordinates and reflection intensity of the point cloud (x, y, z, I), the Manhattan distance between each point in the projection space and the virtual center of the set, the difference in three-dimensional spatial coordinate vectors, the difference in reflection intensity, and the depth information of the point cloud.

6. The real-time semantic segmentation method of point cloud based on distance view according to claim 1, characterized in that Use a multi-layer perceptron to extract global features from the point cloud data after feature fusion.

7. A real-time semantic segmentation method for point clouds based on a distance view according to claim 1, characterized in that, Performing feature dimension compression on the global features to obtain distance view features includes: Using max pooling to obtain the features of each projection space; Performing feature dimension compression to obtain projection space features; The point clouds in each projection space share the projection space features; Then project the projection space features into the distance view features.

8. A real-time semantic segmentation method for point clouds based on distance views according to claim 1, characterized in that, Extracting two-dimensional distance image features from the two-dimensional distance image includes: , is a backbone network composed of a residual convolution structure, the distance view feature is denoted as the network input of the i-th stage, the distance view feature is denoted as the network output of the i-th stage.

9. A real-time semantic segmentation method for point clouds based on a distance view according to claim 1, characterized in that, Updating the point cloud features and projection spaces according to the extracted two-dimensional distance image features, and then obtaining a new two-dimensional distance image using the updated point cloud features and projection spaces includes: Back-project the extracted two-dimensional distance image features into the three-dimensional space and splice them with the point cloud feature vectors; Then perform non-linear transformation and feature fusion on the splicing result, update the point cloud features according to the feature fusion result, and then update the two-dimensional distance image features again using the updated point cloud features; Then project the updated two-dimensional distance image features into the projection space corresponding to the updated point cloud features to obtain a new two-dimensional distance image.

10. A real-time semantic segmentation method for point clouds based on a distance view according to claim 1, characterized in that, During the two-dimensional distance image feature extraction process, high-pass filtering is used to extract high-frequency information from high-resolution features and merge it with residual connections, and low-pass filtering is used to extract low-frequency information from low-resolution features and align it for conventional upsampling.