Railway point cloud adaptive semantic segmentation method, device and equipment and storage medium
Patent Information
- Application Number
- CN202611147858.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-25
AI Technical Summary
基于投影的方法将三维点云转换为二维图像进行特征提取,容易造成深度信息和空间几何信息损失;基于体素的方法将点云划分为规则三维网格,存在内存占用较大、计算量较高以及局部细节损失的问题;基于图结构的方法通过点和边构建邻接关系,能够描述局部拓扑特征,但网络参数和计算开销较大;基于点集网络的方法能够直接处理无序点云,但对点间方向、距离及局部结构关系的表达仍然有限;基于注意力机制的方法能够融合局部特征和全局特征,但在大规模点云处理中通常需要较高计算资源和较多训练数据
本发明提供的面向铁路点云的自适应语义分割方法,通过对铁路场景点云进行分层编码,并根据各点与局部邻域内邻域点之间的空间关系构建局部几何特征,根据融合后的几何特征和点特征确定邻域点的聚合权重,使局部特征聚合能够突出与中心点语义识别相关的结构信息,减少点云采样及常规池化过程中轨道边缘、杆状设施和其他细小结构特征的损失;进一步地,通过对不同编码层级的局部特征进行统一映射和层级加权,并结合局部特征相对于聚类中心的特征残差生成全局上下文特征,将所述全局上下文特征融入逐层解码过程,使各点的语义分类同时利用局部几何信息、不同层级的语义信息以及点云整体的上下文信息。由此,本方法能够改善现有点云语义分割方法局部几何特征表达不足、局部感受范围有限以及长距离空间关联建模能力较弱的问题,在兼顾点云处理规模的同时,提高复杂铁路场景中不同类别目标的综合语义分割准确性和稳定性。
Smart Images

Figure CN122821137A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of point cloud processing, and in particular to an adaptive semantic segmentation method, apparatus, device, and storage medium for railway point clouds. Background Technology
[0002] In existing technologies, railway scene point clouds typically exhibit characteristics such as large spatial span, complex structural types, significant differences in point cloud density, and numerous local occlusions. Traditional point cloud semantic segmentation methods mainly include edge-based, model-fitting, clustering, and region-growing methods. Among these, edge-based methods identify target boundaries based on point cloud intensity, gradient, or normal vectors, but are highly sensitive to scene complexity; model-fitting methods utilize pre-defined geometric models to detect targets with regular shapes, making them difficult to adapt to the diverse forms of railway facilities, and also incurring high computational costs; clustering methods rely on point cloud density or geometric features for classification, and the segmentation results are easily affected by clustering parameters; region-growing methods depend on seed point selection and region merging conditions. Therefore, these methods have certain applicability in small-scale, relatively simple point cloud scenes, but are insufficient to meet the fine semantic segmentation requirements of large-scale, complex railway scenes.
[0003] With the development of deep learning technology, point cloud semantic segmentation methods have gradually formed various technical routes based on projection, voxels, graph structures, point set networks, and attention mechanisms. Projection-based methods convert 3D point clouds into 2D images for feature extraction, which easily leads to the loss of depth and spatial geometric information; voxel-based methods divide point clouds into regular 3D grids, resulting in large memory consumption, high computational cost, and loss of local details; graph structure-based methods construct adjacency relationships through points and edges, which can describe local topological features, but the network parameters and computational overhead are large; point set network-based methods can directly process unordered point clouds, but the expression of inter-point directions, distances, and local structural relationships is still limited; attention mechanism-based methods can fuse local and global features, but in large-scale point cloud processing, they usually require high computational resources and a large amount of training data.
[0004] In summary, existing point cloud semantic segmentation methods still struggle to balance computational efficiency, preservation of local geometric details, and global context modeling when processing large-scale railway scene point clouds. On one hand, geometric information of track edges, pole-like structures, and other small structures is easily lost during point cloud sampling and local feature aggregation. On the other hand, the receptive range of local neighborhood features is limited, making it difficult to fully express the long-distance spatial relationships formed along the railway line's extension direction. Therefore, existing technologies still need to further improve their ability to collaboratively represent local geometric features and global semantic information to enhance the semantic segmentation accuracy and processing efficiency of complex railway scene point clouds. Summary of the Invention
[0005] This invention aims to address at least the technical problems of insufficient local geometric feature representation and limited long-distance spatial association modeling capabilities in existing technologies for large-scale railway point clouds. This invention provides an adaptive semantic segmentation method, apparatus, device, and storage medium for railway point clouds, which can extract local features and model global context for railway scene point clouds, obtaining accurate point-level semantic segmentation results. During the hierarchical encoding process of the point cloud, adaptive local feature aggregation can be performed based on the spatial relationship between points and their neighbors. By weighting and residual aggregation of local features from multiple encoding levels, the fusion of local geometric information and global semantic information is achieved, thereby improving the semantic segmentation accuracy of complex railway scene point clouds.
[0006] In a first aspect, embodiments of the present invention provide an adaptive semantic segmentation method for railway point clouds, the method comprising:
[0007] S1, Obtain the point cloud of the railway scene to be segmented, and perform hierarchical encoding on the point cloud of the railway scene to obtain point features at multiple encoding levels; S2, determine local neighborhoods for points in each coding level, generate local geometric features based on the spatial relationship between the point and each neighboring point in the local neighborhood, fuse the local geometric features with the point features of the neighboring points, determine the aggregation weights corresponding to each neighboring point based on the fused neighboring point features, and aggregate the fused neighboring point features based on the aggregation weights to obtain the local aggregated features of each coding level. S3, perform feature mapping on the local aggregated features of multiple coding levels to obtain multi-layer local features with a unified feature dimension, determine the level weights corresponding to each coding level based on the multi-layer local features, and determine the assigned weights of each cluster center in multiple cluster centers corresponding to the multi-layer local features, and aggregate the feature residuals of the multi-layer local features relative to each cluster center based on the level weights and the assigned weights to obtain global context features. S4. The global context features are fused with the point features in the decoding process to restore the spatial resolution of the railway scene point cloud layer by layer, and the semantic category of each point is determined according to the restored point features to obtain the semantic segmentation result of the railway scene point cloud.
[0008] Optionally, the railway scene point cloud to be segmented is obtained from the original railway scene point cloud through point cloud preprocessing, wherein the point cloud preprocessing includes: The intensity attribute values of each point in the original railway scene point cloud are obtained, points with negative intensity attribute values or values greater than a preset intensity upper limit are filtered out, and the intensity attribute values of the remaining points after filtering are normalized to obtain the attribute-processed point cloud. The attribute processing point cloud is divided into grids according to a preset spatial sampling interval, and the points in each non-empty grid are sampled to obtain a downsampled point cloud as the railway scene point cloud to be segmented. The preset spatial sampling interval is 0.06 meters. A spatial index is established based on the spatial coordinates of each point in the original railway scene point cloud, and a projection index is established based on the spatial correspondence between the points in the original railway scene point cloud and the points in the downsampled point cloud. After obtaining the semantic segmentation result, the semantic segmentation result is mapped to each point in the original railway scene point cloud according to the projection index.
[0009] Optionally, S2 includes local neighborhood determination and local geometric feature generation, specifically including: Using points in each coding level as center points, the Euclidean distances between other points in the same coding level and the center point are determined respectively. Then, a preset number of points are selected from the other points as neighborhood points in order of increasing Euclidean distances to obtain the local neighborhood of the center point. Obtain the spatial coordinates of the center point and the spatial coordinates of each neighboring point. Subtract the spatial coordinates of the center point from the spatial coordinates of the neighboring points to obtain the coordinate offset of each neighboring point relative to the center point. Determine the distance between each neighboring point and the center point based on the coordinate offset. The spatial coordinates of the center point, the spatial coordinates of the neighboring points, the coordinate offset, and the distance are concatenated, and the concatenation result is subjected to point-by-point feature mapping to obtain the local geometric features corresponding to each neighboring point. The local geometric features corresponding to each neighboring point are concatenated with the point features of the corresponding neighboring point, and the concatenation result is subjected to feature mapping to obtain the enhanced neighborhood features corresponding to each neighboring point.
[0010] Optionally, S2 further includes neighborhood feature aggregation, which includes: Feature mapping is performed on each enhanced neighborhood feature to obtain the attention score corresponding to each neighborhood point; The attention scores of each neighboring point are normalized within the local neighborhood corresponding to the same center point to obtain the attention weights corresponding to each neighboring point. The attention weights are used to weight each enhanced neighborhood feature, and the weighted enhanced neighborhood features are summed to obtain the attention aggregation feature. The pooled aggregate features are obtained by comparing the feature values of each enhanced neighborhood feature corresponding to the same center point in each feature channel and retaining the maximum feature value in each feature channel. The attention aggregation feature and the pooling aggregation feature are concatenated, and the concatenation result is subjected to feature mapping to obtain the local aggregation feature corresponding to the center point.
[0011] Optionally, S3 includes determining the hierarchical weights, which includes: Select multiple target coding levels with different spatial resolutions from the plurality of coding levels; Pointwise feature mapping is performed on the local aggregated features of each target encoding level to map the local aggregated features of each target encoding level to the same feature dimension, thereby obtaining the unified dimension features corresponding to each target encoding level; Global pooling is performed on the uniform dimension features corresponding to each target encoding level to obtain the hierarchical description features corresponding to each target encoding level; Feature mapping is performed on the descriptive features of each level to obtain the level score corresponding to each target encoding level; The scores of each target coding level are normalized across the multiple target coding levels to obtain the level weights corresponding to each target coding level.
[0012] Optionally, S3 further includes global context feature generation, which includes: A preset number of trainable cluster centers are set, and the feature dimension of each cluster center is the same as the feature dimension of the unified dimension feature. For each unified dimension feature in each target encoding level, the allocation score of the unified dimension feature corresponding to each cluster center is determined, and the allocation score is normalized among the multiple cluster centers to obtain the soft allocation weight of the unified dimension feature corresponding to each cluster center. The feature residuals of each unified dimension feature relative to each cluster center are obtained by subtracting each unified dimension feature from each cluster center. Based on the hierarchical weights corresponding to each target encoding level and the soft-assigned weights corresponding to each unified dimension feature, the residuals of each feature are weighted, and for each cluster center, the weighted feature residuals corresponding to different target encoding levels and different points are accumulated to obtain the residual aggregated features corresponding to each cluster center. The residual aggregated features corresponding to each cluster center are normalized, and the normalized residual aggregated features are concatenated to obtain the global context features.
[0013] Optionally, S4 includes point feature decoding and semantic classification, specifically including: In each decoding layer, the global context features are copied according to the number of points contained in the decoding layer to obtain the global context point features corresponding to each point in the decoding layer; Upsample the point features output from the previous decoding layer to obtain the upsampled point features corresponding to the current decoding layer; Obtain the coding layer point features corresponding to the current decoding layer during the coding phase, concatenate the global context point features, the upsampled point features, and the coding layer point features, and perform feature mapping on the concatenation result to obtain the fusion point features of the current decoding layer; Following the order of point cloud spatial resolution from low to high, upsampling and feature stitching are performed layer by layer to obtain the restored point features; The recovered point features are mapped to the category features corresponding to the target semantic category, and the semantic category of each point is determined according to the category features to obtain the semantic segmentation result of the railway scene point cloud.
[0014] Secondly, embodiments of the present invention provide an adaptive semantic segmentation device for railway point clouds, the device may include a point cloud encoding module, a local feature aggregation module, a global context generation module, and a semantic decoding module; The point cloud encoding module acquires the point cloud of the railway scene to be segmented, performs hierarchical encoding on the point cloud of the railway scene, and obtains point features at multiple encoding levels. The local feature aggregation module determines local neighborhoods for points in each coding level, generates local geometric features based on the spatial relationship between the point and each neighboring point in the local neighborhood, fuses the local geometric features with the point features of the neighboring points, determines the aggregation weight corresponding to each neighboring point based on the fused neighboring point features, and aggregates the fused neighboring point features based on the aggregation weights to obtain the local aggregated features of each coding level. The global context generation module performs feature mapping on the local aggregated features of multiple coding levels to obtain multi-layer local features with a unified feature dimension. Based on the multi-layer local features, it determines the level weights corresponding to each coding level and the allocation weights of the multi-layer local features to each cluster center in multiple cluster centers. Based on the level weights and the allocation weights, it aggregates the feature residuals of the multi-layer local features relative to each cluster center to obtain global context features. The semantic decoding module fuses the global context features with the point features in the decoding process, restores the spatial resolution of the railway scene point cloud layer by layer, and determines the semantic category of each point based on the restored point features, thereby obtaining the semantic segmentation result of the railway scene point cloud.
[0015] Thirdly, embodiments of the present invention provide a computer device, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the methods described above.
[0016] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer-executable instructions for executing the adaptive semantic segmentation method for railway point clouds as described in the first aspect embodiment above. Since the computer-readable storage medium employs all the technical solutions of the adaptive semantic segmentation method for railway point clouds described in the above embodiments, it possesses at least all the beneficial effects brought about by the technical solutions of the above embodiments.
[0017] The adaptive semantic segmentation method, apparatus, device, and storage medium for railway point clouds according to the present invention have at least the following beneficial effects: The adaptive semantic segmentation method for railway point clouds provided by this invention performs hierarchical encoding on railway scene point clouds and constructs local geometric features based on the spatial relationships between each point and its local neighbors. It then determines the aggregation weights of neighboring points based on the fused geometric features and point features, enabling local feature aggregation to highlight structural information related to the semantic recognition of the center point and reducing the loss of track edges, pole-like facilities, and other small structural features during point cloud sampling and conventional pooling. Furthermore, it performs unified mapping and hierarchical weighting of local features at different encoding levels and generates global context features by combining the feature residuals of local features relative to cluster centers. These global context features are then integrated into the layer-by-layer decoding process, allowing the semantic classification of each point to simultaneously utilize local geometric information, semantic information at different levels, and the overall context information of the point cloud. Therefore, this method can improve upon the shortcomings of existing point cloud semantic segmentation methods, such as insufficient local geometric feature representation, limited local receptive range, and weak long-distance spatial association modeling capabilities. While considering the scale of point cloud processing, it improves the comprehensive semantic segmentation accuracy and stability of different categories of targets in complex railway scenes.
[0018] The adaptive semantic segmentation device for railway point clouds provided by this invention achieves the aforementioned local geometric enhancement, multi-layer global feature aggregation, and point-by-point semantic classification processes through data collaboration between the point cloud encoding module, local feature aggregation module, global context generation module, and semantic decoding module. This enables continuous feature extraction, feature fusion, and segmentation result output for railway scene point clouds. The electronic device executes the above data processing process by calling a computer program in memory via a processor. A computer-readable storage medium stores the computer program capable of implementing this processing, allowing the method to be deployed on appropriate point cloud processing devices or computing platforms, providing feasible software and hardware carriers for automated semantic analysis of railway scene point clouds.
[0019] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. Attached Figure Description
[0020] Figure 1 This is a flowchart of the adaptive semantic segmentation method for railway point clouds in Example 1; Figure 2 This is a flowchart of the point cloud data preprocessing process for a railway scene in Example 2; Figure 3 This is the overall framework diagram of the adaptive semantic segmentation network for railway point clouds in Example 2; Figure 4 This is a structural diagram of the adaptive local feature aggregation module in Example 2; Figure 5 This is a structural diagram of the integrated vector local aggregation descriptor module in Example 2; Figure 6 This is a comparison experiment result of point cloud semantic segmentation in a railway scene, as shown in Example 2. The figure is labeled as follows: (a) shows the results of the semantic segmentation comparison experiment of utility poles. Figure 6 (b) is a comparison of the results of the semantic segmentation experiment of the signal pole. Detailed Implementation
[0021] The present invention will be further described in detail below with reference to experimental examples and specific embodiments. However, this should not be construed as limiting the scope of the above-mentioned subject matter of the present invention to the following embodiments. All technologies implemented based on the content of the present invention fall within the scope of protection of the present invention.
[0022] Unless otherwise specified, the terms "upper," "lower," "left," "right," "center," "inner," "outer," and "side" used in the description of specific embodiments of the present invention to indicate orientation or positional relationships are based on the orientation or positional relationships shown in the accompanying drawings, or the orientation or positional relationship in which the product / equipment / device is usually placed during use. These terms are merely for the purpose of facilitating the description of the present invention or simplifying the description in specific embodiments, and for enabling those skilled in the art to quickly understand the solution, and do not indicate or imply that a particular device / component / element must have a specific orientation, or be constructed and operated in a specific positional relationship. Therefore, they should not be construed as limitations on the present invention.
[0023] In the description of the embodiments of this invention, technical terms such as "first" and "second" only distinguish one entity or operation from another, and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary or secondary relationship of the indicated technical features. In the description of the embodiments of this invention, "multiple" means two or more, unless otherwise explicitly defined.
[0024] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0025] Example 1 This embodiment provides an adaptive semantic segmentation method for railway point clouds.
[0026] In their research on semantic segmentation of point clouds in railway scenes, the inventors discovered that railway scene point clouds typically exhibit characteristics such as large data scale, long spatial extension distance, significant differences in target scale, and complex local structures. Specifically, railway tracks are continuously linearly distributed, signal poles, utility poles, and supporting devices have slender local structures, while track beds, vegetation, buildings, and the ground possess different geometric shapes. To perform point-by-point semantic recognition of different targets in a railway scene, it is necessary to simultaneously extract local geometric information and semantic association information over a large spatial range from the point cloud.
[0027] Existing point cloud segmentation methods based on edge analysis, model fitting, clustering, or region growing are easily affected by geometric models, point cloud density, clustering parameters, seed point selection methods, or region merging conditions, making them difficult to adapt to large-scale and complex railway scenarios. While existing deep learning-based point cloud semantic segmentation methods can automatically extract point cloud features, they are prone to geometric detail loss during hierarchical sampling and local feature aggregation. Furthermore, features extracted based on local neighborhoods cannot fully reflect the long-distance spatial relationships between different areas of the railway line, thus affecting the semantic segmentation results of complex targets and small facilities.
[0028] Based on the above understanding, this embodiment generates local geometric features during the point cloud hierarchical encoding process according to the spatial relationship between each point and its local neighborhood points, and determines the aggregation weights of different neighborhood points based on the neighborhood point features to enhance local structural features; furthermore, it generates global context features based on the local features of multiple encoding levels, and integrates the global context features into the point cloud decoding process, so that the semantic classification of each point utilizes both local geometric information and global semantic information. The embodiment is described according to the structure of "technical background - overall process - intermediate results - final output".
[0029] Figure 1 A flowchart illustrating the adaptive semantic segmentation method for railway point clouds provided in this embodiment is shown. Figure 1As shown, the method includes: acquiring a railway scene point cloud to be segmented; performing hierarchical encoding on the railway scene point cloud to obtain point features at multiple encoding levels; based on the spatial relationship between each point and its neighboring points in the local neighborhood, performing local aggregation on the neighborhood point features to obtain local aggregation features at multiple encoding levels; generating global context features based on the local aggregation features at multiple encoding levels; and fusing the global context features into the point cloud decoding process to restore the spatial resolution of the point cloud and determine the semantic category of each point.
[0030] Specifically, this embodiment first acquires a point cloud of a railway scene to be segmented. The railway scene point cloud includes multiple points, each with corresponding spatial coordinates and point attribute features. The railway scene point cloud is then hierarchically encoded, enabling the point cloud to form point features with different spatial resolutions and semantic levels at different encoding levels. Higher spatial resolution encoding levels can retain more local structural information, while lower spatial resolution encoding levels can form semantic representations covering a larger spatial range.
[0031] After obtaining point features from multiple coding levels, local neighborhoods are determined for points in each coding level. Specifically, the point to be aggregated for local features is taken as the center point, and based on the spatial relationship between the center point and other points in the same coding level, corresponding neighboring points are determined from the other points to form the local neighborhood of the center point.
[0032] Based on the spatial relationship between the center point and its neighboring points, local geometric features are generated for each neighboring point. These local geometric features represent the relative positional relationship between the center point and its neighboring points. Furthermore, the local geometric features are fused with the point features of the neighboring points, so that the fused neighboring point features simultaneously contain the semantic information of the neighboring points themselves and the spatial geometric information between the neighboring points and the center point.
[0033] After forming the fused neighborhood point features, the corresponding aggregation weights are determined based on the fused features of each neighborhood point. Then, based on these aggregation weights, the neighborhood point features within the local neighborhood are aggregated to obtain the local aggregated features corresponding to the center point. This processing method allows for adjusting the proportion of different neighborhood points in the aggregation result according to their contribution to the semantic recognition of the center point, enhancing the expression of point features related to track edges, pole-shaped facilities, and other local structures.
[0034] The above local feature aggregation process is performed on points in each coding level to obtain local aggregated features corresponding to multiple coding levels. These local aggregated features contain local geometric and semantic information at different spatial resolutions, providing a multi-layer feature foundation for subsequent construction of global context features.
[0035] After obtaining local aggregated features from multiple coding levels, feature mapping is performed on the local aggregated features from different coding levels to form multi-layer local features with a unified feature dimension. Based on the multi-layer local features of each coding level, corresponding level weights are determined. These level weights represent the relative contributions of local features from different coding levels to the generation of global context features.
[0036] Further, the assigned weights of the multi-layer local features corresponding to different cluster centers are determined, and the feature residuals of each multi-layer local feature relative to its corresponding cluster center are determined respectively. Based on the level weights corresponding to each encoding level and the assigned weights corresponding to each multi-layer local feature, the feature residuals are aggregated to form the residual aggregation results corresponding to each cluster center; then, global context features are generated based on the residual aggregation results corresponding to multiple cluster centers.
[0037] Through the above processing, local features at different coding levels can participate in global feature aggregation according to their feature contributions, so that the resulting global context features simultaneously contain local geometric details, point features at different semantic levels, and contextual information over a large spatial range, thereby supplementing the long-distance spatial relationships that are difficult to express directly by local neighborhood features.
[0038] After obtaining the global context features, these global context features are fused with the point features generated during the decoding process. Specifically, the point features formed during the encoding stage are decoded layer by layer, and the global context features are fused with the corresponding decoded point features during the decoding process, so that the point features in each decoding layer simultaneously contain local spatial information and global semantic information.
[0039] As the decoding process proceeds layer by layer, the spatial resolution of the point cloud is gradually restored, yielding point features corresponding to the point cloud of the railway scene to be segmented. Based on the restored point features, semantic classification is performed on each point to determine its corresponding semantic category, resulting in the semantic segmentation result of the railway scene point cloud.
[0040] In this embodiment, the point cloud of the railway scene to be segmented sequentially undergoes hierarchical encoding, local geometric feature generation, neighborhood point feature aggregation, multi-layer global context feature generation, and point-by-point decoding and classification. This reduces the loss of local geometric information during point cloud hierarchical sampling and local aggregation, enhances the ability to express long-distance spatial relationships along the railway line's extension direction, and ensures that the semantic classification of each point takes into account both local geometric details and global semantic relationships, thereby improving the semantic segmentation accuracy of large-scale complex railway scene point clouds. The above technical process is consistent with the local feature aggregation, global context encoding, and semantic decoding approach described in the supplementary technical disclosure.
[0041] The adaptive semantic segmentation method for railway point clouds provided in this invention can be applied to railway infrastructure inspection, line safety monitoring, and digital modeling of railway scenes, including semantic recognition of rails, track beds, signal poles, support devices, overhead contact lines, fences, utility poles, vegetation, buildings, and the ground. In the above implementation, when performing semantic segmentation on large-scale railway scene point clouds, local geometric features can be extracted based on the spatial relationship between points and their neighbors. Global contextual features can be generated based on local features from multiple coding levels. By fusing local geometric features and global contextual features into the point cloud decoding process, collaborative expression of local structural information and long-distance semantic association information is achieved, thereby improving the accuracy of semantic segmentation of complex railway scene point clouds.
[0042] Example 2 This embodiment is a specific implementation of the adaptive semantic segmentation method for railway point clouds described in Embodiment 1. This embodiment is based on an encoder-decoder point cloud semantic segmentation network. An adaptive local feature aggregation module is set in the encoding stage, and a comprehensive vector local aggregation descriptor module is set between the encoder and decoder. Specifically, the adaptive local feature aggregation module is used to enhance and weightedly aggregate local neighborhood features by combining the spatial relationships between points; the comprehensive vector local aggregation descriptor module is used to fuse point features from different encoding levels to generate global context features, and introduce these global context features into the decoding process.
[0043] The processing steps in this embodiment include point cloud data preprocessing, point cloud hierarchical encoding, adaptive local feature aggregation, multi-layer global context encoding, point feature decoding, and point-by-point semantic classification. The following combines... Figures 2 to 6 Each processing step is explained.
[0044] I. Point Cloud Data Preprocessing like Figure 2 As shown, before inputting the railway scene point cloud into the semantic segmentation network, the original railway scene point cloud is preprocessed to form the network input point cloud, and the index relationship required for subsequent neighborhood query and segmentation result recovery is established.
[0045] The original railway scene point cloud may exhibit abnormal intensity attributes, excessively dense local point clouds, and significant differences in point density across different regions due to variations in acquisition equipment, acquisition environment, or acquisition time period. Therefore, this embodiment sequentially performs attribute processing, grid downsampling, spatial index construction, and projection index construction.
[0046] Specifically, the original railway scene point cloud is loaded, and the spatial coordinates and intensity attribute values of each point are obtained. Points with negative intensity attribute values or values greater than a preset intensity upper limit are filtered out. The intensity attribute values of the retained points are then normalized to obtain an attribute-processed point cloud. Normalization reduces the impact of differences in intensity attribute value ranges between different point cloud data on subsequent feature extraction.
[0047] Subsequently, the attribute processing point cloud is divided into grids according to a preset spatial sampling interval, and points within each non-empty grid are sampled to obtain a downsampled point cloud. Given that the data is collected by a track-mounted lidar and the point cloud density is relatively high, the grid sampling interval should be set to 0.04 meters to 0.06 meters. In practice, using 0.06 meters or 0.05 meters can effectively reduce the total amount of point cloud, thereby reducing computational costs and improving processing efficiency while better reproducing the railway scene.
[0048] Downsampling is used to reduce redundant points in dense point cloud areas while preserving the main spatial structure of railway tracks and line ancillary facilities, so that the network input point cloud has a relatively balanced point density.
[0049] Furthermore, a KD-tree is constructed based on the spatial coordinates of each point in the original railway scene point cloud. The KD-tree is used to organize the spatial coordinates of the point cloud and provide a spatial index for subsequent neighborhood point retrieval, thereby reducing the computational load generated by traversing the point cloud point by point.
[0050] Simultaneously, a projection index is constructed based on the spatial correspondence between points in the original railway scene point cloud and points in the downsampled point cloud. After the network completes point-by-point classification of the downsampled point cloud, it can map semantic categories to corresponding points in the original railway scene point cloud based on the projection index.
[0051] After completing the above processing, save the downsampled point cloud, KD-tree, and projection index for use in the network training or inference process.
[0052] II. Overall Structure of Semantic Segmentation Networks like Figure 3 As shown, this embodiment uses an encoder-decoder hierarchical UNet network as the basic framework of the semantic segmentation network. The network includes an encoder, a decoder, and a classifier; the encoder extracts point features at different spatial resolutions through layer-by-layer sampling and local feature aggregation, the decoder restores the spatial resolution of the point cloud layer by layer, and the classifier outputs the semantic category of each point based on the restored point features.
[0053] In this embodiment, the adaptive local feature aggregation module is abbreviated as ALFA module, and the comprehensive vector local aggregation descriptor module is abbreviated as C-VLAD module.
[0054] After the preprocessed railway scene point cloud is input into the network, the initial attributes of each point are first mapped through a fully connected layer or a shared multilayer perceptron to obtain initial point features with N points and 16 feature channels, where N represents the number of points input into the network.
[0055] After the initial point features enter the encoder, the current point set is first randomly downsampled in each encoding level, and then the local neighborhood features of the downsampled points are aggregated through the ALFA module. As the encoding level deepens, the number of points decreases layer by layer, while the number of feature channels increases layer by layer.
[0056] In this embodiment, the number of points and feature channels at each level of the encoder are as follows: (N, 16) → (N / 4, 32) → (N / 16, 128) → (N / 64, 256) → (N / 128, 512).
[0057] The first term in the brackets above represents the number of points at the corresponding coding level, and the second term represents the number of feature channels for each point. By progressively reducing the spatial resolution of the point cloud and increasing the number of feature channels layer by layer, the encoder can gradually expand the spatial receptive range corresponding to the point features while preserving multi-layer local information.
[0058] The point features output by different levels of the encoder have different spatial resolutions and semantic levels. Shallow encoded features retain more information about track edges, rod-shaped facilities, and local surface morphology, while deep encoded features contain semantic information over a larger spatial range. This embodiment selects point features from different encoding levels and performs multi-layer feature fusion and global context encoding through the C-VLAD module.
[0059] During the decoding stage, the decoder recovers the spatial resolution of the point cloud layer by layer in the reverse order of the encoder. The number of points and feature channels at each level of the decoder are as follows: (N / 128, 512) → (N / 64, 256) → (N / 16, 128) → (N / 4, 32) → (N, 16).
[0060] In each decoding layer, the point features output from the previous decoding layer are upsampled, and point features with the corresponding spatial resolution in the encoder are obtained through long hop connections. Subsequently, the upsampled point features, the point features of the corresponding encoding layer, and the global context point features generated by the C-VLAD module are fused to obtain the fused point features of the current decoding layer.
[0061] After layer-by-layer decoding, the recovered point features are classified and mapped using a three-layer multilayer perceptron. The output feature channels of the three-layer multilayer perceptron are 64, 32, and c, respectively, where c represents the number of preset semantic categories. The semantic category of each point is determined based on the output of the last layer, resulting in the semantic segmentation result of the downsampled point cloud.
[0062] Finally, based on the projection index established in the preprocessing stage, the semantic segmentation results of the downsampled point cloud are mapped to the original railway scene point cloud to obtain the semantic category corresponding to each point in the original railway scene point cloud.
[0063] III. Adaptive Local Feature Aggregation Different targets in a railway scenario exhibit distinct local geometries. For example, railway tracks are distributed as continuous lines, signal poles, utility poles, and supporting structures have slender rod-like structures, buildings are mainly planar or block-like structures, and the track bed is composed of numerous discrete points. If only a single pooling method is used to aggregate neighborhood features, some discriminative local information may be weakened during the aggregation process.
[0064] To enhance the network's ability to represent local geometric relationships, this embodiment includes an ALFA module in the encoder's feature aggregation stage. The structure of the ALFA module is as follows: Figure 4 As shown, its input includes the spatial coordinates of the point, color or intensity attributes, and point features output from the previous network layer.
[0065] The following combination Figure 4 Formulas (1) to (6) further explain the local neighborhood determination, local geometric encoding, attention weighting and feature aggregation process in the ALFA module.
[0066] For each center point in the input point cloud, let the spatial coordinates of the input point cloud be... The corresponding initial point features are Where N represents the number of input points, and d represents the dimension of the initial point features. Based on Euclidean distance, k neighboring points are determined for the i-th center point, and the resulting local neighborhood point set is represented as... Furthermore, the coordinate offset between the center point and its neighboring points is determined based on their spatial coordinates. The coordinate offset, along with the center point coordinates, neighboring point coordinates, and the Euclidean distance between them, are used for local geometric encoding. The local geometric features corresponding to the k-th neighboring point... Represented as: (1) In equation (1), and These represent the position coordinates of the center point and the k-th neighboring point, respectively. Indicates the join operation. This represents the calculation of the Euclidean distance between neighboring points and the center point. A shared multilayer perceptron is used to map the stitched spatial information to obtain the local geometric features corresponding to the neighboring points. Subsequently, these local geometric features are compared with the point features in the local neighborhood. Semantic concatenation is performed to obtain enhanced local features, represented as follows: (2) After obtaining the enhanced local features, feature mapping is performed on each neighborhood feature, and the obtained scores are normalized within the same k-neighborhood to obtain the attention weights corresponding to each neighborhood point. Then, the enhanced local features are weighted and summed using these attention weights to obtain the attention aggregation features. Attention weights and attention aggregation features They are represented as follows: (3) (4) The attention weights in equation (3) represent the relative contributions of each neighboring point to the aggregation of features at the center point, and equation (4) represents the weighted summation of the neighboring features based on the attention weights. To further preserve the significant responses of each feature channel within the local neighborhood, max pooling is performed within the same k-neighborhood to obtain max pooled features. : (5) Equation (5) represents selecting the maximum feature value in the local neighborhood of each feature channel. The attention-aggregated feature and the max-pooling feature are concatenated, and feature mapping is performed through a multilayer perceptron to obtain the final local aggregated feature corresponding to the center point. : (6) In equation (6), the attention aggregation feature is used to reflect the comprehensive information of each neighborhood point after dynamic weighting, and the max pooling feature is used to retain the significant responses in the local neighborhood. This indicates a concatenation operation. The combined local aggregated features utilize both the relative geometric relationships between neighboring points and local salient features. Repeating the above process for each center point in the current encoding level yields the corresponding local aggregated features for that encoding level. To verify the role of the ALFA module in modeling local geometric structures, this embodiment sets up experiment a with the ALFA module removed, and compares it with a complete network that simultaneously sets the ALFA and C-VLAD modules. The experimental results are shown in Table 1.
[0067] Table 1. Results of ALFA module ablation experiments (%)
[0068] As shown in Table 1, after removing the ALFA module, the average intersection-union ratio (mIoU) decreased from 77.6% to 77.3%. In categories that are more sensitive to local geometry, such as signal poles, support devices, contact wires, and utility poles, the complete network achieved higher segmentation results, indicating that the ALFA module can enhance the expression of local geometry and neighborhood features.
[0069] IV. Multi-layer global context coding The ALFA module primarily extracts and aggregates point features within the local neighborhood corresponding to the center point, enhancing the local representation of track edges, pole-like structures, and other fine details. However, the spatial coverage of the local neighborhood is limited, and relying solely on local aggregation is insufficient to fully represent the long-distance spatial relationships between different locations along a railway line.
[0070] To this end, this embodiment sets up a C-VLAD module, which sequentially performs unified dimension mapping, hierarchical weighting, and cluster residual aggregation on the point features output from different levels of the encoder to generate fixed-dimensional global context features. The structure of the C-VLAD module is as follows: Figure 5 As shown.
[0071] The following explanation, in conjunction with formulas (7) to (12), further elaborates on the process of multi-layer feature selection, hierarchical weight determination, local feature soft allocation, and global context feature generation.
[0072] (1) Multi-layer feature selection. Let S be the set of encoder layers participating in global context coding, and let the point features output by the l-th coding layer be... Represented as: (7) In equation (7), This represents the number of points in the l-th coding level. This indicates the number of feature channels at this encoding level. In this embodiment, shallow, medium, and deep features are selected as inputs to the C-VLAD module, so that the multi-layer features participating in global context encoding simultaneously contain local geometric information and high-level semantic information.
[0073] (2) Channel Mapping and Layer Weighting. Since the number of feature channels may differ at different coding layers, 1×1 convolution or an equivalent shared multilayer perceptron is used to map the features of each layer point by point, mapping them to a unified D-dimensional feature space, resulting in: (8) This represents the point features of the l-th coding layer after channel mapping. Let represent the feature mapping function corresponding to layer l. After unifying the feature dimensions, global pooling and feature mapping are performed on the unified dimension features of each layer to obtain the scale score of the corresponding layer; then, Softmax is used to normalize the encoding layers participating in the fusion to obtain the scale weights corresponding to each layer. : (9) Let represent the scale score corresponding to the t-th coding level, where t iterates through all coding levels participating in the fusion in set S. The uniform dimensional features of the corresponding coding levels are weighted according to the scale weights to obtain the weighted features of the levels: (10) In equations (9) and (10), For the set of encoder layers participating in the fusion, The scale score is obtained after global pooling and feature mapping for the features of the l-th encoding level. The normalized scale weights are used to represent the relative contributions of different encoding levels to the generation of global context features, so that shallow local structural information and deep semantic information participate in subsequent aggregation according to their respective weights.
[0074] (3) Soft assignment of local features. Suppose that the C-VLAD module contains K learnable cluster centers, and the soft assignment weight of the weighted feature of the i-th level in the l-th encoding level corresponding to the k-th cluster center is expressed as: (11) In equation (11), and For trainable parameters, This represents the learnable weight vector corresponding to the k-th cluster center (the superscript T indicates transpose). , Let represent the learnable bias terms corresponding to the k-th and j-th cluster centers, respectively. The soft-assigned weights are normalized among the K cluster centers, so that the same feature can participate in the residual aggregation of multiple cluster centers according to different weights, thereby preserving the continuous relationship between different local semantic patterns.
[0075] (4) VLAD Residual Aggregation. The feature residuals of each level and each point feature relative to each cluster center are calculated, and the feature residuals are weighted according to the scale weights corresponding to the encoding levels and the soft-assignment weights corresponding to the point features. For the k-th cluster center, the weighted feature residuals corresponding to all target encoding levels and each point are summed to obtain: (12) The residual aggregation result in Equation (12) is used to represent the comprehensive offset of point features at different coding levels relative to the k-th cluster center. That is, the residual aggregation vector of the k-th cluster center, which summarizes the weighted residuals of all coding levels and all point features relative to the k-th cluster center, and finally forms a vector of fixed dimensions. Let represent the coordinate vector of the k-th cluster center in the D-dimensional feature space. By performing the above residual aggregation on multiple cluster centers respectively, a global statistical description for different feature patterns can be formed.
[0076] Subsequently, the residual aggregation features corresponding to each cluster center were normalized, and all normalization results were concatenated to obtain a fixed-dimensional global context vector. The global context vector comprehensively includes local structural and semantic information from different encoding levels, which is used to supplement long-distance spatial associations that are difficult to be directly covered by local neighborhood aggregation.
[0077] (5) Global and local feature fusion. Based on the number of points contained in the current decoding layer, the global vector is fused... The features are copied to each point in the current decoding layer and concatenated with the upsampled features of the current decoding layer and the point features passed through skip connections in the corresponding encoding layer. The concatenated results are then feature-mapped to obtain the fused point features of the current decoding layer. The above processing is performed layer by layer in order of increasing spatial resolution to restore the spatial resolution of the point cloud.
[0078] (6) Training Method. The scale scoring parameters, soft assignment parameters, cluster centers, and subsequent fusion layer parameters in the C-VLAD module are jointly trained with the semantic segmentation backbone network and updated end-to-end based on the final point-by-point classification loss. Therefore, the cluster centers are not determined separately through offline clustering before training, but are learned jointly during the training of the semantic segmentation network. In the inference phase, the global context vector is generated by sequentially performing the above processing on the point cloud to be segmented, without the need for re-offline clustering.
[0079] To verify the role of the C-VLAD module in multi-layer semantic information fusion and global context representation, this embodiment sets up experiment b with the C-VLAD module removed, and compares experiment b with the complete network. The experimental results are shown in Table 2.
[0080] Table 2. Ablation Experiment Results of C-VLAD Module (%)
[0081] As shown in Table 2, after removing the C-VLAD module, the average intersection-union ratio (mIoU) decreased from 77.6% to 77.0%. The complete network achieved higher segmentation results in categories such as railway tracks, overhead contact lines, fences, and utility poles, especially improving the segmentation of utility poles from 63.2% to 73.5%. This indicates that the C-VLAD module can supplement global contextual information that is difficult to directly express through local neighborhood aggregation by fusing point features from different encoding levels. The complete network, with both the ALFA and C-VLAD modules enabled, achieved an average intersection-union ratio of 77.6%, demonstrating that local geometric feature enhancement and global contextual modeling can work synergistically within the same network.
[0082] like Figure 6 As shown, the real labels represent the actual distribution of the corresponding categories in the point cloud of the railway scene, Ours represents the semantic segmentation result of the complete network in this embodiment, and LACVNet and RandLANet represent the semantic segmentation results of the corresponding comparison networks. Figure 6 The semantic segmentation results for utility poles and signal poles are shown separately. Compared with the comparison results, this embodiment can more completely preserve the main structure of the pole-shaped facilities and reduce the omission of local structures.
[0083] In this embodiment, railway scene point clouds can sequentially complete data preprocessing, hierarchical encoding, adaptive local feature aggregation, multi-layer global context encoding, and point-by-point decoding and classification. Adaptive local feature aggregation dynamically adjusts the contribution of different neighboring points to the center point's features by combining the spatial relationships and point features of neighboring points, thereby enhancing the expression of local structures such as track edges, overhead contact lines, and pole-like facilities. Multi-layer global context encoding fuses local features from different encoding levels and supplements long-distance semantic relationships that are difficult to cover in local neighborhoods through hierarchical weighting and clustering residual aggregation. Therefore, this embodiment enables the semantic classification process to simultaneously utilize local geometric information, multi-layer semantic information, and overall point cloud context information, thereby improving the comprehensive semantic segmentation accuracy of complex railway scene point clouds.
[0084] Example 3 This embodiment provides an adaptive semantic segmentation device for railway point clouds. This device is used to execute the adaptive semantic segmentation method for railway point clouds described in the foregoing embodiments and can be deployed on a computer, server, workstation, or other computing platform with point cloud data processing capabilities.
[0085] The device includes a point cloud encoding module, a local feature aggregation module, a global context generation module, and a semantic decoding module.
[0086] The point cloud encoding module is used to acquire the point cloud of the railway scene to be segmented, and to perform hierarchical encoding on the point cloud to obtain point features at multiple encoding levels. Specifically, the point cloud encoding module can receive the railway scene point cloud after attribute processing and downsampling, perform initial feature mapping on the spatial coordinates and point attribute features of each point, and gradually reduce the spatial resolution of the point cloud and increase the semantic expression dimension of the point features through multiple encoding levels. The point features output by different encoding levels contain local structural information and semantic information at different spatial scales.
[0087] The local feature aggregation module is used to determine local neighborhoods for points in each coding level, generate local geometric features based on the spatial relationships between the point and its neighboring points, and fuse these local geometric features with the point features of the neighboring points. The local feature aggregation module further determines the aggregation weights corresponding to each neighboring point based on the fused neighboring point features, and aggregates the fused neighboring point features based on these aggregation weights to obtain the local aggregated features for each coding level. This module can employ the aforementioned adaptive local feature aggregation method, fusing attention weighting results with pooling results to enhance the feature representation of track edges, rod-shaped facilities, and other local structures.
[0088] The global context generation module receives local aggregated features from multiple coding levels, performs feature mapping on these features at different coding levels, and obtains multi-layer local features with a unified feature dimension. Subsequently, it determines the level weights corresponding to each coding level based on the multi-layer local features, and determines the assigned weights for each multi-layer local feature corresponding to multiple cluster centers. Based on the level weights and assigned weights, it aggregates the feature residuals of each multi-layer local feature relative to its corresponding cluster center to obtain global context features. These global context features represent the semantic information of different coding levels and the correlation information between different spatial regions in the railway scene point cloud.
[0089] The semantic decoding module fuses the global context features output by the global context generation module with the point features generated during the decoding process, and performs point feature upsampling and feature mapping layer by layer according to the point cloud spatial resolution from low to high. In each decoding layer, the semantic decoding module can also acquire point features from the corresponding encoding level in the point cloud encoding module, and fuse these encoding level point features, upsampled point features, and global context features to compensate for spatial details lost during encoding. After completing layer-by-layer decoding, the semantic decoding module determines the semantic category of each point based on the recovered point features, obtaining the semantic segmentation result of the railway scene point cloud.
[0090] In some implementations, the device can also invoke a pre-established projection index to map the semantic segmentation results of the downsampled point cloud to the corresponding points in the original railway scene point cloud, thereby forming a point-by-point semantic classification result of the original railway scene point cloud.
[0091] Each of the above modules can be implemented by a computer program stored in memory and executed by the processor. Each module can be set up separately, or it can be integrated into one or more software functional units according to actual deployment requirements. The module division is only used to describe the corresponding data processing functions and does not constitute a limitation on the specific physical structure of the device.
[0092] In this embodiment, the point cloud encoding module, local feature aggregation module, global context generation module, and semantic decoding module work together to sequentially complete the hierarchical feature extraction, local geometric feature aggregation, multi-layer global context modeling, and point-by-point semantic classification of railway scene point clouds. This allows the semantic segmentation results to utilize both local structural information and global semantic information, thereby improving the semantic segmentation accuracy of complex railway scene point clouds.
[0093] It should be understood that the various modules of the adaptive semantic segmentation device for railway point clouds provided in the above embodiments are only illustrated by the division of each functional module in the above description when performing adaptive semantic segmentation of point clouds. In practical applications, the above functions can be assigned to different functional modules as needed. That is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0094] The functional modules in the above embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of the embodiments of the present invention.
[0095] Based on the same application concept, embodiments of the present invention also provide a computer device, which may include a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the adaptive semantic segmentation method for railway point clouds as described above.
[0096] Based on the same concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the adaptive semantic segmentation method for railway point clouds as described above.
[0097] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects.
[0098] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0099] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An adaptive semantic segmentation method for railway point clouds, characterized in that, Includes the following steps: S1, Obtain the point cloud of the railway scene to be segmented, and perform hierarchical encoding on the point cloud of the railway scene to obtain point features at multiple encoding levels; S2, determine local neighborhoods for points in each coding level, generate local geometric features based on the spatial relationship between the point and each neighboring point in the local neighborhood, fuse the local geometric features with the point features of the neighboring points, determine the aggregation weights corresponding to each neighboring point based on the fused neighboring point features, and aggregate the fused neighboring point features based on the aggregation weights to obtain the local aggregated features of each coding level. S3, perform feature mapping on the local aggregated features of multiple coding levels to obtain multi-layer local features with a unified feature dimension, determine the level weights corresponding to each coding level based on the multi-layer local features, and determine the assigned weights of each cluster center in multiple cluster centers corresponding to the multi-layer local features, and aggregate the feature residuals of the multi-layer local features relative to each cluster center based on the level weights and the assigned weights to obtain global context features. S4. The global context features are fused with the point features in the decoding process to restore the spatial resolution of the railway scene point cloud layer by layer, and the semantic category of each point is determined according to the restored point features to obtain the semantic segmentation result of the railway scene point cloud.
2. The adaptive semantic segmentation method for railway point clouds according to claim 1, characterized in that, The railway scene point cloud to be segmented is obtained from the original railway scene point cloud through point cloud preprocessing, which includes: The intensity attribute values of each point in the original railway scene point cloud are obtained, points with negative intensity attribute values or values greater than a preset intensity upper limit are filtered out, and the intensity attribute values of the remaining points after filtering are normalized to obtain the attribute-processed point cloud. The attribute processing point cloud is divided into grids according to a preset spatial sampling interval, and the points in each non-empty grid are sampled to obtain a downsampled point cloud as the railway scene point cloud to be segmented. The preset spatial sampling interval is 0.06 meters. A spatial index is established based on the spatial coordinates of each point in the original railway scene point cloud, and a projection index is established based on the spatial correspondence between the points in the original railway scene point cloud and the points in the downsampled point cloud. After obtaining the semantic segmentation result, the semantic segmentation result is mapped to each point in the original railway scene point cloud according to the projection index.
3. The adaptive semantic segmentation method for railway point clouds according to claim 1, characterized in that, S2 includes local neighborhood determination and local geometric feature generation, specifically including: Using points in each coding level as center points, the Euclidean distances between other points in the same coding level and the center point are determined respectively. Then, a preset number of points are selected from the other points as neighborhood points in order of increasing Euclidean distances to obtain the local neighborhood of the center point. Obtain the spatial coordinates of the center point and the spatial coordinates of each neighboring point. Subtract the spatial coordinates of the center point from the spatial coordinates of the neighboring points to obtain the coordinate offset of each neighboring point relative to the center point. Determine the distance between each neighboring point and the center point based on the coordinate offset. The spatial coordinates of the center point, the spatial coordinates of the neighboring points, the coordinate offset, and the distance are concatenated, and the concatenation result is subjected to point-by-point feature mapping to obtain the local geometric features corresponding to each neighboring point. The local geometric features corresponding to each neighboring point are concatenated with the point features of the corresponding neighboring point, and the concatenation result is subjected to feature mapping to obtain the enhanced neighborhood features corresponding to each neighboring point.
4. The adaptive semantic segmentation method for railway point clouds according to claim 3, characterized in that, The S2 further includes neighborhood feature aggregation, which includes: Feature mapping is performed on each enhanced neighborhood feature to obtain the attention score corresponding to each neighborhood point; The attention scores of each neighboring point are normalized within the local neighborhood corresponding to the same center point to obtain the attention weights corresponding to each neighboring point. The attention weights are used to weight each enhanced neighborhood feature, and the weighted enhanced neighborhood features are summed to obtain the attention aggregation feature. The pooled aggregate features are obtained by comparing the feature values of each enhanced neighborhood feature corresponding to the same center point in each feature channel and retaining the maximum feature value in each feature channel. The attention aggregation feature and the pooling aggregation feature are concatenated, and the concatenation result is subjected to feature mapping to obtain the local aggregation feature corresponding to the center point.
5. The adaptive semantic segmentation method for railway point clouds according to claim 1, characterized in that, S3 includes determining hierarchical weights, which includes: Select multiple target coding levels with different spatial resolutions from the plurality of coding levels; Pointwise feature mapping is performed on the local aggregated features of each target encoding level to map the local aggregated features of each target encoding level to the same feature dimension, thereby obtaining the unified dimension features corresponding to each target encoding level; Global pooling is performed on the uniform dimension features corresponding to each target encoding level to obtain the hierarchical description features corresponding to each target encoding level; Feature mapping is performed on the descriptive features of each level to obtain the level score corresponding to each target encoding level; The scores of each target coding level are normalized across the multiple target coding levels to obtain the level weights corresponding to each target coding level.
6. The adaptive semantic segmentation method for railway point clouds according to claim 5, characterized in that, S3 further includes global context feature generation, which includes: A preset number of trainable cluster centers are set, and the feature dimension of each cluster center is the same as the feature dimension of the unified dimension feature. For each unified dimension feature in each target encoding level, the allocation score of the unified dimension feature corresponding to each cluster center is determined, and the allocation score is normalized among the multiple cluster centers to obtain the soft allocation weight of the unified dimension feature corresponding to each cluster center. The feature residuals of each unified dimension feature relative to each cluster center are obtained by subtracting each unified dimension feature from each cluster center. Based on the hierarchical weights corresponding to each target encoding level and the soft-assigned weights corresponding to each unified dimension feature, the residuals of each feature are weighted, and for each cluster center, the weighted feature residuals corresponding to different target encoding levels and different points are accumulated to obtain the residual aggregated features corresponding to each cluster center. The residual aggregated features corresponding to each cluster center are normalized, and the normalized residual aggregated features are concatenated to obtain the global context features.
7. The adaptive semantic segmentation method for railway point clouds according to claim 1, characterized in that, S4 includes point feature decoding and semantic classification, specifically including: In each decoding layer, the global context features are copied according to the number of points contained in the decoding layer to obtain the global context point features corresponding to each point in the decoding layer; Upsample the point features output from the previous decoding layer to obtain the upsampled point features corresponding to the current decoding layer; Obtain the coding layer point features corresponding to the current decoding layer during the coding phase, concatenate the global context point features, the upsampled point features, and the coding layer point features, and perform feature mapping on the concatenation result to obtain the fusion point features of the current decoding layer; Following the order of point cloud spatial resolution from low to high, upsampling and feature stitching are performed layer by layer to obtain the restored point features; The recovered point features are mapped to the category features corresponding to the target semantic category, and the semantic category of each point is determined according to the category features to obtain the semantic segmentation result of the railway scene point cloud.
8. An adaptive semantic segmentation device for railway point clouds, characterized in that, It includes a point cloud encoding module, a local feature aggregation module, a global context generation module, and a semantic decoding module; The point cloud encoding module acquires the point cloud of the railway scene to be segmented, performs hierarchical encoding on the point cloud of the railway scene, and obtains point features at multiple encoding levels. The local feature aggregation module determines local neighborhoods for points in each coding level, generates local geometric features based on the spatial relationship between the point and each neighboring point in the local neighborhood, fuses the local geometric features with the point features of the neighboring points, determines the aggregation weight corresponding to each neighboring point based on the fused neighboring point features, and aggregates the fused neighboring point features based on the aggregation weights to obtain the local aggregated features of each coding level. The global context generation module performs feature mapping on the local aggregated features of multiple coding levels to obtain multi-layer local features with a unified feature dimension. Based on the multi-layer local features, it determines the level weights corresponding to each coding level and the allocation weights of the multi-layer local features to each cluster center in multiple cluster centers. Based on the level weights and the allocation weights, it aggregates the feature residuals of the multi-layer local features relative to each cluster center to obtain global context features. The semantic decoding module fuses the global context features with the point features in the decoding process, restores the spatial resolution of the railway scene point cloud layer by layer, and determines the semantic category of each point based on the restored point features, thereby obtaining the semantic segmentation result of the railway scene point cloud.
9. An electronic device, characterized in that, Includes a processor and a memory communicatively connected to the processor; The memory stores a computer program that can be executed by the processor. When the computer program is executed by the processor, it causes the processor to implement the adaptive semantic segmentation method for railway point clouds as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the adaptive semantic segmentation method for railway point clouds as described in any one of claims 1-7.