A subway tunnel laser point cloud semantic feature fusion enhancement method

By constructing a deep learning network model consisting of a voxel self-attention feature perception module and a spatial geometric relationship feature extraction module, the problem of complex and highly overlapping facility distribution in subway tunnels was solved, achieving more efficient tunnel facility identification and classification.

CN117079091BActive Publication Date: 2025-11-18CHINA UNIV OF MINING & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311052594.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-18
Publication Date
2025-11-18
Estimated Expiration
2043-08-18

AI Technical Summary

Technical Problem

In existing technologies, the facilities in subway tunnels are densely distributed and highly overlapping, resulting in poor isotropy of the edge features of the tunnel facilities, complex neighborhood features with a large amount of information, which makes it difficult for existing deep learning network models to make full use of them and fail to effectively utilize the spatial geometric relationship features of the tunnel facilities.

Method used

A deep learning network model based on a voxel self-attention feature perception module and a spatial geometric relationship feature extraction module is constructed. The deep semantic features of the point cloud are extracted through the voxel self-attention feature perception module and fused with the spatial geometric relationship features of the tunnel facilities to improve the accuracy of point cloud semantic segmentation.

Benefits of technology

It improves the ability to perceive and utilize the feature information of point clouds of subway tunnel facilities, reduces the loss of local features caused by the large number of facilities and high spatial overlap, and improves the accuracy of facility identification and classification precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117079091B_ABST
    Figure CN117079091B_ABST
Patent Text Reader

Abstract

The application discloses a subway tunnel laser point cloud semantic feature fusion enhancement method, constructs a deep learning network model based on voxelization self-attention mechanism and tunnel space geometric features, extracts deep semantic features and spatial geometric relationship features of the tunnel point cloud, constructs a point cloud voxelization index, applies a feature extraction method based on the self-attention mechanism in the voxel grid, extracts the deep semantic features of the point cloud, calculates the spatial geometric features of the tunnel facilities by using a spatial feature module, fuses the deep semantic features and the spatial geometric features of the tunnel facilities, trains the fused features as the encoding features of the deep learning network model, and obtains a semantic segmentation network model of the tunnel facility point cloud. The application can improve the semantic information extraction and utilization efficiency of the subway tunnel laser point cloud, reduce the local feature loss caused by a large number of facilities and high spatial overlap, and improve the facility identification and statistical accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for fusion and enhancement of semantic features of three-dimensional point clouds, specifically a method for fusion and enhancement of semantic features of laser point clouds in subway tunnels, belonging to the field of three-dimensional point cloud processing technology for rail transit. Background Technology

[0002] With the continuous advancement of urbanization in my country, the mileage of subway lines has increased rapidly. The main work of urban subway projects has shifted from the construction phase to the operation and maintenance phase. The inspection and management of tunnel facilities in operating subways is crucial for ensuring passenger safety and the safe operation of the subway system. Traditional tunnel facility inspections and defect detection are usually done manually, with few integrated automated inspection equipment available for detecting tunnel cracks and deformations. These methods suffer from drawbacks such as long inspection times, low efficiency, and significant susceptibility to the subjectivity of inspectors, failing to meet current development needs.

[0003] In recent years, deep learning-based point cloud processing technologies have developed rapidly. In special engineering scenarios such as tunnels and underground spaces, 3D point cloud data not only solves problems such as pose and lighting often encountered when processing 2D images, but also reveals rich spatial information for complex scenes. Deep learning-based point cloud semantic segmentation technology can classify target point cloud data into different categories based on the original and derived feature information of the point cloud, enabling the extraction and recognition of objects. For large-scale point cloud data in subway tunnel scenarios, 3D point cloud semantic segmentation technology can quickly identify and classify facilities such as tunnel segments, rails, tracks, and overhead contact line rails, assisting in tunnel inspection and facility maintenance. However, the application of deep learning point cloud segmentation for tunnel facility identification and inspection still faces the following problems in existing technologies:

[0004] 1. The facilities within subway tunnels are densely distributed and highly overlapping, resulting in poor isotropy of the edge features of these facilities. The neighborhood features of the tunnel facility point cloud are complex and contain a large amount of information. Existing deep learning network models struggle to fully utilize this information.

[0005] 2. The spatial geometric features of subway tunnels have a relatively stable relationship with the spatial location of internal facilities. The relative positions of tunnel facilities on the inner wall are important spatial geometric features, but existing semantic segmentation methods do not utilize these features. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides a method for enhancing the semantic features of laser point clouds in subway tunnels through fusion. This method can improve the shortcomings of existing technologies in the semantic segmentation of tunnel point clouds, such as insufficient utilization of neighborhood features and spatial geometric relationship features of facility point clouds, and low segmentation accuracy of point clouds for complex facilities. It can provide more accurate data support for facility management and safety inspection of subway tunnels.

[0007] To achieve the above objectives, the local subway tunnel laser point cloud semantic feature fusion enhancement method specifically includes the following steps:

[0008] S1: Construct an encoder-decoder deep learning network model based on the voxel self-attention feature perception and extraction module to extract deep semantic features of tunnel laser point clouds;

[0009] S2: Construct a spatial geometric relationship feature extraction module, and use the spatial positional relationship between tunnel facilities and tunnel geometric axes to calculate the spatial geometric feature information of point cloud;

[0010] S3: The spatial geometric features of the point cloud are fused with the deep semantic features extracted by the voxel self-attention feature perception extraction module to obtain point cloud semantic features that include geometric dimension features and deep semantic dimension features. The enhanced point cloud semantic features replace the original semantic features and are used as input features for the decoding layer of the network model.

[0011] S4: The enhanced point cloud semantic features are used as input features to the network model decoding layer. The shallow features of the encoding layer are skipped and connected. The features are propagated through distance difference, and the spatial distribution of the original point cloud is restored layer by layer. A loss function is established to backpropagate the network for training and obtain the semantic segmentation network model of the tunnel facility point cloud.

[0012] In step S1, the encoding network of the deep learning network model constructed based on the voxel self-attention feature perception extraction module consists of a basic framework composed of multiple layers of semantic feature perception extraction modules based on the voxel self-attention mechanism. During the feature extraction process, the segmented voxel size of the voxel self-attention feature perception extraction module is gradually increased, and the semantic feature information output by the previous layer is used as the input feature to extract the deep semantic features of the point cloud. At the same time, the input features are transformed in dimension by a linearization layer and added to the deep semantic features. The voxel self-attention feature perception extraction module with multiple layers is used to complete the point cloud feature extraction of the encoding network and obtain the deep semantic feature information of the point cloud.

[0013] The voxel self-attention feature perception and extraction module consists of a voxelized point cloud retrieval layer, a self-attention feature extraction layer, and a linearized feature connection layer. It performs voxelized segmentation of the point cloud's 3D space to achieve point cloud indexing and neighborhood retrieval. The self-attention mechanism layer extracts semantic feature information of the point cloud, representing attribute-related features between points. The point cloud features extracted by the voxel self-attention perception module are represented as follows:

[0014] F vsa =σ(M A (F i )+γ(F i ))

[0015] In the formula, Fi Voxel v i Features of the point cloud, where γ represents a linear transformation operation, transforming the features F i Dimensional transformation to M A The output dimension is consistent, σ represents mapping the output features to the original voxel space, and M A This indicates that self-attention feature extraction is performed on the point cloud input features within a voxel;

[0016] The voxelized point cloud retrieval layer divides the three-dimensional space of the point cloud into voxel sizes based on the spatial range D, H, and W along the X, Y, and Z axes of the coordinate system. The quantity is D k ′×H k ′×W k Three-dimensional voxels of ′:

[0017]

[0018] In the formula, k is the layer number of the voxel self-attention perception module designed with multiple layers;

[0019] The self-attention feature extraction layer contains voxel V i The point cloud within is used to perform feature extraction point by point using the self-attention principle. j Deep extraction:

[0020]

[0021] σ=ω(p i ,p j )

[0022] In the formula, f j with f i Represent the input features of the query point and the remaining points within the voxel, respectively, ψ and denoted as mlp mapping of the query point and its neighboring points respectively, α represents linear transformation of the features, ω represents mlp mapping of the spatial geometric distance between the two points, τ represents mlp mapping of the neighborhood relationship information of the query point, ξ represents normalization function, and ⊙ represents feature-wise multiplication.

[0023] The voxel space obtained by the voxelization point cloud retrieval layer gradually increases in size, and the point cloud density within each voxel gradually decreases. For the point cloud within the voxel grid obtained by the first layer of voxelization segmentation, downsampling is performed using the farthest point sampling method to ensure that each voxel in the first layer contains an initial number N of point clouds. v1 =t;

[0024] In the multi-layer stacked voxel self-attention feature perception and extraction module, the size of the segmented voxels gradually increases, and the voxel sizes of adjacent (k-1)th and kth layers have the following relationship:

[0025]

[0026] After downsampling, the number of point clouds contained in each voxel has the following relationship:

[0027] N vk =N vk-1 =t.

[0028] In step S2, the tunnel's geometric axis is extracted using a three-dimensional spatial projection method. A query-reference point spatial distance threshold σ is set, and a reference point p on the central axis is determined based on this threshold. c Based on the fitting direction of the tunnel point cloud to be queried, the direction of the reference point pointing to the tunnel clearance, and the coordinates of the two points, a spatial axis relationship reference coordinate system is established, and the spatial geometric related positional features of the tunnel point cloud are obtained through calculation.

[0029] Reference point p c It can be represented as

[0030] ||p q -l i ||2-MAX(x q -x i ,y q -y i ,z q -z i ) < μ

[0031] In the formula, p q For the point to be queried, l i To fit the central axis L={l i =[x i ,y i ,z i ]∈R 3} i=1...n The point on the line with parameter μ = 0.002 satisfies the above relationship. i For reference point p c .

[0032] The spatial geometric relationship features of the point cloud are extracted as follows:

[0033] F eu (p q = [β,γ,σ]∈R 3 p q ∈V1

[0034] β=arccos(v·n t )

[0035]

[0036] σ=arctan(w·n t,u·n t )

[0037] u=n q

[0038]

[0039] w = u × v

[0040] In the formula, p q For the point cloud containing the geometric information to be queried, n q p is the unit vector of the direction of the fitting method for the query point. c Let n be the intersection of the tunnel section where the query point is located and the fitted axis of the tunnel. t For p c A unit vector pointing in the direction of the tunnel's clearance height.

[0041] In step S3, the spatial geometric features of the point cloud are fused with the deep semantic features extracted by the voxel self-attention feature perception extraction module, and represented as follows:

[0042] F df =fc(ο(F,F) sp ))

[0043] F sp =m(F eu )

[0044] In the formula, m represents the mlp layer, F eu F represents the spatial geometric relationship features extracted from the original point cloud, F represents the point cloud features extracted by the voxel self-attention perception layer, ο represents the concatenation of features according to the point cloud index, and fc is a fully connected layer.

[0045] In step S4, the enhanced point cloud semantic features are input into the decoding layer of the network model. The decoding layer completes feature propagation and point cloud spatial distribution recovery in the following order:

[0046] Based on the distance difference algorithm, the point cloud information input to the decoding layer is upsampled to enhance the point cloud semantic features, so that the spatial distribution of the point cloud contained in the decoding layer is consistent with the point cloud contained in the shallow features output by the k-th coding layer.

[0047] The upsampled enhanced semantic features are concatenated with the shallow semantic features, and the features are fused and the feature channel dimensions are adjusted through a self-attention extraction layer and a linearization layer.

[0048] The output features are input into the next layer of the decoding module, upsampled, and then fused with the features of the (k-1)th coding layer.

[0049] By analogy, the spatial distribution of the sparse point cloud containing enhanced features is restored to that of the original point cloud.

[0050] The loss function for building a deep learning network model and performing backpropagation training is established according to the following steps:

[0051] After the decoding layer restores the point cloud to its original spatial distribution, it outputs the predicted distribution of the point cloud's category labels and calculates the loss function value of the point cloud prediction results, expressed as:

[0052]

[0053] In the formula, p represents the true value of the label distribution corresponding to the point cloud, q represents the predicted value of the point cloud label, and L is the number of labels for the point cloud facility type;

[0054] During training, the loss function value of the network model is calculated based on the number of sample point clouds input to the network model for training in one go.

[0055]

[0056] In the formula, N is the number of sample point clouds used to train the input network model.

[0057] Compared with existing technologies, the local subway tunnel laser point cloud semantic feature fusion enhancement method has the following advantages:

[0058] (1) Due to the complex neighborhood features and large amount of information of the point cloud of tunnel facilities, and the poor isotropic edge features of tunnel facilities, the existing deep learning network models are difficult to fully utilize. This invention proposes to use the feature encoding layer of the network model constructed by the voxel self-attention perception module to extract and aggregate the deep semantic features of the point cloud in the ever-expanding voxel space, so that the rich neighborhood feature information of the point cloud can be perceived and utilized.

[0059] (2) Considering that the spatial shape of the subway tunnel and the spatial position of the internal facilities are relatively fixed, the relative position of the tunnel facilities on the inner wall is an important spatial geometric feature. This invention further constructs a spatial geometric relationship feature extraction module, which integrates and enhances the spatial geometric features of the point cloud with the deep semantic features, thereby improving the ability of the deep learning network model to perceive and utilize the feature information of the tunnel facility point cloud. This can improve the efficiency of semantic information extraction and utilization of the laser point cloud of the subway tunnel, reduce the loss of local features caused by the large number of facilities and high spatial overlap, improve the accuracy of facility identification statistics, and improve the problems of poor classification accuracy of tunnel facilities and low utilization rate of spatial distribution and geometric features of facilities in the existing technology. Attached Figure Description

[0060] Figure 1 This is a diagram illustrating the deep learning network model of the present invention;

[0061] Figure 2 This is a flowchart of the present invention;

[0062] Figure 3 This is a diagram illustrating the voxel self-attention feature perception and extraction module of the present invention;

[0063] Figure 4 This is a diagram illustrating the point cloud voxelization retrieval layer of the present invention;

[0064] Figure 5 This is a diagram of the subway tunnel point cloud dataset used to train the network model in this invention. Figure (a) shows the labeled overall tunnel, Figure (b) shows the track bed facilities, Figure (c) shows the track facilities, Figure (d) shows the tunnel wall pipe rack fixing facilities, Figure (e) shows the track facilities, Figure (f) shows the tunnel fire protection pipeline facilities, Figure (g) shows the catenary guide rail, and Figure (h) shows the tunnel wall segments.

[0065] Figure 6 This is a diagram illustrating the spatial geometric relationship features of facility point cloud computing according to the present invention;

[0066] Figure 7 This is a diagram illustrating the feature propagation decoding module of the present invention. Detailed Implementation

[0067] The present invention will be further described below with reference to the accompanying drawings.

[0068] This invention provides a method for semantic feature fusion and enhancement of laser point clouds in subway tunnels, such as... Figure 1 As shown, the method includes the following parts:

[0069] like Figure 1 As shown in the upper dashed box, to improve the perception ability of point cloud neighborhood feature information and the utilization of tunnel spatial geometric features, this invention proposes a feature extraction and encoding layer composed of a voxelized self-attention module and a spatial geometric feature module. The main structure of the network model's feature extraction and encoding structure consists of three layers of voxel self-attention feature perception and extraction modules. The original point cloud information of the tunnel is mapped to the feature space and then input into the voxel self-attention feature extraction layer. The self-attention feature extraction layer extracts the next-layer neighborhood feature information of the point cloud from the gradually expanding voxel grid and generates point cloud depth semantic information. The number of points in the point cloud feature information output by the voxel self-attention feature perception and extraction module gradually decreases, while the number of feature information channels gradually increases. The spatial geometric feature calculation module calculates the spatial geometric position features of the point cloud based on the original point cloud coordinates and fuses them with the depth semantic features extracted by the voxel self-attention feature perception and extraction module to enhance the point cloud feature information. This enhanced point cloud feature information is then used as the input feature of the decoding network.

[0070] like Figure 1As shown in the lower dashed box, the network decoding layer consists of feature propagation decoding modules corresponding to the encoding modules. These modules take enhanced point cloud feature information as input and use distance difference upsampling to restore the sparse feature points to the spatial distribution and feature channel number of the corresponding encoding module's point cloud. The encoding layer uses shallow features for feature addition to solve the gradient vanishing problem in deep feature extraction. A training result loss function is established based on the cross-entropy principle. The loss value between the predicted result and the ground truth of the trained network model is calculated. Backpropagation is used to update the parameters of the trained network model and predict the facility labels corresponding to the tunnel point cloud.

[0071] Through iterative training, a deep learning tunnel point cloud semantic segmentation model is obtained, which can be applied to the identification, extraction, inspection and management of facilities in tunnels.

[0072] Specifically, such as Figure 2 As shown, the semantic feature fusion and enhancement method for local subway tunnel laser point clouds specifically includes the following steps:

[0073] S1: Construct an encoder-decoder deep learning network model based on the voxel self-attention feature perception extraction module to extract deep semantic features of tunnel laser point clouds.

[0074] S2: Construct a spatial geometric relationship feature extraction module, which uses the spatial positional relationship between tunnel facilities and tunnel geometric axes to calculate the spatial geometric feature information of point cloud.

[0075] S3: The spatial geometric features of the point cloud are fused with the deep semantic features extracted by the voxel self-attention feature perception extraction module to obtain point cloud semantic features that include both geometric and deep semantic dimension features. The enhanced point cloud semantic features replace the original semantic features and are used as input features for the decoding layer of the network model.

[0076] S4: The enhanced point cloud semantic features are used as input features to the network model decoding layer. The shallow features of the encoding layer are skipped and connected. The features are propagated through distance difference, and the spatial distribution of the original point cloud is restored layer by layer. A loss function is established to backpropagate the network for training and obtain the semantic segmentation network model of the tunnel facility point cloud.

[0077] In step S1, the encoding network of the deep learning network model constructed based on the voxel self-attention feature perception and extraction module consists of a basic framework of multiple layers of semantic feature perception and extraction modules based on the voxel self-attention mechanism. Taking a four-layer semantic feature perception and extraction module based on the voxel self-attention mechanism as an example, the feature extraction framework is composed of four layers of voxel self-attention feature perception and extraction modules. During the feature extraction process, the voxel self-attention feature perception and extraction module gradually increases the size of the voxel segmentation unit and performs downsampling operations on the point cloud within the segmented voxels, gradually generating a point cloud set with sparse spatial distribution, increased number of feature channels, and richer neighborhood feature information. For the point cloud p input to the network... i ∈R 3 The point cloud features within a single voxel output by the feature extraction module are F1={f i ∈R 8} i=1...t F2={f i ∈R 32} i=1...4t F3 = {f i ∈R 128} i=1...16t F4 = {f h ∈R 512} h=1...t , where f is the feature of a point cloud within a voxel after perceptual extraction, and t is the number of point clouds retained after downsampling of each voxel grid in the semantic feature extraction module.

[0078] like Figure 3 As shown, the voxel self-attention feature perception and extraction module consists of a voxelized point cloud retrieval layer, a self-attention feature extraction layer, and a linearized feature connection layer. It performs voxelized segmentation on the 3D space of the point cloud to achieve point cloud indexing and neighborhood retrieval. The self-attention mechanism layer extracts point cloud semantic features that represent attribute-related features between point clouds. The point cloud features extracted by the voxel self-attention perception module are represented as follows:

[0079] F vsa =σ(M A (F i )+γ(F i ))

[0080] In the formula, F i Voxel v i Features of the point cloud, where γ represents a linear transformation operation, transforming the features F i Dimensional transformation to M A The output dimension is consistent, σ represents mapping the output features to the original voxel space, and M A This indicates that self-attention feature extraction is performed on the point cloud input features within the voxel.

[0081] like Figure 4 As shown, the voxelized point cloud retrieval layer divides the three-dimensional space of the point cloud into voxel sizes based on the spatial range D, H, and W along the X, Y, and Z axes of the coordinate system. The quantity is D′ k ×H′ k ×W′ k Three-dimensional space voxels:

[0082]

[0083] In the formula, k is the layer number of the voxel self-attention perception module in the multi-layer stacked design.

[0084] The voxel space obtained by the voxelization point cloud retrieval layer gradually increases in size, while the point cloud density within each voxel gradually decreases. For the point cloud within the voxel grid obtained from the first layer of voxelization, downsampling is performed using the farthest point sampling method to ensure that each voxel in the first layer contains an initial number N of point clouds. v1 =t.

[0085] In the multi-layer stacked voxel self-attention feature perception and extraction module, the size of the segmented voxels gradually increases. The voxel sizes of adjacent (k-1)th and kth layers have the following relationship:

[0086]

[0087] After downsampling, the number of point clouds contained in each voxel has the following relationship:

[0088] N vk =N vk-1 =t

[0089] The self-attention feature extraction layer contains voxel V i The point cloud within is used to perform feature extraction point by point using the self-attention principle. j Deep extraction:

[0090]

[0091] σ=ω(p i ,p j )

[0092] In the formula, f j with f i Represent the input features of the query point and the remaining points within the voxel, respectively, ψ and denoted as mlp mapping of the query point and its neighboring points respectively, α represents linear transformation of the features, ω represents mlp mapping of the spatial geometric distance between the two points, τ represents mlp mapping of the neighborhood relationship information of the query point, ξ represents normalization function, and ⊙ represents feature-wise multiplication.

[0093] The point cloud dataset was input into the constructed network model for feature extraction. The point cloud dataset was created from measured point cloud data of a subway tunnel. A scalar field named "label" was created for the point cloud data to mark the point clouds corresponding to different facilities. The facilities within the tunnel were divided into seven categories: track bed, track, pipe rack, line, fire protection pipeline, overhead contact line rail, and tunnel segment. A numerical value (0-6) was assigned to the "label" field based on the category of the point cloud as the corresponding facility type label. The point cloud dataset is shown below. Figure 5 As shown.

[0094] The spatial axis relationship reference coordinate system of the spatial geometric relationship feature extraction module in step S2 is as follows: Figure 6 As shown, the geometric axis of the tunnel is extracted using the spatial three-dimensional projection method, and a query-reference point spatial distance threshold σ is set. Based on the distance threshold, a reference point p on the central axis is determined. c p c It can be represented as

[0095] ||p q -l i ||2-MAX(x q -x i ,y q -y i ,z q -z i ) < μ

[0096] In the formula, p q For the point to be queried, l i To fit the central axis L={l i =[x i ,y i ,z i ]∈R 3} i=1...n The point on the line with parameter μ = 0.002 satisfies the above relationship. i That is, the reference point p c .

[0097] Based on the fitting direction of the tunnel point cloud to be queried, the direction of the reference point pointing towards the tunnel's net height, and the coordinates of the two points, a spatial axis reference coordinate system is established. The spatial geometrical positional features of the tunnel point cloud are then calculated. The spatial geometrical relationship features of the point cloud are extracted based on the reference coordinate system and the central axis reference point as follows:

[0098] F eu (p q = [β,γ,σ]∈R 3 p q ∈V1

[0099] β=arccos(v·n t )

[0100]

[0101] σ=arctan(w·n t ,u·n t )

[0102] u=n q

[0103]

[0104] w = u × v

[0105] In the formula, p q For the point cloud containing the geometric information to be queried, n q p is the unit vector of the direction of the fitting method for the query point. c Let n be the intersection of the tunnel section where the query point is located and the fitted axis of the tunnel. t For p c A unit vector pointing in the direction of the tunnel's clearance height.

[0106] In step S3, the spatial geometric features of the point cloud are fused with the deep semantic features extracted by the voxel self-attention feature perception extraction module, and represented as follows:

[0107] F df =fc(ο(F,F) sp ))

[0108] F sp =m(F eu )

[0109] In the formula, m represents the mlp layer, F eu F represents the spatial geometric relationship features extracted from the original point cloud, F represents the point cloud features extracted by the voxel self-attention perception layer, ο represents the concatenation of features according to the point cloud index, and fc is a fully connected layer.

[0110] In step S4, the enhanced point cloud semantic features are input into the network model decoding layer to complete feature propagation and recovery of the point cloud spatial distribution, and a loss function is established. The model nodes are then trained according to the backpropagation algorithm.

[0111] The decoding layer completes feature propagation and point cloud spatial distribution restoration in the following order:

[0112] Based on the distance difference algorithm, the point cloud information input to the decoding layer is upsampled to enhance the point cloud semantic features, so that the spatial distribution of the point cloud contained in the decoding layer is consistent with the point cloud contained in the shallow features output by the k-th coding layer.

[0113] The upsampled enhanced semantic features are concatenated with the shallow semantic features, and the features are fused and the feature channel dimensions are adjusted through a self-attention extraction layer and a linearization layer.

[0114] The output features are input into the next layer of the decoding module, upsampled, and then fused with the features of the (k-1)th coding layer.

[0115] By analogy, the spatial distribution of the sparse point cloud containing enhanced features is restored to that of the original point cloud.

[0116] The decoding layer consists of three feature propagation decoding modules. It upsamples the point cloud information input to the decoding layer using a distance difference algorithm, concatenates it with shallow semantic features, and then fuses the features and adjusts the feature channel dimensions through a self-attention extraction layer and a linearization layer. This restores the spatial distribution of the original point cloud from the sparse enhanced point cloud features.

[0117] Feature propagation decoding module, such as Figure 7 As shown, the point cloud feature information output by the feature propagation decoding module can be represented as: F O ={f i ∈R 128} i=1...N , where f is the point cloud feature obtained within the voxel through feature propagation, and has the same spatial distribution as the original point cloud.

[0118] After the decoding layer restores the point cloud to its original spatial distribution, it calculates the loss function value of the point cloud prediction result based on the predicted distribution of point cloud category labels obtained from the model's forward propagation, which is expressed as:

[0119]

[0120] In the formula, p represents the true value of the label distribution corresponding to the point cloud, q represents the predicted value of the point cloud label, and L is the number of labels for the point cloud facility type.

[0121] During training, the loss function value of the network model is calculated based on the number of sample point clouds input to the network model for training in one go.

[0122]

[0123] In the formula, N is the number of sample point clouds input to the network model training.

[0124] This method for enhancing the semantic features of laser point clouds in subway tunnels utilizes a voxel self-attention perception module constructed using voxelization retrieval and self-attention feature extraction mechanisms. Within an ever-expanding voxel space, it extracts and aggregates the deep semantic features of the point cloud, enabling the perception and utilization of rich neighborhood feature information. Simultaneously, a spatial geometric relationship feature extraction module is constructed to fuse and enhance the spatial geometric features of the point cloud with the deep semantic features. This improves the ability of deep learning network models to perceive and utilize the feature information of subway tunnel facility point clouds, addressing the problems of poor classification accuracy and low utilization rate of spatial distribution and geometric features of existing technologies.

Claims

1. A method for semantic feature fusion and enhancement of laser point clouds in subway tunnels, characterized in that, Specifically, the steps include: S1: Construct an encoder-decoder deep learning network model based on the voxel self-attention feature perception and extraction module to extract deep semantic features of tunnel laser point clouds; The voxel self-attention feature perception and extraction module consists of a voxelized point cloud retrieval layer, a self-attention feature extraction layer, and a linearized feature connection layer. It performs voxelized segmentation of the point cloud's 3D space to achieve point cloud indexing and neighborhood retrieval. The self-attention mechanism layer extracts semantic feature information of the point cloud, representing attribute-related features between points. The point cloud features extracted by the voxel self-attention perception module are represented as follows: F vsa =σ(M A (F i )+γ(F i )) In the formula, F i Voxel v i Features of the point cloud, where γ represents a linear transformation operation, transforming the features F i Dimensional transformation to M A The output dimension is consistent, σ represents mapping the output features to the original voxel space, and M A This indicates that self-attention feature extraction is performed on the point cloud input features within a voxel; The voxelized point cloud retrieval layer divides the three-dimensional space of the point cloud into voxel sizes based on the spatial range D, H, and W along the X, Y, and Z axes of the coordinate system. The quantity is D k ′×H k ′×W k Three-dimensional voxels of ′: In the formula, k is the layer number of the voxel self-attention perception module designed with multiple layers; The self-attention feature extraction layer contains voxel V i The point cloud within is used to perform feature extraction point by point using the self-attention principle. j Deep extraction: σ=ω(p i ,p j ) In the formula, f j with f i Represent the input features of the query point and the remaining points within the voxel, respectively, ψ and denoted as mlp mapping of the query point and its neighboring points respectively, α represents linear transformation of the features, ω represents mlp mapping of the spatial geometric distance between the two points, τ represents mlp mapping of the neighborhood relationship information of the query point, ξ represents normalization function, and ⊙ represents feature-wise multiplication; S2: Construct a spatial geometric relationship feature extraction module, and use the spatial positional relationship between tunnel facilities and tunnel geometric axes to calculate the spatial geometric feature information of point cloud; S3: The spatial geometric features of the point cloud are fused with the deep semantic features extracted by the voxel self-attention feature perception extraction module to obtain point cloud semantic features that include geometric dimension features and deep semantic dimension features. The enhanced point cloud semantic features replace the original semantic features and are used as input features for the decoding layer of the network model. S4: The enhanced point cloud semantic features are used as input features to the network model decoding layer. The shallow features of the encoding layer are skipped and connected. The features are propagated through distance difference, and the spatial distribution of the original point cloud is restored layer by layer. A loss function is established to backpropagate the network for training and obtain the semantic segmentation network model of the tunnel facility point cloud.

2. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 1, characterized in that, In step S1, the encoding network of the deep learning network model constructed based on the voxel self-attention feature perception extraction module consists of a basic framework composed of multiple layers of semantic feature perception extraction modules based on the voxel self-attention mechanism. During the feature extraction process, the segmented voxel size of the voxel self-attention feature perception extraction module is gradually increased, and the semantic feature information output by the previous layer is used as the input feature to extract the deep semantic features of the point cloud. At the same time, the input features are transformed in dimension by a linearization layer and added to the deep semantic features. The voxel self-attention feature perception extraction module with multiple layers is used to complete the point cloud feature extraction of the encoding network and obtain the deep semantic feature information of the point cloud.

3. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 1, characterized in that, The voxel space obtained by the voxelization point cloud retrieval layer gradually increases in size, and the point cloud density within each voxel gradually decreases. For the point cloud within the voxel grid obtained by the first layer of voxelization segmentation, downsampling is performed using the farthest point sampling method to ensure that each voxel in the first layer contains an initial number N of point clouds. v1 =t; In the multi-layer stacked voxel self-attention feature perception and extraction module, the size of the segmented voxels gradually increases, and the voxel sizes of adjacent (k-1)th and kth layers have the following relationship: After downsampling, the number of point clouds contained in each voxel has the following relationship: N vk =N vk-1 =t。 4. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 1, characterized in that, In step S2, the tunnel's geometric axis is extracted using a three-dimensional spatial projection method, and a query-reference point spatial distance threshold σ is set. Based on this threshold, a reference point p on the central axis is determined. c Based on the fitting direction of the tunnel point cloud to be queried, the direction of the reference point pointing to the tunnel clearance, and the coordinates of the two points, a spatial axis relationship reference coordinate system is established, and the spatial geometric related positional features of the tunnel point cloud are obtained through calculation.

5. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 4, characterized in that, Reference point p c It can be represented as ||p q -L i ||2-MAX(x q -x i ,y q -y i ,z q -z i )<μ In the formula, p q For the point to be queried, l i To fit the central axis L = {l i =[x i ,y i ,z i ]∈R 3 } i=1...n The point on the line, with parameter μ = 0.002, satisfies the above relationship. i For reference point p c .

6. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 5, characterized in that, The spatial geometric relationship features of the point cloud are extracted as follows: F eu (p q )=[β,γ,σ]∈R 3 p q ∈V1 β=arccos(v n t ) σ=arctan(w·n t ,u·n t ) u=n q w = u × v In the formula, p q For the point cloud containing the geometric information to be queried, n q p is the unit vector of the direction of the fitting method for the query point. c Let n be the intersection of the tunnel section where the query point is located and the fitted axis of the tunnel. t For p c A unit vector pointing in the direction of the tunnel's clearance height.

7. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 1, characterized in that, In step S3, the spatial geometric features of the point cloud are fused with the deep semantic features extracted by the voxel self-attention feature perception extraction module, and represented as follows: F df =fc(ο(F,F sp )) F sp =m(F eu ) In the formula, m represents the mlp layer, F eu F represents the spatial geometric relationship features extracted from the original point cloud, F represents the point cloud features extracted by the voxel self-attention perception layer, ο represents the concatenation of features according to the point cloud index, and fc is a fully connected layer.

8. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 1, characterized in that, In step S4, the enhanced point cloud semantic features are input into the decoding layer of the network model. The decoding layer completes feature propagation and point cloud spatial distribution recovery in the following order: Based on the distance difference algorithm, the point cloud information input to the decoding layer is upsampled to enhance the point cloud semantic features, so that the spatial distribution of the point cloud contained in the decoding layer is consistent with the point cloud contained in the shallow features output by the k-th coding layer. The upsampled enhanced semantic features are concatenated with the shallow semantic features, and the features are fused and the feature channel dimensions are adjusted through a self-attention extraction layer and a linearization layer. The output features are input into the next layer of the decoding module, upsampled, and then fused with the features of the (k-1)th coding layer. By analogy, the spatial distribution of the sparse point cloud containing enhanced features is restored to that of the original point cloud.

9. The method for semantic feature fusion and enhancement of laser point clouds in subway tunnels according to claim 1, characterized in that, The loss function for building a deep learning network model and performing backpropagation training is established according to the following steps: After the decoding layer restores the point cloud to its original spatial distribution, it outputs the predicted distribution of the point cloud's category labels and calculates the loss function value of the point cloud prediction results, expressed as: In the formula, p represents the true value of the label distribution corresponding to the point cloud, q represents the predicted value of the point cloud label, and L is the number of labels for the point cloud facility type; During training, the loss function value of the network model is calculated based on the number of sample point clouds input to the network model for training in one go. In the formula, N is the number of sample point clouds used to train the input network model.

Citation Information

Patent Citations

  • Deep learning-based point cloud three-dimensional object detection method

    CN113095172A

  • Point cloud semantic segmentation method and device, electronic equipment and storage medium

    CN113516663A