Power line point cloud instance segmentation and power line reconstruction method based on SPA-Group
Through the deep learning model based on SPA-Group, multi-resolution features of power line point clouds are extracted and instance segmented, which solves the problem that traditional methods are difficult to deal with power line point cloud data, realizes efficient identification and reconstruction of power lines, and improves power inspection efficiency.
Patent Information
- Application Number
- CN202510309719.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional clustering algorithms are difficult to deal with long lines, vegetation and obstacles in power line point cloud data, resulting in uneven point cloud density and partial loss, which in turn affects the refined extraction and modeling of power lines.
The deep learning model based on SPA-Group is adopted, including a multi-resolution feature extraction module, a sparse attention feature screening module, a center of mass distance learning module and a center of mass distance constraint module. Point cloud features are extracted through sparse convolution and attention module, and instance segmentation and power line reconstruction are carried out in combination with label constraint conditions.
It has achieved efficient identification and reconstruction of single power lines, solved the problem of refined extraction and modeling of power lines in long-line power corridors, and accelerated the development of power inspection and operation and maintenance actions.
Smart Images

Figure CN120147647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of point cloud processing and instance segmentation, and more specifically, to a method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group. Background Art
[0002] Power line inspection is a systematic work with extremely high requirements for safety, accuracy, and timeliness. Traditional manual inspection is subject to challenges such as complex terrain, wide distribution of equipment, and high risks of working at heights, suffering from problems such as low efficiency, difficult blind area coverage, and single-dimensional data collection. UAV intelligent inspection, by carrying modules such as high-definition cameras, infrared thermal imagers, and lidar, combined with AI algorithms, can achieve power transmission line inspection. The daily inspection volume of a single UAV can reach 8-10 times that of manual inspection, significantly reducing safety hazards such as personnel falling from heights and electromagnetic radiation, and truly realizing the modern inspection revolution of "no power outage, no tower climbing, and no blind area".
[0003] Using UAV lidar to collect power corridor point cloud data and perform refined modeling on a single power line, and using a deep learning instance segmentation network to instantiate it to obtain instance labels of single power transmission lines, ground wires, and towers. The single instance in the power scenario spans dozens to hundreds of meters, and it is difficult to make the point cloud density of each power line completely uniform, and there are also some cases of missing point clouds. Then, traditional clustering algorithms are difficult to complete this task, and directly using instance segmentation networks such as Softgroup and Pointgroup performs poorly on this large-scale point cloud data. At the same time, during the training process, a complete instance needs to be sent into the model for training in the same batch; in the power scenario, some overhead transmission lines have large spans, and there are many surrounding vegetation and obstacles, resulting in a large number of point clouds. If the network is too complex, it requires a large amount of computing power. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention proposes a method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group, including:
[0005] A method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group includes the following steps:
[0006] S1, collection and preprocessing of deep learning training data, scanning the point cloud spatial information and RGB color information of various objects in the power corridor, and using point cloud processing software to label the original point cloud with semantic and instance label formats; downsampling the original point cloud grid and extracting initial point cloud features;
[0007] S2. Establish a deep learning model, including a multi-resolution feature extraction module, a sparse attention feature screening module, a centroid distance learning module, and a centroid distance constraint module. Extract the spatial and channel information of the point cloud through the multi-resolution feature extraction module and the sparse attention feature screening module, and give the features to the centroid distance learning module and the centroid distance constraint module for bias constraint and prediction. Add the bias of each point to the original point cloud to obtain the final spatial position information of the clustered points. The network model composed of the multi-resolution feature extraction module, the sparse attention feature screening module, the centroid distance learning module, and the centroid distance constraint module extracts the spatial position information of the point cloud data, combines the label constraint conditions, and uses the backpropagation algorithm to iteratively optimize the parameters of each module of the network. During the training process, the network selects the optimal model by comparing the verification results and finally outputs the predicted point cloud instance segmentation result.
[0008] S3. Use the power line reconstruction method to reconstruct a single power line.
[0009] Further, in step S1, the downsampling of the original point cloud grid and the extraction of the initial point cloud features are specifically as follows:
[0010] Input point cloud:
[0011]
[0012] c i =(x i ,y i ,z i ,R i ,G i ,B i )
[0013] where c i is a single point cloud containing the spatial position and color information of the point cloud; x i ,y i ,z i ,R i ,G i ,B i represent the coordinate information of the point cloud, including the spatial information x, y, z and the color information red, green, blue;
[0014] Downsample the point cloud through the grid downsampling algorithm, regularize the spatial information of the point cloud, then voxelize the obtained point cloud features, and fuse the features within the voxel through two aggregation methods of taking the mean and summing to input the point cloud:
[0015] V f =V sum +V mean
[0016]
[0017] Among them, V sum and V mean represent the ways of voxel feature fusion, including summation and mean; V f represents the more abundant voxel features obtained by adding the two ways; M represents the number of voxels, and v j represents the feature of the j-th voxel, and the feature dimension is the current channel C dimension.
[0018] Furthermore, in step S2, the multi-resolution feature extraction module is specifically:
[0019] The multi-resolution feature extraction module realizes the extraction and screening of multi-resolution features of the point cloud. Using the v j obtained after point cloud preprocessing is written as the input feature X=(x 1 , x 2 ,..., x n ) in the network, where each x i is a feature vector of the sparse point cloud after voxel feature aggregation, and the input spatial shape is where N is the batch size, C is the number of channels, and H, W, D are the spatial dimensions;
[0020] Perform multi-resolution feature fusion through the strategy of upsampling and restoring in the following sampling insertion stage:
[0021] X down =Conv3d(X, θ down , stride = 2)
[0022] X up =SparseInverseConv3d(X down , θ up , kernel_size = 2)
[0023] X combined =X down +X up (X down )
[0024] Among them, θ down is the downsampling convolution kernel, Conv3d() is the three-dimensional convolution kernel, θ down is the feature after downsampling of the upper layer feature, X is the feature input to the current layer, stride is the step size, SparseInverseConv3d() is the three-dimensional transposed convolution, kernel_size is the convolution kernel size, θ up is the upsampling convolution kernel, X up is the upsampling module, X down is the downsampling module, Xcombined is the feature fusion layer;
[0025] The obtained features of different resolutions are downsampled with a larger stride and upsampled using transposed convolution. While preserving the sparse features after downsampling, the missing points are filled to merge the multi-resolution features. The specific dimensional changes of the features in this module are [32, 64, 96, 128, 160, 192, 224].
[0026] Furthermore, in step S2, the sparse attention feature screening module is specifically as follows:
[0027] The sparse attention feature screening module performs attention screening on the deep point cloud features extracted by the network to obtain relatively independent and strongly expressive features;
[0028] The sparse attention feature screening module includes a feature sequence filter and an attention calculation part; the feature X obtained by the multi-resolution feature extraction module is sent to the fully connected layer A, σ() is the normalization function, and then X is obtained through attention weighting attn , using top k to select k features with the most prominent performance in X attn , and the overall calculation method is as follows:
[0029] A = σ(WX + b)
[0030] X attn = X ⊙ A
[0031] S = top k (X attn ) = topk(X attn , k)
[0032] X out = W fu (S ⊙ X) + X
[0033] Among them, S is the selected feature tensor, ⊙ represents element-wise multiplication, is the weight matrix of the last fully connected layer, responsible for updating the selected features, and finally outputs a more expressive feature X out .
[0034] Furthermore, in step S2, the centroid distance learning module is specifically as follows:
[0035] The centroid distance learning module learns the relative spatial distance between instance centroids, making the model more accurate when predicting offsets and keeping instances away from each other in spatial positions;
[0036] The features output by the sparse attention feature screening module After compression, an adaptive weight is generated using a fully connected layer to assign learnable weights to the distances between each pair of points. Finally, the features are restored to the input dimension n×h, and the input features are transformed through two different fully connected layers into and Taking this feature as the pre-extracted feature of attention, the relative distance weights are assigned in the following way. The original feature h i,t and the attention weight α i,t are dot-producted to obtain the weighted feature This feature is directly used for bias prediction. The specific feature generation is as follows:
[0037]
[0038] where w is the weight matrix, softmax() is the normalization function, and exp() is the exponential function, which is used to amplify the scores.
[0039] Furthermore, in step S2, the centroid distance constraint module is specifically:
[0040] The centroid distance constraint module plays a positive guiding role in the final result of the model by fusing the distance constraint from the instance to the center, the direction constraint, the center distance constraint, the semantic constraint, and the prediction score constraint;
[0041] The loss obtained by calculating the relative spatial position distance between the center of a single instance and the centers of the remaining instances is used to implement the centroid distance constraint for different instances. The specific implementation method is as follows:
[0042] D ij = ||C i - C j || 2
[0043] d min = min(D nonzero )
[0044]
[0045] where C i represents the centroid position information of the current instance, C j represents the centroid position information of the remaining instances, D ij is the Euclidean distance between the center point of each instance and the center points of the remaining instances, D nonzero represents the non-zero distance elements in D ij ; the instance with the smallest distance is taken from D ij as the reference distance d min , which is used to calculate the center distance loss center_loss.
[0046] Furthermore, step S3 is specifically:
[0047] Using the instance segmentation results to reconstruct a single power line. First, rotate and project the power line point cloud onto the YZ plane, select a minimum subset from the point cloud dataset using RANSAC, and then fit the mathematical model of the power line by combining the least squares method.
[0048] Furthermore, in step S1, scanning the point cloud spatial information and RGB color information of various objects in the power corridor specifically means: using an unmanned aerial vehicle equipped with a lidar to scan the point cloud spatial information and RGB color information of various objects in the power corridor.
[0049] The beneficial effects of the present invention are as follows:
[0050] 1. The deep learning model created by the present invention adopts the instance segmentation method, which can efficiently identify a single power line, solves the problem of fine extraction and modeling of power lines in a long-line power corridor, and accelerates the implementation of power inspection, operation and maintenance actions;
[0051] 2. The deep learning model can perform multi-resolution fusion on the features of sparse point clouds and conduct multi-feature fusion learning on the complex spatial features of power line point clouds;
[0052] 3. The sparse attention module can not only retain high expressiveness but also discard other irrelevant interference information at the same time;
[0053] 4. The centroid distance constraint module can enhance the bias cohesion in long-volume instances in the power scenario, aggregating the same instances and dispersing different instances;
[0054] 5. The power line segmentation and reconstruction method of the present invention can be efficiently applied to different environments of different voltage levels. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 It is a flowchart of the steps of the power line point cloud instance segmentation and power line reconstruction method of the present invention;
[0056] Figure 2 It is a network diagram of the instance segmentation of the present invention;
[0057] Figure 3 It is a multi-resolution feature fusion diagram of the present invention;
[0058] Figure 4 It is a schematic diagram of the sparse attention feature screening module of the present invention.
[0059] Figure 5 It is a schematic diagram of the offset output after centroid distance constraint of the present invention.
[0060] Figure 6 It is a schematic diagram of the method of the present invention for power line reconstruction.
[0061] Figure 7 This is a schematic diagram of the power line reconstruction result map of the present invention based on the data of the 110 kV voltage level.
[0062] Figure 8 This is a schematic diagram of the power line reconstruction result map of the present invention based on the data of the 220 kV voltage level.
[0063] Figure 9 This is a schematic diagram of the power line reconstruction result map of the present invention based on the data of the 500 kV voltage level. Detailed implementation manners
[0064] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0065] Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0066] SPA-Group is an instance segmentation network composed of sparse convolution and sparse attention, mainly including a SPA feature extraction network composed of sparse convolution and attention modules, and then using the features output by the network to perform instance segmentation Group. As Figures 1-4 shown, this embodiment provides a method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group. As Figure 1 shown, it includes the following steps:
[0067] S1: Collection and preprocessing of deep learning training data.
[0068] Use a drone to carry a lidar to scan the point cloud spatial information and RGB color information of various objects in the power corridor, and use point cloud processing software such as cloudcompare to label the original point cloud with semantic and instance label formats; downsample the original point cloud grid, and extract initial point cloud features through various fusion methods;
[0069] As a specific implementation manner of this embodiment, for a 220 kV high-voltage line in a certain urban area, the total length of the line collected is about 1.7 km, there are 14 power poles, and after grid downsampling, nearly 20 million points are retained. The grid downsampling rule is to retain one point for every 0.08 m long square grid;
[0070] Input point cloud:
[0071]
[0072] c i =(x i ,y i ,z i ,R i ,G i ,B i )
[0073] where c i is a single point cloud containing the spatial position and color information of the point cloud; x i ,y i ,z i ,R i ,G i ,B i represent the coordinate information of the point cloud, including the spatial information x, y, z and the color information red, green, blue;
[0074] Downsample the point cloud through the grid downsampling algorithm, regularize the spatial information of the point cloud, then voxelize the obtained point cloud features, and fuse the features within the voxelization into the input point cloud through two aggregation methods of taking the mean and summing:
[0075] V f =V sum +V mean
[0076]
[0077] where V sum and V mean represent the ways of voxel feature fusion, including summing and mean; V f represents the richer voxel features obtained by adding the two ways; M represents the number of voxels, and v j represents the feature of the j-th voxel, and the feature dimension is the current channel C dimension.
[0078] S2: Use a deep learning model for instance segmentation to obtain a single power line instance.
[0079] Build a deep learning model, and the overall network structure diagram is as shown in Figure 2 , which mainly includes a multi-resolution feature extraction module (as shown in Figure 3 ), a sparse attention feature screening module (as shown in Figure 4 ), a centroid distance learning module and a centroid distance constraint module;
[0080] The multi-resolution feature extraction module realizes the extraction and screening of multi-resolution features of the point cloud, and uses the v j obtained after point cloud preprocessing as the input feature X=(x1 , x 2 ,..., x n ), where each x i is a feature vector of the sparse point cloud after voxel feature aggregation, and the input spatial shape is where N is the batch size, C is the number of channels, and H, W, D are the spatial dimensions;
[0081] Perform multi-resolution feature fusion through the upsampling reduction strategy in the following sampling insertion stage:
[0082] X down = Conv3d(X, θ down , stride = 2)
[0083] X up = SparseInverseConv3d(X down , θ up , kernel_size = 2)
[0084] X combined = X down + X up (X down )
[0085] where θ down is the downsampling convolution kernel, Conv3d() is the three-dimensional convolution kernel, θ down is the feature after downsampling the upper layer features, X is the feature input to the current layer, stride is the step size, SparseInverseConv3d() is the three-dimensional transposed convolution, kernel_size is the convolution kernel size, θ up is the upsampling convolution kernel, X up is the upsampling module, X down is the downsampling module, and X combined is the feature fusion layer;
[0086] Downsample the obtained features of different resolutions with a larger stride and perform upsampling using transposed convolution. While retaining the sparse features after downsampling, fill in the missing points to merge the multi-resolution features. The specific dimensional changes of the features in this module are [32, 64, 96, 128, 160, 192, 224];
[0087] The sparse attention feature screening module performs attention screening on the deep point cloud features extracted by the network to obtain relatively independent and strongly expressive features;
[0088] The sparse attention feature screening module includes a feature sequence filter and an attention calculation part; the feature X obtained by the multi-resolution feature extraction module is sent to the fully connected layer A, σ() is the normalization function, and then X is obtained through attention weighting attn , and the top k is used to select attn the k most prominent features in X. The overall calculation method is as follows:
[0089] A = σ(WX + b)
[0090] X attn = X ⊙ A
[0091] S = top k (X attn ) = topk(X attn , k)
[0092] X out = W fu (S ⊙ X) + X
[0093] Among them, S is the selected feature tensor, ⊙ represents element-wise multiplication, is the weight matrix of the last fully connected layer, which is responsible for updating the selected features and finally outputting more expressive features X out ;
[0094] The centroid distance learning module learns the relative spatial distance between instance centroids, making the model more accurate when predicting biases and keeping instances away from each other in spatial positions;
[0095] After compressing the features output by the sparse attention feature screening module, an adaptive weight is generated using a fully connected layer to assign learnable weights to the distance between each point and point. Finally, the features are restored to the input dimension n×h, and the input features are transformed through two different fully connected layers into and and This feature is used as the pre-extracted feature of attention, and the relative distance weights are assigned in the following way. The original feature h i,t is dot-producted with the attention weight α i,t to obtain the weighted feature This feature is directly used for bias prediction. The specific feature generation is as follows:
[0096]
[0097] Among them, w is the weight matrix, softmax() is the normalization function, and exp() is the exponential function, which is used to amplify the score;
[0098] The centroid distance constraint module plays a positive guiding role in the final result of the model by fusing the distance constraint from the instance to the center, the direction constraint, the center distance constraint, the semantic constraint, and the prediction score constraint;
[0099] The loss obtained by calculating the relative spatial position distance between the center of a single instance and the centers of the remaining instances is used to implement the centroid distance constraint of different instances. The specific implementation method is as follows:
[0100] D ij = ||C i - C j || 2
[0101] d min = min(D nonzero )
[0102]
[0103] Among them, C i represents the centroid position information of the current instance, C j represents the centroid position information of the remaining instances, D ij is the Euclidean distance between the center points of each instance and the center points of the remaining instances, D nonzero represents the non-zero distance elements in D ij ; The instance with the smallest distance is taken from D ij as the reference distance d min , which is used to calculate the center distance loss center_loss;
[0104] The bias prediction after adding constraints to the original point cloud is as Figure 5 shown;
[0105] The final result of instance segmentation obtained by fusing the multi-resolution feature extraction module, the sparse attention feature screening module, the centroid distance learning module, and the centroid distance constraint module reaches 100% in the AP50, AP25, and AP and other metrics of multi-class instance metrics (AP50 means that when the intersection over union IoU of the predicted box and the ground truth box ≥ 0.5, the prediction is considered correct, which means that the model only needs 50% overlap between the predicted box and the ground truth box to be considered a successful detection, and AP25 is the same, while AP means calculating the intersection over union at multiple intersection ratio thresholds and then taking the average).
[0106] S3: Use an efficient and high-precision power line reconstruction method to reconstruct a single power line.
[0107] After the instance segmentation algorithm, the initial power line point cloud already has the semantic and instance labels of the initial point cloud. Use the predicted instance labels to perform power line mathematical modeling and perform instance segmentation for a single span of power line;
[0108] Using the instance segmentation results for single - file power line reconstruction, first rotate and project the power line point cloud onto the YZ plane, select a minimum subset from the point cloud dataset using RANSAC, and then combine the least - squares method to fit the mathematical model of the power line, as Figure 6 shown;
[0109] The comparison of the fitting results between the least - squares method alone and the RANSAC - combined least - squares algorithm is shown in the following figure:
[0110]
[0111] As can be seen from the above table, after removing noise by the RANSAC algorithm, the root - mean - square error (RSMSE) of the fitting algorithm is significantly reduced. While the error is reduced by at least 4 times, the time cost is 2 to 3 times. Therefore, this method can be used when high - precision detection is required.
[0112] As a specific implementation manner of this embodiment, as Figures 7-9 shown, the generalization results of the reconstruction algorithm are tested in the 110 kV, 220 kV, and 500 kV power corridor data. Through the instance segmentation network SPA - Group and an efficient power line reconstruction method, high - precision power line extraction and reconstruction are finally achieved, providing a certain reference for airborne laser point cloud power inspection and helping to accelerate the development of power inspection, operation, and maintenance operations.
[0113] The above - mentioned are only the specific implementation manners of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claimed rights.
Claims
1. A method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group, characterized in that: The following steps are involved: S1, collection and preprocessing of deep learning training data, scanning the point cloud spatial information and RGB color information of various objects in the power corridor, and using point cloud processing software to mark the original point cloud with semantic and instance label formats; downsampling the original point cloud grid and extracting the initial point cloud features; S2, establish a deep learning model, including a multi-resolution feature extraction module, a sparse attention feature screening module, a centroid distance learning module and a centroid distance constraint module; extract the spatial and channel information of the point cloud through the multi-resolution feature extraction module and the sparse attention feature screening module, and give the features to the centroid distance learning module and the centroid distance constraint module for bias constraint and prediction, and add the bias of each point to the original point cloud to obtain the final spatial position information of the cluster point; The network model composed of multi-resolution feature extraction module, sparse attention feature screening module, centroid distance learning module and centroid distance constraint module extracts the spatial position information of point cloud data, combines label constraints, and uses the back propagation algorithm to iteratively optimize the parameters of each network module; during the training process, the network selects the optimal model by comparing the verification results, and finally outputs the predicted point cloud instance segmentation results; S3, reconstructing a single power line using a power line reconstruction method.
2. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 1, characterized in that: In step S1, the original point cloud grid is downsampled and the initial point cloud features are extracted, specifically: Input point cloud: c i =(x i ,y i ,z i ,R i ,G i ,B i ) Among them, c i is a single point cloud containing the spatial position and color information of the point cloud; x i ,y i ,z i ,R i ,G i ,B i Represents the coordinate information of the point cloud, including spatial information x, y, z and color information red, green, blue; The point cloud is downsampled through the grid downsampling algorithm to regularize the spatial information of the point cloud, and then the obtained point cloud features are voxelized. The features in the voxelization are fused with the input point cloud by two aggregation methods: averaging and summing: V f =V sum +V mean Among them, V sum and V mean Indicates the way of voxel feature fusion, including summation and mean; V f represents the richer voxel features obtained by adding the two methods; M represents the number of voxels, v j Represents the feature of the j-th voxel, and the feature dimension is the C dimension of the current channel.
3. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 2, characterized in that: In step S2, the multi-resolution feature extraction module is specifically: The multi-resolution feature extraction module realizes the multi-resolution feature extraction and screening of point clouds, using the v j Written as the input feature X in the network = (x1, x2, ..., x n ), where each x i is a feature vector of the sparse point cloud after voxel feature aggregation, and the input spatial shape is Where N is the batch size, C is the number of channels, and H, W, D are the spatial dimensions; Multi-resolution feature fusion is performed by inserting the upsampling and restoration strategy in the following sampling stage: X down =Conv3d(X,θ down ,stride=2) X up =SparseInverseConv3d(X down ,θ up ,kernel_size=2) X combined =X down +X up (X down ) Among them, θ down is the downsampling convolution kernel, Conv3d() is the three-dimensional convolution kernel, θ down is the feature after downsampling of the upper layer feature, X is the feature of the current layer input, stride is the step size, SparseInverseConv3d() is the three-dimensional deconvolution, kernel_size is the convolution kernel size, θ up is the upsampling convolution kernel, X up is the upsampling module, X down is the downsampling module, X combined is the feature fusion layer; The obtained features of different resolutions are downsampled with a larger stride and upsampled using deconvolution. While retaining the sparse features after downsampling, the missing points are filled to merge multi-resolution features. In this module, the specific dimension changes of the features are [32, 64, 96, 128, 160, 192, 224].
4. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 3, characterized in that: In step S2, the sparse attention feature screening module is specifically: The sparse attention feature screening module performs attention screening on the deep point cloud features extracted by the network to obtain relatively independent and highly expressive features; The sparse attention feature filtering module includes a feature sequence filter and an attention calculation part. The feature X obtained by the multi-resolution feature extraction module is sent to the fully connected layer A, σ() is a normalization function, and then X is obtained by attention weighting. attn , using top k Select X attn The k most prominent features are calculated as follows: A=σ(WX+b) X attn =X⊙A S=top k (X attn )=topk(X attn ,k) X out =W fu (S⊙X)+X Among them, S is the selected feature tensor, ⊙ represents element-by-element multiplication, It is the weight matrix of the last fully connected layer, which is responsible for updating the selected features and finally outputting more expressive features X out .
5. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 4, characterized in that: In step S2, the centroid distance learning module is specifically: The centroid distance learning module learns the relative spatial distance between instance centroids, making the model more accurate in predicting bias and moving instances farther away from each other in space. The features output by the sparse attention feature filtering module After compression, the fully connected layer is used to generate adaptive weights, and the distance between each point is given a learnable weight. Finally, the features are restored to the input dimension n×h, and the input features are transformed into and This feature is used as the pre-extracted feature of attention, and the relative distance weight is assigned in the following way: the original feature h i,t With the attention weight α i,t Dot product to get weighted features This feature is directly used for bias prediction, and the specific features are generated as follows: Among them, w is the weight matrix, softmax() is the normalization function, and exp() is the exponential function used to amplify the score.
6. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 5, characterized in that: In step S2, the centroid distance constraint module is specifically: The centroid distance constraint module provides positive guidance for the final result of the model by integrating the distance constraint from the instance to the center, the direction constraint, the center distance constraint, the semantic constraint, and the prediction score constraint. The loss obtained by calculating the relative spatial position distance between the center of a single instance and the centers of other instances is used to implement the centroid distance constraint of different instances. The specific implementation method is as follows: D ij =||C i -C j ||2 the min =min(D nonzero ) Among them, C i Represents the centroid position information of the current instance, C j Represents the centroid position information of the remaining instances, D ij is the Euclidean distance between the center point of each instance and the center points of other instances, D nonzero Indicates D ij Non-zero distance elements in D ij Take the instance with the smallest distance as the reference distance d min , used to calculate the center distance loss center_loss.
7. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 1, characterized in that: Step S3 is specifically as follows: The single-speed power line reconstruction is performed using the instance segmentation results. The power line point cloud is first rotated and projected to the YZ plane. A minimum subset is selected from the point cloud dataset using RANSAC, and then the mathematical model of the power line is fitted using the least squares method.
8. The method for power line point cloud instance segmentation and power line reconstruction based on SPA-Group according to claim 1, characterized in that: In step S1, the point cloud spatial information and RGB color information of various objects in the power corridor are scanned, specifically: Use a drone equipped with a lidar to scan the point cloud spatial information and RGB color information of various objects in the power corridor.
Citation Information
Cited By
Multi-modal detection method based on local density perception and dynamic sparse attention
CN121617061A