Semantic segmentation method and system for laser radar point cloud of high-voltage transmission line

By adopting the network structure of the kernel point convolution layer, the graph edge convolution layer and the channel and spatial attention module in the point cloud data processing of high-voltage transmission line, the problem of difficulty in encoding the global context information of large-scale point cloud data in the existing technology is solved, and a high-precision semantic segmentation effect is achieved.

CN120107602AActive Publication Date: 2025-06-06STATE GRID ECONOMIC TECH RES INST CO LTD +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510578265.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The prior art is difficult to effectively encode the global context information of large-scale point cloud data, and it is difficult to meet the semantic segmentation needs of high-voltage transmission line scenarios, especially when processing scales are large, data volumes are huge, and category differences are large.

Method used

Using a network structure based on the kernel point convolution layer, graph edge convolution layer and channel and spatial attention module, an encoder and decoder are built, through these modules, local and global features of point cloud data are fused, and attention mechanism is used for context understanding during the encoding and decoding process.

Benefits of technology

The semantic segmentation of point cloud data of high-voltage transmission line is realized, the accuracy and understanding of semantic segmentation are improved, and the global context information of large-scale point cloud data can be more effectively processed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107602A_ABST
    Figure CN120107602A_ABST
Patent Text Reader

Abstract

The invention provides a semantic segmentation method and system for a laser radar point cloud of a high-voltage transmission line. According to the implementation scheme, based on a high-voltage transmission line point cloud data training set, a first network structure constructed by a kernel point convolution layer, a graph edge convolution layer and a channel and space attention module is trained, and a target semantic segmentation model is obtained; and in response to the semantic segmentation request, based on the target semantic segmentation model, performing semantic segmentation on the high-voltage transmission line point cloud data corresponding to the semantic segmentation request. Wherein the first network structure comprises an encoder and a decoder; the encoder and the decoder comprise at least one first encoding layer and a first decoding layer which are formed by connecting a kernel point convolution layer and a graph edge convolution layer in parallel, and at least one second encoding layer and a second decoding layer which are formed by connecting the kernel point convolution layer and the graph edge convolution layer in parallel and then connecting the kernel point convolution layer and the graph edge convolution layer in series with a channel and a space attention module. By adopting the method, the semantic segmentation accuracy of the point cloud data of the high-voltage transmission line can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a semantic segmentation method and system for a laser radar point cloud of a high-voltage transmission line. Background Art

[0002] High-voltage transmission line inspection is one of the daily tasks of the power system. It is of great significance to the safety and stability of power supply. The main purpose of high-voltage transmission line inspection is to check whether the power infrastructure such as power lines, power towers, hardware, etc. is intact, and whether there are potential tree obstacles, geological disasters, tower tilt and other dangerous hidden dangers.

[0003] In recent years, with the popularization of observation technology, the use of remote sensing technology, especially LiDAR technology, to observe transmission channels has become a mainstream inspection trend. LiDAR uses an active observation method, can be used on multiple platforms, and has the advantages of high precision, high speed, and less affected by weather. Therefore, in more and more power industry work, manned or unmanned aircraft equipped with LiDAR systems are used as one of the important collection methods in high-voltage transmission line inspections. At the same time, the extraction of high-voltage transmission line scene information based on LiDAR three-dimensional point cloud data has become an important part of the inspection.

[0004] In some technologies, for semantic segmentation models in transmission channel scenarios, for example, the PointNet model enhanced by combining the geometric feature extraction module and the neighborhood information aggregation module can achieve the segmentation of railway power lines and towers from railway scene point clouds. For another example, the PointNet++ model is used to segment power lines and towers with high precision, and then the coordinate attention module (CA) is integrated with PointNet++ to achieve an end-to-end CA-PointNet++ model. However, there are still some challenges to be solved for the semantic segmentation of airborne lidar in transmission channel scenarios: 1. Airborne lidar has a larger geographical range. The above methods mainly focus on how to express the local neighborhood information of points, and it is difficult to encode the global context information of large-scale point cloud scenes.

[0005] 2. The transmission channel scene is a long strip with a certain length as a buffer zone, which continues to move along the direction of the transmission line. The transmission channel has the characteristics of large scale, large amount of point cloud data (the number of points between two towers can reach more than 10 million), and large category differences (for example, important factors such as power towers and power lines account for a very small proportion). Therefore, conventional point cloud semantic segmentation processing methods are difficult to meet the task of understanding the transmission channel scene. Summary of the invention

[0006] The present invention provides a semantic segmentation method and system for a lidar point cloud of a high-voltage transmission line, which can solve at least one of the above technical problems.

[0007] According to one aspect of the present invention, a semantic segmentation method for a high-voltage transmission line laser radar point cloud is provided, comprising: The first network structure is constructed based on the kernel point convolution layer, the graph edge convolution layer, and the channel and spatial attention modules; Based on the high-voltage transmission line point cloud data training set, the first network structure is trained to obtain a target semantic segmentation model; In response to the semantic segmentation request, based on the target semantic segmentation model, semantically segment the high-voltage transmission line point cloud data corresponding to the semantic segmentation request; Wherein, the first network structure includes an encoder and a decoder; The encoder includes at least one first encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module; The decoder includes at least one first decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module.

[0008] According to another aspect of the present invention, a semantic segmentation device for a laser radar point cloud of a high-voltage transmission line is provided, the device comprising: A network structure determination module, used to construct a first network structure based on a kernel point convolution layer, a graph edge convolution layer, and a channel and spatial attention module; A model training module, used to train the first network structure based on a training set of high-voltage transmission line point cloud data to obtain a target semantic segmentation model; A semantic segmentation module, configured to respond to a semantic segmentation request and perform semantic segmentation on the high-voltage transmission line point cloud data corresponding to the semantic segmentation request based on the target semantic segmentation model; Wherein, the first network structure includes an encoder and a decoder; The encoder includes at least one first encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module; The decoder includes at least one first decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module.

[0009] The technical solution of the present invention is adopted to construct a first network structure based on a kernel point convolution layer, a graph edge convolution layer, and a channel and space attention module; the first network structure is trained based on a training set of high-voltage transmission line point cloud data to obtain a target semantic segmentation model. In addition, the first network structure includes an encoder and a decoder, the encoder includes at least one first encoding layer composed of a kernel point convolution layer and a graph edge convolution layer in parallel, and at least one second encoding layer composed of a kernel point convolution layer and a graph edge convolution layer in parallel and then connected in series with a channel and space attention module, the decoder includes at least one first decoding layer composed of a kernel point convolution layer and a graph edge convolution layer in parallel, and at least one second decoding layer composed of a kernel point convolution layer and a graph edge convolution layer in parallel and then connected in series with a channel and space attention module. In this way, the target semantic segmentation model can use the kernel point convolution layer and the graph edge convolution layer to perform local and global feature fusion on the point cloud data respectively during encoding and decoding, and can use the channel and spatial attention modules to understand the context of the fused features during encoding and decoding. Finally, the fused features after context understanding obtained by decoding are used to realize the semantic segmentation of the point cloud data of the high-voltage transmission line, thereby improving the accuracy of semantic segmentation.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention. Figure 1 is a flow chart of a semantic segmentation method of a high-voltage transmission line laser radar point cloud according to an embodiment of the present invention; Figure 2 is a network structure diagram of a graph edge convolution layer according to an embodiment of the present invention; Figure 3 is a network structure diagram of a channel and spatial attention module according to an embodiment of the present invention; Figure 4 is a schematic diagram of the structure of the network layer of an embodiment of the present invention; Figure 5 is a schematic diagram of a first network structure according to an embodiment of the present invention; Figure 6 It is a structural block diagram of a semantic segmentation device for a high-voltage transmission line laser radar point cloud according to an embodiment of the present invention; Figure 7 The block diagram is a block diagram of an electronic device for implementing the method according to the embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0013] Figure 1 It is a flow chart of a semantic segmentation method of a high-voltage transmission line laser radar point cloud according to an embodiment of the present invention.

[0014] like Figure 1 As shown, the semantic segmentation method of the high-voltage transmission line laser radar point cloud may include: S110, constructing a first network structure based on a kernel point convolution layer, a graph edge convolution layer, and a channel and spatial attention module; S120, training the first network structure based on the high-voltage transmission line point cloud data training set to obtain a target semantic segmentation model; S130, in response to the semantic segmentation request, performing semantic segmentation on the high-voltage transmission line point cloud data corresponding to the semantic segmentation request based on the target semantic segmentation model; Wherein, the first network structure includes an encoder and a decoder; The encoder includes at least one first encoding layer consisting of a kernel point convolution layer and a graph edge convolution layer connected in parallel, and at least one second encoding layer consisting of a kernel point convolution layer and a graph edge convolution layer connected in parallel and then connected in series with a channel and a spatial attention module; The decoder includes at least one first decoding layer composed of a kernel point convolution layer and a graph edge convolution layer in parallel, and at least one second decoding layer composed of a kernel point convolution layer and a graph edge convolution layer in parallel and then connected in series with a channel and a spatial attention module.

[0015] For example, in the kernel point convolution layer, each convolution kernel is composed of a set of kernel points with fixed coordinates in three-dimensional space. The kernel point convolution layer calculates the weight according to the relative position between each point in the input point cloud data and the kernel point, thereby generating a convolution result and outputting the convolution result. The input data of the kernel point convolution layer includes point position information and point feature information. Point position information is usually expressed as The kernel point convolution layer uses a set of kernel points to represent the specific location of each point in the three-dimensional space. To define the convolution kernel. Each kernel point All have fixed coordinates in 3D space. These kernel points are similar to the filter weights in traditional convolution, but are irregularly distributed in 3D space. The positions of these kernel points are adjusted during training to optimize the extracted features.

[0016] For example, the graph edge convolution layer can construct a graph on a line segment composed of geometrically homogeneous points to capture the relationship between objects. By combining segment features and point features, the network can adaptively encode local and global features, thereby achieving better semantic learning or prediction on airborne lidar datasets. Figure 2 As shown, for edge conditioned convolution (ECC), it is considered as a directed graph or an undirected graph ,in, is a finite set of vertices, ; is a set of edges, .set up is the layer index in the feedforward neural network. Thus, we define a function to assign labels to each vertex, and a function to assign labels to each edge.

[0017] like Figure 2 As shown in FIG. 1 , the point cloud in the directed or undirected graph is segmented into multiple independent point cloud clusters using an unsupervised algorithm. This part is pre-calculated before training using semantic labels and preset geometric features to ensure that the structure of the directed or undirected graph is fixed during the training process, without repeatedly updating the labels of each edge in the graph, thus reducing unnecessary computational overhead.

[0018] Next, we use edge conditions to convolve the graph structure and dynamically generate filter weights to flexibly handle neighborhood relationships. Specifically, given a point cloud Its geometric characteristics (e.g. linearity or normal vector), construct a directed graph , and set the labels of each vertex in the directed graph as follows For point clouds Each point in , create a vertex for it and through Determine its label. Then, for each vertex All vertices in its spatial neighborhood Connected by directed edges. Edge Conditional Convolution (ECC) can dynamically generate filter weights according to directed edges and flexibly handle the number of neighbors to capture contextual information between different edges.

[0019] For example, Figure 3 As shown, for the channel and spatial attention module, it can be a module composed of a channel attention network and a spatial attention network connected in parallel and then connected to a connection layer.

[0020] For example, Figure 3 As shown in the figure, for the channel attention network, the input features are projected into different feature subspaces through different learnable fully connected layers in order to construct queries, keys, and values ​​in the attention function. The attention function has the descriptive ability to encode the global context, and the output of the attention function is the enhanced point cloud features. The channel attention mechanism focuses on selecting and enhancing the most effective features for the network. By studying the relationship between features, the weights of different features are determined, and the obtained weights are multiplied by the features to obtain features that are important for network classification. In the process of calculating the weights, the spatial dimensions of the input features can be compressed to improve the computational efficiency.

[0021] For the channel attention network, the spatial dimension information between different feature channels is aggregated through average pooling and maximum pooling to obtain the average pooling feature and the maximum pooling feature. Then, these two features are input into a multi-layer perceptron (MLP) network to generate the mapping function of the important features. The MLP network can be composed of two fully connected layers, an activation function, and a random dropout mechanism for the output results of the second fully connected layer, which can improve the generalization ability of the network. Finally, the results of the two different features processed by MLP are summed element by element to obtain the output result of the channel attention network.

[0022] like Figure 3 As shown in the figure, for the spatial attention network, it is used to select a neighborhood that is more conducive to expressing the shape information of the point cloud. For the input information, in order to pay more attention to the correlation between different point cloud categories, first, the input features are average pooled and max pooled. The input features are feature matrices that aggregate the edge conditional convolution and kernel point convolution, which is a feature polymer of point cloud and target level features. Then, the pooled features are connected in series and convolution operations are performed to generate different attention coefficients, that is, the output result of the spatial attention network is obtained.

[0023] Finally, the output results of the channel attention network are superimposed with the output results of the spatial attention network, and the superimposed results are calculated point by point through a fully connected layer to obtain the attention score of each point cloud feature.

[0024] In the above example, through the channel and spatial attention modules, the features are updated point by point cloud from a global perspective, and the interactions between complex points are fully learned, which helps to improve the accuracy of subsequent point cloud semantic segmentation.

[0025] For example, Figure 4 As shown in FIG, it shows a network layer composed of a kernel point convolution layer, a graph edge convolution layer, and a channel and space attention module. The input data of the kernel point convolution layer is the spatial coordinates of the three-dimensional point cloud. During training, the input data of the graph edge convolution layer is the annotation information of the three-dimensional point cloud. When applied, its input data is also a three-dimensional point cloud. The channel and space attention module calculates the attention score of the convolved features, further mines and extracts the global features of the point cloud, and its output result is the point cloud feature matrix.

[0026] For example, Figure 5 As shown, the first network structure can use U-Net as the overall framework structure and use the five network layers in the virtual box for encoding. The first layer in the virtual box is used to vectorize the input data to obtain a vector representation of the input data, and the other five network layers are encoding layers. Each encoding layer includes multiple network layers. The network layer in the non-virtual box is a decoding layer. Each network layer in the decoding layer can be composed of at least two of the graph edge convolution layer, the kernel point convolution layer, and the channel and spatial attention module. The network layer composed of these three can be as follows Figure 4 shown.

[0027] like Figure 4 As shown in the figure, during the training phase, the input data first passes through the kernel point convolution coding layer and the graph edge conditional convolution layer to extract the local features and target-level features of the point cloud. Then, the features of the two are connected and input into the channel and spatial attention module, which captures the global features of the input features in depth, and finally extracts the local and global features of the point cloud.

[0028] like Figure 5 As shown in the figure, in order to capture geometric information at multiple scales, downsampling is used to gradually expand the receptive field of the convolution. In the decoder, nearest neighbor upsampling is used to obtain the final point-by-point features. The model extracts local and global features in all encoding layers, and adds channel and spatial attention modules to the first, third, and fifth encoding layers to consider contextual information. Similarly, skip connections are used to pass the intermediate features of each encoding layer in the encoder to each decoding layer in the decoder. In the decoder, the intermediate features are concatenated with the upsampled features and then passed to Figure 5 The last network layer in , the fully connected layer, realizes semantic prediction.

[0029] Exemplarily, for a training data set, that is, a high-voltage transmission line point cloud data training set according to an embodiment of the present invention, it can be acquired and preprocessed in the following manner.

[0030] First, use drones or airplanes to collect lidar data of high-voltage transmission lines in the first power grid, and solve to obtain three-dimensional point cloud data. Then, the three-dimensional point cloud data is cropped and binned. For example, based on the route trajectory, the three-dimensional point cloud data is cropped along the transmission line and divided according to a fixed distance. Usually, the three-dimensional point cloud data between two towers or three towers is divided into a point cloud data set. Then, the cropped and binned point cloud is denoised. For example, outliers are removed. At the same time, grid downsampling is used to downsample the denoised point cloud data, and a KD-Tree is constructed to facilitate the organization and indexing of point cloud data.

[0031] Among them, the three-dimensional KD-Tree is abbreviated as k-dimensional tree, which is a data structure for space partitioning. Specifically, KD-Tree is a binary tree structure for organizing k-dimensional space point data. Each non-leaf node can be divided into two subspaces by a hyperplane, and each corresponding subspace can be recursively divided in the same way. All subspaces are divided into left and right parts or into upper and lower parts. The division of KD-Tree is performed along the coordinate axis, and all hyperplanes are perpendicular to the corresponding coordinate axis. KD-Tree is a relatively effective k-dimensional space point data organization structure, and has its own unique advantages in the field of high-dimensional space search (such as k-neighbor search). The three-dimensional KD-Tree used here can efficiently organize and manage the initial point cloud data with a large amount of data.

[0032] The downsampling is grid sampling. Specifically, the three-dimensional point cloud data is evenly divided into multiple small cubes, and the points in each cube are sampled. The mean of the points in each cube is used as the sampling point cloud of the cube, and the number of each point cloud category in the cube is counted, and the point cloud category with the largest number is used as the category of the sampling point cloud. In this example, using grid sampling can greatly reduce the number of sampling points, reducing data calculation and memory consumption for subsequent model training and testing.

[0033] Finally, the sampled point cloud data are annotated to obtain the label information of each point cloud data, thereby constructing a high-voltage transmission line training dataset.

[0034] In one embodiment, the encoder includes 5 encoding layers. In the direction from the input to the output of the encoder, the first encoding layer, the third encoding layer and the fifth encoding layer in the encoder all adopt the structure of the second encoding layer, and the second encoding layer and the fourth encoding layer in the encoder all adopt the architecture of the first encoding layer.

[0035] In one embodiment, the decoder includes 5 decoding layers. In the direction from the input to the output of the decoder, the first decoding layer, the third decoding layer and the fifth decoding layer in the decoder all adopt the structure of the second decoding layer, and the second decoding layer and the fourth decoding layer in the decoder all adopt the structure of the first decoding layer.

[0036] For example, Figure 5 As shown in , both the encoder and decoder are configured with 5 network layers. Figure 5 As shown, from left to right, the first, third, and fifth coding layers of the encoder all use the second coding layer, and the second and third coding layers of the encoder all use the first coding layer. The first, third, and fifth decoding layers of the decoder all use the second decoding layer, and the second and fourth decoding layers all use the first decoding layer.

[0037] It can be understood that the fifth encoding layer can be used as the first decoding layer of the decoder.

[0038] It can be understood that each encoding layer gradually downsamples the input features, and gradually converts the dimensions of the features from 64 dimensions, 128 dimensions, 256 dimensions, 512 dimensions to 1024 dimensions.

[0039] It can be understood that each decoding layer gradually upsamples the input features, and gradually converts the dimensions of the features from 1024 dimensions, 512 dimensions, 256 dimensions, 128 dimensions to 64 dimensions.

[0040] In this way, a channel and spatial attention module is set every other network layer, and each network layer is provided with a parallel kernel point convolution layer and a graph edge convolution layer.

[0041] Exemplarily, in the direction from the input to the output of the encoder, the second coding layer and the fourth coding layer of the encoder both use the second coding layer, and the first coding layer, the third coding layer, and the fifth coding layer of the encoder all use the first coding layer. In the direction from the input to the output of the decoder, the first decoding layer, the third decoding layer, and the fifth decoding layer of the decoder all use the first decoding layer, and the second and fourth decoding layers all use the second decoding layer.

[0042] According to the above implementation, the point cloud segmentation model uses a network composed of an encoder, a decoder and a fully connected layer, and each network layer is provided with a parallel kernel point convolution layer and a graph edge convolution layer to extract features of semantic orientation and geometric attributes between categories, respectively, and set channel and spatial attention modules in some network layers, which can mine contextual information in the scene and model local and global features. The network structure set in this way can improve the accuracy of semantic segmentation and semantic understanding ability.

[0043] In one embodiment, based on a training set of high-voltage transmission line point cloud data, a first network structure is trained to obtain a target semantic segmentation model, including: based on a training set of high-voltage transmission line point cloud data, the first network structure is trained to obtain a first semantic segmentation model; based on the network parameters of the first semantic segmentation model, the structure type of each coding layer and the structure type of each decoding layer in the first network structure are adjusted to obtain a second network structure, wherein the structure type of the coding layer includes a first coding layer and a second coding layer, and the structure type of the decoding layer includes a first decoding layer and a second decoding layer; based on the structural difference between the first network structure and the second network structure, the network parameters of the first semantic segmentation model are adjusted; based on the adjusted network parameters of the first semantic segmentation model, the network parameters of the second network structure are set; based on the training set of high-voltage transmission line point cloud data, the second network structure after the network parameters have been set is trained to obtain a target semantic segmentation model.

[0044] Exemplarily, the accuracy of the first network structure or the residual of each encoding layer and its corresponding decoding layer is determined using the network parameters obtained by pre-training, and the structural type of the encoding layer and / or decoding layer is adjusted using these parameters to obtain the second network structure.

[0045] Exemplarily, the particle swarm algorithm may be used to adjust the structural type of each encoding layer and the structural type of each decoding layer in the first network structure to obtain the second network structure.

[0046] Exemplarily, the training process for the first network structure and the training process for the second network structure may be the same or different.

[0047] Exemplarily, based on the network parameters of the first semantic segmentation model, the network parameters of the second network structure are set, and then based on the high-voltage transmission line point cloud data training set, the second network structure with the set network parameters is trained to obtain the target semantic segmentation model. This can reduce the number of training times or training time of the second network structure and maintain the accuracy of the target semantic segmentation model.

[0048] Alternatively, the network parameters of the first semantic segmentation model are adjusted based on the structural difference between the first network structure and the second network structure as in the above example; the network parameters of the second network structure are set based on the adjusted network parameters of the first semantic segmentation model; finally, the second network structure with set network parameters is trained based on the training set of high-voltage transmission line point cloud data to obtain the target semantic segmentation model. For example, if the structures of the two corresponding network layers in the first network structure and the second network structure are the same, the network parameters of the corresponding network layer in the first semantic segmentation model are not adjusted; if the structures of the two corresponding network layers in the first network structure and the second network structure are different, the network parameters of the corresponding network layer in the first semantic segmentation model are adjusted.

[0049] For example, during model training, an Adam optimizer may be used. For another example, cross entropy may be used as a loss function. For another example, a warm-up strategy may be used to gradually change the learning rate.

[0050] According to the above implementation, the semantic segmentation model is first preliminarily trained to obtain its network parameters, and then the structural types of each encoding layer and decoding layer in the network structure are adjusted using the preliminarily trained network parameters, so that the network structure can be quickly optimized. Finally, the model is trained using the optimized network structure, which can reduce the number of training times or training time while maintaining or improving the accuracy of the target semantic segmentation model.

[0051] In one embodiment, based on the network parameters of the first semantic segmentation model, the structural type of each coding layer and the structural type of each decoding layer in the first network structure are adjusted to obtain a second network structure, including: based on the first network structure, determining an initial population, wherein the initial population includes multiple particles, each particle corresponds to a first network structure, and the first network structures corresponding to each particle have at least one coding layer or decoding layer with different structural types; performing the following iterative operations starting from the initial population: based on the network parameters of the first semantic segmentation model, setting the network parameters of the first network structure corresponding to each particle in this iteration population, respectively, to obtain a second semantic segmentation model corresponding to each particle in this iteration population; based on a training set of high-voltage transmission line point cloud data, determining the semantic segmentation error of the second semantic segmentation model corresponding to each particle in this iteration population; based on the semantic segmentation error of the second semantic segmentation model corresponding to each particle in this iteration population, updating the individual optimal fitness value; when the individual optimal fitness value meets the preset fitness value condition, determining the second network structure based on the first network structure corresponding to the particle corresponding to the individual optimal fitness value.

[0052] In one embodiment, the above-mentioned iterative operation also includes: when the individual optimal fitness value does not meet the fitness value condition, adjusting the structure type of the encoding layer or decoding layer in the first network structure corresponding to each particle in the current iterative population to obtain the next iterative population.

[0053] Exemplarily, the network parameters of the first semantic segmentation model are used to set the network parameters of the first network structure corresponding to each particle, and the second semantic segmentation model corresponding to each particle is obtained. Then, the semantic segmentation error of the second semantic segmentation model corresponding to each particle is calculated using the high-voltage transmission line point cloud data training set.

[0054] Exemplarily, before the first iteration, the initial value of the individual optimal fitness value is pre-set. In each iteration, if the minimum prediction error among the prediction errors of the second semantic segmentation model corresponding to each particle in the iteration population is less than or equal to the individual optimal fitness value, the individual optimal fitness value is updated with the minimum prediction error; if the minimum prediction error is greater than the individual optimal fitness value, the individual optimal fitness value is kept unchanged.

[0055] Exemplarily, the fitness value condition may be that the individual optimal fitness value is less than a preset fitness value threshold.

[0056] Exemplarily, if the number of iterations exceeds a preset number threshold, the second network structure may also be determined based on the first network structure corresponding to the particle corresponding to the individual optimal fitness value. If the number of iterations does not exceed the preset number threshold, the structure type of the encoding layer or decoding layer in the first network structure corresponding to each particle in the current iteration population is adjusted to obtain the next iteration population.

[0057] According to the above implementation, based on the first network structure, the population is initialized, and then the network parameters of the first semantic segmentation model are used from the initial population to iteratively optimize the particles in the population to obtain the final population and individual optimal fitness values, and the network structure corresponding to the individual optimal fitness values ​​is used to determine the final network structure, so that the network structure can be quickly optimized.

[0058] In one embodiment, the structural type of the coding layer or the decoding layer in the first network structure corresponding to each particle in the current iterative population is adjusted to obtain the next iterative population, including: for each particle in the current iterative population, the following operations can be performed to obtain an iterative population: based on the residual between two adjacent coding layers in the first network structure corresponding to the particle and the respective corresponding decoding layers, the structural types of the two adjacent coding layers are adjusted; based on the adjusted structural types of the two adjacent coding layers, the structural types of the corresponding two decoding layers are adjusted to obtain the adjusted first network structure corresponding to the particle.

[0059] Exemplarily, if the residuals between two adjacent coding layers and their respective corresponding decoding layers are greater than a preset residual threshold, the network structure of the coding layer that is ordered later among the two adjacent coding layers is set to the above-mentioned second coding layer, and the network structure of the decoding layer corresponding to the coding layer that is ordered later is set to the above-mentioned second decoding layer. The coding layer that is ordered earlier can adopt the first coding layer or the second coding layer, and its corresponding decoding layer can adopt the first decoding layer or the second decoding layer.

[0060] According to the above implementation, the optimization method of the first network structure corresponding to each particle in the population is determined by using the residual, thereby improving the accuracy of population iteration and the speed of network structure optimization.

[0061] Figure 6 It is a structural block diagram of a semantic segmentation device for a high-voltage transmission line laser radar point cloud according to an embodiment of the present invention.

[0062] like Figure 6 As shown, the semantic segmentation device may include: A network structure determination module 610, configured to construct a first network structure based on a kernel point convolution layer, a graph edge convolution layer, and a channel and spatial attention module; A model training module 620 is used to train the first network structure based on a high-voltage transmission line point cloud data training set to obtain a target semantic segmentation model; A semantic segmentation module 630, configured to perform semantic segmentation on the high-voltage transmission line point cloud data corresponding to the semantic segmentation request based on the target semantic segmentation model in response to the semantic segmentation request; Wherein, the first network structure includes an encoder and a decoder; The encoder includes at least one first encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module; The decoder includes at least one first decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module.

[0063] In one embodiment, the encoder includes 5 coding layers. From the input to the output direction, the first coding layer, the third coding layer and the fifth coding layer in the encoder all adopt the structure of the second coding layer, and the second coding layer and the fourth coding layer in the encoder all adopt the architecture of the first coding layer.

[0064] In one embodiment, the decoder includes 5 decoding layers. From the input to the output direction, the first decoding layer, the third decoding layer and the fifth decoding layer in the decoder all adopt the structure of the second decoding layer, and the second decoding layer and the fourth decoding layer in the decoder all adopt the structure of the first decoding layer.

[0065] In one embodiment, the model training module 620 includes: A first training unit, configured to train the first network structure based on the high-voltage transmission line point cloud data training set to obtain a first semantic segmentation model; A structure adjustment unit, configured to adjust the structure type of each encoding layer and the structure type of each decoding layer in the first network structure based on the network parameters of the first semantic segmentation model, so as to obtain a second network structure, wherein the structure type of the encoding layer includes the first encoding layer and the second encoding layer, and the structure type of the decoding layer includes the first decoding layer and the second decoding layer; a network parameter adjustment unit, configured to adjust network parameters of the first semantic segmentation model based on a structural difference between the first network structure and the second network structure; A network parameter setting unit, configured to set network parameters of the second network structure based on the adjusted network parameters of the first semantic segmentation model; The second training unit is used to train the second network structure after the network parameters have been set based on the high-voltage transmission line point cloud data training set to obtain the target semantic segmentation model.

[0066] In one embodiment, the structure adjustment unit is specifically used to: Based on the first network structure, an initial population is determined, wherein the initial population includes a plurality of particles, each particle corresponds to one of the first network structures, and the first network structures corresponding to the particles have at least one encoding layer or decoding layer with a different structure type; Starting from the initial population, the following iterative operations are performed: Based on the network parameters of the first semantic segmentation model, the network parameters of the first network structure corresponding to each particle in the current iteration population are respectively set to obtain the second semantic segmentation model corresponding to each particle in the current iteration population; Determining, based on the high-voltage transmission line point cloud data training set, a semantic segmentation error of a second semantic segmentation model corresponding to each particle in the current iteration population; Based on the semantic segmentation error of the second semantic segmentation model corresponding to each particle in the current iteration population, updating the individual optimal fitness value; When the individual optimal fitness value meets the preset fitness value condition, the second network structure is determined based on the first network structure corresponding to the particle corresponding to the individual optimal fitness value.

[0067] In one embodiment, the iterative operation further includes: When the individual optimal fitness value does not meet the fitness value condition, the structure type of the encoding layer or the decoding layer in the first network structure corresponding to each particle in the current iterative population is adjusted to obtain the next iterative population.

[0068] In one implementation, the adjusting the structure type of the encoding layer or the decoding layer in the first network structure corresponding to each particle in the current iteration population includes: Adjusting the structural types of the two adjacent coding layers based on the residuals between the two adjacent coding layers and the respective corresponding decoding layers in the first network structure corresponding to the particle; Based on the adjusted structure types of the two adjacent encoding layers, the structure types of the corresponding two decoding layers are adjusted to obtain the adjusted first network structure corresponding to the particle.

[0069] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0070] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0071] According to an embodiment of the present invention, the present invention also provides a system and a readable storage medium.

[0072] Figure 7 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0073] like Figure 7As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 to a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0074] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0075] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as the semantic segmentation method of the high-voltage transmission line laser radar point cloud. For example, in some embodiments, the semantic segmentation method of the high-voltage transmission line laser radar point cloud may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the semantic segmentation method of the high-voltage transmission line laser radar point cloud described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the semantic segmentation method of the high-voltage transmission line lidar point cloud in any other appropriate manner (eg, by means of firmware).

[0076] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0077] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0078] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0079] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0080] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0081] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0082] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.

[0083] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A semantic segmentation method for high-voltage transmission line laser radar point cloud, characterized in that: include: The first network structure is constructed based on the kernel point convolution layer, the graph edge convolution layer, and the channel and spatial attention modules; Based on the high-voltage transmission line point cloud data training set, the first network structure is trained to obtain a target semantic segmentation model; In response to the semantic segmentation request, based on the target semantic segmentation model, semantically segment the high-voltage transmission line point cloud data corresponding to the semantic segmentation request; Wherein, the first network structure includes an encoder and a decoder; The encoder includes at least one first encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module; The decoder includes at least one first decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module.

2. The method according to claim 1, characterized in that The encoder includes 5 encoding layers. From the input to the output direction, the first encoding layer, the third encoding layer and the fifth encoding layer in the encoder all adopt the structure of the second encoding layer, and the second encoding layer and the fourth encoding layer in the encoder all adopt the architecture of the first encoding layer.

3. The method according to claim 1, characterized in that The decoder includes 5 decoding layers. From the input to the output direction, the first decoding layer, the third decoding layer and the fifth decoding layer in the decoder all adopt the structure of the second decoding layer, and the second decoding layer and the fourth decoding layer in the decoder all adopt the structure of the first decoding layer.

4. The method according to claim 1, characterized in that The first network structure is trained based on the high-voltage transmission line point cloud data training set to obtain a target semantic segmentation model, including: Based on the high-voltage transmission line point cloud data training set, training the first network structure to obtain a first semantic segmentation model; Based on the network parameters of the first semantic segmentation model, adjusting the structure type of each encoding layer and the structure type of each decoding layer in the first network structure to obtain a second network structure, wherein the structure type of the encoding layer includes the first encoding layer and the second encoding layer, and the structure type of the decoding layer includes the first decoding layer and the second decoding layer; Based on the structural difference between the first network structure and the second network structure, adjusting the network parameters of the first semantic segmentation model; Based on the adjusted network parameters of the first semantic segmentation model, setting the network parameters of the second network structure; Based on the high-voltage transmission line point cloud data training set, the second network structure after the network parameters have been set is trained to obtain the target semantic segmentation model.

5. The method according to claim 4, characterized in that The network parameters based on the first semantic segmentation model adjust the structure types of each encoding layer and the structure types of each decoding layer in the first network structure to obtain a second network structure, including: Based on the first network structure, an initial population is determined, wherein the initial population includes a plurality of particles, each particle corresponds to one of the first network structures, and the first network structures corresponding to the particles have at least one encoding layer or decoding layer with a different structure type; Starting from the initial population, the following iterative operations are performed: Based on the network parameters of the first semantic segmentation model, the network parameters of the first network structure corresponding to each particle in the current iteration population are respectively set to obtain the second semantic segmentation model corresponding to each particle in the current iteration population; Determining, based on the high-voltage transmission line point cloud data training set, a semantic segmentation error of a second semantic segmentation model corresponding to each particle in the current iteration population; Based on the semantic segmentation error of the second semantic segmentation model corresponding to each particle in the current iteration population, updating the individual optimal fitness value; When the individual optimal fitness value meets the preset fitness value condition, the second network structure is determined based on the first network structure corresponding to the particle corresponding to the individual optimal fitness value.

6. The method according to claim 5, characterized in that The iterative operation also includes: When the individual optimal fitness value does not meet the fitness value condition, the structure type of the encoding layer or the decoding layer in the first network structure corresponding to each particle in the current iterative population is adjusted to obtain the next iterative population.

7. The method according to claim 6, characterized in that The adjusting the structure type of the encoding layer or the decoding layer in the first network structure corresponding to each particle in the current iteration population includes: Adjusting the structural types of the two adjacent coding layers based on the residuals between the two adjacent coding layers and the respective corresponding decoding layers in the first network structure corresponding to the particle; Based on the adjusted structure types of the two adjacent encoding layers, the structure types of the corresponding two decoding layers are adjusted to obtain the adjusted first network structure corresponding to the particle.

8. A semantic segmentation device for high-voltage transmission line laser radar point cloud, characterized in that: include: A network structure determination module, used to construct a first network structure based on a kernel point convolution layer, a graph edge convolution layer, and a channel and spatial attention module; A model training module, used to train the first network structure based on a training set of high-voltage transmission line point cloud data to obtain a target semantic segmentation model; A semantic segmentation module, configured to respond to a semantic segmentation request and perform semantic segmentation on the high-voltage transmission line point cloud data corresponding to the semantic segmentation request based on the target semantic segmentation model; Wherein, the first network structure includes an encoder and a decoder; The encoder includes at least one first encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second encoding layer consisting of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module; The decoder includes at least one first decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel, and at least one second decoding layer composed of the kernel point convolution layer and the graph edge convolution layer connected in parallel and then connected in series with the channel and spatial attention module.

9. A semantic segmentation system for high-voltage transmission line laser radar point cloud, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the processor, and the processor is used to obtain the instructions from the memory and execute the instructions, so that the processor can execute the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to be provided to a computer, so as to enable the computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Airborne LiDAR urban point cloud semantic segmentation method and system based on recursive residual double-attention kernel point convolutional network

    CN115861619A

  • Unmanned aerial vehicle crop state visual identification method for air-ground cooperation

    CN117409339A

  • Multi-source data fusion ground feature classification method based on multi-scale convolution auto-encoder

    CN117576483A

  • Semantic segmentation method, terminal equipment and storage medium

    CN117671252A

  • Seismic facies identification method and device based on CBAM attention mechanism

    CN119068230A