Transformer substation scene three-dimensional semantic segmentation method based on point cloud enhancement and attention guidance

By constructing an adaptive semantic segmented point cloud data set in the substation scenario, combining the attention-guided coding module and the point cloud enhancement module, the problem of insufficient accuracy of device recognition in the substation scenario is solved, and high-precision device classification and boundary details are achieved.

CN120431325AActive Publication Date: 2025-08-05ANHUI UNIV

Patent Information

Application Number
CN202510504556.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-05
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The existing point cloud processing methods are difficult to effectively deal with the problems of many types and complex shapes in substation scenarios, lack a feature enhancement mechanism for the geometric characteristics of power equipment, it is difficult to integrate local geometric details and global context information, and the ability to deal with uneven density point clouds is limited.

Method used

The three-dimensional semantic segmentation method of substation scenes based on point cloud enhancement and multi-head attention guidance is adopted. By constructing a semantic segmentation point cloud data set that is adapted to power scenes, combining attention guidance encoding module and point cloud enhancement module, geometric features are dynamically enhanced, and the problem of uneven distribution of equipment categories is solved, and collaborative modeling of local details and global context is realized.

Benefits of technology

It significantly improves the identification accuracy and robustness of equipment in substation scenarios, especially the identification ability of elongated structures such as wires and insulators, and improves the clarity and classification accuracy of equipment boundaries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431325A_ABST
    Figure CN120431325A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer substation scene three-dimensional semantic segmentation method based on point cloud enhancement and attention guidance, and the method specifically comprises the steps: carrying out the region division and category labeling of three-dimensional point cloud data in a transformer substation scene, setting a semantic tag based on an equipment type, and constructing a semantic segmentation point cloud data set adaptive to a power scene; a three-dimensional semantic segmentation model is constructed, and collaborative modeling of local details and global context is realized through a point cloud enhancement module and an attention guidance coding module; a weighted loss function and a parameter optimizer training model are adopted to solve the problem of unbalanced equipment category distribution in a substation scene; and performing semantic segmentation on the substation scene point cloud data based on the trained model, and outputting a high-precision equipment classification result. According to the method provided by the invention, through collaborative design of point cloud enhancement and attention guidance, the accuracy and robustness of three-dimensional semantic segmentation in a complex scene are effectively improved, and the method is particularly suitable for a scene with dense substation equipment and a fine structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and artificial intelligence, and specifically relates to a three-dimensional semantic segmentation method for substation scenes based on point cloud enhancement and attention guidance. Background Art

[0002] With the development of three-dimensional laser scanning technology, point cloud data processing has become an important means for intelligent operation and maintenance of substation equipment. Substations contain various equipment such as transformers, circuit breakers, and insulators, and these equipment present complex geometric features and multi-scale spatial distribution characteristics in point cloud data. Accurately identifying and segmenting these equipment is of great significance for the status assessment and fault warning of power facilities.

[0003] Existing point cloud processing methods face multiple challenges in substation scene applications. Traditional methods based on handcrafted features cannot effectively handle the problems of a large variety of equipment and complex shapes; deep learning-based methods such as the PointNet series have made breakthroughs, but still have problems such as insufficient feature expression ability and weak context awareness ability, especially the low recognition accuracy for fine structures. Although the attention mechanism has achieved remarkable results in the field of computer vision in recent years, when directly applied to point cloud data, it lacks consideration of geometric characteristics, and improved methods such as PointNeXt still face problems of accuracy and efficiency in substation scenes.

[0004] The main defects of existing methods include: lack of a feature enhancement mechanism for the geometric characteristics of power equipment; difficulty in fusing local geometric details and global context information; failure to fully utilize the relative position relationship in point clouds; and limited ability to process point clouds with uneven density.

[0005] To solve these problems, the present invention proposes a semantic segmentation method that integrates an attention-guided encoding module (AGE) and a point context enhancement module (PCE). AGE dynamically enhances geometric features through a multi-head attention mechanism and residual connections, and PCE improves the fine structure expression ability through dual-path feature extraction, jointly solving the key challenges of substation point cloud semantic segmentation. Summary of the Invention

[0006] Aiming at the deficiencies existing in the above-mentioned prior art, the present invention provides a three-dimensional semantic segmentation method for substation scenes based on point cloud enhancement and multi-head attention guidance, aiming to extract key features from complex point clouds and achieve accurate semantic segmentation of substation scene point cloud data.

[0007] To achieve the above technical objectives, the present invention provides the following technical solutions:

[0008] A three-dimensional semantic segmentation method for substation scenes based on point cloud enhancement and attention guidance, which specifically includes the following steps:

[0009] S1. Perform regional division and class annotation on the 3D point cloud data in the substation scene, set semantic labels based on equipment types, and construct a semantic segmentation point cloud dataset adapted to the substation scene;

[0010] S2. Construct a 3D semantic segmentation model. Through the point cloud enhancement module, perform density adaptive filtering and reflection intensity correction on the data in the input semantic segmentation point cloud dataset, suppress noise interference and enhance the geometric features of the equipment surface. At the same time, combine the attention-guided encoding module, and use the local convolution and multi-attention head mechanism to dynamically allocate the weights of different neighborhood features, and achieve the collaborative modeling of local details and global context;

[0011] S3. Train the model using a weighted loss function and a parameter optimizer to solve the problem of uneven distribution of equipment categories in the substation scene;

[0012] S4. Perform semantic segmentation on the substation scene point cloud data based on the trained model, and output high-precision equipment classification results.

[0013] Furthermore, step S1 specifically includes:

[0014] S11. Based on the spatial distribution characteristics and geometric attributes of the point cloud data, perform regional division on the substation scene, divide the overall point cloud into blocks of different spatial scales, and divide the training set and the test set; the training set and the test set are respectively composed of multiple independent blocks with different quantities, each block covers a different spatial range and contains point cloud data of all defined categories;

[0015] S12. For the point cloud data in the substation scene, classify the categories according to the manifestation form, functional attributes and spatial distribution information of the objects, and divide the point cloud data into non-substation equipment and substation equipment; among them, the non-substation equipment is further subdivided into busbars, wires, overheads, sundries, buildings, floors and walls;

[0016] S13. For the substation equipment point cloud data, further divide the substation equipment into small equipment, medium equipment and large equipment according to the physical scale of the equipment and the number of point clouds it contains;

[0017] S14. Use professional 3D annotation software to annotate the point cloud data classified in steps S12 and S13 with unique semantic labels, and integrate them under a unified semantic label system, so as to construct a labeled semantic segmentation point cloud dataset adapted to the substation scene;

[0018] S15. Then, use the mainstream 3D semantic segmentation algorithms to evaluate the performance of the constructed semantic segmentation point cloud dataset; verify the applicability of the obtained dataset in the 3D semantic segmentation task by recording and analyzing the precision, recall, and IoU metrics of each mainstream 3D semantic segmentation algorithm in differentiating different point cloud categories; if the obtained dataset is verified to be applicable, it is used for subsequent model training and 3D semantic segmentation tasks, and if it is verified not to have general applicability, it is necessary to re - perform region division, category division, and label annotation.

[0019] Further, the construction of the 3D semantic segmentation model in step S2 is specifically as follows:

[0020] Use the PointMetaBase network as the backbone network for learning the features of point cloud data to construct the 3D semantic segmentation model; the backbone network adopts an encoder - decoder structure, which consists of a multi - layer perceptron (MLP) layer, four down - sampling and abstraction modules, and three up - sampling and propagation modules.

[0021] The 3D semantic segmentation model takes the semantic segmentation point cloud dataset obtained in step S1 as input, and first uses a multi - layer perceptron (MLP) layer to enrich the information of the input point cloud data.

[0022] The down - sampling and abstraction modules are used to extract multi - level features; the input of the first down - sampling and abstraction module is the output of the MLP layer, and the input of the remaining down - sampling and abstraction modules is the output of the previous layer; each down - sampling and abstraction module includes an attention - guided encoding module (AGE), a point cloud enhancement module (PCE), and an inverted bottleneck MLP module.

[0023] Among them, the attention - guided encoding module (AGE) and the point cloud enhancement module (PCE) are connected in series as feature extraction modules within the down - sampling and abstraction module: the AGE module takes the central point feature and neighborhood feature as input, dynamically aggregates neighborhood information through the multi - head attention mechanism, and focuses on capturing global context relationships and geometric structure features; the output feature of the AGE module is then passed into the PCE module for further enhancement. The PCE module strengthens the local detail expression and fine geometric structure features through dual - path convolution and channel - space dual - path compression mechanisms; the inverted bottleneck MLP module receives the enhanced features output by the PCE module and further enriches the feature expression through the channel transformation of expansion - compression - expansion.

[0024] The three up - sampling and propagation modules gradually restore the multi - level features extracted by the down - sampling and abstraction modules to the original point cloud resolution through feature interpolation, feature fusion, and residual refinement mechanisms, and are connected to the corresponding - level down - sampling and abstraction modules through skip connections to achieve the fusion of high - and low - resolution features.

[0025] The output of the last upsampling propagation module is then fused with the output of the MLP layer as the final output features of the model; the final output features are passed through a classification head to output the class probability distribution of each point, completing the semantic segmentation task.

[0026] More specifically, the input of the backbone network is divided into two parts: one part is the spatial information of the point cloud data, denoted as P, which includes the spatial coordinates x, y, z of the point cloud data; the other part is the feature information of the point cloud, denoted as F, which includes the color information r, g, b and the reflection intensity i of the point cloud; the spatial information P is used for position embedding, and the feature information F is fed into the backbone network to extract global and local features. After feature learning by the backbone network, high-dimensional feature representations are obtained for subsequent semantic segmentation tasks.

[0027] Furthermore, the specific process of the attention-guided encoding module AGE for feature extraction is as follows:

[0028] First, the central point feature and the neighborhood feature are obtained; the central point feature directly comes from the output feature of the previous module, denoted as F center ∈R B×C×N ; then, a neighborhood is constructed around the central point through ball query, and the neighborhood feature is extracted from the global feature matrix according to the neighborhood index, denoted as F neighbor ∈R B×C×N×K , where B is the batch size, C is the number of feature channels, N is the number of central points, and K is the number of neighborhood points for each central point;

[0029] Secondly, batch normalization and linear transformation are performed on the input central point feature and neighborhood feature, and the relative position is encoded into three-dimensional coordinates and input into the multi-head attention layer. The specific calculation formula is:

[0030] P rel =P neighbor -P center ;

[0031] F′ center =BN(F center );

[0032] F′ neighbor =BN(F neighbor );

[0033] Q=W Q ·F′ center ;

[0034] K=W K ·F′ neighbor ;

[0035] V=W v ·F′ neighbor ;

[0036] Among them, P rel is the relative position encoding; BN represents the batch normalization operation, which is used to standardize the feature distribution; F' center and F' neighbor respectively represent the central point feature and the neighborhood feature after batch normalization; Q, K, and V are the query matrix, key matrix, and value matrix respectively; W Q , W K , W V are the corresponding weight matrices, which are used for linear transformation;

[0037] In the multi-head attention layer, the geometrically aware attention weights are calculated to dynamically aggregate the neighborhood features; the calculation formula of the geometrically aware attention weight Attention(Q, K, V) is:

[0038]

[0039] Among them, PE is the position encoding term, which is obtained by non-linear mapping of the relative position P rel ; d is the square root of the feature dimension, which is used for numerical stability; then the output of the multi-head attention layer MultiHead(Q, K, V) is:

[0040] MultiHead(Q, K, V) = W o ·Contact(head1,..., head h );

[0041] Among them, head1 and head h are the outputs of the first attention head and the h-th attention head respectively. There are h attention heads in total. Among them, the output head i of the i-th attention head = Attention(Q i , K i , V i ); W o is the linear transformation matrix;

[0042] Then, through the first residual connection, the original geometric information of the central feature is retained; we get

[0043] F attn = MultiHead(Q, K, V) + F center ;

[0044] After that, it is non-linearly enhanced by a two-layer feed-forward neural network with an expansion-compression structure; specifically:

[0045] First, in the expansion layer of the feed-forward neural network, the feature dimension is expanded from C to 2C through linear transformation to increase the feature expression ability, and the expanded feature F expand is obtained. The formula is expressed as:

[0046] F expand = W1·F attn + b1;

[0047] Then, the GELU activation function is adopted to introduce non-linearity, and through the feed-forward neural network compression layer, the feature dimension is compressed back from 2C to C again, obtaining the output F of the two-layer feed-forward neural network compress , which is expressed by the formula:

[0048] F compress = W2·GELU(F expand ) + b2;

[0049] Finally, multi-level features are fused through a second-order residual connection, and layer normalization is applied to obtain the output feature F of the attention-guided encoding module AGE AGE ; which is expressed by the formula:

[0050] F AGE = LayerNorm(F compress + F attn );

[0051] Among them, W1 and W2 are linear transformation matrices, b1 and b2 are bias terms; GELU() represents the GELU activation function, and LayerNorm() represents layer normalization

[0052] Furthermore, the specific process of the point cloud enhancement module PCE to extract features is as follows:

[0053] First, the central point is selected, and an adaptive neighborhood of the central point is constructed by the KNN method, dividing the input point cloud into local neighborhoods containing relative coordinate differences and feature differences; among them, the relative coordinate difference matrix X rel = Concat(Δx, x c ) is obtained by concatenating the coordinate difference Δx = x neighbor - x center between the neighborhood points and the central point and the central point coordinate x c ; the feature difference matrix X feat = Concat(Δf, f c ) is obtained by concatenating the feature difference Δf = f neighbor - f center between the neighborhood points and the central point feature f c ;

[0054] Secondly, the neighborhood geometric features and feature difference information are respectively extracted through a dual-path convolutional network, specifically:

[0055] In the geometric feature path of the dual-path convolutional network, the relative coordinate difference matrix is processed through convolution The calculation formula is as follows:

[0056] F xyz = ReLU(BN(Conv2D(X rel )));

[0057] Then, the feature difference matrix is processed by convolution in the feature difference path The calculation formula is as follows:

[0058] F feat = ReLU(BN(Conv2D(X feat )));

[0059] The fused feature is obtained through feature concatenation and max pooling operations:

[0060] F cat = MaxPool(Concat(F xyz , F feat ));

[0061] Then, a channel-spatial dual-path compression mechanism is adopted to generate global regularization features, and dimensionality reduction is performed through two independent convolutional paths respectively;

[0062] F ch\reduced = ReLU(Conv1D(F cat ));

[0063] F sp\reduced = ReLU(Conv1D(F cat ));

[0064] Among them, F ch\reduced is the dimensionality reduction feature of the channel path, and F sp\reduced is the dimensionality reduction feature of the spatial path; ReLU(.) is the ReLU activation function;

[0065] Through average pooling in the channel dimension and spatial dimension, the channel feature F ch and the spatial feature F sp are obtained, and the formula is expressed as:

[0066] F ch = Meanch(F ch\reduced );

[0067] F sp = Meansp(F sp\reduced );

[0068] Among them, Meanch() and Meansp() respectively represent average pooling in the channel dimension and average pooling in the spatial dimension;

[0069] The global regularization feature F global is generated through cross-operation and feature fusion:

[0070]

[0071] Finally, dynamic calibration of local features and global context is achieved through residual difference connection, and the Mish activation function is used to output highly discriminative enhanced features, that is, the output features of the PCE module. The calculation formula is:

[0072] F PCE = Mish(F cat - ReLU(BN(Conv1D(F global ))));

[0073] Where, B is the batch size, N is the number of points, k is the neighborhood size, C is the feature dimension, ∈ is a numerical stability constant, and F cat is the fused feature; the residual difference connection uses subtraction operation.

[0074] Furthermore, in S3:

[0075] The parameter optimizer is the AdamW optimizer;

[0076] The weighted loss function uses the cross-entropy loss function, and the formula is:

[0077]

[0078] Where, n represents the number of categories; y i is the one-hot encoding of the true label, and p i is the probability distribution of the model's class prediction for each point.

[0079] Based on the above technical solutions, the present invention has at least the following beneficial effects:

[0080] 1. Through the organic combination of the attention-guided encoding module AGE and the point cloud enhancement module PCE, the present invention gives full play to the respective advantages of the multi-head attention mechanism and local convolution, significantly improves the model's ability to extract local details and global context features, enhances the accuracy of semantic segmentation of substation data, and solves the technical problem of insufficient accuracy of existing three-dimensional semantic segmentation methods when facing complex point cloud data in substation scenarios;

[0081] 2. The present invention innovatively designs the PCE module. Through the dual-path feature extraction and channel-space dual-path compression mechanism, high-quality feature expression of the fine structure of substation equipment is achieved. In particular, the recognition ability of slender structures such as wires and insulators is significantly improved. At the same time, the application of the Mish activation function further enhances the discriminability of features;

[0082] 3. The AGE module of the present invention adopts a multi-head attention mechanism and a double residual connection structure, which can dynamically adjust the importance of different neighborhood features. At the same time, the understanding of the point cloud spatial structure by the model is enhanced through positional encoding. In the calculation of attention weights and feature fusion, positional encoding is organically combined, enabling the model to better perceive the geometric relationships of substation equipment;

[0083] 4. Based on the PointMetaBase architecture, the present invention effectively addresses the challenge of uneven point cloud density in the substation scenario through hierarchical downsampling and feature learning, and adopts an adaptive neighborhood construction method, improving the robustness of the model in the real environment. Compared with existing methods, the present invention achieves higher classification accuracy and better retention of boundary details in the substation equipment recognition task. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0085] Figure 1 is a flowchart of the method proposed by the present invention;

[0086] Figure 2 is a schematic diagram of the point cloud enhancement module PCE designed by the present invention;

[0087] Figure 3 is a schematic diagram of the attention-guided encoding module AGE designed by the present invention;

[0088] Figure 4 is a schematic diagram of the feature extraction module after combining the point cloud enhancement module and the attention-guided encoding module of the present invention;

[0089] Figure 5 is a schematic diagram of the three-dimensional semantic segmentation network framework proposed by the present invention;

[0090] Figure 6 is a schematic diagram of the three-dimensional semantic segmentation result of the point cloud data in the substation dataset of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0091] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the attached Figures 1-6 drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0092] Although the steps in this invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0093] In view of the technical deficiencies of the prior art in dealing with complex point cloud data in substation scenarios, an embodiment of the present invention proposes a three-dimensional semantic segmentation method for substation scenarios based on point cloud enhancement and attention guidance, as Figure 1 shown, which specifically includes the following steps:

[0094] S1. Perform regional division and category annotation on the three-dimensional point cloud data in the substation scenario, set semantic labels based on equipment types, and construct a semantic segmentation point cloud data set adapted to the substation scenario;

[0095] As a preferred implementation manner, step S1 specifically includes:

[0096] S11. Based on the spatial distribution characteristics and geometric attributes of the point cloud data, perform regional division on the substation scenario, divide the overall point cloud into blocks of different spatial scales, and divide the training set and the test set; the training set and the test set are respectively composed of multiple independent blocks with different quantities, each block covering a different spatial range and containing point cloud data of all defined categories; in this embodiment, a total of 6 fixed regions are planned, with region 5 as the test set, containing 16 blocks of different sizes and containing all categories; the other regions are used as the training set, containing a total of 29 blocks of different sizes and containing all categories;

[0097] S12. For the point cloud data in the substation scenario, classify according to the manifestation form, functional attributes and spatial distribution information of the objects, and divide the point cloud data into non-substation equipment and substation equipment; among them, the non-substation equipment is further divided into busbars, wires, overhead structures, sundries, buildings, ground and walls;

[0098] S13. For the point cloud data of substation equipment, further divide the substation equipment into small equipment, medium equipment and large equipment according to the physical scale of the equipment and the number of point clouds it contains;

[0099] S14. Use professional three-dimensional annotation software to annotate the point cloud data classified in steps S12 and S13 with unique semantic labels and integrate them into a unified semantic label system, thereby constructing a labeled semantic segmentation point cloud data set adapted to the substation scenario;

[0100] S15. Then, use the mainstream 3D semantic segmentation algorithms to evaluate the performance of the constructed semantic segmentation point cloud dataset; by recording and analyzing the precision, recall, and IoU metrics of each mainstream 3D semantic segmentation algorithm in differentiating different point cloud categories, verify the applicability of the obtained dataset in the 3D semantic segmentation task; if the obtained dataset is verified to be applicable, it is used for subsequent model training and 3D semantic segmentation tasks. If it is verified not to have general applicability, it is necessary to re - perform regional division, category division, and label annotation.

[0101] S2. Construct a 3D semantic segmentation model. Through the point cloud enhancement module, perform density - adaptive filtering and reflection intensity correction on the input data to suppress noise interference and enhance the geometric features of the device surface. At the same time, combine the attention - guided encoding module, and use the local convolution and multi - attention - head mechanism to dynamically allocate the weights of different neighborhood features, realizing the collaborative modeling of local details and global context.

[0102] As a preferred implementation, the 3D semantic segmentation model constructed in step S2 is specifically:

[0103] Use the PointMetaBase network as the backbone network for learning the features of point cloud data to construct a 3D semantic segmentation model; as Figure 5 shown, the backbone network adopts an encoder - decoder structure, which consists of a multi - layer perceptron MLP layer, four down - sampling abstraction modules, and three up - sampling propagation modules.

[0104] The 3D semantic segmentation model takes the semantic segmentation point cloud dataset obtained in step S1 as input, and first passes through a multi - layer perceptron MLP layer to enrich the information of the input point cloud data.

[0105] The down - sampling abstraction modules are used to extract multi - level features; the input of the first down - sampling abstraction module is the output of the MLP layer, and the input of the remaining down - sampling abstraction modules is the output of the previous layer; each down - sampling abstraction module includes an attention - guided encoding module AGE, a point cloud enhancement module PCE, and an inverted bottleneck MLP module.

[0106] Among them, the attention - guided encoding module AGE and the point cloud enhancement module PCE are connected in series as feature extraction modules within the down - sampling abstraction module: the AGE module takes the central point feature and neighborhood features as input, dynamically aggregates neighborhood information through the multi - head attention mechanism, focusing on capturing global context relationships and geometric structure features; the output features of the AGE module are then passed into the PCE module for further enhancement. The PCE module uses dual - path convolution and channel - space dual - path compression mechanisms to strengthen local detail expression and fine geometric structure features; as Figure 4As shown, the entire feature extraction module realizes the collaborative modeling of local details and global context by concatenating AGE and PCE and combining local features and global features; then, it receives the enhanced features output by the PCE module through the inverted bottleneck MLP module and performs channel transformation of expansion-compression-expansion to further enrich the feature representation;

[0107] In this application, the input of the entire backbone network is divided into two parts: one part is the spatial information of the point cloud data, denoted as P, which includes the spatial coordinates x, y, z of the point cloud data, and its dimension is (B, N, 3), where B represents the batch size and N represents the number of points; the other part is the feature information of the point cloud, denoted as F, which includes the color information r, g, b and the reflection intensity i of the point cloud, and its dimension is (B, N, C); the spatial information P is used for position embedding, and the feature information F is fed into the backbone network to extract global and local features, and after the feature learning of the backbone network, a high-dimensional feature representation is obtained for subsequent semantic segmentation tasks.

[0108] Next, the AGE module and the PCE module will be introduced separately;

[0109] In this embodiment, as Figure 3 shown, the process of the attention-guided encoding module AGE extracting features is specifically as follows:

[0110] First, the central point feature and the neighborhood feature are obtained; the central point feature directly comes from the output feature of the previous module, denoted as F center ∈R B×C×N ; then, a neighborhood is constructed around the central point through ball query, and the neighborhood feature is extracted from the global feature matrix according to the neighborhood index, denoted as F neighbor ∈R B×C×N×K , where B is the batch size, C is the feature channel number, N is the number of central points, and K is the number of neighborhood points for each central point;

[0111] Here, it should be noted that the so-called output feature of the previous module for the AGE module in the first downsampling and abstraction module is the point cloud data information enriched by the MLP layer, and for the AGE module in the remaining downsampling and abstraction modules, it is the output feature of the previous downsampling and abstraction module; in this application, before the AGE module of the downsampling and abstraction model, multiple central points will be extracted from the input point cloud data information by the farthest point sampling method, and the corresponding features are the central point features; the features corresponding to the neighborhood points of the central point are the neighborhood features; the global feature refers to the complete feature set containing all point cloud data, which is continuously transformed and enriched in the overall network;

[0112] Secondly, batch normalization and linear transformation are performed on the input central point feature and neighborhood feature, encoding the relative position into three-dimensional coordinates and inputting them into the multi-head attention layer. The specific calculation formula is:

[0113] P rel = P neighbor - P center ;

[0114] F′ center = BN(F center );

[0115] F′ neighbor = BN(F neighbor );

[0116] Q = W Q · F′ center ;

[0117] K = W K · F′ neighbor ;

[0118] V = W v · F′ neighbor ;

[0119] Where, P rel is the relative position encoding; BN represents the batch normalization operation for standardizing the feature distribution; F' center and F' neighbor respectively represent the central point feature and the neighborhood feature after batch normalization; Q, K, and V are the query matrix, the key matrix, and the value matrix respectively; W Q , W K , W V are the corresponding weight matrices for linear transformation;

[0120] In the multi-head attention layer, the geometric perception attention weights are calculated to dynamically aggregate the neighborhood features; the calculation formula for the geometric perception attention weights Attention(Q, K, V) is:

[0121]

[0122] Where, PE is the position encoding term obtained by non-linearly mapping the relative position P rel ; d is the square root of the feature dimension for numerical stability; then the output of the multi-head attention layer MultiHead(Q, K, V) is:

[0123] MultiHead(Q, K, V) = W o · Contact(head1,..., head h );

[0124] Where, head1, head hThey are the outputs of the 1st attention head and the hth attention head respectively. There are h attention heads in total, and the output of the ith attention head is head i = Attention(Q i , K i , V i ); W o is a linear transformation matrix;

[0125] Then, through the first residual connection, the original geometric information of the central feature is retained; and we get

[0126] F attn = MultiHead(Q, K, V) + F center ;

[0127] After that, it is non-linearly enhanced by a two-layer feed-forward neural network with an expansion-compression structure; specifically:

[0128] First, in the expansion layer of the feed-forward neural network, the feature dimension is expanded from C to 2C through linear transformation to increase the feature expression ability, and the expanded feature F expand is obtained. The formula is expressed as:

[0129] F expand = W1 · F attn + b1;

[0130] Then, the GELU activation function is adopted to introduce non-linearity, and through the compression layer of the feed-forward neural network, the feature dimension is compressed back from 2C to C again, and the output F compress of the two-layer feed-forward neural network is obtained. The formula is expressed as:

[0131] F compress = W2 · GELU(F expand ) + b2;

[0132] In this embodiment, the purpose of first increasing the dimension is to enhance the expression ability. By expanding the intermediate dimension, the feature has more space for non-linear transformation and feature interaction. Then reducing the dimension is to reduce the computational overhead;

[0133] Finally, the multi-level features are fused through a secondary residual connection, and layer normalization is applied to obtain the output feature F AGE of the attention-guided encoding module AGE; The formula is expressed as:

[0134] F AGE = LayerNorm(F compress + F attn );

[0135] Among them, W1 and W2 are linear transformation matrices, b1 and b2 are bias terms; GELU() represents the GELU activation function, and LayerNorm() represents layer normalization.

[0136] For another example Figure 2 As shown, the specific process of the point cloud enhancement module PCE to extract features is as follows:

[0137] First, select the center point, and construct an adaptive neighborhood of the center point through the KNN method, dividing the input point cloud into local neighborhoods containing relative coordinate differences and feature differences; where the relative coordinate difference matrix X rel = Concat(Δx, x c ) is obtained by concatenating the coordinate difference Δx = x neighbor - x center between the neighborhood points and the center point and the center point coordinate x c ; the feature difference matrix X feat = Concat(Δf, f c ) is obtained by concatenating the feature difference Δf = f neighbor - f center between the neighborhood points and the center point feature f c ;

[0138] It should be noted here that the main purpose of the PCE module in this application is to enhance the local context information of the point cloud, and a new feature representation needs to be constructed based on the spatial structure of the original point cloud. Therefore, even though the AGE module has processed the center point and neighborhood points, it still needs to be selected again in the PCE module. To reflect the difference from the AGE module, the neighborhood is constructed through the KNN method in the PCE module.

[0139] Secondly, extract the neighborhood geometric features and feature difference information through a dual-path convolutional network, specifically:

[0140] In the geometric feature path of the dual-path convolutional network, the relative coordinate difference matrix is processed through convolution The calculation formula is:

[0141] F xyz = ReLU(BN(Conv2D(X rel )));

[0142] Then, in the feature difference path, the feature difference matrix is processed through convolution The calculation formula is:

[0143] F feat = ReLU(BN(Conv2D(X feat )));

[0144] Obtain the fused feature through feature concatenation and max pooling operations:

[0145] F cat = MaxPool(Concat(Fxyz , F feat ));

[0146] Then, a channel-spatial dual-path compression mechanism is adopted to generate global regularization features, and the dimensionality is reduced through two independent convolutional paths respectively;

[0147] F ch\reduced = ReLU(Conv1D(F cat ));

[0148] F sp\reduced = ReLU(Conv1D(F cat ));

[0149] Among them, F ch\reduced is the dimensionality reduction feature of the channel path, and F sp\reduced is the dimensionality reduction feature of the spatial path; ReLU(.) is the ReLU activation function;

[0150] Through average pooling in the channel dimension and the spatial dimension, the channel feature F ch and the spatial feature F sp are obtained, and the formula is expressed as:

[0151] F ch = Meanch(F ch\reduced );

[0152] F sp = Meansp(F sp\reduced );

[0153] Among them, Meanch() and Meansp() respectively represent average pooling in the channel dimension and average pooling in the spatial dimension;

[0154] The global regularization feature F global is generated through cross-operation and feature fusion:

[0155]

[0156] Finally, dynamic calibration of local features and global context is achieved through residual difference connection, and the Mish activation function is adopted to output highly discriminative enhanced features, that is, the output features of the PCE module. The calculation formula is:

[0157] F PCE = Mish(F cat - ReLU(BN(Conv1D(F global ))));

[0158] Among them, B is the batch size, N is the number of points, k is the neighborhood size, C is the feature dimension, ∈ is a numerical stability constant, and F cat is the fused feature; the residual difference connection adopts a subtraction operation;

[0159] In this embodiment, after passing through four downsampling and abstraction modules, rich and comprehensive multi-level features are extracted. Then, through three upsampling and propagation modules, the multi-level features extracted by the downsampling and abstraction modules are gradually restored to the original point cloud resolution through feature interpolation, feature fusion, and residual refinement mechanisms, and are connected to the corresponding-level downsampling and abstraction modules through skip connections (that is, as shown in Figure 5 Figure, the input of upsampling propagation module 3 comes from downsampling abstraction module 1 and upsampling propagation module 2, the input of upsampling propagation module 2 comes from downsampling abstraction module 2 and upsampling propagation module 1, and the input of upsampling propagation module 1 comes from downsampling propagation module 3 and downsampling propagation module 4), thereby achieving the fusion of high- and low-resolution features;

[0160] Finally, the output of the last upsampling and propagation module is fused with the output of the MLP layer as the final output feature of the model; the final output feature outputs the category probability distribution of each point through the classification head to complete the semantic segmentation task.

[0161] S3. Use a weighted loss function and a parameter optimizer to train the model to solve the problem of uneven distribution of equipment categories in the substation scenario;

[0162] As a preferred implementation, the parameter optimizer in step S3 is the AdamW optimizer; in this embodiment, the learning rate of the optimizer is set to 0.01, and the weight decay rate is 1e -4 , and the model parameters are iteratively updated until the loss function L of the semantic segmentation model CE converges;

[0163] The weighted loss function uses the cross-entropy loss function, and the formula is:

[0164]

[0165] where n represents the number of categories; y i is the one-hot encoding of the true label, and p i is the category prediction probability distribution of the model for each point; through the loss function and the optimizer, the model can be trained to achieve the best prediction effect.

[0166] S4. Perform semantic segmentation on the substation scenario point cloud data based on the trained model to output a high-precision equipment classification result; in this embodiment, through the method proposed in this invention, the semantic segmentation result of the substation dataset is as shown in Figure 6 Figure; Figure 6It shows the comparison between the true labels of the original point cloud in the substation scenario and the prediction results processed by the method of the present invention. As can be seen from the results, the method of the present invention can accurately identify various types of equipment in the substation. In particular, the recognition accuracy for slender structures (such as wires) and small equipment is significantly better than that of the existing methods, and the equipment boundaries are also clearer and more accurate.

[0167] In summary, through the method combining point cloud enhancement and attention guidance, the present invention fully integrates local geometric features and global context information, effectively solves the technical problem of insufficient accuracy of the existing 3D semantic segmentation method when facing complex point cloud data in the substation scenario, and provides strong technical support for the intelligent operation and maintenance of power facilities.

[0168] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0169] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A 3D semantic segmentation method for substation scenes based on point cloud enhancement and attention guidance, characterized by: The following steps are involved: S1. Perform regional division and category labeling on the 3D point cloud data in the substation scene, set semantic labels based on device types, and construct a semantic segmentation point cloud dataset adapted to the substation scene. S2. Build a 3D semantic segmentation model. Use the point cloud enhancement module to perform density adaptive filtering and reflection intensity correction on the input semantic segmentation point cloud dataset to suppress noise interference and enhance the device surface geometric features. Combined with the attention-guided encoding module, it utilizes local convolution and multiple attention heads to dynamically assign weights to different neighborhood features, achieving collaborative modeling of local details and global context. S3, using weighted loss function and parameter optimizer to train the model to solve the problem of uneven distribution of equipment categories in substation scenarios; S4. Perform semantic segmentation on the substation scene point cloud data based on the trained model and output high-precision equipment classification results.

2. The method for 3D semantic segmentation of substation scenes based on point cloud enhancement and attention guidance according to claim 1 is characterized in that: Step S1 specifically includes: S11. Based on the spatial distribution characteristics and geometric properties of the point cloud data, the substation scene is divided into regions, the entire point cloud is divided into blocks of different spatial scales, and the training set and test set are divided. The training set and test set are composed of multiple independent blocks of different numbers, each covering a different spatial range and containing point cloud data of all defined categories. S12. Classify the point cloud data in the substation scene into non-substation equipment and substation equipment based on the object's form, functional attributes, and spatial distribution information. Non-substation equipment is further subdivided into busbars, power lines, overhead structures, debris, buildings, ground, and walls. S13. Based on the point cloud data of the substation equipment, the substation equipment is further subdivided into small equipment, medium equipment, and large equipment according to the physical size of the equipment and the number of point clouds it contains; S14. Use professional 3D annotation software to label the point cloud data classified in steps S12 and S13 with unique semantic labels, and integrate them into a unified semantic labeling system, thereby constructing a labeled semantic segmentation point cloud dataset suitable for the substation scene; S15. Use mainstream 3D semantic segmentation algorithms to perform performance evaluation on the constructed semantic segmentation point cloud dataset. Verify the applicability of the obtained dataset in the 3D semantic segmentation task by recording and analyzing the precision, recall rate, and IoU indicators of each mainstream 3D semantic segmentation algorithm in distinguishing different point cloud categories. If the obtained dataset is verified to be applicable, it will be used for subsequent model training and 3D semantic segmentation tasks. If it is verified to be not generally applicable, it will need to be re-divided into regions, categories, and labels.

3. The method for 3D semantic segmentation of substation scenes based on point cloud enhancement and attention guidance according to claim 1 is characterized in that: The three-dimensional semantic segmentation model is constructed in step S2 as follows: The PointMetaBase network is used as the backbone network for point cloud data feature learning to build a 3D semantic segmentation model. The backbone network adopts an encoder-decoder structure, consisting of a multi-layer perceptron (MLP) layer, four downsampling abstraction modules, and three upsampling propagation modules. The 3D semantic segmentation model takes the semantic segmentation point cloud dataset obtained in step S1 as input and first passes it through a multi-layer perceptron (MLP) layer to enrich the input point cloud data information. The downsampling abstraction module is used to extract multi-level features; The input of the first downsampling abstraction module is the output of the MLP layer, and the input of the remaining downsampling abstraction modules is the output of the previous layer; each downsampling abstraction module includes an attention-guided encoding module AGE, a point cloud enhancement module PCE, and an inverted bottleneck MLP module; Among them, the attention-guided encoding module AGE and the point cloud enhancement module PCE are connected in series as feature extraction modules within the downsampling abstraction module: the AGE module takes the center point feature and neighborhood feature as input, dynamically aggregates the neighborhood information through the multi-head attention mechanism, and focuses on capturing the global contextual relationship and geometric structure features; the output features of the AGE module are then passed to the PCE module for further enhancement. The PCE module strengthens the expression of local details and fine geometric structure features through dual-path convolution and channel-space dual-path compression mechanism; the inverted bottleneck MLP module receives the enhanced features output by the PCE module and further enriches the feature expression through the expansion-compression-expansion channel transformation; The three upsampling propagation modules gradually restore the multi-level features extracted by the downsampling abstraction module to the original point cloud resolution through feature interpolation, feature fusion and residual refinement mechanisms, and connect with the downsampling abstraction modules of the corresponding levels through skip connections to achieve the fusion of high- and low-resolution features. The output of the last upsampling propagation module is then fused with the output of the MLP layer as the final output feature of the model; The final output features are output through the classification head to output the category probability distribution of each point, completing the semantic segmentation task.

4. The method for 3D semantic segmentation of substation scenes based on point cloud enhancement and attention guidance according to claim 3 is characterized in that: The input of the backbone network is divided into two parts: one part is the spatial information of the point cloud data, represented by P, which contains the spatial coordinates x, y, and z of the point cloud data; the other part is the feature information of the point cloud, represented by F, which contains the color information r, g, b and reflection intensity i of the point cloud; the spatial information P is used for position embedding, and the feature information F is passed to the backbone network to extract global and local features. After the backbone network feature learning, a high-dimensional feature representation is obtained for subsequent semantic segmentation tasks.

5. The method for 3D semantic segmentation of substation scenes based on point cloud enhancement and attention guidance according to claim 4 is characterized in that: The specific process of feature extraction in the attention-guided encoding module AGE is as follows: First, obtain the center point features and neighborhood features; The center point feature comes directly from the output feature of the previous module, denoted as F center ∈R B×C×N ; Then, a neighborhood is constructed around the center point through ball query, and the neighborhood features are extracted from the global feature matrix according to the neighborhood index, which is recorded as F neighbor ∈R B×C×N×K , where B is the batch size, C is the number of feature channels, N is the number of center points, and K is the number of neighborhood points for each center point; Secondly, the input center point features and neighborhood features are batch normalized and linearly transformed, and the relative positions are encoded into three-dimensional coordinates and input into the multi-head attention layer. The specific calculation formula is: P rel =P neighbor -P center ; F′ center =BN(F center ); F′ neighbor =BN(F neighbor ); Q=W Q ·F′ center ; K=W K ·F′ neighbor ; V=W v ·F′ neighbor ; Among them, P rel is the relative position encoding; BN represents the batch normalization operation, which is used to standardize the feature distribution; F' center and F' neighbor Represent the center point features and neighborhood features after batch normalization respectively; Q, K, V are the query matrix, key matrix and value matrix respectively; W Q 、W K 、W V is the corresponding weight matrix for linear transformation; In the multi-head attention layer, the geometrically perceived attention weight is calculated to dynamically aggregate neighborhood features. The calculation formula of the geometrically perceived attention weight Attention(Q,K,V) is: Among them, PE is the position encoding item, which is composed of the relative position P rel Obtained through nonlinear mapping; d is the square root of the feature dimension, used for numerical stability; the output of the multi-head attention layer MultiHead(Q,K,V) is: MultiHead(Q,K,V)=W o ·Contact(head1,...,head h ); Among them, head1, head h are the outputs of the 1st and hth attention heads respectively. There are h attention heads in total, and the output of the i-th attention head is i =Attention(Q i ,K i ,V i );W o is the linear transformation matrix; Then, through the first residual connection, the original geometric information of the central feature is retained; we get F attn =MultiHead(Q,K,V)+F center ; Then, it is subjected to nonlinear enhancement through a double-layer feedforward neural network with an expansion-compression structure; specifically: First, in the feedforward neural network expansion layer, the feature dimension is expanded from C to 2C through linear transformation to increase the feature expression ability and obtain the expanded feature F expand , the formula is: F expand =W1·F attn +b1; Then the GELU activation function is used to introduce nonlinearity, and after the feedforward neural network compression layer, the feature dimension is compressed from 2C back to C to obtain the output F of the double-layer feedforward neural network. compress , the formula is: <h2 style=";text-align:left;direction:ltr">F<h2 style=";text-align:left;direction:ltr"> compress <h2 style=";text-align:left;direction:ltr"> =W2·GELU(F<h2 style=";text-align:left;direction:ltr"> expand <h2 style=";text-align:left;direction:ltr"> )+b2; Finally, the multi-level features are fused through the secondary residual connection and layer normalization is applied to obtain the output feature F of the attention-guided encoding module AGE. AGE ; The formula is: F AGE =LayerNorm(F compress +F attn ); Among them, W1 and W2 are linear transformation matrices, b1 and b2 are bias terms; GELU() represents the GELU activation function, and LayerNorm() represents layer normalization.

6. The method for 3D semantic segmentation of substation scenes based on point cloud enhancement and attention guidance according to claim 4 is characterized in that: The specific process of feature extraction in the point cloud enhancement module PCE is as follows: First, the center point is selected and the adaptive neighborhood of the center point is constructed by the KNN method to divide the input point cloud into local neighborhoods containing relative coordinate differences and feature differences; the relative coordinate difference matrix X rel =Concat(Δx,x c ) The coordinate difference between the neighborhood point and the center point Δx=x neighbor -x center and the center point coordinate x c Splicing to obtain; feature difference matrix X feat =Concat(Δf,f c ) The characteristic difference between the neighborhood point and the center point Δf=f neighbor -f center With the center point feature f c Splicing obtained; Secondly, the neighborhood geometric features and feature difference information are extracted through a dual-path convolutional network, specifically: Processing the relative coordinate difference matrix by convolution in the geometric feature path of the dual-path convolutional network The calculation formula is: F xyz =ReLU(BN(Conv2D(X rel ))); Then, the feature difference matrix is processed by convolution in the feature difference path The calculation formula is: F feat =ReLU(BN(Conv2D(X feat ))); The fusion features are obtained through feature concatenation and maximum pooling operations: F cat =MaxPool(Concat(F xyz ,F feat )); Then, a channel-space dual-path compression mechanism is used to generate global regularized features, and dimensionality reduction is performed through two independent convolution paths; F ch\reduced =ReLU(Conv1D(F cat )); F sp\reduced =ReLU(Conv1D(F cat )); Among them, F ch\reduced is the channel path dimension reduction feature, F sp\reduced is the spatial path dimensionality reduction feature; ReLU(.) is the ReLU activation function; The channel feature F is obtained by averaging the channel dimension and spatial dimension. ch and spatial characteristics F sp , the formula is: F ch =Meanch(F ch\reduced ); F sp =Meansp(F sp\reduced ); Among them, Meanch() and Meansp() represent channel dimension average pooling and spatial dimension average pooling respectively; Generate global regularization feature F through cross operation and feature fusion global : Finally, the dynamic calibration of local features and global context is achieved through residual differential connection, and the Mish activation function is used to output high-discrimination enhanced features, that is, the output features of the PCE module. The calculation formula is: F PCE =Mish(F cat -ReLU(BN(Conv1D(F global )))); Where B is the batch size, N is the number of points, k is the neighborhood size, C is the feature dimension, ∈ is a numerical stability constant, and F cat To fuse features; residual differential connection uses subtraction operation.

7. The method for 3D semantic segmentation of substation scenes based on point cloud enhancement and attention guidance according to claim 1 is characterized in that: In S3: The parameter optimizer is AdamW optimizer; The weighted loss function adopts the cross entropy loss function, and the formula is: Where n represents the number of categories; y i is the one-hot encoding of the true label, p i is the model's predicted probability distribution for each point.

Citation Information

Patent Citations

  • Digital twinborn scene-oriented point cloud automatic semantic modeling method

    CN116844004A

  • Point cloud semantic segmentation method based on feature fusion and attention mechanism

    CN116894940A

  • Three-dimensional indoor scene semantic segmentation method based on residual pulse neural network

    CN116958557A

  • Transformer substation scene three-dimensional semantic segmentation method based on combination of convolution and attention

    CN119559402A

  • Multi-modal point cloud quality evaluation method based on adaptive block projection

    CN119693934A

Cited By

  • Human body intelligent modeling method and system based on three-dimensional dot matrix

    CN120894506A

  • Transformer substation point cloud data processing method for constructing three-dimensional model

    CN120951609A

  • A Substation Cloud Data Processing Method for Constructing 3D Models

    CN120951609B

  • 3D semantic scene graph construction method based on neural network

    CN121121724A

  • A neural network-based 3D semantic scene graph construction method

    CN121121724B