Substation scene-oriented equipment component-level 3D instance segmentation method and related device
Through the combination of spherical progressive segmentation model and dynamic graph convolution network, the topological correlation fracture and boundary modeling distortion problems in substation equipment component-level diagnosis are solved, and accurate identification and healthy state modeling are achieved at the equipment component-level.
Patent Information
- Application Number
- CN202511042289.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-09-02
AI Technical Summary
The prior art cannot realize accurate identification and healthy state modeling of substation equipment components, and there are topological correlation fractures, fine-grained feature loss and boundary modeling distortion, resulting in untraceable component-level diagnosis.
The device component-level 3D instance segmentation method is adopted for substation scenarios, and the spherical progressive segmentation model and dynamic graph convolution network are combined with the virtual offset mechanism and the operation tree supervision space to achieve accurate segmentation of device instances and component levels.
It significantly improves defect positioning accuracy and equipment topological integrity, ensures traceability of the attribution relationship between components and parent equipment, and provides a reliable basis for assessing equipment health status.
Smart Images

Figure CN120580253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of substation equipment diagnosis, and in particular to a component-level 3D instance segmentation method and related devices for substation scene equipment. Background Art
[0002] In the construction of new power systems, accurate identification of millimeter-level defects in substation equipment components (such as insulator shed defects and grading ring displacement) is a core requirement for intelligent operations and maintenance. Traditional 2D vision technologies (such as Mask R-CNN) struggle to accurately quantify 3D deformation parameters due to perspective distortion and occlusion. Furthermore, 3D laser scanning suffers from bottlenecks such as low acquisition efficiency and insufficient point cloud density, hindering the implementation of precise component-level diagnosis.
[0003] Existing technologies use drone-based oblique photography to generate submillimeter-level color point clouds, combined with general 3D segmentation models for device identification. These methods include: data layer: multi-view image fusion to construct high-precision point clouds; algorithm layer: PointNet++ hierarchical random sampling or PV-RCNN point cloud-voxel fusion; and output layer: direct device-level semantic segmentation (e.g., locating the entire transformer) or segmenting individual components. However, these existing solutions suffer from the following major technical flaws, hindering traceable component-level diagnosis: 1) Broken topological associations: Direct component segmentation cannot establish "component-device" attribution (e.g., associating a high-voltage bushing cap with a specific transformer), resulting in a lack of structural support for defect localization and health assessment; 2) Loss of fine-grained features: Random sampling strategies dilute the point cloud features of millimeter-level components such as arrester grading rings, resulting in inadequate spatial occlusion modeling for densely assembled devices (e.g., insulator strings); and 3) Boundary modeling distortion: Axis-aligned bounding boxes (AABBs) create redundant space, making it difficult to match irregular geometries such as cylindrical bushings. Furthermore, metal oxidation artifacts under electromagnetic interference exacerbate missegmentation. These defects make it impossible for existing technologies to achieve traceable component-level diagnosis and difficult to meet the needs of accurate modeling of equipment health status. Therefore, it is urgent to solve the core problem of device-component hierarchical decoupling and establish a complete "component-equipment" mapping relationship. Summary of the Invention
[0004] The present invention provides a 3D instance segmentation method and related devices for substation scene equipment component level, which are used to solve the problem that the existing technology cannot achieve traceable component-level diagnosis due to topological association breaks, loss of fine-grained features and distorted boundary modeling, and is difficult to meet the needs of accurate modeling of equipment health status.
[0005] In view of this, a first aspect of the present invention provides a method for 3D instance segmentation of equipment components in a substation scene, the method comprising:
[0006] S1. Collect data from substation equipment to obtain oblique image data, preprocess the oblique image data to obtain preprocessed point cloud data, and divide the preprocessed point cloud data into point cloud grids;
[0007] S2. Input the segmented point cloud into the spherical progressive segmentation model to perform device instance segmentation, including:
[0008] The 3D backbone network extracts features from the input point cloud and generates several voting points. A spherical coordinate system is constructed with the voting points as the center. The point cloud space is evenly divided into several sectors according to the horizontal and vertical angles. The radial distance of the farthest foreground point in each sector is predicted as the instance boundary.
[0009] Optimize misclassified points through a virtual offset mechanism: learn positive radial offsets for false negative points to move them into the sector boundary, and learn negative offsets for false positive points to move them out of the boundary;
[0010] Combining the point positions after virtual migration with the predicted ray boundary distance to generate an instance binary mask, and determining the device point cloud after instance segmentation based on the instance binary mask;
[0011] S3. Input the device point cloud after instance segmentation into the component-level segmentation model for fine-grained segmentation, including:
[0012] Generate intermediate supervision features by operating the combined point cloud geometric features of the tree and the real part labels of the device point cloud after the instance segmentation;
[0013] The training domain is divided into mutually exclusive subdomains, and a leave-one-out validation strategy is used to optimize the intermediate supervision features to obtain supervision distribution parameters;
[0014] Sampling candidate supervisory features from the supervisory distribution parameters, gradually combining several supervisory features and evaluating their performance, and selecting the supervisory combination that achieves the greatest generalization improvement;
[0015] A dynamic graph convolutional network is used as the backbone network to simultaneously optimize the component classification task and the supervision combination, and output component-level segmentation results.
[0016] Optionally, step S1 includes:
[0017] Use drones to collect data from substation equipment and obtain oblique image data;
[0018] Preprocessing the oblique image data to obtain preprocessed point cloud data, wherein the preprocessing includes aerial triangulation encryption processing, denoising processing, and data labeling processing;
[0019] The pre-processed point cloud data is divided into point cloud meshes, and the division methods include the equidistant fixed mesh division strategy and the equal-step bidirectional extension method.
[0020] Optionally, the virtual offset mechanism includes:
[0021] Define the misclassification correction loss function , used to drive the migration of false negative points or false positive points:
[0022]
[0023]
[0024] in, represents the misclassification correction loss, represents the total number of misclassified points, is the direction control parameter, is the original spherical coordinate radius of the i-th misclassified point, is the predicted radial offset, is the length of the rough detection ray boundary of the i-th sector area, is the hyperbolic tangent function, yes Function that imparts a gentle gradient to points near the boundary.
[0025] Optionally, the optimizing of misclassified points by a virtual offset mechanism: learning a positive radial offset for a false negative point to move it into the sector boundary, and learning a negative offset for a false positive point to move it out of the boundary, further includes:
[0026] Define sector cohesion loss function , used to drive the true positive points to gather towards the centroid:
[0027] ;
[0028] Where, represents the loss of sector cohesion, is the number of true positive points, It is The coordinate radius of the true positive point sphere, It is The radial offset of the point, It is the hyperbolic tangent function to prevent gradient explosion. yes function, replacing hard threshold constraints and allowing gradual optimization.
[0029] Optionally, the process of processing data by the 3D backbone network includes:
[0030] The sparse convolution-based U-Net is used to encode point cloud features and generate voting points through a set abstraction layer.
[0031] Optionally, the process of processing data in the spherical coordinate system includes:
[0032] Point Cloud Convert to instance-centered Spherical coordinates centered on , the conversion formula is:
[0033] ;
[0034] in, Indicates that the Cartesian coordinate system Convert to spherical coordinates Mathematical conversion function of ;
[0035] ;
[0036] ;
[0037] ;
[0038] Where, Represent the radius, horizontal angle and vertical angle respectively, They represent the Cartesian coordinates of points in three-dimensional space.
[0039] Optionally, the optimization process in the task-conditional supervised distribution learning includes:
[0040] Based on the reinforcement learning optimization algorithm, the supervision distribution parameters are adjusted according to the generalization error reward to obtain the optimized supervision distribution parameters.
[0041] A second aspect of the present invention provides a 3D instance segmentation system for equipment components in a substation scene, the system comprising:
[0042] A data processing unit is used to collect data from substation equipment to obtain oblique image data, preprocess the oblique image data to obtain preprocessed point cloud data, and perform point cloud meshing on the preprocessed point cloud data;
[0043] The first segmentation unit is used to input the segmented point cloud into the spherical progressive segmentation model for device instance segmentation, including:
[0044] The 3D backbone network extracts features from the input point cloud and generates several voting points. A spherical coordinate system is constructed with the voting points as the center. The point cloud space is evenly divided into several sectors according to the horizontal and vertical angles. The radial distance of the farthest foreground point in each sector is predicted as the instance boundary.
[0045] Optimize misclassified points through a virtual offset mechanism: learn positive radial offsets for false negative points to move them into the sector boundary, and learn negative offsets for false positive points to move them out of the boundary;
[0046] Combining the point positions after virtual migration with the predicted ray boundary distance to generate an instance binary mask, and determining the device point cloud after instance segmentation based on the instance binary mask;
[0047] The second segmentation unit is used to input the device point cloud after instance segmentation into the component-level segmentation model for fine-grained segmentation, including:
[0048] Generate intermediate supervision features by operating the combined point cloud geometric features of the tree and the real part labels of the device point cloud after the instance segmentation;
[0049] The training domain is divided into mutually exclusive subdomains, and a leave-one-out validation strategy is used to optimize the intermediate supervision features to obtain supervision distribution parameters;
[0050] Sampling candidate supervisory features from the supervisory distribution parameters, gradually combining several supervisory features and evaluating their performance, and selecting the supervisory combination that achieves the greatest generalization improvement;
[0051] A dynamic graph convolutional network is used as the backbone network to simultaneously optimize the component classification task and the supervision combination, and output component-level segmentation results.
[0052] A third aspect of the present invention provides a device for component-level 3D instance segmentation of equipment in a substation scene, the device comprising a processor and a memory:
[0053] The memory is used to store program code and transmit the program code to the processor;
[0054] The processor is configured to execute the steps of the method for 3D instance segmentation of equipment component level for substation scenes as described in the first aspect according to the instructions in the program code.
[0055] A fourth aspect of the present invention provides a computer-readable storage medium for storing program code, wherein the program code is used to execute the 3D instance segmentation method for substation scene equipment component level described in the first aspect.
[0056] It can be seen from the above technical solutions that the present invention has the following advantages:
[0057] The embodiment of the present invention provides a 3D instance segmentation method for substation equipment component level, which includes the following key features: 1) Adopting a hierarchical segmentation architecture process: First, the substation equipment individual (such as transformer A) is obtained through instance segmentation, and then the component level segmentation is performed based on the instance mask to ensure the traceability of the component ownership relationship (such as identifying that the high-voltage bushing cap belongs to transformer A). This dual segmentation mechanism significantly improves the defect location accuracy by sacrificing some computational efficiency in exchange for the topological integrity of the equipment. 2) It is proposed to use the center point and radial distance in the spherical coordinate system ( and ) defines the instance boundaries, replacing the traditional axis-aligned bounding box (AABB). By evenly dividing the horizontal angle ( ) and vertical angle ( ) Generate sector areas (sectors), using the distance of the farthest foreground point in each sector as the boundary to reduce redundant space. 3) Design of two loss functions, misclassification correction loss ( ) For false positive and false negative points, a soft boundary loss with hyperbolic tangent (tanh) is used to force misclassified points to migrate to the correct area. Sector cohesion loss ( ) drives true positive points to cluster toward the center of the fan, enhancing semantic consistency within instances. 4) By constructing a supervision space (Parametric Supervision Space) containing geometric priors, it automatically searches for task-relevant intermediate supervision signals, rather than relying on manual design. 5) A greedy search process generates a specific algorithm for supervision combination step by step. 6) Reinforcement learning (REINFORCE) optimizes the supervision distribution model. This invention addresses the existing problems of inability to achieve traceable component-level diagnosis and meet the requirements for accurate equipment health status modeling due to broken topological associations, loss of fine-grained features, and distorted boundary modeling. It achieves the following technical advantages: 1) Spatial attribution traceability: The attribution relationship between components and parent equipment is clearly defined (for example, accurately locating a high-voltage bushing cap as belonging to a specific transformer A); 2) Improved defect localization accuracy: Device-level instance constraints effectively suppress the propagation of false positives during component segmentation; 3) Topological integrity assurance: A complete "component-equipment-scenario" three-dimensional semantic graph is established. This dual segmentation mechanism significantly improves defect localization accuracy while maintaining the integrity of the equipment topology, while moderately increasing computational complexity. This provides a reliable digital foundation for power equipment condition assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 A schematic diagram of a process flow of a component-level 3D instance segmentation method for substation scenes provided by an embodiment of the present invention;
[0060] Figure 2 An overall network diagram provided for an embodiment of the present invention;
[0061] Figure 3 A schematic structural diagram of a component-level 3D instance segmentation system for substation scenes provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0063] See also Figure 1 , an embodiment of the present invention provides a 3D instance segmentation method for substation scene equipment component level, comprising:
[0064] Step 101: collect data from substation equipment to obtain oblique image data, preprocess the oblique image data to obtain preprocessed point cloud data, and divide the preprocessed point cloud data into point cloud grids.
[0065] In one embodiment, step 101 includes:
[0066] Step 1011: collect data from substation equipment using a drone to obtain oblique image data;
[0067] It should be noted that drones are used to collect data from substation equipment. Specifically, the drones, equipped with RGB camera platforms, capture images from multiple angles along a preset route to obtain oblique image data. Note: Oblique image data is a set of image data acquired using oblique photography technology. Oblique photography involves equipping multiple sensors on the same flight platform to simultaneously capture images from different angles, such as vertical and oblique angles. Compared to traditional vertical orthophotos, oblique image data can provide richer side texture and geometric information of land features. It can realistically reflect the appearance, location, height, and other attributes of land features, constructing three-dimensional models. It has a wide range of applications in many fields, including urban planning, surveying and mapping, real estate, tourism, and disaster assessment. For example, in urban planning, it can help planners more intuitively understand the current state of the city; in surveying and mapping, it can improve measurement accuracy and efficiency.
[0068] Step 1012: Preprocess the oblique image data to obtain preprocessed point cloud data, where the preprocessing includes aerial triangulation encryption processing, denoising processing, and data labeling processing.
[0069] It should be noted that:
[0070] First, the oblique image obtained in step 1011 is subjected to aerial triangulation encryption processing (aerial triangulation encryption processing is an important process in photogrammetry for determining the exterior orientation elements of the image and the ground coordinates of the encryption points. Its specific process generally includes steps such as internal orientation, relative orientation, absolute orientation and regional block adjustment. Aerial triangulation encryption processing is widely used in surveying and mapping, remote sensing, geographic information systems and other fields. It can provide high-precision control point data for subsequent topographic mapping, orthophoto production, 3D modeling and other tasks, thereby improving work efficiency and quality of results.), a geometric constraint network is constructed, camera parameters and 3D coordinates are optimized, a high-precision digital surface model and elevation model are generated, and a high-density 3D color point cloud is generated by matching the same-name points between multiple views. At the same time, a dense matching algorithm is used to convert the calculation results into a high-density 3D color point cloud.
[0071] Then, 1) the point cloud data is denoised using a straight-through filter method. The original point cloud data is loaded and preprocessed to remove invalid points. 2) The target coordinate axis is selected according to the scene requirements, and the threshold range is set to retain the valid area. 3) The straight-through filter algorithm is called, the parameters are input, and the filtering is performed to remove noise points outside the range. If multi-dimensional filtering is required, this process can be repeated in the X and Y axes to further constrain the spatial range. 4) The filtered point cloud is output.
[0072] Finally, the point cloud, denoised by a direct-pass filter, is imported into CloudCompare software. This involves: 1) creating global labels at the device level (e.g., for the entire voltage transformer), using a 3D oriented bounding box to calibrate the spatial extent of the device's main body; 2) performing step-by-step component annotation within the calibrated device space based on a hierarchical label tree structure. For the compression bushing, the frame parameters are adaptively adjusted to match its geometric shape based on the axial cylindrical features; for the porcelain insulator string, an axially symmetric segmentation model is used to extract the shed contour layer by layer along the central axis; and for the lightning arrester instrument panel, a plane detection algorithm is combined to calibrate rectangular or circular feature areas. A spatial registration module is introduced for the device CAD model, and the point cloud is aligned using the ICP algorithm for accurate annotation. 3) Structured annotation data containing semantic relationships at the device-component level is output to ensure that the training and validation sets retain a complete hierarchical feature distribution.
[0073] Step 1013: performing point cloud mesh division on the pre-processed point cloud data, wherein the division methods include an equidistant fixed mesh division strategy and an equal-step bidirectional extension method.
[0074] It should be noted that based on the preprocessed point cloud data, an equidistant fixed grid division strategy was adopted to establish an XY axis orthogonal grid coordinate system with reference point O as the origin. The X axis was aligned with the main direction of the transmission corridor, and the Y axis was orthogonal to the axis of the main transformer group. Using an equal-step bidirectional extension method, multi-axial extension was performed along the positive and negative directions of the X / Y axes. Using a spatial hash mapping algorithm, the point cloud data was dynamically allocated to standard 40m×40m spatial blocks, thus completing the division of the substation site cloud.
[0075] Notes: 1) Equidistant fixed gridding is a regular and uniform division method that divides the space containing the point cloud data into equally sized grid cells based on a uniform spacing standard. The advantages of this division method lie in its simplicity and standardization. It ensures that each grid cell has a consistent spatial scale, facilitating the subsequent unified analysis and processing of the point cloud data within each grid cell. By constructing an orthogonal grid coordinate system based on a reference point, the division process becomes more directional and systematic, ensuring that the entire division result closely aligns with the layout characteristics of the transmission corridor and main transformer group, providing a clear and organized foundation for subsequent annotation and analysis. 2) The equal-step bidirectional extension method is an efficient and reasonable method for extending point cloud data. It performs multi-axial extension along the positive and negative directions of the X and Y axes, maintaining a consistent step size throughout the process. This consistency ensures uniformity and standardization of the data during the extension process, avoiding the potential confusion in the division caused by varying step sizes. Through bidirectional extension with equal step length, point cloud data can be reasonably distributed to a wider spatial range, allowing the originally limited point cloud data to cover a larger area. Especially for scenes with a certain scale and complex layout such as substations, their spatial information can be obtained more comprehensively.
[0076] For steps 102 to 103, please refer to Figure 2 :
[0077] Step 102: Input the segmented point cloud into the spherical progressive segmentation model to perform device instance segmentation, including:
[0078] It should be noted that in order to improve the overall accuracy and robustness of multi-instance segmentation in complex 3D scenes, the present invention designs a spherical progressive segmentation model to overcome the limitations of the traditional coarse-to-fine method and achieve more accurate instance segmentation. Among them, the spherical progressive segmentation model includes two major modules: 3D backbone network and instance mask estimation. The 3D backbone network consists of a 3D encoder and a voting module to extract point cloud features and generate voting points. The instance mask estimation module generates accurate instance segmentation results through a two-stage process of coarse detection and refinement. The specific steps are as follows:
[0079] Step 1021: Extract features of the input point cloud using a 3D backbone network and generate several voting points. A spherical coordinate system is constructed with the voting points as the center. The point cloud space is evenly divided into several sectors based on horizontal and vertical angles. The radial distance of the farthest foreground point in each sector is predicted as the instance boundary.
[0080] It should be noted that for the 3D backbone network module specifically: given an input point cloud and its corresponding color information ,in The point cloud enters the 3D encoder, which uses a sparse convolution-based U-Net to encode the input point cloud into deep features. The voting module then receives and , generated through collection abstraction Voting points with coordinates of , characterized by ,in is the feature dimension.
[0081] Specifically for the instance mask estimation module: the point cloud and voting characteristics All inputs are fed into the instance mask estimation module for radial instance detection, aiming to represent instances through 3D polygons in spherical coordinates. Here, "radial" refers to rays radiating from the instance center along different angles in the spherical coordinate system. The spherical coordinate system uses radial distances to represent the instance. , horizontal angle and vertical angle Three parameters to describe the three-dimensional position of a point.
[0082] In order to overcome the inherent defects of the traditional axis-aligned bounding box and achieve more accurate instance boundary modeling, the concept of ray is introduced. Each instance is defined as a and multiple rays The 3D polygons are formed. Rays are emitted from the center to form a spherical sector. Each ray determines the boundary distance of its corresponding sector through a preset angle. The MLP network is used to process each voting feature. , predict the center offset. Add the offset to the voting point On, get the instance center Then the point cloud Convert to Spherical coordinates centered on :
[0083] ;
[0084] in Indicates that the Cartesian coordinate system Convert to spherical coordinates Mathematical conversion function.
[0085] , the conversion formula is as follows:
[0086] ;
[0087] ;
[0088] ;
[0089] in Represent the radius, horizontal angle and vertical angle respectively, They represent the Cartesian coordinates of points in three-dimensional space.
[0090] Then, the spherical coordinate point cloud In the horizontal and vertical directions, and is divided into even intervals, and finally forms sectors. Each sector corresponds to a unique ( Combination. Predict the maximum radius of each sector through RayHead (MLP network) , if the radius of the point in the sector is less than It is considered as prospect.
[0091] The loss function of its rough detection is as follows:
[0092] ;
[0093] Among them, the ray loss use Loss Comparison Prediction Ray Distance to the farthest foreground point :
[0094] ;
[0095] in Set to The furthest foreground point in the sector If there is no foreground point in the sector, Set to minimum value .
[0096] Center point loss use Loss Comparison and Forecasting Center The mean of the true foreground points :
[0097] ;
[0098] in, is the mean Cartesian coordinate of the foreground points of the ground-truth instance.
[0099] Step 1022: Optimize misclassified points through a virtual offset mechanism: learn a positive radial offset for false negative points to move them into the sector boundary, and learn a negative offset for false positive points to move them out of the boundary;
[0100] It should be noted that due to the inherent defects of the rough detection stage, the boundary may contain misjudged points of background or other instances (false positives), and some real instance points may be missed (false negatives). This paper proposes a dual optimization mechanism to dynamically optimize the accuracy of the instance attribution by predicting the virtual offset vector of the point cloud in radial space. Specifically, this method learns a radial increment for each point, so that it moves along the ray direction to the instance centroid while maintaining the vertical angle and horizontal angle Its essence is to use the offset to implicitly correct the spatial division boundary of the coarse detection sector without changing the geometric distribution of the original point cloud, thereby generating clear instance segmentation labels.
[0101] By for each point Estimate an offset value so that these misclassified points can be virtually migrated to the correct area. Construct a point migration head to predict point-by-point parameters. . The voting feature As query, with point features To maintain symbol consistency, we use express An output of corresponds to a voting process. We divide these points into two groups, one group learns radial increments in case of misclassification, and the other group enhances the compactness and structural cohesion of instances by migrating points to the sector centroid.
[0102] A set of radial increments in the case of misclassification is learned, defining the set of false negative points, i.e. points that actually belong to the current instance but were missed by the coarse detection stage:
[0103] ={ : ;
[0104] in, Indicates the spherical coordinate system foreground points (i.e., points that actually belong to the current instance). The predicted The radial offset of the foreground point (used to virtually adjust the radius of the point). Indicates the predicted ray length of the sector where the point is located ), which means the azimuth of the point Determine the index of the sector to which it belongs. It means that the adjusted point radius exceeds the sector boundary ray length, that is, the point is still outside the sector and is therefore misclassified as background (false negative).
[0105] Define the set of false positive points, which are points that do not actually belong to the instance but are mistakenly included in the predicted sector after migration:
[0106] ={ : ;
[0107] in, Indicates the spherical coordinate system background points (i.e., points that do not actually belong to the current instance), The predicted The radial offset of the background points, Indicates the predicted ray length of the sector where the point is located, = It is to determine the sector to which the background point belongs. It means that the adjusted point radius is smaller than the sector boundary ray length, that is, the point is incorrectly included in the sector (false positive).
[0108] Misclassified point set:
[0109] ;
[0110] Its misclassification correction loss function is as follows:
[0111] ;
[0112] ;
[0113] in, represents the misclassification correction loss, represents the total number of misclassified points, is the direction control parameter, is the original spherical coordinate radius of the i-th misclassified point, is the predicted radial offset, is the coarse detection ray boundary, It is the hyperbolic tangent function, which limits the gradient range of the offset to prevent unstable training. yes The function is to give a gentle gradient to points near the boundary.
[0114] The other group moves the true positive points to the centroid of the sector to enhance the compactness and structural cohesion of the instance. This approach encourages the model to learn the common and shared features of the instances because the foreground features will gradually move closer to each other. In addition, this mechanism can also provide learning signals for the true positive points, making up for the Only considering the insufficiency of false negatives and false positives. Similarly, using the index of the foreground point , screening true positive points , which is defined as:
[0115] ={ : ;
[0116] in, Indicates the spherical coordinate system foreground points (i.e., points that actually belong to the current instance). The predicted The radial offset of each foreground point. Indicates the predicted ray length of the sector where the point is located. ), which means the azimuth of the point Determine the index of the sector to which it belongs. It means that the adjusted point radius is smaller than the sector boundary ray length. The point is located within the boundary after offset and is therefore marked as a true positive point (correctly classified as an instance point).
[0117] Its sector cohesion loss function is:
[0118] ;
[0119] in, represents the loss of sector cohesion, is the number of true positive points, It is The coordinates of the true positive points, It is The radial offset of the point, It is the hyperbolic tangent function to prevent gradient explosion. yes function, replacing hard threshold constraints and allowing gradual optimization.
[0120] So the refinement loss function is:
[0121] ;
[0122] Step 1023: Generate an instance binary mask by combining the point position after virtual migration and the predicted ray boundary distance, and determine the device point cloud after instance segmentation based on the instance binary mask;
[0123] It should be noted that the Mask Assembly module is used to integrate the rough detection and refinement results to generate an accurate instance binary mask. , use the threshold to determine whether each point belongs to an instance:
[0124] ;
[0125] in, point After the offset adjustment, it is still within the predicted boundary (ray length) of the current instance and is considered a true positive (mask is 1). Points that exceed the boundary after offset may be false positives or background points and need to be excluded (mask is 0).
[0126] Step 103: Input the device point cloud after instance segmentation into the component-level segmentation model for fine-grained segmentation, including:
[0127] It should be noted that to address the refined identification and condition monitoring of complex power equipment, a component-level segmentation model for power equipment was constructed. This model consists of three core modules: parameterized supervision space modeling, task-conditional supervision distribution learning, and a greedy supervision selection strategy. These three modules collaborate to automatically discover optimal intermediate supervision. Through geometric prior-guided feature space construction and task-adaptive supervision selection, they effectively suppress shortcut features and improve cross-domain segmentation generalization capabilities. The specific steps are as follows:
[0128] Step 1031, parameterizing the supervision space: generating intermediate supervision features by combining the point cloud geometric features of the operation tree and the real component labels of the device point cloud after instance segmentation.
[0129] It should be noted that the operation tree contains grouping operators, unary operators, and binary operators;
[0130] It should be noted that we first design a parameterized supervision space that contains geometric prior knowledge. The supervision space is defined by an operation tree. Through the tree structure, geometric features and real component labels are combined to generate a variety of intermediate supervision features. The goal is to find supervision signals that can effectively suppress shortcut features and enhance real component clues through automated search. An operation tree is randomly sampled in the
[15] , which consists of three basic operators: grouping operations (such as summation and SVD) are used to extract local or global geometric statistical information, unary operations (such as square and normalization) generate diverse intermediate features, and binary operations (addition and subtraction) generate complex combination features.
[0131] Combine the original geometric features of the point cloud (such as normal vectors, curvature, etc.) with the real component labels after instance segmentation, convert them into component-aware features through pre-coding, input the component-aware features into the sampled operation tree, perform grouping operations, unary operations and binary operations layer by layer according to the tree structure, and finally generate intermediate supervision features For each point , calculate the corresponding supervision truth value through the same operation tree according to its input features .
[0132] Step 1032: Task-Conditioned Supervised Distribution Learning: Divide the training domain into mutually exclusive subdomains, and use a leave-one-out validation strategy to optimize the intermediate supervisory features to obtain supervisory distribution parameters, thereby minimizing the out-of-distribution generalization error.
[0133] It should be noted that the training domain is divided into multiple mutually exclusive subdomains, and a leave-one-out validation strategy is adopted: one subdomain is used as the validation domain at a time, and the remaining subdomains are used as the training domain. By calculating the average generalization error under all possible training-validation partitions, the performance of the sampled supervision feature in the out-of-distribution (OOD) scenario is evaluated. The proposed supervision feature s is input into the network together with the standard segmentation task supervision for training. By comparing the reduction in the generalization error of the model before and after the introduction of s, the generalization gain of the supervision feature is quantified. The reinforcement learning (REINFORCE) algorithm is used to adjust the distribution parameters according to the generalization error reward. , increasing the probability of high-value supervision.
[0134] Step 1033, greedy supervision selection: sampling candidate supervision features from the supervision distribution parameters, gradually combining several supervision features and evaluating the performance, and selecting the supervision combination with the greatest generalization improvement;
[0135] It should be noted that the greedy search strategy is used to optimize the supervision space (through the distribution parameters of reinforcement learning) ), sample multiple candidate supervisions, evaluate their performance, retain the best single supervision, and based on the single supervision results, gradually combine multiple supervision features (1-3), evaluate the combination performance, and finally select the supervision combination with the largest generalization improvement.
[0136] Step 1034, dynamic graph convolutional segmentation: Use the dynamic graph convolutional network as the backbone network to simultaneously optimize the component classification task and supervision combination, and output the component-level segmentation result.
[0137] It should be noted that the dynamic graph convolutional network (DGCNN) is finally used as the backbone network. The local geometric features and global shape features of the point cloud are extracted through multiple layers of dynamic graph convolution layers, and multi-scale information is fused using jump connections. The backbone network is followed by a classification-based segmentation module, which consists of a fully connected layer and maps the features of each point to the corresponding component category label. The intermediate supervision module automatically searches for task-related geometric prior features through the operation tree to generate supervision signals. The backbone network simultaneously optimizes the primary segmentation task (component classification) and the intermediate supervision task (predicting the searched geometric features). During the inference phase, the intermediate supervision module is removed, retaining only the optimized backbone network and classification layer. After feature extraction from the input test point cloud, the component labels are directly output to form the final segmentation result.
[0138] The embodiment of the present invention provides a 3D instance segmentation method for substation equipment component level, which includes the following key features: 1) Adopting a hierarchical segmentation architecture process: First, the substation equipment individual (such as transformer A) is obtained through instance segmentation, and then the component level segmentation is performed based on the instance mask to ensure the traceability of the component ownership relationship (such as identifying that the high-voltage bushing cap belongs to transformer A). This dual segmentation mechanism significantly improves the defect location accuracy by sacrificing some computational efficiency in exchange for the topological integrity of the equipment. 2) It is proposed to use the center point and radial distance in the spherical coordinate system ( and ) defines the instance boundaries, replacing the traditional axis-aligned bounding box (AABB). By evenly dividing the horizontal angle ( ) and vertical angle ( ) Generate sector areas (sectors), using the distance of the farthest foreground point in each sector as the boundary to reduce redundant space. 3) Design of two loss functions, misclassification correction loss ( ) For false positive and false negative points, a soft boundary loss with hyperbolic tangent (tanh) is used to force misclassified points to migrate to the correct area. Sector cohesion loss ( ) drives true positive points to cluster toward the center of the fan, enhancing semantic consistency within instances. 4) By constructing a supervision space (Parametric Supervision Space) containing geometric priors, it automatically searches for task-relevant intermediate supervision signals, rather than relying on manual design. 5) A greedy search process generates a specific algorithm for supervision combination step by step. 6) Reinforcement learning (REINFORCE) optimizes the supervision distribution model. This invention addresses the existing problems of inability to achieve traceable component-level diagnosis and meet the requirements for accurate equipment health status modeling due to broken topological associations, loss of fine-grained features, and distorted boundary modeling. It achieves the following technical advantages: 1) Spatial attribution traceability: The attribution relationship between components and parent equipment is clearly defined (for example, accurately locating a high-voltage bushing cap as belonging to a specific transformer A); 2) Improved defect localization accuracy: Device-level instance constraints effectively suppress the propagation of false positives during component segmentation; 3) Topological integrity assurance: A complete "component-equipment-scenario" three-dimensional semantic graph is established. This dual segmentation mechanism significantly improves defect localization accuracy while maintaining the integrity of the equipment topology, while moderately increasing computational complexity. This provides a reliable digital foundation for power equipment condition assessment.
[0139] The above is a method for 3D instance segmentation of equipment components at the substation scene provided in an embodiment of the present invention. The following is a system for 3D instance segmentation of equipment components at the substation scene provided in an embodiment of the present invention.
[0140] See also Figure 3 , an embodiment of the present invention provides a 3D instance segmentation system for substation scene equipment component level, including:
[0141] The data processing unit 201 is used to collect data from the substation equipment to obtain oblique image data, pre-process the oblique image data to obtain pre-processed point cloud data, and perform point cloud meshing on the pre-processed point cloud data;
[0142] The first segmentation unit 202 is configured to input the segmented point cloud into a spherical progressive segmentation model for device instance segmentation, including:
[0143] The 3D backbone network extracts features from the input point cloud and generates several voting points. A spherical coordinate system is constructed with the voting points as the center. The point cloud space is evenly divided into several sectors according to the horizontal and vertical angles. The radial distance of the farthest foreground point in each sector is predicted as the instance boundary.
[0144] Optimize misclassified points through a virtual offset mechanism: learn positive radial offsets for false negative points to move them into the sector boundary, and learn negative offsets for false positive points to move them out of the boundary;
[0145] Combine the point positions after virtual migration and the predicted ray distance to generate an instance binary mask, and determine the device point cloud after instance segmentation based on the instance binary mask;
[0146] The second segmentation unit 203 is configured to input the device point cloud after instance segmentation into the component-level segmentation model for fine-grained segmentation, including:
[0147] Generate intermediate supervision features by combining the point cloud geometric features of the operation tree with the real part labels of the device point cloud after instance segmentation;
[0148] The training domain is divided into mutually exclusive subdomains, and the leave-one-out validation strategy is used to optimize the intermediate supervision features to obtain the supervision distribution parameters.
[0149] Sample candidate supervisory features from the supervisory distribution parameters, gradually combine several supervisory features and evaluate their performance, and select the supervisory combination that has the greatest generalization improvement;
[0150] A dynamic graph convolutional network is used as the backbone network to simultaneously optimize the component classification task and supervision combination, and output component-level segmentation results.
[0151] Furthermore, an embodiment of the present invention also provides a device for component-level 3D instance segmentation of equipment in a substation scene, the device including a processor and a memory:
[0152] The memory is used to store program code and transmit the program code to the processor;
[0153] The processor is configured to execute the steps of the substation scene device component-level 3D instance segmentation method as described in the above method embodiment according to the instructions in the program code.
[0154] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the substation scene device component-level 3D instance segmentation method described in the above method embodiment.
[0155] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0156] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0157] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0158] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0159] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0160] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A 3D instance segmentation method for equipment component level in substation scenes, characterized by: include: S1. Collect data from substation equipment to obtain oblique image data, preprocess the oblique image data to obtain preprocessed point cloud data, and divide the preprocessed point cloud data into point cloud grids; S2. Input the segmented point cloud into the spherical progressive segmentation model to perform device instance segmentation, including: The 3D backbone network extracts features from the input point cloud and generates several voting points. A spherical coordinate system is constructed with the voting points as the center. The point cloud space is evenly divided into several sectors according to the horizontal and vertical angles. The radial distance of the farthest foreground point in each sector is predicted as the instance boundary. Optimize misclassified points through a virtual offset mechanism: learn positive radial offsets for false negative points to move them into the sector boundary, and learn negative offsets for false positive points to move them out of the boundary; Combining the point positions after virtual migration with the predicted ray boundary distance to generate an instance binary mask, and determining the device point cloud after instance segmentation based on the instance binary mask; S3. Input the device point cloud after instance segmentation into the component-level segmentation model for fine-grained segmentation, including: Generate intermediate supervision features by operating the combined point cloud geometric features of the tree and the real part labels of the device point cloud after the instance segmentation; The training domain is divided into mutually exclusive subdomains, and a leave-one-out validation strategy is used to optimize the intermediate supervision features to obtain supervision distribution parameters; Sampling candidate supervisory features from the supervisory distribution parameters, gradually combining several supervisory features and evaluating their performance, and selecting the supervisory combination that achieves the greatest generalization improvement; A dynamic graph convolutional network is used as the backbone network to simultaneously optimize the component classification task and the supervision combination, and output component-level segmentation results.
2. The method for 3D instance segmentation of equipment components at the substation scene according to claim 1 is characterized in that: Step S1 includes: Use drones to collect data from substation equipment and obtain oblique image data; Preprocessing the oblique image data to obtain preprocessed point cloud data, wherein the preprocessing includes aerial triangulation encryption processing, denoising processing, and data labeling processing; The pre-processed point cloud data is divided into point cloud meshes, and the division methods include the equidistant fixed mesh division strategy and the equal-step bidirectional extension method.
3. The method for 3D instance segmentation of equipment components at the substation scene according to claim 1 is characterized in that: The virtual offset mechanism includes: Define the misclassification correction loss function , used to drive the migration of false negative points or false positive points: in, represents the misclassification correction loss, represents the total number of misclassified points, is the direction control parameter, is the original spherical coordinate radius of the i-th misclassified point, is the predicted radial offset, is the length of the rough detection ray boundary of the i-th sector area, is the hyperbolic tangent function, yes Function that assigns a gentle gradient to points near the boundary.
4. The method for 3D instance segmentation of equipment components at the substation scene according to claim 1 is characterized in that: The method of optimizing misclassified points by using a virtual offset mechanism includes: learning a positive radial offset for false negative points to move them into the sector boundary, and learning a negative offset for false positive points to move them out of the boundary, and further includes: Define sector cohesion loss function , used to drive the true positive points to gather towards the centroid: ; Where, represents the loss of sector cohesion, is the number of true positive points, It is The spherical coordinate radius of the true positive point, It is The radial offset of the point, It is the hyperbolic tangent function to prevent gradient explosion. yes function, replacing hard threshold constraints and allowing gradual optimization.
5. The method for 3D instance segmentation of equipment components at the substation scene according to claim 1 is characterized in that: The process of processing data of the 3D backbone network includes: The sparse convolution-based U-Net is used to encode point cloud features and generate voting points through a set abstraction layer.
6. The method for 3D instance segmentation of equipment components at the substation scene according to claim 1 is characterized in that: The process of processing data in the spherical coordinate system includes: Point Cloud Convert to instance-centered Spherical coordinates centered on , the conversion formula is: ; in, Indicates that the Cartesian coordinate system Convert to spherical coordinates Mathematical conversion function of ; ; ; ; Where, Represent the radius, horizontal angle and vertical angle respectively, They represent the Cartesian coordinates of points in three-dimensional space.
7. The method for 3D instance segmentation of equipment components at the substation scene according to claim 1 is characterized in that: The optimization process in the task-conditional supervised distribution learning includes: Based on the reinforcement learning optimization algorithm, the supervision distribution parameters are adjusted according to the generalization error reward to obtain the optimized supervision distribution parameters.
8. A 3D instance segmentation system for substation scene equipment component level, characterized by: include: A data processing unit is used to collect data from substation equipment to obtain oblique image data, preprocess the oblique image data to obtain preprocessed point cloud data, and perform point cloud meshing on the preprocessed point cloud data; The first segmentation unit is used to input the segmented point cloud into the spherical progressive segmentation model for device instance segmentation, including: The 3D backbone network extracts features from the input point cloud and generates several voting points. A spherical coordinate system is constructed with the voting points as the center. The point cloud space is evenly divided into several sectors according to the horizontal and vertical angles. The radial distance of the farthest foreground point in each sector is predicted as the instance boundary. Optimize misclassified points through a virtual offset mechanism: learn positive radial offsets for false negative points to move them into the sector boundary, and learn negative offsets for false positive points to move them out of the boundary; Combining the point positions after virtual migration with the predicted ray boundary distance to generate an instance binary mask, and determining the device point cloud after instance segmentation based on the instance binary mask; The second segmentation unit is used to input the device point cloud after instance segmentation into the component-level segmentation model for fine-grained segmentation, including: Generate intermediate supervision features by operating the combined point cloud geometric features of the tree and the real part labels of the device point cloud after the instance segmentation; The training domain is divided into mutually exclusive subdomains, and a leave-one-out validation strategy is used to optimize the intermediate supervision features to obtain supervision distribution parameters; Sampling candidate supervisory features from the supervisory distribution parameters, gradually combining several supervisory features and evaluating their performance, and selecting the supervisory combination that achieves the greatest generalization improvement; A dynamic graph convolutional network is used as the backbone network to simultaneously optimize the component classification task and the supervision combination, and output component-level segmentation results.
9. A 3D instance segmentation device for substation scene equipment components, characterized by: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the substation scene equipment component-level 3D instance segmentation method according to any one of claims 1 to 7 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the substation scene device component-level 3D instance segmentation method according to any one of claims 1 to 7.
Citation Information
Cited By
Three-dimensional scene instance segmentation method combined with boundary perception loss
CN122089761A
Point cloud completion methods, devices, media and equipment for underground parking lot scenes
CN122415639A