A power transmission line working condition prediction method and system based on three-dimensional point cloud and multi-source data fusion

CN122548614APending Publication Date: 2026-08-11JIANGSU FRONTIER ELECTRIC TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]本发明要解决的技术问题是:克服现有输电线路状态监测中存在的部件识别精度不足、多源数据融合不充分、时序演变分析不准确以及缺乏空间关联性分析等缺陷

Benefits of technology

[0049]本发明达到的有益效果:本发明的方法通过多视图融合和注意力机制,显著提升了细小部件的识别精度,实现了点云几何特征与物理监测数据的深度跨模态融合,采用鲁棒的时序配准算法,准确提取工况演变规律,通过时空图神经网络模型,实现了考虑空间关联的准确预测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548614A_ABST
    Figure CN122548614A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data. The method includes: acquiring 3D laser point cloud data of the transmission line corridor and dividing it into multiple independent line segment units; using a multi-view feature fusion network that integrates the original point cloud view, 3D voxel view, and 2D profile projection view to identify 3D instance information of towers, conductors, and insulators; quantifying the operating condition feature parameters based on the 3D instance information and using a self-attention mechanism to perform feature-level fusion with multi-source monitoring data to generate a comprehensive feature vector; using a registration algorithm based on the maximum correlation entropy criterion to perform temporal registration of point clouds from different periods and extract operating condition evolution features; inputting the comprehensive feature vector and operating condition evolution features into a spatiotemporal graph neural network prediction model to output the future operating condition prediction result. This invention achieves accurate perception and prediction of the operating conditions of transmission lines through multi-view fusion and cross-modal attention mechanisms.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power equipment monitoring and digital twin technology, specifically relating to a method and system for predicting the operating conditions of transmission lines based on the fusion of three-dimensional laser point clouds and multi-source data. Background Technology

[0002] As a critical component of the power system, the operating status of transmission lines directly affects the safety and stability of the power grid. Due to long-term exposure to complex natural environments, transmission lines and their components are susceptible to natural factors such as wind vibration, icing, and lightning strikes, leading to potential hazards such as tower tilting, conductor sag changes, and insulator deterioration. Therefore, achieving accurate prediction of transmission line operating conditions is of great significance for preventing faults and ensuring power supply reliability.

[0003] Traditional power transmission line inspections primarily rely on manual patrols or drone aerial photography. These methods are inefficient, subjective, and struggle to obtain accurate 3D quantitative data. In recent years, lidar technology has been widely applied in power line inspections due to its ability to quickly and accurately acquire 3D point cloud data of target objects. By processing point cloud data, the 3D scene of the power transmission line corridor can be reconstructed, and the geometric parameters of key components can be extracted.

[0004] However, existing point cloud-based analysis methods still have significant limitations: First, most methods rely on point cloud data from a single view, making it difficult to simultaneously consider detailed features and global context, resulting in insufficient accuracy in identifying small components such as conductors and insulators; second, there is a lack of effective mechanisms to deeply integrate the geometric features extracted from the point cloud with the physical parameters (such as wind speed and load) monitored by the line, limiting the comprehensiveness of the operating condition assessment; furthermore, the analysis of time-series changes in transmission lines often uses simple difference comparisons, which are sensitive to noise and local deformation, making it difficult to reliably extract effective evolution patterns; finally, existing prediction models often ignore the spatial correlation between different segments of the line, resulting in prediction results that lack integrity and accuracy.

[0005] Therefore, there is an urgent need for a transmission line condition prediction method that can integrate multi-source data, achieve accurate time-series analysis, and have spatial reasoning capabilities, in order to address the shortcomings of existing technologies in component identification, feature fusion, and evolution analysis. Summary of the Invention

[0006] The technical problem to be solved by this invention is to overcome the shortcomings of existing power transmission line condition monitoring, such as insufficient component identification accuracy, inadequate fusion of multi-source data, inaccurate time-series evolution analysis, and lack of spatial correlation analysis.

[0007] To address the aforementioned technical problems, this invention provides a method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data, comprising:

[0008] S1. Obtain three-dimensional laser point cloud data of the transmission line corridor, and divide the three-dimensional laser point cloud data into multiple independent line segment units according to the spatial position of the towers.

[0009] S2. A multi-view feature fusion network is used to identify the point cloud data of each line segment unit and output three-dimensional instance information of towers, conductors and insulators. The multi-view feature fusion network is used to fuse the original point cloud view, the three-dimensional voxel view and the two-dimensional sectional projection view.

[0010] S3. Based on the three-dimensional instance information, the working condition feature parameters are quantified, and a self-attention mechanism is used to fuse the working condition feature parameters with the multi-source monitoring data of the corresponding segments at the feature level to generate a comprehensive feature vector.

[0011] S4. Use a registration algorithm based on the maximum correlation entropy criterion to perform temporal registration of point clouds of the same segment unit collected at different times and extract working condition evolution features.

[0012] S5. Input the comprehensive feature vector and the operating condition evolution features into the prediction model, and output the future operating condition prediction results.

[0013] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data includes, in step S1:

[0014] Three-dimensional laser point cloud data of the power transmission line corridor were obtained by LiDAR scanning;

[0015] Identify and locate the center three-dimensional coordinates of all towers from the point cloud data;

[0016] Using the center of the adjacent tower as a reference, a preset safety distance is extended to both sides to define the three-dimensional boundary. The point cloud data in each boundary box constitutes a line segmentation unit.

[0017] Each line segment unit is assigned a unique identifier and associated with the multi-source monitoring data of the line segment unit.

[0018] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data includes, in step S2:

[0019] The input point cloud is converted into a raw point cloud view, a 3D voxel view, and a 2D profile projection view in parallel.

[0020] PointNet++ network, 3D sparse convolutional network and U-Net structured convolutional network are used respectively to extract specific features from the original point cloud view, 3D voxel view and 2D profile projection view;

[0021] The channel attention module maps the features of the original point cloud view, the 3D voxel view and the 2D profile projection view to a unified feature space and performs adaptive weighted fusion to generate fused features.

[0022] Based on the fused features, the three-dimensional instance segmentation results and bounding box information of towers, conductors and insulators are output through the segmentation and regression head output layer.

[0023] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data includes a method for generating a 2D profile projection view, comprising:

[0024] A three-dimensional reference profile of dynamic bending is constructed based on the tower center point and the conductor sag model;

[0025] The point cloud adjacent to the reference profile is vertically projected onto the reference profile to generate a two-dimensional image. The pixel values ​​of the two-dimensional image are obtained by encoding the three-dimensional coordinates, intensity, and echo information of the projection points.

[0026] Adjust the thickness and width of the three-dimensional reference profile to generate a series of multi-scale sectional projection views.

[0027] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data includes:

[0028] The PointNet++ network employs a hybrid sampling method that combines farthest point sampling with instance-aware sampling, and integrates local feature aggregation based on vector attention mechanism in the Set Abstraction module.

[0029] The three-dimensional sparse convolutional network consists of alternating sub-manifold sparse convolutional blocks and regular sparse convolutional blocks;

[0030] The U-Net structured convolutional network employs coordinate-based skip connections between the encoder and decoder.

[0031] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data, wherein the channel attention module achieves multi-view feature fusion through the following steps:

[0032] In the cross-view interaction layer, the original features are modulated by calculating the correlation of feature channels between views, so that the features of each view are aware of the information context of other views before fusion.

[0033] The global descriptors of the modulated features of each view are concatenated and input into a shared multilayer perceptron to generate channel weight vectors for all views simultaneously.

[0034] The generated channel weight vector is multiplied with the corresponding modulated feature channel by channel for recalibration, and all recalibrated features are added element by element in three-dimensional space coordinates to obtain the final fused feature.

[0035] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data includes, in step S3:

[0036] The quantified operating condition characteristic parameters and multi-source monitoring data are projected onto a unified feature dimension to obtain the corresponding embedded representation;

[0037] The embedded representations are concatenated and input into the Transformer encoder, and cross-modal deep interaction between geometric features and physical features is performed through a self-attention mechanism;

[0038] The joint feature sequence output by the Transformer encoder is segmented according to its source. The physical feature part is then subjected to global average pooling and weighted summation with the geometric feature part to generate the comprehensive feature vector.

[0039] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data includes, in step S4:

[0040] Local point cloud blocks were extracted from the point cloud at different time periods, including the tower body, insulator string connection points, and conductor suspension points.

[0041] Coarse registration is performed using the tower body point cloud block as a rigid reference, and then non-rigid fine registration is performed between the insulator string and the conductor hanging point point cloud block.

[0042] Based on the non-rigid fine registration results, rigid displacement characteristics, local deformation characteristics, and dynamic response characteristics are calculated, and the rigid displacement characteristics, local deformation characteristics, and dynamic response characteristics are aggregated into a time-series evolution characteristic vector.

[0043] The aforementioned method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data, in step S5, the prediction model is a spatiotemporal graph neural network, and the prediction process of the spatiotemporal graph neural network includes:

[0044] Each tower is treated as a graph node, and the node features are the comprehensive feature vectors. A topology graph is constructed based on the electrical and spatial relationships to form spatial features.

[0045] Graph convolutional networks are used to aggregate information about neighboring nodes to capture spatial dependencies, and temporal convolutional networks or gated recurrent units are used to learn the temporal dynamics of nodes themselves to form temporal features.

[0046] By integrating spatial and temporal features, a fully connected layer is used to output the probability distribution of future multi-time point operating condition safety levels and predicted values ​​of key parameters. A computer system includes a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the transmission line operating condition prediction method based on the fusion of 3D point cloud and multi-source data, as described above.

[0047] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the transmission line condition prediction method based on the fusion of three-dimensional point cloud and multi-source data as described above.

[0048] A computer program product includes a computer program / instructions that, when executed by a processor, implement the steps of the transmission line condition prediction method based on the fusion of three-dimensional point cloud and multi-source data as described above.

[0049] The beneficial effects achieved by this invention are as follows: The method of this invention significantly improves the recognition accuracy of small parts through multi-view fusion and attention mechanism, realizes deep cross-modal fusion of point cloud geometric features and physical monitoring data, adopts a robust temporal registration algorithm to accurately extract the working condition evolution law, and realizes accurate prediction considering spatial correlation through spatiotemporal graph neural network model. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of a method for predicting the operating conditions of transmission lines based on the fusion of three-dimensional point clouds and multi-source data, provided in Embodiment 1 of the present invention.

[0051] Figure 2 This is a schematic diagram of the processing flow of the multi-view feature fusion network provided in Embodiment 1 of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0053] It should be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than two" refers to two or more. Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0054] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] Example 1

[0056] like Figure 1 As shown, this embodiment provides a method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data, including:

[0057] S1. Acquire three-dimensional laser point cloud data of the transmission line corridor, and divide the three-dimensional laser point cloud data into multiple independent line segment units according to the spatial location of the towers. Each line segment unit includes the point cloud of adjacent towers and the conductors, insulators and corridor environment between adjacent towers.

[0058] The line segment unit is the basic processing unit of the analysis method. A transmission line corridor of tens or even hundreds of kilometers is divided into independently managed digital segments according to the physical structure of the towers, facilitating analysis. For example, on a line, all the conductors, insulators, and other equipment and the surrounding environment located between tower #25 and tower #26 together constitute a line segment unit.

[0059] In this embodiment, the data acquisition and segmentation steps in step S1 specifically include:

[0060] S11 uses airborne or ground-based lidar systems to scan the transmission line corridor and acquire three-dimensional laser point cloud data including towers, conductors, insulators, and surrounding features.

[0061] LiDAR can efficiently and accurately acquire the three-dimensional coordinates of millions or even hundreds of millions of points along power transmission line corridors. Information such as reflection intensity is used to generate realistic 3D laser point cloud data.

[0062] S12, using a clustering recognition algorithm to identify and locate the center three-dimensional coordinates of all towers from the three-dimensional laser point cloud data.

[0063] The three-dimensional coordinates of the tower's center are the key spatial reference point and the basis for segmentation. The tower is the most stable and prominent structure in a transmission line, making segmentation centered on the tower the most reasonable approach. Using point cloud density and geometric features as input, a clustering recognition algorithm automatically calculates the accurate position of each tower in three-dimensional space. ).

[0064] S13, using the line connecting the center points of two adjacent towers as the central axis, expand outwards by a preset safety distance to generate a three-dimensional bounding box. This three-dimensional bounding box covers all equipment and passageways extending from the current tower to the next, and is defined as an independent line segmentation unit. The preset safety distance is a buffer distance set according to safety regulations and experience. The three-dimensional bounding box is a cuboid region defined in three-dimensional space. It transforms the abstract concept of segmentation into a concrete, calculable three-dimensional spatial region. Expanding from the tower as the center, it completely includes the tower itself and all conductors, insulators, hardware, and passageways that may endanger line safety between adjacent towers.

[0065] S14. Establish a local coordinate system for each line segment unit and assign a unique segment identifier to associate it with the multi-source physical monitoring data of the corresponding line segment unit, thereby digitizing the segment unit. This identifier is named Line1_Segment_25-26 for storage, querying, and tracking. Simultaneously, by establishing this association, features extracted from the point cloud of this segment unit in subsequent steps can be accurately matched and fused with the physical monitoring data of the same segment unit. This physical monitoring data includes micro-weather station data installed on poles and line load data.

[0066] S2. The original point cloud view, three-dimensional voxel view and two-dimensional profile projection view of each line segment unit are fused using a multi-view feature fusion network to output three-dimensional instance information of towers, conductors and insulators.

[0067] An adaptive weighted fusion of multi-view features using a channel attention mechanism is employed to identify and extract 3D instance information of towers, conductors, and insulators. The multi-view feature fusion network is a deep learning model capable of simultaneously analyzing point cloud data from three different views at the same angle. The output of the multi-view feature fusion network is 3D instance information, which is structured and computable data. This 3D instance information includes at least semantic labels, instance identifiers, and their 3D oriented bounding boxes for each component.

[0068] Figure 2 This is a schematic diagram of the processing flow of the multi-view feature fusion network provided in Embodiment 1 of the present invention. In step S2, the processing flow of the multi-view feature fusion network includes:

[0069] S21. The input line segment unit point cloud data is converted into three view representations in parallel, including:

[0070] The original point cloud view retains all original points and features;

[0071] A 3D voxel view that divides the 3D space into a regular grid and encodes the point cloud distribution;

[0072] And a two-dimensional cross-sectional projection view generated by projecting along the conductor's direction or the tower's central axis.

[0073] The original point cloud view directly uses the set of three-dimensional coordinates of the points, preserving the most original and accurate geometric information. The three-dimensional voxel view divides the three-dimensional space into a regular, tiny cubic grid, converting the point cloud data into a regular, three-dimensional pixel structure. The two-dimensional cross-sectional projection view projects the three-dimensional point cloud data onto a two-dimensional plane, generating a two-dimensional image; the two-dimensional plane is a vertical cross-section along the direction of the guide wire. Because point cloud data itself is disordered and sparse, direct processing is difficult. Converting point cloud data into multiple regularized representations allows for the extraction of complementary features using different neural networks.

[0074] In step S21, the specific method for generating the two-dimensional cross-sectional projection view includes:

[0075] S211. Based on the identified tower center point and conductor sag curve model, construct a dynamically curved cylindrical or fan-shaped three-dimensional reference profile along the line's direction. The conductor sag curve model is a mathematical model describing the natural sag of the conductor under its own weight and tension, such as the catenary equation. The dynamically curved three-dimensional reference profile is a curved surface in three-dimensional space that closely matches the actual direction of the conductor. Existing methods use fixed-plane projection, which causes the curved conductor to be compressed and deformed during projection, resulting in feature distortion. The curved surface model created in this step closely follows the natural direction of the line, ensuring that the projected image retains the true shape of the conductor and greatly reducing geometric distortion.

[0076] S212. The point cloud data in the vicinity of the three-dimensional reference profile is vertically projected onto the three-dimensional reference profile to generate a two-dimensional image. The pixel values ​​of the two-dimensional image are obtained by multi-dimensional encoding of the projected points. The encoded information of the two-dimensional image includes at least the original three-dimensional coordinates of the projected points, the vertical distance relative to the two-dimensional reference profile, the point cloud intensity, and the echo count. This step not only completes the conversion from 3D to 2D to utilize an efficient 2D convolutional network, but also preserves and encodes the three-dimensional attributes into the pixels of the two-dimensional image. The three-dimensional coordinates preserve the spatial location information of the points. The reflection intensity reflects the material characteristics of the object's surface, such as the different reflection intensities of ceramic insulators and metal wires. The echo information helps distinguish between foreground and occlusion objects, such as a wire behind leaves. Each pixel is not a single numerical value, but a feature vector containing multiple original attributes of the projected point.

[0077] S213. By adjusting the thickness and width of the three-dimensional reference profile, a series of multi-scale profile projection views with different spatial containment are generated to capture the features of the conductor body, the insulator string structure, and large obstacles within the channel, respectively. These multi-scale profile projection views are used to achieve cross-scale feature capture. By generating profiles of different thicknesses, the network model has the ability to simultaneously observe details and the global picture. Thin profiles can accurately capture the fine contours and surface textures of components such as conductors and insulators. Thick profiles can include environmental objects within a certain range around the conductor in the projection range, thereby capturing the channel's safety relationship features. This step enables the network model to possess both perspectives simultaneously, solving the geometric distortion problem in existing projections and maximizing the preservation of the depth and breadth of three-dimensional information. It provides a high-quality, lossless input source for subsequent network models such as U-Net, and is a key step in improving the recognition accuracy of small targets such as insulators.

[0078] S22. A PointNet++ network model is used as the backbone to perform deep local feature extraction on the original point cloud view. A 3D sparse convolutional network is used to perform regularized feature extraction on the 3D voxel view to obtain spatial context information. A U-Net-structured convolutional network model is used to extract pixel-level semantic features from the 2D profile projection view. For each network best suited for processing the corresponding view data, the unique information under that view is maximized. PointNet++ learns fine local features of points from the original points, 3D sparse convolution learns the shape and structural features of 3D space from the voxel mesh, and U-Net learns 2D texture and contour features from the projection image.

[0079] In step S22, when extracting features from the original point cloud view using the PointNet++ network, a hybrid sampling method combining farthest point sampling and instance-aware sampling is adopted. The instance-aware mechanism is embedded in the selection logic of the farthest point sampling, including:

[0080] First, the PointNet++ network predicts an instance perception score for each point. The instance perception score is the confidence score of the key components of the corresponding point. The key components include insulators, vibration dampers, and conductors.

[0081] Secondly, based on spatial distance and using instance-aware scores as weighting coefficients, the selection priority of candidate points is calculated. A set proportion of points are retained according to their selection priority from high to low, completing the iterative selection of the farthest point sampling. Points with high selection priority scores, even if not spatially farthest, still have a higher retention probability; background points with low selection priority scores, even if spatially far, may be discarded first. The final output set of sampling points satisfies two conditions: first, the spatial distribution is relatively uniform, preserving the global coverage characteristic of farthest point sampling; second, key component points are densely retained, achieving the semantic enhancement characteristics of instance-awareness, effectively solving the problem of insufficient sampling of small targets in traditional farthest point sampling.

[0082] The spatial distance refers to the shortest Euclidean distance between the current candidate point and the set of selected points. During the iterative sampling process, the shortest distance from all unselected points to the current set of selected points is calculated each time to ensure spatial coverage of the sampling points.

[0083] Priority of selection The calculation formula is as follows:

[0084] in: The shortest spatial distance between the candidate point and the set of selected points is normalized to [0,1]. Let be the instance perception score of the candidate point, and be the confidence score of the key component, ranging from [0,1]. λ is a balancing factor, 0≤λ≤1, used to adjust the weights of spatial uniformity and semantic importance. In each point selection process, both spatial distribution uniformity and the retention requirements of key components are considered, thereby achieving the goal of hybrid sampling.

[0085] Farthest-point sampling is a method that ensures the uniform spatial distribution of sampling points, but it is prone to missing points of small targets (such as insulators); instance-aware sampling prioritizes important points based on the confidence level of each point belonging to a key component. This invention combines the two, considering both spatial distance and instance score in each sampling, ensuring uniform coverage of the overall point cloud while densely retaining key component points, thereby improving the ability to identify small targets.

[0086] Furthermore, the Set Abstraction module integrates local feature aggregation based on vector attention. This Set Abstraction module is a core unit of the PointNet++ network, used for hierarchical downsampling and local feature extraction of point clouds, achieving a progressive abstraction from raw geometric coordinates to higher-order semantic features.

[0087] The Set Abstraction module comprises three sub-layers: A sampling layer: Selects keypoints from the input point cloud as local region centers. This invention replaces the original farthest point sampling with the aforementioned hybrid sampling method to enhance small target perception; A grouping layer: Uses each keypoint as a center and searches for neighboring points within a radius to form a local point set; A PointNet layer: Uses a small PointNet to extract and aggregate features from the local point set. This invention integrates local feature aggregation based on a vector attention mechanism in this layer to replace the original max pooling.

[0088] This invention utilizes a vector attention mechanism in the PointNet layer of the Set Abstraction module to achieve weighted aggregation of local features. This includes generating an attention weight vector equal to the number of feature channels for each point in the local neighborhood, instead of a single scalar weight; and achieving fine-grained feature recalibration by performing channel-wise weighted summation on the local point features. This method enables the network to adaptively focus on discriminative structures within local regions, such as the edge contour of an insulator string, the lowest point of the conductor's sag, and the installation position of a vibration damper. The original PointNet layer of the Set Abstraction module uses max pooling, retaining only single-point extrema, which easily loses local morphological details. The vector attention aggregation method in this invention significantly enhances the ability to extract continuous deformation features and structural features of small-sized components, providing more accurate geometric input for subsequent quantification of working condition features. This enhances the feature perception capability of small-sized key components such as insulators and vibration dampers.

[0089] When using a 3D sparse convolutional network to extract features from a 3D voxel view, the main body of the 3D sparse convolutional network consists of alternating sub-manifold sparse convolutional blocks and regular sparse convolutional blocks. The sub-manifold convolutional blocks are used to maintain feature resolution on non-empty voxels and extract fine geometric features, including the thickness of the tower angle steel, thus preserving the resolution of the feature map. The regular sparse convolutional blocks are used for downsampling to expand the receptive field and capture spatial contextual information, including the overall structure of the tower and the relationship between the tower and the conductor.

[0090] A voxel, short for volumetric pixel, is the smallest unit of a regular grid in three-dimensional space, similar to a pixel in a two-dimensional image. Dividing three-dimensional space into a grid of equal-sized cubes results in a voxel for each small cube. For example, dividing a 10m x 10m x 10m space into a 1m x 1m x 1m grid yields 1000 voxels. In point cloud processing, a non-empty voxel is a voxel containing at least one laser point cloud; an empty voxel is a voxel containing no points. In 3D sparse convolution, computation is performed only on non-empty voxels, avoiding invalid computations on empty voxels and significantly improving computational efficiency. Submanifold convolution blocks maintain feature resolution on non-empty voxels, meaning no downsampling is performed; information is only transferred between voxels containing data; and fine geometric features such as the thickness of tower angle steel and the surface texture of insulators can be extracted.

[0091] When using a convolutional network with a U-Net structure to extract features from a 2D sectional projection view, a coordinate-based skip connection is employed between the encoder and decoder. This skip connection concatenates the 3D spatial coordinates of pixels in the projection image with the upsampled features from the decoder, strengthening the correlation between 2D features and 3D spatial positions. The coordinate-based skip connection concatenates coordinate information, addressing the problem of lost spatial information in the projection view. Standard U-Net skip connections only transmit feature maps, but 2D projection itself loses depth information. This method, while transmitting features, also transmits the original 3D spatial coordinates (x, y, z) of each pixel in the image to the decoder, solving problems related to small target preservation, local feature extraction, efficiency and context balance, and 2D-3D spatial alignment, thus improving the accuracy of the model's component recognition.

[0092] S23. Features from different views are transformed into a unified feature space through a preset mapping relationship, and then input into the channel attention module. The channel attention module generates weight vectors for each view feature channel through global average pooling and a multilayer perceptron, and performs weighted summation on the original features based on the weight vectors to achieve adaptive fusion of multi-view features. The original features are the output features from the previous stage cross-view interaction layer that will be processed in the current weighted summation step. This step addresses the problem of how to combine features from different views more effectively, rather than simply concatenating or adding features. Instead, the network model determines which features provided by which views are more important in the current task.

[0093] In step S23, the channel attention module achieves adaptive fusion of multi-view features through the following steps:

[0094] S231. Before inputting the features of each view into the weight generation network, information preheating is performed using a cross-view interaction layer, including: denoting the feature maps of any two different views as features respectively. Figure 1 and characteristics Figure 2 First, the features Figure 1 Reshape into a two-dimensional matrix and calculate features. Figure 1 With features Figure 2 Matrix multiplication by transpose yields the correlation matrix. Correlation matrix Element representation view Channels and Views The correlation strength of the channels; then the Softmax function is used to refine the correlation matrix. Row-by-row normalization is performed to obtain the normalized matrix, and the normalized matrix is ​​then compared with the feature matrix. Figure 2 Multiply to obtain the feature Figure 1 Features of context modulation Figure 2 Enhanced features. The above process is performed symmetrically across all view pairs, and the enhanced features are compared with the original features. The learnable weights are added together to complete the initial modulation.

[0095] S232. Perform global average pooling on the features of each view after interactive modulation to obtain a global descriptor, and concatenate the obtained global descriptors along the channel dimension to form a unified context descriptor; input the descriptor into the weight generation network to generate a set of channel weight vectors corresponding to all input views.

[0096] The weight generation network employs a shared multilayer perceptron with a bottleneck structure: the number of neurons in the input layer equals the total dimension of the global descriptors of all views concatenated; the first fully connected layer reduces the dimension to 1 / 4 of the total input dimension, followed by a ReLU activation function; the second fully connected layer restores the dimension to a specific value, which is equal to the number of views multiplied by the number of feature channels in each view; finally, the output is reshaped into a channel weight vector with the corresponding number of channels through a Sigmoid activation function.

[0097] S233. Multiply the set of weight vectors generated in step S232 with the corresponding original view features after interactive modulation channel by channel to complete feature recalibration; finally, add all the recalibrated view features element by element on the corresponding three-dimensional spatial coordinates to obtain the final fused features.

[0098] The above steps employ a three-stage fusion mechanism of interaction, collaboration, and consistency. Information preheating is achieved through cross-view interaction, collaborative decision-making is realized through shared weight generation, and geometrically consistent fusion is achieved through spatial alignment and summation. This surpasses simple weighted summation, enabling multi-view fusion to achieve collaborative and accurate information integration.

[0099] S24. Based on the final fusion features, the three-dimensional instance segmentation results and bounding box information of towers, conductors and insulators are output using the segmentation output layer and the regression head output layer.

[0100] The segmentation output layer and the regression head output layer are two parallel output layers at the end of the network. The segmentation output layer is responsible for classifying each point and completing the segmentation, while the regression head output layer is responsible for calculating the bounding box. The 3D instance segmentation result can distinguish the category of each point, such as a conductor, and can also distinguish different individuals, such as a first-phase conductor and a second-phase conductor. The bounding box information is a 3D cuboid containing the target position, size, and orientation, generating the final quantifiable information and decoding the fused abstract features into specific, structured data that can be used for engineering analysis. The segmentation result includes the component name and the affiliation of each point. The bounding box information includes the specific location and extent of each component in space, with specific parameters including the center point (x, y, z), length, width, height, and rotation angle.

[0101] This step utilizes different data representations and attention mechanisms to achieve high-precision, instance-level identification and localization of key components in complex scenes.

[0102] S3. Based on the three-dimensional instance information, the working condition feature parameters are quantified, and a self-attention mechanism is used to fuse the working condition feature parameters with the multi-source monitoring data of the corresponding segments at the feature level to generate a comprehensive feature vector.

[0103] The operating condition characteristic parameters are specific numerical indicators calculated from the three-dimensional instance information in step S2. Multi-source monitoring data comes from other sensors, including meteorological monitoring data such as wind speed and temperature, and line load current. The geometric shape obtained from the point cloud data is fused with the physical state obtained from the monitoring sensors to provide a more comprehensive assessment of the operating conditions.

[0104] In this embodiment, the process of feature-level fusion using a self-attention mechanism includes:

[0105] S31. Project the quantized operating condition feature parameters and multi-source monitoring data onto a unified feature dimension to obtain the corresponding embedding representations. Specifically, denote the set of operating condition feature parameters quantized based on 3D instance information as the geometric feature vector G, and denote the corresponding segmented multi-source monitoring data as the physical feature vector P; map the geometric feature vector G and the physical feature vector P onto a unified feature dimension through linear projection layers respectively to obtain the geometric feature embedding representation. and physical feature embedding representation 。 The geometric parameters of point cloud quantization (such as sag and deflection) and the physical parameters of multi-source monitoring (such as temperature, wind speed, and load) are completely different in numerical range, physical meaning, and dimensionality, and cannot be directly compared or fused. This step generates a feature vector of uniform dimension through linear projection, which can be operated on in the same vector space.

[0106] S32. The embedded representations are concatenated and then input into the Transformer encoder to achieve cross-modal deep interaction between geometric features and physical features through a self-attention mechanism.

[0107] Specifically, the geometric features are embedded in the representation. and physical feature embedding representation The sequence is concatenated along the sequence dimension to form a joint feature sequence. This joint feature sequence is then fed into a multi-layer Transformer encoder. Inside the Transformer encoder, the core self-attention mechanism calculates the association weights between any two elements in the geometric and physical feature sequences, achieving cross-modal deep information interaction and context modeling. This step represents the complex and deep correlation between geometric and physical parameters, not a simple concatenation or weighting, but rather the discovery of complex patterns using the Transformer encoder's self-attention mechanism, such as the normal sag range when the wind speed reaches a certain value.

[0108] S33. The joint feature sequence output by the encoder is segmented according to its source. The physical feature part is then subjected to global average pooling and weighted summation with the geometric feature part to generate the comprehensive feature vector.

[0109] Specifically, the joint feature sequence output by the Transformer encoder is segmented according to its source; corresponding to the physical feature embedding representation. A portion of the output sequence undergoes global average pooling to obtain an enhanced physical feature vector with global context; finally, the enhanced physical feature vector is coupled with the geometric feature embedding representation. The output sequences are weighted and summed to generate the final comprehensive feature vector used for prediction.

[0110] This step enables information distillation and focusing, generating the final fused features used for prediction. Physical monitoring data (such as meteorological and load data) typically describe the global state of the entire segmented unit. Pooling operations refine the enhanced information from the global state through deep interaction into a compact global context vector. Geometric features from point clouds are inherently spatially distributed. Weighted summing of the global context vector with a sequence of geometric features that retains spatial information is equivalent to modulating the geometric features of each spatial location using the physical global context. The final composite feature vector primarily contains the geometric features of conductors and insulators with spatial details, but each detailed feature has been enhanced by a global physical context including current wind speed (X), load (Y), and temperature (Z). This allows the prediction model to determine, when judging excessive sag, whether it is a normal phenomenon caused by high-temperature thermal expansion or an abnormal signal that may indicate mechanical failure, by combining the physical context.

[0111] The above steps utilize an advanced cross-modal fusion process involving alignment, deep interaction, and context modulation to uncover the deep causal relationships hidden between geometric shapes and physical states using the Transformer model, ultimately generating a comprehensive feature vector deeply modulated by physical laws.

[0112] S4. Use a registration algorithm based on the maximum correlation entropy criterion to perform time-series registration of point cloud data of the same segment unit collected at different times, and extract the working condition evolution characteristics.

[0113] The maximum correlation entropy criterion is a mathematical measure that is insensitive to noise and outliers, making it more robust than the traditional least squares method. Through robust registration algorithms, it accurately detects minute changes in components over time, such as the slow tilting of towers or the gradual deflection of insulator strings, and quantifies these changes to form evolution characteristics. For example, comparing point clouds scanned from the same tower in January and July, MCC (Maximizing Correlation Coefficient) registration reveals a 2-centimeter shift in the X-direction at the tower top; this 2-centimeter shift represents the evolution characteristic of the operating conditions.

[0114] Specifically, it includes:

[0115] S41. Extract local high-precision point cloud blocks from the point cloud of different periods, including the tower body, insulator string connection points, and conductor suspension points. This is not a registration of the entire line segment unit point cloud, but rather the extraction of local high-precision point cloud blocks from the point cloud of different periods for the tower body, insulator string connection points, and conductor suspension points. Using a registration algorithm based on the maximum correlation entropy criterion, an initial overall coarse registration is performed using the tower body point cloud blocks as a rigid reference. Then, non-rigid fine registration is performed on the point cloud blocks for the insulator string connection points and conductor suspension points respectively.

[0116] S42. On the corresponding point cloud blocks after registration, calculate multi-level working condition evolution characteristics, including rigid displacement characteristics, local deformation characteristics, and dynamic response characteristics. The rigid displacement characteristics are calculated based on the registration results of the tower body point cloud to determine the overall settlement, tilt, and horizontal displacement of the tower base. The local deformation characteristics are calculated based on the registration results of the insulator string connection points to determine the change in the tilt angle and the three-dimensional spatial vector offset of the insulator string. The dynamic response characteristics are calculated based on the registration results of the conductor hanging points to determine the relative change in the conductor sag, and combined with the concurrent wind speed monitoring data, to determine the trend of the dynamic swing amplitude change of the conductor under wind load.

[0117] S43. The rigid displacement characteristics, local deformation characteristics and dynamic response characteristics are aggregated to form a comprehensive and quantitative time-series operating condition evolution feature vector, which is used to characterize the overall operating condition change process of the corresponding line segment unit from macro to micro and from static to dynamic.

[0118] The above steps, through a time-series analysis process involving component separation, hierarchical registration, and multi-feature quantification, extract stable, accurate, and clearly interpretable evolutionary patterns from the coarse original point cloud changes through refined operations, providing the most critical time-series input for the final accurate prediction.

[0119] S5. Input the comprehensive feature vector and the operating condition evolution feature vector into the prediction model, and output the probability distribution of operating condition safety level and the predicted value of key parameters at multiple future time points.

[0120] Using the current overall status provided in step S3 and the historical evolution patterns provided in step S4, predict future development trends.

[0121] In this embodiment, in step S5, the prediction model employs a spatiotemporal graph neural network, and the workflow includes:

[0122] S51. In the spatial module, each tower is treated as a graph node. The node characteristics of the graph node are the comprehensive feature vectors of the tower and adjacent segment units. A topology graph is constructed based on the electrical and spatial relationships to form spatial features.

[0123] Each tower is defined as a graph node, and the node's features are the comprehensive feature vector. Edges between nodes are constructed based on the electrical connections and spatial proximity relationships of the transmission line, thus forming a line-level topology graph. The features of each node at different time steps constitute a dynamic node feature sequence. A graph node represents an entity in the graph data structure; each tower is a node. The topology graph describes the connections between nodes. Edges are constructed based on electrical connections (e.g., passing through the same substation) and spatial proximity relationships (e.g., adjacent towers). This step abstracts the physical power grid structure into a computer-processable mathematical graph model, no longer treating each segment unit as an isolated individual, but placing each segment unit within the context of the entire line network. For example, a line containing towers #24, #25, and #26 can construct a graph with three nodes, with an edge established between nodes Seg_24-25 and Seg_25-26 because they share tower #25.

[0124] S52. In the time module, graph convolutional networks are used to aggregate information of adjacent nodes to capture spatial dependencies, and temporal convolutional networks or gated recurrent units are used to learn the temporal dynamics of the nodes themselves to form temporal features.

[0125] The topology graph is processed using a graph convolutional network (GCN). By aggregating the operating condition information of adjacent segmented units through message passing between nodes, the propagation and correlation characteristics of faults or potential hazards in space are captured. The GCN is a neural network that operates on a graph structure, updating the features of the central node by aggregating information from neighboring nodes. Each node's dynamic feature sequence is input into a temporal convolutional network or a gated recurrent unit to learn the long-term patterns and short-term fluctuations of its operating condition over time; the dynamic feature sequence includes comprehensive features and operating condition evolution features.

[0126] S53 integrates spatial and temporal features, and outputs the probability distribution of future operating condition safety levels and predicted values ​​of key parameters at multiple time points through a fully connected layer.

[0127] By integrating the outputs of the spatial and temporal modules, and passing through a fully connected prediction layer, the system simultaneously outputs the probability distribution of the operating condition safety level and the predicted values ​​of key operating condition parameters at multiple future time points, thus completing a risk assessment of the multi-step evolution from the current state to the future. The key operating condition parameters include sag and offset.

[0128] This embodiment presents a method for predicting the operating conditions of transmission lines based on the fusion of 3D point clouds and multi-source data, enabling static analysis and dynamic prediction of the operating conditions of transmission lines.

[0129] Example 2

[0130] A computer system includes a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the method as described in Embodiment 1.

[0131] Example 3

[0132] A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that the computer program / instructions, when executed by a processor, implement the steps of the method as described in Embodiment 1.

[0133] Example 4

[0134] A computer program product includes a computer program / instructions, characterized in that, when the computer program / instructions are executed by a processor, they implement the steps of the method as described in Example 1.

[0135] It should be noted that a portion of the electronic device described above can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.

[0136] It should be noted that the computer mentioned here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, computer-readable recording media refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage systems such as hard drives built into the computer.

[0137] Furthermore, computer-readable recording media can include: media that dynamically stores programs for short periods of time, such as communication lines used when transmitting programs via networks like the Internet or communication lines like telephone lines; and media that store programs for fixed periods of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining them with programs already recorded in the computer.

[0138] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (system group) composed of multiple systems. Each system constituting the system group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a system group, it is sufficient to have all the functions or functional blocks of the electronic device.

[0139] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.

Claims

1. A method for predicting the operating conditions of transmission lines based on the fusion of 3D point cloud and multi-source data, characterized in that, Includes the following steps: S1. Obtain three-dimensional laser point cloud data of the transmission line corridor, and divide the three-dimensional laser point cloud data into multiple independent line segment units according to the spatial position of the towers. S2. A multi-view feature fusion network is used to identify the point cloud data of each line segment unit and output three-dimensional instance information of towers, conductors and insulators. The multi-view feature fusion network is used to fuse the original point cloud view, the three-dimensional voxel view and the two-dimensional sectional projection view. S3. Based on the three-dimensional instance information, the working condition feature parameters are quantified, and a self-attention mechanism is used to fuse the working condition feature parameters with the multi-source monitoring data of the corresponding segment at the feature level to generate a comprehensive feature vector. S4. Use a registration algorithm based on the maximum correlation entropy criterion to perform temporal registration of point clouds of the same segment unit collected at different times and extract working condition evolution features. S5. Input the comprehensive feature vector and the operating condition evolution features into the prediction model, and output the future operating condition prediction results. 2.The power line working condition prediction method based on three-dimensional point cloud and multi-source data fusion according to claim 1, characterized in that, Step S1 includes: Three-dimensional laser point cloud data of the power transmission line corridor were obtained by LiDAR scanning; Identify and locate the center three-dimensional coordinates of all towers from the point cloud data; Using the center of the adjacent tower as a reference, a preset safety distance is extended to both sides to define the three-dimensional boundary. The point cloud data in each boundary box constitutes a line segmentation unit. Each line segment unit is assigned a unique identifier and associated with the multi-source monitoring data of the line segment unit.

3. The method for predicting the operating conditions of transmission lines based on the fusion of three-dimensional point clouds and multi-source data according to claim 1, characterized in that, Step S2 includes: The input point cloud is converted into a raw point cloud view, a 3D voxel view, and a 2D profile projection view in parallel. PointNet++ network, 3D sparse convolutional network and U-Net structured convolutional network are used respectively to extract specific features from the original point cloud view, 3D voxel view and 2D profile projection view; The channel attention module maps the features of the original point cloud view, the 3D voxel view and the 2D profile projection view to a unified feature space and performs adaptive weighted fusion to generate fused features. Based on the fused features, the three-dimensional instance segmentation results and bounding box information of towers, conductors and insulators are output through the segmentation and regression head output layer.

4. The power line working condition prediction method based on three-dimensional point cloud and multi-source data fusion according to claim 3, characterized in that, The method for generating the two-dimensional sectional projection view includes: A three-dimensional reference profile of dynamic bending is constructed based on the tower center point and the conductor sag model; The point cloud adjacent to the reference profile is vertically projected onto the reference profile to generate a two-dimensional image. The pixel values ​​of the two-dimensional image are obtained by encoding the three-dimensional coordinates, intensity, and echo information of the projection points. Adjust the thickness and width of the three-dimensional reference profile to generate a series of multi-scale sectional projection views.

5. The power line working condition prediction method based on three-dimensional point cloud and multi-source data fusion according to claim 3, characterized in that, include: The PointNet++ network employs a hybrid sampling method that combines farthest point sampling with instance-aware sampling, and integrates local feature aggregation based on vector attention mechanism in the Set Abstraction module. The three-dimensional sparse convolutional network consists of alternating sub-manifold sparse convolutional blocks and regular sparse convolutional blocks; The U-Net structured convolutional network employs coordinate-based skip connections between the encoder and decoder.

6. The power line working condition prediction method based on three-dimensional point cloud and multi-source data fusion according to claim 3, characterized in that, The channel attention module achieves multi-view feature fusion through the following steps: In the cross-view interaction layer, the original features are modulated by calculating the correlation of feature channels between views, so that the features of each view are aware of the information context of other views before fusion. The global descriptors of the modulated features of each view are concatenated and input into a shared multilayer perceptron to generate channel weight vectors for all views simultaneously. The generated channel weight vector is multiplied with the corresponding modulated feature channel by channel for recalibration, and all recalibrated features are added element by element in three-dimensional space coordinates to obtain the final fused feature.

7. The power line working condition prediction method based on fusion of three-dimensional point cloud and multi-source data according to claim 1, characterized in that, Step S3 includes: The quantified operating condition characteristic parameters and multi-source monitoring data are projected onto a unified feature dimension to obtain the corresponding embedded representation; The embedded representations are concatenated and input into the Transformer encoder, and cross-modal deep interaction between geometric features and physical features is performed through a self-attention mechanism; The joint feature sequence output by the Transformer encoder is segmented according to its source. The physical feature part is then subjected to global average pooling and weighted summation with the geometric feature part to generate the comprehensive feature vector. 8.The power line working condition prediction method based on fusion of three-dimensional point cloud and multi-source data according to claim 1, characterized in that, Step S4 includes: Local point cloud blocks were extracted from the point cloud at different time periods, including the tower body, insulator string connection points, and conductor suspension points. Coarse registration is performed using the tower body point cloud block as a rigid reference, and then non-rigid fine registration is performed between the insulator string and the conductor hanging point point cloud block. Based on the non-rigid fine registration results, rigid displacement characteristics, local deformation characteristics, and dynamic response characteristics are calculated, and the rigid displacement characteristics, local deformation characteristics, and dynamic response characteristics are aggregated into a time-series evolution characteristic vector.

9. The power line working condition prediction method based on fusion of three-dimensional point cloud and multi-source data according to claim 1, characterized in that, In step S5, the prediction model is a spatiotemporal graph neural network, and the prediction process of the spatiotemporal graph neural network includes: Each tower is treated as a graph node, and the node features are the comprehensive feature vectors. A topology graph is constructed based on the electrical and spatial relationships to form spatial features. Graph convolutional networks are used to aggregate information about neighboring nodes to capture spatial dependencies, and temporal convolutional networks or gated recurrent units are used to learn the temporal dynamics of nodes themselves to form temporal features. By integrating spatial and temporal characteristics, the system outputs the probability distribution of safety levels for future operating conditions and predicted values ​​of key parameters at multiple time points through a fully connected layer.

10. A computer system comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the transmission line condition prediction method based on the fusion of three-dimensional point cloud and multi-source data as described in any one of claims 1-9.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the transmission line condition prediction method based on the fusion of three-dimensional point cloud and multi-source data as described in any one of claims 1-9.

12. A computer program product, comprising a computer program, characterized in that, When the computer program / instruction is executed by the processor, it implements the steps of the transmission line condition prediction method based on the fusion of three-dimensional point cloud and multi-source data as described in any one of claims 1-9.