Method for monitoring galloping of high-voltage transmission line based on binocular vision and deep learning

By combining binocular vision and deep learning methods with a dual-branch U-net network and triangulation, the problems of high sensor cost and difficulty in obtaining three-dimensional displacement in high-voltage transmission line galloping monitoring are solved, achieving high-precision galloping displacement monitoring and trend prediction.

CN122435535APending Publication Date: 2026-07-21STATE GRID JIBEI ELECTRIC POWER CO LTD TANGSHAN POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID JIBEI ELECTRIC POWER CO LTD TANGSHAN POWER SUPPLY CO
Filing Date
2026-04-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies for monitoring high-voltage transmission line galloping suffer from problems such as high sensor cost and low accuracy. Video image-based methods cannot accurately acquire three-dimensional displacement and amplitude, while monocular vision methods suffer from scale uncertainty and poor engineering adaptability.

Method used

A method based on binocular vision and deep learning is adopted. Video is acquired by binocular cameras, and a dual-branch U-net network is constructed for transmission line segmentation. Combined with a cross-view attention fusion module and triangulation method, the three-dimensional galloping displacement of the transmission line is obtained, and the galloping displacement is predicted by a time-frequency domain dual-branch physical parameter regression model.

Benefits of technology

It achieves high-precision monitoring of transmission line galloping displacement, accurately acquires three-dimensional displacement and amplitude, reduces monitoring costs, improves model stability and engineering adaptability, and has the ability to predict short-term trends.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435535A_ABST
    Figure CN122435535A_ABST
Patent Text Reader

Abstract

The application discloses a high-voltage transmission line galloping monitoring method based on binocular vision and deep learning. First, a binocular camera is used to collect transmission line galloping videos, key frame images are extracted, and a binocular transmission line image dataset is constructed; then, a transmission line segmentation model is constructed based on cross-view attention fusion and a double-branch U-Net network, and the transmission line segmentation model is trained by using the binocular transmission line image dataset; the trained transmission line segmentation model is used to segment left and right transmission line images to obtain left and right transmission line semantic segmentation images; the left and right transmission line semantic segmentation images are processed to obtain transmission line galloping displacement; finally, a galloping displacement prediction model is constructed and trained; and a historical transmission line galloping displacement sequence is input into the trained galloping displacement prediction model to obtain a transmission line galloping displacement prediction sequence. The method improves the precision of transmission line segmentation through binocular vision fusion, and makes the galloping displacement extraction more reliable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-voltage transmission line galloping monitoring technology, specifically a high-voltage transmission line galloping monitoring method based on binocular vision and deep learning. Background Technology

[0002] High-voltage transmission lines play a crucial role in transmitting electrical energy. Galloping, a low-frequency, large-amplitude self-excited vibration phenomenon caused by icing or uneven stress on transmission lines under wind excitation, can lead to electrical faults and / or mechanical damage, thus affecting the normal operation and stability of the power system. Galloping has become one of the main factors threatening the safety of transmission lines; therefore, monitoring transmission line galloping is of great significance for line safety and the normal operation of the power system.

[0003] Currently, transmission line galloping is mainly monitored through sensor data and video images. The sensor-based monitoring method involves installing sensors on transmission towers to collect various parameters during transmission line galloping, and then remotely processing and analyzing these parameters to achieve online galloping monitoring. This method can accurately measure parameters such as galloping amplitude, frequency, and half-wave number, facilitating the acquisition of complete transmission line galloping waveforms. However, deploying more sensors increases investment and maintenance costs, and also increases the burden on the lines, further increasing safety risks. Using fewer sensors can reduce costs, but accuracy will decrease, and it will be unable to accurately fit the line galloping trajectory.

[0004] Video-based monitoring methods acquire real-time video of power line galloping, extract keyframe images, and perform image analysis to determine if the line is galloping. This method eliminates the need for numerous sensors on the transmission line, saving on installation and maintenance costs, but it still has shortcomings. First, most methods remain at the two-dimensional image level, identifying and classifying pixel displacements, morphological changes, or galloping states of the transmission line in the image, but failing to directly obtain its three-dimensional displacement, amplitude, and trajectory in real space, leading to a significant discrepancy between the monitoring results and the actual galloping behavior. Second, some methods use optical flow estimation networks or end-to-end regression networks for galloping displacement calculation, but optical flow methods are susceptible to changes in illumination, background interference, occlusion, and the slender structural characteristics of the line, resulting in insufficient stability in complex environments. Simultaneously, end-to-end direct regression of galloping amplitude or frequency lacks clear physical constraints, making its output difficult to interpret, limiting its generalization ability, and easily leading to a significant decrease in accuracy when operating conditions change. Third, deep learning methods based on monocular vision generally suffer from scale uncertainty. The amplitude of the dancing motion often needs to be obtained through manual calibration, empirical scaling, or post-correction, making it difficult to operate stably in the long term. When the camera installation position or viewing angle changes, the model needs to be recalibrated or retrained, resulting in poor engineering adaptability. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the technical problem this invention aims to solve is to provide a method for monitoring high-voltage transmission line galloping based on binocular vision and deep learning.

[0006] The present invention solves the aforementioned technical problem by adopting the following technical solution: A method for monitoring high-voltage transmission line galloping based on binocular vision and deep learning, characterized by the following steps: Step 1: Use a stereo camera to acquire video of power line galloping, extract keyframe images from the video to obtain left and right power line image sequences; perform synchronous annotation and data augmentation on the left and right power line images to obtain a stereo power line image dataset; Step 2: Construct a transmission line segmentation model and train it using a binocular transmission line image dataset; use the trained transmission line segmentation model to segment the left and right transmission line images to obtain semantic segmentation images of the left and right transmission lines. The transmission line segmentation model includes a dual-branch U-net network, which comprises a dual-branch encoder and a dual-branch decoder. The dual-branch encoder and decoder are connected via a cross-view attention fusion module. The dual-branch encoder includes left and right encoding branches, with the left encoding branch including… n The right encoding branch includes one left encoding module and one self-attention module, with the self-attention module embedded between the last two left encoding modules. A downsampling operation is embedded between adjacent modules. n There is one right encoding module and one self-attention module, and the connection relationship is the same as that of the left encoding branch; The dual-branch decoder includes left and right decoding branches. The left decoding branch includes... n The left decoding module has an upsampling operation embedded between adjacent decoding modules; the right decoding branch includes... n The connection relationship of the right decoding module is the same as that of the left decoding branch; No. n The left and right view encoded features output by the left and right encoding modules are processed by the first... n Interact with the cross-perspective attention fusion module to obtain the first... n Enhancement features for the left and right views; n The left and right view enhancement features are used as the input features of the first left decoding module and the right decoding module, respectively; n -1 The output of the left and right encoding modules n- One left and right view encoding feature is processed by the first n -1 cross-view attention fusion module interacts to obtain the first... n-One left and right view enhancement feature; the left view decoding feature output by the first left decoding module is upsampled and compared with the first... n -1 left view enhancement features are concatenated, and the resulting feature is used as the input feature of the second left decoding module; the right view decoding feature output by the first right decoding module is upsampled and then combined with the -1 left view enhancement feature. n -1 right-view enhancement features are concatenated, and the resulting feature serves as the input feature for the second right-view decoding module; similarly, the connection relationships between the remaining cross-view attention fusion modules and their corresponding encoding and decoding modules are obtained; among them... n It is a positive integer greater than 2; In the cross-view attention fusion module, the left and right view encoding features are respectively normalized by layers and then concatenated to form key vectors and value vectors. The normalized left view encoding features are used as query vectors. The query vector, key vector, and value vector are processed by a multi-head attention mechanism. The output features of the multi-head attention mechanism are normalized by layers and then residually connected with the left view encoding features to obtain the left view enhancement features. Similarly, the right view enhancement features are obtained. Step 3: Process the semantic segmentation images of the left and right transmission lines to obtain the galloping displacement of the transmission lines; Binarization is performed on the semantic segmentation images of the left and right transmission lines. The largest connected regions in the binarized images are retained to obtain binary masks for the left and right transmission lines. Based on these binary masks, the transmission line regions are extracted as regions of interest (ROIs) from the original transmission line images. Reference object identification is performed within each ROI in the left and right transmission line images to obtain reference object feature points. The geometric distance from the reference object feature points in the left transmission line image to the corresponding epipolar line is calculated. If the geometric distance is greater than a tolerance threshold, the left and right transmission line images are considered... The pairing of reference feature points in the images is invalid and discarded; if the geometric distance is less than or equal to the tolerance threshold, the pairing of reference feature points in the left and right transmission line images is valid, and a matching reference feature point pair is obtained; for the matching reference feature point pair, the three-dimensional coordinates of the reference feature points in the corresponding key frame images are calculated by triangulation using the calibration parameters of the binocular camera; the three-dimensional coordinates of the reference feature points in all key frame images are obtained by traversing the transmission line image sequence, thereby obtaining the motion displacement of the reference feature points in each key frame image, i.e., the galloping displacement of the transmission line; Step 4: Construct and train the galloping displacement prediction model; input the historical galloping displacement sequence of the transmission line into the trained galloping displacement prediction model to obtain the transmission line galloping displacement prediction sequence; The galloping displacement prediction model includes a time-frequency domain dual-branch physical parameter regression model and a galloping displacement prediction model. In the time-frequency domain dual-branch physical parameter regression model, the historical transmission line galloping displacement sequence is processed by convolution and ReLU activation function, and then processed by time domain feature extraction branch and frequency domain feature extraction branch respectively to obtain time domain features and frequency domain features. The time domain features and frequency domain features are concatenated to obtain time-frequency domain joint features. The time-frequency domain joint features are then mapped by fully connected layer and ReLU activation function respectively to obtain galloping amplitude and galloping dominant frequency. The galloping displacement prediction model includes an encoder and a decoder. The encoder extracts galloping displacement context encoding features from historical transmission line galloping displacement sequences. The galloping amplitude and dominant galloping frequency are concatenated to obtain a galloping physical parameter vector. This vector is then mapped using a fully connected layer and a ReLU activation function to obtain an initial hidden state. The decoder includes a recurrent neural network and a fully connected layer. The recurrent neural network starts from the initial hidden state and updates the hidden state for the next time step by combining the predicted transmission line galloping displacement value at the current time step. The galloping displacement context encoding features are concatenated with the hidden state at the current time step. The resulting features are then mapped using a fully connected layer and a ReLU activation function to obtain the predicted transmission line galloping displacement value for the next time step, thus gradually generating a transmission line galloping displacement prediction sequence.

[0007] Furthermore, the temporal feature extraction branch includes multiple convolutional modules, with a global pooling layer connected after the last convolutional module. Each convolutional module includes a 1D convolutional layer, a ReLU activation layer, and a pooling layer connected in sequence.

[0008] Furthermore, the frequency domain feature extraction branch includes a Fast Fourier Transform (FFT) layer and multiple fully connected layers, with a ReLU activation function connected after each fully connected layer. The FFT layer converts the historical transmission line galloping displacement sequence into a frequency domain amplitude spectrum, which is then dimensionality-reduced by the fully connected layer and nonlinearly transformed by the ReLU activation function.

[0009] Furthermore, the encoder of the dancing displacement prediction model includes a convolutional layer, each convolutional layer is followed by a ReLU activation function, and the last convolutional layer is followed by a bidirectional LSTM layer.

[0010] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention proposes a transmission line segmentation model that combines a cross-view attention fusion module with a dual-branch U-Net network. By effectively fusing spatial and parallax information between binocular views, the transmission line region is accurately identified, overcoming the limitations of monocular views under complex backgrounds and lighting changes, and significantly improving the accuracy and robustness of transmission line segmentation.

[0011] 2. Unlike traditional monocular vision or optical flow methods that can only obtain the pixel displacement of feature points in an image, this invention is based on the principle of binocular stereo vision. Through epipolar geometry verification and triangulation, it directly calculates the three-dimensional coordinates of the reference object feature points in the real world from the binocular view. This achieves a precise mapping from two-dimensional images to three-dimensional physical quantities, making the obtained dancing displacement data more realistic and reliable, and laying a solid foundation for accurate dancing analysis.

[0012] 3. Through a two-stage galloping displacement prediction model, the key physical parameters of galloping (including the dominant frequency and amplitude of galloping) are accurately regressed from the displacement sequence. On the other hand, the key physical parameters are used as strong guiding information to predict the future short-term displacement sequence. This extends the real-time monitoring of transmission line galloping to trend prediction, and has the foresight of anomaly monitoring. It provides a reference for timely formulation of targeted measures to prevent galloping, thereby reducing the risk of line galloping-induced faults. Attached Figure Description

[0013] Figure 1 This is an overall flowchart of the present invention; Figure 2 This is a schematic diagram of the transmission line segmentation model of the present invention; Figure 3 This is a flowchart of the dance displacement prediction process of the present invention; Figure 4 This is a schematic diagram of the time-frequency domain dual-branch physical parameter regression model of the present invention; Figure 5 This is a schematic diagram of the structure of the dancing displacement prediction model of the present invention. Detailed Implementation

[0014] Specific embodiments are given below with reference to the accompanying drawings. These specific embodiments are only used to describe the technical solution of the present invention in detail and are not intended to limit the scope of protection of this application.

[0015] This invention provides a method for monitoring galloping of high-voltage transmission lines based on binocular vision and deep learning (hereinafter referred to as the method, see [link]). Figures 1-5 ), including the following steps: Step 1: Use a stereo camera to acquire video of power line galloping, extract keyframe images from the video to obtain time-aligned left and right power line image sequences; perform synchronous annotation and data augmentation on the left and right power line images to obtain a stereo power line image dataset; Install a reference object (such as a small sphere) on the power transmission line, mount a binocular camera on the high-voltage power pole, and use the binocular camera to collect videos of the power transmission line galloping. The videos should include the reference object and multiple video clips with different amplitudes of the power transmission line galloping to ensure a sufficiently large dataset. Using video processing tools, keyframe images from the binocular transmission line galloping video were extracted synchronously at a rate of 1 frame / second, forming time-aligned left and right transmission line image sequences. The foreground (transmission lines) and background of the left and right transmission line images were synchronously labeled using the Image Labeler in Matlab software. Data enhancement, including random rotation, translation, and brightness transformation, was then performed synchronously on the labeled left and right transmission line images to form several transmission line image pairs (including left and right transmission line images), thus obtaining a binocular transmission line image dataset.

[0016] Step 2: Construct a transmission line segmentation model based on cross-view attention fusion and a dual-branch U-Net network, and train the transmission line segmentation model using a binocular transmission line image dataset; use the trained transmission line segmentation model to segment the left and right transmission line images to obtain semantic segmentation images of the left and right transmission lines. like Figure 2 As shown, the transmission line segmentation model includes a dual-branch U-net network, which comprises a dual-branch encoder and a dual-branch decoder. The dual-branch encoder and the dual-branch decoder are connected by a cross-view attention fusion module (CVAFM). The dual-branch encoder includes left and right encoding branches, and the dual-branch decoder includes left and right decoding branches. The left and right transmission line images are processed by the left and right encoding branches to extract multi-scale left and right view encoding features. The left and right view encoding features are then decoded by the left and right decoding branches to obtain the left and right view decoding features. The left coding branch includes n One left encoding module ( n The code uses three positive integers greater than 2 (in this embodiment, three of them) and a self-attention module. The self-attention module is embedded between the last two left encoding modules to capture long-range context information, and a downsampling operation is embedded between adjacent modules; the right encoding branch includes... n There is one right-encoding module and one self-attention module. The connection relationships between the modules are similar to those of the left-encoding branch. The left decoding branch includes n There are three left decoding modules (three in this embodiment), with an upsampling operation embedded between adjacent decoding modules; the right decoding branch includes... n The connection relationship of the right decoding module is the same as that of the left decoding branch; No. n The left and right view encoded features output by the left and right encoding modules are processed by the first... n Interact with the cross-perspective attention fusion module to obtain the first... n Enhancement features for the left and right views; n The left and right view enhancement features are used as the input features of the first left decoding module and the right decoding module, respectively; n -1 The output of the left and right encoding modules n- One left and right view encoding feature is processed by the first n -1 cross-view attention fusion module interacts to obtain the first... n- One left and right view enhancement feature; the left view decoding feature output by the first left decoding module is upsampled and compared with the first... n -1 left view enhancement features are concatenated, and the resulting feature is used as the input feature of the second left decoding module; similarly, the right view decoding feature output by the first right decoding module is upsampled and then concatenated with the first left view enhancement feature. n -1 right view enhancement features are concatenated, and the concatenated features are used as input features for the second right decoding module; similarly, the connection relationships between the remaining cross-view attention fusion modules and the corresponding encoding and decoding modules are obtained.

[0017] The left and right encoding modules have the same structure, both including depthwise separable convolution and ReLU activation function; the left and right decoding modules have the same structure, both including depthwise separable convolution and ReLU activation function; depthwise separable convolution significantly reduces the number of parameters and computational cost while ensuring feature expressiveness, in order to meet real-time requirements.

[0018] In the cross-view attention fusion module, the left and right view encoded features are each layer-normalized before being concatenated to obtain shared key and value vectors. The layer-normalized left view encoded features are used as the query vector. The query vector, key vector, and value vector are then processed through a multi-head attention mechanism to achieve efficient fusion and interaction of the left and right view encoded features. The output features of the multi-head attention mechanism are layer-normalized and then residually connected with the left view encoded features to obtain the left view enhanced features. Similarly, the right view enhanced features are obtained. Through this cross-view attention interaction method, the model can capture the correlation and disparity consistency of corresponding regions under different viewpoints, effectively improving the geometric alignment ability and spatial structure information fusion between the left and right views.

[0019] Step 3: Process the semantic segmentation images of the left and right transmission lines to obtain the galloping displacement of the transmission lines; To facilitate the extraction of transmission line galloping displacement, the semantic segmentation images of the left and right transmission lines need to be binarized, retaining only the largest connected regions in the binarized images to obtain binary masks for the left and right transmission lines containing only the transmission line regions. Based on the binary masks, the transmission line regions are extracted as regions of interest in the original transmission line images to remove interference from complex backgrounds. Since the reference objects are located on the transmission lines and gallop synchronously with them, calculating the galloping displacement of the transmission lines is transformed into calculating the motion displacement of the reference objects. Within the regions of interest in the left and right transmission line images, a target detection algorithm is used to identify uniquely identified reference objects, and the geometric center points of these reference objects are selected as reference object feature points, thus obtaining the reference object feature points in the left and right transmission line images. The reference feature points in the left and right transmission line images are denoted as... and The '1' indicates a placeholder used to extend the coordinates of feature points from two dimensions to three dimensions; the reference feature points in the left and right transmission line images must satisfy the following polar geometric constraint equations: (1) In the formula, Indicates the transpose operation; The fundamental matrix is ​​determined by the intrinsic and extrinsic parameters of the stereo camera. To quantify the matching error, the points are calculated according to equation (2). to its corresponding polar line geometric distance : (2) In the formula, This indicates taking the absolute value. , represent the coefficients of the first and second terms of the polar equation, respectively; Set a tolerance threshold (Usually 1-3 pixels), if If the pairing of reference feature points in the left and right transmission line images is deemed invalid and removed, the geometric accuracy of the reference feature point pairing is ensured; if If the pairing of reference feature points in the left and right transmission line images is deemed valid, then a matching reference feature point pair is obtained. For the matching reference feature point pair, the three-dimensional coordinates of the reference feature points in the corresponding keyframe image are calculated using triangulation based on the calibration parameters of the binocular camera. By traversing the transmission line image sequence, the three-dimensional coordinates of the reference feature points in all keyframe images are obtained, thus yielding the motion displacement of the reference feature points in each keyframe image. This refers to the galloping displacement of power transmission lines.

[0020] Step 4: Construct and train a galloping displacement prediction model. Input the historical galloping displacement sequence of transmission lines into the trained galloping displacement prediction model to obtain the transmission line galloping displacement prediction sequence, thereby realizing intelligent analysis of transmission line galloping displacement from physical perception to trend prediction. like Figure 3 As shown, the galloping displacement prediction model consists of two stages. The first stage uses a time-frequency domain dual-branch physical parameter regression model to predict key physical parameters, including galloping amplitude and galloping dominant frequency. The second stage uses the galloping amplitude and galloping dominant frequency as strong guiding signals and uses the galloping displacement prediction model to process the historical transmission line galloping displacement sequence to obtain the predicted value of the transmission line galloping displacement.

[0021] like Figure 4 As shown, in the time-frequency domain dual-branch physical parameter regression model, the historical transmission line galloping displacement sequence is processed by convolution and ReLU activation function, and then processed by time domain feature extraction branch and frequency domain feature extraction branch respectively to obtain time domain features and frequency domain features, ensuring the completeness of feature expression; the time domain features and frequency domain features are concatenated to obtain time-frequency domain joint features; the time-frequency domain joint features are then mapped by fully connected layer and ReLU activation function respectively to obtain galloping amplitude and galloping dominant frequency.

[0022] The temporal feature extraction branch includes multiple convolutional modules, with a global pooling layer connected after the last convolutional module to map temporal features to a fixed length. Each convolutional module includes a 1D convolutional layer, a ReLU activation layer, and a pooling layer connected in sequence to capture local features and short-term dependencies of the displacement sequence in the time dimension.

[0023] The frequency domain feature extraction branch includes a Fast Fourier Transform (FFT) layer and multiple fully connected layers, with each fully connected layer followed by a ReLU activation function. The FFT layer converts the historical transmission line galloping displacement sequence into a frequency domain amplitude spectrum, the fully connected layers reduce the dimensionality of the frequency domain amplitude spectrum, and the ReLU activation function performs a nonlinear transformation to accurately extract the frequency components and energy distribution in the displacement sequence.

[0024] like Figure 5As shown, the galloping displacement prediction model includes an encoder and a decoder. The encoder comprises multiple convolutional layers, each followed by a ReLU activation function, and the last convolutional layer is followed by a bidirectional LSTM layer. These convolutional layers act as local feature extractors, capturing fine-grained local fluctuation patterns and short-term temporal dependencies in the displacement sequence. The bidirectional LSTM layer comprehensively analyzes and encodes the historical and future context information at each moment through forward and backward propagation paths, thereby comprehensively capturing the complete cycle and long-term dependencies in the galloping displacement and obtaining the forward hidden state of the historical transmission line galloping displacement sequence. and backward hidden state The forward and backward hidden states are concatenated to obtain the dancing displacement context encoding features. ;in, This indicates a splicing operation.

[0025] The dancing amplitude and the dominant dancing frequency are concatenated to obtain the dancing physical parameter vector; this vector is then mapped through a fully connected layer and a ReLU activation function to obtain the initial hidden state. The decoder consists of a recurrent neural network (RNN) and fully connected layers. The RNN uses an autoregressive approach, starting from the initial hidden state and combining it with the predicted transmission line galloping displacement value at the current time step to update the hidden state for the next time step; the galloping displacement context encoding features are then applied. The hidden state at the current time step is concatenated with the hidden state. The concatenated features are then mapped through a fully connected layer and a ReLU activation function to obtain the predicted value of the transmission line galloping displacement at the next time step, thereby gradually generating the transmission line galloping displacement prediction sequence.

[0026] In summary, firstly, the transmission line segmentation model combining a cross-view attention fusion module and a dual-branch U-Net network significantly improves the segmentation accuracy of transmission lines in complex scenarios. Secondly, by directly calculating the three-dimensional coordinates of reference feature points through binocular matching, epipolar geometric verification, and triangulation, a precise mapping from two-dimensional images to three-dimensional coordinates is achieved, resulting in more accurate transmission line galloping displacement. Finally, by constructing a two-stage galloping displacement prediction model, high-precision regression of key physical parameters and prediction of galloping displacement are achieved while effectively reducing computational resource consumption. Furthermore, during the galloping displacement prediction process, the galloping amplitude and dominant frequency are innovatively fused into the decoding process as strong guiding signals, ensuring the physical reliability of the prediction results. This invention extends real-time monitoring of transmission line galloping to trend prediction, enabling early prediction of short-term galloping displacement trends and achieving early warning and risk assessment.

[0027] Any aspects not covered in this invention are applicable to existing technologies.

Claims

1. A method for monitoring galloping of high-voltage transmission lines based on binocular vision and deep learning, characterized in that, Includes the following steps: Step 1: Use a stereo camera to acquire video of power line galloping, extract keyframe images from the video to obtain left and right power line image sequences; perform synchronous annotation and data augmentation on the left and right power line images to obtain a stereo power line image dataset; Step 2: Construct a transmission line segmentation model and train the transmission line segmentation model using a binocular transmission line image dataset; The trained transmission line segmentation model is used to segment the left and right transmission line images to obtain semantic segmentation images of the left and right transmission lines. The transmission line segmentation model includes a dual-branch U-net network, which comprises a dual-branch encoder and a dual-branch decoder. The dual-branch encoder and decoder are connected via a cross-view attention fusion module. The dual-branch encoder includes left and right encoding branches, with the left encoding branch including… n The right encoding branch includes one left encoding module and one self-attention module, with the self-attention module embedded between the last two left encoding modules. A downsampling operation is embedded between adjacent modules. n There is one right encoding module and one self-attention module, and the connection relationship is the same as that of the left encoding branch; The dual-branch decoder includes left and right decoding branches. The left decoding branch includes... n Each left decoding module has an upsampling operation embedded between adjacent decoding modules; The right decoding branch includes n The connection relationship of the right decoding module is the same as that of the left decoding branch; No. n The left and right view encoded features output by the left and right encoding modules are processed by the first... n Interact with the cross-perspective attention fusion module to obtain the first... n Enhanced features in both left and right views; No. n The left and right view enhancement features are used as the input features of the first left decoding module and the right decoding module, respectively; No. n -1 The output of the left and right encoding modules n- One left and right view encoding feature is processed by the first n -1 cross-view attention fusion module interacts to obtain the first... n- One left and right view enhancement feature; The left-view decoding features output by the first left decoding module are upsampled and compared with those of the second left decoding module. n -1 left view enhancement features are concatenated, and the concatenated features are used as the input features of the second left decoding module; The right view decoding features output by the first right decoding module are upsampled and compared with those of the second right decoding module. n -1 right view enhancement features are concatenated, and the concatenated features are used as the input features of the second right decoding module; Similarly, the connection relationships between the remaining cross-view attention fusion modules and their corresponding encoding and decoding modules are obtained; among them, n It is a positive integer greater than 2; In the cross-view attention fusion module, the left and right view encoding features are respectively normalized by layers and then concatenated to form key vectors and value vectors. The normalized left view encoding features are used as query vectors. The query vector, key vector, and value vector are processed by a multi-head attention mechanism. The output features of the multi-head attention mechanism are normalized by layers and then residually connected with the left view encoding features to obtain the left view enhancement features. Similarly, the right view enhancement features are obtained. Step 3: Process the semantic segmentation images of the left and right transmission lines to obtain the galloping displacement of the transmission lines; Binarization is performed on the semantic segmentation images of the left and right transmission lines. The largest connected regions in the binarized images are retained to obtain binary masks for the left and right transmission lines. Based on these binary masks, the transmission line regions are extracted as regions of interest (ROIs) from the original transmission line images. Reference object identification is performed within each ROI in the left and right transmission line images to obtain reference object feature points. The geometric distance from the reference object feature points in the left transmission line image to the corresponding epipolar line is calculated. If the geometric distance is greater than a tolerance threshold, the left and right transmission line images are considered... The pairing of reference feature points in the images is invalid and discarded; if the geometric distance is less than or equal to the tolerance threshold, the pairing of reference feature points in the left and right transmission line images is valid, and a matching reference feature point pair is obtained; for the matching reference feature point pair, the three-dimensional coordinates of the reference feature points in the corresponding key frame images are calculated by triangulation using the calibration parameters of the binocular camera; the three-dimensional coordinates of the reference feature points in all key frame images are obtained by traversing the transmission line image sequence, thereby obtaining the motion displacement of the reference feature points in each key frame image, i.e., the galloping displacement of the transmission line; Step 4: Construct and train the galloping displacement prediction model; input the historical galloping displacement sequence of the transmission line into the trained galloping displacement prediction model to obtain the transmission line galloping displacement prediction sequence; The galloping displacement prediction model includes a time-frequency domain dual-branch physical parameter regression model and a galloping displacement prediction model. In the time-frequency domain dual-branch physical parameter regression model, the historical transmission line galloping displacement sequence is processed by convolution and ReLU activation function, and then processed by time domain feature extraction branch and frequency domain feature extraction branch respectively to obtain time domain features and frequency domain features. The time domain features and frequency domain features are concatenated to obtain time-frequency domain joint features. The time-frequency domain joint features are then mapped by fully connected layer and ReLU activation function respectively to obtain galloping amplitude and galloping dominant frequency. The galloping displacement prediction model includes an encoder and a decoder. The encoder extracts galloping displacement context encoding features from historical transmission line galloping displacement sequences. The galloping amplitude and dominant galloping frequency are concatenated to obtain a galloping physical parameter vector. This vector is then mapped using a fully connected layer and a ReLU activation function to obtain an initial hidden state. The decoder includes a recurrent neural network and a fully connected layer. The recurrent neural network starts from the initial hidden state and updates the hidden state for the next time step by combining the predicted transmission line galloping displacement value at the current time step. The galloping displacement context encoding features are concatenated with the hidden state at the current time step. The resulting features are then mapped using a fully connected layer and a ReLU activation function to obtain the predicted transmission line galloping displacement value for the next time step, thus gradually generating a transmission line galloping displacement prediction sequence.

2. The high-voltage transmission line galloping monitoring method based on binocular vision and deep learning according to claim 1, characterized in that, The temporal feature extraction branch includes multiple convolutional modules, with a global pooling layer connected after the last convolutional module. Each convolutional module includes a 1D convolutional layer, a ReLU activation layer, and a pooling layer connected in sequence.

3. The high-voltage transmission line galloping monitoring method based on binocular vision and deep learning according to claim 1, characterized in that, The frequency domain feature extraction branch includes a Fast Fourier Transform (FFT) layer and multiple fully connected layers, with a ReLU activation function following each fully connected layer. The FFT layer converts the historical transmission line galloping displacement sequence into a frequency domain amplitude spectrum, which is then dimensionality-reduced by the fully connected layer and nonlinearly transformed by the ReLU activation function.

4. The high-voltage transmission line galloping monitoring method based on binocular vision and deep learning according to claim 1, characterized in that, The encoder of the dancing displacement prediction model includes several convolutional layers, each followed by a ReLU activation function, and the last convolutional layer is followed by a bidirectional LSTM layer.