Method, system and equipment for measuring icing volume of power transmission line, medium and program product

By combining binocular stereo vision with a deep learning model, non-contact three-dimensional automated measurement of icing on transmission lines has been achieved, solving the problems of insufficient accuracy and reliance on manual labor in existing technologies, and providing an efficient, accurate and safe solution for measuring icing volume.

CN121708074APending Publication Date: 2026-03-20XIDIAN UNIV +1

Patent Information

Application Number
CN202511876792.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing technologies for detecting icing on power transmission lines suffer from insufficient accuracy, reliance on manual operation, and the inability to obtain three-dimensional information due to the contact-based arrangement of sensors.

Method used

A non-contact 3D reconstruction method is achieved by combining depth estimation based on binocular stereo vision with YOLO object detection and SAM2 semantic segmentation. High-precision ice volume data is obtained through accurate modeling and automated segmentation using deep learning models.

Benefits of technology

It achieves high-precision three-dimensional icing volume measurement under non-contact conditions, improves measurement accuracy and safety, reduces human intervention, has strong adaptability and robustness, and can provide reliable decision-making basis for power operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708074A_ABST
    Figure CN121708074A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of three-dimensional measurement, and discloses a power transmission line icing volume measurement method, system and device, a medium and a program product, and the method comprises the steps: employing a binocular camera to shoot and correct an icing power transmission line view, and obtaining internal and external parameters, performing depth estimation on the power transmission line view based on the corrected power transmission line view and the internal and external parameters, training a conductor detection model and inputting the corrected power transmission line view to obtain a detection frame and a corresponding confidence coefficient, using the detection frame with the highest confidence coefficient as a prompt, inputting the prompt and the corrected power transmission line view into an SAM2 network for segmentation, and obtaining a segmentation result; de-noising and morphological processing are carried out on the masks of the ice-coated wires, and finally the ice-coated area of the power transmission line is calculated. The system, the equipment and the medium are used for implementing the method. The program product comprises a computer program for implementing the method; the method for measuring the icing volume of the power transmission line has the effects of stability, reliability, high efficiency, accuracy and high safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of three-dimensional measurement technology, specifically relating to a method, system, equipment, medium, and program product for measuring the ice volume of transmission lines. Background Technology

[0002] Power transmission lines play a crucial role in my country's power system, ensuring the safe and reliable transmission of electricity. They also contribute significantly to my country's high-quality economic development. In recent years, with the continuous changes in the global ecological environment, extreme cold and cold waves have become more frequent, leading to icing on power transmission lines. Icing prevents transmission lines from functioning properly and reliably. When icing exceeds design standards, it can cause ice-related accidents such as tower collapse, wire breakage, conductor galloping, and switchgear malfunctions.

[0003] In the field of transmission line icing detection, there are various existing research methods, each with its own characteristics, as follows: Simulated conductor method: A conductor of the same material and diameter as the actual transmission line is installed near the line to simulate icing. The advantage is that it reduces workload compared to manual line inspection; however, because the current in the simulated conductor differs from that of the actual line, the accuracy of icing detection is significantly inaccurate. Fiber optic sensing method: This method uses fiber optic sensors (fixed to insulators) to detect icing by utilizing the change in sensor stress caused by temperature changes before and after icing. Its advantages include high automation and significantly reduced workload; however, its accuracy is affected by sensor temperature drift and time lag, and the fiber optic material is fragile, easily damaged, and difficult to maintain.

[0004] Capacitance method: Parallel induction conductors are erected along the transmission line. Utilizing the capacitance effect, the capacitance between the line under test and the induction conductor (which varies with dielectric constant and ice thickness) is used to detect ice accumulation. The principle is simple, but it requires the installation of a complete induction line, making operation difficult and costly, thus hindering its widespread adoption. Image detection method: This is a non-contact method that uses images of icing on power transmission lines (sources include tower cameras, drones, and manual inspections) and image processing algorithms to estimate the ice thickness. It is simple to operate, low in cost, and highly accurate, making it one of the mainstream methods.

[0005] Wei Yewen et al. proposed an invention patent, CN114677428A, entitled "A Method for Detecting Icing Thickness of Power Transmission Lines Based on UAV Image Processing." The method first converts the original RGB images of icing power transmission lines acquired by the UAV into grayscale images; then, it uses the maximum inter-class variance method for initial segmentation of the grayscale images, completing the preprocessing of the original images; finally, it extracts the icing information of the power transmission line by combining the icing information and the connected component feature parameters of the background noise; and finally, it proposes a vertical approximation method to obtain the thickness value of the iced power transmission line in the vertical direction, thus calculating the icing thickness. This method requires high positioning accuracy from the UAV, cannot operate in adverse weather conditions, and cannot represent three-dimensional information.

[0006] Wang Shuai et al. proposed an invention patent, CN118052760A, entitled "An Automated Method for Detecting the Thickness of Iced Transmission Lines." This method inputs an image of an iced transmission line into an iced transmission line thickness detection model, obtaining coded features through an encoding path. These coded features are then fused with upsampled features and input into a dual-attention fusion module. First, a coordinate attention mechanism extracts and concatenates the horizontal and vertical features of the input features to obtain the output features of the coordinate attention mechanism. This output feature is then concatenated with the input features and input into a channel attention mechanism. Feature extraction is performed on each channel and its multiple neighboring channels, capturing local cross-channel information of the concatenated features. After a decoding path, a segmentation result image of the iced transmission line is output. However, this method suffers from inaccuracies due to the angle between the image captured by the camera and the iced transmission line. Summary of the Invention

[0007] To overcome the shortcomings of the existing technologies, the present invention aims to provide a method and system for measuring the ice volume of power transmission lines based on binocular stereo vision. This method integrates binocular stereo vision depth estimation (using advanced deep learning models such as FoundationStereo), YOLO object detection, and SAM2 semantic segmentation into the field of power line ice detection, forming a non-contact, fully automated measurement system capable of 3D reconstruction. This system possesses key features such as accurate modeling, fully automated segmentation, and point cloud 3D reconstruction, which can improve the efficiency and accuracy of power line ice monitoring. It solves a series of problems inherent in traditional technologies, such as limitations imposed by manual operation, contact-based sensor placement, and the inability to obtain 3D information. The system architecture is flexible, facilitating on-site deployment and widespread application.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for measuring the ice volume of a power transmission line includes the following steps: Obtain the intrinsic and extrinsic parameters of the binocular camera, use the binocular camera to capture a view of the icy transmission line, correct the transmission line view, and obtain new intrinsic and extrinsic parameters. Based on the corrected transmission line view and the new intrinsic and extrinsic parameters, perform depth estimation on the corrected transmission line view. Based on the object detection dataset, a conductor detection model was trained using YOLO 11. The corrected transmission line view was then input into the conductor detection model to obtain the detection boxes and confidence scores of the iced conductors. The detection box with the highest confidence score was used as a cue. The cue and the corrected transmission line view were then input into the Segment Anything Model 2 (SAM2) network for accurate segmentation to obtain the mask of the iced conductor. The mask of the iced conductor was then denoised and morphologically processed. The icing volume of the transmission line is calculated based on the mask of the iced conductor after denoising and morphological processing and the depth estimation of the corrected transmission line view.

[0009] The intrinsic and extrinsic parameters of the binocular camera specifically include the first focal length, the first optical center, the first rotation matrix, the first translation vector, and the distortion coefficients in the x-axis and y-axis directions of the image coordinate system.

[0010] The specific steps for capturing and correcting views of icy power transmission lines using a binocular camera are as follows: Using a binocular camera, images of ice-covered power transmission lines are captured to obtain RGB left and RGB right views of the transmission lines. The intrinsic and extrinsic parameters are input into the OpenCV stereo calibration algorithm to calibrate the RGB left and RGB right views onto the same plane, resulting in calibrated RGB left and RGB right views. The depth information corresponding to the calibrated RGB left view is saved as an npy format file. The new intrinsic and extrinsic parameters include the second focal length, the second optical center, and the second translation vector in the x-axis and y-axis directions of the image coordinate system.

[0011] The process of depth estimation of the corrected transmission line view based on the corrected transmission line view and new intrinsic and extrinsic parameters is as follows: The corrected RGB left view and corrected RGB right view are processed by the Side-Tuning Adapter (STA) network of the FoundationStereo network to obtain deep features and contextual features: The corrected RGB left and right views are input into the Side-Tuning CNN and the pre-trained DepthAnythingV2 network, respectively. Through feature fusion, deep features that combine monocular priors and multi-scale spatial details are extracted, enhancing the feature representation capability. Specifically, the corrected RGB left and right views are first processed by the Side-Tuning CNN with shared weights to extract multi-scale spatial features, and then downsampled at a 1 / 4 scale through convolution. These features are then fused with the high semantic and rich geometric priors output by DepthAnythingV2 to obtain deep features. The corrected RGB left view is further input into a context network to obtain context features; The deep features of the corrected RGB left and right views at the 1 / 4 scale are processed in two ways: Method 1: Group-wise correlation is used to calculate the group correlation volume, which characterizes the matching similarity of binocular features across different channel groups. Method 2: The deep features of the corrected RGB left and right views are directly stitched together along the channel dimension under the assumption of disparity shift, resulting in a stitched feature volume. This preserves rich monocular prior and spatial detail information. Subsequently, the group correlation volume and the stitched feature volume are merged along the channel dimension to construct a 4D hybrid cost volume. ; Attentive Hybrid Cost Filtering (AHCF) network for 4D hybrid cost volumes The AHCF network, which performs processing, consists of two sub-networks: Axial-Planar Convolution (APC) and Disparity Transformer (DT). Their functions are as follows: The APC network decomposes 3D convolution into separate convolutions in two directions: spatial and parallax. The DT network models the 4D hybrid cost volume using a global self-attention mechanism in the parallax dimension; the outputs of APC and DT are upsampled to restore the original 4D hybrid cost volume scale and then added and fused. For the 4D mixture cost volume output by AHCF, the initial disparity estimation results are obtained using the soft-argmin method: in, Indicates parallax. Indicates the initial disparity. This represents the maximum value of the parallax. This represents the hybrid cost volume after passing through the APC network; Based on the initial disparity, the network is recursively updated in multiple steps using ConvGRU (Convolutionally Gated Recurrent Unit) recursive refinement. Each refinement step is as follows: First, voxel lookup is performed using the latest parallax, and features are retrieved from the 4D hybrid cost volume and the grouped correlation volume, respectively. The concatenated features are then input into the convolutional gated recurrent unit (GRU) network. The GRU network simultaneously integrates contextual features and recursively refines disparity estimation. A multi-level GRU network recursively refines the output to achieve the final high-precision parallax. ; The depth is obtained based on the triangulation principle of a binocular camera: in, Indicates depth, Indicates the baseline distance of the binocular cameras. Indicates the focal length of a binocular camera. Indicates parallax.

[0012] The method involves training a conductor detection model using YOLO 11 based on the target detection dataset. The corrected transmission line view is then input into the model to obtain the detection boxes and their confidence scores for iced conductors. The detection box with the highest confidence score is used as a cue. The cue and the corrected transmission line view are then input into a Segment Anything Model 2 (SAM2) network for precise segmentation, resulting in a mask for the iced conductor. The specific steps for denoising and morphological processing of the mask for the iced conductor are as follows: A target detection dataset was constructed using multiple images. The target detection dataset includes icy guide lines and normal guide lines, with icy guide lines accounting for 45-55% and the remainder being normal guide lines. YOLO 11 was used as the pre-training algorithm for the object detection dataset. The object detection dataset was input into YOLO 11 to obtain the wire detection model. The corrected RGB left view is input into the conductor detection model to obtain the detection box and confidence level of the icy conductor. Using the detection box with the highest confidence as a cue, the cue and the corrected RGB left view are input into the SegmentAnything Model 2 network for accurate segmentation to obtain the mask of the icy wire; First, remove minor noise: Assume the mask for the icing conductor is a two-dimensional array. , Describe the height of the image. Indicates the width of the image. Indicates the first Line 1 Column pixels belong to the target. Indicates the first Line 1 The column pixels belong to the background; The mask of the icy conductor is labeled with connected regions in 4-pixel neighborhoods, and the label matrix is ​​as follows: Each connected target is assigned a unique integer label; Count all connected regions The total number of pixels, denoted as : Let the area threshold be Then for all Set the entire area to background 0: Processed icing wire mask Only targets with an area greater than or equal to the threshold are retained; Then perform morphological operations: Let the structural element be , It is a 3×3 all-1 square matrix: First, process the icing-covered wire mask. An opening operation is performed to remove small foreground noise and burrs. The process involves erosion followed by dilation, denoted as: in, Represents corrosion. Represents expansion The corrosion formula is: The expansion formula is: After the re-opening operation, the icing wire mask is processed. To perform a closing operation to fill small holes in the foreground, the process involves expansion followed by corrosion, denoted as: .

[0013] The specific steps for calculating the icing volume of the transmission line based on the mask of the iced conductor after noise reduction and morphological processing and the depth estimation of the corrected transmission line view are as follows: The mask obtained after the closing operation Perform skeletonization operations, using median transformation, to obtain the target backbone pixel set: in The first part represents the skeleton. pixel coordinates, N This represents the total number of sampling points for the skeleton. For the target main pixel set S Every point in Select the front and back points within the local window. , Fit the tangent direction and calculate the normal direction angle: Therefore, the unit vector of the normal direction is obtained: Based on skeleton points Centered on the mask, sample along the positive and negative normal directions until the mask boundary is encountered, thus obtaining the left and right endpoints. , ; Depth value at the endpoint , This is denoted as the median of the depth information during the sampling process; Endpoint pixel coordinates Using a pinhole camera model, second focal length Second focal length and the second optical center According to the endpoint depth Calculate its three-dimensional coordinates: in, The three-dimensional coordinates of the endpoints; Calculate the first The spatial diameter of the cross section is obtained by the Euclidean distance between the endpoints of the normal line segment at each point. : average diameter Represented as: No. The center point of each cross section The three-dimensional coordinates of the two endpoints of the spatial diameter of this cross section and The arithmetic mean is determined as follows: To reduce accumulated errors and computational redundancy, the sampling factor is defined as... ; Then the total length of the object Defined as: Among them, the remaining items This indicates that if the tail end is less than one sampling rate, distance compensation is performed from the last valid sampling point to the end point. For each length interval Let the actual length of this segment be: This segment contains sr local cross-sectional diameters. ; Define the representative diameter of this segment as the arithmetic mean: The radius of this segment is: The volume of the cylinder in that section is: Total volume The sum of the volumes of all the smaller cylinders: Among them, the remaining items This indicates that if the tail end is less than one sampling rate, volume compensation is performed from the last valid sampling point to the end point. The volume of ice accumulation is the total volume minus the volume of the corresponding portion of the conductor: in, Given the diameter of the wire, This represents the total length of the measured conductor.

[0014] A system for measuring the ice volume of a power transmission line, comprising: The depth estimation module acquires the intrinsic and extrinsic parameters of the binocular camera, uses the binocular camera to capture a view of the icy transmission line, corrects the transmission line view, acquires new intrinsic and extrinsic parameters, and performs depth estimation on the corrected transmission line view based on the corrected transmission line view and the new intrinsic and extrinsic parameters. The mask generation and processing module trains a conductor detection model using YOLO 11 based on the target detection dataset. The corrected transmission line view is then input into the conductor detection model to obtain the detection boxes and confidence scores of the icy conductors. The detection box with the highest confidence score is used as a cue. The cue and the corrected transmission line view are then input into the SegmentAnything Model 2 (SAM2) network for accurate segmentation to obtain the mask of the icy conductor. The mask of the icy conductor is then subjected to denoising and morphological processing. The icing volume calculation module calculates the icing zone of the transmission line based on the mask of the iced conductor after noise reduction and morphological processing, and the depth estimation of the corrected transmission line view.

[0015] A device for measuring the icing volume of a power transmission line, comprising: Memory: Used to store the computer program that implements the method for measuring the ice volume of transmission lines; Processor: Used to implement the method for measuring the ice volume of transmission lines when executing the computer program.

[0016] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for measuring the ice volume of transmission lines.

[0017] A computer program product includes a computer program that, when executed by a processor, implements a method for measuring the ice volume of transmission lines.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention uses binocular stereo vision combined with advanced deep learning models (FoundationStereo, YOLO11, SAM2) to achieve automated three-dimensional measurement of icing on transmission lines. Compared with solutions that rely on manual inspection or contact sensors, it can obtain high-precision three-dimensional icing volume data under non-contact conditions, significantly improving the accuracy and safety of the measurement.

[0019] 2. This invention employs a collaborative detection and segmentation strategy of YOLO11 and SAM2, which can automatically extract icing areas in complex backgrounds, exhibiting strong adaptability and robustness, reducing manual intervention, and improving processing efficiency.

[0020] 3. This invention, through morphological processing, skeletonization, and a three-dimensional reconstruction algorithm based on the normal direction, can accurately calculate the diameter, length, and total volume of the icing cross section. The results are stable and reliable, providing accurate decision-making basis for power operation and maintenance.

[0021] In summary, the method for measuring the ice volume of transmission lines according to the present invention is stable, reliable, efficient, accurate, and highly safe. Attached Figure Description

[0022] Figure 1 This is a flowchart of the measurement method of the present invention.

[0023] Figure 2(a) is a left view of the binocular camera calibration of the present invention.

[0024] Figure 2(b) is a right view of the binocular camera calibration of the present invention.

[0025] Figure 3(a) is a left view of the power transmission line taken using a binocular camera according to the present invention.

[0026] Figure 3(b) is a right view of the power transmission line taken using a binocular camera according to the present invention.

[0027] Figure 4(a) is a left view of the transmission line after correction according to the present invention.

[0028] Figure 4(b) is a right view of the transmission line after correction according to the present invention.

[0029] Figure 5 This is a schematic diagram of the FoundationStereo network structure used in this invention.

[0030] Figure 6 This is a visualization point cloud diagram of the experiment of this invention. Detailed Implementation

[0031] The present invention will now be described in detail with reference to the accompanying drawings.

[0032] like Figure 1 A method for measuring the ice volume of transmission lines based on binocular stereo vision includes the following steps: Depth estimation based on binocular stereo vision As shown in Figures 2(a) and 2(b), the binocular camera is calibrated using the checkerboard method to obtain the intrinsic and extrinsic parameters of the binocular camera. The intrinsic and extrinsic parameters include the first focal length, the first optical center, the first rotation matrix, the first translation vector, and the distortion coefficients in the x-axis and y-axis directions of the image coordinate system. As shown in Figures 3(a) and 3(b), the icy power transmission line is photographed using a binocular camera to obtain the RGB left and RGB right views of the power transmission line; as shown in Figures 4(a) and 4(b), the intrinsic and extrinsic parameters are input into the OpenCV stereo correction algorithm to calibrate the RGB left and RGB right views onto the same plane, obtaining the corrected RGB left and RGB right views, as well as new intrinsic and extrinsic parameters, including the second focal length, the second optical center, and the second translation vector in the x and y directions of the image coordinate system; Depth estimation is performed using the FoundationStereo network proposed by NVIDIA to read the corrected RGB left and right views, and the depth information corresponding to the corrected RGB left view is saved as depth.npy; Figure 5 The FoundationStereo network depth estimation process is as follows: The corrected RGB left view and the corrected RGB right view are processed by the Side-Tuning Adapter (STA) network to obtain deep features and contextual features: The corrected RGB left and right views are input into a Side-Tuning CNN and a pre-trained DepthAnythingV2 network, respectively. Through feature fusion, deep features that combine monocular priors and multi-scale spatial details are extracted, enhancing feature representation capabilities. Specifically, the corrected RGB left and right views are first processed by a Side-Tuning CNN with shared weights to extract multi-scale spatial features, and then downsampled at a 1 / 4 scale through convolution. These features are then fused with the high semantic and rich geometric priors output by DepthAnythingV2 to obtain deep features.

[0033] The corrected RGB left view is then input into a context network to obtain context features.

[0034] The deep features of the corrected RGB left and right views at the 1 / 4 scale are processed in two ways: Method 1: Group-wise correlation is used to calculate the group correlation volume, which characterizes the matching similarity of binocular features across different channel groups. Method 2: The deep features of the corrected RGB left and right views are directly stitched together along the channel dimension under the assumption of disparity shift, resulting in a stitched feature volume to preserve rich monocular prior and spatial detail information. Subsequently, the group correlation volume and the stitched feature volume are merged along the channel dimension to construct a 4D hybrid cost volume. .

[0035] Attentive Hybrid Cost Filtering (AHCF) network for 4D hybrid cost volumes The AHCF network, which performs processing, consists of two sub-networks: Axial-Planar Convolution (APC) and Disparity Transformer (DT). Their functions are as follows: The APC network decomposes 3D convolution into separate convolutions in two directions: spatial and parallax, effectively expanding the receptive field and achieving efficient feature aggregation.

[0036] The DT network models the 4D hybrid cost volume using a global self-attention mechanism along the disparity dimension, further enhancing the global expressive power of the features. The outputs of APC and DT are then added and fused after being upsampled back to the original 4D hybrid cost volume scale.

[0037] For the 4D mixture cost volume output by AHCF, the initial disparity estimation results are obtained using the soft-argmin method: in, Indicates parallax. Indicates the initial disparity. This represents the maximum value of the parallax. This represents the hybrid cost body after passing through the APC network.

[0038] Based on the initial disparity, a multi-step recursive update is performed through a convolutionally gated recurrent single ConvGRU recursive refinement network, with each step of refinement as follows: First, voxel lookup is performed using the latest parallax, and features are retrieved from the 4D hybrid cost volume and the grouped correlation volume, respectively. The concatenated features are then input into the convolutional gated recurrent unit (GRU) network. The GRU network simultaneously integrates contextual features, recursively refines disparity estimation, avoids getting trapped in local optima, and improves global consistency.

[0039] A multi-level GRU network recursively refines the output to achieve the final high-precision parallax. .

[0040] Depth can be obtained based on the triangulation principle of binocular cameras: in, Indicates depth, Indicates the baseline distance of the binocular cameras. Indicates the focal length of a binocular camera. Indicates parallax.

[0041] SAM2-based automated segmentation Target detection A target detection dataset containing two types of conductors (iced conductors and normal conductors) was constructed. This embodiment contains a total of 2061 images, including 1036 images of iced conductors and 1025 images of normal conductors. This scheme uses YOLO 11 as the pre-training algorithm for the object detection dataset. The object detection dataset is input into YOLO 11 to obtain the wire detection model, whose network structure mainly consists of three parts: Backbone: Used to extract multi-level image features from input images. YOLO 11's Backbone uses convolutional (Conv) and feature aggregation block (C3k2) networks for feature extraction, and finally uses a spatial pyramid pooling fast structure (SPPF) network to further fuse information from different receptive fields, improving the ability to recognize multi-scale targets.

[0042] Neck: This component performs multi-scale feature fusion to enhance the detection network's ability to perceive targets of different sizes. YOLO 11's Neck employs upsampling, feature concatenation, and a multi-level C3k2 network, combined with a C2PSA attention mechanism, to construct rich feature representations. The Neck network effectively transfers and integrates features from different levels, contributing to improved overall detection accuracy and robustness.

[0043] Head: Used to output detection results, including class prediction and bounding box regression. YOLO 11 employs multiple V11Detect heads to achieve target detection at different scales, which is helpful for detecting both small and large targets. Each detection head can perform classification and localization prediction on the output of different feature layers, ensuring the comprehensiveness and accuracy of the detection results.

[0044] The corrected RGB left view is input into the conductor detection model to obtain the detection box and confidence level of the icy conductor.

[0045] SAM2 semantic segmentation Using the detection box with the highest confidence as a cue, the cue and the corrected RGB left view are input into the SegmentAnything Model 2 (SAM2) network for accurate segmentation to obtain the mask of the icy wire.

[0046] Segment Anything Model 2 (SAM2) is a next-generation general-purpose segmentation model from Meta. Building upon the original SAM, it achieves unified segmentation of images and videos, significantly improving accuracy and interactivity. SAM2 introduces an innovative streaming memory mechanism, enabling it to continuously track and segment targets in videos, maintaining segmentation consistency even when objects are occluded or moved. Users can interactively segment any frame using simple prompts such as dots, bounding boxes, or masks, and SAM2 automatically performs consistent object tracking and segmentation in subsequent frames. Trained on the massive SA-V video segmentation dataset, the model boasts powerful zero-shot generalization capabilities, adapting to unfamiliar objects without retraining. SAM2 not only possesses real-time inference capabilities, several times faster than its predecessor, but is also widely applicable to scenarios requiring efficient segmentation and object tracking in video editing, augmented reality, autonomous driving, and medical imaging, making it one of the most representative foundational models in the field of visual segmentation.

[0047] Image processing Remove small noises Assume the mask for the icing conductor is a two-dimensional array. , Describe the height of the image. Indicates the width of the image. Indicates the first Line 1 Column pixels belong to the target. Indicates the first Line 1 The column pixels belong to the background; The mask of the icy conductor is labeled with connected regions in 4-pixel neighborhoods, and the label matrix is ​​as follows: Each connected target is assigned a unique integer label; Count all connected regions The total number of pixels, denoted as : Let the area threshold be (In this embodiment, it is set to 100), then for all Set the entire area to background 0: Processed icing wire mask Only targets with an area greater than or equal to the threshold are retained.

[0048] Morphological operations Let the structuring element (convolution kernel) be... , It is a 3×3 all-1 square matrix: First, process the icing-covered wire mask. An opening operation is performed to remove small foreground noise and burrs. The process involves erosion followed by dilation, denoted as: in, It represents corrosion (Erosion). Represents dilation. The corrosion formula is: The expansion formula is: After the re-opening operation, the icing wire mask is processed. To perform a closing operation to fill small holes in the foreground, the process involves expansion followed by corrosion, denoted as: Calculation of icing area The mask obtained after the closing operation Perform skeletonization operations using a medial axis transform to obtain the target backbone pixel set: in The first part represents the skeleton. pixel coordinates, N This represents the total number of sampling points for the skeleton.

[0049] For the target main pixel set S Every point in Select the front and back points within the local window. , Fit the tangent direction and calculate the normal direction angle: Therefore, the unit vector of the normal direction is obtained: Based on skeleton points Centered on the mask, sample along the positive and negative normal directions until the mask boundary is encountered, thus obtaining the left and right endpoints. , ; During sampling, using the initial point as a reference, retain points with depth changes within a 10cm range, and delete the rest; Depth value at the endpoint , This is denoted as the median of the depth information during the sampling process; Endpoint pixel coordinates Using a pinhole camera model and a new intrinsic parameter (second focal length) Second focal length Second optical center ), based on endpoint depth Calculate its three-dimensional coordinates: in, The three-dimensional coordinates of the endpoints.

[0050] Calculate the first The spatial diameter of the cross section is obtained by the Euclidean distance between the endpoints of the normal line segment at each point. : Set the diameter range [ , ], retain the diameters belonging to this interval, and filter out the rest. After filtering, there are a total of Bar diameter; average diameter Represented as: No. The center point of each cross section The three-dimensional coordinates of the two endpoints of the spatial diameter of this cross section and The arithmetic mean is determined as follows: To reduce accumulated errors and computational redundancy, the sampling rate (step size rate) is defined as follows: .

[0051] The total length L of the object is defined as: Among them, the remaining items This indicates that if the distance from the last valid sampling point to the end point is less than one sampling rate, the compensation is applied.

[0052] For each length interval Let the actual length of this segment be: This segment contains sr local cross-sectional diameters. .

[0053] Define the representative diameter of this segment as the arithmetic mean: The radius of this segment is: The volume of the cylinder in that section is: Total volume The sum of the volumes of all the smaller cylinders: Among them, the remaining items This indicates that if the tail end is less than one sampling rate, volume compensation is performed from the last valid sampling point to the end point. The volume of ice accumulation is the total volume minus the volume of the corresponding portion of the conductor: in, Given the diameter of the wire, This represents the total length of the measured conductor.

[0054] A system for measuring the ice volume of a power transmission line, comprising: The depth estimation module acquires the intrinsic and extrinsic parameters of the binocular camera, uses the binocular camera to capture a view of the icy transmission line, corrects the transmission line view, acquires new intrinsic and extrinsic parameters, and performs depth estimation on the corrected transmission line view based on the corrected transmission line view and the new intrinsic and extrinsic parameters. The mask generation and processing module trains a conductor detection model using YOLO 11 based on the target detection dataset. The corrected transmission line view is then input into the conductor detection model to obtain the detection boxes and confidence scores of the icy conductors. The detection box with the highest confidence score is used as a cue. The cue and the corrected transmission line view are then input into the SegmentAnything Model 2 (SAM2) network for accurate segmentation to obtain the mask of the icy conductor. The mask of the icy conductor is then subjected to denoising and morphological processing. The icing volume calculation module calculates the icing zone of the transmission line based on the mask of the iced conductor after noise reduction and morphological processing, and the depth estimation of the corrected transmission line view.

[0055] A device for measuring the icing volume of a power transmission line, comprising: Memory: Used to store the computer program that implements the method for measuring the ice volume of transmission lines; Processor: Used to implement the method for measuring the ice volume of transmission lines when executing the computer program.

[0056] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of a method for measuring the ice volume of transmission lines.

[0057] A computer program product includes a computer program that, when executed by a processor, implements a method for measuring the ice volume of transmission lines.

[0058] Experimental Analysis Steel-cored aluminum stranded wire was placed in a long polyethylene (PE) plastic bag, a suitable amount of water was added, the bag was sealed, and then frozen in a freezer. A binocular camera system consisting of two cameras was used in the experiment. The camera focal length was adjusted at a distance of approximately 6 meters from the binocular cameras to obtain suitable imaging results. First, the two cameras of the binocular camera were calibrated using a checkerboard method. Then, the pre-frozen ice was placed 6 meters in front of the cameras for imaging, and corresponding measurements were taken.

[0059] like Figure 6As shown in the figure, this is a visualization of the point cloud of the icing conductor to be measured. The conductor diameter is known to be 3 cm. The ice diameter was measured 10 times using calipers, and the average value was 5.85 cm. The ice length was measured to be 51 cm using a measuring tape. Substituting these values ​​into the formula, the total volume is calculated to be 1370.79 mL, the conductor volume is 360 mL, and the actual ice volume is 1010.79 mL. Using the method of this invention, the average diameter is calculated to be 5.91 cm, the length 51.5 cm, the total volume 1412.77 mL, and the ice volume 1052.77 mL. Therefore, the measurement error of the ice volume is 4.15%. In comparison, Dong Weifeng mentioned in "Research on Equivalent Icing Detection Technology for Transmission Lines Based on Deep Learning" that the error in the ice volume of regular-shaped lines is 6%. The measurement accuracy of this method is significantly better than the results of that study.

Claims

1. A method for measuring the ice volume of a power transmission line, characterized in that, Includes the following steps: Obtain the intrinsic and extrinsic parameters of the binocular camera, use the binocular camera to capture a view of the icy transmission line, correct the transmission line view, and obtain new intrinsic and extrinsic parameters. Based on the corrected transmission line view and the new intrinsic and extrinsic parameters, perform depth estimation on the corrected transmission line view. Based on the object detection dataset, a conductor detection model was trained using YOLO 11. The corrected transmission line view was then input into the conductor detection model to obtain the detection boxes and confidence scores of the iced conductors. The detection box with the highest confidence score was used as a cue. The cue and the corrected transmission line view were then input into the Segment Anything Model 2 network for accurate segmentation to obtain the mask of the iced conductor. The mask of the iced conductor was then denoised and morphologically processed. The icing volume of the transmission line is calculated based on the mask of the iced conductor after denoising and morphological processing and the depth estimation of the corrected transmission line view.

2. The method according to claim 1, characterized in that, The intrinsic and extrinsic parameters of the binocular camera specifically include the first focal length, the first optical center, the first rotation matrix, the first translation vector, and the distortion coefficients in the x-axis and y-axis directions of the image coordinate system.

3. The method according to claim 1, characterized in that, The specific steps for capturing and correcting views of icy power transmission lines using a binocular camera are as follows: Using a binocular camera, images of ice-covered power transmission lines are captured to obtain RGB left and RGB right views of the transmission lines. The intrinsic and extrinsic parameters are input into the OpenCV stereo calibration algorithm to calibrate the RGB left and RGB right views onto the same plane, resulting in calibrated RGB left and RGB right views. The depth information corresponding to the calibrated RGB left view is saved as an npy format file. The new intrinsic and extrinsic parameters include the second focal length, the second optical center, and the second translation vector in the x-axis and y-axis directions of the image coordinate system.

4. The method according to claim 1, characterized in that, The process of depth estimation of the corrected transmission line view based on the corrected transmission line view and new intrinsic and extrinsic parameters is as follows: The corrected RGB left view and corrected RGB right view are processed by the Side-Tuning Adapter network of the FoundationStereo network to obtain deep features and contextual features: The corrected RGB left and right views are input into a Side-Tuning CNN and a pre-trained DepthAnythingV2 network, respectively. Through feature fusion, deep features that combine monocular priors and multi-scale spatial details are extracted, enhancing the feature representation capability. Specifically, the corrected RGB left and right views are first processed by a Side-Tuning CNN with shared weights to extract multi-scale spatial features, and then downsampled at a 1 / 4 scale through convolution. These features are then fused with the high semantic and rich geometric priors output by DepthAnythingV2 to obtain deep features. The corrected RGB left view is further input into a context network to obtain context features; The deep features of the corrected RGB left and right views at the 1 / 4 scale are processed in two ways: Method 1, a group correlation volume is obtained through group correlation calculation to characterize the matching similarity of binocular features across different channel groups; Method 2, the deep features of the corrected RGB left and right views are directly stitched together in the channel dimension under the assumption of disparity shift to obtain a stitched feature volume, thus preserving rich monocular prior and spatial detail information. Subsequently, the group correlation volume and the stitched feature volume are merged in the channel dimension to construct a 4D hybrid cost volume. ; Attentive Hybrid Cost Filtering Network for 4D Hybrid Cost Volumes The AHCF network, which performs processing, consists of two sub-networks: Axial-Planar Convolution and Disparity Transformer. Their functions are as follows: The APC network decomposes 3D convolution into separate convolutions in two directions: spatial and parallax. The DT network models the 4D hybrid cost volume using a global self-attention mechanism in the parallax dimension; the outputs of APC and DT are upsampled to restore the original 4D hybrid cost volume scale and then added and fused. For the 4D mixture cost volume output by AHCF, the initial disparity estimation results are obtained using the soft-argmin method: in, Indicates parallax. Indicates the initial disparity. This represents the maximum value of the parallax. This represents the hybrid cost volume after passing through the APC network; Based on the initial disparity, the network is recursively updated in multiple steps using ConvGRU (Convolutionally Gated Recurrent Unit) recursive refinement. Each refinement step is as follows: First, voxel lookup is performed using the latest parallax, and features are retrieved from the 4D hybrid cost volume and the grouped correlation volume, respectively. The concatenated features are then input into the convolutional gated recurrent unit (GRU) network. The GRU network simultaneously integrates contextual features and recursively refines disparity estimation. A multi-level GRU network recursively refines the output to achieve the final high-precision parallax. ; The depth is obtained based on the triangulation principle of a binocular camera: in, Indicates depth, Indicates the baseline distance of the binocular cameras. Indicates the focal length of a binocular camera. Indicates parallax.

5. The method according to claim 1, characterized in that, The method involves training a conductor detection model using YOLO 11 based on the target detection dataset. The corrected transmission line view is then input into the model to obtain the detection boxes and their confidence scores for iced conductors. The detection box with the highest confidence score is used as a cue. The cue and the corrected transmission line view are then input into the Segment Anything Model 2 network for precise segmentation, resulting in a mask for the iced conductor. The specific steps for denoising and morphological processing of the mask for the iced conductor are as follows: A target detection dataset was constructed using multiple images. The target detection dataset includes icy guide lines and normal guide lines, with icy guide lines accounting for 45-55% and the remainder being normal guide lines. YOLO 11 was used as the pre-training algorithm for the object detection dataset. The object detection dataset was input into YOLO 11 to obtain the wire detection model. The corrected RGB left view is input into the conductor detection model to obtain the detection box and confidence level of the icy conductor. Using the detection box with the highest confidence as a cue, the cue and the corrected RGB left view are input into the SegmentAnything Model 2 network for accurate segmentation to obtain the mask of the icy wire; First, remove minor noise: Assume the mask for the icing conductor is a two-dimensional array. , Describe the height of the image. Indicates the width of the image. Indicates the first Line number Column pixels belong to the target. Indicates the first Line number The column pixels belong to the background; The mask of the icy conductor is labeled with connected regions in 4-pixel neighborhoods, and the label matrix is ​​as follows: Each connected target is assigned a unique integer label; Count all connected regions The total number of pixels, denoted as : Let the area threshold be Then for all Set the entire area to background 0: Processed icing wire mask Only targets with an area greater than or equal to the threshold are retained; Then perform morphological operations: Let the structural element be , It is a 3×3 all-1 square matrix: First, process the icing-covered wire mask. An opening operation is performed to remove small foreground noise and burrs. The process involves erosion followed by dilation, denoted as: in, Represents corrosion. Represents expansion The corrosion formula is: The expansion formula is: After the re-opening operation, the icing wire mask is processed. To perform a closing operation to fill small holes in the foreground, the process involves expansion followed by corrosion, denoted as: 。 6. The method according to claim 1, characterized in that, The specific steps for calculating the icing volume of the transmission line based on the mask of the iced conductor after noise reduction and morphological processing, and the depth estimation of the corrected transmission line view are as follows: The mask obtained after the closing operation Perform skeletonization operations, using median transformation, to obtain the target backbone pixel set: in, The first part represents the skeleton. pixel coordinates, N This represents the total number of sampling points for the skeleton. For the target main pixel set S Every point in Select the front and back points within the local window. , Fit the tangent direction and calculate the normal direction angle: Therefore, the unit vector of the normal direction is obtained: Based on skeleton points Centered on the mask, sample along the positive and negative normal directions until the mask boundary is encountered, thus obtaining the left and right endpoints. , ; Depth value at the endpoint , This is denoted as the median of the depth information during the sampling process; Endpoint pixel coordinates Using a pinhole camera model, second focal length Second focal length and the second optical center According to the endpoint depth Calculate its three-dimensional coordinates: in, The three-dimensional coordinates of the endpoints; Calculate the first The spatial diameter of the cross section is obtained by the Euclidean distance between the endpoints of the normal line segment at each point. : average diameter Represented as: No. The center point of each cross section The three-dimensional coordinates of the two endpoints of the spatial diameter of this cross section and The arithmetic mean is determined as follows: To reduce accumulated errors and computational redundancy, the sampling factor is defined as... ; Then the total length of the object Defined as: Among them, the remaining items This indicates that if the tail end is less than one sampling rate, distance compensation is performed from the last valid sampling point to the end point. For each length interval Let the actual length of this segment be: This segment contains sr local cross-sectional diameters. ; Define the representative diameter of this segment as the arithmetic mean: The radius of this segment is: The volume of the cylinder in that section is: Total volume The sum of the volumes of all the smaller cylinders: Among them, the remaining items This indicates that if the tail end is less than one sampling rate, volume compensation is performed from the last valid sampling point to the end point. The volume of ice accumulation is the total volume minus the volume of the corresponding portion of the conductor: in, Given the diameter of the wire, This represents the total length of the measured conductor.

7. A transmission line icing volume measurement system based on the method of any one of claims 1 to 6, characterized in that, include: The depth estimation module acquires the intrinsic and extrinsic parameters of the binocular camera, uses the binocular camera to capture a view of the icy transmission line, corrects the transmission line view, acquires new intrinsic and extrinsic parameters, and performs depth estimation on the corrected transmission line view based on the corrected transmission line view and the new intrinsic and extrinsic parameters. The mask generation and processing module trains a conductor detection model using YOLO 11 based on the target detection dataset. The corrected transmission line view is then input into the conductor detection model to obtain the detection boxes and confidence scores of the icy conductors. The detection box with the highest confidence score is used as a cue. The cue and the corrected transmission line view are then input into the Segment AnythingModel 2 network for accurate segmentation to obtain the mask of the icy conductor. The mask of the icy conductor is then denoised and morphologically processed. The icing volume calculation module calculates the icing zone of the transmission line based on the mask of the iced conductor after noise reduction and morphological processing, and the depth estimation of the corrected transmission line view.

8. A device for measuring the ice volume of a power transmission line, characterized in that, include: Memory: for storing a computer program that implements the method for measuring the ice volume of transmission lines as described in any one of claims 1 to 6; Processor: Used to implement the method for measuring the ice volume of transmission lines as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for measuring the ice volume of transmission lines as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for measuring the ice volume of transmission lines as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Power transmission line icing thickness detection method based on unmanned aerial vehicle image processing

    CN114677428A

  • Automatic icing power transmission line thickness detection method

    CN118052760A

Cited By

  • Method, system, medium and device for calculating ice thickness of power transmission line

    CN122134787A

  • Method, system, medium and device for calculating ice thickness of power transmission line

    CN122134787B