A power transmission line safety hazard detection method and system
By constructing a three-dimensional ranging benchmark system and an uncalibrated monocular vision three-dimensional ranging model, the safe distance between potential hazards and conductors can be accurately calculated using existing monocular cameras. This solves the problem of the inability to quantify potential hazard risks in existing technologies and enables precise early warning and handling of transmission lines.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- POWERCHINA JIANGXI ELECTRIC POWER ENGINEERING CO LTD
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing power transmission line hazard monitoring technologies based on fixed cameras on poles lack three-dimensional depth information, making it impossible to accurately calculate the actual safe distance between the hazard and the conductor, or quantify the risk level of the hazard. This makes it difficult to meet the actual needs of power operation and maintenance for accurate early warning and precise handling of hazards.
By acquiring time-series images of transmission lines captured by a monocular camera, the preset marking dimensions of the towers, the geometric parameters of the conductor catenary, and the camera installation and fixing parameters are extracted to construct a three-dimensional ranging benchmark system. Combined with a lightweight convolutional neural network and geometric projection, an uncalibrated monocular vision three-dimensional ranging model is constructed to calculate the actual vertical and horizontal distances between potential hazards and the conductors.
It enables precise 3D ranging and risk quantification assessment using existing monocular cameras without increasing additional hardware costs, solving the problem of the inability to quantify hidden risks in existing technologies and meeting the needs of precise early warning and handling in power operation and maintenance.
Smart Images

Figure CN122435547A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of line testing technology, and in particular to a method and system for detecting safety hazards in power transmission lines. Background Technology
[0002] High-voltage transmission lines are the core backbone of the power system, and their safe and stable operation is directly related to public safety. Therefore, accurate monitoring and timely early warning of potential hazards around transmission lines are crucial to ensuring reliable power supply. With the deep integration of artificial intelligence and power monitoring technologies, transmission line hazard monitoring based on fixed cameras on power poles has become one of the mainstream online monitoring methods currently available.
[0003] In existing technologies, this type of monitoring system typically installs bullet cameras on the top of transmission line towers, deployed along the transmission line direction. It transmits a sequence of on-site images captured at intervals by the cameras via wireless communication to a remote server. Then, it uses manual image analysis or intelligent image recognition technology to detect potential hazards in the images and outputs monitoring results indicating the presence and type of hazard. This monitoring method eliminates the need for personnel to carry equipment on-site, enabling routine monitoring of transmission lines over a large area. It effectively overcomes the shortcomings of traditional manual, robotic, and drone inspections, which rely on manpower, are inefficient, and lack timely coverage. Therefore, it has been widely adopted in the field of transmission line hazard monitoring.
[0004] However, the current hazard monitoring technology based on fixed cameras on power poles still has a core and unresolved technical defect: most of the fixed cameras used for monitoring are monocular bullet cameras, which can only collect two-dimensional images of power transmission lines and the surrounding environment. They lack three-dimensional depth information, cannot accurately calculate the actual safe distance between the hazard and the power transmission line, and cannot quantify the key dimensional parameters of the hazard itself (such as the height of trees, the extension length of construction robotic arms, etc.).
[0005] In actual power transmission line operation and maintenance, the core criterion for determining whether a potential hazard poses a threat to line safety is whether the distance between the hazard and the conductor is less than a safety threshold, rather than simply determining whether the hazard exists. Current technologies, lacking three-dimensional depth information, rely on two-dimensional images to vaguely assess the distance between the hazard and the conductor, unable to provide specific distance values or quantify the hazard's risk level. For example, regarding slow-growing trees, only their presence can be identified, but the annual growth rate, current height, and actual vertical and horizontal distance from the conductor cannot be accurately calculated. This makes it difficult to predict whether the tree will exceed the safety distance and cause short circuits or other accidents. Similarly, for construction machinery near the line, the real-time distance between the robotic arm and the conductor cannot be accurately calculated, hindering accurate early warning and potentially leading to accidents caused by machinery touching the conductor.
[0006] To address the aforementioned 3D distance measurement problem, some existing technologies offer improvements, such as adding additional 3D measurement equipment like binocular cameras and LiDAR to the tower, enabling distance measurement through multi-device collaboration. However, these solutions have significant limitations: Firstly, binocular cameras and LiDAR are expensive and require integration with existing monocular cameras, necessitating large-scale hardware modifications to the existing monitoring system. This not only increases the costs of equipment purchase, installation, and maintenance but also makes it difficult to achieve compatibility with existing monocular camera monitoring equipment, hindering implementation. Secondly, the additional 3D measurement equipment consumes significant power and is incompatible with the power supply system of the existing monocular cameras, potentially leading to insufficient power and affecting the continuous 24 / 7 operation of the monitoring system.
[0007] In addition, some existing technologies attempt to obtain three-dimensional information through drone inspections and gimbal multi-angle shooting. However, drone inspections cannot achieve routine, all-weather monitoring, are greatly limited by factors such as weather and battery life, and rely on manual operation. Gimbal multi-angle shooting requires additional gimbal equipment, which also increases costs and power consumption, and cannot achieve continuous monitoring of fixed points, making it difficult to meet the actual needs of all-weather, full-coverage, and high-precision hidden danger monitoring of power transmission lines.
[0008] In summary, existing transmission line hazard monitoring technologies based on fixed monocular cameras on poles lack three-dimensional depth information, making it impossible to accurately determine the safe distance between hazards and transmission lines. These technologies can only achieve qualitative identification of hazards and cannot complete quantitative risk assessment, thus failing to meet the actual needs of power operation and maintenance for accurate early warning and handling of hazards. Summary of the Invention
[0009] In view of this, the purpose of the present invention is to provide a method and system for detecting safety hazards in transmission lines, which aims to solve the problem that the existing technology cannot complete the quantitative assessment of hazard risks during the detection of safety hazards in transmission lines, and is difficult to meet the actual needs of power operation and maintenance for accurate early warning and accurate handling of hazards.
[0010] This invention provides a method for detecting safety hazards in power transmission lines, the method comprising: A sequence of time-series images of the power transmission line site captured at intervals by a monocular camera installed on the tower is obtained, and the time-series image sequence is preprocessed to obtain the target monitoring image; Feature extraction is performed on the target monitoring image to obtain the inherent reference features of the transmission line. The inherent reference features include the preset identification size of the tower, the geometric parameters of the catenary of the conductor, and the installation and fixing parameters of the camera. A three-dimensional ranging benchmark system is constructed by correlating and calibrating the actual physical size with the corresponding image pixel size based on inherent benchmark features. A pre-trained target detection model is used to identify potential hazards in the target monitoring image, locate the potential hazard target area, and extract the two-dimensional image features of the potential hazard target area; Based on the aforementioned three-dimensional ranging benchmark system, and combined with the two-dimensional image features of the target area of the potential hazard, an uncalibrated monocular vision three-dimensional ranging model is constructed. Through the mapping relationship between image pixel scale and actual physical scale, the actual vertical and horizontal distances between the potential hazard target and the power transmission line are calculated. Based on the safety distance threshold standard for transmission lines, the actual vertical and horizontal distances are assessed for risk, and the hazard safety distance value and corresponding risk level are output.
[0011] Furthermore, in the above-mentioned method for detecting safety hazards in transmission lines, the step of extracting features from the target monitoring image to obtain the inherent baseline features of the transmission line includes: An algorithm combining template matching and shape context descriptors is used to locate the pre-defined marking area of the tower in the target monitoring image, and the marking outline is extracted by a sub-pixel edge detection algorithm to calculate the marking center coordinates and pixel size; The edge point set of the conductor is extracted by using the Canny operator combined with nonmaximum suppression, and the sag, span, and tangent slope of the conductor are calculated by fitting the catenary equation of the conductor based on the least squares method. Based on the ratio of the actual physical size of the tower marker to the image pixel size, and combined with the geometric parameters of the catenary, the camera installation tilt angle, focal length, and pose parameters relative to the tower coordinate system are inverted through the perspective projection transformation model. The extracted marker dimensions, catenary parameters, and inverted camera parameters are subjected to consistency verification, abnormal feature values are removed, and a fused transmission line inherent reference feature vector is generated.
[0012] Furthermore, in the aforementioned method for detecting safety hazards in transmission lines, the step of correlating and calibrating the actual physical dimensions based on inherent reference features with the corresponding image pixel dimensions to construct a three-dimensional ranging reference system includes: Based on the actual height of the pole / tower markings With pixel height in the image Calculate the longitudinal scale factor ; Using the catenary fitting results of the conductor, a two-dimensional image coordinate system with the bottom of the tower as the origin is constructed in the image, and the key points of the conductor are projected onto this coordinate system; Based on the camera installation height and tilt angle parameters, a ground plane equation is constructed. The projection lines in the image serve as a depth reference plane; Combined with scale factor Image coordinate system and depth reference plane, constructing from image pixel coordinates to actual physical coordinates mapping function ,in The direction is constrained by the direction of gravity, forming a three-dimensional ranging reference system with physical significance.
[0013] Furthermore, in the aforementioned method for detecting safety hazards in transmission lines, the step of constructing an uncalibrated monocular vision three-dimensional ranging model based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the hazard target area includes: A lightweight convolutional neural network is used to extract two-dimensional image feature maps of the target area of potential hazards. It includes the target bounding box, center point coordinates, and contour features; center point of the target Projecting onto the ground plane in the three-dimensional ranging reference system yields preliminary ground projection points. ; Based on the relative vertical pixel distance between the bottom of the target and the guide wire in the image By combining the geometric parameters of the catenary and the camera tilt angle, the actual height of the target is calculated using the principle of triangulation. ,in, The camera's tilt angle; By integrating deep learning features with geometric projection results, an end-to-end calibration-free monocular vision 3D ranging model is constructed. Output the three-dimensional coordinates of the target in physical space. .
[0014] Furthermore, in the aforementioned method for detecting safety hazards in power transmission lines, the step of calculating the actual vertical and horizontal distances between the hazard target and the power transmission line by mapping the image pixel scale to the actual physical scale includes: Based on catenary parameters and a three-dimensional ranging benchmark system, the three-dimensional spatial curve of the transmission line is reconstructed in the area where the potential hazard is located. ; Three-dimensional coordinates of the potential hazard target With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. ; Calculate the target projection point With point Projected distance on the horizontal plane ; Calculate target height With point height absolute value of the difference ; Output the actual horizontal distance between the potential hazard target and the power transmission line. and vertical distance .
[0015] Furthermore, in the aforementioned method for detecting safety hazards in transmission lines, the combination of scale factors... Image coordinate system and depth reference plane, constructing from image pixel coordinates to actual physical coordinates mapping function The steps include: Select at least four non-collinear corner points of the tower markings in the image as control points, and construct the homography matrix between the two-dimensional image plane and the three-dimensional physical space ground plane by combining their corresponding actual physical coordinates. The homography matrix is normalized using the longitudinal scale factor to eliminate scale blur caused by the unknown focal length of the camera. Based on the normalized homography matrix, a linear mapping relationship between image pixel coordinates and ground plane physical coordinates is established. By introducing the slope of the catenary tangent as a nonlinear correction factor in the vertical direction, distortion compensation is performed on the linear mapping relationship, resulting in a complete mapping function containing both horizontal and vertical dimensions.
[0016] Furthermore, in the aforementioned method for detecting safety hazards in transmission lines, the step of fusing deep learning features and geometric projection results to construct an end-to-end uncalibrated monocular vision 3D ranging model includes: A multimodal feature fusion network is constructed, which includes a geometric prior branch and a visual feature branch; The coordinates of the preliminary ground projection points and the target height in the geometric projection results are used as the input of the geometric prior branch, and the two-dimensional image feature map extracted by the lightweight convolutional neural network is used as the input of the visual feature branch. In the feature fusion layer, a channel attention mechanism is used to weight the visual feature branches, and the output of the geometric prior branch is used to constrain the spatial position of the weighted visual features. The fused features constrained by spatial location are input into the fully connected regression layer, and the depth-corrected three-dimensional coordinates of the target in physical space are output.
[0017] Furthermore, in the aforementioned method for detecting safety hazards in transmission lines, the step of determining the three-dimensional coordinates of the hazard target... With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. The steps include: A discretization sampling strategy is adopted to generate several spatial sampling points along the reconstructed three-dimensional spatial curve of the transmission line at a preset step size; Construct a spatial search sphere centered on the three-dimensional coordinates of the potential hazard target, and set an initial search radius; Calculate the Euclidean distance between each spatial sampling point and the three-dimensional coordinates of the potential hazard target, and filter out the set of candidate sampling points that fall within the spatial search sphere; If the candidate sampling point set is not empty, then the point with the smallest Euclidean distance in the candidate sampling point set is selected as the nearest Euclidean distance point. If the candidate sampling point set is empty, increase the search radius and repeat the sampling and filtering steps until the nearest Euclidean distance point is found. Furthermore, in the aforementioned method for detecting safety hazards in transmission lines, the step of using the output of geometric prior branches to constrain the spatial position of the weighted visual features includes: The initial ground projection point coordinates output by the geometric prior branch are converted into a Gaussian heatmap distribution, where the center of the Gaussian heatmap distribution corresponds to the initial ground projection point coordinates and the variance corresponds to the uncertainty of the geometric projection. The Gaussian heatmap distribution is upsampled to the same size as the visual feature map to generate a spatial attention mask; The spatial attention mask is multiplied element-wise with the weighted visual features to suppress redundant feature responses that deviate from the geometric prior position and to enhance the feature weights of the target's true landing area.
[0018] Another object of the present invention is to provide a power transmission line safety hazard detection system, the system comprising: The acquisition module is used to acquire a sequence of time-series images of the power transmission line at intervals captured by a monocular camera set on the tower, and to preprocess the time-series image sequence to obtain the target monitoring image; The extraction module is used to extract features from the target monitoring image to obtain the inherent reference features of the transmission line. The inherent reference features include the preset identification size of the tower, the geometric parameters of the conductor catenary, and the installation and fixing parameters of the camera. The module is used to correlate and calibrate the actual physical size and corresponding image pixel size based on the inherent reference features to build a three-dimensional ranging reference system; The localization module is used to identify potential hazards in the target monitoring image using a pre-trained target detection model, locate the potential hazard target area, and extract the two-dimensional image features of the potential hazard target area. The mapping module is used to construct an uncalibrated monocular vision three-dimensional ranging model based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the target area of the hidden danger. Through the mapping relationship between the image pixel scale and the actual physical scale, the actual vertical and horizontal distances between the target danger and the transmission line are calculated. The identification module is used to determine the risk of the actual vertical and horizontal distances by combining the safety distance threshold standard for transmission lines, and output the safety distance value of the hidden danger and the corresponding risk level.
[0019] This invention obtains target monitoring images by acquiring and preprocessing a sequence of time-series images of a power transmission line captured at intervals by a monocular camera mounted on a tower. Feature extraction is performed on these images to obtain inherent reference features of the power transmission line, including the pre-defined dimensions of the tower markers, the geometric parameters of the conductor catenary, and the camera's installation and fixing parameters. A three-dimensional ranging reference system is constructed by correlating and calibrating the actual physical dimensions of the inherent reference features with the corresponding image pixel dimensions. A pre-trained target detection model is used to identify potential hazards in the target monitoring images and extract two-dimensional image features. Based on the three-dimensional ranging reference system and the two-dimensional image features of the potential hazard area, an uncalibrated monocular vision three-dimensional ranging model is constructed. The actual vertical and horizontal distances between the potential hazard and the power transmission conductor are calculated through the mapping relationship between image pixel scale and actual physical scale. Finally, the actual vertical and horizontal distances are assessed using a power transmission line safety distance threshold standard, and the resulting hazard safety distance value and corresponding risk level are output. This approach avoids the limitations of existing technologies, which rely solely on blurry two-dimensional images to determine the distance to potential hazards due to a lack of 3D depth information, thus failing to provide specific distance values and quantified risk levels. It also avoids the high costs, high power consumption, and significant hardware modification challenges associated with adding expensive 3D measurement equipment such as binocular cameras and LiDAR. Ultimately, it achieves precise 3D ranging and risk quantification assessment of power transmission line hazards using existing monocular cameras without incurring additional hardware costs. This solves the problem that existing technologies cannot quantitatively assess the risk of power transmission line safety hazards, failing to meet the practical needs of power operation and maintenance for accurate early warning and handling of hazards. Attached Figure Description
[0020] Figure 1 This is a flowchart of the method for detecting safety hazards in power transmission lines in the first embodiment of the present invention; Figure 2 This is a structural block diagram of the power transmission line safety hazard detection system in the third embodiment of the present invention.
[0021] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0022] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0023] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0025] Example 1 Please see Figure 1 The figure shows a method for detecting safety hazards in power transmission lines according to the first embodiment of the present invention, the method including steps S10 to S15.
[0026] Step S10: Obtain a sequence of time-series images of the power transmission line at intervals captured by a monocular camera set on the tower, and preprocess the time-series image sequence to obtain the target monitoring image.
[0027] First, using monocular cameras pre-installed at the locations of the transmission line towers, images are captured at set time intervals, continuously obtaining a sequence of time-series images containing various on-site scenes, including the transmission line itself, the surrounding ground environment, vegetation, and buildings. These image sequences are arranged chronologically according to the time of capture, providing a complete record of the actual on-site conditions at different monitoring points along the transmission line.
[0028] After completing the large-scale acquisition of time-series raw image data, standardized preprocessing was performed on all acquired raw images. This preprocessing included basic image processing operations such as overall brightness equalization, noise and artifact removal, distortion correction, cropping of invalid edge areas, image sharpness enhancement, and removal of redundant background interference. Through a series of standardized image preprocessing steps, the raw time-series images were uniformly rectified and optimized, ultimately selecting and generating target monitoring images that were clear, had complete scene information, highlighted target areas, and met the standards for subsequent feature extraction and intelligent recognition operations. Simultaneously, irrelevant interference factors such as shooting shake, lighting changes, and environmental clutter in the original acquired images were removed, further improving the accuracy and operational stability of subsequent overall detection work.
[0029] Step S11: Extract features from the target monitoring image to obtain inherent reference features of the transmission line. The inherent reference features include the preset marking size of the tower, the geometric parameters of the conductor catenary, and the installation and fixing parameters of the camera.
[0030] Among them, the inherent reference features of transmission lines are the core foundation for the subsequent construction of a three-dimensional ranging reference system and the realization of three-dimensional ranging. The extraction process must ensure accuracy and stability. In specific implementation, corresponding extraction algorithms are adopted for the three types of inherent reference features to ensure the accuracy of feature extraction. Specifically, by extracting the size parameters of the tower's preset markings, the correspondence between image pixels and actual physical dimensions can be obtained; by extracting the geometric parameters of the conductor's catenary, the spatial morphology of the conductor can be understood; and by extracting the camera's installation and fixing parameters, the camera's shooting angle and spatial position can be clarified. The three types of features together constitute the reference basis for subsequent three-dimensional ranging.
[0031] For example, an algorithm based on template matching and shape context descriptor is used to locate the pre-defined marking area of the tower in the target monitoring image, and the marking outline is extracted by a sub-pixel edge detection algorithm to calculate the marking center coordinates and pixel size; Pre-set pole markers are markers with fixed physical dimensions that are pre-installed on poles. Rectangular markers are typically used, with known height and width parameters, for example, an actual height of 0.5 meters and an actual width of 0.3 meters. First, a template image of the pre-set pole marker is constructed, ensuring its shape and size proportions match the actual marker. Then, a normalized cross-correlation matching algorithm within the template matching algorithm is used to search for the region in the target monitoring image most similar to the template image. This region is identified as the pre-set pole marker area. To improve positioning accuracy, a shape context descriptor is used to optimize the positioning results. The shape context descriptor describes the shape characteristics of the marker by calculating the angle distribution histogram of the marker area's edge points, further confirming the accuracy of the marker area and avoiding positioning errors caused by similar areas in the image. After locating the marker area, a sub-pixel edge detection algorithm is used to extract the marker contour. The Zernike matrix sub-pixel edge detection algorithm can be used. After obtaining the complete contour of the marker through sub-pixel edge detection, the pixel size of the marker is calculated using the minimum bounding rectangle method, and the center coordinates of the marker are calculated using the centroid of the contour. Next, the edge point set of the traverse is extracted using the Canny operator combined with nonmaximum suppression, and the equation of the catenary of the traverse is fitted based on the least squares method. ,in, The catenary coefficient is used to calculate the conductor sag, span, and tangent slope. Under natural gravity, the conductor exhibits a catenary shape; extracting its geometric parameters is fundamental for subsequent 3D ranging and spatial curve reconstruction. In the implementation process, the target monitoring image is first converted to grayscale, then edge detection is performed using the Canny operator. After obtaining the conductor edge point set, the catenary equation is fitted using the least squares method to obtain the catenary coefficient. The optimal solution is obtained by further calculating the sag, span, and tangent slope of the conductor after obtaining the catenary equation: the span is the horizontal distance between two adjacent towers, which is calculated by extracting the coordinates of the suspension points of the conductor at the top of the two towers; the sag is the vertical distance between the lowest point of the catenary and the line connecting the two suspension points; the tangent slope is obtained by differentiating the catenary equation. Next, based on the ratio of the actual physical size of the tower marker to the image pixel size, and combined with the geometric parameters of the catenary, the camera's installation tilt angle, focal length, and pose parameters relative to the tower coordinate system are inverted using a perspective projection transformation model. This perspective projection transformation model describes the mapping relationship between points in three-dimensional physical space and two-dimensional image pixels, expressed as: , in Image pixel coordinates, For three-dimensional physical space coordinates, For camera focal length, The coordinates of the principal point in the image; During implementation, the ratio is first calculated based on the actual height of the tower's preset markings and the pixel height extracted from the image. This ratio is the preliminary proportional relationship between the image pixels and the actual physical size.
[0032] Then, several key points on the catenary of the conductor are selected. The actual three-dimensional coordinates of these key points can be initially estimated based on parameters such as tower height, span, and sag. These key points are then combined with their image pixel coordinates and substituted into the perspective projection transformation model to construct a system of equations. Simultaneously, the camera's installation tilt angle includes the pitch angle θ and azimuth angle φ. The pitch angle is the angle between the camera's optical axis and the horizontal plane, and the azimuth angle is the angle between the projection of the camera's optical axis onto the horizontal plane and the tower axis. These parameters are all included in the pose matrix of the perspective projection transformation model. By solving the system of equations, the camera's focal length is obtained through inversion. Pitch angle Azimuth The camera's position coordinates relative to the tower coordinate system, along with these parameters, together constitute the camera's installation and fixing parameters.
[0033] Step S12: Based on the correlation calibration between the actual physical size of the inherent reference features and the corresponding image pixel size, a three-dimensional ranging reference system is constructed.
[0034] Specifically, firstly, the scale factor is calculated by using the known actual physical dimensions of the tower's pre-marked markings and their pixel dimensions in the target monitoring image, thus achieving a preliminary correlation between pixel scale and physical scale; then, by combining the geometric parameters of the catenary and the camera's installation and fixing parameters, the correspondence between the image coordinate system and the actual physical coordinate system is constructed, the depth reference plane is determined, and finally a complete three-dimensional ranging benchmark system is formed.
[0035] In practice, the actual height of the pole / tower markings will be used. With pixel height in the image Calculate the longitudinal scale factor ; During implementation, the actual height of the pole markers is a pre-known fixed parameter, while the pixel height of the markers in the image is calculated data. It should be noted that the calculation of the longitudinal scale factor must be based on the pre-processed target monitoring image to ensure accurate pixel height extraction. Furthermore, if multiple pole markers exist, multiple values can be calculated and averaged to further improve the reliability of the scale factor.
[0036] Then, using the catenary fitting results, a two-dimensional image coordinate system is constructed in the image with the base of the tower as the origin. The key points of the conductor are projected onto this coordinate system. The construction of the two-dimensional image coordinate system is fundamental to mapping pixel coordinates to physical coordinates. In implementation, the base of the tower is first located in the target monitoring image. The base of the tower is usually the lowest point of the tower in the image, and its coordinates can be determined through edge detection and morphological processing. This location is then set as the origin O(0,0) of the two-dimensional image coordinate system. The horizontal axis of the coordinate system is... The axis is horizontal, with the positive direction to the right; the vertical axis is... The axis, with the vertical direction upwards as the positive direction, can be represented as the pixel coordinates of any point in the image. .
[0037] After constructing the two-dimensional image coordinate system, the extracted key points of the catenary, including the suspension points of the conductors at the top of the two towers and the lowest point of the catenary, are projected onto the two-dimensional image coordinate system according to their actual positions in the image to obtain the pixel coordinates of each key point, which provides a basis for the subsequent construction of the depth reference plane and the establishment of coordinate mapping relationships. Next, based on the camera's installation height and tilt angle parameters, a ground plane equation is constructed. The projection lines in the image serve as the depth reference plane. This depth reference plane is a crucial benchmark for 3D ranging, used to determine the depth information of each pixel in the image. During implementation, the camera's installation height is first determined. The camera's installation height is its actual vertical distance from the ground, which can be calculated using the known height of the tower and the camera's installation position. This is then combined with the inverted camera pitch angle. The angle between the optical axis of the camera and the horizontal plane is The equation of the ground plane is z = 0, representing the ground in the actual physical space. Projecting this plane onto a two-dimensional image coordinate system yields projection lines, the equations of which can be derived using a perspective projection transformation model.
[0038] The specific derivation process is as follows: the three-dimensional physical coordinates of any point on the ground plane are... Substitute into the perspective projection transformation formula ,because The equation of the projection line can be obtained through limit derivation as follows: ,in Principal point of the image coordinate, For camera focal length, This is the camera's tilt angle. This projection line is the projection of the ground plane onto the image, serving as a depth reference plane for subsequent pixel depth information calculations.
[0039] Finally, combining the scale factor Image coordinate system and depth reference plane, constructing from image pixel coordinates to actual physical coordinates mapping function ,in The direction is constrained by the direction of gravity, forming a physically meaningful three-dimensional ranging reference system. The construction of the mapping function is the core of this three-dimensional ranging reference system, aiming to achieve a precise mapping from image pixel coordinates to actual physical coordinates. First, based on the vertical scale factor, the vertical dimension of the image pixels is converted into the actual physical height. Then, combining the two-dimensional image coordinate system and the depth reference plane, the horizontal and vertical physical coordinates of the pixels are determined. Because... The direction is constrained by the direction of gravity, that is... The direction is vertical, perpendicular to the ground, so the mapping relationship can be constrained by the direction of gravity to ensure the rationality of the physical coordinates; the final constructed mapping function can map any pixel in the image. Mapped to actual physical space The coordinates, combined with the depth reference plane and the catenary parameters, can be used to further determine the pixel point. Coordinates (depth information) are used to form a complete three-dimensional ranging benchmark system.
[0040] Additionally, in some optional embodiments of the present invention, the combination of scale factor Image coordinate system and depth reference plane, constructing from image pixel coordinates to actual physical coordinates mapping function The steps include: Select at least four non-collinear corner points of the tower markings in the image as control points, and construct the homography matrix between the two-dimensional image plane and the three-dimensional physical space ground plane by combining their corresponding actual physical coordinates. The homography matrix is normalized using the longitudinal scale factor to eliminate scale blur caused by the unknown focal length of the camera. Based on the normalized homography matrix, a linear mapping relationship between image pixel coordinates and ground plane physical coordinates is established. By introducing the slope of the catenary tangent as a nonlinear correction factor in the vertical direction, distortion compensation is performed on the linear mapping relationship, resulting in a complete mapping function containing both horizontal and vertical dimensions.
[0041] The homography matrix, a core matrix describing the mapping relationship between two planes, has a dimension of 3×3 and enables precise mapping between two-dimensional image coordinates and three-dimensional ground plane coordinates. In implementation, four non-collinear corner points of the pole marker are first selected in the target monitoring image. These non-collinear corner points ensure the uniqueness and validity of the homography matrix; typically, the four vertices of the marker are selected as control points. Then, the actual physical coordinates corresponding to these four corner points are obtained. Since the actual physical dimensions of the pole marker are known, and the marker is fixedly installed on the pole, its actual physical coordinates can be determined based on the pole's position and the marker's installation height. By substituting the coordinates of the four control points into the formula, eight linear equations are constructed, and the values of each element of the homography matrix are obtained by solving them.
[0042] Then, a vertical scaling factor is used to normalize the homography matrix, eliminating scale ambiguity caused by the unknown camera focal length. During the homography matrix calculation, scale ambiguity arises due to the unknown camera focal length, meaning the scale of matrix elements is not unique, affecting mapping accuracy. In implementation, a vertical scaling factor is used to normalize the homography matrix. Normalization is performed by dividing all elements of the homography matrix by... ,in Homography matrix The element in the second row and second column, The vertical scale factor. The normalized homography matrix. This process eliminates scale ambiguity, giving the elements of the homography matrix a clear physical meaning and ensuring the accuracy of the mapping relationship.
[0043] Next, based on the normalized homography matrix, a linear mapping relationship between image pixel coordinates and ground plane physical coordinates is established. The core of this linear mapping relationship is the use of the normalized homography matrix. Image pixel coordinates Actual physical coordinates converted to ground plane During implementation, the image pixel coordinates are... Convert to homogeneous coordinates Substitute into the homography matrix formula The homogeneous coordinates of the physical coordinates of the ground plane are obtained through matrix operations. Then convert to non-homogeneous coordinates This linear mapping relationship allows the coordinates of any pixel in an image to be converted into actual physical coordinates on the ground plane, providing a basis for subsequent depth information calculation.
[0044] Finally, the tangent slope of the catenary is introduced as a nonlinear correction factor in the vertical direction to compensate for distortion in the linear mapping relationship, resulting in a complete mapping function containing both horizontal and vertical dimensions. Due to perspective distortion in monocular camera images, the physical coordinates obtained solely through linear mapping will have certain errors, especially in the vertical direction. Therefore, a nonlinear correction factor is needed for distortion compensation. During implementation, multiple key points on the catenary are selected, and the tangent slope at each key point is calculated. The tangent slope is then adjusted based on the catenary equation. Differentiation yields, i.e. Using the tangent slope as a nonlinearity correction factor in the vertical direction, a correction function is constructed as follows: ,in pixel coordinates in directional deviation, For physical coordinates in The direction correction amount. This correction function is incorporated into the linear mapping relationship, and applied to the calculated physical coordinates. After correction, the complete mapping function is finally obtained. This function can achieve accurate mapping from image pixel coordinates to actual physical coordinates, taking into account the accuracy of both horizontal and vertical dimensions, and providing a reliable mapping basis for the three-dimensional ranging benchmark system.
[0045] Step S13: Use a pre-trained target detection model to identify potential hazards in the target monitoring image, locate the potential hazard target area, and extract the two-dimensional image features of the potential hazard target area.
[0046] The core objective of target detection is to identify potential safety hazards in power transmission lines within images and accurately locate their positions within the images, providing target objects for subsequent 3D ranging. The pre-trained target detection model can utilize lightweight models commonly used in this field, balancing detection accuracy and computational efficiency, such as the YOLOv5s model. This model, trained on a large number of power transmission line hazard samples, can effectively identify common hazard types such as trees, construction machinery, illegal structures, and floating foreign objects. During implementation, the pre-processed target monitoring image is input into the pre-trained YOLOv5s model. The model, through feature extraction, feature fusion, and target prediction, outputs the bounding box coordinates, category probability, and confidence score of the hazard target. When the confidence score is greater than a preset threshold (usually set to 0.5), it is determined to be a valid hazard target, and the hazard target area is then located based on the bounding box coordinates. After locating the hazard target area, a convolutional neural network is used to extract the 2D image features of the area. The extracted features include the bounding box size, center point coordinates, contour features, and texture features of the hazard target. These 2D image features will serve as important inputs for subsequently constructing the 3D ranging model and calculating the 3D coordinates.
[0047] Step S14: Based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the target area of the hidden danger, construct an uncalibrated monocular vision three-dimensional ranging model. Through the mapping relationship between image pixel scale and actual physical scale, calculate the actual vertical and horizontal distances between the target danger and the power transmission line.
[0048] Specifically, firstly, by utilizing the extracted two-dimensional image features of the potential hazard target and combining parameters such as scale factor and coordinate mapping relationship in the three-dimensional ranging benchmark system, a calibration-free monocular vision three-dimensional ranging model is constructed. This model integrates deep learning features and geometric projection principles, and can accurately output the three-dimensional coordinates of the potential hazard target in the actual physical space. Then, based on the three-dimensional coordinates of the potential hazard target and the three-dimensional spatial curve of the power transmission line, the actual vertical and horizontal distances between the potential hazard target and the power transmission line are calculated using a spatial distance calculation algorithm.
[0049] Step S15: Based on the transmission line safety distance threshold standard, assess the risk of the actual vertical and horizontal distances, and output the hazard safety distance value and the corresponding risk level.
[0050] The safety distance threshold standard for transmission lines is implemented in accordance with relevant national power industry standards. Different voltage levels of transmission lines have different safety distance thresholds; for example, the safety distance threshold for a 110kV transmission line is 1.5 meters, for a 220kV transmission line it is 3.0 meters, and for a 500kV transmission line it is 5.0 meters. During implementation, the voltage level of the currently monitored transmission line is first obtained to determine the corresponding safety distance threshold. Then, the calculated actual vertical and horizontal distances are compared with the safe distance threshold. If the actual distances are both greater than the safe distance threshold, it is determined to be no risk. If one or both of the actual distances are close to the safe distance threshold (for example, the actual distance is 1.0 to 1.2 times the safe distance threshold), it is determined to be a general risk. If one or both of the actual distances are less than the safe distance threshold but greater than 0.5 times the safe distance threshold, it is determined to be a key risk. If one or both of the actual distances are less than or equal to 0.5 times the safe distance threshold, it is determined to be an emergency high-risk risk.
[0051] Finally, the system outputs the specific category of the hazard, the actual vertical distance between the hazard and the conductor, the actual horizontal distance between the hazard and the conductor, and the corresponding risk level.
[0052] In summary, the transmission line safety hazard detection method in the above embodiments of the present invention obtains a target monitoring image by acquiring and preprocessing a sequence of time-series images of the transmission line taken at intervals by a monocular camera installed on the tower; extracts features from the target monitoring image to obtain inherent reference features of the transmission line, including the preset identification size of the tower, the geometric parameters of the conductor catenary, and the installation and fixing parameters of the camera; aligns and calibrates the actual physical size of the inherent reference features with the corresponding image pixel size to construct a three-dimensional ranging reference system; uses a pre-trained target detection model to identify hazards in the target monitoring image and extract two-dimensional image features; based on the three-dimensional ranging reference system and combined with the two-dimensional image features of the hazard target area, constructs an uncalibrated monocular vision three-dimensional ranging model, and calculates the actual vertical and horizontal distances between the hazard target and the transmission conductor through the mapping relationship between the image pixel scale and the actual physical scale; and, combined with the transmission line safety distance threshold standard, determines the risk of the actual vertical and horizontal distances, and outputs the hazard safety distance value and the corresponding risk level. This approach avoids the limitations of existing technologies, which rely solely on blurry two-dimensional images to determine the distance to potential hazards due to a lack of 3D depth information, thus failing to provide specific distance values and quantified risk levels. It also avoids the high costs, high power consumption, and significant hardware modification challenges associated with adding expensive 3D measurement equipment such as binocular cameras and LiDAR. Ultimately, it achieves precise 3D ranging and risk quantification assessment of power transmission line hazards using existing monocular cameras without incurring additional hardware costs. This solves the problem that existing technologies cannot quantitatively assess the risk of power transmission line safety hazards, failing to meet the practical needs of power operation and maintenance for accurate early warning and handling of hazards.
[0053] Example 2 This embodiment also proposes a method for detecting safety hazards in transmission lines. The difference between the method for detecting safety hazards in transmission lines in this embodiment and the method for detecting safety hazards in transmission lines in Embodiment 1 is as follows: The steps for constructing an uncalibrated monocular vision three-dimensional ranging model based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the potential hazard target area include: A lightweight convolutional neural network is used to extract two-dimensional image feature maps of the target area of potential hazards. It includes the target bounding box, center point coordinates, and contour features; center point of the target Projecting onto the ground plane in the three-dimensional ranging reference system yields preliminary ground projection points. ; Based on the relative vertical pixel distance between the bottom of the target and the guide wire in the image By combining the geometric parameters of the catenary and the camera tilt angle, the actual height of the target is calculated using the principle of triangulation. ,in, The camera's tilt angle; By integrating deep learning features with geometric projection results, an end-to-end calibration-free monocular vision 3D ranging model is constructed. Output the three-dimensional coordinates of the target in physical space. .
[0054] First, a lightweight convolutional neural network is used to extract two-dimensional image feature maps of the target area of the potential hazard. The data includes the target bounding box, center point coordinates, and contour features. The selection of a lightweight convolutional neural network needs to balance feature extraction accuracy and computational efficiency, suitable for the real-time computational requirements of a monocular camera monitoring system. This embodiment uses the MobileNetV3-small model, which, through depthwise separable convolution and attention mechanisms, reduces the number of model parameters and computational load while ensuring feature extraction accuracy. In implementation, the image of the potential hazard target area located in Embodiment 1 is cropped and used as the input to the MobileNetV3-small model. The input image size is set to 224×224. The model's feature extraction network, including convolutional layers, pooling layers, and attention mechanism layers, extracts the two-dimensional image feature map of the potential hazard target. Feature map The number of channels is 512, and the size is 7×7. It contains various feature information of the potential hazard target, including the size information of the target bounding box (width, height) and the pixel coordinates of the target center point. The target's contour features, texture features, etc., will serve as important inputs for subsequent model construction and 3D coordinate calculation. To further improve the accuracy of feature extraction, a feature fusion module can be added to the model's output layer to fuse features from different levels, thereby enhancing the expressive power of the target features.
[0055] Then, the target center point Projecting onto the ground plane in the three-dimensional ranging reference system yields preliminary ground projection points. The ground projection of the target's center point is fundamental to calculating the three-dimensional coordinates of potential hazards. Its core lies in using a mapping function within the constructed three-dimensional ranging benchmark system to convert the pixel coordinates of the target's center point into actual physical coordinates on the ground plane. In the implementation process, the pixel coordinates of the target's center point are first obtained. These coordinates are from the feature map extracted from the MobileNetV3-small model. The value obtained represents the center position of the potential hazard target in the two-dimensional image coordinate system. Then, the pixel coordinates... Substitute into the constructed mapping function Since the Z-coordinate of the ground plane is 0, the initial projection coordinates of the target center point on the ground plane can be directly obtained through the mapping function. These coordinates represent the approximate location of the potential target's projection onto the ground in the actual physical space, providing a basis for subsequent calculations of the target's actual height and determination of its three-dimensional coordinates.
[0056] Next, based on the relative vertical pixel distance Δh between the bottom of the target and the guide wire in the image, combined with the geometric parameters of the catenary and the camera tilt angle, the actual height of the target is calculated using the principle of triangulation. ,in The camera's tilt angle is used. The core of calculating the actual height of the target is to utilize the principle of triangulation, combined with a longitudinal scale factor and the camera's tilt angle, to convert the vertical pixel distance in the image into the actual physical height. In practice, firstly, the pixel coordinates of the bottom of the hazard target and the corresponding position of the guide wire are determined in the target monitoring image. The pixel coordinates of the bottom of the target can be determined by the midpoint of the lower edge of the target's bounding box, and the pixel coordinates of the corresponding position of the guide wire can be determined by the guide wire's catenary equation, i.e., finding the guide wire pixel point on the same vertical line as the bottom of the target. Then, the vertical pixel distance between the two is calculated. That is, two points at The absolute value of the pixel difference along the axial direction. Combined with the calculated vertical scale factor. The inverted camera pitch angle Substitute into the formula The actual height of the potential hazard can then be calculated. For example, longitudinal scale factor. meters per pixel, relative vertical pixel distance Pixels, camera tilt angle = 30°, If ≈ 0.577, then the actual height of the target is... = 0.01×100×0.577 ≈ 0.577 meters, which is the actual vertical distance from the ground to the top of the potential hazard.
[0057] Finally, by fusing deep learning features with geometric projection results, an end-to-end calibration-free monocular vision 3D ranging model is constructed. Output the three-dimensional coordinates of the target in physical space. Uncalibrated monocular vision 3D ranging model The core is to integrate visual features extracted by deep learning with preliminary coordinate and height information obtained from geometric projection to achieve accurate output of 3D coordinates. The model adopts a multimodal feature fusion structure, taking into account the 2D image feature maps extracted by deep learning. Preliminary ground projection points obtained by geometric projection Actual height of the target By employing feature fusion and depth correction, errors in monocular vision imaging are eliminated, improving the accuracy of 3D coordinate calculation. The model training process utilizes a large number of transmission line hazard samples, including images and actual 3D coordinates of the hazard targets. Backpropagation is used to optimize model parameters, ensuring the model can accurately output the 3D coordinates of the hazard targets. Ultimately, the model... Output the complete three-dimensional coordinates of the potential hazard in the actual physical space. ,in and Let be the projected coordinates of the target on the ground plane. The actual height of the target, these three-dimensional coordinates provide core data support for subsequent calculations of the safe distance between potential hazards and the conductor.
[0058] Furthermore, the step of fusing deep learning features and geometric projection results to construct an end-to-end uncalibrated monocular vision 3D ranging model includes: A multimodal feature fusion network is constructed, which includes a geometric prior branch and a visual feature branch; The coordinates of the preliminary ground projection points and the target height in the geometric projection results are used as the input of the geometric prior branch, and the two-dimensional image feature map extracted by the lightweight convolutional neural network is used as the input of the visual feature branch. In the feature fusion layer, a channel attention mechanism is used to weight the visual feature branches, and the output of the geometric prior branch is used to constrain the spatial position of the weighted visual features. The fused features constrained by spatial location are input into the fully connected regression layer, and the depth-corrected three-dimensional coordinates of the target in physical space are output.
[0059] For example, the step of using the output of the geometric prior branch to constrain the spatial position of the weighted visual features includes: The initial ground projection point coordinates output by the geometric prior branch are converted into a Gaussian heatmap distribution, where the center of the Gaussian heatmap distribution corresponds to the initial ground projection point coordinates and the variance corresponds to the uncertainty of the geometric projection. The Gaussian heatmap distribution is upsampled to the same size as the visual feature map to generate a spatial attention mask; The spatial attention mask is multiplied element-wise with the weighted visual features to suppress redundant feature responses that deviate from the geometric prior position and to enhance the feature weights of the target's true landing area.
[0060] In addition, in some optional embodiments of the present invention, the step of calculating the actual vertical and horizontal distances between the potential hazard target and the power transmission line by means of the mapping relationship between image pixel scale and actual physical scale includes: Based on catenary parameters and a three-dimensional ranging benchmark system, the three-dimensional spatial curve of the transmission line is reconstructed in the area where the potential hazard is located. ; Three-dimensional coordinates of the potential hazard target With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. ; Calculate the target projection point With point Projected distance on the horizontal plane ; Calculate target height With point height absolute value of the difference ; Output the actual horizontal distance between the potential hazard target and the power transmission line. and vertical distance .
[0061] First, based on the catenary parameters and the three-dimensional ranging benchmark system, the three-dimensional spatial curve of the transmission line is reconstructed in the area where the potential hazard is located. The reconstruction of the three-dimensional spatial curve of the conductor is the basis for calculating the distance between the hazard and the conductor. Its core is to use the extracted catenary parameters of the conductor and the constructed three-dimensional ranging benchmark system to map the catenary of the conductor in the two-dimensional image to the actual physical space.
[0062] During implementation, the fitted catenary equation is first obtained. and catenary coefficient Gear spacing sag Then, by combining the mapping function in the three-dimensional ranging reference system, the pixel coordinates of multiple key points on the catenary of the conductor (such as suspension point, lowest point, intermediate point, etc.) are converted into actual physical coordinates; finally, a cubic spline interpolation algorithm is used to interpolate and fit the actual physical coordinates of these key points to obtain the complete three-dimensional space curve of the transmission conductor. This curve can accurately describe the shape of the conductor in actual physical space, and its equation can be expressed as follows: ,in The wires are respectively in direction, direction, The coordinate function of direction.
[0063] Then, the three-dimensional coordinates of the potential hazard target With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. The closest Euclidean distance point This is the point on the three-dimensional curve of the traverse line that is closest to the target hazard. The accuracy of its measurement directly affects the accuracy of subsequent distance calculations. Specifically: First, a discretization sampling strategy is adopted, along the three-dimensional curve of the traverse line... A large number of spatial sampling points are generated according to a preset step size. The preset step size can be set according to the length of the traverse and the required measurement accuracy, and is usually set to 0.1 meters to 0.5 meters. The smaller the step size, the denser the sampling points and the higher the measurement accuracy. Then, a three-dimensional coordinate system based on the target of the hazard is constructed. For a spatial search sphere centered on a target hazard, an initial search radius is set, typically 10 meters, to ensure coverage of the nearest possible area between the guide wire and the hazard target. Next, the Euclidean distance between each spatial sampling point and the hazard target's three-dimensional coordinates is calculated. The formula for calculating the Euclidean distance is... ,in The three-dimensional coordinates of the spatial sampling points are used; finally, the sampling point with the smallest distance is found through filtering, which is the nearest Euclidean distance point. Its three-dimensional coordinates are .
[0064] Next, calculate the target projection point. With point Projected distance on the horizontal plane Horizontal distance This is the distance between the projection of the potential hazard onto the ground and the projection of the nearest point of the traverse line onto the ground, reflecting the horizontal distance relationship between the potential hazard and the traverse line. During implementation, the coordinates of the ground projection point of the potential hazard are first determined as follows: This point represents the three-dimensional coordinates of the potential hazard on the ground plane. Project the image onto the surface; then determine the nearest Euclidean distance point. The coordinates of the ground projection point are This point is The projection of the target onto the ground plane; finally, substituting into the Euclidean distance formula, the distance between the two points is calculated, which is the actual horizontal distance between the potential hazard target and the power transmission line. .
[0065] Then, calculate the target height. With point height absolute value of the difference Vertical distance This is the vertical difference between the actual height of the potential hazard and the actual height of the nearest point on the guide wire, reflecting the vertical distance between the potential hazard and the guide wire. During implementation, obtaining the actual height of the potential hazard is crucial. and the nearest Euclidean distance point actual height Calculate the absolute value of the difference between the two, which is the actual vertical distance between the potential hazard and the power transmission line. For example, the actual height of the potential hazard target. meters, the nearest point of the conductor actual height meters, then vertical distance Meters, this distance reflects the vertical distance between the potential hazard and the conductor.
[0066] Finally, output the actual horizontal distance between the potential hazard target and the power transmission line. and vertical distance During implementation, the calculated... and The data is processed and rounded to two decimal places to ensure accuracy. It is also output along with information about the type of potential hazard, providing crucial data support for subsequent risk assessment. For example, if a hazard is identified as a tree, the output shows the actual horizontal distance between the tree and the power line as 8.52 meters and the actual vertical distance as 9.36 meters. This allows maintenance personnel to accurately determine the distance relationship between the hazard and the power line and assess whether the hazard poses a safety risk.
[0067] Furthermore, the three-dimensional coordinates of the hazard target... With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. The steps include: A discretization sampling strategy is adopted to generate several spatial sampling points along the reconstructed three-dimensional spatial curve of the transmission line at a preset step size; Construct a spatial search sphere centered on the three-dimensional coordinates of the potential hazard target, and set an initial search radius; Calculate the Euclidean distance between each spatial sampling point and the three-dimensional coordinates of the potential hazard target, and filter out the set of candidate sampling points that fall within the spatial search sphere; If the candidate sampling point set is not empty, then the point with the smallest Euclidean distance in the candidate sampling point set is selected as the nearest Euclidean distance point. If the candidate sampling point set is empty, increase the search radius and repeat the sampling and filtering steps until the nearest Euclidean distance point is found. First, the core of discretization sampling is to convert the continuous three-dimensional spatial curve of the conductor into discrete sampling points, facilitating subsequent distance calculation and nearest-point selection. The setting of the sampling step size directly affects the measurement accuracy and computational efficiency. The smaller the step size, the denser the sampling points and the higher the measurement accuracy, but the greater the computational load; conversely, the accuracy decreases and the computational load decreases. In implementation, based on the actual conditions of the transmission line, a suitable preset step size is set. Spatial sampling points are generated sequentially along the three-dimensional spatial curve of the conductor from one suspension point to the other, according to the preset step size. The three-dimensional coordinates of each sampling point are calculated using the equation of the three-dimensional spatial curve of the conductor. All sampling points are arranged in sequence to form a complete set of conductor spatial sampling points. The number of sampling points in the set is determined according to the conductor span, ensuring that the sampling points can fully cover the entire conductor spatial curve, providing sufficient data support for subsequent nearest-point selection.
[0068] Then, a spatial search sphere centered on the 3D coordinates of the potential hazard is constructed, and an initial search radius is set. The purpose of the spatial search sphere is to narrow the search range for the nearest point, reduce computational load, improve the efficiency of nearest point measurement, avoid calculating the distance to all sampling points, and ensure that no possible nearest points are missed. In implementation, the initial search radius is set using the 3D coordinates of the potential hazard as the center, combined with the distance range between common potential hazards and conductors in transmission lines. This initial radius can cover the possible distance between potential hazards and conductors in most scenarios, while effectively narrowing the search range and balancing computational efficiency and measurement accuracy. If the monitoring scenario is a high-incidence area of potential hazards, the initial search radius can be appropriately increased; if it is a regular monitoring area, the initial radius can be reduced, flexibly adapting to the needs of different monitoring scenarios.
[0069] Next, the Euclidean distance between each spatial sampling point and the three-dimensional coordinates of the potential hazard target is calculated, and a set of candidate sampling points falling within the spatial search sphere is selected. Euclidean distance accurately reflects the straight-line distance between two spatial points and is the core method for measuring spatial point distances. During implementation, all generated spatial sampling points are traversed, and the Euclidean distance between each sampling point and the three-dimensional coordinates of the potential hazard target is calculated one by one. Simultaneously, the calculated Euclidean distance is compared with the initial search radius. If the Euclidean distance of a sampling point is less than or equal to the initial search radius, the sampling point is determined to fall within the spatial search sphere and included in the candidate sampling point set; if the Euclidean distance is greater than the initial search radius, the sampling point is discarded and not included in the candidate set. Through this selection process, a set of candidate sampling points containing only the closest possible points is obtained, significantly reducing the computational workload of subsequent distance comparisons.
[0070] If the candidate sampling point set is not empty, the point with the smallest Euclidean distance in the candidate sampling point set is selected as the nearest Euclidean distance point. During implementation, if the candidate sampling point set obtained after screening is not empty, it means that the initial search radius can cover the closest area between the guide wire and the hazard target. In this case, all sampling points in the candidate sampling point set are traversed, and the Euclidean distances between each sampling point and the hazard target are compared. The sampling point with the smallest distance value is selected and determined as the nearest Euclidean distance point, and its three-dimensional coordinates are recorded. To ensure the accuracy of the selection results, the Euclidean distances of the candidate sampling point set can be sorted in ascending order. The first sampling point after sorting is taken as the nearest Euclidean distance point. Simultaneously, the distance difference between this point and the second sampling point after sorting is calculated. If the difference is too small, it means that the two sampling points are close. The sampling step size can be further reduced, and additional sampling can be performed between the two points. The distance is then recalculated to determine the final nearest point, avoiding measurement errors caused by the sampling step size.
[0071] If the candidate sampling point set is empty, increase the search radius and repeat the sampling and filtering steps until the nearest Euclidean distance point is found. An empty candidate sampling point set indicates that the initial search radius is too small and does not cover the closest area between the guide wire and the potential hazard. In this case, the search radius needs to be increased, typically by 1.5 times. If it is still empty after expansion, continue to gradually increase the radius proportionally until a candidate sampling point set is found. Simultaneously, to avoid a surge in computational load due to an excessively large radius, the sampling step size can be adjusted appropriately after each increase in the search radius to balance computational efficiency while ensuring measurement accuracy. Repeat the above discretization sampling, spatial search sphere construction, Euclidean distance calculation, and filtering steps until a non-empty candidate sampling point set is obtained. Then, select the point with the smallest Euclidean distance from this set as the nearest Euclidean distance point to ensure accurate location of the closest spatial point between the guide wire and the potential hazard, providing accurate basic data for subsequent horizontal and vertical distance measurements.
[0072] In summary, the transmission line safety hazard detection method in the above embodiments of the present invention obtains a target monitoring image by acquiring and preprocessing a sequence of time-series images of the transmission line taken at intervals by a monocular camera installed on the tower; extracts features from the target monitoring image to obtain inherent reference features of the transmission line, including the preset identification size of the tower, the geometric parameters of the conductor catenary, and the installation and fixing parameters of the camera; aligns and calibrates the actual physical size of the inherent reference features with the corresponding image pixel size to construct a three-dimensional ranging reference system; uses a pre-trained target detection model to identify hazards in the target monitoring image and extract two-dimensional image features; based on the three-dimensional ranging reference system and combined with the two-dimensional image features of the hazard target area, constructs an uncalibrated monocular vision three-dimensional ranging model, and calculates the actual vertical and horizontal distances between the hazard target and the transmission conductor through the mapping relationship between the image pixel scale and the actual physical scale; and, combined with the transmission line safety distance threshold standard, determines the risk of the actual vertical and horizontal distances, and outputs the hazard safety distance value and the corresponding risk level. This approach avoids the limitations of existing technologies, which rely solely on blurry two-dimensional images to determine the distance to potential hazards due to a lack of 3D depth information, thus failing to provide specific distance values and quantified risk levels. It also avoids the high costs, high power consumption, and significant hardware modification challenges associated with adding expensive 3D measurement equipment such as binocular cameras and LiDAR. Ultimately, it achieves precise 3D ranging and risk quantification assessment of power transmission line hazards using existing monocular cameras without incurring additional hardware costs. This solves the problem that existing technologies cannot quantitatively assess the risk of power transmission line safety hazards, failing to meet the practical needs of power operation and maintenance for accurate early warning and handling of hazards.
[0073] Example 3 Please see Figure 2 The figure shows a power transmission line safety hazard detection system proposed in the third embodiment of the present invention, the system comprising: The acquisition module 100 is used to acquire a sequence of time-series images of the power transmission line site captured at intervals by a monocular camera set on the tower, and to preprocess the time-series image sequence to obtain the target monitoring image. The extraction module 200 is used to extract features from the target monitoring image to obtain the inherent reference features of the transmission line. The inherent reference features include the preset identification size of the tower, the geometric parameters of the conductor catenary, and the installation and fixing parameters of the camera. Module 300 is used to perform correlation calibration between the actual physical size and the corresponding image pixel size based on the inherent reference features, and to construct a three-dimensional ranging reference system. The positioning module 400 is used to identify potential hazards in the target monitoring image using a pre-trained target detection model, locate the potential hazard target area, and extract the two-dimensional image features of the potential hazard target area. The mapping module 500 is used to construct an uncalibrated monocular vision three-dimensional ranging model based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the target area of the hidden danger. Through the mapping relationship between the image pixel scale and the actual physical scale, the actual vertical distance and horizontal distance between the target danger and the transmission line are calculated. The identification module 600 is used to determine the risk of the actual vertical and horizontal distances by combining the transmission line safety distance threshold standard, and output the hidden danger safety distance value and the corresponding risk level.
[0074] The functions or operation steps implemented by the above modules are largely the same as those in the above method embodiments, and will not be repeated here.
[0075] Example 4 In another aspect, the present invention provides a readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the method described in any one of Embodiments 1 to 2 above.
[0076] Example 5 In another aspect, the present invention provides an electronic device, the electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of any one of the methods described in Embodiments 1 to 2 above.
[0077] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0078] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequential list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0079] More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable storage media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0080] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0081] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0082] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for detecting safety hazards in power transmission lines, characterized in that, The method includes: A sequence of time-series images of the power transmission line site captured at intervals by a monocular camera installed on the tower is obtained, and the time-series image sequence is preprocessed to obtain the target monitoring image; Feature extraction is performed on the target monitoring image to obtain the inherent reference features of the transmission line. The inherent reference features include the preset identification size of the tower, the geometric parameters of the catenary of the conductor, and the installation and fixing parameters of the camera. A three-dimensional ranging benchmark system is constructed by correlating and calibrating the actual physical size with the corresponding image pixel size based on inherent benchmark features. A pre-trained target detection model is used to identify potential hazards in the target monitoring image, locate the potential hazard target area, and extract the two-dimensional image features of the potential hazard target area; Based on the aforementioned three-dimensional ranging benchmark system, and combined with the two-dimensional image features of the target area of the potential hazard, an uncalibrated monocular vision three-dimensional ranging model is constructed. Through the mapping relationship between image pixel scale and actual physical scale, the actual vertical and horizontal distances between the potential hazard target and the power transmission line are calculated. Based on the safety distance threshold standard for transmission lines, the actual vertical and horizontal distances are assessed for risk, and the hazard safety distance value and corresponding risk level are output.
2. The method for detecting safety hazards in transmission lines according to claim 1, characterized in that, The step of extracting features from the target monitoring image to obtain the inherent reference features of the transmission line includes: An algorithm combining template matching and shape context descriptors is used to locate the pre-defined marking area of the tower in the target monitoring image, and the marking outline is extracted by a sub-pixel edge detection algorithm to calculate the marking center coordinates and pixel size; The edge point set of the conductor is extracted by using the Canny operator combined with nonmaximum suppression, and the sag, span, and tangent slope of the conductor are calculated by fitting the catenary equation of the conductor based on the least squares method. Based on the ratio of the actual physical size of the tower marker to the image pixel size, and combined with the geometric parameters of the catenary, the camera installation tilt angle, focal length, and pose parameters relative to the tower coordinate system are inverted through the perspective projection transformation model. The extracted marker dimensions, catenary parameters, and inverted camera parameters are subjected to consistency verification, abnormal feature values are removed, and a fused transmission line inherent reference feature vector is generated.
3. The method for detecting safety hazards in transmission lines according to claim 2, characterized in that, The steps for constructing a three-dimensional ranging reference system by correlating and calibrating the actual physical dimensions based on inherent reference features with the corresponding image pixel dimensions include: Based on the actual height of the pole / tower markings With pixel height in the image Calculate the longitudinal scale factor ; Using the catenary fitting results of the conductor, a two-dimensional image coordinate system with the bottom of the tower as the origin is constructed in the image, and the key points of the conductor are projected onto this coordinate system; Based on the camera installation height and tilt angle parameters, a ground plane equation is constructed. The projection lines in the image serve as a depth reference plane; Combined with scale factor Image coordinate system and depth reference plane, constructing from image pixel coordinates to actual physical coordinates mapping function ,in The direction is constrained by the direction of gravity, forming a three-dimensional ranging reference system with physical significance.
4. The method for detecting safety hazards in transmission lines according to claim 1, characterized in that, The steps for constructing an uncalibrated monocular vision three-dimensional ranging model based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the potential hazard target area include: A lightweight convolutional neural network is used to extract two-dimensional image feature maps of the target area of potential hazards. It includes the target bounding box, center point coordinates, and contour features; center point of the target Projecting onto the ground plane in the three-dimensional ranging reference system yields preliminary ground projection points. ; Based on the relative vertical pixel distance between the bottom of the target and the guide wire in the image By combining the geometric parameters of the catenary and the camera tilt angle, the actual height of the target is calculated using the principle of triangulation. ,in, The camera's tilt angle; By integrating deep learning features with geometric projection results, an end-to-end calibration-free monocular vision 3D ranging model is constructed. Output the three-dimensional coordinates of the target in physical space. .
5. The method for detecting safety hazards in transmission lines according to claim 4, characterized in that, The steps for calculating the actual vertical and horizontal distances between the potential hazard target and the power transmission line by mapping the image pixel scale to the actual physical scale include: Based on catenary parameters and a three-dimensional ranging benchmark system, the three-dimensional spatial curve of the transmission line is reconstructed in the area where the potential hazard is located. ; Three-dimensional coordinates of the potential hazard target With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. ; Calculate the target projection point With point Projected distance on the horizontal plane ; Calculate target height With point height absolute value of the difference ; Output the actual horizontal distance between the potential hazard target and the power transmission line. and vertical distance .
6. The method for detecting safety hazards in transmission lines according to claim 3, characterized in that, The combined scale factor Image coordinate system and depth reference plane, constructing from image pixel coordinates to actual physical coordinates mapping function The steps include: Select at least four non-collinear corner points of the tower markings in the image as control points, and construct the homography matrix between the two-dimensional image plane and the three-dimensional physical space ground plane by combining their corresponding actual physical coordinates. The homography matrix is normalized using the longitudinal scale factor to eliminate scale blur caused by the unknown focal length of the camera. Based on the normalized homography matrix, a linear mapping relationship between image pixel coordinates and ground plane physical coordinates is established. By introducing the slope of the catenary tangent as a nonlinear correction factor in the vertical direction, distortion compensation is performed on the linear mapping relationship, resulting in a complete mapping function containing both horizontal and vertical dimensions.
7. The method for detecting safety hazards in transmission lines according to claim 4, characterized in that, The steps for constructing an end-to-end uncalibrated monocular vision 3D ranging model by fusing deep learning features and geometric projection results include: A multimodal feature fusion network is constructed, which includes a geometric prior branch and a visual feature branch; The coordinates of the preliminary ground projection points and the target height in the geometric projection results are used as the input of the geometric prior branch, and the two-dimensional image feature map extracted by the lightweight convolutional neural network is used as the input of the visual feature branch. In the feature fusion layer, a channel attention mechanism is used to weight the visual feature branches, and the output of the geometric prior branch is used to constrain the spatial position of the weighted visual features. The fused features constrained by spatial location are input into the fully connected regression layer, and the depth-corrected three-dimensional coordinates of the target in physical space are output.
8. The method for detecting safety hazards in transmission lines according to claim 5, characterized in that, The three-dimensional coordinates of the hidden danger target With conductor curve Perform spatial matching and calculate the nearest Euclidean distance point from the target to the traverse. The steps include: A discretization sampling strategy is adopted to generate several spatial sampling points along the reconstructed three-dimensional spatial curve of the transmission line at a preset step size; Construct a spatial search sphere centered on the three-dimensional coordinates of the potential hazard target, and set an initial search radius; Calculate the Euclidean distance between each spatial sampling point and the three-dimensional coordinates of the potential hazard target, and filter out the set of candidate sampling points that fall within the spatial search sphere; If the candidate sampling point set is not empty, then the point with the smallest Euclidean distance in the candidate sampling point set is selected as the nearest Euclidean distance point. If the candidate sampling point set is empty, increase the search radius and repeat the sampling and filtering steps until the nearest Euclidean distance point is found.
9. The method for detecting safety hazards in transmission lines according to claim 7, characterized in that, The step of using the output of the geometric prior branch to constrain the spatial position of the weighted visual features includes: The initial ground projection point coordinates output by the geometric prior branch are converted into a Gaussian heatmap distribution, where the center of the Gaussian heatmap distribution corresponds to the initial ground projection point coordinates and the variance corresponds to the uncertainty of the geometric projection. The Gaussian heatmap distribution is upsampled to the same size as the visual feature map to generate a spatial attention mask; The spatial attention mask is multiplied element-wise with the weighted visual features to suppress redundant feature responses that deviate from the geometric prior position and to enhance the feature weights of the target's true landing area.
10. A power transmission line safety hazard detection system, characterized in that, The system includes: The acquisition module is used to acquire a sequence of time-series images of the power transmission line at intervals captured by a monocular camera set on the tower, and to preprocess the time-series image sequence to obtain the target monitoring image; The extraction module is used to extract features from the target monitoring image to obtain the inherent reference features of the transmission line. The inherent reference features include the preset identification size of the tower, the geometric parameters of the conductor catenary, and the installation and fixing parameters of the camera. The module is used to correlate and calibrate the actual physical size and corresponding image pixel size based on the inherent reference features to build a three-dimensional ranging reference system; The localization module is used to identify potential hazards in the target monitoring image using a pre-trained target detection model, locate the potential hazard target area, and extract the two-dimensional image features of the potential hazard target area. The mapping module is used to construct an uncalibrated monocular vision three-dimensional ranging model based on the three-dimensional ranging benchmark system and combined with the two-dimensional image features of the target area of the hidden danger. Through the mapping relationship between the image pixel scale and the actual physical scale, the actual vertical and horizontal distances between the target danger and the transmission line are calculated. The identification module is used to determine the risk of the actual vertical and horizontal distances by combining the safety distance threshold standard for transmission lines, and output the safety distance value of the hidden danger and the corresponding risk level.