Power transmission line corridor monocular image ranging method and system based on depth vision
By employing a depth vision-based monocular image ranging method for power transmission line corridors, and utilizing monocular image acquisition equipment and a lightweight model, hazardous objects can be identified and located in real time. This solves the problem of inaccurate identification of hazardous objects in existing technologies, enables dynamic distance monitoring and hierarchical alarms, and improves the safety of power transmission.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU HAOHAN INFORMATION TECH
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies cannot accurately identify hazards in power transmission line corridors in real time, and lack dynamic distance monitoring and graded alarm mechanisms, making it difficult to ensure the safety of power transmission.
A monocular image ranging method based on depth vision is adopted for power transmission line corridors. Two-dimensional image information is acquired using monocular image acquisition equipment, a lightweight target detection model is constructed, and a mapping relationship between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system is established. The spatial coordinates of the target hazard are located through the mapping relationship matrix, the real-time spatial distance to the power transmission line is determined, and a graded alarm is triggered.
It enables real-time monitoring and tiered alarm of hazardous materials along power transmission line corridors, improving the efficiency and safety of power inspections and ensuring the stability of power transmission.
Smart Images

Figure CN121978667A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image ranging technology, specifically to a method and system for monocular image ranging of power transmission line corridors based on depth vision. Background Technology
[0002] With the advancement of smart grid construction, the safety monitoring of transmission line corridors faces higher requirements. Traditional inspection methods suffer from low efficiency, high cost, and are limited by terrain and climate, making it difficult to achieve all-weather, all-area coverage. Existing automated monitoring technologies mostly rely on complex deep learning models, which suffer from bottlenecks such as insufficient computing power and response latency when deployed on edge computing devices, failing to meet the real-time detection requirements of hazardous materials. At the same time, some solutions can only identify target categories, lacking accurate quantification of the spatial distance between hazardous materials and transmission lines, making it difficult to trigger graded alarms in conjunction with safety thresholds.
[0003] Existing technologies have technical problems such as the inability to accurately identify hazards in power transmission line corridors in real time, the lack of dynamic distance monitoring and graded alarm mechanisms, and difficulty in ensuring the safety of power transmission. Summary of the Invention
[0004] This application provides a method and system for monocular image ranging of power transmission line corridors based on depth vision, which is used to address the technical problems in the prior art that are unable to accurately identify dangerous objects in power transmission line corridors in real time, lack dynamic distance monitoring and hierarchical alarm mechanisms, and are difficult to ensure the safety of power transmission.
[0005] In view of the above problems, this application provides a method and system for monocular image ranging of power transmission line corridors based on depth vision.
[0006] The first aspect of this application provides a method for ranging power transmission line corridors using monocular images based on depth vision, the method comprising: Using monocular image acquisition equipment deployed on transmission towers, two-dimensional image information of the transmission line corridor area is acquired, and a lightweight target detection model for edge computing devices is constructed. Hazard source identification is performed on the two-dimensional image information, outputting a set of two-dimensional pixel coordinates with hazard source category labels. Three-dimensional point cloud data of the transmission line corridor is imported into the edge computing device, establishing a mapping matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system. Using the mapping matrix, the spatial coordinate information of the target hazard in the three-dimensional point cloud space is located through the set of two-dimensional pixel coordinates with hazard source category labels. The spatial coordinate information is associated with each element in the two-dimensional pixel coordinate set. Based on the spatial coordinate information, the real-time spatial distance between the target hazard and the transmission line is determined, and it is judged whether the real-time spatial distance is less than a preset safety threshold. If the judgment result is yes, a local early warning mechanism is triggered to generate a graded alarm signal.
[0007] A second aspect of this application provides a monocular image ranging system for power transmission line corridors based on depth vision, the system comprising: The system includes the following modules: a hazard identification module (using monocular image acquisition equipment deployed on transmission towers to acquire two-dimensional image information of the transmission line corridor area, constructing a lightweight target detection model for edge computing devices, identifying hazard sources in the two-dimensional image information, and outputting a set of two-dimensional pixel coordinates labeled with hazard categories); a relation matrix establishment module (importing three-dimensional point cloud data of the transmission line corridor into the edge computing device to establish a mapping relationship matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system); a spatial coordinate information positioning module (using the mapping relationship matrix to locate the spatial coordinate information of the target hazard in the three-dimensional point cloud space through the set of two-dimensional pixel coordinates labeled with hazard categories, wherein the spatial coordinate information is associated with each element in the two-dimensional pixel coordinate set); a spatial distance determination module (determining the real-time spatial distance between the target hazard and the transmission line based on the spatial coordinate information, and judging whether the real-time spatial distance is less than a preset safety threshold); and a graded alarm signal generation module (triggering a local early warning mechanism to generate graded alarm signals if the judgment result is yes).
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: Using monocular image acquisition equipment deployed on transmission towers, two-dimensional image information of the transmission line corridor area is acquired, and a lightweight target detection model for edge computing devices is constructed to identify hazard sources and output a set of two-dimensional pixel coordinates with hazard source category labels. A mapping matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system is established. Using the mapping matrix, the spatial coordinate information of the target hazard in the three-dimensional point cloud space is located through the set of two-dimensional pixel coordinates with hazard source category labels. The real-time spatial distance between the target hazard and the transmission line is determined, and it is judged whether the real-time spatial distance is less than a preset safety threshold. If the judgment result is yes, a local early warning mechanism is triggered to generate a graded alarm signal. This achieves the technical effect of real-time monitoring and graded alarm of hazards in the transmission line corridor, improving the efficiency and safety of power inspection. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1A schematic flowchart of a monocular image ranging method for power transmission line corridors based on depth vision provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of a monocular image ranging system for power transmission line corridors based on depth vision, provided in an embodiment of this application.
[0011] Explanation of reference numerals in the attached diagram: Hazard source identification module 10, Relationship matrix establishment module 20, Spatial coordinate information positioning module 30, Spatial distance determination module 40, and Graded alarm signal generation module 50. Detailed Implementation
[0012] This application provides a method and system for monocular image ranging of power transmission line corridors based on depth vision, which is used to address the technical problems in the prior art that are unable to accurately identify dangerous objects in power transmission line corridors in real time, lack dynamic distance monitoring and hierarchical alarm mechanisms, and are difficult to ensure the safety of power transmission.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] Example 1, as Figure 1 As shown, this application provides a method for ranging from monocular images of transmission line corridors based on depth vision, the method comprising: Step S100: Use a monocular image acquisition device deployed on the transmission tower to acquire two-dimensional image information of the transmission line corridor area, and construct a lightweight target detection model for edge computing devices to identify hazards in the two-dimensional image information and output a set of two-dimensional pixel coordinates with hazard category labels.
[0015] Specifically, monocular image acquisition devices are deployed on transmission towers to continuously and stably acquire two-dimensional image information of the transmission line corridor area. To efficiently and accurately identify hazards in the images on edge computing devices, a lightweight target detection model specifically designed for edge computing devices is constructed. This model uses MobileNetV3 as the feature extraction backbone network, leveraging its lightweight design to reduce computational load. A spatial attention unit is added after the last convolutional layer to enhance the ability to capture key features. During construction, the pre-trained model parameters are transferred to the lightweight model, and then the model is quantized during deployment on edge computing devices using the spatial attention unit to further optimize performance. This constructed model processes the acquired two-dimensional image information, accurately identifying potential external obstacles or hazards in complex image scenes, such as trees near transmission lines or construction machinery, and outputting a set of two-dimensional pixel coordinates with accurate hazard category labels.
[0016] Step S200: Import the three-dimensional point cloud data of the transmission line corridor into the edge computing device and establish a mapping relationship matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system.
[0017] Specifically, to establish the mapping matrix between the pixel coordinate system and the 3D point cloud spatial coordinate system, the following algorithm flow is adopted. First, the intrinsic parameters of the monocular image acquisition device are calibrated using Zhang's calibration method to obtain accurate camera focal length and principal point offset parameters. Next, the spatial distribution characteristics of the transmission conductors are extracted from the 3D point cloud data of the transmission line corridor, and a distortion correction algorithm based on the least squares method is applied, combined with the camera intrinsic parameters, to correct the distortion error of the image. Then, based on the pinhole camera model, the transformation relationship from the pixel coordinate system to the camera coordinate system is constructed. Next, based on the installation position parameters of the transmission towers, the rotation matrix and translation vector from the camera coordinate system to the world coordinate system (consistent with the 3D point cloud spatial coordinate system) are determined, and the pixel coordinates are transformed to the world coordinate system using the coordinate transformation formula. On this basis, the Iterative Closest Point (ICP) algorithm is used to continuously optimize the transformation parameters, so that the points in the pixel coordinate system and the corresponding points in the 3D point cloud spatial coordinate system achieve the best match, ultimately generating an accurate mapping matrix to achieve precise mapping between the pixel coordinate system and the 3D point cloud spatial coordinate system.
[0018] Step S300: Using the mapping matrix, the spatial coordinate information of the target hazard in the three-dimensional point cloud space is located through the two-dimensional pixel coordinate set with hazard source category markings. The spatial coordinate information is associated and mapped with each element in the two-dimensional pixel coordinate set.
[0019] Specifically, for each coordinate point in the obtained set of two-dimensional pixel coordinates, it is used as input data. A mapping matrix is used to perform spatial transformation of the coordinates through matrix operations. Specifically, according to the transformation rules defined by the matrix, the two-dimensional pixel coordinates are transformed from the pixel coordinate system to the three-dimensional point cloud spatial coordinate system. This accurately maps the coordinates of the hazard location marked on the two-dimensional plane to the three-dimensional space, thereby determining the specific location of the target hazard in the three-dimensional point cloud space and obtaining its spatial coordinate information. Furthermore, due to the deterministic nature of the mapping relationship, there is a one-to-one correspondence between each two-dimensional pixel coordinate and its corresponding three-dimensional spatial coordinate. This provides a reliable data foundation for accurately calculating the spatial distance between the target hazard and the transmission line, as well as for timely assessment of the safety status of the transmission line.
[0020] Step S400: Based on the spatial coordinate information, determine the real-time spatial distance between the target hazard and the transmission line, and determine whether the real-time spatial distance is less than a preset safety threshold.
[0021] Specifically, by acquiring the spatial coordinates of the target hazardous object, the real-time spatial distance between it and the transmission line is calculated, using the Euclidean distance between the target hazardous object and the transmission line as the basic distance measurement method. Simultaneously, considering that the conductors of the actual transmission line may sway due to factors such as wind during operation, a conductor sway parameter is introduced to dynamically compensate for the Euclidean distance. The compensation amount is dynamically adjusted based on real-time wind speed data; the higher the wind speed, the greater the compensation amount, to more accurately reflect the actual distance changes between the target hazardous object and the transmission line. After obtaining the real-time spatial distance, it is compared with a preset safety threshold. The preset safety threshold is a distance value pre-set based on the voltage level of the transmission line, equipment characteristics, and relevant safety standards. If the real-time spatial distance is less than this threshold, it means that the target hazardous object is too close to the transmission line and may pose a threat to the normal operation of the transmission line, requiring timely corresponding measures; if it is greater than or equal to the threshold, it indicates that the current transmission line is in a safe state, and there is no need to worry about the impact of the target hazardous object for the time being.
[0022] Step S500: If the judgment result is yes, trigger the local early warning mechanism to generate a graded alarm signal.
[0023] Specifically, this emergency measure is immediately activated when the real-time spatial distance between the target hazard and the transmission line is determined to be less than a preset safety threshold. A local early warning mechanism is rapidly triggered. This mechanism is an alarm system pre-installed near the transmission towers, which generates different levels of alarm signals based on the degree of hazard. For example, when the hazard is close to the transmission line but not yet at a level that could immediately cause an accident, a lower-level yellow alarm signal is generated, reminding maintenance personnel to pay attention and arrange for further inspection. If the hazard is very close to the transmission line and could potentially cause serious accidents such as short circuits or discharges, a higher-level red alarm signal is generated. Emergency alerts are sent to maintenance personnel and relevant management departments through various means such as audible and visual alarms, SMS notifications, and system pop-ups, ensuring that hazard information is obtained immediately and timely measures such as notifying nearby construction units to stop work and organizing personnel to clear obstacles are taken. This effectively avoids power outages caused by external hazards approaching the transmission line, ensuring the stable operation of the transmission line and the reliability of the power supply.
[0024] In one possible implementation, step S200 further includes: Step S210: Based on the camera focal length and principal point offset parameters of the monocular image acquisition device, and utilizing the spatial distribution characteristics of the transmission conductors in the three-dimensional point cloud data of the transmission line corridor, combined with the camera focal length and principal point offset parameters, the distortion error is corrected.
[0025] Step S220: Simultaneously, select M uniformly distributed transmission tower feature points from the three-dimensional point cloud data of the transmission line corridor, obtain the first coordinate information in the three-dimensional point cloud spatial coordinate system and the second coordinate information in the pixel coordinate system, and construct a nonlinear least squares equation system, where M is greater than or equal to 6.
[0026] Step S230: Perform projection error control based on the nonlinear least squares equations.
[0027] Specifically, the calibration parameter file of the monocular image acquisition device is first called to extract intrinsic parameters such as camera focal length f and principal point coordinates (cx, cy). The radial distortion coefficients (k1, k2, k3) and tangential distortion coefficients (p1, p2) derived by Zhang Zhengyou's calibration method are used to establish an initial distortion model. At the same time, the spatial straight line equation of the transmission conductor is fitted from the three-dimensional point cloud data of the transmission line corridor using the RANSAC algorithm. The direction vector and sag curve parameters of the conductor in three-dimensional space are extracted, and these features are used as the spatial reference for distortion correction. Next, the pixel coordinates of the conductor in the 2D image are substituted into the distortion model to calculate the theoretical distortion offset. Then, combined with the actual spatial position of the conductor in the 3D point cloud, the distortion coefficient is adjusted through an iterative optimization algorithm to make the pixel distribution of the conductor in the image consistent with the spatial distribution characteristics in the 3D point cloud. For example, the difference between the slope of the pixel connection of the conductor in the image and the projection slope of the conductor's direction vector on the imaging plane in 3D space is controlled within 0.5°. This completes the correction of distortion error. Finally, the corrected camera intrinsic parameter matrix and distortion correction lookup table are output.
[0028] Simultaneously with distortion error correction, M uniformly distributed feature points of the transmission towers are selected from the 3D point cloud data of the transmission line corridor, where M is not less than 6. These feature points can be selected from locations with obvious spatial identifiers, such as the top of the lightning rod, the corner connection point of the tower body, and the end point of the crossarm. Then, the first coordinate information of these feature points in the 3D point cloud spatial coordinate system is read using point cloud data processing tools, and the positions of these feature points in the corresponding 2D image are determined using image annotation tools, obtaining their second coordinate information in the pixel coordinate system. Based on the perspective projection principle of camera imaging, a nonlinear least squares equation system is constructed using these paired coordinate information. This equation system aims to minimize the sum of squared errors between the calculated value of each feature point projected from the 3D point cloud spatial coordinate system to the pixel coordinate system and the actual annotated second coordinate information. This ensures the uniqueness and accuracy of the equation system solution through a sufficient number of constraint points, laying a solid foundation for the accurate solution of the subsequent mapping relationship matrix.
[0029] The Levenberg-Marquardt iterative algorithm is used to solve the nonlinear least squares equations to control projection error. First, the initial values of the camera extrinsic matrix (rotation matrix R and translation vector T) are initialized, and the coordinates of feature points in the 3D point cloud spatial coordinate system are converted into predicted coordinates in the pixel coordinate system using the perspective projection formula. Then, the pixel deviation between the predicted coordinates of each feature point and the actual labeled second coordinate information is calculated, and the objective function is constructed by the sum of the squares of all deviations. Then, the incremental equation of the current iteration step is solved by the Jacobian matrix, and the damping factor is dynamically adjusted to control the iteration step size, so that the objective function value gradually decreases. After every 10 iterations, the average projection error is calculated. When the average error change of three consecutive iterations is less than 0.5 pixels and the overall average error is less than 2 pixels, the iteration is stopped. Finally, the optimized camera extrinsic matrix is output, completing the precise control of projection error and ensuring that the coordinate mapping error from 3D to 2D meets the engineering accuracy requirements.
[0030] In one possible implementation, step S210 further includes: Step S211: Based on the monocular visual image stream acquired by the monocular image acquisition device, perform real-time dehazing processing and perform super-resolution reconstruction on the identified hazard source areas.
[0031] Step S212: Simultaneously, coordinate jump errors caused by jitter are eliminated by time series filtering to obtain the two-dimensional image information.
[0032] Specifically, the environment along power transmission line corridors is complex and variable, with frequent occurrences of fog and other weather conditions. This severely reduces image clarity and discernibility, impacting the identification of hazard sources. Therefore, based on monocular visual image streams acquired by monocular image acquisition equipment, real-time dehazing algorithms, such as those based on dark channel priors, are used. By analyzing the characteristics and distribution patterns of fog in the image, pixel-by-pixel processing is performed to remove blur and noise caused by fog, restoring the original clear details of the image. Simultaneously, considering that image details in hazard source areas are crucial for accurate identification and location, super-resolution reconstruction is performed on the identified hazard source areas. Using deep learning algorithms, such as super-resolution reconstruction models based on convolutional neural networks, the mapping relationships between a large number of high-resolution and low-resolution images are learned. This processes the low-resolution images of hazard source areas, increasing the number of pixels and improving image resolution, thereby more clearly presenting the shape and features of the hazard source and providing a more reliable basis for subsequent analysis and judgment.
[0033] Because transmission towers are affected by factors such as wind and vibration, monocular image acquisition equipment experiences jitter, causing abrupt changes in the coordinates of target objects in the acquired images. This introduces significant errors into subsequent target localization and ranging. A time-series filtering method, such as the Kalman filter algorithm, is employed. This algorithm utilizes the correlation of image sequences in the time dimension, predicting the current state based on the state of the previous moment, and combining this with the current observation value for optimal estimation, thereby effectively eliminating coordinate jump errors caused by jitter. This processing step yields stable and accurate two-dimensional image information, providing a high-quality image data foundation for subsequently building a lightweight target detection model for edge computing devices and the entire depth vision-based monocular image ranging process for transmission line corridors. This ensures that the system can more accurately identify hazard sources, locate target hazards, and calculate their distances to transmission lines, improving the reliability and effectiveness of transmission line safety monitoring.
[0034] In one possible implementation, step S220 further includes: Step S221: Based on the installation location parameters of the transmission tower, establish a rigid transformation relationship between the pixel coordinate system and the transmission line corridor area.
[0035] Step S222: Correct the spatial distribution characteristics of the transmission conductors using the rigid transformation relationship from the pixel coordinate system to the transmission line corridor area.
[0036] Specifically, the installation location parameters of transmission towers encompass crucial information such as their latitude and longitude coordinates, height, and orientation. This data precisely determines the tower's spatial position within the actual transmission line corridor area. Based on these parameters, the first step is to clarify the benchmark differences between the pixel coordinate system and the actual spatial coordinate system. Using trigonometric functions and vector operations, a rotation matrix is calculated. This matrix describes the angle of rotation of the pixel coordinate system around different coordinate axes, ensuring that its coordinate axis directions align with the spatial coordinate axes of the transmission line corridor area. Simultaneously, a translation vector is determined based on the tower's spatial location information to compensate for the difference in the origin position between the pixel coordinate system and the actual spatial coordinate system. Combining the rotation matrix and translation vector, a rigid transformation relationship is constructed from the pixel coordinate system to the transmission line corridor area. This rigid transformation relationship allows for accurate mapping of two-dimensional pixels in the pixel coordinate system to the three-dimensional space of the transmission line corridor. This provides a reliable spatial transformation basis for subsequent operations such as accurately processing the spatial distribution characteristics of transmission conductors and determining the spatial coordinates of target hazards, ensuring the accuracy and effectiveness of the entire ranging system.
[0037] After establishing a rigid transformation relationship from the pixel coordinate system to the transmission line corridor area, this relationship is used to correct the spatial distribution characteristics of the transmission conductors. Due to measurement errors, equipment installation deviations, and environmental factors, the initially acquired spatial distribution characteristics of the transmission conductors may contain deviations. The rigid transformation relationship contains accurate rotation and translation information between the pixel coordinate system and the actual transmission line corridor area. Based on this relationship, the point cloud data of the transmission conductors in the pixel coordinate system is angularly adjusted using a rotation matrix to conform to the actual spatial orientation. Then, translation vectors are used to offset the position of these point cloud data to match the actual spatial location. In this way, deviations caused by various factors can be effectively eliminated, making the spatial distribution characteristics of the transmission conductors more consistent with the actual situation.
[0038] In one possible implementation, step S222 further includes: Step S2221: Obtain the raw point cloud data of the power transmission corridor.
[0039] Step S2222: Filter the original point cloud data to remove vegetation interference points and terrain interference points.
[0040] Step S2223: Divide the filtered point cloud data into multiple independent blocks according to the location of the power transmission tower, and convert them into an encrypted and compressed format that can be parsed by the edge computing device.
[0041] Specifically, using lidar equipment, the emitted laser beams perform a comprehensive scan of the power transmission corridor area at extremely high frequencies. The laser beams reflect off surfaces, and based on the time difference between emission and reception, as well as the laser's propagation speed, the spatial coordinates of each reflection point are precisely calculated. These coordinates are aggregated to form richly detailed raw point cloud data. This data not only covers key information such as the location and shape of the power transmission lines, but also includes data on vegetation, topography, and other potential objects within the transmission corridor. During the scanning process, the lidar operates according to a pre-set scanning path and parameters, ensuring that the acquired data is comprehensive and accurate.
[0042] After acquiring the raw point cloud data of the power transmission corridor, it is filtered to improve data quality and ensure the accuracy of subsequent analysis. The raw point cloud data contains point cloud information of vegetation and terrain, which can interfere with the analysis of the spatial distribution characteristics of the transmission lines and therefore must be removed. Filtering is performed using various characteristics of the point cloud data. Based on spatial geometric features, vegetation point clouds exhibit a scattered and irregular distribution, which is significantly different from the relatively regular and concentrated distribution of the transmission lines. Terrain point clouds show specific patterns in height and slope variations, differing considerably from the spatial location of the transmission lines. Clustering algorithms, such as the DBSCAN algorithm, are used to classify the data based on point cloud density, effectively distinguishing low-density vegetation point cloud clusters. A height threshold-based filtering method is employed to filter out terrain point clouds that are significantly below or above the normal erection height range of the transmission lines. Through the comprehensive application of these methods, vegetation and terrain interference points are effectively removed from the raw point cloud data, providing a more reliable data foundation and ensuring the efficient operation of the entire power transmission line safety monitoring system.
[0043] Based on the precise location information of the transmission towers, the filtered point cloud data is divided into multiple independent blocks. This division ensures that each block corresponds to a specific transmission tower section, facilitating subsequent detailed analysis of each section independently and improving processing efficiency and focus. Subsequently, to ensure efficient storage, transmission, and security of the data on edge computing devices, these independent blocks need to be converted into an encrypted and compressed format that can be parsed by the edge computing devices. In the encryption stage, encryption algorithms suitable for the processing capabilities of edge computing devices, such as the lightweight AES algorithm, are used to encrypt the data, preventing theft or tampering during transmission and storage. In the compression stage, a wavelet transform-based compression algorithm is used to significantly reduce the data volume while preserving key data features as much as possible, thus reducing data transmission bandwidth requirements and storage space requirements. The encrypted and compressed data meets the data security requirements of edge computing devices and can be quickly parsed and processed by the devices, laying a data foundation for subsequent tasks such as target detection and distance calculation on the edge computing devices.
[0044] In one possible implementation, step S400 further includes: Step S410: Use the Euclidean distance between the target hazard and the transmission line as the real-time spatial distance.
[0045] Step S420: Simultaneously, the conductor swing parameter is introduced to dynamically compensate for the Euclidean distance, and the compensation amount is dynamically adjusted according to the real-time wind speed data.
[0046] Specifically, after obtaining the coordinates of the target hazardous object in a three-dimensional point cloud space and the coordinates of the corresponding reference point on the transmission line, the Euclidean distance formula is used for calculation. The Euclidean distance formula in three-dimensional space is expressed as follows: ,in, This represents the Euclidean distance between the target hazardous object and the corresponding reference point on the transmission line, that is, the straight-line distance between these two points in three-dimensional space. , , ) represents the spatial coordinates of the target hazardous object, ( , , () represents the coordinates of a reference point selected on the transmission line. By substituting these coordinate values into the calculation, the straight-line distance between the target hazard and the transmission line can be obtained. This distance is the preliminary real-time spatial distance.
[0047] A machine learning algorithm based on Long Short-Term Memory (LSTM) networks is employed to dynamically compensate for Euclidean distance. First, historical wind speed data, conductor sway amplitude, and corresponding Euclidean distances are collected as training datasets. The data is organized into multi-dimensional feature vectors according to time series. The input layer of the LSTM network receives vectors containing features such as real-time wind speed, historical wind speed sequences, and historical compensation distances. Through the memory units in the hidden layers capturing the long-short-term dependencies, the complex mapping between wind speed changes and conductor sway amplitude is learned. During training, the actual measured conductor sway compensation distance is used as the label, and the mean squared error (MSE) is used as the loss function. The network parameters are optimized using a backpropagation algorithm. When new real-time wind speed data is input, the trained LSTM model predicts the corresponding conductor sway compensation amount, which is then superimposed on the Euclidean distance to achieve dynamic correction of the real-time spatial distance. This algorithm can adaptively learn the nonlinear relationship between different meteorological conditions and conductor sway, and compared to traditional empirical models, it can more accurately adjust the compensation amount according to real-time wind speed, improving the accuracy of transmission line safety monitoring.
[0048] In one possible implementation, step S100 further includes: Step S110: Use MobileNetV3 as the backbone network for feature extraction, and add a spatial attention unit after the last convolutional layer.
[0049] Step S120: Based on the feature extraction backbone network, transfer the pre-trained model parameters to the lightweight model, and use the spatial attention unit to quantize the model when it is deployed on the edge computing device.
[0050] Specifically, when constructing the hazard detection model for power transmission line corridors, considering the limited computing resources and storage capacity of edge computing devices, MobileNetV3 was selected as the backbone network for feature extraction. MobileNetV3, through its depthwise separable convolutional structure, decomposes standard convolution into depthwise convolution and pointwise convolution, significantly reducing computational cost and parameter count. Simultaneously, it employs optimized activation functions such as h-swish to maintain feature extraction performance while reducing power consumption. To further enhance the model's sensitivity to hazard sources in power transmission scenarios, a Spatial Attention Module was added after the last convolutional layer of MobileNetV3. This module analyzes the spatial dimensions of the feature map, calculates weight coefficients at different locations, automatically focuses on the hazard source area, suppresses background noise interference, and enhances the feature extraction capability for hazards such as foreign objects and tree growth around the transmission lines, thus providing more discriminative feature representations for subsequent target detection tasks.
[0051] After completing the construction of the feature extraction backbone network (MobileNetV3 combined with spatial attention units), the focus shifted to adapting the model to edge computing devices. First, transfer learning techniques were employed to transfer model parameters pre-trained on large-scale general datasets (such as ImageNet) to a lightweight model for hazard detection in power transmission line corridors. This parameter transfer strategy leverages the basic image features (such as edges and textures) already learned by the pre-trained model, avoiding overfitting risks under limited power transmission scenario data, while significantly reducing training time and resource consumption. During deployment, the model quantization operation was performed using the feature focusing capabilities of the spatial attention units. Model parameters were converted from their original high-precision data types (such as 32-bit floating-point numbers) to low-precision representations (such as 8-bit integers), compressing the model size and reducing computational complexity. During this process, the spatial attention units continuously strengthened the feature weights of hazard source areas, ensuring that the accuracy loss caused by quantization was mainly concentrated in non-critical background areas, while the impact on the detection accuracy of dangerous targets such as foreign objects and vegetation remained controllable. Finally, a lightweight model adapted to edge computing devices was generated, meeting real-time detection requirements while operating efficiently in the hardware environment of tower ends with limited computing power and storage.
[0052] In one possible implementation, step S120 further includes: Step S121: Establish a shared memory channel between the first chip associated with the edge computing device and the second chip associated with the lightweight target detection model, and determine the first ranging delay factor.
[0053] Step S122: Based on task priority, dynamically allocate computing resources between the first chip and the second chip, determine the second ranging delay factor, and combine it with the first ranging delay factor to perform ranging task response optimization processing under delay compensation.
[0054] Specifically, a shared memory channel is established between the first chip mounted on the edge computing device and the second chip running the lightweight target detection model. The first chip can be any chip with hardware acceleration capabilities, such as an FPGA, while the second chip can be any chip adept at neural network operations, such as a GPU or NPU. This channel utilizes a high-speed interconnect bus (such as PCIe Gen4). This channel employs a dual-port RAM architecture, allowing both the first and second chips to access the same memory space simultaneously, avoiding the overhead of multiple copying in traditional data transmission. After the channel is established, metrics such as data throughput, bus arbitration latency, and memory access conflict frequency are monitored in real time to quantify the latency caused by data interaction between the chips. For example, by recording the transmission time of the first chip's preprocessed data before the lightweight model's inference and the waiting period for the second chip to read shared memory, the first ranging latency factor is comprehensively determined. This factor reflects the impact of data communication between heterogeneous chips on the overall ranging response speed, providing key benchmark parameters for subsequent system latency optimization.
[0055] The transmission line monitoring tasks are pre-classified into three categories: emergency alarms (such as foreign objects touching the conductor), high-priority inspections (tracking nearby hazardous objects), and routine inspections (periodic environmental scanning), with priority weights assigned to each task. During runtime, based on this priority system, the computing resources of the first chip (including FPGA) and the second chip (including GPU and NPU) are dynamically allocated through the resource scheduling module of the edge computing device: the proportion of hardware acceleration logic units for high-priority tasks is configured for the FPGA, and the allocation ratio of its internal DSP array and BRAM cache is adjusted; the number of computing cores, memory bandwidth, and inference threads corresponding to the priority tasks are allocated to the GPU / NPU. During the resource allocation process, the switching time, core utilization, and data waiting time of different tasks between chips are continuously monitored, and the second ranging delay factor is obtained by quantifying these parameters. Subsequently, the gradient descent iterative method was adopted, with the weighted sum of the first and second ranging delay factors as the optimization objective. Each iteration cycle was 100ms. By adjusting the cache prefetch strategy of the shared memory channel and the chip resource allocation ratio, the total delay fluctuation was controlled within ±5ms. After 3-5 iterations, the dynamic balance of the two types of delay factors was achieved, ensuring that the real-time response performance of the ranging system remained stable within the preset threshold.
[0056] In one possible implementation, step S2223 further includes: Step S22231: When the edge computing device loads point cloud data in encrypted and compressed format, it uses hash check to compare the topological structure features of the corresponding independent block and M adjacent independent blocks, where M is greater than or equal to 3.
[0057] Step S22232: If the hash value of the consistency between the spatial curvature distribution and the normal vector is less than the preset deviation value, then continue to perform the decompression operation of the point cloud data.
[0058] Specifically, when edge computing devices process transmission corridor monitoring data, they focus on verifying the integrity of encrypted compressed point cloud data. When the device loads such data, a hash verification mechanism is automatically initiated to compare the topological features of the current independent block and at least three adjacent independent blocks. In practice, by analyzing the spatial distribution of the point cloud within each block, geometric features such as distances and angles between points are extracted to construct a topological descriptor for the block. Subsequently, these topological features are encoded using a hash function (such as SHA-256) to generate a fixed-length hash value. This process not only verifies the accuracy of data in a single block but also ensures the spatial continuity and consistency of the point cloud data by associating it with the topological features of surrounding blocks. This effectively mitigates the risks of data loss, tampering, or compression distortion during data transmission, laying a reliable foundation for subsequent point cloud data decompression and spatial distribution analysis of transmission lines.
[0059] In the processing flow of encrypted compressed point cloud data by edge computing devices, data availability is ensured through geometric feature verification. First, the spatial curvature distribution (quantifying morphological changes such as conductor bends and terrain undulations) of the current independent block and adjacent areas of the point cloud is calculated to match the normal vector (evaluating surface smoothness). These features are then encoded into fixed-length hash values using a hash function (such as SHA-256). Subsequently, this hash value is compared with a pre-stored benchmark value. If the deviation is less than a preset threshold (e.g., Hamming distance < 0.05%), the data is determined not to have undergone significant distortion during encrypted transmission, and decompression can continue to release the data for subsequent analysis. Conversely, if the deviation exceeds the threshold, a data retransmission mechanism is immediately triggered, bypassing the decompression step and preventing misjudgments of hazardous objects in power transmission lines due to distorted data. This process balances data integrity and processing efficiency through a dual strategy, ensuring that edge computing devices acquire accurate and reliable point cloud data.
[0060] Example 2 is based on the same inventive concept as the monocular image ranging method for power transmission line corridors based on depth vision in the previous examples, such as... Figure 2 As shown, this application provides a monocular image ranging system for power transmission line corridors based on depth vision. The system and method embodiments in this application are based on the same inventive concept. The system includes: The hazard identification module 10 is used to acquire two-dimensional image information of the transmission line corridor area using a monocular image acquisition device deployed on the transmission tower, and to construct a lightweight target detection model for edge computing devices. The module identifies hazards in the two-dimensional image information and outputs a set of two-dimensional pixel coordinates with hazard category labels.
[0061] The relation matrix establishment module 20 is used to import the three-dimensional point cloud data of the transmission line corridor into the edge computing device and establish a mapping relation matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system.
[0062] The spatial coordinate information positioning module 30 is used to locate the spatial coordinate information of the target hazard in the three-dimensional point cloud space by using the mapping relationship matrix and a two-dimensional pixel coordinate set with hazard source category markings. The spatial coordinate information is associated and mapped with each element of the two-dimensional pixel coordinate set.
[0063] The spatial distance determination module 40 is used to determine the real-time spatial distance between the target hazard and the transmission line based on the spatial coordinate information, and to determine whether the real-time spatial distance is less than a preset safety threshold.
[0064] The graded alarm signal generation module 50 is used to trigger the local early warning mechanism to generate a graded alarm signal if the judgment result is yes.
[0065] Furthermore, the system is also used to implement the following functions: Based on the camera focal length and principal point offset parameters of the monocular image acquisition device, the spatial distribution characteristics of the transmission conductors in the three-dimensional point cloud data of the transmission line corridor are used, combined with the camera focal length and principal point offset parameters, to correct distortion errors. Simultaneously, M uniformly distributed transmission tower feature points are selected from the three-dimensional point cloud data of the transmission line corridor, and the first coordinate information in the three-dimensional point cloud spatial coordinate system and the second coordinate information in the pixel coordinate system are obtained. A nonlinear least squares equation system is constructed, where M is greater than or equal to 6. Projection error control is performed based on the nonlinear least squares equation system.
[0066] Furthermore, the system is also used to implement the following functions: Based on the monocular visual image stream acquired by the monocular image acquisition device, real-time dehazing processing is performed, and super-resolution reconstruction is performed on the identified hazard source areas; at the same time, coordinate jump errors caused by jitter are eliminated by time series filtering to obtain the two-dimensional image information.
[0067] Furthermore, the system is also used to implement the following functions: Based on the installation location parameters of the transmission towers, a rigid transformation relationship is established from the pixel coordinate system to the transmission line corridor area; the spatial distribution characteristics of the transmission conductors are corrected using the rigid transformation relationship from the pixel coordinate system to the transmission line corridor area.
[0068] Furthermore, the system is also used to implement the following functions: The raw point cloud data of the power transmission corridor is acquired; the raw point cloud data is filtered to remove vegetation interference points and terrain interference points; the filtered point cloud data is divided into multiple independent blocks according to the location of the power transmission towers and converted into an encrypted and compressed format that can be parsed by the edge computing device.
[0069] Furthermore, the system is also used to implement the following functions: The Euclidean distance between the target hazard and the transmission line is used as the real-time spatial distance; at the same time, the conductor swing parameter is introduced to dynamically compensate for the Euclidean distance, and the compensation amount is dynamically adjusted according to the real-time wind speed data.
[0070] Furthermore, the system is also used to implement the following functions: MobileNetV3 is used as the feature extraction backbone network, and a spatial attention unit is added after the last convolutional layer. Based on the feature extraction backbone network, the pre-trained model parameters are transferred to the lightweight model, and the spatial attention unit is used to quantize the model when it is deployed on the edge computing device.
[0071] Furthermore, the system is also used to implement the following functions: A shared memory channel is established between the first chip associated with the edge computing device and the second chip associated with the lightweight target detection model to determine a first ranging delay factor. Based on task priority, computing resources of the first chip and the second chip are dynamically allocated to determine a second ranging delay factor. Combined with the first ranging delay factor, ranging task response optimization processing under delay compensation is performed.
[0072] Furthermore, the system is also used to implement the following functions: When the edge computing device loads point cloud data in encrypted and compressed format, it uses hash verification to compare the topological features of the corresponding independent block and M adjacent independent blocks, where M is greater than or equal to 3. If the hash value of the spatial curvature distribution and the normal vector are consistent with the pre-stored reference value, the deviation value is less than the preset deviation value, then the decompression operation of the point cloud data continues.
[0073] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
[0074] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
[0075] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.
Claims
1. A monocular image ranging method for power transmission line corridors based on depth vision, characterized in that, include: Using monocular image acquisition equipment deployed on transmission towers, two-dimensional image information of the transmission line corridor area is acquired, and a lightweight target detection model for edge computing devices is constructed to identify hazard sources in the two-dimensional image information and output a set of two-dimensional pixel coordinates with hazard source category labels. Import the 3D point cloud data of the power transmission line corridor into the edge computing device and establish a mapping matrix between the pixel coordinate system and the 3D point cloud spatial coordinate system; Using the mapping matrix, the spatial coordinate information of the target hazard in the three-dimensional point cloud space is located through the two-dimensional pixel coordinate set with hazard source category markings. The spatial coordinate information is associated and mapped with each element of the two-dimensional pixel coordinate set. Based on the spatial coordinate information, determine the real-time spatial distance between the target hazard and the transmission line, and determine whether the real-time spatial distance is less than a preset safety threshold. If the judgment result is yes, the local early warning mechanism is triggered to generate a graded alarm signal.
2. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 1, characterized in that, The method further includes establishing a mapping matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system, and further comprising: Based on the camera focal length and principal point offset parameters of the monocular image acquisition device, the distortion error is corrected by utilizing the spatial distribution characteristics of the transmission conductors in the three-dimensional point cloud data of the transmission line corridor, combined with the camera focal length and principal point offset parameters. Meanwhile, M uniformly distributed transmission tower feature points are selected from the three-dimensional point cloud data of the transmission line corridor, and the first coordinate information in the three-dimensional point cloud spatial coordinate system and the second coordinate information in the pixel coordinate system are obtained to construct a nonlinear least squares equation system, wherein M is greater than or equal to 6. Projection error control is performed based on the aforementioned nonlinear least squares equations.
3. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 2, characterized in that, The method further includes: Based on the monocular visual image stream acquired by the monocular image acquisition device, real-time dehazing processing is performed, and super-resolution reconstruction is performed on the identified hazardous source areas. Meanwhile, time series filtering is used to eliminate coordinate jump errors caused by jitter, thus obtaining the two-dimensional image information.
4. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 3, characterized in that, In conjunction with the camera focal length and principal point offset parameters, distortion errors are corrected, and the method further includes: Based on the installation location parameters of the transmission towers, a rigid transformation relationship is established from the pixel coordinate system to the transmission line corridor area; The spatial distribution characteristics of the power transmission lines are corrected by using the rigid transformation relationship from the pixel coordinate system to the transmission line corridor area.
5. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 4, characterized in that, Before correcting the spatial distribution characteristics of transmission lines, the method further includes: Obtain the raw point cloud data of the power transmission corridor; The original point cloud data is filtered to remove vegetation interference points and terrain interference points. The filtered point cloud data is divided into multiple independent blocks according to the location of the power transmission towers, and then converted into an encrypted and compressed format that can be parsed by the edge computing device.
6. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 1, characterized in that, Based on the spatial coordinate information, the real-time spatial distance between the target hazardous object and the transmission line is determined, the method comprising: The real-time spatial distance is defined as the Euclidean distance between the target hazardous object and the transmission line. Meanwhile, a conductor swing parameter is introduced to dynamically compensate for the Euclidean distance, and the compensation amount is dynamically adjusted according to real-time wind speed data.
7. The method for monocular image ranging of transmission line corridors based on depth vision as described in claim 1, characterized in that, Constructing a lightweight target detection model for edge computing devices, the method includes: MobileNetV3 is used as the backbone network for feature extraction, and a spatial attention unit is added after the last convolutional layer. Based on the feature extraction backbone network, the pre-trained model parameters are transferred to the lightweight model, and the spatial attention unit is used to quantize the model when it is deployed on the edge computing device.
8. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 7, characterized in that, The method includes: A shared memory channel is established between the first chip associated with the edge computing device and the second chip associated with the lightweight target detection model to determine the first ranging delay factor. Based on task priority, the computing resources of the first chip and the second chip are dynamically allocated, the second ranging delay factor is determined, and the ranging task response is optimized under delay compensation in combination with the first ranging delay factor.
9. The monocular image ranging method for power transmission line corridors based on depth vision as described in claim 8, characterized in that, The method includes converting the data into an encrypted compression format that the edge computing device can parse, and the conversion includes: When the edge computing device loads point cloud data in encrypted and compressed format, it uses hash check to compare the topological structure features of the corresponding independent block and M adjacent independent blocks, where M is greater than or equal to 3. If the hash value of the spatial curvature distribution and the normal vector are consistent with the pre-stored reference value, the deviation value is less than the preset deviation value, then the decompression operation of the point cloud data continues.
10. A monocular image ranging system for power transmission line corridors based on depth vision, characterized in that, The system is used to implement the monocular image ranging method for power transmission line corridors based on depth vision as described in any one of claims 1-9, and the system comprises: The hazard identification module is used to acquire two-dimensional image information of the transmission line corridor area using a monocular image acquisition device deployed on the transmission tower, and to construct a lightweight target detection model for edge computing devices. The module then identifies hazard sources in the two-dimensional image information and outputs a set of two-dimensional pixel coordinates with hazard category labels. The relation matrix establishment module is used to import the three-dimensional point cloud data of the transmission line corridor into the edge computing device and establish a mapping relation matrix between the pixel coordinate system and the three-dimensional point cloud spatial coordinate system. The spatial coordinate information positioning module is used to locate the spatial coordinate information of the target hazard in the three-dimensional point cloud space by using the mapping relationship matrix and a two-dimensional pixel coordinate set with hazard source category markings. The spatial coordinate information is associated and mapped with each element of the two-dimensional pixel coordinate set. The spatial distance determination module is used to determine the real-time spatial distance between the target hazard and the transmission line based on the spatial coordinate information, and to determine whether the real-time spatial distance is less than a preset safety threshold. The graded alarm signal generation module is used to trigger the local early warning mechanism to generate graded alarm signals if the judgment result is yes.