Intelligent identification method and device for power transmission line construction disturbance
By using the improved Unet3P model and binocular vision technology, the efficiency and accuracy issues of construction disturbance identification for power transmission lines have been resolved. This enables construction disturbance identification and real-time alarms in complex scenarios, thereby enhancing the safe operation and maintenance capabilities of power transmission lines.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-03-27
AI Technical Summary
Traditional methods for identifying construction disturbances in power transmission lines are inefficient and lack accuracy. They are unable to distinguish between construction disturbances and changes in the natural environment in complex scenarios and cannot meet the needs of real-time inspection.
An improved Unet3P model is adopted, which combines binocular vision data acquisition, image preprocessing and stereo correction. Through depth calculation and feature extraction, combined with threshold segmentation, morphological processing and semantic segmentation methods, the construction disturbance area is identified, and quantitative analysis and risk assessment are carried out to achieve real-time alarm.
It improves the accuracy and efficiency of construction disturbance identification, can work stably in complex environments, achieves minute-level response, reduces operation and maintenance costs, and improves the safety of transmission lines.
Smart Images

Figure CN121746900A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent monitoring of power systems, in particular to an intelligent identification method and device for construction disturbance of a power transmission line, and belongs to the technical field of operation state monitoring and intelligent analysis of power transmission lines. BACKGROUND
[0002] Generally, as a core component of the power system, the safe and stable operation of the power transmission line is directly related to the social production and living order and the safety of energy supply. In the operation and maintenance process of the power transmission line, the construction disturbance (such as illegal excavation, mechanical operation, and load operation) in the line corridor range is an important hidden danger causing line faults, which may lead to problems such as line tower tilting, conductor wear, and insulation failure, and even cause large-area power outage accidents, causing huge economic losses and social impact.
[0003] Traditional identification of construction disturbance of the power transmission line mainly relies on manual inspection and conventional remote sensing image analysis, but both of these two methods have obvious defects: manual inspection is limited by factors such as terrain (e.g. high mountains, Gobi), climate (e.g. sand, rain and snow), and is not only low in efficiency and limited in coverage, but also has a problem of response lag, which is difficult to meet the needs of large-scale and high-frequency inspection; conventional remote sensing image analysis mostly uses traditional image processing methods such as threshold segmentation and edge detection, which have low recognition accuracy of disturbance targets in complex scenes, and are difficult to effectively distinguish between construction disturbance and natural environmental changes (such as vegetation growth and soil erosion), which cannot meet the requirements of modern power transmission line lean operation.
[0004] With the rapid development of deep learning technology in the field of computer vision, semantic segmentation models provide a new technical path for high-precision target recognition. Unet3P, as an advanced semantic segmentation network, through multi-scale feature fusion and deep semantic mining, has shown excellent performance in medical image segmentation and remote sensing target recognition. However, directly applying the original Unet3P to the identification of construction disturbance of the power transmission line still has three key problems: first, the feature extraction of small-scale disturbance targets (such as small construction machinery and initial excavation marks) is insufficient, which is prone to missed detection; second, the interference suppression capability of complex background (such as vegetation obstruction and terrain undulation) is weak, which is prone to misjudgment; third, the model reasoning speed is slow, which is difficult to meet the real-time inspection requirements, and needs to be improved and optimized in view of the power transmission line scene. SUMMARY
[0005] The application proposes a power transmission line construction disturbance intelligent identification method based on improved Unet3P, and the core goal is to solve the defects of traditional identification methods and improve the accuracy and efficiency of construction disturbance identification in complex scenes. Through optimization and improvement of the Unet3P model, the extraction and identification ability of the model to the construction disturbance features of the power transmission line in complex terrain is enhanced, and the precise positioning and type discrimination of the construction disturbance area are realized, such as illegal excavation area, mechanical operation area and material stacking area.
[0006] In order to solve the above problems, according to the intelligent identification method for power transmission line construction disturbance provided by the application, the method comprises the following steps:
[0007] S1, binocular vision data acquisition step, the unmanned aerial vehicle carrying binocular camera flies along the power transmission line corridor, and high-resolution binocular image data is collected;
[0008] S2, image preprocessing and stereoscopic correction step, denoising, contrast enhancement and stereoscopic correction are performed on the binocular image, so as to eliminate distortion and improve matching accuracy;
[0009] S3, depth calculation and feature extraction step, the improved Unet3P model is used to extract multi-scale features of the stereoscopic corrected binocular image, establish the parallax relationship and calculate the depth map;
[0010] S4, construction disturbance area identification step, based on the depth map and feature map, threshold segmentation, morphological processing and semantic segmentation method are used to identify the construction disturbance area, and the disturbance type, contour and spatial boundary are output;
[0011] S5, analysis and processing step, according to the identification result, the quantitative parameters such as the area, depth or height, distance from the tower of the disturbance area are calculated, and the risk level is divided according to the line safety threshold;
[0012] S6, alarm processing step, when the risk level is medium risk or high risk, the alarm is triggered and the disturbance information is pushed to the operation and maintenance system, realizing the alarm and disposal closed loop, the disturbance information includes disturbance type, latitude and longitude coordinates, area, risk level and corresponding image.
[0013] According to another embodiment of the present invention, an intelligent identification device for construction disturbances of power transmission lines is provided, characterized in that it includes: a binocular vision data acquisition module for acquiring high-resolution binocular image data by flying along the power transmission line corridor using a drone equipped with a binocular camera; an image preprocessing and stereo correction module for denoising, contrast enhancement, and stereo correction of the binocular images to eliminate distortion and improve matching accuracy; and a depth calculation and feature extraction module for utilizing an improved Unet3P... The model performs multi-scale feature extraction on stereo-corrected binocular images, establishes disparity relationships, and calculates depth maps. A construction disturbance area identification module identifies construction disturbance areas based on the depth and feature maps, using threshold segmentation, morphological processing, and semantic segmentation methods, and outputs the disturbance type, outline, and spatial boundary. An analysis and processing module calculates quantitative parameters such as the area, depth or height, and distance from the tower of the disturbance area based on the identification results, and classifies the risk level according to the line safety threshold. An alarm processing module triggers alarms and pushes disturbance information to the operation and maintenance system when the risk level is medium or high, achieving a closed loop of alarm and response. The disturbance information includes disturbance type, latitude and longitude coordinates, area, risk level, and corresponding image.
[0014] This invention has several advantages: First, it significantly improves the efficiency and accuracy of disturbance identification, reducing the time from hazard discovery to alarm from several hours or even days in traditional manual inspections to minutes. Second, it provides timely and reliable decision support for the safe operation and maintenance of transmission lines, helping maintenance personnel to respond quickly to hazards and reduce the risk of safety accidents. Third, it reduces the blindness and workload of manual inspections, lowers operation and maintenance costs, and has significant engineering application value and economic and social benefits. Attached Figure Description
[0015] Figure 1 This paper illustrates a method for intelligent identification of construction disturbances in transmission lines according to an embodiment of the present invention.
[0016] Figure 2 An intelligent identification device for construction disturbances of transmission lines according to an embodiment of the present invention is shown.
[0017] Figure 3 An improved Unet3P model according to an embodiment of this disclosure is shown. Detailed Implementation
[0018] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings, so that those skilled in the art can readily understand the present invention.
[0019] Only the parts necessary for understanding the technical contents of the present application will be described herein, and the description of the remaining parts will be omitted so as not to obscure the gist of the present application, and attention should be paid to this. Also, in this process, the thickness of the lines or the size of the constituent elements shown in the drawings can be exaggerated for the sake of clarity and convenience of description.
[0020] The terms used herein are used to describe the embodiments and are not intended to limit and / or define the present application. When a certain constituent element is described as being "connected", "coupled", or "joined" to another constituent element, this includes not only a direct connection relationship but also an indirect connection relationship in which other constituent elements are present therebetween. Also, terms such as "include", "comprise", or "have" are intended to mean that the features, numbers, steps, operations, constituent elements, components or combinations thereof described in the specification are present, and do not exclude the presence or possibility of addition of one or more other features, numbers, steps, operations, constituent elements, components or combinations thereof. In addition, the terms "first", "second", and the like that can be referred to herein are used only for the purpose of distinguishing a certain constituent element from another constituent element, and do not limit the order or importance of the constituent elements in the absence of a specific mention. Therefore, in this document, a first constituent element in one embodiment can be referred to as a second constituent element in another embodiment, and similarly, a second constituent element in one embodiment can be referred to as a first constituent element in another embodiment.
[0021] Figure 1 A power transmission line construction disturbance intelligent identification method according to an embodiment of the present application is shown.
[0022] As Figure 1 shown, the present application provides a power transmission line construction disturbance intelligent identification method based on an improved Unet3P (UNet3+) model, which is suitable for remote areas (such as Xinjiang) with wide area, complex terrain, and variable climate, and can realize automatic detection, risk assessment, and alarm management of construction disturbance within the power transmission line corridor. The method comprises the following steps:
[0023] Step S1: binocular vision data acquisition
[0024] A drone equipped with a binocular camera is used to fly along the power transmission line corridor to collect high-resolution binocular image data.
[0025] The camera uses a parallel optical axis design, and the baseline length is accurately calibrated (10-20 cm).
[0026] The flight path is planned according to the line terrain to achieve high coverage rate acquisition.
[0027] Data is collected during sunny periods, and the collection frequency is increased for key areas (such as around the tower).
[0028] Step S2: Image preprocessing and stereo rectification
[0029] Preprocess and rectify the original binocular images to improve image quality and eliminate distortion.
[0030] Use the CNN denoising algorithm (DnCNN) to remove sensor noise and dust particle interference;
[0031] Use the adaptive histogram equalization method (CLAHE) to enhance image contrast;
[0032] Calculate the epipolar equation and perform re-projection based on the intrinsic and extrinsic matrices of the binocular camera, so that the corresponding points of the left and right images are located on the same horizontal line, achieving stereo rectification.
[0033] Step S3: Depth calculation and feature extraction
[0034] Based on the improved Unet3P model, perform depth calculation and feature extraction on the stereo rectified binocular images.
[0035] The model generates depth-enhanced feature maps through multi-scale feature extraction, residual connection, and attention mechanism;
[0036] Use the improved SIFT algorithm to establish the correspondence relationship between the feature points of the left and right images;
[0037] Based on the principle of triangulation and the disparity-depth conversion formula Z=f×B / d, calculate the pixel depth value and generate a depth map.
[0038] Step S4: Construction disturbance area identification
[0039] Identify the construction disturbance area based on the depth map and feature map.
[0040] Use the depth threshold segmentation method to screen potential disturbance areas;
[0041] Optimize the region boundary by combining morphological operations (dilation, erosion, opening and closing operations);
[0042] Use the improved Unet3P model for semantic segmentation to identify disturbance types (illegal excavation, mechanical operation, and material stacking);
[0043] Extract the region boundary through edge detection and contour fitting algorithms, and calculate quantitative features such as area and perimeter.
[0044] Step S5: Analysis and processing (disturbance data quantitative analysis and risk assessment)
[0045] Perform quantitative analysis and risk assessment on the identified disturbance areas.
[0046] Calculate disturbance parameters according to disturbance types and geometric characteristics (area, depth / height, distance from tower);
[0047] Divide into low, medium and high risk levels according to line safety thresholds;
[0048] Output disturbance category, quantitative index and risk level results.
[0049] Step S6: Alarm processing (risk alarm and disposal closed loop)
[0050] Convert risk information into alarm instructions to realize real-time operation and maintenance closed loop.
[0051] Automatically trigger an alarm when the disturbance is determined to be medium / high risk;
[0052] Implement mandatory alarm for critical line sections (such as 30 m around the tower);
[0053] Alarm information includes: disturbance occurrence time, latitude and longitude, type, risk level, area / depth parameters and visual image;
[0054] The system pushes the alarm to the cloud platform in real time through the unmanned aerial vehicle communication link to realize the closed loop of "detection-alarm-disposal-feedback".
[0055] Compared with the prior art, the method of the present application has the following significant advantages:
[0056] 1. High-precision identification: Achieve high-precision identification of construction disturbance targets within the transmission line corridor, including illegal excavation areas, construction machinery, and stacked materials, with an identification accuracy (mIoU) of not less than 85%.
[0057] 2. Small target optimization: Improve the model's ability to identify small-scale disturbance targets (such as small construction machinery and initial excavation marks) with a recall rate of not less than 80%.
[0058] 3. Real-time guarantee: Optimize model inference efficiency, with a single remote sensing image (1024x1024 pixels) inference time of not more than 0.5 seconds, meeting real-time inspection requirements.
[0059] 4. Environmental adaptability: Enhance the model's adaptability to complex environments (such as vegetation obstruction, light changes, and terrain differences) to ensure stable operation in different scenarios.
[0060] The transmission line construction disturbance intelligent identification device will be described in more detail below.
[0061] Figure 2 The transmission line construction disturbance intelligent identification device according to an embodiment of the present application is shown.
[0062] The embodiment provides a power transmission line construction disturbance intelligent identification device based on an improved Unet3P. The device comprises a binocular vision data acquisition module, an image preprocessing and stereoscopic correction module, a depth calculation and feature extraction module, a construction disturbance area identification module, and an analysis and alarm module. Each module can be independently optimized and work cooperatively to ensure the stability and reliability of the overall performance.
[0063] Binocular vision data acquisition module
[0064] In the embodiment, the binocular vision data acquisition module adopts a combination of "drone + binocular imaging equipment" to adapt to the characteristics of wide geographical distribution and complex terrain of the power transmission line in Xinjiang, and realizes large-range and high-precision data acquisition.
[0065] Equipment selection and calibration
[0066] The drone is equipped with a binocular camera with parallel optical axes, and the baseline length, that is, the distance between the two cameras, is accurately calibrated. The baseline length can be adjusted according to the inspection height, and the range is usually 10cm to 20cm, so as to ensure that enough parallax information is collected for subsequent depth calculation processing.
[0067] Data acquisition strategy
[0068] According to the spatial distribution range of the power transmission line corridor, the drone can perform inspection based on a preset autonomous flight path. The path can be realized by visual navigation, so as to adapt to areas with insufficient signal coverage. The drone collects high-resolution raw binocular images during flight, and the image resolution is usually 4K or higher, so as to ensure that the details of the distant construction equipment and the disturbance area are clear and visible, and provide high-quality input data for subsequent identification.
[0069] Data coverage guarantee
[0070] In view of the characteristics of frequent sand-dust weather and complex climate in Xinjiang region, the data acquisition module performs acquisition task in sunny period to avoid the influence of climate condition on image quality. Meanwhile, key areas (such as the periphery of power transmission tower, line crossing section, etc.) are collected with emphasis, and the collection frequency is appropriately increased, so as to ensure the integrity of data and the sufficiency of coverage range.
[0071] Through the above design, the binocular vision data acquisition module of the embodiment can provide high-precision and full-coverage visual data basis for subsequent stereoscopic correction, depth calculation and disturbance identification, so as to significantly improve the robustness and identification accuracy of the whole system.
[0072] Image preprocessing and stereoscopic correction module
[0073] In this embodiment, the image preprocessing and stereo rectification module is used to optimize the quality and correct the geometry of the original binocular images acquired by the binocular vision data acquisition module, so as to eliminate the adverse effects of noise, distortion and uneven illumination on image quality and feature matching accuracy, thereby providing high-quality input data for subsequent depth calculation and feature extraction.
[0074] Image denoising
[0075] The original binocular images often contain random noise introduced by the imaging sensor and particle noise caused by environmental dust. Therefore, this embodiment uses a denoising algorithm based on a convolutional neural network (such as DnCNN) to effectively remove noise while preserving the edge and texture details of the image, avoiding noise interference in the feature extraction process and improving the signal-to-noise ratio and clarity of the overall image.
[0076] Image enhancement
[0077] To address the brightness unevenness problem caused by the drastic changes in lighting conditions in Xinjiang (such as strong light at noon and shadows in the morning), this embodiment uses a self-adaptive histogram equalization algorithm (CLAHE) for image enhancement processing. This algorithm can adaptively adjust the local contrast and improve the visibility of details in over-bright or over-dark areas, thereby improving the consistency of target recognition under different lighting conditions and ensuring the stability of subsequent network feature extraction.
[0078] Stereo rectification
[0079] To achieve accurate matching of binocular images, this embodiment calculates the epipolar constraint equation based on the intrinsic matrix of the binocular camera (including focal length, pixel size, distortion coefficient) and the extrinsic matrix (including the relative position and attitude parameters between the two cameras), and performs re-projection processing on the left and right images to make the corresponding points lie on the same horizontal line. Then, a bilinear interpolation method is used to resample the image to obtain the stereo rectified binocular image. Through this stereo rectification process, the search dimension of subsequent feature matching can be significantly reduced, improving the efficiency and accuracy of matching operations.
[0080] Through the above processing, the image preprocessing and stereo rectification module can effectively improve the geometric consistency and visual quality of the input data, providing a reliable foundation for the accurate operation of the depth calculation and feature extraction module.
[0081] Depth calculation and feature extraction module
[0082] In this embodiment, the depth calculation and feature extraction module is used to calculate the pixel depth information from the stereo rectified binocular images to generate a high-precision depth map, providing data support for three-dimensional quantitative analysis of the construction disturbed area (such as excavation depth, pile height, etc.). This module constitutes the core part of the binocular vision system, and its specific processing flow is as follows:
[0083] Feature extraction and enhancement
[0084] In this embodiment, a deep feature extraction model based on improved Unet3P is adopted to perform multi-scale feature extraction on the stereo corrected binocular images and obtain deep enhanced feature maps. By introducing residual connection mechanism and attention mechanism into the Unet3P structure, the significant features related to depth estimation (such as target edge contour, texture difference) can be strengthened, and the interference of irrelevant background features can be suppressed, thereby improving the robustness and accuracy of depth estimation.
[0085] The mathematical expression of multi-scale feature extraction is as follows:
[0086]
[0087] where f represents the depth feature map, l represents the network layer number, w(l) is the weight parameter of the lth layer, b(l) is the bias term, i(l) represents the feature representation of the image at the lth layer, the symbol "*" represents convolution operation, relu(·) is the ReLU activation function, w_r(l) and b_r(l) are the weight and bias parameters of the residual connection respectively.
[0088] Feature matching
[0089] In order to establish the feature correspondence between the left and right eye images, this embodiment adopts an improved SIFT feature matching algorithm. This algorithm selects robust feature points based on the deep enhanced feature map, and determines the best matching pair based on the similarity of feature descriptors. By introducing feature constraint and outlier rejection mechanism, the mis-matching phenomenon caused by environmental factors such as dust, vegetation and shadow can be effectively reduced. The mathematical expression of feature matching process is as follows:
[0090]
[0091] where M(x_l) represents the best matching point corresponding to the left eye feature point x_l, D_l and D_r are the feature descriptors of the left and right eye feature points respectively, and ||·|| represents the Euclidean distance.
[0092] Depth calculation
[0093] After completing the feature matching, this embodiment calculates the disparity information between the left and right eye feature points based on the principle of triangulation, and uses the "disparity-depth" conversion formula to obtain the pixel depth value, thereby generating a high-precision depth map. The conversion formula is as follows:
[0094]
[0095] Wherein, Z is the pixel depth value (unit: meter), f is the camera focal length (unit: pixel), B is the baseline length (unit: meter), and d is the parallax value (unit: pixel).
[0096] Through the above steps, the depth calculation and feature extraction module can accurately restore the three-dimensional spatial structure of the construction scene, providing reliable depth information basis for intelligent identification and quantitative analysis of the construction disturbance area.
[0097] Construction disturbance area identification module
[0098] In this embodiment, the construction disturbance area identification module is used to accurately identify, type distinguish and quantitatively analyze the construction disturbance area along the transmission line based on the depth information output by the depth calculation and feature extraction module, combined with the improved Unet3P model. This module realizes automatic and intelligent disturbance detection through multi-stage processing, significantly improving the identification accuracy and environmental adaptability. The specific steps are as follows:
[0099] Potential disturbance area screening
[0100] This embodiment first classifies the pixels in the scene based on depth map data using a depth threshold segmentation algorithm. The system sets the depth threshold parameters based on historical statistical data and field measurement results, for example, marking areas with depth values exceeding 0.5m as potential excavation areas and areas higher than 1m as loading areas. Through the above threshold segmentation, potential engineering disturbance areas can be preliminarily screened out, effectively excluding normal terrain areas and non-construction activity areas.
[0101] Region optimization processing
[0102] To eliminate noise and morphological discontinuity in the preliminary screening results, this embodiment uses various morphological operations (including dilation, erosion, opening operation and closing operation) to optimize the potential disturbance area. Through this process, the following goals can be achieved: removing small isolated noise areas (such as single pixel points), filling disturbance boundary holes (such as excavation area edge notches), smoothing region boundaries and maintaining overall contour continuity. The optimized disturbance area has higher spatial consistency and recognition stability.
[0103] Accurate identification and quantitative analysis
[0104] After completing the region optimization, the system further uses the improved Unet3P model to perform semantic segmentation on the disturbance area, achieving accurate identification of different disturbance types. This model can distinguish various disturbance types such as illegal excavation, mechanical operation and material loading through multi-scale feature fusion and deep residual mechanism.
[0105] In addition, by combining edge detection algorithms (such as Canny edge detection) with contour fitting methods, the system can extract the boundary information of the disturbance region, and use integral calculation methods to obtain the area, perimeter, and volume of the disturbance region. The parameters can be used for risk assessment and construction supervision decision-making by the power transmission line operation and maintenance department, providing quantifiable basis for subsequent governance.
[0106] To verify the performance advantages of the improved Unet3P-based power transmission line construction disturbance intelligent identification system described in this embodiment, the performances of various construction disturbance identification methods in terms of identification accuracy, processing speed, environmental adaptability, and automation degree are compared and analyzed. The results are shown in Table 1 as follows:
[0107] Method type Recognition accuracy Processing speed Environmental adaptability Automation degree Manual inspection 60-70% Low (hours to days) Limited by accessibility Low Monocular visual recognition 75-85% Medium (minute level) Greatly affected by light Medium Traditional binocular vision 85-90% Medium-high (second level) Greatly affected by texture Medium-high Invention method 89-95% High (millisecond level) Strong, adaptable to various environments High
[0108] Table 1 Comparison of performance of different construction disturbance identification methods
[0109] Analysis and alarm module
[0110] The analysis and alarm module of the present application is located in the post-processing and decision-making layer of the system, forming a closed loop from identification results to operation and maintenance decision-making. The analysis module is the data basis of the alarm module, and the alarm module is the value landing of the analysis module, and the two form a dependent relationship.
[0111] The overall workflow of the system is as follows:
[0112] The prelude module (binocular vision data acquisition → image preprocessing and stereoscopic correction → depth calculation and feature extraction → construction disturbance region identification) outputs the segmentation results of the disturbance region;
[0113] The analysis module performs type discrimination, quantitative calculation, and risk assessment on the identification results;
[0114] The alarm module performs risk warning and closed-loop disposal according to the analysis results, thereby achieving the overall goal of "reducing line faults and reducing operation and maintenance costs".
[0115] The analysis and alarm module includes an analysis module and an alarm module.
[0116] Analysis module: quantitative interpretation and risk assessment of disturbance data
[0117] The core task of the analysis module is to convert the semantic segmentation results (including disturbance region contour, depth value, etc.) output by the construction disturbance identification module into quantifiable information and risk levels for operation and decision-making. This module relies on the output results of the depth calculation and feature extraction module, and mainly includes the following three links:
[0118] Accurate disturbance type discrimination
[0119] The improved Unet3P model has the ability of semantic segmentation, which can accurately classify the disturbance types through full-scale feature aggregation and local attention mechanism.
[0120] The analysis logic is as follows:
[0121] The model classifies the optimized disturbance area at the pixel level and outputs the type label of each disturbance area.
[0122] For example:
[0123] When the disturbance area depth value is greater than 0.5m and the boundary is irregular, it is determined as a violation of excavation area;
[0124] When the area contains metal texture features and presents a continuous moving track, it is determined as a mechanical operation area;
[0125] When the area presents a stacking pattern and elevation anomaly, it is determined as a stacked material area.
[0126] This process solves the type confusion problem between natural changes and construction disturbances, and improves the value of operational decision-making of the recognition results.
[0127] Disturbance parameter quantitative calculation
[0128] Based on the depth map and contour extraction algorithm (using Canny edge detection and contour fitting method), the shape and geometric parameters of the disturbance area are accurately calculated.
[0129] The calculation content includes:
[0130] Area and perimeter:
[0131] Through the integral method, the pixels inside the disturbance area contour are counted, and the area value in square meters (㎡) and the perimeter value in meters (m) are output, which are used to evaluate the disturbance scale.
[0132] Depth / height:
[0133] Based on the parallax-depth conversion formula of binocular vision
[0134]
[0135] Where Z is the depth value (m), f is the camera focal length (pixel), B is the baseline length (m), and d is the parallax value (pixel).
[0136] The system calculates the maximum and average depth of the excavation area and the maximum and average height of the stacked area, forming a quantifiable disturbance intensity index.
[0137] Disturbance risk level evaluation
[0138] The analysis module classifies the disturbance into three risk levels: low, medium, and high, based on the disturbance type, geometric parameters, and line safety threshold:
[0139] Risk level Example of determination condition High risk Excavation depth ≥ 1m or area ≥ 50㎡; large machinery (crane / excavator) enters the line corridor within 50m range; pile height ≥ 2m and close to the tower Medium risk Excavation depth 0.5-1m or area 10-50㎡; small machinery enters the line corridor within 50-100m range; pile height 1-2m Low risk Excavation depth < 0.5m or area < 10㎡; temporary load height < 1m and away from the tower.
[0140] Through this hierarchical system, the potential threat of the disturbance to the safety of the transmission line can be quantified and visualized, providing a decision basis for subsequent alarm and disposal.
[0141] Alarm module: real-time delivery and disposal of risk information
[0142] The core goal of the alarm module is to convert the risk level and quantitative results output by the analysis module into perceptible and executable operation and maintenance alarms, and to achieve a closed-loop management of "alarm-disposal-feedback". Its functions rely on the interface of real-time transmission technology and operation and maintenance management system, and include the following:
[0143] 1. Alarm triggering conditions
[0144] The system only triggers alarms for medium and high-risk disturbances, and records low-risk disturbances in the database but does not push them, avoiding interference from invalid information.
[0145] If the disturbance occurs in a critical line segment (such as within 30m of a tower or a cross-over section), even if it is assessed as low-risk, a forced alarm will be triggered.
[0146] 2. Real-time alarm guarantee
[0147] Model level:
[0148] Through deep separable convolution and model pruning techniques, the inference time of a single 1024x1024 pixel image is ≤0.05 seconds, ensuring real-time output of recognition and analysis results.
[0149] Data transmission level:
[0150] The UAV uses a visual navigation system, and the high-definition images and depth data collected can be transmitted back to the cloud in real time through edge computing nodes, ensuring no delay in alarm transmission.
[0151] 3. Alarm information content and format
[0152] The alarm information pushed by the system contains five core elements, ensuring that operation and maintenance personnel can quickly locate and make decisions:
[0153] Category Content description Basic information Disturbance occurrence time, latitude and longitude coordinates (based on UAV GPS positioning), belonging to the power transmission line section Analysis result Disturbance type (illegal excavation / mechanical operation / pile material), risk level (medium / high), quantitative parameters (area, depth / height, distance from tower Supporting data Original binocular image, depth map and semantic segmentation mask image (highlighted disturbance area)
[0154] Through real-time pushing, visual display, and closed-loop recording, the alarm module realizes immediate linkage from recognition results to operation and maintenance actions, forming a complete automated monitoring and decision-making system.
[0155] Core technical principles and improvement motivation
[0156] To facilitate understanding of the technical implementation principles and improvement motivation of the present application, the evolution process and improvement direction of the deep segmentation network Unet series model adopted will now be described.
[0157] Evolution and limitations of Unet series model
[0158] The Unet model was originally proposed by Ronneberger et al. in 2015, and its network structure is a typical encoder-decoder architecture with a symmetrical U-shaped structure overall. The model realizes the fusion of low-level features and high-level features through skip connection, and can utilize high-level semantic features while preserving details such as target edges and textures, achieving remarkable results in the field of medical image segmentation.
[0159] However, the standard Unet model still has several limitations: first, the skip connection directly splices features of different semantic levels, which is prone to semantic gap problems; second, the fixed-depth feature fusion method lacks flexibility when dealing with different scale targets, and has limited adaptability to multi-scale targets in complex scenes.
[0160] Improvement direction of Unet++ and Unet3P model
[0161] To overcome the above shortcomings, the Unet++ and Unet3P (U-Net3+) models were subsequently proposed:
[0162] Unet++ model
[0163] Unet++ enhances the fusion of multi-scale features by introducing dense skip paths and nested structure design, thereby significantly improving segmentation accuracy. However, the model structure is complex, with large computational load and memory occupation, making it unsuitable for deployment on resource-constrained edge devices or unmanned aerial vehicle platforms.
[0164] Unet3P model
[0165] The Unet3P model is further improved based on Unet and Unet++, introducing a full-scale skip connection mechanism and progressive semantic learning strategy, and constructing a dense feature interaction network. This structure can integrate features from each layer of the encoder, enabling cross-scale semantic fusion and effectively addressing recognition challenges such as target boundary blurring and scale diversity, making it particularly suitable for segmentation tasks in complex terrain and multiple disturbance areas in power line construction disturbance scenarios.
[0166] Main structure and functions of Unet3P model
[0167] The Unet3P model adopts an encoder-decoder architecture overall, with the following functions for each module:
[0168] Encoder module
[0169] The input image is down-sampled layer by layer through the stacked convolutional and pooling layers to extract multi-scale features from low-order to high-order. The low-order features contain detailed information such as target edges and textures, and the high-order features reflect the semantic information and overall contour of the target.
[0170] Decoder module
[0171] The spatial resolution of the feature map is gradually restored by upsampling operation, and the feature information from the corresponding layer of the encoder is fused by the skip connection mechanism to compensate for the details that may be lost in the upsampling process.
[0172] Multi-scale fusion mechanism
[0173] Unlike the traditional Unet which only fuses the features of adjacent layers, the Unet3P directly fuses the features of each layer of the encoder with each layer of the decoder through the cross-scale feature fusion module, realizes the full interaction of multi-scale information, and thus improves the accuracy and robustness of semantic segmentation.
[0174] Adaptability requirements in special environments
[0175] In this embodiment, considering the characteristics of the power line construction scene, such as wide geographical area, complex terrain, harsh climate conditions, and multi-ethnic population, the construction disturbance intelligent identification technology needs to have strong environmental adaptability and robustness. Traditional monocular vision methods lack depth information, making it difficult to realize quantitative analysis of the three-dimensional space of the construction scene (such as the calculation of excavation depth and load height); at the same time, complex lighting conditions (such as strong light, shadow area) and shielding factors (such as vegetation, terrain undulations) easily lead to fluctuations in the identification results, resulting in false positives or false negatives, thereby reducing the stability and reliability of the system identification.
[0176] To further clarify the advantages of the improved Unet3P model used in the present application in complex environments, the core performance indicators of the Unet series models are compared and analyzed as shown in Table 2 below:
[0177] Feature UNet UNet++ Unet3P(UNet3+) Skip connection method Direct connection Densely nested Full-scale aggregation Feature fusion Same level fusion Multi-level fusion Cross-scale fusion Parameter quantity (M) 7.8 9.1 12.4 Inference time (ms) 45 68 58 Dice coefficient 0.815 0.834 0.847 Applicable scenario Small-scale data set, high real-time requirement Medium-sized data, pursuit of accuracy Large-scale data, complex boundary segmentation task
[0178] Table 2: Performance comparison of Unet series models
[0179] As can be seen from Table 2, the Unet3P (UNet3+) model is superior to UNet and UNet++ in terms of cross-scale feature fusion, semantic consistency, and complex boundary segmentation capability, and has stronger environmental adaptability while balancing accuracy and inference speed.
[0180] Based on the above analysis, this embodiment selects Unet3P as the core infrastructure and makes the following three improvements and optimizations for the special needs of the power transmission line construction environment in Xinjiang region:
[0181] Lightweight design: By introducing deep separable convolution and model pruning techniques, the model parameter and computation amount are significantly reduced, making it adaptable to unmanned aerial vehicle on-board edge computing platforms and meeting the performance requirements of real-time detection;
[0182] Attention mechanism enhancement: Local attention mechanism modules are introduced into the feature extraction network to enhance the response capability to key areas, improve the feature discrimination ability in complex backgrounds (such as terrain undulations and vegetation obstruction), and reduce misidentification;
[0183] Multi-modal data fusion: Combining binocular vision depth information and two-dimensional feature maps, joint analysis is realized at the three-dimensional spatial level, thereby improving the reliability and quantitative analysis accuracy of disturbance area identification.
[0184] Through the above improvements, the system can still maintain high recognition accuracy and stability in complex lighting, shielding and terrain environments, meeting the application requirements of power transmission line construction disturbance identification.
[0185] Core innovation of improved Unet3P (UNet3+) model
[0186] To adapt to the problems of complex environment, diverse target scale, and fuzzy boundary in the power transmission line construction disturbance identification task in special areas (such as Xinjiang region), this embodiment has made multiple improvements and optimizations based on the original Unet3P (UNet3+) model. The improved model has achieved performance improvement in feature aggregation, attention mechanism, semantic consistency, and noise suppression, making it more suitable for construction disturbance identification in complex scenarios.
[0187] Enhanced full-scale feature aggregation mechanism
[0188] Although the original Unet3P model introduces a full-scale skip connection mechanism to realize the fusion of multi-level features, there is still room for optimization in the feature aggregation method. The traditional simple feature splicing method is prone to introduce redundant information and noise, which adversely affects the segmentation accuracy.
[0189] Therefore, this embodiment designs an efficient feature selection module that calculates the attention weights of features from different sources to achieve selective feature fusion, reducing the interference of invalid features and enhancing the model's response capability to key areas.
[0190] In the improved model, each stage of the decoder aggregates feature information from five different sources, and the specific processing method is as follows:
[0191] Encoder small-scale feature: The feature map size is reduced by the max-pooling operation, and then the channel number is unified by the convolution operation to retain high-order semantic information.
[0192] Encoder same-scale feature: The channel number is adjusted directly by the convolution operation to retain detailed information such as target edges and textures.
[0193] Decoder large-scale feature: The feature map size is expanded by bilinear interpolation upsampling, and then the channel number is unified by the convolution operation to supplement the spatial resolution information.
[0194] Bottleneck layer global context feature: The global context information is extracted by global average pooling to improve the model's semantic understanding ability of the overall scene.
[0195] Previous decoding layer feature: Inherit the decoding result of the previous layer to realize the continuity and consistency of feature transmission.
[0196] After the scale unification of the above-mentioned multi-source features, the model does not simply concatenate each feature, but uses an attention weight matrix to weight and combine different source features. The attention weight is dynamically calculated according to the semantic content of the feature, so that the model can preferentially pay attention to features useful for the disturbance identification task (such as construction machinery outlines and excavation area boundaries), and effectively suppress irrelevant backgrounds (such as sparse vegetation and surface noise).
[0197] The feature aggregation process can be represented in the following mathematical form:
[0198]
[0199] Where w_i represents the attention weight of the i-th source feature, F_i represents the feature map after scale and channel unification, and bias is the bias term.
[0200] Through the improvement of the above full-scale feature aggregation mechanism, the embodiment can realize higher feature utilization rate and information fusion degree, thereby significantly improving the segmentation accuracy and robustness of the model in complex environments.
[0201] Lightweight design for edge deployment
[0202] In the embodiment, considering that the power transmission lines in Xinjiang are mostly located in remote areas, the on-site computing resources are limited, such as the unmanned aerial vehicle on-board processing unit and the edge monitoring station along the line, and the computing power, storage and energy consumption of the equipment are limited. In order to ensure the real-time operation ability of the model in such edge environment, the improved Unet3P model is optimized in a lightweight manner under the premise of maintaining the recognition accuracy, thereby significantly reducing the model parameter quantity and computational complexity.
[0203] The specific optimization measures are as follows:
[0204] Depthwise Separable Convolution replaces standard convolution
[0205] Traditional convolution operation needs to perform convolution calculation in both spatial and channel directions at the same time, which is computationally intensive and parameter redundant. This embodiment decomposes the standard convolution into two independent steps:
[0206] (1) Depthwise Convolution: independent convolution operation is performed on each channel to extract local spatial features;
[0207] (2) Pointwise Convolution: 1x1 convolution is used to fuse different channel features to realize cross-channel information interaction.
[0208] Through the above replacement, the model calculation amount can be greatly reduced without significantly reducing the recognition accuracy. Experimental results show that this optimization scheme can reduce about 70% of the parameter amount, and the inference performance only decreases by less than 2%, achieving a good balance between accuracy and efficiency.
[0209] Introduction of model pruning technology
[0210] To further improve the running efficiency of the model, this embodiment performs pruning optimization on the network structure after training. Specifically, by evaluating the gradient contribution of each neuron or connection, redundant nodes and connections that have little effect on the model output are identified and removed. After pruning, the model size is reduced by about 40%, and the inference speed is improved by about 60%, significantly reducing the calculation resource occupation while maintaining the accuracy.
[0211] Through the above lightweight design, the improved Unet3P model of this embodiment can stably run on edge devices with limited computing resources, realize real-time construction disturbance detection and identification, and meet the on-site deployment needs of unmanned aerial vehicles and monitoring terminals.
[0212] Local attention mechanism enhancement
[0213] In view of the problems of complex natural environment, obvious terrain undulation, frequent sand weather, and serious vegetation obstruction in the construction scene of power transmission lines in Xinjiang region, this embodiment further introduces a local attention mechanism in the improved Unet3P model to enhance the focusing ability of the model on the construction disturbance area and reduce the interference of complex background. This mechanism combines "channel attention" and "spatial attention" two feature weighting strategies, so that the model can realize adaptive feature enhancement in both channel and spatial dimensions, thereby improving the recognition accuracy of the disturbance area.
[0214] Channel attention mechanism
[0215] The channel attention module is used to learn the importance relationships between different feature channels, thereby highlighting significant channel features related to construction disturbances. Specifically, global average pooling is used to obtain statistical information of the input feature map in each channel dimension, and two 1×1 convolutional layers are used to perform dimensionality reduction and expansion operations to learn the dependencies between channels. Finally, a channel weight vector is generated using the sigmoid function to weight the original features. This mechanism can highlight features directly related to construction disturbances, such as mechanical metal textures and soil color changes in the excavated area, thereby improving the model's recognition ability in the channel dimension.
[0216] Spatial attention mechanism
[0217] The spatial attention module is used to determine the spatial importance of key regions. Specifically, it compresses multi-channel feature maps into single-channel representations using 1×1 convolutions and generates spatial weight maps using the sigmoid function to focus on disturbed areas and suppress background areas. This module guides the model to focus more intently on construction disturbance areas (such as excavation areas, loading areas, and areas of mechanical activity), reducing interference caused by vegetation, shadows, or terrain undulations.
[0218] Joint attention mechanism fusion
[0219] During model execution, the channel attention weights and spatial attention weights are multiplied and applied together to the input features, achieving a dual enhancement of key information. Its output can be represented as:
[0220] F out =F in ×W c ×W s
[0221] Among them, F in For the input feature map, W c W represents the channel attention weights. s This represents the spatial attention weights. Through this dual-weighting strategy, the model can adaptively focus on key perturbation regions and effectively suppress irrelevant background noise.
[0222] By introducing the aforementioned local attention mechanism, this embodiment enables the improved Unet3P model to have stronger robustness and stability in construction disturbance identification tasks with "small target area and complex background", significantly improving identification accuracy and boundary segmentation quality.
[0223] Category guidance module prevents false alarms
[0224] In the construction disturbance identification task, due to the influence of natural terrain changes, shadows, vegetation distribution, and surface texture and other factors, non-disturbance areas (such as normal terrain, non-construction vegetation belt) are often misidentified as disturbance targets, resulting in a high false alarm rate. In order to reduce such misidentification, the embodiment introduces a classification guided module (Classification-Guided Module) in the improved Unet3P model, through a "classification-segmentation" joint training mechanism, realizing multi-layer supervision guidance from image-level semantic judgment to pixel-level segmentation, thereby effectively suppressing the over-segmentation of non-disturbance areas.
[0225] Classification branch design
[0226] The embodiment adds an independent classification branch in the model structure, which is used for global disturbance judgment of the input image. The branch takes the deepest layer features of the encoder (with global semantic information) as input, and goes through the following steps in turn:
[0227] Global Average Pooling: extract global semantic features and compress spatial dimensions;
[0228] Dropout layer: prevent network overfitting and improve generalization ability;
[0229] Fully connected layer + Sigmoid activation function: output image-level binary classification results to determine whether there is a disturbance target in the input scene ("with disturbance" or "without disturbance").
[0230] Joint training mechanism
[0231] During model training, the embodiment adopts a joint optimization strategy to minimize the segmentation loss function (such as Dice loss) and the classification loss function (such as binary cross-entropy loss). This joint training method enables the model to learn pixel-level boundary features while being globally supervised by image-level labels, forming a "coarse-to-fine" disturbance identification logic.
[0232] In the inference stage, the classification results output by the model will be fused with the side outputs of each stage of the segmentation decoder:
[0233] When the classification result is "no disturbance", the segmentation output is suppressed to avoid misjudging non-disturbance areas as disturbance targets;
[0234] When the classification result is "with disturbance", the segmentation output is retained and strengthened to ensure complete identification of the disturbance area.
[0235] The mechanism realizes dynamic guidance of image-level supervision to pixel-level segmentation, effectively reduces the false positive rate in complex scenes, and is especially suitable for the typical environmental characteristics of "large area non-disturbance background and small area disturbance target" in the Xinjiang power transmission line scene. Through the addition of the classification guidance module, the system has been significantly improved in recognition accuracy and stability.
[0236] To further illustrate the actual value of the innovations in the improved Unet3P (UNet3+) model, the main innovations, the technical means used, the key problems solved, and the expected technical benefits are summarized as shown in Table 3 below:
[0237] Innovation point Technical means Problems solved Expected benefit Full-scale feature aggregation enhancement Attention weight feature selection Redundancy and noise problem in feature fusion Improve target boundary clarity by 10-15% Lightweight design Depth separable convolution + model pruning Limited computing resources on edge devices Model size reduced by 40%, inference speed increased by 60% Local attention mechanism Channel attention + spatial attention Complex background interference Background misjudgment rate reduced by 20-30% Classification guide module Classification-segmentation joint training Over-segmentation of non-disturbance area False positive rate reduced by 15-25%
[0238] Table 3: Main innovations and technical benefits of the improved Unet3P (UNet3+) model
[0239] As can be seen from Table 3 above, the present embodiment makes systematic improvements from four aspects of feature aggregation, computation optimization, attention guidance, and output false alarm prevention for the complex application environment of power transmission line construction disturbance identification. Through the coordinated application of these innovative technologies, the model has achieved significant improvements in recognition accuracy, computational efficiency, environmental adaptability, and output reliability compared to traditional network structures, providing high-performance algorithm support for intelligent identification and accurate evaluation of power transmission line construction disturbances.
[0240] Training process and parameter setting
[0241] To ensure the stable performance of the improved Unet3P-based intelligent identification system for power transmission line construction disturbances in complex environments, the present embodiment designs a systematic model training process and parameter configuration strategy. The training process includes data preparation and preprocessing, model training parameter setting, and optimization strategy selection.
[0242] Data preparation and preprocessing strategy
[0243] High-quality data training is the foundation of ensuring model identification performance. In view of the complexity and diversity of power transmission line image data in Xinjiang region, the present embodiment designs a strict data division and preprocessing strategy to improve the robustness and generalization ability of the model.
[0244] Training set and test set division method
[0245] To ensure the balance of data distribution and avoid performance degradation due to sample bias, the present embodiment uses stratified sampling strategy (Stratified Sampling) for data division, which includes the following steps:
[0246] (1) The original data is divided into training set, validation set and test set according to the ratio of 8:1:1. Among them:
[0247] · The training set is used for learning and optimization of model parameters;
[0248] · The validation set is used for hyperparameter tuning and network structure selection;
[0249] · The test set is used for independent evaluation of the final model performance.
[0250] (2) During the division process, ensure that each data set contains typical terrain features of Xinjiang region (including Gobi, desert, high mountain, etc.) and different light conditions (such as morning, noon, dusk), and covers complex weather scenes such as sand, rain and snow, so as to improve the adaptability of the model under multiple environmental conditions.
[0251] (3) In the specific implementation process, the train_test_split function in the Python scikit-learn library is used for two-layered division operations. The first division is used to separate the training set and the temporary test set (including the validation set and the test set), and the second division is used to further divide the temporary test set into the validation set and the test set to ensure distribution consistency. The schematic process is as follows:
[0252]
[0253] Among them, terrain_labels and terrain_temp_labels represent the terrain type labels of the original data and the temporary test set respectively. Through this layered strategy, the distribution of terrain, light and climate between different data sets can be ensured, thereby improving the representativeness and fairness of the model training process.
[0254] Data augmentation technology
[0255] In view of the problem of limited amount of labeled data available in the Xinjiang power transmission line construction monitoring scene, the embodiment introduces on-the-fly data augmentation technology in the model training process, that is, dynamically generating augmented samples during the training phase without additional storage space, thereby effectively improving the diversity of training data and enhancing the robustness and generalization ability of the model.
[0256] In each training cycle, the system will randomly apply one or more enhancement operations to each training image. Specifically, the following three types of enhancement strategies are included:
[0257] Geometric transformation
[0258] Random rotation: randomly rotate within the range of−10° to 10°, which is used to simulate the slight change of the unmanned aerial vehicle shooting angle;
[0259] Horizontal flip (probability 0.5) and vertical flip (probability 0.3): These are used to simulate different flight directions and different viewpoints of the drone, respectively.
[0260] Random cropping: The cropping ratio is 0.8 to 1.0, used to simulate local magnification and close-up scenes, and to enhance the scale adaptability of the model.
[0261] Photometric transformation
[0262] Brightness adjustment: Factor range 0.5 to 1.5, used to simulate image differences between day and night and under different lighting conditions;
[0263] Contrast and saturation adjustment: Factor range 0.5 to 1.5, used to simulate visual differences caused by sunny and cloudy weather, light reflection, etc.
[0264] Add Gaussian noise: mean 0, variance 0.01, to simulate sensor noise and transmission interference, and improve the model's noise resistance.
[0265] Special Environment Simulation
[0266] Dust storm simulation: By superimposing brown tones and particle noise, the dust storm environment in Xinjiang in spring is simulated to improve the model's recognition ability in low-contrast scenes;
[0267] Rain and snow weather simulation: Winter rain and snow weather conditions are simulated by adding white stripes and reducing contrast, enhancing the model's all-weather monitoring capabilities.
[0268] To visually demonstrate the application scenarios and technical effects of various data augmentation methods, Table 3 lists typical augmentation strategies, parameter ranges, and their applicability in special terrain scenarios.
[0269] Enhancement type Specific method Parameter range Simulation environment Implementation effect Geometric transformation Random rotation ±10° Different UAV shooting angles Improve rotation invariance Geometric transformation Horizontal / vertical flip Probability 0.3-0.5 Different flight directions Improve symmetry recognition ability Lightness transformation Brightness adjustment Factor 0.5-1.5 Light changes in different time periods Adapt to day and night light differences Lightness transformation Add Gaussian noise Variance 0.01 Sensor noise Improve noise resistance Environment simulation Sand and dust weather simulation Particle density 0.001-0.01 Spring sandstorm Improve the robustness of bad weather Environment simulation Rain and snow weather simulation Stripe density 0.001-0.005 Winter snow weather Improve all-weather monitoring capability
[0270] Table 4: Data augmentation methods and their application effects in special terrain scenarios
[0271] The data augmentation module in this embodiment is implemented based on the imgaug library, which supports batch image parallel processing and multi-threaded acceleration. It can significantly improve data diversity without reducing training speed, thereby ensuring the model's stable performance and generalization performance in complex environments.
[0272] Data normalization processing
[0273] In this embodiment, since the input images can come from different types of UAVs and various imaging sensors, each device has differences in exposure mode, photosensitive characteristics and dynamic range, resulting in inconsistent pixel distribution ranges of the input images. To ensure data consistency and feature comparability during model training, this embodiment introduces normalization and standardization processing strategies in the data loading stage to perform uniform pixel scale conversion and channel-level correction on all input samples.
[0274] The specific steps are as follows:
[0275] Pixel value normalization
[0276] First, the pixel values of the input image are linearly scaled from the original [0, 255] range to the [0, 1] range to reduce the training instability caused by numerical differences and ensure the consistency of input amplitudes between different samples.
[0277] Channel-level mean and variance standardization
[0278] Then, the channel-level mean and variance standardization method is used to eliminate the illumination bias and color difference under different sensor imaging conditions. This method calculates the pixel statistics of each channel in the Xinjiang transmission line data set to standardize the input data, so that the model can converge faster and improve the recognition accuracy during the training process.
[0279] The normalization processing of this embodiment is implemented based on the torchvision.transforms module of the PyTorch framework, and the specific process is as follows:
[0280]
[0281] Among them, the mean and std parameters represent the channel mean and standard deviation based on the statistics of the Xinjiang transmission line special data set, which is different from the standardization parameters of the general ImageNet data set. By using regional statistical features for normalization processing, it can better fit the local data features, thereby improving the convergence speed and final recognition accuracy of the model.
[0282] Model architecture configuration details
[0283] In this embodiment, the improved Unet3P (UNet3+) model is based on the original network architecture, combined with the characteristics of the Xinjiang transmission line construction disturbance recognition scene, and optimized the encoder, decoder, skip connection and classification guide module to balance the feature expression ability, environmental adaptability and computational efficiency.
[0284] Figure 3An improved Unet3P model according to an embodiment of the present disclosure is shown.
[0285] Encoder and decoder settings
[0286] (1) Encoder design
[0287] In this embodiment, ResNet-34 is selected as the backbone network. The gradient vanishing problem of deep networks is effectively alleviated through the residual connection structure, and the encoder has stronger feature extraction capability than the traditional VGG type encoder. The encoder is divided into five stages. Each stage is down-sampled by convolution with a step of 2 (instead of traditional maximum pooling) to preserve more spatial detail information. The configuration of each stage is shown in Table 5:
[0288] Stage Number of convolution layers Output feature map size Channel number Downsampling method 1 3 256×256 64 None 2 3 128×128 128 Step 2 convolution 3 4 64×64 256 Step 2 convolution 4 6 32×32 512 Step 2 convolution 5 3 16×16 1024 Step 2 convolution
[0289] (2) Decoder design
[0290] The decoder uses a full-scale feature aggregation mechanism. Each decoding layer receives feature input from all encoder layers and fuses them after uniform scale and channel number processing.
[0291] For example, when the output size of the third layer of the decoder is 64x64, the input features include:
[0292] • From encoder stage 1 (256x256, down-sampled by maximum pooling);
[0293] • From encoder stage 2 (128x128, down-sampled by maximum pooling);
[0294] • From encoder stage 3 (64x64, direct connection);
[0295] • From encoder stage 4 (32x32, up-sampled by bilinear interpolation);
[0296] • From encoder stage 5 (16x16, up-sampled by bilinear interpolation).
[0297] After the above features are aligned in scale, they are fused using an attention weight matrix to output a rich feature representation containing multi-level information, effectively improving the recognition accuracy of complex boundary regions.
[0298] Implementation of skip connection
[0299] Different from the direct feature concatenation of the original Unet model, the embodiment introduces an attention gate mechanism in the skip connection structure to dynamically adjust the weights of features from each layer of the encoder, so as to suppress background interference and strengthen the target feature region.
[0300] Specifically:
[0301] Input: features from the encoder as “Key” and “Value”, and features from the upper layer of the decoder as “Query”;
[0302] Calculation: generate an attention weight matrix through matrix operation, and the specific expression is as follows:
[0303]
[0304] Wherein, d_k is the dimension of the key, and sqrt(d_k) is a scaling factor, which is used to avoid the problem of unstable gradient caused by too large weight value. The attention gate mechanism can adaptively allocate feature weights according to the context information, thereby enhancing the semantic discrimination ability of the model in complex scenes.
[0305] Classification guidance module
[0306] In order to suppress excessive segmentation of non-disturbance regions, the embodiment introduces a classification guidance module (CGM) in the network. The structure configuration is as follows:
[0307] Input: the deepest layer feature of the encoder (stage 5, size 16x16x1024);
[0308] Structure: sequentially pass through global average pooling (compress 16x16x1024 to 1x1x1024), Dropout (proportion 0.5, used to prevent overfitting), fully connected layer (1024→256→1) and Sigmoid activation function;
[0309] Output: output scalar value range [0, 1], representing the probability of the existence of construction disturbance in the image;
[0310] Training and inference: in the training stage, a binary cross-entropy loss function is used to optimize the classification branch; in the inference stage, the classification result is multiplied by the decoding output of the segmentation mask, and when the classification result is “no disturbance”, the segmentation output is suppressed, thereby reducing false positives.
[0311] Overall parameter configuration of the model
[0312] For the purpose of understanding the network scale and design points of the present embodiment, the core parameter configurations of each module are shown in Table 6 as follows:
[0313] Module name Core component Parameter quantity Output scale Special design Encoder ResNet-34 main body 21.3M 256×256→8×8 Residual connection, step convolution downsampling Decoder Full-scale aggregation 4.2M 8×8→256×256 Multi-source fusion, attention gate Skip connection Attention gate 0.8M Self-corresponding encoder layer Query-key-value attention mechanism Classification guide module Global pooling + FC 0.1M 1×1 Dropout (0.5) to prevent overfitting Deep supervision head 3x3 convolution + up-sampling 0.4M 0.4M Same as input image Side output loss supervision
[0314] Table 6: Parameter configurations of each module of the improved Unet3+ model
[0315] Through the above configurations, the present embodiment significantly improves the computing efficiency and environmental adaptability while maintaining high segmentation accuracy, enabling efficient and reliable real-time construction disturbance recognition on edge computing devices.
[0316] Training strategy and hyperparameter setting
[0317] To ensure the recognition performance and generalization ability of the improved Unet3P (UNet3+) model in the complex scenarios of Xinjiang power transmission lines, the present embodiment adopts a multi-stage training strategy and fine-grained hyperparameter configuration to achieve efficient and stable model convergence. Specifically, the following contents are included:
[0318] Learning rate scheduling strategy
[0319] The present embodiment adopts a dynamic learning rate adjustment strategy combining Cosine Annealing and Warm Restart mechanism to balance the convergence speed and parameter optimization ability of the model.
[0320] Initial learning rate setting:
[0321] A lower initial learning rate of 0.0001 is set for the pre-trained encoder (ResNet-34 backbone) to train in a fine-tuning manner, preventing the destruction of original features;
[0322] An initial learning rate of 0.001 is set for the randomly initialized decoder and classification guide module to accelerate the parameter convergence.
[0323] Learning rate scheduler configuration:
[0324] The CosineAnnealingWarmRestarts method of the PyTorch framework is adopted, with the following parameter settings:
[0325] Initial restart period T_0=50 (i.e., the first restart period is 50 epochs);
[0326] Period multiplication factor T_mult=1 (maintain the period length unchanged);
[0327] Minimum learning rate eta_min=1e-6 to avoid gradient vanishing.
[0328] Warm restart mechanism:
[0329] The learning rate is restarted every 50 epochs, and the maximum learning rate is decayed to 0.8 times the previous value after each restart. This strategy helps the model escape local optima, explore better parameter space, and improve overall convergence quality.
[0330] The code example is as follows:
[0331]
[0332] Batch size and number of iterations
[0333] Batch size settings:
[0334] Based on a GPU device with 16GB of video memory, set the batch size to 8. Although a smaller batch size increases training fluctuations, it can improve model generalization and avoid video memory overflow.
[0335] Training phase division:
[0336] The total training period is 200 epochs, divided into three stages:
[0337] Warm-up stage (first 50 epochs): the learning rate is linearly increased from 0.0001 to 0.001 to help the model quickly adapt to data distribution;
[0338] Fine-tuning stage (middle 100 epochs): enable cosine annealing scheduling to gradually reduce the learning rate to refine parameter optimization;
[0339] Fine-tuning stage (last 50 epochs): fix the learning rate to 1e-6 for stable convergence and performance fine-tuning.
[0340] Regularization techniques
[0341] To prevent the model from overfitting in the case of limited data, this embodiment uses multiple regularization methods, as follows:
[0342] Weight decay (L2 regularization): the coefficient is set to 0.0001 to suppress large parameters and stabilize model training;
[0343] Dropout technique: use Dropout (ratio 0.5) in the classification guide module to randomly discard 50% of the neurons to reduce feature dependency;
[0344] Early Stopping: Automatically stop training and restore the best model parameters when the validation set Dice coefficient does not improve for 15 consecutive epochs.
[0345] Label Smoothing: In the classification task, set the smoothing coefficient to 0.1, adjust the true label from "0 / 1" to "0.1 / 0.9", to reduce the overfitting of the model to the label and improve the generalization ability.
[0346] Hyperparameter optimization results
[0347] In this embodiment, various methods such as Bayesian optimization, grid search and random search are used to optimize key hyperparameters. The final hyperparameter configuration and its optimization range are shown in Table 7:
[0348] Hyperparameters Search range Final selection Tuning method Influence analysis Initial learning rate [1e-4, 1e-2] 1e-3 Bayesian optimization Too high causes oscillation, too low converges slowly Batch size [4, 32] 8 Grid search Trade-off between memory efficiency and training stability Weight decay [0, 1e-3] 1e-4 Random search Effectively control overfitting without affecting convergence Dropout rate [0, 0.7] 0.5 Manual tuning Effectively reduce overfitting of classification module Cycle number [100, 500] 200 Early stopping mechanism Avoid overfitting and save training time
[0349] Table 7: Hyperparameter optimization range and final selection
[0350] Loss function and optimizer selection
[0351] In this embodiment, to solve the class imbalance and boundary ambiguity problems in the construction disturbance identification task, while ensuring the robustness of the model in complex topographic environments, a hybrid loss function with multiple loss fusion is designed, and a high-stability optimizer is used to achieve efficient parameter convergence.
[0352] Hybrid loss function design
[0353] In the construction disturbance scenario, the number of foreground pixels (disturbance area) is much less than the number of background pixels, and the disturbance boundary is ambiguous, so it is difficult for a single loss function to balance class balance and boundary accuracy. This embodiment constructs a hybrid loss function composed of three loss functions weighted combination, its expression is as follows:
[0354]
[0355] where, = 0.5, = 0.3, = 0.2, and each weight parameter is determined by grid search. The functions of each loss function are as follows:
[0356] Dice Loss
[0357] Used to alleviate the class imbalance problem, by calculating the overlap between the predicted mask and the true label to measure the prediction accuracy of the foreground area, focusing on optimizing the identification effect of small area disturbance targets.
[0358] The calculation of DiceLoss is as follows:
[0359]
[0360] Focal Loss
[0361] To solve the problem of imbalance between easy and difficult samples. By reducing the weight of easy-to-classify samples (such as large-area background pixels), the model pays more attention to difficult-to-classify areas (such as boundary disturbance targets), improving the overall segmentation robustness.
[0362] The formula of Focal Loss is as follows:
[0363]
[0364] MS-SSIM Loss (Multi-Scale Structural Similarity Loss)
[0365] To enhance the structural consistency of boundary segmentation. This loss measures the local structural similarity between the predicted result and the true label through the multi-scale structural similarity index (MS-SSIM), thereby improving the model's ability to recognize boundary contours.
[0366] The calculation of MS-SSIM Loss (Multi-Scale Structural Similarity Loss) is as follows:
[0367]
[0368] Through the weighted combination of the above three loss functions, the model can simultaneously consider global class balance, local structural accuracy, and small target recognition performance, significantly improving the disturbance segmentation effect in complex terrain conditions in Xinjiang.
[0369] Optimizer selection and parameter configuration
[0370] To further improve the training stability and parameter update efficiency, this embodiment selects AdamW optimizer. Compared with the traditional Adam optimizer, AdamW decouples the weight decay term and the learning rate schedule, effectively avoiding over-regularization interference with parameter updates and improving model generalization performance.
[0371] The optimizer parameter configuration is as follows:
[0372]
[0373] In addition, to ensure the stability of pre-trained features, this embodiment adopts the freeze statistics strategy for the BatchNormalization layer in the encoder, that is, the mean and variance during pre-training on the ImageNet dataset are maintained, and the statistics are not updated in the current task. This strategy can avoid statistical bias caused by the small number of Xinjiang power transmission line samples, thereby improving the stability of the training process and the final performance of the model.
[0374] In summary, by combining the mixed loss function with the AdamW optimizer, the model of the present embodiment achieves higher training stability and better boundary segmentation accuracy in the construction disturbance identification task.
[0375] Training process monitoring and early stopping
[0376] To ensure the stability of the model training process and the reliability of the final performance, the present embodiment introduces a multi-index monitoring mechanism and a patience early stopping strategy (Early Stopping) during the training phase, achieving real-time tracking and dynamic control of the training convergence state.
[0377] Training index monitoring
[0378] During the model training process, the system monitors multiple performance indicators in real time to comprehensively evaluate the training state and model performance. The monitoring indicators include:
[0379] (1) Main indicators:
[0380] Dice Coefficient: used to measure the overlap between the predicted segmentation result and the true label, it is the core evaluation indicator of this task;
[0381] Mean Intersection over Union (mIoU): used to evaluate the overall accuracy of semantic segmentation, measure the consistency of segmentation of different classes.
[0382] (2) Secondary indicators:
[0383] Precision: reflects the false alarm control ability of the model;
[0384] Recall: used to evaluate the detection ability of the model to the disturbance area;
[0385] F1 Score: balances precision and recall, measures the overall identification performance of the model.
[0386] (3) Auxiliary indicators:
[0387] Training loss and validation loss (Loss Trend);
[0388] Learning rate schedule.
[0389] After each training cycle (epoch) is completed, the system calculates the above indicators on the training set and the validation set, and uses the TensorBoard tool for visual display, achieving dynamic tracking of the training process and performance trend analysis.
[0390] Through real-time monitoring, problems such as model overfitting, underfitting, or convergence anomalies can be identified in a timely manner, allowing for the adjustment of hyperparameters or the termination of training.
[0391] Early stopping strategy
[0392] To prevent model overfitting and reduce invalid calculations, the present embodiment uses a patience early stopping strategy. When the performance of the validation set does not improve for a certain number of cycles, the system automatically suspends training and restores the best model weight.
[0393] (1) Patience setting:
[0394] Set the patience value to 15 epochs, and when the Dice coefficient of the validation set does not improve for 15 consecutive epochs, the system triggers the early stopping mechanism.
[0395] (2) Indicator smoothing mechanism:
[0396] To avoid the early triggering of early stopping due to accidental fluctuations in the performance of the validation set, the performance indicators of the validation set are smoothed using an exponential moving average. The calculation formula is as follows:
[0397] {smoothed_val_score = 0.7 \times smoothed_val_score + 0.3 \timescurrent_val_score}
[0398] (3) Minimum training cycle constraint:
[0399] Set the minimum training cycle number to 50 epochs, meaning that even if the early stopping conditions are met, the model must complete at least 50 epochs of training to avoid being misjudged as converged during fluctuations in early-stage performance.
[0400] After training is complete, the system automatically loads the model weight that performs best on the validation set and performs final evaluation on the test set to ensure the objectivity and reliability of the results.
[0401] Through the above training monitoring and early stopping strategy, the present embodiment can effectively balance model training efficiency and performance, prevent overfitting, and improve the stability and generalization ability of the model in complex environments.
[0402] Model validation and evaluation method
[0403] To verify the applicability and stability of the power line construction disturbance intelligent identification system based on the improved Unet3P (UNet3+) model in the complex environment of Xinjiang region, a systematic model verification and evaluation scheme is designed.
[0404] Cross-validation strategy
[0405] Due to the relatively limited size of the Xinjiang power line image dataset (about 2000 original samples), to comprehensively evaluate the stability and generalization ability of the model performance, this embodiment adopts a five-fold cross-validation (5-Fold Cross Validation) strategy for model verification. The specific steps are as follows:
[0406] 80% of the samples in the original data set are used as the training set, and the remaining 20% are used as the independent test set.
[0407] Within the training set, the data is evenly divided into five subsets (Fold), and each time four subsets are selected as the training set and one subset is selected as the validation set.
[0408] Repeat the training five times, each time changing the validation set, and the rest of the training process remains the same, and the test set remains fixed.
[0409] Take the average of the main performance indicators (Dice coefficient, IoU, precision, recall, etc.) of the five experimental results as the final evaluation result to reduce the random bias caused by data division.
[0410] Through the above cross-validation strategy, not only can the accidental error in model evaluation be effectively reduced, but also the sensitivity of the model to specific sample distribution (such as sandstorm weather, low light scene) can be identified, providing a basis for further optimization of the model.
[0411] Performance evaluation indicators
[0412] To comprehensively measure the performance of the model from multiple dimensions, this embodiment uses multiple semantic segmentation and classification related indicators for evaluation. The calculation formula, focus and meaning of each indicator in the Xinjiang power line scene are shown in Table 8:
[0413] Evaluation index Calculation formula Focus direction Best value Significance in Xinjiang scenarios Dice coefficient 2TP / (2TP+FP+FN) Segmentation overlap The closer to 1, the better Measure the completeness of disturbance area identification IoU TP / (TP+FP+FN) Segmentation accuracy The closer to 1, the better Measure the accuracy of area identification Precision TP / (TP+FP) False alarm control The closer to 1, the better Reduce false alarms in non-disturbance areas Recall TP / (TP+FN) False alarm control The closer to 1, the better Reduce false alarms in non-disturbance areas Boundary F1 Boundary matching degree Boundary clarity The closer to 1, the better Evaluate the quality of boundary segmentation AUC-ROC Area under ROC curve Overall performance The closer to 1, the better Comprehensive evaluation of classification performance
[0414] Table 8: Comparison of model performance evaluation indicators
[0415] Among them, TP (true positive) represents the number of pixels correctly identified as disturbance, FP (false positive) represents the number of non-disturbance pixels incorrectly identified as disturbance, and FN (false negative) represents the number of disturbance pixels that cannot be identified.
[0416] The embodiment uses the above multi-dimensional indicators to comprehensively evaluate the model performance from the aspects of recognition integrity, accuracy, boundary definition, and category balance, etc., to ensure that the model has high reliability and practicability under the special geographical and climatic conditions in Xinjiang.
[0417] Specific tuning in special topographic scenes
[0418] In view of the complex natural environment (including gobi, desert, high mountain, etc.) and extreme climate conditions (such as sandstorm, ice and snow, strong light, etc.) of the power transmission line in Xinjiang region, the embodiment further performs scene-specific fine-tuning after the model training is completed, to improve the environmental adaptability and recognition stability of the model under multiple terrain and climate conditions.
[0419] Adaptive learning rate adjustment
[0420] To improve the learning ability of the model for samples under specific topography and climate conditions, the embodiment uses an adaptive learning rate adjustment strategy to dynamically optimize in combination with the validation set feedback information and sample category distribution.
[0421] Scene-aware tuning:
[0422] When the validation set performance (such as Dice coefficient or mIoU) appears to be in a plateau during the training process, the system analyzes the distribution characteristics of the misclassified samples (such as whether they are concentrated in sandstorm weather, high mountain shadow, or reflection area).
[0423] The training weight of such scene samples is specifically improved, for example, the weight of sandstorm scene samples is improved from 1.0 to 1.5, and the learning rate is temporarily improved from 1e-5 to 5e-5, and fine-tuning is performed for about 20 epochs to strengthen the adaptability of the model to specific scenes.
[0424] Seasonal tuning:
[0425] According to the construction characteristics and inspection priorities of different seasons in Xinjiang region, a targeted tuning scheme is designed as shown in the data strategy of Table 9:
[0426] Season Tuning focus Data strategy Spring Line inspection after sand weather Line inspection after sand weather Summer Monitoring during peak construction period Increase mechanical operation and illegal excavation samples Autumn Increase mechanical operation and illegal excavation samples Increase samples of scenes around farmland Winter Ice and snow and low temperature impact Increase ice and snow simulation samples and depth feature weight
[0427] Table 9: Seasonal adjustment
[0428] This strategy periodically fine-tunes the model parameters to make the model maintain high recognition accuracy and stability in different seasons, ensuring all-weather monitoring capability throughout the year.
[0429] Data augmentation strategy optimization
[0430] According to the unique geographical environment and climate characteristics of Xinjiang region, on the basis of the basic data enhancement strategy, a variety of environmental simulation enhancement methods are designed to improve the robustness of the model under extreme scenarios.
[0431] Dust weather enhancement:
[0432] According to the intensity characteristics of spring dust in Xinjiang, the particle density (0.001-0.01) and brown tone intensity (0.1-0.3) are adjusted to generate different concentration of dust scene samples, which can improve the recognition ability of the model under low contrast conditions.
[0433] Day and night change enhancement:
[0434] By adjusting the brightness (0.3-1.8), color temperature (2000K-10000K) and contrast (0.4-1.6), the morning, noon, dusk and night infrared imaging conditions are simulated to improve the robustness of the model to day and night light changes.
[0435] Terrain simulation enhancement:
[0436] For high mountain and undulating terrain samples, random slope adjustment (−15° to +15°) is performed to simulate the changes in shooting angle of the unmanned aerial vehicle on the slope terrain, so as to improve the spatial adaptability of the model to the terrain undulation.
[0437] Extreme cold weather enhancement:
[0438] By superimposing white frost effect (density 0.001-0.005) and reducing saturation, the image degradation under winter extreme cold weather is simulated to improve the segmentation stability of the model under low temperature and low visibility conditions.
[0439] Through the joint application of the above enhancement means, the performance of the model of the embodiment in Xinjiang extreme environment (including sandstorm, low temperature, complex terrain) is significantly improved, and the stable identification of power transmission line construction disturbance can be realized in all-weather and all-terrain.
[0440] Performance verification and test
[0441] In order to verify the effectiveness and engineering practicability of the power transmission line construction disturbance intelligent identification system based on the improved Unet3P (UNet3+) model, systematic experimental test and performance evaluation are carried out.
[0442] Experimental environment
[0443] In order to ensure the stability and reproducibility of the experimental results, the hardware and software environment configuration of the embodiment is as follows:
[0444] Hardware environment:
[0445] Intel Xeon Gold 6248 (24 cores) processor, NVIDIA Tesla V100 (32 GB video memory) GPU, and 128 GB system memory are used to support large-scale data parallel training and deep model optimization.
[0446] Software Environment
[0447] The operating system is Ubuntu 20.04, the deep learning framework is PyTorch 1.12.0, and the Python version is 3.8.10. The main dependent libraries include:
[0448] OpenCV: for image reading and basic processing;
[0449] imgaug: for image data augmentation;
[0450] scikit-learn: for dataset division and statistical analysis;
[0451] TensorBoard: for training process visualization and index monitoring.
[0452] Experimental Dataset
[0453] The experiment uses a self-built remote sensing dataset of Xinjiang power transmission line construction disturbance, which is derived from power transmission line inspection images in different terrain and climate regions in northern Xinjiang. The specific composition is as follows:
[0454] Raw data:
[0455] A total of 2000 high-resolution images (size 512×512) covering 4 typical scenarios:
[0456] Earth excavation area: 420 (including samples of different depths and areas);
[0457] Mechanical operation area: 380 (including excavators, cranes, transport vehicles, and other equipment);
[0458] Material stacking area: 300 (including sand, steel, cable, and other types of materials);
[0459] Non-disturbance area: 900 (including Gobi, desert, mountain, vegetation, and other natural background samples).
[0460] After data augmentation:
[0461] Through the aforementioned scenario-based enhancement and special environment simulation strategy, the dataset size is expanded to 8000, covering different terrains, climates, and lighting conditions in Xinjiang.
[0462] Annotation method:
[0463] Pixel-level semantic annotation is performed using the LabelMe tool to generate semantic segmentation mask files. The annotation accuracy is verified by manual review to reach more than 98%, ensuring the consistency and reliability of the training and testing data.
[0464] Performance evaluation indicators
[0465] The model performance evaluation uses the common indicator system for semantic segmentation tasks, combined with the special application scenarios in Xinjiang, focusing on the recognition accuracy, integrity, and small target detection rate of the model. The main evaluation indicators are as follows:
[0466] Precision:
[0467] The number of correctly identified disturbance pixels accounts for the proportion of all identified disturbance pixels, used to measure the model's false alarm control ability.
[0468] Recall:
[0469] The number of correctly identified disturbance pixels accounts for the proportion of all real disturbance pixels, reflecting the model's capture ability of disturbance targets.
[0470] F1 Score:
[0471] The harmonic mean of precision and recall, used to comprehensively evaluate the overall recognition performance of the model. The formula is as follows:
[0472] 2×(Precision×Recall) / (Precision+Recall)
[0473] Mean Intersection over Union (mIoU):
[0474] The average value of the Intersection over Union (IoU) of each class, is the core evaluation indicator of the semantic segmentation task, used to measure the degree of overlap between the predicted results and the true labels. The formula is as follows:
[0475] (\sum_{c=1}^{C} IoU_c) / C
[0476] Where C is the number of classes.
[0477] Experimental results and analysis
[0478] Comparison of experimental results
[0479] To verify the performance advantages of the model in this embodiment, the improved Unet3P model is compared with the mainstream semantic segmentation network (SegNet, Unet++, and the original Unet3P) in the comparison experiment. The performance results of each model on the test set are shown in Table 10 as follows:
[0480] Model Precision Recall F1 score mIoU SegNet 0.721 0.685 0.702 0.653 Unet++ 0.783 0.756 0.769 0.712 Original Unet3P 0.825 0.801 0.813 0.765 Improved Unet3P 0.88 0.872 0.879 0.836
[0481] Table 10: Performance comparison of different models on the Xinjiang power transmission line construction disturbance data set
[0482] Result analysis
[0483] From the experimental results, the improved Unet3P model described in the application is significantly better than the comparison model in various indicators, and the specific analysis is as follows:
[0484] Overall performance improvement:
[0485] Compared with the original Unet3P, the mIoU is improved by 7.1 percentage points, and the F1 score is improved by 6.6 percentage points, indicating that the application realizes higher precision segmentation and identification of construction disturbance areas by optimizing the encoder (ResNet-34 backbone network), embedding the attention mechanism, and jointly designing the hybrid loss function.
[0486] Small target recognition optimization:
[0487] In complex terrain (such as the junction of Gobi and mountainous areas) test samples, the recall rate of the improved model for small area mechanical operation area and fine material stacking area is more than 10% higher than that of the original Unet3P, benefiting from the synergistic effect of full-scale feature aggregation and local attention mechanism.
[0488] Anti-interference ability enhancement:
[0489] Under the background of sand weather and vegetation shelter, the precision of the improved model is improved by 5.5 percentage points compared with the original Unet3P, and the background misjudgment rate is reduced by about 25%, indicating that the classification guide module and the attention gate mechanism effectively suppress the complex background interference.
[0490] Real-time guarantee:
[0491] After lightweight optimization, the average inference time of the improved model is 42ms / picture (input size 1024x1024), which is improved by 27.6% compared with the original Unet3P of 58ms / picture, meeting the requirements of unmanned aerial vehicle real-time inspection on calculation speed.
[0492] In addition, in the special environment scene of Xinjiang, the improved model shows excellent environmental adaptability:
[0493] Sand weather: mIoU only decreased by 3.2 percentage points (original Unet3P decreased by 8.5 percentage points);
[0494] Winter snow scene: recall rate remains above 85% (original Unet3P decreased to 72%).
[0495] High mountain terrain scene: boundary F1 score reaches 0.84 (original Unet3P is 0.76), significantly improving the boundary segmentation accuracy.
[0496] In summary, the improved Unet3P model has shown significant advantages in recognition accuracy, robustness and real-time performance under complex topography and extreme climate conditions of Xinjiang power transmission line, fully verifying the innovativeness and engineering application value of the present application scheme.
[0497] Advantages and challenges in practical application in special topography remote areas
[0498] The power transmission line construction disturbance intelligent identification method based on the improved Unet3P (UNet3+) model has shown significant advantages in the application in special topography remote areas such as Xinjiang. Xinjiang is a vast region with complex terrain, extreme climate conditions and multi-ethnic population, which puts higher requirements on the environmental adaptability, algorithm robustness and engineering deployment capability of the construction disturbance intelligent identification system.
[0499] Application advantages
[0500] The improved Unet3P model combined with binocular vision technology has the following outstanding advantages in the identification of power transmission line construction disturbance in special topography remote areas in Xinjiang:
[0501] Strong adaptability to complex environment
[0502] (1) Terrain adaptability:
[0503] The unmanned aerial vehicle equipped with a binocular camera can effectively overcome the limitations of complex terrains such as gobi, desert and high mountains in Xinjiang, realize high-altitude large-scale data collection, and avoid the problem of limited accessibility of manual inspection. At the same time, the model can stably process image distortion caused by terrain undulation through terrain simulation and adaptive learning rate adjustment.
[0504] (2) Climate adaptability:
[0505] In view of the extreme climate conditions such as sandstorm, snow, extreme cold and the like in Xinjiang region, the model adopts scene data enhancement and multi-modal data fusion (RGB + depth + infrared) strategy, which still maintains high recognition accuracy (sandstorm weather mIoU ≥ 80%) even in low visibility or bad weather, realizing all-weather monitoring.
[0506] (3) Adaptation to areas without signal:
[0507] The unmanned aerial vehicle system uses visual navigation and inertial measurement attitude technology, and does not need to rely on ground communication signals, so it can fly and collect data autonomously in remote areas without signal, effectively solving the problem of "signal blind area" inspection.
[0508] Balancing recognition accuracy and real-time performance
[0509] (1) High-precision recognition:
[0510] The improved Unet3P model of the application has an mIoU of 83.6% on a Xinjiang test data set, a false positive rate reduced by 20-30%, and can accurately identify small-scale perturbation targets (such as initial excavation marks and small machines), effectively reducing the adverse effects of “missed reports” and “false positives” on operation and maintenance decisions.
[0511] (2) Real-time response:
[0512] After lightweight design, the inference time of a single 1024x1024 image is not more than 0.05 seconds. Combined with the real-time transmission module of the unmanned aerial vehicle, the response time from discovery to alarm of the hidden danger can be shortened from several hours or even several days of traditional manual inspection to minutes, gaining valuable time for emergency disposal.
[0513] Operation and maintenance cost and efficiency optimization
[0514] (1) Reduce labor costs:
[0515] The system supports an “unattended + centralized monitoring” operation mode, significantly reducing the frequency and coverage of manual inspection, especially suitable for the characteristics of “long line span and few operation and maintenance personnel” in Xinjiang, and is expected to reduce labor costs by more than 60%.
[0516] (2) Improve inspection efficiency:
[0517] The system can accurately locate and classify the disturbed area, and the operation and maintenance personnel only need to check the identified high-risk areas, avoiding “full coverage and no focus” blind inspection, and the overall inspection efficiency can be improved by 3-5 times.
[0518] (3) Data value extension:
[0519] The depth and segmentation results collected by the system can be further used to construct a power transmission line digital twin model, providing data support for risk prediction and intelligent scheduling, and realizing the transition from “passive operation and maintenance” to “active prevention”.
[0520] The scope of the claims of the application is not limited to the specific embodiments described above. Various other embodiments that can be made by those of ordinary skill in the art without departing from the essential technical idea of the application described in the claims should also be included in the scope of the claims of the application.
Claims
1. A method for intelligent identification of construction disturbances in transmission lines, characterized in that, Includes the following steps: S1. Binocular vision data acquisition step: A drone equipped with a binocular camera flies along the power transmission line corridor to acquire high-resolution binocular image data. S2, Image preprocessing and stereo correction steps: Denoising, contrast enhancement and stereo correction are performed on the binocular images to eliminate distortion and improve matching accuracy. S3, Depth Calculation and Feature Extraction Steps: The improved Unet3P model is used to extract multi-scale features from stereo-corrected binocular images, establish disparity relationships, and calculate depth maps. S4. Construction disturbance area identification step: Based on the depth map and feature map, threshold segmentation, morphological processing and semantic segmentation methods are used to identify the construction disturbance area, and the disturbance type, contour and spatial boundary are output. S5. Analysis and processing steps: Calculate quantitative parameters such as the area, depth or height, and distance from the tower of the disturbance area based on the identification results, and classify the risk level according to the line safety threshold. S6. Alarm processing steps: When the risk level is medium or high, an alarm is triggered and the disturbance information is pushed to the operation and maintenance system to realize the alarm and handling closed loop. The disturbance information includes disturbance type, latitude and longitude coordinates, area, risk level and corresponding image.
2. The intelligent identification method for construction disturbance of transmission lines according to claim 1, characterized in that, The improved Unet3P model includes: The encoder module adopts the ResNet-34 network structure, which alleviates the gradient vanishing problem through residual connections and replaces the max pooling operation with stride convolution to preserve spatial features. The decoder module is equipped with a full-scale feature aggregation mechanism, which is used to perform weighted fusion of encoder features from different scales at each decoding stage to improve the recognition accuracy of small target perturbations. The skip connection module introduces an attention gating mechanism between the encoder and decoder, dynamically adjusting the fusion weights based on feature relevance to suppress irrelevant background information; The classification guidance module guides pixel-level segmentation output based on image-level classification results during the model training phase, thereby reducing over-segmentation of non-perturbed regions.
3. The intelligent identification method for construction disturbance of transmission lines according to claim 2, characterized in that, The full-scale feature aggregation mechanism fuses features from different scales at each decoding stage, including: The low-level features of the encoder are input after being processed by max pooling downsampling, and the high-level features of the encoder are input after being processed by bilinear interpolation upsampling. The features at each scale are then weighted and combined using an attention weight matrix to reduce feature redundancy and enhance the recognition accuracy of perturbed boundary regions.
4. The intelligent identification method for construction disturbance of transmission lines according to claim 2, characterized in that, The classification guidance module includes: The classification branch is used to take the deepest features of the encoder as input and output image-level classification results through global average pooling, Dropout layers and fully connected layers. The joint training unit is used to simultaneously optimize the segmentation loss and classification loss during the training phase, and multiply the classification result with the segmentation output during the inference phase to suppress over-segmentation of non-perturbed regions, reduce false alarm rate and improve overall recognition accuracy.
5. The intelligent identification method for construction disturbance of transmission lines according to claim 2, characterized in that, The improved Unet3P model employs a lightweight design, including: Depthwise separable convolution is used instead of standard convolution. Spatial features and channel information are extracted by decomposing them into channel-wise convolution and point-wise convolution, thereby reducing the number of model parameters and computational cost while maintaining recognition accuracy. By removing redundant connections and nodes that contribute little to the output through model pruning techniques, the model size is reduced and the inference speed is improved, enabling the model to run efficiently in edge computing environments such as UAV onboard devices.
6. An intelligent identification device for construction disturbance of transmission lines, characterized in that, include: A binocular vision data acquisition module is used to acquire high-resolution binocular image data by flying a drone equipped with a binocular camera along the power transmission line corridor. The image preprocessing and stereo correction module is used to denoise, enhance contrast, and correct stereo for the binocular images in order to eliminate distortion and improve matching accuracy. The depth calculation and feature extraction module is used to perform multi-scale feature extraction on stereo-corrected binocular images using the improved Unet3P model, establish disparity relationships, and calculate depth maps. The construction disturbance area identification module is used to identify construction disturbance areas based on the depth map and feature map, using threshold segmentation, morphological processing and semantic segmentation methods, and output disturbance type, contour and spatial boundary; The analysis and processing module is used to calculate quantitative parameters such as the area, depth or height, and distance from the tower of the disturbance area based on the identification results, and to classify the risk level according to the line safety threshold. The alarm processing module is used to trigger an alarm and push the disturbance information to the operation and maintenance system when the risk level is medium or high, so as to realize the alarm and handling closed loop. The disturbance information includes disturbance type, latitude and longitude coordinates, area, risk level and corresponding image.
7. The intelligent identification device for construction disturbance of transmission lines according to claim 6, characterized in that, The depth calculation and feature extraction module includes an improved Unet3P model. The improved Unet3P model includes: The encoder module adopts the ResNet-34 network structure, alleviates the gradient vanishing problem through residual connections, and replaces the max pooling operation with stride convolution to preserve spatial features; The decoder module is equipped with a full-scale feature aggregation mechanism, which is used to perform weighted fusion of encoder features from different scales at each decoding stage to improve the recognition accuracy of small target perturbations. The skip connection module introduces an attention gating mechanism between the encoder and decoder, dynamically adjusting the fusion weights based on feature relevance to suppress irrelevant background information; The classification guidance module guides pixel-level segmentation output based on image-level classification results during the model training phase, thereby reducing over-segmentation of non-perturbed regions.
8. The intelligent identification device for construction disturbance of transmission lines according to claim 7, characterized in that, The full-scale feature aggregation mechanism fuses features from different scales at each decoding stage, including: The low-level features of the encoder are input after being processed by max pooling downsampling, and the high-level features of the encoder are input after being processed by bilinear interpolation upsampling. The features at each scale are then weighted and combined using an attention weight matrix to reduce feature redundancy and enhance the recognition accuracy of perturbed boundary regions.
9. The intelligent identification device for construction disturbance of transmission lines according to claim 7, characterized in that, The classification guidance module includes: The classification branch is used to take the deepest features of the encoder as input and output image-level classification results through global average pooling, Dropout layers and fully connected layers. The joint training unit is used to simultaneously optimize the segmentation loss and classification loss during the training phase, and multiply the classification result with the segmentation output during the inference phase to suppress over-segmentation of non-perturbed regions, reduce false alarm rate and improve overall recognition accuracy.
10. The intelligent identification device for construction disturbance of transmission lines according to claim 7, characterized in that, The improved Unet3P model employs a lightweight design, including: Depthwise separable convolution is used instead of standard convolution. Spatial features and channel information are extracted by decomposing them into channel-wise convolution and point-wise convolution, thereby reducing the number of model parameters and computational cost while maintaining recognition accuracy. By removing redundant connections and nodes that contribute little to the output through model pruning techniques, the model size is reduced and the inference speed is improved, enabling the model to run efficiently in edge computing environments such as UAV onboard devices.