Power transmission line image video analyzing and monitoring system

The transmission line monitoring system, which combines multimodal data acquisition, edge computing, and cloud analysis, solves the challenges of real-time defect detection and early warning in existing technologies, enables real-time, accurate monitoring and efficient operation and maintenance of transmission lines, and improves system reliability and data security.

CN120750004APending Publication Date: 2025-10-03NANJING NEMIN POWER TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510906827.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing transmission line monitoring systems have significant technical bottlenecks in data collection, processing and analysis, making it difficult to achieve real-time and accurate defect detection and early warning. In particular, the imaging quality is poor in harsh environments, the data fusion accuracy is insufficient, and the algorithm adaptability and system coordination are insufficient, which cannot meet the rapid response needs of the power grid.

Method used

A multimodal acquisition module integrates 4K cameras, infrared thermal imagers, and lidars. Edge computing nodes perform lightweight model detection. Cloud analysis modules perform distributed computing and prediction. Combined with drone collaboration modules and adaptive learning modules, real-time defect detection and early warning are achieved. The system integrates an anti-shake gimbal and an automatic zoom lens, optimizes data transmission and storage, and builds a digital twin for real-time mapping and analysis.

Benefits of technology

It realizes high-frequency collection of multi-dimensional data of transmission lines and real-time detection of various defects. The system establishes a three-level alarm system to reduce defect discovery time, improve operation and maintenance efficiency, reduce operation and maintenance costs, ensure system reliability and data security in complex environments, and support the reliable operation of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120750004A_ABST
    Figure CN120750004A_ABST
Patent Text Reader

Abstract

The invention discloses a power transmission line image video analyzing and monitoring system, and relates to the field of power transmission line monitoring. The multi-modal acquisition module adopts a 4K camera, a thermal infrared imager and the like to acquire data and perform time-space synchronization; the edge processing module operates YOLOv6s to detect defects and upload compressed features; the cloud analysis module analyzes and predicts by means of 3D-CNN and digital twinning technologies; the early warning linkage module gives a multi-stage alarm and dispatches a work order; the storage security module encrypts the storage data. The power transmission line is intelligently monitored, defects are detected in a multi-mode mode, and real-time early warning and quick response are achieved; the cloud edge collaboratively improves the operation and maintenance efficiency and shortens the work order processing period; reliable operation in a complex environment reduces storage and annual average operation and maintenance cost, and ensures power grid safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power transmission line monitoring, and in particular to a power transmission line image and video analysis monitoring system. Background Art

[0002] Transmission lines are critical infrastructure for the power system, and their safe operation directly impacts grid stability. Traditionally, transmission line monitoring relies on manual inspections, requiring operators to carry equipment such as binoculars and infrared thermometers while conducting inspections on foot or by vehicle. Limited by inspection frequency and observation distance, these inspections can make it difficult to detect hidden defects such as minor conductor breaks and premature insulator degradation.

[0003] With the development of intelligent power grids, some systems have introduced fixed cameras and drone inspections, but significant technical bottlenecks remain. Fixed cameras are limited by their mounting locations, making it difficult to cover blind spots on towers (such as above crossarms and behind insulators). Image quality also degrades significantly in harsh environments such as rain, fog, and at night. Traditional drone inspections rely on manual control by pilots, lacking systematic route planning. Individual inspections are time-consuming, and data transmission requires manual image interpretation, resulting in low processing efficiency and inability to meet real-time monitoring needs. Furthermore, multi-source data (visible light, infrared, and lidar) lacks unified spatiotemporal calibration, resulting in insufficient data fusion accuracy and making it difficult to construct a three-dimensional health model of transmission lines.

[0004] The development of artificial intelligence and edge computing technologies has brought new directions to power transmission line monitoring, but existing solutions suffer from shortcomings in algorithm adaptability and system interoperability. For example, deep learning-based defect detection models often utilize centralized cloud-based computing, requiring the upload of large amounts of raw video. This results in high network bandwidth usage and latency, making it difficult to meet the demand for rapid fault response. Edge computing nodes lack the ability to dynamically update models, resulting in reduced detection accuracy when faced with new defects (such as new types of metal corrosion). Furthermore, algorithm upgrades rely on manual on-site deployment, making them difficult to adapt to the complex and changing power transmission environment. Summary of the Invention

[0005] The present invention proposes a power transmission line image and video analysis monitoring system to solve the problems mentioned in the above-mentioned prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: a transmission line image video analysis and monitoring system, comprising: Multimodal acquisition module: Deploys 4K cameras, infrared thermal imagers, and lidar, mounted on tower tops and drone platforms. Utilizes CMOS sensors with built-in GPS, BeiDou, and inertial measurement units to synchronously capture images, videos, and spatiotemporal data, enabling 5G and Mesh ad hoc network communications. Edge processing module: The edge computing node uses an ARM Cortex-A76 processor and an NVIDIA Jetson NXNPU, running a lightweight YOLOv6s object detection model for real-time defect detection. An integrated image enhancement algorithm optimizes image quality in low-light, rainy, and foggy scenes, and compressed feature data is uploaded to the cloud via the MQTT protocol. Cloud-based analysis module: Builds a distributed computing cluster based on Kubernetes, uses a 3D convolutional neural network to analyze the spatiotemporal characteristics of video sequences, and combines it with the Transformer model to predict defect development trends. It also establishes a digital twin of the transmission line, maps equipment status in real time, and conducts retrospective historical data and multi-period comparative analysis. The analysis results are pushed to the client via the WebSocket protocol. Early warning linkage module: Set multi-level alarm thresholds and push notifications via SMS, voice, and APP; integrate automatic work order distribution system, link the nearest operation and maintenance team with the spare parts library, and combine GIS maps to generate the optimal inspection route; link with the intelligent switch of the transmission line to trigger tripping in the event of a fault; Storage security module: uses a distributed file system to store original images and videos, and a time series database to store feature data; uses the national secret SM4 algorithm to encrypt and transmit data, and uses blockchain technology to store key operation records.

[0007] Furthermore, it also includes: Drone collaboration module: The drone fleet is equipped with dual optical payloads and automatically plans inspection routes based on the cloud-based defect distribution heat map. SLAM technology is used to build a three-dimensional model of the transmission corridor, mark defect locations in real time, and capture blind spot images of towers, integrating data from fixed cameras.

[0008] Furthermore, it also includes: Adaptive learning module: Based on the federated learning framework, edge nodes anonymously upload encrypted defect data, and the cloud aggregates and updates the detection model. The ResNet50 model is compressed using knowledge distillation technology and sent to the edge side through OTA differential upgrades.

[0009] Furthermore, the multimodal acquisition module integrates an anti-shake gimbal and an automatic zoom lens, allowing remote control to adjust the field of view angle.

[0010] Furthermore, the edge processing module develops a real-time semantic segmentation algorithm based on the FastFCN model to automatically identify vegetation and buildings around transmission lines.

[0011] Furthermore, the cloud-based analysis module uses a stream computing framework to process real-time video streams and synchronously analyze the video source.

[0012] Furthermore, the early warning linkage module establishes a defect knowledge base and automatically generates maintenance suggestions based on the GPT-3.5 large model.

[0013] Furthermore, it also includes: Energy consumption optimization module: Edge nodes use dynamic voltage and frequency adjustment technology, and the drone fleet uses path optimization algorithm to optimize energy consumption.

[0014] Furthermore, it also includes: Remote interaction module: Develop a web-based visualization platform to zoom and rotate the digital twin of the transmission line through gestures, call historical image and video data, and trigger alarm review through voice commands.

[0015] Furthermore, the storage security module deploys an intelligent retrieval engine to retrieve image content through the engine.

[0016] Compared with the existing technology, the beneficial effects of the present invention are: The system integrates a 4K visible light camera, an infrared thermal imager, and a LiDAR (LiDAR) to enable high-frequency acquisition of multi-dimensional data from transmission lines. The visible light camera supports ultra-low-light imaging, the infrared thermal imager can detect abnormal conductor temperature rise, and the LiDAR accurately measures conductor sag and insulator tilt. Edge computing nodes, equipped with lightweight models, detect various defects in real time, with excellent results for typical defects. The system establishes a three-level alarm system, rapidly triggering tripping of critical defects and notifying operations and maintenance personnel through multiple channels. This reduces defect detection time from several days of manual inspections to real-time.

[0017] Edge computing nodes perform data preprocessing and only upload defect characteristics to the cloud, reducing data transmission and network bandwidth usage. The cloud analyzes defect development trends based on models and provides early warning of potential risks. The intelligent work order system automatically generates optimal inspection routes based on maps and operation and maintenance team locations, shortening the response and defect resolution cycle. The drone collaborative inspection module plans routes based on defect heat maps, increasing the average daily number of towers inspected, reducing coverage blind spots, and improving operation and maintenance efficiency.

[0018] The system utilizes hybrid networking to ensure data transmission in signal-blind areas. Edge nodes can cache and retransmit data during network outages. Anti-shake gimbals and image enhancement algorithms ensure stable imaging even in harsh conditions, ensuring system reliability in complex environments. Data compression and federated learning technologies reduce storage and computing costs, as well as edge node power consumption, reducing annual operation and maintenance costs. Digital twin technology simulates fault scenarios to assist in developing emergency response plans, improve training efficiency, and reduce operational errors among operators.

[0019] The entire process, from data collection to operations and maintenance, is fully automated, reducing manual intervention. Blockchain technology stores key operation records, ensuring data security and meeting security requirements. An adaptive learning module updates detection models, improving the recognition of new defect types and the speed of model inference. In transmission line applications, the system improves the overall health of equipment after commissioning, providing technical support for reliable grid operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 This is a schematic block diagram of a power transmission line image and video analysis and monitoring system proposed by the present invention; Figure 2 This is a schematic diagram of the comparison of multimodal defect detection accuracy; Figure 3 This is a schematic diagram for comparing system response time levels; Figure 4 This is a comparison diagram for edge computing resource optimization; Figure 5 This is a schematic diagram comparing system reliability under different climate conditions; Figure 6 This is a diagram of long-term operation and maintenance cost-benefit analysis. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.

[0023] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be a connection between the two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below with reference to the accompanying drawings.

[0024] Reference Figures 1 to 6 : A transmission line image video analysis and monitoring system, comprising: Multimodal acquisition module: An integrated monitoring device is installed at the top of every 2 kilometers of transmission line towers. The main device uses a Hikvision DS-2CD8643G0-I 4K camera, equipped with a 1 / 1.8-inch starlight-class CMOS sensor and supporting H.265 encoding. It delivers clear color images even in ultra-low-light conditions of 0.001 Lux. The camera features a built-in 12x optical zoom lens (focal length 20-120mm), motorized to achieve ±360° pan and ±90° tilt, covering a sector with a radius of 50 meters. The FLIRT 1040 thermal imager uses an uncooled vanadium oxide detector with 640×480 pixels and a temperature resolution of 0.03°C, capable of detecting surface temperature changes of 0.1°C on the conductor. The device is rigidly connected to the camera via an M12 connector, and the optical axis angle between the two is factory-calibrated to 15°, ensuring spatial consistency between visible and infrared images. The anti-shake gimbal is driven by a high-precision servo motor and features a built-in three-axis gyroscope (±2000° / s range) and an accelerometer (±16g range). A Kalman filter algorithm provides real-time compensation for tower vibration (frequency range 0.5-20Hz). In wind speeds of level 6 (10.8-13.8m / s), image stabilization accuracy is ≤0.01°, and video blur is ≤3%. The drone platform is a DJI Matrice 350 RTK, equipped with dual-redundant IMUs and GNSS modules (supporting GPS, Beidou, and GLONASS positioning). Positioning accuracy is ≤1cm + 1ppm in RTK mode. The Zenmuse H20T multispectral payload includes a 20-megapixel visible light camera, a 640×512 pixel infrared camera, and a 200-meter ranging lidar, enabling simultaneous trimodal data collection. The data acquisition equipment integrates a BeiDou-2 timing module, receives the B1I frequency signal (1561.098 MHz), and achieves sub-microsecond time synchronization via PPS (pulse per second) signals. The equipment automatically performs internal calibration upon startup, and the camera uses Zhang's calibration method to calculate distortion coefficients (radial distortion ≤ 0.1%, tangential distortion ≤ 0.05%). The lidar and camera perform extrinsic calibration using an AprilTag checkerboard grid, achieving rotation matrix error ≤ 0.5° and translation vector error ≤ 2 cm. Data transmission utilizes a hybrid networking strategy, with tower equipment prioritizing data upload via the 5G network (SA mode, band n78 3.5 GHz), achieving a peak throughput of 1.8 Gbps. In 5G signal blind spots, it automatically switches to a mesh ad hoc network (operating at 470 MHz, using LoRa modulation, and a spreading factor of 12), with a single-hop distance of 2 km and support for level 6 relays. Critical alarm data is transmitted via 5G sliced ​​networks, with end-to-end latency ≤ 80 ms and a packet loss rate ≤ 0.1%.

[0025] Edge processing module: The edge computing node uses Advantech UNO-2483G industrial computer, equipped with Intel Atomx6425E processor (4 cores 2.0GHz), 8GB DDR4 memory and 128GB eMMC storage on board. The NPU uses NVIDIA Jetson Orin, which has 1792 CUDA cores, a computing power of 200TOPS, and a power consumption of only 45W. The device supports a wide operating temperature of -20℃~60℃, with a protection level of IP67, and can adapt to harsh outdoor environments. The image preprocessing adopts a three-stage pipeline architecture, and uses the CLAHE algorithm to enhance the contrast (the contrast gain factor is limited to 4.0), uses a 3×3 median filter to remove salt and pepper noise, and uses a Gaussian pyramid to downsample to a 640×640 input size. Assume that the time it takes for the CLAHE algorithm to process a single frame of image is t CLAHE , median filter processing time is t Median , Gaussian pyramid downsampling processing time is t Pyramid , pipeline synchronization and other minor overhead time is t Sync , introduce the single frame delay quantization formula T=t CLAHE +t Median +t Pyramid +t SyncWhere T is the single-frame latency of the entire preprocessing process. Through hardware scheduling optimization and algorithm parallel execution, T is guaranteed to be ≤ 30ms, achieving efficient image preprocessing. The landmark detection model uses a lightweight YOLOv6s. After pre-training on the COCO dataset, transfer learning is performed on the transmission line scenario. During the training phase, MixUp data augmentation (mixing coefficient λ = 0.5) and Mosaic data splicing are used, and the CBAM attention mechanism is introduced to improve small object detection capabilities. After quantization to INT8 precision, the model achieves an inference speed of 40 FPS on Jetson Orin, with an accuracy of ≥95% for broken strand detection of conductors with a diameter of ≥5mm and an F1 score of 0.92 for zero-value defects on insulators. Defect feature extraction utilizes a multi-scale fusion strategy. Shallow features (such as color and texture) are used to identify rust and damage, while deep semantic features (such as HOG and SIFT) are used to determine deformation and displacement. After the feature vector is compressed to 512 dimensions, defect classification is achieved through cosine similarity calculation (with a threshold of 0.75). For complex defects that are difficult to identify in real time (such as early discharge marks), the system automatically tags and uploads the complete video clip to the cloud. Data compression uses a dual-track strategy. The video stream is compressed using the SVT-AV1 encoder (configured with CRF=30 and a keyframe interval of 1 second), with a bitrate controlled within 4Mbps. Structured feature data is serialized using Protobuf and further compressed using the LZ4 algorithm, achieving a compression ratio of 15:1. Edge nodes have a built-in 256GB SSD cache, which can store 72 hours of processed data in the event of a network interruption. Once the network is restored, data is retransmitted via breakpoint resume (supporting HTTP / 2 Range requests), with a retransmission success rate of ≥99.9%.

[0026] Cloud-based analysis module: A Kubernetes cluster was built on Alibaba Cloud Container Service (ACK) in the cloud, deploying 10 ecs.g7.4xlarge instances (24 cores / 96GB of memory) as compute nodes and 3 ecs.r7.2xlarge instances (8 cores / 64GB of memory) as storage nodes. The cluster was configured with a high-availability architecture across three availability zones, using an SLB load balancer for traffic distribution and supporting elastic scalability (minimum 5 nodes, maximum 30 nodes). Real-time stream processing was performed using the Apache Flink 1.15 framework with a parallelism of 200 and a RocksDB state backend (70% memory usage). The system supports concurrent analysis of 200 video streams, with a single-stream processing latency of ≤800ms. Batch processing jobs, running on Spark 3.3.0, perform in-depth mining of historical data (e.g., seasonal defect distribution patterns), completing jobs in ≤2 hours (processing 10TB of data). The digital twin platform, developed using the Unity 2022 LTS engine, builds 1:1 accurate 3D models of transmission lines. The model imports a 10cm-precision DEM (digital elevation model) and a 0.5m-resolution DOM (digital orthophoto). Pole towers, conductors, and other equipment are modeled parametrically (e.g., conductor tension and sag conform to actual mechanical properties). Defect coordinates (latitude and longitude accuracy to 0.0001°) uploaded from the edge are received in real time via the WebSocket protocol (with a heartbeat interval of 30 seconds) and displayed in the virtual scene as colored markers (yellow: general defects, red: emergency defects). The 3D-CNN model uses an I3D architecture and inputs a 10-frame video sequence (with a time step of 0.5 seconds). It extracts spatiotemporal features using a 3D convolution kernel (3×3×3). After pre-training on ImageNet, the model is fine-tuned on a transmission line defect dataset, introducing a Temporal Shift Module (TSM) to enhance temporal modeling capabilities. Let the spatiotemporal feature sequence extracted by the 3D-CNN be X = {x1, x2, ..., x N} (where N=10 is the number of input frames, x i is the feature vector corresponding to the i-th frame), in the Transformer prediction module, the weight of the feature sequence processed by the attention mechanism is W={w1,w2,...,w N}, build defect development prediction formula Where X is the spatiotemporal feature sequence corresponding to the 10-frame video sequence extracted by 3D-CNN, and each element x i is a high-dimensional feature vector used to characterize the spatiotemporal information of the transmission line in the frame video; W is the feature sequence weight calculated by the attention mechanism in the Transformer prediction module, w i Represents the feature x of the i-th frame iThe attention weights represent the importance of different frame features to the prediction result. ⊙ represents the element-by-element multiplication operation, which integrates the attention weights with the features of the corresponding frame, highlighting the role of key frame features in prediction. The Transformer prediction module's computational process maps the fused feature sequence into the prediction space through multiple layers of self-attention and feedforward networks. The prediction module uses the Transformer architecture to map the feature sequence extracted by the 3D-CNN into the defect development space for the next seven days. Output parameters include crack growth rate (mm / day) and ice growth thickness (cm / day).

[0027] Early Warning Linkage Module: A three-level alarm system is established, with thresholds dynamically adjusted based on machine learning. A Level 1 (yellow) alarm is triggered when a conductor temperature ≥ 60°C (for more than 10 minutes), an insulator surface temperature gradient ≥ 1°C / cm, ice thickness ≥ 5mm, or a conductor sag change rate ≥ 5%. Alarm information is pushed to the web client via WebSocket (with a defect image and location KML file), simultaneously triggering a vibration alert in the app (supporting iOS and Android). A Level 2 (orange) alarm is triggered when the insulator damage area ≥ 10%, the shock absorber displacement distance ≥ 5cm, or 1-2 conductor strands are broken. The alarm is sent to the operation and maintenance personnel's mobile phone via text message (including a GPS location link) within 0.5 seconds, and a work order is automatically created and pushed to the work order system. A Level 3 (red) alarm is triggered when the insulator damage area ≥ 20%, 3 or more conductor strands are broken, or hardware deformation ≥ 15%. The system immediately sends a voice alarm (speech synthesis clarity ≥ 95%) to the three nearest operation and maintenance teams, and simultaneously triggers the tripping of nearby smart switches (tripping time ≤ 100ms) through the OPC UA protocol. The work order system automatically generates standardized task orders based on RPA technology, including defect type, location (accurate to tower number + GPS coordinates), processing priority (high / medium / low), and recommended processing solutions. The work order dispatch uses the ant colony optimization algorithm, which comprehensively considers factors such as the location of the operation and maintenance personnel (GPS positioning accuracy ≤ 5m), skill matching (such as whether they hold a high altitude permit), and task load, to define the task-personnel fit formula. . Among them F i,j D is the compatibility between the i-th work order task and the j-th operation and maintenance personnel; i,j S is the GPS distance between the operation and maintenance personnel j and the work order task i; i,j is the skill matching coefficient; L j is the current workload of operator j; α, β, and γ are weight coefficients (satisfying α + β + γ = 1) used to balance the impact of location, skill, and workload on dispatch. The GIS route planning module combines real-time traffic data from Amap (update frequency ≤ 30 seconds) to avoid congested roads and calculate the theoretical arrival time (error ≤ 10%).

[0028] Storage Security Module: Raw video data is stored in Alibaba Cloud OSS object storage, using standard storage (read and write response time ≤ 100ms) for 90 days. Upon expiration, data is automatically archived to cold storage (reducing costs by 70%). Feature data is written to the InfluxDB time series database, partitioned into three levels: device ID, parameter type, and time slice (time slices are 1 hour). Three replicas are configured for redundancy, ensuring data durability of ≥ 99.999%. To improve retrieval efficiency, the system builds a multi-level index. The primary index uses an inverted index based on Elasticsearch, enabling full-text searches for metadata such as defect type, time range, and severity, with a response time of ≤1.5 seconds. The secondary index uses a B-Tree-based spatiotemporal index, using an R-Tree spatial partitioning of GPS coordinates (spatial granularity of 100m × 100m), supporting complex queries such as "all defects within 500 meters of a certain point." The tertiary index uses an LSM-Tree-based time series index, creating a timestamp index for time series data such as temperature and sag, supporting fast filtering within time windows (configurable window sizes of 1 hour, 1 day, or 1 week). Data transmission is encrypted using the SM4 algorithm (128-bit key length), a 16-byte encryption block size, and CTR (Counter Mode) to ensure efficient parallel encryption. Key management complies with NIST SP800-131A standards, with the master key stored in an HSM hardware security module (PCIe interface, random number generation rate ≥ 100Mbps). The data encryption key (DEK) is automatically rotated every 24 hours. The blockchain platform is built on a consortium chain based on Hyperledger Fabric 2.4, whose members include operations and maintenance departments, equipment manufacturers, and regulatory agencies. Each block contains the first 100 operation records (such as alarm confirmation, work order processing, and equipment replacement). The block size is ≤2MB and the generation interval is ≤2 seconds. Transaction signatures are signed using the ECDSA algorithm (P-256 curve), and the consensus mechanism is PBFT (Practical Byzantine Fault Tolerance), ensuring that the system can continue to operate normally even if ≤1 / 3 of the nodes fail.

[0029] The present invention also includes the following modules: Drone Collaboration Module: Drone inspections utilize a "fixed + dynamic" two-tier planning strategy. Fixed routes cover routine inspection points (such as special sections crossing railways and rivers) and are generated using a grid algorithm (grid size 5m×5m), ensuring an image overlap rate of ≥70%. Dynamic routes are adjusted in real time based on cloud-based defect heat maps, prioritizing inspections of high-risk areas (such as historical fault-prone areas and newly discovered potential hazards). The drone's LiDAR point cloud is fused with the tower camera imagery through feature matching. Planar features (such as insulator surfaces) are extracted from the point cloud, and their normal vectors are calculated. SIFT feature points are extracted from the image, and the transformation matrix is ​​estimated using the RANSAC algorithm (500 iterations, inlier threshold 0.01m). The fused data is then reconstructed using a Poisson surface reconstruction algorithm to generate a 3D model with a surface accuracy of ≤5cm.

[0030] The present invention also includes the following modules: Adaptive Learning Module: The system adopts a horizontal federated learning architecture, with edge nodes acting as clients and the cloud as an aggregation server. During training, edge nodes calculate gradients using local data (batch size = 16, learning rate 0.001), which are encrypted using homomorphic encryption (Paillier encryption scheme). After uploading the encrypted gradients to the cloud, the server aggregates model parameters using a weighted average algorithm (weighted by the data volume of each node) and decrypts them to generate a global model. During knowledge distillation, ResNet50 serves as the teacher model and MobileNetV3 as the student model. The softmax probability distribution (temperature parameter T = 4) output by the teacher model serves as the soft label, which, together with the true label, guides the training of the student model. An attention transfer mechanism is introduced to enable the student model to learn the distribution of the intermediate feature maps of the teacher model. Ultimately, the number of parameters is compressed from 25.6M to 3.4M and the computational load is reduced from 4.1GMACs to 0.22GMACs, while maintaining 92% of the original accuracy.

[0031] In this invention, the multimodal acquisition module integrates an anti-shake pan / tilt (PTZ) and an auto-zoom lens, creating a highly stable and adaptive acquisition hardware system. The anti-shake pan / tilt utilizes high-precision servo control technology and a built-in angular displacement sensor with a resolution of 0.001°, enabling sensitive capture of subtle posture changes. Coupled with an adaptive PID control algorithm, it dynamically adjusts parameters based on wind speed and equipment motion, compensating for external disturbances in real time with a stability accuracy of ≤0.01°. The motor can quickly fine-tune the pan / tilt to withstand disturbances such as gusts and base vibration, providing a stable foundation for image acquisition. The auto-zoom lens covers focal lengths from 20-120mm and utilizes a combination of ultra-low dispersion optical glass and aspherical lenses. The former suppresses light refraction deviation and reduces chromatic aberration, while the latter optimizes spherical aberration and enhances edge definition. A built-in stepper motor drive system supports 0.1mm focal length adjustment. Combined with the PTZ, the field of view can be flexibly and remotely adjusted, balancing local details (such as hardware and insulator defects) with wide-angle coverage, adapting to diverse inspection scenarios. In terms of environmental adaptability, the module has undergone rigorous wind load and stability testing. Its streamlined design, combined with a high-strength lightweight alloy casing, effectively reduces wind resistance. In wind speeds of level 6 (10.8-13.8m / s), the anti-shake gimbal and lens optical image stabilization work together, supplemented by image blur assessment (edge ​​gradient detection, clarity entropy calculation) and post-processing compensation (inter-frame motion estimation, image registration and fusion), to control image blur to ≤5%. Even in strong winds that cause the device to shake, it still produces clear, usable images and videos, providing a high-quality data source for accurate identification during transmission line inspections.

[0032] In this invention, the edge processing module deeply develops a real-time semantic segmentation algorithm based on the FastFCN model, focusing on transmission line channel clearing scenarios. The model structure is optimized, replacing some standard convolutional layers with depthwise separable convolutions to reduce computational complexity. A multi-scale feature fusion module is also constructed, leveraging a channel attention mechanism to weightedly fuse convolutional feature maps from different levels, accurately capturing the contours and details of vegetation and buildings in the environment. A dedicated dataset is constructed for power transmission inspection data, collecting line images under different geographical environments, seasons, and lighting conditions. After annotation, a training set of over 100,000 images containing categories such as vegetation and buildings is formed. Transfer learning is employed, with basic capabilities first pre-trained on general datasets (such as Cityscapes) before fine-tuning on a dedicated dataset to adapt to the visual characteristics of power transmission scenarios. During inference deployment, the algorithm leverages edge computing device hardware acceleration (GPU and NPU collaboration), model quantization (floating point to INT8 precision), and operator fusion optimization to improve edge operational efficiency. In actual applications, vegetation identification can accurately distinguish types with an accuracy of ≥92%; it can accurately frame the outlines and determine the categories of buildings, providing information such as the scope of vegetation encroachment, the safe distance between buildings and lines for channel clearing, helping operation and maintenance personnel to make scientific clearing decisions and realize intelligent identification and active prevention and control of channel hidden dangers.

[0033] In this invention, the cloud-based analysis module builds a real-time video stream processing system based on the stream computing framework (Flink). With its low latency and high throughput, Flink is well-suited for power transmission line monitoring scenarios. Through customized optimization, it supports multiple protocol inputs, such as RTMP and RTSP, at the data access layer. Its adaptive stream parsing component intelligently identifies codecs like H.264 and H.265, rapidly decapsulating and extracting frames. Distributed access nodes are used to distribute the burden of over 200 video sources. Preprocessing units remove redundancy and filter static frames, reducing subsequent load. The computing layer, based on Flink's stateful stream processing, features a dedicated operator chain for anomaly detection. A custom discriminant operator, integrating CNN feature extraction, rules, and machine learning, extracts video frame features and detects anomalies such as icing. Multi-frame data is aggregated using a sliding time window (1 second) to improve detection accuracy, ensuring an average processing latency of ≤1 second. Incremental learning is implemented, allowing operators to update models online based on labeled data for optimal performance. For quality assurance, a multi-layered verification system is implemented. Parallel detection is implemented by primary and backup operators, with cross-window backtracking and secondary confirmation of anomalies, ensuring a false negative rate of ≤0.5%. The module also has monitoring and scheduling functions, which can monitor Flink cluster resources in real time, automatically schedule and expand capacity when the load is uneven or tight, ensure stable analysis of more than 200 video sources, and provide reliable monitoring and early warning for intelligent power transmission operation and maintenance.

[0034] In this invention, the early warning linkage module builds an intelligent decision-making system that integrates a defect knowledge base and a large-scale model. The defect knowledge base collects over 100,000 historical transmission operation and maintenance cases, covering defect type, location, environment, and the entire handling process. After cleaning and constructing a knowledge graph, it forms a structured resource pool. The module deeply integrates the GPT-3.5 large-scale model to build an inference engine. When a new defect is identified (such as a broken insulator on a tower), similar cases are first retrieved from the knowledge base, key features are extracted, converted into natural language, and input into the large-scale model. The large-scale model, in conjunction with industry standards, generates maintenance recommendations through reasoning, including handling priority, action plan, and timeline (e.g., recommending replacement within three days, integrating operation windows and material lifecycles). To ensure effectiveness, a closed-loop feedback mechanism is established. After recommendations are pushed to the operation and maintenance end, adoption and handling results are tracked, and data is fed back to the knowledge base and model optimization module. The knowledge base incrementally learns and enriches cases, while the large-scale model is fine-tuned and trained to continuously improve the quality of recommendations. A proven recommendation adoption rate of ≥85% has been achieved, helping O&M transition from passive fault elimination to proactive early warning and intelligent decision-making.

[0035] The present invention also includes the following modules: Energy Optimization Module: Edge nodes utilize Dynamic Voltage and Frequency Scaling (DVFS) technology and a built-in intelligent power management unit to monitor computing load in real time. During idle time (load factor <10%), a hardware regulator reduces the core voltage to 0.8V and the CPU frequency to 500MHz, reducing power consumption to ≤5W and static power consumption by over 60%. During active mode (load factor ≥30%), voltage and frequency are dynamically adapted based on load, ensuring performance while maintaining power consumption ≤25W and precisely controlling energy consumption. The drone fleet uses an improved ant colony algorithm to optimize routes. During mission planning, it collects geographic (tower coordinates, terrain, etc.), meteorological (wind speed, temperature), and drone status (battery charge, payload) data. Aiming to achieve "minimum energy consumption and optimal coverage," the algorithm utilizes "artificial ants" and pheromone simulations to search for routes, combined with heuristic factors for node selection. Multiple iterations of route generation reduce inefficient flight distance by 20%. Energy consumption feedback is also incorporated, collecting real-time power and attitude data to dynamically adjust pheromone strategies, reducing overall energy consumption by 15%, achieving efficient and energy-saving inspections and helping the system reach industry-leading energy efficiency standards.

[0036] The present invention also includes the following modules: Remote Interaction Module: The front-end is built using the React framework, integrated with Three.js for 3D scene rendering. During scene initialization, vertex and fragment shaders are created using WebGL 2.0, supporting instanced rendering (capable of rendering 10,000 tower models in a single batch). The interaction layer implements gesture recognition: pinch-to-zoom (zoom ratio range: 1:1-1:500), rotate with two fingers for viewing angle adjustment (rotation sensitivity: 0.5° / pixel), and slide with one finger for panning (movement speed: 0.1m / pixel). Voice interaction integrates the Baidu Speech Recognition SDK and supports Mandarin Chinese and multiple dialects (such as Cantonese and Sichuanese). The system uses a wake-up word, "Xiaodian Assistant," with an accuracy rate of ≥98%. Upon wakeup, various commands can be executed. Command parsing utilizes a BERT-based intent classification model, fine-tuned on specialized power transmission line data, achieving a classification accuracy rate of ≥97%. System response time is ≤1.5 seconds, and complex queries (such as cross-month data comparisons) are ≤3 seconds.

[0037] In the present invention, the storage security module deploys an intelligent retrieval engine based on Elasticsearch. For the transmission line image data, a multi-layer index system is constructed, and coarse-grained indexes are set according to the tower area, voltage level, etc. CNN is used to extract the target visual features such as bird nests and convert them into 128-dimensional vectors, and an inverted index is established for them. During retrieval, visual and text features are integrated, the visual features are extracted by a pre-trained model, and the text features are semantically matched using the BM25 algorithm to ensure retrieval requirements such as "finding tower images with bird nests", with an accuracy rate of ≥90%. At the same time, relying on Elasticsearch sharding and parallel query, combined with GPU using the CUDA framework to accelerate feature vector similarity calculations, the efficiency is improved by 5-8 times, and with a high-frequency retrieval cache strategy, it is ensured that the response time is ≤3 seconds for tens of millions of data volumes, laying a solid technical foundation for the efficient use of transmission line inspection data and abnormality detection.

[0038] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A power transmission line image video analysis and monitoring system, characterized in that: Includes the following modules: Multimodal acquisition module: Deploys 4K cameras, infrared thermal imagers, and lidar, mounted on tower tops and drone platforms. Utilizes CMOS sensors with built-in GPS, BeiDou, and inertial measurement units to synchronously capture images, videos, and spatiotemporal data, enabling 5G and Mesh ad hoc network communications. Edge processing module: The edge computing node uses an ARM Cortex-A76 processor and an NVIDIA Jetson NX NPU, running a lightweight YOLOv6s object detection model for real-time defect detection. An integrated image enhancement algorithm optimizes image quality in low-light, rainy, and foggy scenes, and compressed feature data is uploaded to the cloud via the MQTT protocol. Cloud-based analysis module: Builds a distributed computing cluster based on Kubernetes, uses a 3D convolutional neural network to analyze the spatiotemporal characteristics of video sequences, and combines it with the Transformer model to predict defect development trends. It also establishes a digital twin of the transmission line, maps equipment status in real time, and conducts retrospective historical data and multi-period comparative analysis. The analysis results are pushed to the client via the WebSocket protocol. Early warning linkage module: Set multi-level alarm thresholds and push notifications via SMS, voice, and APP; integrate automatic work order distribution system, link the nearest operation and maintenance team with the spare parts library, and combine GIS maps to generate the optimal inspection route; link with the intelligent switch of the transmission line to trigger tripping in the event of a fault; Storage security module: uses a distributed file system to store original images and videos, and a time series database to store feature data; uses the national secret SM4 algorithm to encrypt and transmit data, and uses blockchain technology to store key operation records.

2. The power transmission line image video analysis and monitoring system according to claim 1, characterized in that: Also includes: Drone collaboration module: The drone fleet is equipped with dual optical payloads and automatically plans inspection routes based on the cloud-based defect distribution heat map; SLAM technology is used to build a three-dimensional model of the transmission corridor, mark defect locations in real time, retake images of blind spots in towers, and integrate fixed camera data.

3. The power transmission line image video analysis and monitoring system according to claim 1, characterized in that: Also includes: Adaptive learning module: Based on the federated learning framework, edge nodes anonymously upload encrypted defect data, and the cloud aggregates and updates the detection model; The ResNet50 model is compressed using knowledge distillation technology and delivered to the edge through OTA differential upgrades.

4. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: The multimodal acquisition module integrates an anti-shake pan / tilt and an automatic zoom lens, allowing remote control to adjust the field of view.

5. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: The edge processing module develops a real-time semantic segmentation algorithm based on the FastFCN model to automatically identify vegetation and buildings around transmission lines.

6. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: The cloud analysis module uses a stream computing framework to process real-time video streams and synchronously analyze video sources.

7. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: The early warning linkage module establishes a defect knowledge base and automatically generates maintenance suggestions based on the GPT-3.5 large model.

8. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: Also includes: Energy consumption optimization module: Edge nodes use dynamic voltage and frequency adjustment technology, and the drone fleet uses path optimization algorithm to optimize energy consumption.

9. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: Also includes: Remote interaction module: Develop a web-based visualization platform to zoom and rotate the digital twin of the transmission line through gestures, call historical image and video data, and trigger alarm review through voice commands.

10. The power transmission line image and video analysis monitoring system according to claim 1, characterized in that: The storage security module deploys an intelligent retrieval engine to retrieve image content.

Citation Information

Cited By

  • Power transmission line typical defect image big data analysis system and method based on big data

    CN121258976A

  • Photovoltaic dust thickness real-time inversion method and system

    CN121520988A