Multi-modal data enhancement and intelligent fault identification method for power transmission line under extreme weather
By combining multimodal data collaborative acquisition and deep data augmentation pipeline with LoRA fine-tuned visual-language large model, the problems of multimodal data governance and long-tailed fault identification in transmission line inspection under extreme weather conditions are solved, and efficient and accurate fault diagnosis of transmission lines is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUBEI ELECTRIC POWER TRANSMISSION & DISTRIBUTION ENG
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-07
AI Technical Summary
Under extreme weather conditions, existing technologies struggle to achieve efficient and automated inspections of power transmission lines, especially given the degradation of images at disaster sites and the long-tailed distribution of fault types. They also lack support from multimodal data governance, structured diagnostics, and natural language descriptions.
By adopting multimodal disaster data collaborative collection and standardized governance, combined with a deep data augmentation pipeline and a LoRA-based fine-tuned visual-language large model, we can achieve structured management and intelligent fault identification of multi-source heterogeneous inspection data.
Under extreme weather conditions, it has achieved standardized governance of multimodal data and accurate fault identification, and output structured natural language diagnostic results, meeting the real-time diagnostic needs of power faults.
Smart Images

Figure CN122347768A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to automated inspection technology, and more particularly to a method for multimodal data enhancement and intelligent fault identification of transmission lines under extreme weather conditions. Background Technology
[0002] Against the backdrop of global climate change, extreme weather events are becoming more frequent, intense, and concurrent. Disasters such as ice storms, freezing rain, torrential rain, typhoons, and dense fog pose severe threats to power transmission lines, tower structures, and ancillary facilities, leading to frequent accidents such as line breaks, tower collapses, and insulator damage. Post-disaster emergency inspections are crucial for ensuring power grid safety and minimizing power outages. However, traditional inspection methods rely primarily on manual on-site visual interpretation, which is not only inefficient and subjective but also struggles to cover large areas of power lines under extreme weather conditions. In recent years, automated inspection technology based on drone aerial photography has gradually become more widespread, but it still faces technical bottlenecks in complex post-disaster environments, including unclear visibility, inaccurate judgments, and slow calculations.
[0003] At the data level, image degradation caused by extreme weather disasters is particularly prominent. The scattering and absorption of light by atmospheric particles in freezing rain and dense fog leads to severe contrast attenuation, color shift, and loss of high-frequency details in inspection images, directly affecting subsequent algorithms' ability to identify minute defects such as broken conductor strands and cracked insulators. Furthermore, disaster samples are highly sporadic and unreproducible; severe fault types (such as tower collapse and line breakage) account for a very low percentage of the total sample, exhibiting a typical long-tail distribution. While existing research has made rapid progress in image enhancement and dehazing algorithms, multi-layered enhancement methods for extreme environments such as dense fog and smoke in power disaster inspection scenarios remain incomplete, lacking a standardized end-to-end processing system from data acquisition to model input.
[0004] At the fault identification level, existing methods mostly employ single-modal image classification or object detection frameworks based on convolutional neural networks, which have several inherent limitations. Single-modal visual models cannot output natural language semantic descriptions of faults, only providing category labels and confidence scores, which is insufficient to meet the needs of engineering applications for structured diagnostic results regarding fault type, spatial location, and damage extent. Traditional models lack prior understanding of power industry expertise, making them prone to misjudgment and missed detection in disaster sites with severe background interference and frequent occlusion. Low-frequency fault types suffer from insufficient model recall due to scarce samples, and existing data augmentation strategies are mostly general designs, lacking targeted augmentation mechanisms for the fault morphology and long-tail distribution characteristics of power disasters. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies, such as the lack of standardization in multimodal disaster data governance, insufficient enhancement of long-tailed fault samples, and lack of structured natural language diagnostic methods. This invention provides a method for multimodal data enhancement and intelligent fault identification of transmission lines under extreme weather conditions, enabling structured management of multi-source heterogeneous inspection data, effective expansion of low-frequency fault samples at the tail end, and accurate identification of power faults and output of structured natural language diagnostic results based on a large vision-language model.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: This invention provides a method for multimodal data augmentation and intelligent fault identification of transmission lines under extreme weather conditions. S1. Multimodal disaster data collaborative collection, standardized governance, and spatiotemporal alignment: The pre-disaster-post-disaster comparative grouping strategy with a single fault as the smallest governance unit, the five-tuple unique identifier coding, the 17 types of typical fault classification and three-level labeling system, the structured naming and location indexing mechanism, and the local geometric mapping model based on similarity transformation realize the cross-temporal tailoring of pre-disaster-post-disaster defect areas. S2, Depth Data Augmentation Pipeline for Long-Tail Distributions: A unified enhancement operator framework based on Bernoulli random variable control, a combination of multiple geometric transformations and photometric perturbation operators, a physical dehazing enhancement step based on dark channel priors, and channel-based statistical normalization processing. S3. Deployment of intelligent fault identification and reasoning for large visual-language models based on LoRA fine-tuning: This includes targeted LoRA low-rank adaptation for the three main targets of Qwen2.5-VL, a differentiated matrix initialization strategy, a combined loss function of weighted cross-entropy and multimodal consistency, a third-order training pacing, and an inference acceleration deployment scheme integrating Paged Attention and Continuous Batching.
[0007] Furthermore, in S3, the LoRA low-rank adaptation and differential matrix initialization strategy includes using high-variance Gaussian initialization to enhance fault feature exploration. And using low-variance Gaussian initialization to protect pre-trained features .
[0008] Furthermore, in S3, the three-stage training rhythm includes: a warm-up phase, i.e., 1 to 3 rounds, with a learning rate of 1 / 10 of the initial value; an optimization phase, i.e., 4 to 15 rounds, using a differentiated learning rate and adding difficult high-weight samples; and a convergence phase, i.e., 16 to 20 rounds, with the learning rate decaying to 1e-6 and early stopping triggered by the validation set.
[0009] Furthermore, in S3, the Paged Attention paging KV cache and Continuous Batching iterative dynamic batch processing are integrated into the vLLM inference framework, and 12GB GPU multi-path parallel inference is achieved by combining 8-bit quantization.
[0010] Furthermore, in S2, the channel-based statistical normalization at the end of the enhanced pipeline works in conjunction with the Batch Normalization within the network. The channel-based statistical normalization at the end of the enhanced pipeline controls the numerical scale at the input end; the Batch Normalization within the network dynamically calibrates intermediate features, thereby reducing the training instability caused by the large enhancement.
[0011] The beneficial effects of this invention are as follows: In terms of multimodal data governance, a standardized governance method for multimodal disaster data under extreme weather disaster conditions has been established. Through pre-disaster-post-disaster comparative grouping, unique five-tuple identifiers, a 17-category fault classification system, and a unified naming index mechanism, structured management of multi-source heterogeneous inspection data is achieved. Furthermore, by combining a spatiotemporal alignment method based on local geometric mapping, paired samples with consistent fields of view are obtained, solving the problems of missing metadata and inability to perform comparative analysis in traditional inspection data. In terms of data augmentation, a deep data augmentation pipeline oriented towards long-tail distribution was designed. Under a unified augmentation operator framework, geometric transformation, photometric perturbation, physical dehazing based on dark channel priors, and statistical normalization processing are integrated. By constructing an equivalence class multi-view sample family through Bernoulli random combination, the low-frequency fault samples at the tail are effectively expanded, solving the problem of insufficient recall rate of key low-frequency faults in the deep model under imbalanced data. In the area of intelligent fault identification, a vision-language large-scale model-based intelligent power fault identification method based on LoRA low-rank adaptation is proposed. A differentiated matrix initialization strategy is designed for the three major targets of Qwen2.5-VL, combined with a weighted cross-entropy and multimodal consistency loss function and a third-order training pacing. This achieves accurate power fault adaptation with only about 0.1% parameter updates, while retaining general multimodal understanding capabilities. Furthermore, an inference acceleration deployment scheme integrating Paged Attention and Continuous Batching is constructed, enabling the model to meet real-time diagnostic response requirements on mainstream industry hardware and directly output structured natural language fault diagnosis results. Attached Figure Description
[0012] Figure 1 A flowchart illustrating a method for multimodal data augmentation and intelligent fault identification of transmission lines under extreme weather conditions; Figure 2 This is a technical roadmap for multimodal data enhancement and intelligent fault identification methods for transmission lines under extreme weather conditions. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0014] Specifically, S1 is: Please see Figure 1 Using a single fault as the smallest governance unit, a pre-disaster and post-disaster comparative sample group was constructed. Each comparative unit consisted of pre-disaster and post-disaster images of the same line, tower, phase, and orientation, minimizing unstructured interference. A five-tuple unique identifier coding rule (line name, tower number, phase, orientation, and component name) was established and applied throughout the entire sample retrieval and training process. A classification system covering 17 typical fault modes was constructed (broken wire, broken strand, dropped wire, twisted wire, stepping wire, spacer bar failure, protective wire damage, sagging, loosening, surge arrester failure, insulator damage, shielding ring deformation, slippage, wire clamp loosening, collapse, jumper abnormality, foreign object attachment). The labeling process employs a three-tiered process: initial labeling, review, and arbitration.
[0015] The file naming uses a structured field concatenation strategy to embed location information, and at the same time, a location index table is built at the underlying database level to establish a stable mapping between data samples and physical space; At the spatiotemporal alignment level, a simplified geometric mapping is used to achieve the cross-temporal projection of the clipping box, assuming the pre-disaster image is... Post-disaster images are The mapping relationship is .in, Images of the aftermath of the disaster The homogeneous coordinates of any pixel in the array. For this pixel in the pre-disaster image The homogeneous coordinates of the corresponding pixels. This is a two-dimensional planar transformation matrix describing the local geometric correspondence between two temporal phases. This mapping establishes a pixel-level correspondence between pre-disaster and post-disaster images, where any pixel in the post-disaster image... After transformation After the operation, the image can be restored to its corresponding location in the pre-disaster image in homogeneous coordinates. This provides a mathematical basis for the subsequent reverse projection of post-disaster defect boxes onto pre-disaster images, and thus a similarity transformation model is adopted: ; in, It is the relative rotation angle; These are translation parameters; As a scaling factor, a set of stable structural points is selected near the defect. The transformation parameters are solved by least-squares fitting to obtain the defect boxes in the post-disaster images. After mapping the four corner points, construct the corresponding clipping box before the disaster using the circumscribed rectangle to obtain paired close-up samples with consistent field of view.
[0016] Specifically, S2 is: To address the long-tail distribution characteristics of fault samples, a deep data augmentation pipeline is constructed, treating augmentation as a random transformation operator. ,in, To augment the parameters of the random variable, the augmented sample is... Within the empirical risk minimization framework, the training objective is expanded to an expected fit to the perturbation distribution: ; in, These are the model parameters to be optimized; Indicated by model parameters To optimize the variable to its minimum value; Let be the outer expectation, representing the expectation of training samples. Follows the true data distribution Sampling and calculating the expected value; The inner expectation represents the enhancement parameter. Follows its prior distribution Sampling and calculating the expected value; For A random augmentation operator with parameters is applied to the input sample. The resulting enhanced sample; For A prediction model with parameters; The true labels for the samples; Let be the loss function. The overall meaning of this objective function is to minimize the loss between the model's predicted output for augmented samples and the true label under the dual expectations of the true data distribution and the augmented perturbation distribution. This extends the empirical risk minimization framework from single-point fitting to expected fitting of the perturbation distribution, so that each labeled sample is no longer regarded as a single data point, but constitutes an equivalence class induced by augmentation. The model must maintain consistent output on this equivalence class, thereby obtaining stronger invariance and equivariance.
[0017] The enhancement method consists of a random composite of multiple operators, and is provided with If there are several enhancement operators, then the composite enhancement formula is: ; Whether each operator is enabled is determined by a Bernoulli random variable. Decide, The core mathematical model of the enhancement operator is as follows: Horizontal flip: Simulate the left and right positions of the equipment under different flight paths. Assume the image width is... The coordinates are mapped as follows: ; in, The x-coordinate of a pixel in the original image; This represents the x-coordinate of the corresponding pixel in the horizontally flipped image. The ordinate of a pixel in the original image; This represents the ordinate of the corresponding pixel in the horizontally flipped image, and its value is... Equal, demonstrating that horizontal flipping does not change the pixel's vertical coordinate; This transformation belongs to the dihedral group generator operation, preserving local distances and angles. The bounding box is updated synchronously. The segmentation mask is transformed into ; Random rotation: Simulates image rotation caused by drone gusts and gimbal pitch. Centered on the image... For anchor point, rotation angle Sampling from a uniform distribution, the affine transformation is: ; After rotation, the coordinates fall on a non-integer grid, using bilinear interpolation. To obtain pixel values, the bounding box is the outer rectangle of the rotated four corner points: ; Random brightness adjustment: To simulate changes in illumination such as solar altitude angle and cloud cover, the linear model includes additive offset and multiplicative gain. ; in, It is a multiplicative gain; To achieve additive shift, gamma correction is introduced to simulate the nonlinear response of the camera. , ; but ,in The dark areas are compressed, simulating underexposure; Dark details are brightened to simulate low dynamic range scenes; Random contrast adjustment: To simulate dynamic range changes caused by smog and lens grease, the average grayscale value of each channel is defined. Contrast coefficient is ; ; Gradient magnitude scaled proportionally From the perspective of atmospheric scattering models, low contrast is approximated as Reduce contrast to simulate transmittance The process of decreasing.
[0018] HSV color space perturbation: The RGB three channels are strongly correlated, and direct perturbation is prone to producing artifacts. Therefore, decoupling and independent perturbation are performed in the HSV space:
[0019] Hue is represented by modulo operations to reflect periodicity, while saturation and brightness simulate purity and illumination. After transformation, it is inversely transformed back to RGB space.
[0020] Gaussian blur: To simulate the loss of sharpness caused by focusing lag, lens contamination, etc., a two-dimensional Gaussian kernel convolution is used: ; Controlling the fuzzy scale. Gaussian kernel Fourier transform. It decays exponentially with increasing frequency, making it a typical low-pass filter.
[0021] Sharpening: The USM algorithm is used to enhance high-frequency components. Calculate high-frequency residuals Superimposed by intensity: ; threshold To prevent noise amplification in flat areas, combined with blur enhancement, it covers the full spectrum of sharpness from soft to hard image quality.
[0022] Physical dehazing based on dark channel prior: Extreme weather conditions such as freezing rain and dense fog can degrade images. Physical restoration of images is performed based on atmospheric scattering models. ; in, Radiation in fog-free scenarios; For global atmospheric light; For transmittance, the dark channel is defined as... In a haze-free image, the dark channel approaches zero. Global atmospheric light is located and estimated in the original image through the brightest 0.1% region of the dark channel. The coarse estimate of transmittance is: ; A value of 0.95 is used to retain a small amount of residual fog, and the solution is obtained through soft matting. To obtain the fine transmittance, a lower bound threshold is introduced for scene reconstruction: ; After completing the composite enhancement, perform channel-based statistical normalization, assuming the enhanced pixel is... For each channel : ; The statistics use a global benchmark from the dataset, emphasizing consistent scale across samples. Input standardization and Batch Normalization within the network work together; the former controls the numerical scale at the very beginning, while the latter dynamically calibrates intermediate features. Using them together can significantly reduce training instability caused by large-scale augmentations.
[0023] Specifically, S3 is: Model selection: Qwen2.5-VL was selected as the base model. Its dynamic resolution processing supports native input of 4 to 16384 visual tokens. The ViT architecture, combined with window attention and Swiglu activation, has advantages in fine-grained visual feature extraction. It supports LoRA lightweight fine-tuning and the deployment resources meet the constraints of the power industry. LoRA low-rank adaptation LoRA in pre-trained weights Add a low-rank bypass , The weights after fine-tuning are: ; Freeze the entire process, training only. , LoRA bypasses are applied to the three main targets of Qwen2.5-VL: In the high-level layers of the visual encoder (ViT layers 10-16), the Q / K / V and FFN layers are responsible for fine-grained fault feature extraction; the cross-modal attention layer Query / Key projection is responsible for binding terms to image regions; and the fault classification output layer is responsible for mapping to a dedicated classification space. All targets are unified. The total number of parameters is approximately 6 million (accounting for 0.12%).
[0024] Differential matrix initialization: High-variance Gaussian initialization is used. Enhance fault characteristic exploration by employing low-variance Gaussian initialization. Protect pre-training features.
[0025] Combined loss function: The core is weighted cross-entropy (with minor faults weighted 2-3 times), and the auxiliary is multimodal consistency loss. Total losses: ; The three-stage training rhythm is as follows: In the warm-up phase (rounds 1-3), the learning rate is 1 / 10 of the initial value, and basic associations are established using core fault samples; In the optimization phase (rounds 4-15), a differentiated learning rate is adopted (5e-4 for the visual layer, 3e-4 for the cross-modal layer, and 1e-4 for the output layer), and difficult samples are added and given high weights; In the convergence phase (rounds 16-20), the learning rate decays to 1e-6, and early stopping is triggered by the validation set.
[0026] Inference acceleration and deployment: Paged Attention borrows from operating system paging memory management, dividing the key-value cache into fixed-size memory pages for distributed storage. Efficient addressing is achieved through page table mapping, eliminating GPU memory fragmentation. For power scenarios, the memory pages are adapted to 768-dimensional feature dimensions, with each page storing 128 visual tokens in a key-value cache. Continuous Batching employs iterative dynamic batch processing, checking the queue and replenishing new requests after each token generation round to achieve seamless resource integration, and a priority scheduling strategy is designed. These two technologies are integrated into the vLLM inference framework, combined with 8-bit quantization, enabling a 12GB GPU to support multi-path parallel inference, with end-to-end latency controlled to within 1 second, outputting structured natural language fault diagnosis results.
[0027] The embodiments described above are merely illustrative of implementation methods of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.
Claims
1. A method for multimodal data augmentation and intelligent fault identification of transmission lines under extreme weather conditions, characterized by: S1. Multimodal disaster data collaborative collection, standardized governance, and spatiotemporal alignment: The pre-disaster-post-disaster comparative grouping strategy with a single fault as the smallest governance unit, the five-tuple unique identifier coding, the 17 types of typical fault classification and three-level labeling system, the structured naming and location indexing mechanism, and the local geometric mapping model based on similarity transformation realize the cross-temporal tailoring of pre-disaster-post-disaster defect areas. S2, Depth Data Augmentation Pipeline for Long-Tail Distributions: A unified enhancement operator framework based on Bernoulli random variable control, a combination of multiple geometric transformations and photometric perturbation operators, a physical dehazing enhancement step based on dark channel priors, and channel-based statistical normalization processing. S3. Deployment of intelligent fault identification and reasoning for large visual-language models based on LoRA fine-tuning: This includes targeted LoRA low-rank adaptation for the three main targets of Qwen2.5-VL, a differentiated matrix initialization strategy, a combined loss function of weighted cross-entropy and multimodal consistency, a third-order training pacing, and an inference acceleration deployment scheme integrating Paged Attention and Continuous Batching.
2. The method for multimodal data enhancement and intelligent fault identification of transmission lines under extreme weather conditions according to claim 1, characterized in that: In S3, the LoRA low-rank adaptation and differential matrix initialization strategy includes using high-variance Gaussian initialization to enhance fault feature exploration. And using low-variance Gaussian initialization to protect pre-trained features .
3. The method for multimodal data enhancement and intelligent fault identification of transmission lines under extreme weather conditions according to claim 2, characterized in that: In S3, the three-stage training rhythm includes: a warm-up phase (rounds 1-3) with an initial learning rate of 1 / 10; an optimization phase (rounds 4-15) with a differentiated learning rate and the addition of difficult, high-weight samples; and a convergence phase (rounds 16-20) where the learning rate decays to 1e-6 and early stopping is triggered on the validation set.
4. The method for multimodal data enhancement and intelligent fault identification of transmission lines under extreme weather conditions according to claim 3, characterized in that: In S3, Paged Attention paging KV cache and Continuous Batching iterative dynamic batch processing are integrated into the vLLM inference framework, and 12GB GPU multi-path parallel inference is achieved by combining 8-bit quantization.
5. The method for multimodal data enhancement and intelligent fault identification of transmission lines under extreme weather conditions according to claim 4, characterized in that: In S2, the channel-based statistical normalization at the end of the enhanced pipeline works in conjunction with the Batch Normalization within the network. The channel-based statistical normalization at the end of the enhanced pipeline controls the numerical scale at the input end. The Batch Normalization within the network dynamically calibrates intermediate features, thereby reducing the training instability caused by large enhancements.