Method for constructing power grid equipment defect training sample set and defect detection method thereof

CN121903945BActive Publication Date: 2026-09-29安徽明生恒卓科技有限公司 +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511929259.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-09-29
Estimated Expiration
2045-12-19

AI Technical Summary

Technical Problem

本发明的目的在于针对现有电网设备缺陷检测中普遍存在的数据不完备、缺陷样本稀缺、类别分布不平衡以及细粒度检测精度不足等问题,提供一种电网设备缺陷训练样本集的构建方法及其相关的电网设备细粒度缺陷检测方法

Benefits of technology

本发明的目的在于针对现有电网设备缺陷检测中普遍存在的数据不完备、缺陷样本稀缺、类别分布不平衡以及细粒度检测精度不足等问题,提供一种结合扩散模型与RT-DETR架构的生成式电网细粒度缺陷检测方法。通过引入可控生成的扩散模型以增强稀缺缺陷类别样本、结合传统与对抗式生成方法提高样本多样性,解决现有电网设备缺陷检测中普遍存在的数据不完备、缺陷样本稀缺、类别分布不平衡以及细粒度检测精度不足等问题,并利用RT-DETR的高效特征表达与端到端检测能力,解决现有方法在复杂巡检环境下检测性能不稳定、难以识别细微缺陷特征等不足。本发明旨在实现对电网设备缺陷的高精度定位、细粒度分类与多属性识别,并提高模型在不同环境条件下的鲁棒性和实际应用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903945B_ABST
    Figure CN121903945B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of power grid equipment defect training sample set construction method and its defect detection method in the field of computer vision and power inspection.The construction method includes: latent space mapping;Noise and condition injection;Condition denoising;Defect directional generation.By diffusion model being guided, no-defect background and structural information are retained, and defect features consistent with text vector description are generated in specified areas, thereby converting one by one no-defect image into image of known fault state, for constituting power grid equipment defect training sample set.The present application introduces controllable generation diffusion model to enhance scarce defect class samples, combines traditional and adversarial generation method to improve sample diversity, solves the problems of data incompleteness, defect sample scarcity, class distribution imbalance and insufficient fine-grained detection accuracy in existing power grid equipment defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and power line inspection, specifically to a method for constructing a training sample set of power grid equipment defects and a related fine-grained defect detection method for power grid equipment, which is particularly suitable for scenarios where samples are incomplete or defect samples are scarce. Background Technology

[0002] With the widespread deployment of drone inspection technology in power grid operation and maintenance, a large number of equipment images are continuously collected. However, due to the limited types and extremely low frequency of defects in power grid equipment, coupled with the high cost of professional annotation and the difficulty in obtaining real defects, the defect samples exhibit significant incompleteness and imbalance. Traditional deep learning detection models such as YOLO and Faster R-CNN typically rely on large-scale training data with balanced class distribution, thus their performance degrades significantly in scenarios with scarce defect samples. Existing methods face the following challenges: 1) There are few defective samples and many normal samples, resulting in a significant imbalance in data categories.

[0003] 2) The types of defects are complex. There are many fine-grained defects in the same parts, which are difficult to record completely, and the labeling depends on the experience of experts, which is costly.

[0004] 3) Fine-grained defect features are highly homogeneous, making it difficult for traditional visual methods to fully extract detailed features and limiting recognition accuracy.

[0005] To alleviate the problems of incompleteness and class imbalance in power grid defect data, generative modeling methods have become an important technological direction. Generative models learn the conditional probability distribution of samples and can synthesize simulated defect samples with high similarity under given class, scenario, or structural constraints, thereby effectively supplementing the amount of data in scarce categories and improving the generalization ability of detection models.

[0006] In the existing generative technology system, GAN (Generative Adversarial Network), as an early mainstream method, generated images through adversarial training and was widely used in the field of data augmentation. However, GAN is susceptible to instability during training, often suffering from problems such as pattern collapse and gradient vanishing, making it difficult to guarantee the stable generation of fine-grained defective samples. In contrast, diffusion models, as a rapidly developing generative framework in recent years, possess higher reliability and generation quality, with key advantages including: 1) Diffusion models have higher training efficiency, while GANs suffer from mode collapse during training. 2) Diffusion models typically outperform GANs in high-resolution image generation. Diffusion models refine images progressively through multi-step denoising, resulting in richer details. 3) Diffusion models exhibit stronger generalization ability and greater versatility. The training objective of diffusion models is based on probability density estimation, allowing them to better learn the full picture of the data distribution.

[0007] Diffusion models are a class of deep learning models based on probabilistic generation mechanisms. They generate, reconstruct, and simulate original data by progressively perturbing and reconstructing the data distribution. These models have been widely used in image generation tasks in recent years and, due to their stable training characteristics and high generation quality, have gradually become the mainstream alternative to traditional generative adversarial networks.

[0008] The basic idea of ​​the diffusion model is to gradually transform the original sample into near-isotropic Gaussian noise through a pre-defined noise addition process, and then train a neural network to learn the inverse process, thereby gradually recovering the original data from the noise. The diffusion model mainly consists of two parts: a forward diffusion process and a backward denoising process. The forward diffusion process gradually adds Gaussian noise to the input sample according to a fixed or learnable noise scheduling sequence, causing it to gradually degenerate over multiple time steps. This process is usually constructed as a Markov chain, and its noise addition method does not require training and can be directly defined. After a sufficient number of time steps, the sample will degenerate into an approximately standard Gaussian distribution. The backward denoising generation process is modeled by a parameterized deep neural network, which learns to denoise the noisy sample at each time step and approximate the true data distribution. The model achieves backward inference by predicting the noise components or predicting the clean sample at each step, enabling the gradual generation of a clear image from random noise. This process is completed through stepwise sampling, effectively restoring the detailed structure of the image.

[0009] The training objective of diffusion models typically involves regressing the introduced noise using the mean squared error loss function, thereby enabling the model to learn its denoising capabilities at different noise time steps. Because its training does not involve adversarial game processes, diffusion models exhibit high training stability, smooth convergence, and are less prone to mode collapse.

[0010] While existing generative data augmentation methods such as GANs or Diffusion models can expand the sample size to some extent, they fail to combine detection task features with contextual information for targeted generation. Furthermore, the defect performance of power grid equipment is significantly influenced by environmental factors such as weather and lighting, as well as operating conditions such as power load and temperature—indirect attributes that existing detection models generally neglect the fusion of these multimodal information. Summary of the Invention

[0011] (1) Technical problems to be solved The purpose of this invention is to address the common problems in existing power grid equipment defect detection, such as incomplete data, scarce defect samples, unbalanced category distribution, and insufficient fine-grained detection accuracy, by providing a method for constructing a power grid equipment defect training sample set and a related fine-grained defect detection method for power grid equipment.

[0012] (2) Technical solution A method for constructing a training sample set of defects in power grid equipment, comprising: Latent space mapping: Encode and map a defect-free image of a power grid device from the pixel space to the latent space of the diffusion model to obtain its latent representation; Noise and Conditional Injection: A certain amount of noise is added to the latent representation to bring it into the intermediate state of the diffusion process; at the same time, the known defect information is transformed into text vectors as conditional guidance through the CLIP text encoder and combined with the LoRA model. Conditional denoising: The noisy latent table and defective text vectors are input together into the U-Net model combined with the LoRA model. The U-Net model uses a cross-attention mechanism to fuse text guidance information and perform the DDIM sampling inverse denoising process. Targeted Defect Generation: The diffusion model is guided to preserve the background and structural information of defect-free areas, while generating defect features consistent with the text vector description in the specified areas. This transforms defect-free images into images of known fault states, which are then used to construct a training sample set of defects for power grid equipment.

[0013] As a further improvement to the above scheme, the diffusion model is constructed using a controllable diffusion framework based on forward noise addition and reverse noise reduction: Noise scheduling: Gaussian noise is gradually added to the real defect image according to a linear or cosine distribution strategy to generate feature representation sequences of multiple noise levels; this sequence is used as the training input of the diffusion model to learn the potential distribution characteristics of the defect structure of the power grid equipment under different noise states. Region guidance: In the de-diffusion stage, the region guidance method is introduced to use the extracted defect region mask, edge guidance tensor or structural attention weight as conditional input, so that the de-diffusion network of the diffusion model maintains the geometric consistency and texture continuity of the defect morphology during the step-by-step generation process. Temporal coding: Noise stage identifiers are embedded using temporal coding methods. The ability of the diffusion model to model dynamic information of time steps is enhanced by sine-cosine coding or learnable embedding methods, so as to achieve step-by-step controllable target image reconstruction.

[0014] As a further improvement to the above scheme, an adversarial discrimination and generative quality screening method is adopted to filter images from known fault states that only meet the validity criteria to generate samples: A local-global dual-scale discriminant network is used to determine the validity of the generated samples: the overall optical features and the fine-grained texture features of the defect local area are evaluated respectively. Multi-index sample screening rules are set, and a comprehensive quality evaluation function is used for sample screening. The screening rules include: defect texture continuity threshold, edge structure signal-to-noise ratio threshold, target area discernibility threshold, and artifact suppression score. For samples that meet the threshold requirements, label inheritance and structural consistency correction are performed, including automatic copying and correction of defect type labels, location labels, and environmental attribute labels. Only generated samples that meet the validity criteria are dynamically added to the power grid equipment defect training sample set to form an enhanced sample set that can be used to expand the model training.

[0015] As a further improvement to the above scheme, the images of power grid equipment are preprocessed before potential spatial mapping: The original inspection images of power grid equipment are made uniform in image size and spatial structure; at the same time, the images of different inspection batches have a consistent dynamic range; sensor noise and compression artifacts are suppressed to ensure that the structural information of the defect texture area in the image is not destroyed; and then the preprocessed image is generated through random rotation, scale perturbation, light cropping, local occlusion, random noise superposition and brightness and color perturbation operations; the perturbation parameters are dynamically sampled within the set range.

[0016] A second aspect of the present invention also provides a generative power grid fine-grained defect detection method for power grid equipment, which detects fine-grained defects through a fine-grained feature detection model, wherein the samples used for training the fine-grained feature detection model are diverse defect training samples constructed using the above-mentioned method for constructing any power grid equipment defect training sample set.

[0017] As a further improvement to the above scheme, the construction method of the fine-grained feature detection model includes: Backbone network construction: The classic object detection backbone network is used as the basic feature extraction unit. Through multi-scale convolution and feature pyramid structure, deep feature mining is performed on the input power grid equipment image to generate multi-scale features with hierarchical expression capabilities. Multi-scale image feature fusion module design: Taking the multi-scale features output by the backbone network as input, feature complementarity is achieved through a three-level architecture of "high-level attention enhancement - multi-level spatial alignment - semantic collaborative fusion"; Cross-modal fusion module design: to achieve deep fusion of visual features and latent features; Decoder and Header Module Design: The decoder adopts a deformable attention mechanism, which performs sparse interaction based on the filtered Top-K query vectors and multimodal fusion features; Cross-modal fusion design: First, spatial resolution alignment and channel dimension unification are used to match the dimensions of latent features and direct features after multi-scale fusion. Then, vector multiplication is used to achieve fusion of the two. By leveraging the physical rules contained in the latent features, physical constraints are imposed on the fused features to reduce cross-modal semantic bias. The fused feature sequence enters the minimum uncertainty query selection module, which filters high-confidence samples and completes the feature details of low-confidence regions by calculating feature entropy values.

[0018] Furthermore, in the design of the multi-scale image feature fusion module, multi-head self-attention is first applied to the highest-level features to strengthen global semantic association; then, spatial and channel alignment of features at different levels is completed through upsampling, channel unification and RepBlock enhancement; finally, through multi-branch parallel processing, fused features for adaptive classification, regression and confidence prediction are output respectively, taking into account both fine-grained defect details and global contextual information. In the design of the cross-modal fusion module, environmental factors and power load are first standardized and vectorized to match their dimensions with visual fusion features; then, modal interaction is achieved through vector dot product, and semantic bias is reduced by physical rule constraints; finally, a minimum uncertainty query selection module is introduced to filter high-confidence features and complete the details of low-confidence areas. In the design of the decoder and head module, the head module adopts a multi-task parallel design: the classification head is responsible for fine-grained defect category determination, the regression head outputs the defect bounding box coordinates, and the target confidence head evaluates the prediction reliability; through the splicing of the results output by the multi-scale feature branches and channel compression, the integrated output of "category-location-confidence" is finally realized.

[0019] Furthermore, the backbone network adopts the ResNet50_vd architecture as the core of basic feature extraction for power grid equipment images: input layer → Stem layer → residual block stacking → multi-scale feature output; Input layer: Receives RGB images of power grid equipment of the target size, where 3 represents the RGB three-channel dimension; Stem layer: As the network entry point, it uses a 7×7 large convolutional kernel and downsampling with a stride of 2 to quickly compress the image size, increase the number of channels, and generate a preliminary feature map; Residual block stacking: After the Stem layer output, it goes through 4 sets of residual blocks stacked. Each set of residual blocks contains multiple sub-residual units. Each sub-residual unit adopts a serial structure of "convolution-normalization-activation" and combines shortcut connections to achieve feature reuse. Multi-scale feature output: By controlling the step size of different groups of residual blocks, the final output features maps at 3 levels are obtained.

[0020] Furthermore, the fine-grained feature detection model is used for: 1) Defect localization: Based on the RT-DETR regression branch, the defects in the power grid equipment images are spatially located to generate two-dimensional coordinate information, including the defect bounding box and pixel-level mask, so as to accurately identify the location and range of the defect on the equipment surface and provide a clear target area for inspection operations; 2) Defect category identification: The detected defects are classified by classification branches to achieve fine-grained defect classification; 3) Defect attribute output: Outputs attributes related to defects, enabling a comprehensive analysis of defect status; 4) Multimodal fusion and robustness enhancement: joint training of visual feature branches and environmental feature branches; 5) Visualization and application: Visualize and annotate the location, category and attribute information in the inspection image.

[0021] Furthermore, the two-dimensional coordinate information includes the defect bounding box and pixel-level mask; And / or, the type of judgment includes broken wire, rust, loosening, and corrosion categories; And / or, multidimensional attribute information includes defect severity, area, length, and probability of occurrence.

[0022] (3) Beneficial effects The purpose of this invention is to address the common problems in existing power grid equipment defect detection, such as incomplete data, scarce defect samples, imbalanced category distribution, and insufficient fine-grained detection accuracy. It provides a generative fine-grained defect detection method for power grids that combines a diffusion model with an RT-DETR architecture. By introducing a controllable diffusion model to enhance the number of scarce defect category samples and combining traditional and adversarial generation methods to improve sample diversity, this invention solves the problems of incomplete data, scarce defect samples, imbalanced category distribution, and insufficient fine-grained detection accuracy commonly found in existing power grid equipment defect detection. Furthermore, it utilizes the efficient feature representation and end-to-end detection capabilities of RT-DETR to overcome the shortcomings of existing methods, such as unstable detection performance in complex inspection environments and difficulty in identifying subtle defect features. This invention aims to achieve high-precision localization, fine-grained classification, and multi-attribute identification of power grid equipment defects, and to improve the robustness and practical application value of the model under different environmental conditions. Attached Figure Description

[0023] Figure 1 This is a flowchart of the fine-grained defect detection method for power grid equipment according to the present invention.

[0024] Figure 2 This is a schematic diagram of the diffusion model principle used when constructing a training sample set of defects in power grid equipment.

[0025] Figure 3 yes Figure 2 A schematic diagram of the diffusion model.

[0026] Figure 4 This is a schematic diagram of the structure of the fine-grained feature detection model used in the fine-grained defect detection method for power grid equipment.

[0027] Figure 5 for Figure 4 A schematic diagram of the structure of the multi-scale image feature fusion module in the medium-fine granularity feature detection model.

[0028] Figure 6 This is a diagram showing the detection effect of a fine-grained defect detection method for power grid equipment. Detailed Implementation

[0029] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] It should be noted that when a component is said to be "installed on" another component, it can be directly on the other component or it may be in a component that is centered on it. When a component is said to be "set on" another component, it can be directly set on the other component or it may also be in a component that is centered on it. When a component is said to be "fixed to" another component, it can be directly fixed to the other component or it may also be in a component that is centered on it.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "or / and" as used herein includes any and all combinations of one or more of the associated listed items.

[0032] This invention addresses the shortcomings of existing methods, such as unstable detection performance and difficulty in identifying subtle defect features in complex inspection environments, by introducing a controllable diffusion model to enhance scarce defect category samples, combining traditional and adversarial generation methods to improve sample diversity, and utilizing the efficient feature representation and end-to-end detection capabilities of RT-DETR. The invention aims to achieve high-precision location, fine-grained classification, and multi-attribute identification of defects in power grid equipment, and improve the robustness and practical application value of the model under different environmental conditions. Based on the core logic of data completion, feature fusion, and accurate detection, this invention constructs an end-to-end power grid equipment defect detection system. Each module operates collaboratively based on the overall architecture, achieving fine-grained defect identification in scenarios with incomplete data. The overall process and module relationships of the generative fine-grained defect detection method for power grid equipment are as follows: Figure 1 As shown, the process involves first acquiring inspection data using drones, then performing adaptive generative data enhancement on images of defect-free power grid equipment to form a training sample set of power grid equipment defects. This training is used to train a fine-grained feature detection model, and finally, the fine-grained feature detection model is used to achieve multimodal fine-grained defect detection.

[0033] I. Acquisition of UAV Inspection Data The data acquired by drone inspections refers to images of power grid equipment. These raw inspection images are preprocessed as much as possible after acquisition. The acquisition and preprocessing of raw inspection images are as follows.

[0034] Raw image data of power grid equipment is acquired using an unmanned aerial vehicle (UAV) inspection system. This system includes components such as an airborne high-definition camera, attitude stabilization device, and remote control terminal. The acquired raw images may contain quality issues such as blurring due to jitter, uneven lighting, and noise interference. To improve the input quality of subsequent generative enhancement and detection models, preprocessing operations are performed on the images, including: 1) Adjusting images obtained from different acquisition devices and inspection distances to a preset model input resolution to ensure data format consistency and facilitate batch training. 2) Removing random and compressed noise from the images using bilateral filtering, Gaussian smoothing, and non-local means. If necessary, sharpening algorithms are combined to enhance structural edges and improve the recognizability of defect areas. 3) For images with insufficient lighting or uneven exposure, algorithms such as histogram equalization and adaptive contrast limiting are used to enhance local and global contrast, making equipment surface details more prominent. 4) Color shift correction is performed on images significantly affected by weather and time of day to improve style consistency across different inspection batches. 5) Correct edge distortion caused by lens distortion, or automatically crop the region of interest according to the device position to improve the learning efficiency of the generative model and the detection model on the features of the target region.

[0035] The preprocessed images serve as the basic input data for subsequent traditional data augmentation, diffusion model generation enhancement, and RT-DETR detection modules, ensuring image quality and data consistency throughout the entire process.

[0036] II. Constructing a training sample set for power grid equipment defects 1. Traditional Data Augmentation Traditional image enhancement methods such as rotation, cropping, flipping, and brightness adjustment are employed to expand the base sample set. To improve the diversity of the base samples and the model's adaptability to changes in the real inspection environment, limited but crucial traditional enhancement operations are performed on the preprocessed inspection images, including: 1) Viewpoint and structural change enhancement: By randomly rotating the image at a small angle (±5° to ±15°) and performing moderate random cropping, the viewpoint differences that may occur when the UAV photographs the same equipment under slight attitude changes are simulated, enabling the model to adapt to defect morphology under different viewing angles. 2) Illumination condition simulation enhancement: The brightness and contrast of the image are perturbed to maintain visual consistency under different lighting conditions such as dark, strong light, or shadow coverage, thereby improving the model's stability under various inspection conditions such as morning, night, and backlight. 3) Mild noise and blur enhancement: Low-intensity Gaussian noise or slight blur is added to simulate sensor noise and slight jitter that may occur during UAV flight, improving the model's robustness to real-world noisy inspection environments.

[0037] 2 Generative Data Augmentation Diffusion models are a type of probabilistic model that learns the data distribution by progressively denoising variables from a normal distribution. p ( x These models consist of two processes: a forward process and a backward process. The diffusion model can be viewed as a series of equally weighted denoising autoencoders. ε θ ( x t They learn to predict inputs. x t The denoising results are used for training, where x t It is the raw input x The version with added noise. The specific process is shown in Figure 2, which is a schematic diagram of the diffusion model principle.

[0038] Diffusion models, a promising class of generative models based on probability theory in the current field of generative artificial intelligence, are built upon a bidirectional probabilistic process of forward noise addition and inverse noise removal. In the forward process, the model does not introduce noise all at once, but rather adds Gaussian noise gradually and controllably to the real image over multiple discrete time steps using a predefined fixed Markov chain. Ultimately, the features of the original image are completely masked by the noise, transforming it into a pure noise image conforming to a standard normal distribution. This process is essentially a smoothing transformation of the data distribution, laying a clear probabilistic foundation for subsequent inverse learning. The core learning objective of the model is to master the inverse process—that is, starting from completely random noise, gradually removing noise, and ultimately restoring a high-quality image highly consistent with the original data distribution. While diffusion models can generate high-quality images, they often consume a large amount of GPU memory and have slow inference speeds when generating complex images. Stable Diffusion models can effectively solve these problems.

[0039] The Stable Diffusion model first relies on a well-trained autoencoder. The encoder ε compresses the input image into a latent space, performing a diffusion operation within that space, while the decoder D is responsible for restoring the latent representation to the original pixel-space image. This image compression process using the encoder ε is called perceptual compression. Perceptual compression discards high-frequency details in the image, retaining only core and key semantic features, thus significantly reducing the computational cost of the model during training and sampling.

[0040] Furthermore, Stable Diffusion introduces a conditional mechanism to achieve conditional image generation. This mechanism is primarily implemented through a U-Net model with cross-attention layers, denoted as... ε θ ( z t , t , y ).

[0041] The model also uses a domain-specific encoder, specifically the CLIP text encoder, which maps conditional y, such as text or image cues, into intermediate representations. τ θ ( y This representation is then injected into the cross-attention module of U-Net, enabling the model to control the final image generation process based on the input condition (y). Finally, the training objective function of the Stable Diffusion model can be expressed as:

[0042] In this study, focusing on the specific scenario of power grid equipment defect detection, the core value of the diffusion model lies in accurately learning the unique data distribution characteristics of power grid defect images, thereby generating defect image samples with both high fidelity and sample diversity. In actual power grid inspection work, due to the randomness and low probability of equipment defects, and the difficulty in collecting large quantities of some serious defects due to safety risks, various defect image samples are generally scarce and unevenly distributed. This significantly limits the training effect and generalization ability of data-driven defect detection models. The powerful generation capability of the diffusion model provides an effective solution to this industry pain point. By manually generating high-quality defect samples, the scale and coverage of the training dataset can be significantly expanded.

[0043] The training process of the diffusion model strictly follows a two-stage logic of forward noise addition and reverse denoising. Specifically, in the power grid defect image task, it can be further divided into two core stages. In the forward training stage, the model takes collected real power grid defect images, such as typical defect images like insulator damage, conductor strand breakage, and hardware corrosion, as input. Following a pre-defined noise scheduling strategy, it adds Gaussian noise of increasing intensity to the image at each time step. This process generates a series of continuous image sequences with noise levels ranging from low to high. Each intermediate image contains some features of the original defect and the noise information of the current stage. By learning the conditional probability relationships between these images with different noise levels, the model builds a deep understanding of the feature distribution of defect images. In the reverse generation stage, the model needs to use the probabilistic knowledge learned in the forward process to complete the journey from noise to image restoration. At this point, the model's input becomes a completely random noise image. Through iterative denoising operations—that is, predicting and subtracting noise components based on the current image's noise level—the blurred noise image is gradually transformed into a defect image with clear outlines and complete details. Compared with traditional generative models, the progressive denoising mechanism of the diffusion model has significant advantages. It can accurately capture the fine-grained texture of power grid defects, such as the mottled texture of rusted areas, the irregular shape of complex shapes, such as the broken strands of wires, and the sharp edge features, such as the clear boundaries of insulator cracks. This ensures that the generated images not only fully retain the core visual features of the original defects, but also generate diverse samples under different angles and lighting conditions through the randomness of noise, effectively avoiding the overfitting problem caused by the homogeneity of samples.

[0044] Regarding the data augmentation and generation module described above, in this embodiment, the data augmentation and generation module includes a sample preprocessing submodule, a traditional augmentation submodule, a diffusion generation submodule, an adversarial discrimination submodule, and a quality screening and label inheritance submodule. This module is used to automatically construct a high-quality, diverse defect training sample set when there are insufficient defect samples, class imbalance, or insufficient fine-grained feature expression, thereby improving the training stability and generalization ability of the subsequent detection model. Its specific implementation process is as follows.

[0045] I. Sample Preprocessing and Primary Enhancement Procedure 1) Input the original inspection images into the sample preprocessing submodule, and achieve uniformity of image size and spatial structure through resolution standardization and geometric correction; adopt a brightness and color normalization method based on statistical modeling to ensure that images of different inspection batches have a consistent dynamic range.

[0046] 2) Multi-level detail-preserving denoising methods, including Gaussian filtering, nonlocal mean denoising, and guided filtering, are used to suppress sensor noise and compression artifacts, ensuring that the structural information of defective texture regions is not destroyed.

[0047] 3) Primary enhanced samples are generated through operations such as random rotation, scale perturbation, light cropping, local occlusion, random noise superposition, and brightness and color perturbation; the perturbation parameters are dynamically sampled within a set range to ensure that the structural stability of the defect area is maintained under weak perturbation conditions, and to avoid the samples being disturbed to the point of being unrecognizable.

[0048] II. Diffusion Model Training and Data Augmentation In this embodiment, the diffusion generation submodule is constructed based on a controllable diffusion framework of progressive noise addition and reverse denoising, including a noise scheduling unit, a timing coding unit, a region guiding unit, and a defect structure-preserving reverse diffusion network. The specific process is as follows: 1) Using a noise scheduling unit with a linear or cosine distribution strategy, Gaussian noise is progressively added to the real defect image to generate feature representation sequences with multiple noise levels; the noise addition process can be described as follows:

[0049] in x 0 represents a true defect image. x t This is the image representation after the noise step. α t These are the noise scheduling coefficients. This sequence serves as the training input for the diffusion model, used to learn the potential distribution characteristics of the defect structure under different noise conditions.

[0050] 2) In the de-diffusion stage, a region guidance unit is introduced, using the extracted defect region mask, edge guidance tensor, or structural attention weight as conditional inputs to ensure that the de-diffusion network maintains the geometric consistency and texture continuity of the defect morphology during the gradual generation process. Its conditional de-diffusion process can be expressed as follows.

[0051]

[0052] in This represents a parameterized inverse diffusion prediction network.

[0053] 3) By embedding noise stage identifiers into time-series coding units, and enhancing the model's ability to model dynamic information at time steps through sine-cosine coding or learnable embedding methods, it can be represented as:

[0054] This enables progressively controllable reconstruction of the target image. Through the above mechanism, the diffusion generation described in this embodiment can achieve class-conditional generation, structure-preserving generation, and fine-grained texture enhancement generation, and can expand sample diversity while maintaining the authenticity of defects.

[0055] To further achieve efficient, targeted defect image generation that meets practical detection needs, this study abandons the traditional unconstrained generation mode from noise to arbitrary images and innovatively adopts an image-to-image targeted generation strategy. Its core process is as follows: Figure 3 The diagram shown illustrates the diffusion model, also known as the amplification module. This strategy uses a defect-free image of a normal power grid device as a guiding image. By introducing a conditional attention mechanism into the model, it guides the diffusion model to reconstruct the image only in specified areas or according to preset defect types during the generation process. This targeted generation method not only significantly improves generation efficiency and avoids wasting resources by generating irrelevant images, but also precisely controls the type, location, and morphology of defects. This makes the generated samples more closely match the defect distribution characteristics in actual inspection scenarios, providing more targeted and high-quality data support for the training of subsequent defect detection models.

[0056] The key steps are as follows: 1) Latent space mapping: First, a normal power grid equipment image, such as a normal insulator, is encoded from the pixel space and mapped to the latent space of the model to obtain its latent representation.

[0057] 2) Noise and Conditional Injection: Adding a certain amount of noise to this latent representation. z noise This allows it to enter the intermediate state of the diffusion process. Simultaneously, the CLIP text encoder, combined with a LoRA model for fine-tuning, transforms specific defect information, such as discharge contamination, into text vectors. τ θ( y (This serves as a conditional guide.)

[0058] 3) Conditional denoising, which removes noisy latent representations. z noise and defective text vectors τ θ ( y These are input together into the LoRA-tuned U-Net model. U-Net uses a cross-attention mechanism to fuse textual guidance information and performs inverse denoising processes such as DDIM sampling (approximately 100 steps).

[0059] 4) Targeted defect generation: The final model is guided to retain the normal background and structural information, while generating realistic defect features consistent with the text description in the specified area, thereby efficiently transforming a normal image into an image of a specific fault state.

[0060] This method can provide defect samples with greater structural consistency and background diversity, providing rich data support for subsequent deep detection model training. Furthermore, the images generated by the diffusion model can be combined with samples generated by traditional data augmentation methods to further expand the dataset and improve the model's robustness and detection accuracy in scenarios with few samples and incomplete data.

[0061] Therefore, this embodiment provides a method for constructing a training sample set of defects in power grid equipment, which includes: Latent space mapping: Encode and map a defect-free image of a power grid device from the pixel space to the latent space of the diffusion model to obtain its latent representation; Noise and Conditional Injection: A certain amount of noise is added to the latent representation to bring it into the intermediate state of the diffusion process; at the same time, the known defect information is transformed into text vectors as conditional guidance through the CLIP text encoder and combined with the LoRA model. Conditional denoising: The noisy latent table and defective text vectors are input together into the U-Net model combined with the LoRA model. The U-Net model uses a cross-attention mechanism to fuse text guidance information and perform the DDIM sampling inverse denoising process. Targeted Defect Generation: The diffusion model is guided to preserve the background and structural information of defect-free areas, while generating defect features consistent with the text vector description in the specified areas. This transforms defect-free images into images of known fault states, which are then used to construct a training sample set of defects for power grid equipment.

[0062] III. Adversarial Discrimination and Generative Quality Screening Strategies To ensure the usability of the output images from the diffusion model, this embodiment sets up an adversarial discrimination submodule and a quality screening submodule to jointly execute a quality assessment and removal strategy, specifically including: 1) A local-global dual-scale discriminant network is used to determine the validity of generated samples, evaluating both overall optical features and fine-grained texture features of local defects. This can be represented as:

[0063] 2) Set multi-index sample selection rules, including but not limited to defect texture continuity threshold, edge structure signal-to-noise ratio threshold, target area discernibility threshold, and artifact suppression score. A comprehensive quality evaluation function is used for sample selection, which can be expressed as:

[0064] Among them, satisfying S≥ δ.

[0065] 3) Perform label inheritance and structural consistency correction on samples that meet the threshold requirements, including automatic copying and correction of defect type labels, location labels, and environmental attribute labels. The label inheritance process can be represented as:

[0066] Ultimately, only generated samples that meet the validity criteria are dynamically added to the training data pool to form an augmented sample set that can be used for model training.

[0067] 3. Construction of fine-grained feature detection model Please combine Figure 3 This fine-grained feature detection model uses a classic target detection structure as its backbone and is designed for defect detection needs in industrial scenarios such as power grids. It integrates multimodal information and a refined feature processing mechanism to achieve accurate defect identification in scenarios with incomplete data. The overall architecture of the model follows the core logic of "feature extraction - multimodal fusion - fine-grained detection", and is specifically constructed as follows.

[0068] Backbone network construction: The classic object detection backbone network is used as the basic feature extraction unit. Through multi-scale convolution and feature pyramid structure, deep feature mining is performed on the input power grid equipment image to generate visual basic features with hierarchical expression capabilities, providing core feature support for fine-grained defect recognition.

[0069] The multi-scale image feature fusion module is designed as follows: Using the S3, S4, and S5 multi-scale features output from the backbone network as input, feature complementarity is achieved through a three-level architecture of "high-level attention enhancement - multi-level spatial alignment - semantic collaborative fusion". First, multi-head self-attention is applied to the highest-level S5 features to strengthen global semantic association. Then, upsampling, channel unification, and RepBlock enhancement are used to complete the spatial and channel alignment of features at different levels. Finally, through multi-branch parallel processing, fused features for adaptive classification, regression, and confidence prediction are output respectively, taking into account both fine-grained defect details and global contextual information.

[0070] The cross-modal fusion module design focuses on achieving deep fusion of visual and latent features. First, indirect attributes such as environmental factors (lighting, weather) and power load are standardized and vectorized to match their dimensions with the visual fusion features. Then, modal interaction is achieved through vector dot product, leveraging physical rules to reduce semantic bias. Finally, a minimum uncertainty query selection module is introduced to filter high-confidence features and supplement details in low-confidence regions, providing robust input for subsequent detection.

[0071] Decoder and Header Module Design: The decoder employs a deformable attention mechanism, using sparse interaction between the filtered Top-K query vectors and multimodal fusion features to accurately capture spatial correlations of defects and reduce computational costs. The head module adopts a multi-task parallel design: the classification head is responsible for fine-grained defect category determination, the regression head outputs defect bounding box coordinates, and the target confidence head evaluates prediction reliability; through the concatenation and channel compression of the results output by multi-scale feature branches, an integrated output of "category-location-confidence" is finally achieved.

[0072] Cross-modal fusion design: First, spatial resolution alignment and channel dimension unification are used to match the dimensions of latent features (encoded vectors of information such as environmental factors and power load) with the direct features (hierarchical representation of visual features) obtained after multi-scale fusion. Then, vector multiplication is used to fuse the two. By leveraging the physical rules inherent in the latent features (such as the impact of lighting on the visibility of equipment defects and the correlation between load changes and equipment aging), physical constraints are imposed on the fused features to reduce cross-modal semantic bias. The fused feature sequence enters the minimum uncertainty query selection module, which uses feature entropy values ​​to filter high-confidence samples and fill in feature details in low-confidence regions, further reducing the feature loss problem caused by incomplete data and providing more robust input for subsequent detection heads.

[0073] This embodiment proposes a multimodal feature fusion mechanism suitable for fine-grained defect identification of power grid equipment. By synchronously modeling and fusing multi-source heterogeneous data such as inspection images, equipment operating parameters, and environmental status information, it achieves a more comprehensive, reliable, and refined intelligent analysis of the operating status of power grid equipment. Compared with traditional methods that rely solely on visual image features for defect identification, this mechanism effectively overcomes problems such as insufficient image information representation, missing equipment status information, and lack of modeling of complex external interference factors, thereby significantly improving the stability, generalization ability, and interpretability of the identification. In real power grid inspection scenarios, equipment defects are often hidden and cumulative. A single image may not reflect factors such as long-term load, harsh weather, and equipment aging. Therefore, multimodal joint modeling can more accurately characterize potential health risks of equipment, reducing problems such as high false detection rates, uncontrollable false detection rates, and fluctuations in judgment due to environmental sensitivity in traditional methods.

[0074] I. Environmental and Equipment Operation Feature Coding Method External environmental parameters, such as meteorological characteristics like air humidity, temperature, and wind speed, and equipment load operating parameters, such as operating current, operating voltage, and operating time, are first input in structured numerical form into a parameter embedded coding network constructed by a multilayer perceptron (MLP). Let the input vector be: X= [ E•O ],in E Represents an environment parameter vector. O This represents a vector of device operating parameters. Through nonlinear mapping in the MLP, the joint environment-operation feature vector is obtained: F env-op = MLP ( X ) F env-op In order to work with visual feature vectors F vis Alignment, where the joint feature vector undergoes a simple linear transformation or dimension alignment: , making F env-op and Maintaining consistency across feature dimensions ensures that subsequent multimodal fusion can directly perform weighted combination or interactive operations.

[0075] This process ensures that environmental and operational information can be effectively integrated into the detection model, improving the ability to identify fine-grained defects and adaptability to different inspection scenarios.

[0076] The encoding network achieves low-redundancy, high-representational-capability vectorization of the original parameters through nonlinear feature mapping and inter-layer weight learning, thereby obtaining an environment-running joint feature vector with statistical feature compression and semantic representation capabilities. This feature vector is then adjusted to the same semantic space and dimensional scale as the visual feature vector through linear mapping or a dimension alignment function, ensuring additivity, comparability, and correlation at the feature level during subsequent multimodal fusion, thus providing consistent and learnable input conditions for multimodal joint feature fusion.

[0077] II. Network Structure and Multimodal Feature Fusion Strategy (1) Backbone network The backbone network adopts the ResNet50_vd architecture as the core for basic feature extraction of power grid equipment images. This architecture extracts low-, medium-, and high-dimensional visual features from the input image layer by layer through stacked residual blocks and multi-scale downsampling operations, forming a hierarchical feature representation. Its core advantage lies in balancing the depth and efficiency of feature extraction. It alleviates the gradient vanishing problem in deep network training through residual connections, while adapting to a 640×640 input resolution, meeting the dual requirements of real-time performance and fine-grained feature capture in power grid defect detection.

[0078] The overall process of the backbone network is as follows: input image → Stem layer → stacked residual blocks (ResBlock) → multi-scale feature output, with the specific structure as follows: l Input layer: Receives RGB images of power grid equipment with dimensions of 640×640×3, where 3 represents the RGB three-channel dimension.

[0079] Stem layer: As the network entry point, it uses a 7×7 large convolutional kernel and downsampling with a stride of 2 to quickly compress the image size, increase the number of channels, and generate preliminary feature maps. The calculation formula for this layer is:

[0080] in, I For the input image, k The kernel size is [size]. s Step size, p For the number of fillers, Conv (•) BN (•) RELU (•) represent convolution, batch normalization, and activation functions, respectively. This indicates that the operation is executed serially.

[0081] l Residual Block Stacking: After the Stem layer output, four sets of residual blocks (ResBlocks) are stacked, each containing multiple sub-residual units. Each sub-residual unit adopts a sequential structure of "convolution-normalization-activation," combined with shortcut connections to achieve feature reuse, as shown in the following formula:

[0082]

[0083] in, F in Input features for residual units, Conv 1. Conv 2 are 1×1 and 3×3 convolutional layers, used for dimensionality compression and feature extraction, respectively. F res For the residual fusion result, F out This is the final output of the residual unit.

[0084] Multi-scale feature output: By controlling the step size of different groups of residual blocks, three levels of feature maps are finally output, corresponding to different resolutions and receptive fields, to meet the detection requirements of fine-grained defects in the power grid (such as small cracks and loose components) and large-scale defects (such as equipment deformation and large-area damage). The specific output features are as follows: S3 feature map: resolution 80×80 (1 / 8 downsampled from the input image), 256 channels, focusing on fine-grained local features; S4 feature map: resolution 40×40 (1 / 16 downsampled from the input image), 512 channels, balancing local and global features; S5 feature map: resolution 20×20 (1 / 32 downsampled from the input image), 1024 channels, capturing global contextual features.

[0085] (2) Multi-scale image feature fusion module 1) High-level feature attention Please combine Figure 5 The diagram shows a multi-scale image feature fusion module. After the S5 features of the backbone network are flattened, they pass through an adaptive feature interaction fusion module to obtain F5, as shown in the following formula: X flat = Q=K=V=Flatten ( S 5)

[0086] MHSA ( Q , K ,V =Linear(Concat( Attn 1, Attn 2, ..., Attn h )) Where h is the number of attention heads, d model For the attention dimension.

[0087] Input the self-attention results into the residual and normalization layer: X res = X flat + MHSA ( Q , K , V ) F 5= Reshape ( Linear ( SiLU ( Linear ( X res , d ffn )), C out )) in, This is a scaling factor used to mitigate the "curse of dimensionality". yes K i The transpose of the matrix ensures dimension matching during matrix multiplication.

[0088] 2) Multi-level spatial alignment Multi-level spatial alignment is an efficient fusion mechanism for multi-scale feature collaboration. Its core functionality utilizes a CNN structure to achieve semantic complementarity of features at different levels. Its operational logic revolves around scale alignment, channel unification, dual-path enhancement, and fusion output.

[0089] First, regarding the characteristics of high-level... F 5 Channel adjustment and upsampling are performed to adjust its resolution and number of channels to match the mid-layer features. S 4The two are then matched; subsequently, they are concatenated along the channel dimension and processed through two parallel paths—one directly compressing the channels using 1×1 convolutions, and the other enhancing local feature interactions through RepBlocks (reparameterized convolutional blocks); finally, the outputs of the two paths are fused by element-wise addition, preserving the global correlation of high-level semantics while integrating the detailed information of mid-level features, forming a fused feature with both semantic integrity and spatial accuracy, providing more robust feature support for subsequent detection tasks. This mechanism avoids the high computational cost of cross-scale fusion in Transformers, achieving efficient multi-scale feature collaboration through a pure CNN structure, improving feature representation capabilities while ensuring real-time performance. The specific formula is as follows:

[0090]

[0091] in, C For the target number of channels (and) S (4 channels consistent); upsampling makes Resolution from Upgraded to ,and S 4. Alignment.

[0092] right S 4. (Mid-layer features) Adjust the channels to ensure consistency with... Uniform number of channels:

[0093]

[0094] Will and The input is stitched together along the channel dimension to form a fused input:

[0095] Feature alignment is the process of merging features. F concat The features processed by the two parallel paths are then added element-wise to achieve feature information fusion. The formula is as follows:

[0096]

[0097]

[0098] Where C is the target number of channels, and N is the number of repetitions of RepBlock. RepBlock consists of multiple convolutional branches, allowing the network to learn richer feature representations.

[0099] F 5 and S 4. Fusion characteristics obtained after fusion Then and S 3. Fusion characteristics are obtained after fusion. .

[0100] The following will Perform downsampling to obtain downsampled features. After undergoing the same Fusion operation, and Fusion yields downsampled features The formula is as follows:

[0101]

[0102] Similarly, and Fusion .

[0103] Will , and As the output of the multi-level spatial alignment module.

[0104] 3) Multi-level semantic fusion module right The feature maps are sequentially processed through convolutional layers, batch normalization layers, and SiLU activation layers for feature extraction and enhancement. Then, an upsampling layer is applied to improve resolution before the final data is input into the classification head to output the classification prediction result. The formula is as follows:

[0105]

[0106] Feature maps enable parallel processing of multiple tasks through a branching structure: a sub-branch undergoes convolution, normalization, activation, and upsampling before being input into the regression head and outputting the regression prediction result. The other sub-branch, after convolution, normalization, and activation, is input into the target confidence header and outputs the confidence prediction result. Meanwhile, there is also a sub-branch that, after convolution, normalization, activation, and upsampling, is input into the classification head and outputs the classification prediction result. The formula is as follows: l Regression and confidence branch:

[0107]

[0108]

[0109] l Classification branches:

[0110]

[0111] right The feature map is sequentially processed through a convolutional layer, a batch normalization layer, and a SiLU activation layer for feature extraction and enhancement. The subsequent sub-branch is upsampled and input into the regression head to output the regression prediction result. The other sub-branch takes the target confidence header as input and outputs the confidence prediction result. The formula is as follows:

[0112]

[0113]

[0114] Ultimately , , Combining along the channel dimensions and then performing a reshape operation yields a vector. ; , , After the same operation, a vector is obtained. The formula is as follows:

[0115] in, σ It is the Sigmoid activation function. Concatenate (•) represents concatenation along the channel dimension. Conv (•) represents a 1×1 convolutional layer.

[0116] Finally, the two different scale fusion vectors output by Neck are fused together. (dimension) B × C × H A × W A )and (dimension) B × C × H B × W B By concatenating the data along the channel dimension, the fused features are obtained. F concat (dimension) B ×2 C ×H A × W A Then, it is compressed back to its original size using a 1×1 convolution channel. C The formula is as follows:

[0117]

[0118] in, Concatenate (•) represents concatenation along the channel dimension. Conv (•) represents a 1×1 convolutional layer. This represents the dot product operation. F fusion This is the final fused multimodal feature map.

[0119] (3) Design of latent feature extraction module To address the impact of environmental interference (lighting, weather) and equipment operating status (power load) on defect identification in power grid inspection scenarios, a multimodal latent feature fusion module is designed. This module encodes indirect attribute information into structured vectors, deeply fuses them with the visual features output by the backbone network, supplements contextual information, and improves the model's robustness in scenarios with incomplete data.

[0120] Data standardization: Environmental factors (light intensity, weather conditions) and power load are normalized to eliminate dimensional differences, as shown in the following formula:

[0121] in, x For the original potential attributes, μ The mean of the data. σ The standard deviation of the data. X nrom This is the standardized data.

[0122] l Vector encoding: The standardized multidimensional indirect attributes are converted into fixed-dimensional auxiliary feature vectors through a fully connected layer. The encoding formula is as follows:

[0123] in, X nrom The standardized latent attribute feature matrix (dimension 1) B × M , B For batch size, M (Indirect attribute feature dimension). FC (•) is a fully connected layer, and its output dimension is the same as the number of channels in the Neck fusion vector (let's call it C). V latentThe final encoded latent feature vector (dimension 1) B × M ×1×1).

[0124] (4) Cross-modal fusion module It includes a vector fusion module and a minimum uncertainty query selection module. After fusing direct and indirect attribute features, a fused vector is obtained. The fused vector is then fed into the minimum uncertainty query selection module to obtain the target query subset with the lowest prediction uncertainty.

[0125] l Vector fusion module Compressed F compress With quantity V latent The formula for combining dot products is as follows:

[0126] in, This represents the dot product operation. F fusion This is the final fused multimodal feature map.

[0127] l Minimum Uncertainty Query Selection Module By integrating a minimum uncertainty query selection module and a deformable decoder, the fused multimodal features are filtered and accurately analyzed, reducing prediction bias caused by incomplete data and improving the accuracy and reliability of defect detection.

[0128] By calculating the prediction uncertainty at each location in the feature map, the Top-K query vectors with the lowest uncertainty are selected, focusing on high-reliability feature regions, as shown in the following formula:

[0129]

[0130] in, U Uncertainty score for each feature location, Max ( P cls ) represents the maximum classification probability, and k is the number of query vectors (set to 300). TopK (•) generates a set of query vectors by selecting the k feature positions with the smallest uncertainty scores.

[0131] (5) Decoder and Header Module Design The decoder employs a deformable attention mechanism to perform sparse interactions between the query vector and the fused features, accurately capturing the spatial correlation of defective features, as shown in the following formula: Q enhance = DeformableAttention ( Query , F fusion , R , S ) P final = FC ( Q enhance ) in, R As the reference point coordinates, S For sparse sampling locations, DeformableAttention (•) represents a deformable attention function that reduces computation by replacing global attention with sparse sampling. Q enhance For the enhanced query vector, P final This represents the final predicted feature output by the decoder.

[0132] In the header module, P final The data will be fed into the classification head and the regression head: the classification head is responsible for determining the defect category, and the regression head is responsible for predicting the bounding box (or segmentation mask, etc.) of the defect. Finally, the detection results such as "category + location" of the defect are output, completing the closed loop from feature to final prediction.

[0133] III. Training a Fine-Grained Feature Detection Model Model training and loss optimization employ a joint training strategy, simultaneously optimizing both the visual feature branch and the environmental feature branch. This allows the model to extract fine-grained information about target defects while fully integrating environmental contextual information to enhance detection performance. Specifically, the visual feature branch focuses on capturing texture, edges, shapes, and local defect structures in the image, while the environmental feature branch models image background, lighting variations, the surrounding environment of the device, and shooting conditions, providing environmental constraints and semantic supplements for defect detection. Through this branch collaboration, the model can maintain sensitivity to defect targets in complex scenes while reducing the risk of false detections due to environmental interference.

[0134] During training, a multi-task loss function is used for joint optimization. The losses include: 1) Detection loss, which constrains the visual feature branch to improve the accuracy of defect target localization and classification, including bounding box regression loss and class cross-entropy loss. 2) Environmental consistency loss, which constrains the environmental feature branch to learn environmental information complementary to the visual features, ensuring the consistency and stability of feature representation under different lighting, angles, or background conditions. 3) Joint constraint loss, used to coordinate the feature spaces of the two branches, enabling visual and environmental information to jointly optimize the final prediction performance after fusion, avoiding overfitting of a single branch or feature redundancy.

[0135] The entire network achieves joint parameter updates through end-to-end training, with visual and environmental features guiding and fusing each other in the feature space to form a complementary and stable representation. The advantages of this strategy are: 1) In scenarios with few or incomplete samples, environmental features provide additional constraints, reducing overfitting caused by data scarcity.

[0136] 2) The model can maintain high-precision detection under complex lighting, background interference, or changing viewing angle conditions.

[0137] 3) Through joint training, the model learns to automatically weigh the contributions of visual and environmental information in different scenarios, thereby improving its sensitivity to fine-grained defects.

[0138] This joint training strategy provides a solid feature foundation for subsequent generative data augmentation and deep detection models, making the few-sample power grid defect detection system more stable and reliable in real inspection environments. The model training uses a mixed training dataset containing both real and generated samples, and the training process includes a feature learning phase and a joint optimization phase. The model training adopts a phased learning strategy: the first phase focuses on training the detection task by locking onto the visual backbone structure; the second phase introduces multimodal features and completes joint optimization. The optimization objective function includes localization loss, classification loss, and sample consistency constraint loss, used to ensure the model's adaptability and reliability under different defect types and environmental conditions.

[0139] After model training, functional verification is performed using standardized test data, including verification of defect identification accuracy, defect location visualization, and uncertainty sample handling. The model described in this embodiment can be deployed on a ground workstation or server of an inspection drone for online or offline defect identification.

[0140] Fourth, a fine-grained feature detection model is used to achieve multimodal fine-grained defect detection.

[0141] This invention achieves high-precision defect location and fine-grained identification of power grid equipment through the above-mentioned generative data augmentation and joint feature optimization methods: defect location and fine-grained identification output.

[0142] 1) Defect localization module: The module performs spatial localization of defects in power grid equipment images based on RT-DETR regression branch, generates two-dimensional coordinate information, including defect bounding boxes and pixel-level masks, thereby accurately identifying the location and range of defects on the equipment surface and providing a clear target area for inspection operations.

[0143] 2) Defect Category Recognition Module: This module determines the type of detected defects through classification branches, including but not limited to categories such as broken wires, rust, loosening, and corrosion, achieving fine-grained defect classification. During the classification process, contextual information provided by the environmental feature branch is incorporated to enhance the model's recognition accuracy under complex lighting, occlusion, and background interference conditions.

[0144] 3) Defect attribute output module: This module can output multi-dimensional attribute information related to defects, including defect severity, area, length and probability of occurrence, etc., to provide quantitative basis for equipment maintenance and risk assessment, and realize comprehensive analysis of defect status.

[0145] 4) Multimodal fusion and robustness enhancement: Through joint training of visual feature branches and environmental feature branches, the model can maintain detection accuracy and result stability under different lighting, weather, viewing angle and background conditions, thereby improving the applicability of the system in real inspection scenarios.

[0146] 5) Visualization and application module: The model visualizes and annotates the location, category and attribute information on the inspection image, which makes it easy for maintenance personnel to quickly understand the defect status and realize a closed-loop application from detection, analysis to decision-making.

[0147] Through the modular approach described above, this invention achieves precise positioning, fine-grained classification, and multi-attribute recognition of power grid equipment. It fully leverages the advantages of generative data augmentation and joint feature optimization, making it suitable for power grid defect detection and intelligent inspection in environments with limited samples and complex conditions. The model's detection performance is as follows: Figure 6 The example shown is a visualization of the defect detection results.

[0148] This invention combines generative data augmentation with a fine-grained detection model to effectively improve detection performance in situations where power grid equipment defect samples are scarce or imbalanced. Specifically, this invention utilizes a diffusion model and a generative adversarial network to jointly generate defect samples, thus fully expanding the scarce categories, alleviating the dependence of traditional detection models on large-scale labeled data, and improving training efficiency and model generalization ability in scenarios with few samples. Simultaneously, by incorporating multimodal information from environmental factors and power operation parameters and fusing it with visual features, the model achieves robust detection of defects under different lighting conditions, weather conditions, shooting angles, and complex backgrounds, thereby significantly reducing the incidence of false positives and false negatives.

[0149] Furthermore, this invention, based on the direct and indirect attribute prediction mechanism of RT-DETR, can not only accurately locate defects and determine their types, but also output multi-dimensional attribute information of defects, such as area, length, severity, and probability of occurrence, achieving high-precision, fine-grained identification of minute defects in power grid equipment. The collaborative training of generative augmentation and joint feature optimization allows visual and environmental information to complement each other in end-to-end learning, improving the model's stability and robustness in scenarios with few samples and complex environments.

[0150] The overall solution of this invention features a unified structure and a reasonable modular design, enabling data preprocessing, generative augmentation, feature extraction, defect identification, and multi-attribute output within a single framework. It supports end-to-end training and inference, facilitating direct deployment in UAV inspection systems or edge computing devices. Through these technical measures, this invention achieves high-precision location, fine-grained classification, and multi-attribute identification of defects in power grid equipment. It is particularly suitable for intelligent power grid inspection in environments with incomplete data, scarce defect samples, and complex conditions, significantly improving the practicality and reliability of the detection system.

[0151] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0152] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A generative fine-grained defect detection method for power grid equipment, which detects fine-grained defects through a fine-grained feature detection model, characterized in that... Methods for constructing fine-grained feature detection models include: Backbone network construction: The classic object detection backbone network is used as the basic feature extraction unit. Through multi-scale convolution and feature pyramid structure, deep feature mining is performed on the input power grid equipment image to generate multi-scale features with hierarchical expression capabilities. Multi-scale image feature fusion module design: Taking the multi-scale features output by the backbone network as input, feature complementarity is achieved through a three-level architecture of "high-level attention enhancement - multi-level spatial alignment - semantic collaborative fusion"; Cross-modal fusion module design: to achieve deep fusion of visual features and latent features; Decoder and Header Module Design: The decoder adopts a deformable attention mechanism, which performs sparse interaction based on the filtered Top-K query vector and multimodal fusion features; Cross-modal fusion design: First, spatial resolution alignment and channel dimension unification are used to match the dimensions of latent features and direct features after multi-scale fusion. Then, vector multiplication is used to achieve fusion of the two. By leveraging the physical rules contained in the latent features, physical constraints are imposed on the fused features to reduce cross-modal semantic bias. The fused feature sequence enters the minimum uncertainty query selection module, which filters high-confidence samples and completes the feature details of low-confidence regions by calculating feature entropy values.

2. The generative power grid fine-grained defect detection method for power grid equipment according to claim 1, characterized in that, The training sample set used to train the fine-grained feature detection model is obtained through the following steps: Latent space mapping: Encode and map a defect-free image of a power grid device from the pixel space to the latent space of the diffusion model to obtain its latent representation; Noise and Conditional Injection: A certain amount of noise is added to the latent representation to bring it into the intermediate state of the diffusion process; at the same time, the known defect information is transformed into text vectors as conditional guidance through the CLIP text encoder and combined with the LoRA model. Conditional denoising: The noisy latent table and defective text vectors are input into the U-Net model combined with the LoRA model. The U-Net model uses a cross-attention mechanism to fuse text guidance information and perform the DDIM sampling inverse denoising process. Targeted Defect Generation: The diffusion model is guided to preserve the background and structural information of defect-free areas, while generating defect features consistent with the text vector description in the specified areas. This transforms defect-free images into images of known fault states, which are then used to construct a training sample set of defects for power grid equipment.

3. The generative power grid fine-grained defect detection method for power grid equipment according to claim 2, characterized in that, The diffusion model is constructed using a controllable diffusion framework based on forward noise addition and reverse denoising: Noise scheduling: Gaussian noise is gradually added to the real defect image according to a linear or cosine distribution strategy to generate feature representation sequences of multiple noise levels; this sequence is used as the training input of the diffusion model to learn the potential distribution characteristics of the defect structure of the power grid equipment under different noise states. Region guidance: In the de-diffusion stage, the region guidance method is introduced to use the extracted defect region mask, edge guidance tensor or structural attention weight as conditional input, so that the de-diffusion network of the diffusion model maintains the geometric consistency and texture continuity of the defect morphology during the step-by-step generation process. Temporal coding: Noise stage identifiers are embedded using temporal coding methods. The ability of the diffusion model to model dynamic information of time steps is enhanced by sine-cosine coding or learnable embedding methods, so as to achieve step-by-step controllable target image reconstruction.

4. The generative power grid fine-grained defect detection method for power grid equipment according to claim 2, characterized in that, An adversarial discrimination and generative quality screening method is used to select only images that meet the validity criteria from images of known fault states to generate samples: A local-global dual-scale discriminant network is used to determine the validity of the generated samples: the overall optical features and the fine-grained texture features of the defect local area are evaluated respectively. Multi-index sample screening rules are set, and a comprehensive quality evaluation function is used for sample screening. The screening rules include: defect texture continuity threshold, edge structure signal-to-noise ratio threshold, target area discernibility threshold, and artifact suppression score. For samples that meet the threshold requirements, label inheritance and structural consistency correction are performed, including automatic copying and correction of defect type labels, location labels, and environmental attribute labels. Only generated samples that meet the validity criteria are dynamically added to the power grid equipment defect training sample set to form an enhanced sample set that can be used to expand the model training.

5. The generative power grid fine-grained defect detection method for power grid equipment according to claim 2, characterized in that, During latent spatial mapping, the images of power grid equipment are preprocessed first: The original inspection images of power grid equipment are made uniform in image size and spatial structure; at the same time, the images of different inspection batches have a consistent dynamic range; sensor noise and compression artifacts are suppressed to ensure that the structural information of the defect texture area in the image is not destroyed; and then the preprocessed image is generated through random rotation, scale perturbation, light cropping, local occlusion, random noise superposition and brightness and color perturbation operations; the perturbation parameters are dynamically sampled within the set range.

6. The method for generating fine-grained defects in power grid equipment according to any one of claims 2-5, characterized in that, In the design of the multi-scale image feature fusion module, multi-head self-attention is first applied to the highest-level features to strengthen global semantic association; then, spatial and channel alignment of features at different levels is completed through upsampling, channel unification and RepBlock enhancement; finally, through multi-branch parallel processing, fused features for adaptive classification, regression and confidence prediction are output respectively, taking into account fine-grained defect details and global context information. In the design of the cross-modal fusion module, environmental factors and power load are first standardized and vectorized to match their dimensions with visual fusion features; then, modal interaction is achieved through vector dot product, and semantic bias is reduced by physical rule constraints; finally, a minimum uncertainty query selection module is introduced to filter high-confidence features and complete the details of low-confidence areas. In the design of the decoder and head module, the head module adopts a multi-task parallel design: the classification head is responsible for fine-grained defect category determination, the regression head outputs the defect bounding box coordinates, and the target confidence head evaluates the prediction reliability; through the splicing of the results output by the multi-scale feature branches and channel compression, the integrated output of "category-location-confidence" is finally realized.

7. The method for generating fine-grained defects in power grid equipment according to any one of claims 2-5, characterized in that, The backbone network adopts the ResNet50_vd architecture as the core of basic feature extraction for power grid equipment images: input layer → Stem layer → residual block stacking → multi-scale feature output; Input layer: Receives RGB images of power grid equipment of the target size, where 3 represents the RGB three-channel dimension; Stem layer: As the network entry point, it uses a 7×7 large convolutional kernel and downsampling with a stride of 2 to quickly compress the image size, increase the number of channels, and generate a preliminary feature map; Residual block stacking: After the Stem layer output, it goes through 4 sets of residual blocks stacked. Each set of residual blocks contains multiple sub-residual units. Each sub-residual unit adopts a serial structure of "convolution-normalization-activation" and combines shortcut connections to realize feature reuse. Multi-scale feature output: By controlling the step size of different groups of residual blocks, the final output features maps at 3 levels are obtained.

8. The method for generating fine-grained defects in power grid equipment according to any one of claims 2-5, characterized in that, Fine-grained feature detection models are used for: 1) Defect localization: Based on the RT-DETR regression branch, the defects in the power grid equipment images are spatially located to generate two-dimensional coordinate information, including the defect bounding box and pixel-level mask, so as to accurately identify the location and range of the defect on the equipment surface and provide a clear target area for inspection operations; 2) Defect category identification: The detected defects are classified by classification branches to achieve fine-grained defect classification; 3) Defect attribute output: Outputs attributes related to defects, enabling a comprehensive analysis of defect status; 4) Multimodal fusion and robustness enhancement: joint training of visual feature branches and environmental feature branches; 5) Visualization and application: Visualize and annotate the location, category and attribute information in the inspection image.

9. The generative power grid fine-grained defect detection method for power grid equipment according to any one of claims 2-5, characterized in that, Two-dimensional coordinate information includes the defect bounding box and pixel-level mask; And / or, the type of judgment includes broken wire, rust, loosening, and corrosion categories; And / or, multidimensional attribute information includes defect severity, area, length, and probability of occurrence.

Citation Information

Patent Citations

  • Model training method, defect image generation method and related device

    CN117058490A

  • Power equipment defect enhancement and generation method and system based on stable diffusion model

    CN119785143A