Weak texture insulator discharge trace identification method based on large model driving

By using multi-angle drone photography, image feature enhancement, and an improved visual base model, combined with a task transfer plugin and Monte Carlo Dropout technology, the problem of weak texture feature processing in insulator discharge trace recognition was solved, achieving high-precision and robust recognition results, reducing annotation costs, and improving automation levels.

CN121962971APending Publication Date: 2026-05-01STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2025-12-01
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively handle weak texture features when identifying discharge traces on insulators, especially under complex lighting conditions or cluttered backgrounds, which can easily lead to false positives or false negatives. Furthermore, they lack effective methods for estimating prediction uncertainty and enhancing images.

Method used

By using multi-angle drone photography, spatial and frequency domain feature enhancement, improved visual base model, task transfer plugin, and Monte Carlo Dropout technology, combined with a multi-stage post-processing optimization module, low-confidence images are identified and verified, and insulator structure semantic rules are applied for accurate segmentation.

Benefits of technology

It improves the accuracy and robustness of discharge trace identification, reduces the influence of background noise and illumination, achieves high-precision insulator discharge trace identification, reduces labeling costs, and improves the level of automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962971A_ABST
    Figure CN121962971A_ABST
Patent Text Reader

Abstract

The invention discloses a weak texture insulator discharge trace identification method based on large model driving, and the method comprises the steps: carrying out the multi-angle and multi-illumination-condition directional shooting of an insulator through an unmanned plane, and obtaining an insulator image set; and performing spatial domain and frequency domain feature enhancement on each image in the image set to generate a multi-channel feature map. An improved visual basic model is constructed, a task migration plug-in is combined, a multi-channel feature map is processed, a discharge trace segmentation mask is generated, the Monte Carlo Dropout technology is adopted to calculate and predict uncertainty, a low-confidence insulator image is automatically identified, and manual verification is carried out. A verification result is fed back to the task migration plug-in for incremental training, and the model performance is optimized. And completing automatic identification of insulator discharge traces by using the trained improved visual basic model. According to the method, the recognition precision of the weak texture discharge trace can be remarkably improved, manual intervention is reduced, and the method has relatively high engineering practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of insulator detection technology, specifically relating to a method for identifying discharge traces in weakly textured insulators based on a large model-driven approach. Background Technology

[0002] Insulators, as the most critical insulation and mechanical support components in overhead transmission lines and substations, are exposed to the outdoor environment for extended periods. They are susceptible to factors such as flashover, arc discharge, aging, and external contamination, resulting in discharge marks. Discharge marks are a significant indicator of abnormal insulator deterioration. Failure to detect and identify them promptly can lead to decreased insulation performance and, in severe cases, tripping or power outages. Therefore, accurate identification and location of insulator discharge marks is one of the core tasks of power grid operation and maintenance.

[0003] Currently, power line inspection is gradually shifting from manual ground patrols to drone inspections. Drones can acquire high-resolution images from different angles and under different lighting conditions, significantly improving inspection efficiency. However, due to complex shooting environments, strong background interference, weak surface texture of insulators, small target areas, and blurred edges, the existing discharge trace recognition methods based on traditional image processing methods or shallow machine learning methods have limited effectiveness. For example, in the prior art, Chinese patent CN114820567A discloses a deep learning-based method for power line insulator detection and classification. The steps include drone image acquisition, image preprocessing, sample expansion, and construction of a deep learning-based insulator detection model, specifically using a convolutional neural network (CNN) for feature extraction and target detection. This method uses a multi-layer convolutional neural network (CNN) for feature extraction and target detection, combines CBAM (convolutional block attention module) for feature fusion, and finally uses RoI (region of interest) pooling and softmax for classification. However, this method has the following technical problems in practical applications: 1. Discharge traces often have weak texture features, and their shapes are usually irregular with blurred edges. Although existing methods can extract image features to some extent, deep learning models rely on strong texture information in images and still cannot effectively handle the weak texture features of discharge traces. This leads to the possibility of false detections or missed detections under complex lighting conditions or cluttered backgrounds.

[0004] 2. Existing image enhancement techniques mainly focus on multiple processing steps and feature fusion through convolutional layers, failing to fully extract detailed features. Especially for images with strong noise or complex backgrounds, existing methods have limited effectiveness in enhancing low-contrast areas and highlighting minute discharge traces.

[0005] Therefore, there is an urgent need for a large-model-driven insulator discharge trace recognition method that can adapt to complex inspection scenarios, maintain high robustness under weak texture conditions, combine active learning and automatic reasoning capabilities, and make full use of domain prior knowledge, so as to improve recognition accuracy, stability and automation level. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a method for identifying discharge traces in weakly textured insulators based on a large model-driven approach.

[0007] The objective of this invention can be achieved through the following technical solutions: This invention provides a method for identifying discharge traces in weakly textured insulators based on a large model, comprising the following steps: A set of insulator images was obtained by taking directional photos of insulators from multiple angles and under various lighting conditions using drones; Spatial and frequency domain feature enhancement processing is performed on each insulator image in the insulator image set to obtain multi-channel feature maps of each insulator image; An improved visual base model is constructed, and the multi-channel feature maps of each insulator image are processed by the improved visual base model to obtain a discharge trace segmentation mask; the improvement includes embedding a task transfer plugin into the visual base model. During model processing, Monte Carlo Dropout technology is introduced to calculate prediction uncertainty and identify insulator images with low confidence. Low-confidence insulator images and their discharge trace segmentation masks are selected for manual verification, and the verification results are fed back to the task transfer plugin for incremental training. A method for identifying discharge traces in insulators is implemented by training an improved general visual foundation model.

[0008] Furthermore, the step of using a drone to take directional photos of the insulator from multiple angles and under various lighting conditions to obtain an insulator image set specifically includes: The flight path is simulated and optimized using the three-dimensional model of the tower, generating a flight trajectory that can cover different perspectives of the insulator; The flight control system automatically controls the drone's gimbal to acquire high-definition images from multiple perspectives, including top, side, and bottom views. The pre-trained convolutional autoencoder is used to coarsely screen the acquired multi-view high-definition images, calculate the reconstruction error of the images, and exclude conventional background images. For the remaining images, an active learning strategy is used to filter out images with high uncertainty based on Monte Carlo Dropout variance, and the filtered images are used as the insulator image set.

[0009] Furthermore, the step of performing spatial and frequency domain feature enhancement processing on each insulator image in the insulator image set to obtain a multi-channel feature map for each insulator image specifically includes: Spatial enhancement of insulator images is achieved by performing adaptive histogram equalization on the V or L channels to obtain the enhanced image, as shown in the formula: in, This indicates the enhanced image in pixels. Pixel values ​​on; For the original image in pixels Pixel values ​​on; , These represent the minimum and maximum pixel values ​​within the current local region, respectively. , These represent the minimum and maximum pixel values ​​of the enhanced image, respectively, and are set as the contrast enhancement range of the image. The enhanced image is subjected to frequency domain analysis, and the image is decomposed using wavelet transform to obtain multiple frequency domain sub-bands, including four frequency domain sub-bands: LL, LH, HL, and HH. Nonlinear enhancement is applied to multiple frequency domain subbands; Multiple frequency domain subbands after nonlinear enhancement are reconstructed by inverse wavelet transform to obtain the reconstructed and enhanced high-frequency feature map; A saliency detection model is used to generate a saliency map from the high-frequency feature map to highlight potential abnormal regions. The saliency detection model is based on deep learning, including the U2-Net model. The specific channels, high-frequency feature maps, and saliency maps of the original image are stitched together to generate a multi-channel feature map.

[0010] Furthermore, embedding the task transfer plugin in the visual base model specifically involves inserting the task transfer plugin after each multilayer perceptron module of the visual base model via residual connections.

[0011] Furthermore, the task migration plugin includes a settling layer, a dimensionality reduction layer, an activation layer, an attention mechanism, and a dimensionality increase layer connected in sequence.

[0012] Furthermore, the training process for the improved visual base model includes: With the image encoder weights of the visual base model frozen, only the task transfer plugin and mask generator are trained. The multi-channel feature maps corresponding to the acquired insulator image set are input into the improved visual base model for processing to obtain a preliminary prediction mask. The total loss function is calculated based on the preliminary predicted mask and the real labeled mask, and the parameters of the task transfer plugin are updated by gradient. Regularization strategies are applied during training, including random weight averaging and label smoothing; The AdamW optimizer was used during training, and the learning rate was set much lower than that of the pre-trained model.

[0013] Furthermore, the total loss function is expressed as: in, For the total loss function, , , Preset weights; The weighted cross-entropy loss; Indicates Dice loss; This represents the boundary constraint loss.

[0014] Furthermore, in the model processing, the prediction uncertainty is calculated by introducing Monte Carlo Dropout technology to identify insulator images with low confidence, specifically including: For each insulator image, during the forward inference phase of the visual base model and task transfer plugin, a Dropout layer is enabled to preserve randomly dropped neurons. T The next forward sampling is performed to obtain the prediction set for each pixel. ; Calculate the prediction variance for each pixel based on the prediction set for each pixel. : in, Indicates the first t Subsampling of pixels i The predicted value; i Indicates the pixel number; For pixels i The average predicted value; For pixels i The prediction variance; Based on the prediction variance of each pixel, calculate the overall prediction uncertainty of the insulator image: in, This represents the total number of pixels in the insulator image. This indicates the overall prediction uncertainty of the insulator image; when At that time, the insulator image was identified as a low-confidence insulator image; where The set uncertainty threshold.

[0015] Furthermore, the process of selecting low-confidence insulator images and their discharge trace segmentation masks for manual verification, and feeding the verification results back to the task transfer plugin for incremental training, specifically includes: For insulator images identified as having low confidence and their corresponding discharge trace segmentation masks, a manual verification interface is provided, where professionals correct and confirm the pixel annotations of the discharge traces. Generate verification labels from the manually verified mask results. and the original predicted mask One-to-one correspondence, forming a small batch incremental training dataset. ,in This represents a low-confidence image. M Indicates the incremental number of training samples; For the incremental training dataset, input it into the improved visual base model, freeze the encoder weights of the visual base model, and only update the parameters of the task transfer plugin to achieve incremental fine-tuning; During incremental training, the total loss function is used for optimization.

[0016] Furthermore, the method for identifying insulator discharge traces through a trained, improved general visual basic model specifically includes: Acquire images of the insulators to be inspected for discharge trace identification; A pre-trained insulator instance segmentation model is used to detect and segment the insulator image to be detected, and outputs the pixel-level mask, bounding box and class confidence of each detected insulator instance; the instance segmentation model is a model pre-trained on a large power component dataset, including Mask R-CNN or YOLOv8-Seg. For each insulator instance, a bounding box is used as a box cue, and several foreground cue points are generated by sampling within a uniform grid within the pixel-level mask of the insulator instance. The improved visual base model trained by the image of the insulator to be detected, the bounding box, and the foreground cue input is used for forward inference to obtain the initial discharge trace segmentation mask and pixel-level confidence map. Morphological cleansing was performed on the initial segmentation mask; Connectivity analysis is performed on the morphologically processed mask to calculate the appearance attributes of each connected component. The connected components are then filtered based on their appearance attributes. The appearance attributes include area, perimeter, bounding rectangle, and aspect ratio. The retained connected components are further filtered using insulator structure semantic rules. If a connected component intersects with multiple skirt parts of the initial segmentation mask and has a long and thin shape, it is judged as a false detection and removed. The average confidence of the finally retained connected components is calculated based on the pixel-level confidence map, and the optimized binary segmentation mask and the corresponding average confidence are output as the final discharge trace recognition result.

[0017] Compared with the prior art, the present invention has the following advantages: (1) The main problem faced by existing technologies in processing discharge traces is that the surface texture of insulators is weak, and the edges of discharge traces are blurred and the contrast is low, resulting in unsatisfactory recognition results of traditional methods, especially under complex lighting or strong background interference. This invention improves the contrast and detail of the image by performing spatial and frequency domain feature enhancement on the image of the insulator to be detected during the image processing stage, thereby enhancing the features of the weak texture area and making the discharge traces more clearly visible. This technology improves the image quality while effectively reducing background noise and the influence of lighting by performing adaptive histogram equalization and wavelet transform enhancement on the image, thus improving the recognition accuracy.

[0018] (2) In existing technologies, deep learning models often rely on strong texture features, making it difficult to handle subtle, weak texture features such as discharge traces, leading to missed or false detections of discharge traces. This invention optimizes the model's feature extraction capabilities, especially for the recognition of weak texture regions, by constructing an improved visual base model and introducing a task transfer plugin. The task transfer plugin is embedded into the model through residual connections, which, while retaining the powerful expressive capabilities of the base visual model, increases its adaptability to specific features of insulator discharge traces. Through this technical feature, the model can more accurately identify the subtle textures of discharge traces, achieving high-precision segmentation results.

[0019] (3) Existing deep learning methods mostly rely on single image enhancement or simple feature extraction methods, failing to fully consider the spatial structure in the image and the relationship between discharge traces and the insulator body. To overcome this technical problem, this invention proposes to post-process connected components by introducing semantic rules based on insulator structure. This rule further filters out falsely detected connected components by judging the spatial relationship between the morphological features of discharge traces and various components of the insulator (such as skirts, fittings, etc.). Through this technical feature, this invention can effectively remove falsely detected areas, improve segmentation accuracy, and make the discharge trace recognition results more consistent with physical reality.

[0020] (4) Existing technologies lack effective estimation of prediction uncertainty when processing low-confidence samples, making it difficult to detect and correct errors in a timely manner. This invention introduces Monte Carlo Dropout technology to calculate the prediction uncertainty of each pixel, identify low-confidence images, and use them as incremental training data for manual verification and fine-tuning. This technical feature enables the model to dynamically adjust and optimize the recognition effect of low-confidence samples. By introducing incremental training, the robustness and generalization ability of the model are gradually improved, ensuring efficient recognition in complex and dynamic environments.

[0021] (5) Existing technologies mostly rely on traditional target detection and semantic segmentation methods, which are difficult to handle complex situations such as small size and high noise in discharge trace regions. To address these challenges, this invention uses a multi-stage post-processing optimization module to perform morphological purification, connected component analysis, and region attribute-based filtering on the preliminary segmentation results. This technology uses morphological operations to purify the mask, effectively removing isolated noise points and micro-holes, and through connected component analysis, accurately filters out effective discharge trace regions, thereby improving the accuracy and robustness of the recognition results. Attached Figure Description

[0022] Figure 1 This is a flowchart of the insulator discharge trace identification method according to an embodiment of the present invention; Figure 2 This is a general structural diagram of the insulator discharge trace identification method according to an embodiment of the present invention; Figure 3 This is a structural model diagram of the Task Migration Plugin TMP according to an embodiment of the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] Example 1: This embodiment specifically provides a method for identifying discharge traces in weakly textured insulators based on a large model-driven approach, such as... Figure 1 , Figure 2 As shown, it includes the following steps: Step S1: Use a drone to take directional photos of the insulator from multiple angles and under various lighting conditions to obtain an image set of the insulator, specifically including: Flight path simulation and optimization are performed based on the 3D model of the tower to ensure that the UAV gimbal can accurately align with the insulator string and acquire high-definition images from multiple perspectives, including top, side, and bottom views. The flight control system automatically records and embeds high-precision GPS, IMU attitude angle, light intensity, temperature, and humidity multimodal metadata, laying the foundation for data traceability and multi-factor correlation analysis.

[0025] A cascaded screening mechanism based on pre-trained models and active learning employs a two-stage pipeline of "coarse screening" and "fine screening." The coarse screening stage utilizes a convolutional autoencoder pre-trained on a common anomaly dataset to calculate image reconstruction error, quickly filtering out over 95% of normal background images. The fine screening stage introduces an active learning loop, using the prediction uncertainty of the trained model (such as Monte Carlo Dropout variance) as a measure of sample value, prioritizing the labeling of samples where the model is "undecided," maximizing the return on investment of labeling resources.

[0026] The coarse screening module loads a pre-trained CAE model, reconstructs the returned image, and calculates pixel-level MSE loss. It sets a dynamic threshold to automatically move "simple" samples with low reconstruction errors into the cold storage repository.

[0027] The fine-sieving module integrates an active learning strategy, calling a pre-trained segmentation model (which can initially use the uncertainty approximation of a pre-trained CAE during the first cold start, replacing it after the first round of model training) to infer the coarse-screened candidate set and calculate the average prediction uncertainty for each image. Samples are sorted in descending order of uncertainty score, with the top-K% being prioritized for submission to the annotation platform.

[0028] A refined annotation standard and quality control system for weak texture features: A detailed "Guideline for Annotating Weak Texture Discharge Traces" was developed, clearly defining pixel-level definition standards for different morphological traces (dot-like, line-like, and sheet-like), methods for handling blurred boundaries, and spectral and morphological differentiation criteria from common interfering substances such as water droplets, stains, and reflections. A three-tiered quality control system was implemented, consisting of annotator certification, sample cross-validation (IoU>0.85), and final expert arbitration, to ensure annotation consistency.

[0029] Step S2: Perform spatial and frequency domain feature enhancement processing on each insulator image in the insulator image set to obtain multi-channel feature maps for each insulator image, specifically including: Spatial enhancement and color space conversion module: First, convert the RGB image to HSV or Lab color space, and perform a CLAHE operation on the V or L channel to better separate luminance and color information. Then, optional gamma correction can be performed to adjust the overall grayscale distribution.

[0030] Frequency domain analysis and high-frequency feature extraction module: The grayscale image is subjected to discrete wavelet transform to decompose it into multiple sub-bands such as LL, LH, HL, and HH. The coefficients of the LH, HL, and HH sub-bands, which contain detailed information, are nonlinearly enhanced (e.g., using sigmoid stretching), followed by inverse wavelet transform to reconstruct the enhanced high-frequency feature map. Simultaneously, DoG filters can be used in parallel to extract edge features at specific scales.

[0031] Saliency detection and multi-channel fusion module: A saliency detection model based on frequency domain residuals or deep learning (such as U2-Net) is used to generate saliency maps that highlight potential anomalous regions. Finally, specific channels of the original image, the enhanced high-frequency feature map, and the saliency map are concatenated (Stacked) to generate the final multi-channel feature tensor, which serves as the input for subsequent models.

[0032] To address the weak texture and low contrast of discharge traces, this invention employs a multi-layered, feedforward-feedback combined enhancement mechanism for optimization. Before the model, frequency domain feature enhancement and visual cues explicitly amplify the edges and details of the target. Within the model, a TMP (Time-Based Programming) attention mechanism is integrated to guide the model to focus on key feature channels and spatial locations. After the model, post-processing based on structural prior ROI constraints and semantic rules further eliminates background interference. This systematic design ensures that the model maintains stable and high-precision output even when facing complex backgrounds and varying lighting conditions.

[0033] Step S3: Construct an improved visual base model. Process the multi-channel feature maps of each insulator image using the improved visual base model to obtain a discharge trace segmentation mask. The improvement includes embedding a task transfer plugin into the visual base model. This embodiment inserts a pluggable task-specific transfer plug-in (TMP) into the frozen general vision foundation model (GVFM). By training only this plug-in and the mask generator, efficient domain transfer in small-sample scenarios is achieved. TMP works in a "light parameter, residual injection" manner, which retains the general representation of GVFM while injecting specific features for weakly textured targets, thereby controlling the number of labeled samples required to be on the order of tens.

[0034] Existing deep learning methods typically require thousands or even tens of thousands of precisely labeled images to train high-performance models, resulting in high annotation costs. This invention introduces the Parameter Efficient Fine-Tuning (PEFT) paradigm, freezing the general visual representation capabilities of the large GVFM model and training only the lightweight TMP module. This design enables the model to quickly transfer knowledge learned from large amounts of general data to specific domain tasks. Therefore, only about 20 carefully labeled small sample sets are needed to achieve high-precision segmentation of weak-texture discharge traces, reducing the annotation workload by 1-2 orders of magnitude and significantly solving the problem of excessively high annotation costs.

[0035] like Figure 3 As shown, the detailed structure of the TMP is designed as follows: Normalization layer -> Dimensionality reduction layer -> Activation layer -> Attention mechanism -> Dimensionality increase layer. This TMP is then inserted into each multilayer perceptron of the GVFM image encoder using residual connections. It is ensured that the injection points are located on the critical path of feature processing.

[0036] The improved visual base model, the training process includes: With the image encoder weights of the visual base model frozen, only the task transfer plugin and mask generator are trained. The multi-channel feature maps corresponding to the acquired insulator image set are input into the improved visual base model for processing to obtain a preliminary prediction mask. The total loss function is calculated based on the preliminary predicted mask and the real labeled mask, and the parameters of the task transfer plugin are updated by gradient. Regularization strategies are applied during training, including random weight averaging and label smoothing; The AdamW optimizer was used during training, and the learning rate was set much lower than that of the pre-trained model.

[0037] The total loss function is expressed as: in, For the total loss function, , , Preset weights; The weighted cross-entropy loss; Indicates Dice loss; This represents the boundary constraint loss.

[0038] Step S4: During model processing, the prediction uncertainty is calculated by introducing Monte Carlo Dropout technology to identify insulator images with low confidence, specifically including: During model inference, Dropout is enabled and T (e.g., 10) random forward propagations are performed to obtain T prediction results. The prediction variance for each pixel is calculated as cognitive uncertainty and fused with the confidence map of a single prediction (random uncertainty) to obtain a comprehensive uncertainty map. The system sorts the inference results based on the global average of this map, prioritizing high-uncertainty samples to the validation queue.

[0039] For each insulator image, during the forward inference phase of the visual base model and task transfer plugin, a Dropout layer is enabled to preserve randomly dropped neurons. T The next forward sampling is performed to obtain the prediction set for each pixel. ; Calculate the prediction variance for each pixel based on the prediction set for each pixel. : in, Indicates the first t Subsampling of pixels i The predicted value; i Indicates the pixel number; For pixels i The average predicted value; For pixels i The prediction variance; Based on the prediction variance of each pixel, calculate the overall prediction uncertainty of the insulator image: in, This represents the total number of pixels in the insulator image. This indicates the overall prediction uncertainty of the insulator image; when At that time, the insulator image was identified as a low-confidence insulator image; where The set uncertainty threshold.

[0040] Step S5: Select low-confidence insulator images and their discharge trace segmentation masks, perform manual verification, and feed the verification results back to the task transfer plugin for incremental training. Specifically, this includes: An interactive platform developed based on web technology. The verifier interface is divided into: an image display area (overlaying various information), a tool panel (providing annotation tools), and a sample queue list. Annotators can confirm, correct, or reject automatic annotation results. All operations (including modifying trajectories) are recorded and associated with the user ID and timestamp.

[0041] The system monitoring and verification platform automatically initiates the incremental training pipeline when the number of manually corrected samples accumulates to a set quantity (e.g., 15-20 images) or reaches a fixed period. This process loads the current production version model from the model library, performs one round of training on the TMP parameters using only the newly added sample set (keeping the backbone frozen), verifies the performance improvement on the test set after training, and automatically releases the new model version and updates the online inference service once it passes the test.

[0042] Step S6: Implement the insulator discharge trace recognition method using the trained and improved general visual basic model, specifically including: Insulator instance segmentation module: Load a high-precision instance segmentation model (such as MaskR-CNN or YOLOv8-Seg) pre-trained on a large power component dataset. The model takes the original image and outputs the precise pixel-level mask, bounding box, and class confidence for each insulator instance in the image.

[0043] Automatic suggestion generation and model inference module: For each detected insulator instance, its bounding box is used as a box prompt (BoxPrompt). Simultaneously, within its mask region, a set (e.g., 20) of evenly distributed positivepoint prompts (PositivePointPrompts) is generated using uniform grid sampling or a distance-transform-based skeleton point sampling algorithm. These automatically generated prompts, along with the original image, are then fed into a trained GVFM-TMP framework for forward inference to obtain the initial segmentation mask.

[0044] Multi-stage post-processing optimization module: Morphological cleanup: The initial mask is first closed to fill small holes, and then opened to remove tiny isolated noise points.

[0045] Connected component analysis and filtering: Calculate the area, perimeter, circumscribed rectangle, and other properties of all connected regions, and filter out non-significant regions with too small an area (e.g., <50 pixels) or abnormal aspect ratios.

[0046] Semantic rule filtering: Applying domain-specific knowledge rules, for example: if a connected component intersects multiple skirt sections of the insulator mask and its shape is elongated, it is likely a falsely detected shadow or gap and should be removed (this rule can be relaxed in the first iteration and tightened after expert review). The final output is the optimized binary mask and its corresponding average confidence score.

[0047] The native GVFM model requires manual provision of points or bounding boxes as prompts, making full automation impossible. This invention integrates a high-precision insulator instance segmentation model to automatically acquire component-level spatial prior knowledge and generate optimal point and bounding box prompts, completely replacing manual interaction. This allows the system to be seamlessly integrated into the post-processing data workflow of UAV inspections, enabling batch, unattended automatic annotation operations, significantly improving the automation level of the inspection process and its feasibility for engineering implementation.

[0048] This embodiment also includes the following steps: The construction, version management, and knowledge tracing of the annotation database: This step is the core of system results persistence and asset management, aiming to build a unified, structured, and traceable annotation knowledge base. It not only stores the original images and annotation results, but also connects all derived data, model versions, and operation logs through a refined data model, forming a complete digital asset archive. This provides powerful data infrastructure support for historical data queries, model iteration analysis, audit tracing, and potential hazard trend analysis.

[0049] A fully interconnected star schema data model: It adopts a fact table with "detection task" as the core, and is associated with multiple dimension tables such as "original image", "annotation result", "model version" and "operation log" through foreign keys to form an efficient star schema data warehouse structure that supports complex historical data slice and dice and roll-up analysis.

[0050] Full lifecycle version control for data and models: Strict version management is implemented for datasets and model files (e.g., following semantic versioning specifications). Every data update, annotation correction, and model iteration corresponds to a unique version number, supporting data backtracking, model reproduction, and performance comparison for any version, ensuring the reproducibility of the research process.

[0051] Powerful retrieval and statistical analysis capabilities based on metadata: Supports multi-dimensional combined queries and rapid retrieval based on spatiotemporal attributes (line, tower, time range), equipment information, environmental conditions (sunlight, weather), model version, confidence interval, etc. It can automatically generate heat maps of potential hazards and defect rate statistical reports, providing data insights for operation and maintenance decisions.

[0052] Multi-dimensional performance benchmarking and comparative analysis: On the test set reserved in the design, not only segmentation accuracy metrics are evaluated, but also efficiency metrics such as inference speed and memory usage are assessed. Through horizontal comparison with various baseline methods (such as untuned GVFM, traditional CV algorithms, and other state-of-the-art segmentation networks), the technical advancement and engineering practicality of this solution are comprehensively demonstrated with data.

[0053] Containerization and microservice architecture: Docker container technology is used to encapsulate the inference service and all its dependencies into lightweight, self-contained images. Kubernetes is used for container orchestration to achieve rapid service deployment, elastic scaling, high availability, and rolling updates, providing powerful cloud-native features.

[0054] End-to-end real-time monitoring and business insights: The monitoring system not only covers infrastructure metrics (CPU, memory) but also delves into business logic, monitoring the time consumption, confidence distribution, and changes in model uncertainty for each annotation task. By establishing baselines and setting intelligent alerts, performance degradation or data drift can be detected in a timely manner, preventing problems before they occur.

[0055] Example 2: This embodiment also provides a large model-driven system for identifying discharge traces in weakly textured insulators, including: Image acquisition module: Used to acquire images of insulators from different angles and under different lighting conditions using a drone. This module can obtain high-definition images from multiple perspectives, ensuring sufficient image data can be acquired in various environments.

[0056] Image preprocessing module: Performs spatial and frequency domain feature enhancement processing on the acquired insulator images. Through methods such as adaptive histogram equalization and frequency domain wavelet transform, it enhances the texture details and edge information of the image, thereby improving the segmentation accuracy in subsequent recognition processes.

[0057] Target Detection and Segmentation Module: This module, based on a pre-trained insulator instance segmentation model such as Mask R-CNN or YOLOv8-Seg, performs target detection and instance segmentation on the insulator image to be detected, outputting a pixel-level mask, bounding box, and class confidence score for each detected insulator instance. This module provides an initial segmentation mask for discharge traces through bounding box cues and uniform grid sampling.

[0058] The Visual Foundation Model (GVFM) module takes the image to be detected, bounding boxes, and foreground cues as input to the improved GVFM. The model obtains a preliminary discharge trace segmentation mask and corresponding pixel-level confidence maps through forward inference. By introducing a task transfer plugin and Monte Carlo Dropout technology, the model further improves its robustness and accuracy in weakly textured images, calculates prediction uncertainty, and identifies low-confidence images.

[0059] The post-processing module performs morphological cleansing on the initial segmentation mask to remove noise and enhance segmentation edges. Then, connected component analysis is performed to calculate the appearance attributes (such as area and perimeter) of each connected component, filtering based on structural priors to eliminate false positives. Finally, the optimized discharge trace segmentation mask is output by calculating the average confidence level, and the results are further refined using semantic rules.

[0060] Manual verification and incremental training module: For low-confidence images and their discharge trace segmentation masks, the system automatically submits them to the manual verification interface for expert annotation and correction. The corrected results are then fed back as incremental training data to the visual base model for periodic fine-tuning, thereby improving the system's recognition accuracy.

[0061] Results Output and Data Storage Module: The system ultimately outputs a structured segmentation mask and related metadata. All annotation results, images, model versions, and annotation history will be stored in a high-quality database for subsequent hazard analysis, lifetime prediction, and risk assessment.

[0062] Production Environment Monitoring and Alarm Module: In the production environment, the system integrates the Prometheus client and ELKStack for data collection, monitoring key metrics such as inference performance, processing speed, and memory usage in real time. Simultaneously, it utilizes tools like Grafana to create visual monitoring dashboards, displaying various performance data in real time and setting intelligent alarm rules to ensure the system maintains high efficiency and stability during long-term operation.

[0063] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0064] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for identifying discharge traces in weakly textured insulators based on a large model-driven approach, characterized in that, Includes the following steps: A set of insulator images was obtained by taking directional photos of insulators from multiple angles and under various lighting conditions using drones; Spatial and frequency domain feature enhancement processing is performed on each insulator image in the insulator image set to obtain multi-channel feature maps for each insulator image; An improved visual base model is constructed, and the multi-channel feature maps of each insulator image are processed by the improved visual base model to obtain a discharge trace segmentation mask; the improvement includes embedding a task transfer plugin into the visual base model. During model processing, Monte Carlo Dropout technology is introduced to calculate prediction uncertainty and identify insulator images with low confidence. Low-confidence insulator images and their discharge trace segmentation masks are selected for manual verification, and the verification results are fed back to the task transfer plugin for incremental training. A method for identifying discharge traces in insulators is implemented by training an improved general visual foundation model.

2. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The process of obtaining an insulator image set by taking directional photos of the insulator from multiple angles and under various lighting conditions using a drone specifically includes: The flight path is simulated and optimized using the three-dimensional model of the tower, generating a flight trajectory that can cover different perspectives of the insulator; The flight control system automatically controls the drone's gimbal to acquire high-definition images from multiple perspectives, including top, side, and bottom views. The pre-trained convolutional autoencoder is used to coarsely screen the acquired multi-view high-definition images, calculate the reconstruction error of the images, and exclude conventional background images. For the remaining images, an active learning strategy is used to filter out images with high uncertainty based on Monte Carlo Dropout variance, and the filtered images are used as the insulator image set.

3. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The step of performing spatial and frequency domain feature enhancement processing on each insulator image in the insulator image set to obtain a multi-channel feature map for each insulator image specifically includes: Spatial enhancement of insulator images is achieved by performing adaptive histogram equalization on the V or L channels to obtain the enhanced image, as shown in the formula: in, This indicates the enhanced image in pixels. Pixel values ​​on; For the original image in pixels Pixel values ​​on; , These represent the minimum and maximum pixel values ​​within the current local region, respectively. , These represent the minimum and maximum pixel values ​​of the enhanced image, respectively, and are set as the contrast enhancement range of the image. The enhanced image is subjected to frequency domain analysis, and the image is decomposed using wavelet transform to obtain multiple frequency domain sub-bands, including four frequency domain sub-bands: LL, LH, HL, and HH. Nonlinear enhancement is applied to multiple frequency domain subbands; Multiple frequency domain subbands after nonlinear enhancement are reconstructed by inverse wavelet transform to obtain the reconstructed and enhanced high-frequency feature map; A saliency detection model is used to generate a saliency map from the high-frequency feature map to highlight potential abnormal regions. The saliency detection model is based on deep learning, including the U2-Net model. The specific channels, high-frequency feature maps, and saliency maps of the original image are stitched together to generate a multi-channel feature map.

4. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The embedding of the task transfer plugin in the visual base model specifically involves inserting the task transfer plugin after each multilayer perceptron module of the visual base model via residual connections.

5. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The task migration plugin includes a normalization layer, a dimensionality reduction layer, an activation layer, an attention mechanism, and a dimensionality increase layer connected in sequence.

6. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The training process for the improved visual base model includes: With the image encoder weights of the visual base model frozen, only the task transfer plugin and mask generator are trained. The multi-channel feature maps corresponding to the acquired insulator image set are input into the improved visual base model for processing to obtain a preliminary prediction mask. The total loss function is calculated based on the preliminary predicted mask and the real labeled mask, and the parameters of the task transfer plugin are updated by gradient. Regularization strategies are applied during training, including random weight averaging and label smoothing; The AdamW optimizer was used during training, and the learning rate was set much lower than that of the pre-trained model.

7. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 6, characterized in that, The total loss function is expressed as: in, For the total loss function, , , Preset weights; The weighted cross-entropy loss; Indicates Dice loss; This represents the boundary constraint loss.

8. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, During model processing, the Monte Carlo Dropout technique is introduced to calculate prediction uncertainty and identify low-confidence insulator images. Specifically, this includes: For each insulator image, during the forward inference phase of the visual base model and task transfer plugin, a Dropout layer is enabled to preserve randomly dropped neurons. T The next forward sampling is performed to obtain the prediction set for each pixel. ; Calculate the prediction variance for each pixel based on the prediction set for each pixel. : in, Indicates the first t Subsampling of pixels i The predicted value; i Indicates the pixel number; For pixels i The average predicted value; For pixels i The prediction variance; Based on the prediction variance of each pixel, calculate the overall prediction uncertainty of the insulator image: in, This represents the total number of pixels in the insulator image. This indicates the overall prediction uncertainty of the insulator image; when At that time, the insulator image was identified as a low-confidence insulator image; where The set uncertainty threshold.

9. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The process of selecting low-confidence insulator images and their discharge trace segmentation masks for manual verification, and then feeding the verification results back to the task transfer plugin for incremental training, specifically includes: For insulator images identified as having low confidence and their corresponding discharge trace segmentation masks, a manual verification interface is provided, where professionals correct and confirm the pixel annotations of the discharge traces. Generate verification labels from the manually verified mask results. and the original predicted mask One-to-one correspondence, forming a small batch incremental training dataset. ,in This represents a low-confidence image. M Indicates the incremental number of training samples; For the incremental training dataset, input it into the improved visual base model, freeze the encoder weights of the visual base model, and only update the parameters of the task transfer plugin to achieve incremental fine-tuning; During incremental training, the total loss function is used for optimization.

10. The method for identifying discharge traces in weakly textured insulators based on a large model-driven approach according to claim 1, characterized in that, The method for identifying insulator discharge traces through a trained, improved general visual model specifically includes: Acquire images of the insulators to be inspected for discharge trace identification; A pre-trained insulator instance segmentation model is used to detect and segment the insulator image to be detected, and outputs the pixel-level mask, bounding box and class confidence of each detected insulator instance; the instance segmentation model is a model pre-trained on a large power component dataset, including Mask R-CNN or YOLOv8-Seg. For each insulator instance, a bounding box is used as a box cue, and several foreground cue points are generated by sampling within a uniform grid within the pixel-level mask of the insulator instance. The improved visual base model trained by the image of the insulator to be detected, the bounding box, and the foreground cue input is used for forward inference to obtain the initial discharge trace segmentation mask and pixel-level confidence map. Morphological cleansing was performed on the initial segmentation mask; Connectivity analysis is performed on the morphologically processed mask to calculate the appearance attributes of each connected component. The connected components are then filtered based on their appearance attributes. The appearance attributes include area, perimeter, bounding rectangle, and aspect ratio. The retained connected components are further filtered using insulator structure semantic rules. If a connected component intersects with multiple skirt parts of the initial segmentation mask and has a long and thin shape, it is judged as a false detection and removed. The average confidence of the finally retained connected components is calculated based on the pixel-level confidence map, and the optimized binary segmentation mask and the corresponding average confidence are output as the final discharge trace recognition result.

Citation Information

Patent Citations

  • Insulator detection method based on deep learning

    CN114820567A