Distribution network unmanned aerial vehicle image acquisition diagnosis method combining content adaptive mask and visual retrieval enhancement

By combining content-adaptive masking and visual retrieval enhancement, high-quality feature representations are generated and external knowledge is integrated, solving the problems of low feature learning efficiency and insufficient diagnostic capability in existing image diagnosis methods for power distribution network drones, and achieving efficient and reliable intelligent diagnosis.

CN120932133APending Publication Date: 2025-11-11STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510982079.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing image diagnostic methods for power distribution networks rely on manual annotation, which is costly, has low feature learning efficiency, lacks specificity, struggles to learn key details of power equipment, and lacks the ability to conduct in-depth analysis by combining historical data and expert knowledge, resulting in insufficient diagnostic accuracy and reliability.

Method used

We employ a content-adaptive masking and visual retrieval enhancement approach. By training an autoencoder and combining it with a feature deduplication loss term, we generate high-quality feature representations. We then retrieve knowledge related to the image recognition results from a multimodal knowledge base for fusion diagnosis.

Benefits of technology

It improves the automation level and accuracy of image diagnosis for power distribution network UAV inspections, effectively handles complex or rare faults, generates reliable and interpretable diagnostic reports, and significantly improves the accuracy of small target detection and the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932133A_ABST
    Figure CN120932133A_ABST
Patent Text Reader

Abstract

The invention discloses a method for diagnosing images acquired by a distribution network unmanned aerial vehicle in combination with content adaptive mask and visual retrieval enhancement. The method comprises the following steps: acquiring and preprocessing the images acquired by the distribution network unmanned aerial vehicle; extracting edge strength and texture information from the preprocessed images, and calculating an information density map of each image according to the edge strength and the texture information, shielding an area with a specified proportion in each image according to the information density map to obtain a corresponding mask image, inputting the mask image into an auto-encoder, and performing training by using a loss function added with a feature redundancy elimination loss item; constructing a downstream model for image recognition by using an encoder in the trained auto-encoder, preprocessing a new image collected by the distribution network unmanned aerial vehicle, and inputting the preprocessed new image into the trained downstream model to obtain an image recognition result; and retrieving knowledge related to the image recognition result in the multi-modal knowledge base, and fusing the knowledge with the image recognition result to obtain a diagnosis report. The automation level, diagnosis accuracy and reliability of distribution network unmanned aerial vehicle inspection can be comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to artificial intelligence technology, specifically to a method for diagnosing images collected by power distribution drones that combines content-adaptive masking with visual retrieval enhancement. Background Technology

[0002] With the widespread application of power distribution network drones in power line inspection, developing high-efficiency and high-precision intelligent vision models has become a core industry requirement. Traditional supervised learning methods heavily rely on manual annotation, which is costly and time-consuming. Self-supervised learning has emerged to address this, but existing methods often employ random masking strategies, which lack specificity in addressing image content occlusion, making it difficult for models to efficiently learn key details of power equipment (such as insulators and fittings). Furthermore, these methods have relatively singular learning objectives, typically limited to pixel reconstruction, which may lead to redundancy in learned features and affect the robustness of downstream tasks. More importantly, after completing detection, existing models generally lack the ability to combine massive historical data and expert knowledge for in-depth analysis and reliable diagnosis, limiting their application value when dealing with complex or rare faults. Therefore, a novel intelligent diagnostic method is urgently needed to solve the technical challenges of low feature learning efficiency and insufficient diagnostic capabilities of existing methods. Summary of the Invention

[0003] The technical problem to be solved by this invention is to provide a method for diagnosing images collected by distribution network drones that combines content adaptive masking and visual retrieval enhancement, thereby comprehensively improving the automation level, diagnostic accuracy, and reliability of images collected by distribution network drones during inspections.

[0004] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A diagnostic method for images collected by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement, includes the following steps: Acquire images captured by the power distribution network drone and perform preprocessing; Edge intensity and texture information are extracted from the preprocessed image to obtain the edge intensity map and texture information map of each image. The information density map of each image is calculated based on the edge intensity map and texture information map. The corresponding mask image is obtained by occluding a specified proportion of the region in each image based on the information density map. Each mask image is input into the autoencoder and the autoencoder is trained using a loss function with added feature redundancy loss term. The downstream model for image recognition is constructed using the encoder in the trained autoencoder and trained. The new images collected by the distribution network drone are preprocessed and input into the trained downstream model to obtain the image recognition results of the downstream model. Knowledge related to the image recognition result is retrieved from a multimodal knowledge base, and then the image recognition result is fused with the retrieved knowledge to obtain a diagnostic report.

[0005] Furthermore, during preprocessing, the images collected by the distribution network drone are adjusted to the same size and then normalized.

[0006] Furthermore, when calculating the information density map of each image based on the edge intensity map and texture information map of each image, the edge intensity map and texture information map of the same image are weighted and summed to obtain the information density map of each image.

[0007] Furthermore, when occluding a specified proportion of a region in the corresponding image based on the information density map of each image, the specific steps include: The current image is divided into image blocks of a specified size. Based on the position of each image block, the information density of the corresponding block in the information density map of the current image is obtained and used as the information density of each image block. Image blocks are occluded sequentially in descending order of information density until the occluded image blocks reach a specified proportion among all image blocks.

[0008] Furthermore, the expression for the loss function with the added feature redundancy removal loss term is as follows:

[0009] in, It is the mean square error between the reconstructed image and the original image of the mask image. It's a hyperparameter. This is the feature redundancy removal loss term, expressed as follows:

[0010] in, It is the first element in the cross-correlation matrix of the feature representation of the encoder output in an autoencoder. i Line 1 j The elements of the column represent the correlation between different sample features, expressed as follows:

[0011] in, and The first feature representation in the same batch output by the encoder is the second feature representation. b The first sample i The and the first j Normalized representation of each feature dimension.

[0012] Furthermore, in the downstream model, the trained encoder is then connected to a task head that matches the inspection task of the power distribution drone. The task head is either a target detection task head or a semantic segmentation task head, so that the image recognition result of the trained downstream model includes the identified target region and the corresponding inference result.

[0013] Furthermore, in the downstream model, a multi-scale feature fusion module is provided between the trained encoder and the task head. The multi-scale feature fusion module is used to perform weighted fusion of feature maps from different levels of the encoder using weights learned through model training, and input the fused feature map into the task head for object detection or semantic segmentation.

[0014] Furthermore, the multimodal knowledge base includes historical images and corresponding diagnostic text. Retrieving knowledge related to the image recognition result from the multimodal knowledge base, and then fusing the image recognition result with the retrieved knowledge, specifically includes: Extract the features of the target region, then calculate the feature similarity between all historical images and the target region, and select a specified number of historical images in descending order of feature similarity, and obtain the diagnostic text corresponding to the selected historical images. The inference result from the image recognition result is used as the query, and the diagnostic text corresponding to the selected historical image is used as the key and value. A cross-attention mechanism is used to calculate the similarity score between the query and each key. Then, the similarity score is used as the weight to sum all values. Finally, inference is performed on the weighted sum of values ​​and the new inference result is added to the diagnostic report.

[0015] Furthermore, the autoencoder is an asymmetric autoencoder, which uses a VisionTransformer as the encoder and a Transformer with a smaller number of layers and a smaller width as the decoder.

[0016] Furthermore, when acquiring images collected by distribution network drones, specifically, the first set of distribution network drone images is acquired to train the autoencoder, while the second set of distribution network drone images is acquired and labeled to train the downstream model.

[0017] Compared with the prior art, the advantages of the present invention are as follows: This invention achieves content-adaptive masking by occluding a specified proportion of the image based on the information density map of each image. Furthermore, it uses a loss function with added feature redundancy removal loss term during autoencoder training, enabling the pre-trained visual encoder to learn more efficient, robust, and discriminative feature representations. This results in higher accuracy in downstream tasks, with significant improvements, especially in refined tasks such as small object detection.

[0018] When generating a diagnostic report, this invention retrieves knowledge related to the image recognition results of the downstream model from a multimodal knowledge base, and then integrates the image recognition results with the retrieved knowledge, giving the model the ability to reason and diagnose by combining historical knowledge. This enables the model to use external knowledge bases for reasoning, effectively overcoming the uncertainty of traditional deep learning models when dealing with rare and ambiguous cases, and making the diagnostic results more reliable and interpretable. Attached Figure Description

[0019] Figure 1 This is a schematic diagram illustrating the steps of a method according to an embodiment of the present invention.

[0020] Figure 2 This is a detailed flowchart of the adaptive mask self-supervised pre-training steps in an embodiment of the present invention.

[0021] Figure 3 This is a detailed flowchart of the downstream model construction and training steps in an embodiment of the present invention.

[0022] Figure 4 This is a detailed flowchart of the steps for intelligent diagnosis based on visual retrieval enhancement in an embodiment of the present invention. Detailed Implementation

[0023] The present invention will be further described below with reference to the accompanying drawings and specific preferred embodiments, but this does not limit the scope of protection of the present invention.

[0024] This embodiment proposes a diagnostic method for images collected by distribution network drones that combines content-adaptive masking and visual retrieval enhancement. It uses a content-adaptive masking strategy and a feature redundancy removal learning objective to significantly improve the feature representation capability and training efficiency of the visual base model. At the same time, it implements a visual retrieval enhanced generation (Visual RAG) diagnostic framework, which gives the model the ability to reason and diagnose by combining historical knowledge, thereby comprehensively improving the automation level, diagnostic accuracy and reliability of distribution network drone inspections.

[0025] like Figure 1 As shown, the method in this embodiment includes the following steps: S1) Data Acquisition and Preprocessing: Acquire images collected by the distribution network drone and perform preprocessing; S2) Content-adaptive mask self-supervised pre-training: Extract edge intensity and texture information from the preprocessed image to obtain the edge intensity map and texture information map of each image. Calculate the information density map of each image based on the edge intensity map and texture information map. Obscure a specified proportion of the region in each image based on the information density map to obtain the corresponding mask image. Input each mask image into the autoencoder and train the autoencoder using a loss function with added feature redundancy removal loss term. S3) Downstream model construction and training: The downstream model for image recognition is constructed using the encoder in the trained autoencoder and the downstream model is trained. The new images collected by the distribution network drone are preprocessed and input into the trained downstream model to obtain the image recognition results of the downstream model. The image recognition results include the identified target area and the corresponding inference results. S4) Visual retrieval-enhanced intelligent diagnosis: Retrieve knowledge related to the image recognition results from a preset multimodal knowledge base containing historical images and corresponding diagnostic texts, and then fuse the reasoning results in the image recognition results with the retrieved knowledge to obtain a diagnostic report.

[0026] Through the above steps, a complete closed-loop framework from model learning to intelligent diagnosis is realized. This framework has good versatility. The encoder in the pre-trained autoencoder can serve as the basic model for various power grid inspection visual detection tasks of power distribution network drones, while the visual retrieval enhancement generative diagnostic framework can also be easily connected to various industry knowledge bases and continuously expand the content of the knowledge base to realize the continuous learning and capability evolution of the model.

[0027] The following is a detailed explanation of each step.

[0028] In step S1 of this embodiment, when acquiring images collected by the power distribution network drone, specifically, images collected by the power distribution network drone in the first set are acquired to train the autoencoder, and images collected by the power distribution network drone in the second set are acquired and labeled to train the downstream model. The number of images in the first set is much larger than that in the second set to ensure the reliability of the encoder used as the base model.

[0029] In step S1, the preprocessing involves adjusting all the images collected by the distribution network drones to the same size and then normalizing them to unify the numerical distribution of all images, which facilitates subsequent model training.

[0030] Step S2 of this embodiment uses a novel self-supervised pre-training method to obtain an encoder with powerful general feature extraction capabilities. The innovation of this pre-training method is reflected in two aspects: first, it adopts a content-adaptive masking strategy instead of traditional random masks, making the learning task more targeted; second, it introduces a feature deredundancy learning objective instead of a single pixel reconstruction objective, aiming to improve the intrinsic quality of the learned features. Figure 2 As shown, it includes the following steps: S21) Adaptive mask generation: Analyze the content of the input image, combine the edge information and texture information of the image, and dynamically generate a mask for occluding the image; The most information-rich areas in an image are often those with complex structures and textures. By weighted and fused these two types of information, an "information density map" can be generated. When generating a mask, prioritize selecting and masking. Regions with higher values ​​present the model with a more challenging "cloze test" task: not only must it complete the simple background, but it must also infer the complex structures and textures that are being occluded based on the context. This targeted "hard sample learning" strategy, compared to random masks, can more effectively drive the model to learn deeper knowledge about the internal structure and state of power equipment. Based on the above analysis, the adaptive mask generation steps include: The preprocessed image is processed using an edge detection algorithm to extract edge intensity, resulting in an edge intensity map. Simultaneously, a filter bank is used to extract texture information, yielding a texture information map. The edge intensity map and texture information map of the same image are then weighted and summed to obtain the information density map for each image. The formula for the information density map is as follows:

[0031] in, I For the input image data, E(I) This is the edge intensity map obtained through edge detection algorithms. Edge detection algorithms can employ operators such as Canny and Sobel to capture high-frequency structural information in the image, such as tower outlines and conductor edges. T(I) To obtain the texture feature map through the filter bank, a Gabor filter bank can be used. By using filters of different directions and scales, texture information of different patterns in the image, such as dirt on the surface of insulators and corrosion of metal parts, can be captured. and These are preset weighting coefficients for edge information and texture information, respectively. and It can be preset or adaptively adjusted according to the characteristics of the inspection scenario. For example, in scenarios where the main focus is on detecting structural damage, the speed can be appropriately increased. .

[0032] Next, when occluding a specified proportion of a region in the corresponding image based on the information density map of each image, the specific steps include: The current image is divided into image blocks of a specified size. Based on the position of each image block, the information density of the corresponding block in the information density map of the current image is obtained and used as the information density of each image block. Image blocks are occluded sequentially in descending order of information density until the occluded image blocks reach a specified proportion among all image blocks.

[0033] S22) Asymmetric Feature Extraction and Redundancy Removal Learning: An autoencoder with an asymmetric architecture is used to train a visual encoder with image feature extraction capabilities. The encoder in the autoencoder only obtains the unoccluded part of the mask image and extracts feature representations, while the decoder reconstructs the original information of the occluded area in the mask image based on the feature representations output by the encoder.

[0034] Asymmetric architecture is key to achieving efficient self-supervised pre-training. In this embodiment, the autoencoder uses a VisionTransformer (ViT) as the encoder, but only processes unoccluded image blocks (tokens). For example, when the mask ratio is 75%, the encoder only needs to process 25% of the input data. The decoder's task is relatively simple; it only needs to reconstruct the complete image based on a small number of features from the encoder output and mask location information. A lightweight decoder with a simpler structure than the encoder can be used, such as a lightweight Transformer structure with fewer layers and a narrower width than the encoder. This asymmetric design of "heavy encoding, light decoding," compared to the symmetrical encoder and decoder structure in traditional autoencoders, can reduce the computational and time costs of pre-training by several times without sacrificing model performance, making pre-training on large-scale unlabeled network image data possible.

[0035] During the training of the autoencoder, the autoencoder is trained using images from the first preprocessed set, and the model parameters are adjusted using a composite loss function that includes a feature deredundancy loss term. This loss function aims to reduce the correlation between feature representations of different samples. The expression for the loss function with the feature deredundancy loss term is as follows:

[0036] in, This is the mean square error (MSE) between the reconstructed image from the decoder and the original image from the mask image. Its calculation formula is well-known to those skilled in the art and therefore will not be elaborated upon. It's a hyperparameter. This is a feature redundancy removal loss term, a core method for improving feature quality. In deep learning, an ideal feature representation should have different dimensions that are as independent as possible, i.e., "decoupled" or "redundant," with each dimension capturing a unique variation factor in the data. Traditional reconstruction loss can only guarantee that the features contain enough information to reconstruct the original image, but it cannot prevent high correlations between different feature dimensions, i.e., information redundancy. This embodiment introduces... The cross-correlation matrix of the feature representations output by the encoder during training. C The off-diagonal elements (i.e., the correlations between different feature dimensions) are explicitly penalized. Minimizing this loss term drives the optimizer to adjust the model parameters, making the different dimensions of the finally learned feature vectors tend to be orthogonal, thereby obtaining a feature space with higher information density, lower redundancy, and stronger generalization ability. The expression for the feature deredundancy loss term is as follows:

[0037] in, It is the first element in the cross-correlation matrix of the feature representation of the encoder output in an autoencoder. i Line 1 j The elements of the column represent the correlation between different sample features, expressed as follows:

[0038] in, and The first feature representation in the same batch output by the encoder is the second feature representation. b The first sample i The and the first j The normalized representation of the nth feature dimension. In this embodiment, the normalized representation is obtained by standardizing the original feature representation output by the encoder along the batch dimension (subtracting the mean and dividing by the standard deviation). This formula calculates the nth feature dimension of all samples within the batch. i The feature dimension and the first j The Pearson correlation coefficients for each feature dimension are calculated. By calculating the pairwise correlation coefficients between all dimensions, a complete cross-correlation matrix is ​​formed. C The diagonal elements of this matrix Represents the correlation of the feature dimension itself, always equal to 1; off-diagonal elements (When i=j) This measures the degree of linear correlation between different feature dimensions. The feature deduplication loss term operates precisely on these off-diagonal elements.

[0039] In step S3 of this embodiment, application adaptation is performed based on a high-quality encoder to construct a model oriented towards specific tasks, such as... Figure 3As shown, when constructing the downstream model, the encoder weights pre-trained in step S2 are frozen or fine-tuned with a small learning rate. These weights are then connected to a task head matching the inspection task of the power distribution drone, thus constructing a downstream model for a specific inspection task. The task head uses an object detection task head (such as a Faster R-CNN detection head) or a semantic segmentation task head (such as a U-Net structure decoder head), thereby achieving image recognition functionality at the object detection granularity or semantic segmentation granularity of the downstream model. This ensures that the image recognition results of the trained downstream model include the identified target regions and the corresponding inference results. After constructing the downstream model, it is trained using labeled images from the second set.

[0040] In power distribution network inspection scenarios, the size range of targets to be inspected is extremely wide, from macroscopic tower structures to microscopic bolts and pins. To address this challenge, such as... Figure 3 As shown, in the downstream model of this embodiment, a multi-scale feature fusion module is also provided between the trained encoder and the task head. The multi-scale feature fusion module is used to perform weighted fusion of feature maps from different levels of the encoder using weights learned through model training. The fused feature map is then input into the task head for object detection or semantic segmentation. The formula for weighted fusion of feature maps from different levels of the encoder is as follows:

[0041] in, The fused feature map For the first i Feature maps at various scales This is the weight matrix learned through training.

[0042] For feature maps of different depths within the pre-trained encoder Generally, shallow feature maps have high spatial resolution and are rich in detailed information, making them suitable for detecting small targets; deep feature maps have strong semantic information and a large receptive field, making them suitable for detecting large targets. Multi-scale feature fusion modules (e.g., implemented using a Feature Pyramid Network, FPN) effectively combine these advantageous features from different levels through operations such as upsampling, downsampling, and lateral connections. Weight matrix These are learnable parameters, which means that the model can automatically learn during training how to optimally combine information of different scales for a specific task (such as insulator string detection), thereby achieving high-precision detection of targets of various sizes.

[0043] Step S4 of this embodiment introduces a Visual Retrieval Enhancement (Visual RAG) mechanism, enabling the system to transcend the scope of the traditional "perception" model and possess the ability to combine external knowledge for advanced "cognition" and "diagnosis," evolving from a "detector" into an "auxiliary diagnostic expert," thus solving the pain point of insufficient reliability of traditional models when facing difficult and rare cases.

[0044] Specifically, such as Figure 4 As shown, after analyzing the new images acquired by the UAV using the downstream model trained in step S3 and obtaining the image recognition results, knowledge related to the image recognition results is retrieved from the multimodal knowledge base. Then, when the image recognition results are fused with the retrieved knowledge, the specific steps include: S41) Retrieval: Extract the features of the target region from the image recognition results, then calculate the feature similarity between all historical images and the target region, and select a specified number of historical images in descending order of feature similarity, and obtain the diagnostic text corresponding to the selected historical images. S42) Enhancement and Generation: The direct inference results of the downstream model and the retrieved knowledge are fused through a cross-attention mechanism module. Specifically, the inference results in the image recognition results are used as the query, the diagnostic text corresponding to the selected historical image is used as the key and value, the cross-attention mechanism is used to calculate the similarity score between the query and each key, and then the similarity score is used as the weight to perform a weighted summation of all values. Finally, the weighted summation of values ​​is used for inference and the inference result is added to the diagnostic report.

[0045] The cross-attention mechanism is the core of visual retrieval enhancement. When the model's inference result for the current target region (which can be represented as a feature vector) is used as the Query, it represents the model's intent to "want more information about this situation." Knowledge from multiple similar cases retrieved from the multimodal knowledge base (such as historical image features and the embedding vector of diagnostic text) serves as the Key and Value. The cross-attention mechanism calculates the similarity score between the current Query and all Keys, and uses this as weight to perform a weighted summation on the corresponding Values. The physical meaning of this process is that the model can actively and selectively focus on and "absorb" the historical knowledge most relevant to the current situation. For example, if the current situation is suspected to be a rare icing fault, the mechanism will assign higher weights to the retrieved icing cases. Finally, the output, which integrates this weighted knowledge, forms a well-reasoned and highly interpretable enhanced diagnostic report, such as: "Target A was detected (model inference), its features are highly similar to historical case B (icing fault) (retrieval and attention judgment), diagnosed as severe icing, and immediate action is recommended (report generation)."

[0046] The following are two specific examples of applying the methods of this embodiment.

[0047] Example 1: Fine detection and diagnosis of insulator cracks, specifically high-precision detection and intelligent diagnosis of minute cracks in ceramic insulators on distribution network lines.

[0048] Following step S1, a drone equipped with a 45-megapixel camera is used to hover and capture high-resolution images of the insulator strings at a distance of 10-20 meters from the tower. A dataset containing 100,000 unlabeled images of various insulator types is collected as the first set for encoder pre-training, and a dataset containing 2,000 finely labeled images of crack locations is collected as the second set for downstream model training and evaluation. The images are then uniformly cropped and scaled to 224×224 pixels and normalized.

[0049] Following step S2, for each preprocessed insulator image, the Canny operator (with high and low thresholds set to 100 and 200) is used to extract the edge intensity map. E(I) Texture feature maps were extracted using a set of Gabor filters with four orientations (0°, 45°, 90°, 135°). T(I) Set weights =0.6, =0.4, substitute into formula (1) to calculate the information density map. .according to The grayscale values ​​are used to determine the region priority from high to low, and a mask with an occlusion ratio of 75% is generated.

[0050] Then, ViT-Base is used as the encoder in the autoencoder, which processes only the 25% of the image that is not occluded. The decoder in the autoencoder uses a lightweight 4-layer Transformer. The composite loss function is... ,in The mean square error (MSE) between the reconstructed image and the original image is calculated. The hyperparameters are calculated according to formulas (3) and (4). Set to 0.05. Use the AdamW optimizer and pre-train for 300 epochs with unlabeled images from the first set.

[0051] Following step S3, the weights of the pre-trained ViT-Base encoder are frozen or fine-tuned with a small learning rate. A Feature Pyramid Network (FPN) is then connected as a multi-scale feature fusion module, along with a Faster R-CNN detection head to achieve target detection of insulator cracks. The model is trained for 50 epochs using a second set of 2000 labeled crack images to obtain the final crack detection model. The trained model is then deployed in the inspection system to detect insulator images collected by a drone and obtain image recognition results.

[0052] Following step S4, a multimodal knowledge base containing 5000 historical insulator faults is pre-constructed. Each fault includes a fault image, fault type (e.g., "pollution flashover traces," "mechanical stress cracks," "ice-induced cracks"), severity level, and expert diagnostic annotations. Feature vectors of the target region are extracted from the image recognition results. A similarity search is performed on all fault images in the knowledge base, returning the three most similar cases. Then, a cross-attention module is used to receive the model's preliminary results as the query, and the knowledge from the three retrieved cases as the key / value pair. An enhanced diagnostic report is generated after cross-attention calculation.

[0053] For example, the downstream model detects a suspected crack with a confidence level of 0.68, which is considered uncertain. After performing a similarity search on all fault images, the case with the highest similarity (0.92) is tagged "crack caused by ice damage," with expert annotations stating that "the crack is an irregular network, commonly seen during the winter ice melting period." The cross-attention module receives the initial result from the downstream model ("crack, confidence level 0.68") as the query, and the knowledge of the three most similar cases as the key / value pair. The module calculates that the current crack characteristics are highly correlated with the feature key of the "crack caused by ice damage" case. Output: "Diagnostic conclusion: Highly suspected crack caused by ice damage (confidence level increased to 0.95). Diagnostic basis: Irregular network cracks were detected, and their visual characteristics highly match those of historically diagnosed 'crack caused by ice damage' cases. Recommended measures: Increase inspection level and arrange replacement." Insulators, as key components in distribution network lines, are relatively small in size and have diverse defect morphologies. Therefore, the accuracy of their detection is a crucial indicator for evaluating the performance of inspection models. Example 1, through content-adaptive masking and feature redundancy removal learning in step S2, enables the model to learn features more sensitive to fine structures and minute anomalies. Furthermore, multi-scale fusion in step S3 further enhances the detection capability for small targets. Therefore, compared to the baseline method using only traditional random masks, a significant and quantifiable improvement in average accuracy can be achieved. Comparatively, when applied to the detection of insulator targets in distribution network lines, the method of this embodiment improves the average accuracy (AP) by at least 15% compared to the baseline method using a random masking strategy.

[0054] Example 2: Assessment of corrosion degree of transmission towers, specifically, the corrosion status of the angle steel of the transmission towers is segmented and its severity level is assessed in conjunction with a knowledge base.

[0055] Following step 1, use a drone to take images of different parts of the tower during a circumferential flight. Add 50,000 unlabeled tower images to the first set, and add 1,000 images labeled with rust areas (pixel-level) and rust levels (1-4) by experts to the second set. Then perform the same preprocessing operations as in Example 1.

[0056] Following step S2, for each preprocessed tower image, the edge intensity map is extracted using the Canny operator (with height and low thresholds set to 100 and 200). E(I) Texture feature maps were extracted using a set of Gabor filters with four orientations (0°, 45°, 90°, 135°). T(I) Since corrosion is mainly manifested as texture changes, the weights were adjusted to we=0.3, wt=0.7, focusing more on occluding areas with complex textures. The adjusted weights were substituted into formula (1) to calculate the information density map. .according to The grayscale values ​​are used to determine the region priority from high to low, and the mask ratio is still 75%, generating a mask with an occlusion ratio of 75%.

[0057] Then, ViT-Base is used as the encoder in the autoencoder, which processes only the 25% of the image that is not occluded. The decoder in the autoencoder uses a lightweight 4-layer Transformer. The composite loss function is... ,in The mean square error (MSE) between the reconstructed image and the original image is calculated. The hyperparameters are calculated according to formulas (3) and (4). Set to 0.05. Use the AdamW optimizer and pre-train for 300 epochs with unlabeled images from the first set.

[0058] Following step S3, the pre-trained ViT-Base encoder is used as the backbone, followed by a U-Net decoder head for semantic segmentation of the corrosion status of the angle steel of the transmission tower. The downstream model is trained using 1000 labeled corrosion images from the second set. It not only learns to segment the corrosion region but also learns to classify the segmented region (corrosion level 1-4).

[0059] Following step S4, a database containing a large number of images of corroded towers rated according to official standards and corresponding effective cross-sectional analysis data of residual metal is pre-constructed. Feature vectors of target areas are extracted from the image recognition results. Similarity retrieval is performed on all fault images in the knowledge base, returning the five most similar cases. Then, a cross-attention module is used to receive the model's preliminary results as the query, and the five retrieved case knowledge as the key / value pair. An enhanced diagnostic report is generated after cross-attention calculation.

[0060] For example, the downstream model analyzes the base of a certain tower, segments out 15% of the corroded area, and initially classifies it as "Level 3 corrosion," but the model's confidence in this classification is low. Then, features are extracted based on the color distribution, texture characteristics, and coverage percentage of the current corroded area. Following this, a similarity search is performed on all fault images based on the extracted features, retrieving the five most similar historical cases. Of these, four were rated "Level 3 corrosion" and one was rated "Level 4 corrosion." The cross-attention module receives the initial result from the downstream model ("Level 3 corrosion") as the query, and the knowledge of the five most similar cases as the key / value pair. A cross-attention mechanism is then used to calculate and fuse the information from the retrieved cases, particularly the level information of the case most closely resembling the current corrosion texture. Output: "Diagnosis: Confirmed as Grade 3 rust, trending towards Grade 4. Diagnostic basis: The current rust coverage is 15%, with a dark brown, flaky appearance, consistent with the characteristics of multiple 'Grade 3 rust' cases in the knowledge base, and approaching the critical standard for 'Grade 4 rust'. Recommended action: Include in next year's overhaul plan." As can be seen from the two examples above, the method proposed in this embodiment not only has superior performance in specific detection or segmentation tasks, but also, through its unique visual RAG framework, elevates the application value of the model from simple defect discovery to a new level of intelligent, reliable, and interpretable diagnosis that combines historical experience.

[0061] In summary, the image diagnosis method for power distribution network UAVs combining content-adaptive masking and visual retrieval enhancement proposed in this invention first pre-trains the model through content-adaptive masking self-supervised learning. This model fuses edge and texture information of the image when generating the mask and employs an asymmetric architecture and a feature redundancy removal loss function to learn high-quality visual feature representations. In the diagnosis stage, this invention uniquely introduces a visual retrieval enhancement (Visual RAG) mechanism. After the model identifies the target, it automatically retrieves relevant historical cases from a multimodal knowledge base and integrates the retrieved knowledge with the model's own reasoning results to generate an interpretable enhanced diagnostic report. Compared with traditional methods, this invention significantly improves the feature quality of the visual model and the performance of downstream tasks, offering the following advantages: Significantly improved model performance and feature quality: By employing a content-adaptive masking strategy and feature redundancy removal learning objectives, the pre-trained visual encoder of this invention can learn more efficient, robust, and highly discriminative feature representations, thereby achieving higher accuracy in downstream tasks, especially in refined tasks such as small object detection.

[0062] The reliability and intelligence of diagnosis have reached new heights: the original visual RAG diagnostic framework enables the model to use external knowledge bases for reasoning, effectively overcoming the uncertainty of traditional deep learning models when dealing with rare and ambiguous cases, and enabling the system to evolve from a "detector" to an "auxiliary diagnostic expert", with more reliable and interpretable diagnostic results.

[0063] High training efficiency: It adopts an asymmetric encoder-decoder architecture, which processes only a portion of the image data during the pre-training stage, greatly reducing computational overhead and shortening the model training cycle.

[0064] High practicality and scalability: The framework proposed in this invention has good versatility. The pre-trained encoder can be used as the basic model for various power inspection vision tasks, and the RAG diagnostic framework can also be easily connected to the ever-expanding industry knowledge base to realize the continuous learning and capability evolution of the model.

[0065] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

Claims

1. A method for diagnosing images acquired by a UAV in power distribution networks, combining content-adaptive masking and visual retrieval enhancement, characterized in that... Includes the following steps: Acquire images captured by the power distribution network drone and perform preprocessing; Edge intensity and texture information are extracted from the preprocessed image to obtain the edge intensity map and texture information map of each image. The information density map of each image is calculated based on the edge intensity map and texture information map. The corresponding mask image is obtained by occluding a specified proportion of the region in each image based on the information density map. Each mask image is input into the autoencoder and the autoencoder is trained using a loss function with added feature redundancy loss term. The downstream model for image recognition is constructed using the encoder in the trained autoencoder and trained. The new images collected by the distribution network drone are preprocessed and input into the trained downstream model to obtain the image recognition results of the downstream model. Knowledge related to the image recognition result is retrieved from a multimodal knowledge base, and then the image recognition result is fused with the retrieved knowledge to obtain a diagnostic report.

2. The method for diagnostic analysis of images acquired by UAVs for power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... During preprocessing, the images collected by the distribution network drone are adjusted to the same size and then normalized.

3. The method for diagnostic analysis of images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... When calculating the information density map of each image based on its edge intensity map and texture information map, the edge intensity map and texture information map of the same image are weighted and summed to obtain the information density map of each image.

4. The method for diagnosing images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... When occluding a specified proportion of a region in a corresponding image based on the information density map of each image, the specific steps include: The current image is divided into image blocks of a specified size. Based on the position of each image block, the information density of the corresponding block in the information density map of the current image is obtained and used as the information density of each image block. Image blocks are occluded sequentially in descending order of information density until the occluded image blocks reach a specified proportion among all image blocks.

5. The method for diagnostic analysis of images acquired by UAVs for power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... The expression for the loss function with the added feature redundancy removal loss term is as follows: in, It is the mean square error between the reconstructed image and the original image of the mask image. It's a hyperparameter. This is the feature redundancy removal loss term, expressed as follows: in, It is the first element in the cross-correlation matrix of the feature representation of the encoder output in an autoencoder. i Line number j The elements of the column represent the correlation between different sample features, expressed as follows: in, and The first feature representation in the same batch output by the encoder is the second feature representation. b The first sample i The and the first j Normalized representation of each feature dimension.

6. The method for diagnostic analysis of images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... In the downstream model, the trained encoder is then connected to a task head that matches the inspection task of the power distribution drone. The task head is either a target detection task head or a semantic segmentation task head, so that the image recognition result of the trained downstream model includes the identified target region and the corresponding inference result.

7. The method for diagnostic analysis of images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 6, is characterized in that... In the downstream model, a multi-scale feature fusion module is also provided between the trained encoder and the task head. The multi-scale feature fusion module is used to perform weighted fusion of feature maps from different levels of the encoder using weights learned through model training, and input the fused feature map into the task head for object detection or semantic segmentation.

8. The method for diagnosing images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 6, is characterized in that... The multimodal knowledge base includes historical images and corresponding diagnostic texts. Retrieving knowledge related to the image recognition results from the multimodal knowledge base, and then fusing the image recognition results with the retrieved knowledge, specifically includes: Extract the features of the target region, then calculate the feature similarity between all historical images and the target region, and select a specified number of historical images in descending order of feature similarity, and obtain the diagnostic text corresponding to the selected historical images. The inference result from the image recognition result is used as the query, and the diagnostic text corresponding to the selected historical image is used as the key and value. A cross-attention mechanism is used to calculate the similarity score between the query and each key. Then, the similarity score is used as the weight to sum all values. Finally, inference is performed on the weighted sum of values ​​and the new inference result is added to the diagnostic report.

9. The method for diagnostic analysis of images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... The autoencoder is an asymmetric autoencoder, which uses a VisionTransformer as the encoder and a Transformer with a smaller number of layers and a smaller width as the decoder.

10. The method for diagnostic analysis of images acquired by UAVs in power distribution networks, combining content-adaptive masking and visual retrieval enhancement as described in claim 1, is characterized in that... When acquiring images collected by distribution network drones, specifically, the first set of distribution network drone images is acquired to train the autoencoder, while the second set of distribution network drone images is acquired and labeled to train the downstream model.