Power distribution unmanned aerial vehicle inspection defect intelligent identification method, system, device and medium
Patent Information
- Application Number
- CN202511848903.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2045-12-09
AI Technical Summary
无人机巡检技术因效率高(人均日巡检里程提升至20km)、安全性强(避免高空作业风险),已成为配电巡检的主流手段,但现有的缺陷辨识技术仍存在三大核心痛点:1.样本标注成本高、效率低
针对配电无人机巡检缺陷辨识中标注成本高、推理不可解释、端侧实时性差的核心痛点,本发明通过多模态弱监督对比学习与缺陷语义层次化标签树相结合,实现了少量文本引导下缺陷区域自动分割,采用多模态特征对齐精准定位缺陷区域,生成缺陷区域结构化样本。对缺陷区域结构化样本进行多模态思维链推理,生成可解释的推理链,实现缺陷类型与属性的精准判定,提升推理可解释性与可靠性,推理过程可追溯、可验证。通过跨模态轻量化蒸馏,将多模态思维链推理判定的缺陷类型与属性进行知识迁移,获得轻量化模型后部署到无人机端侧推理引擎,实现实时缺陷辨识推理,满足配电无人机实时巡检需求。经验证,本发明在样本标注上,多模态弱监督对比学习结合缺陷语义层次化标签树,使标注效率提升4.3倍,单缺陷标注成本从30元降至6.9元,10万帧数据标注周期从100天缩短至23天,缺陷区域分割IoU达87.2%,较传统弱监督算法提升了17.2个百分点;在缺陷推理方面,多模态思维链推理实现“分步推理+语义依据”的可视化输出,将巡检人员对结果信任度从65%提升至92%,人工复核率从100%降至10%,使缺陷类型准确率达到95.2%,鸟巢与树枝误判率为3.1%,较传统单模态模型降低了16.9个百分点;端侧部署性能上,跨模态轻量化蒸馏后模型参数量达到7.8M,较教师模型压缩了99.94%,端侧存储占用从50GB降至30MB,端侧推理延迟为112ms,总延迟≤200ms,较云端推理提升57.6%;在巡检应用中,无人机巡检效率达到22km/日,是人工巡检的4.4倍,100杆线路巡检时间从20天缩短至5天,能够实现“机巡替代人巡”,将缺陷漏检率从15%降至4.8%,有效解决了配电无人机巡检缺陷辨识的核心痛点。
Smart Images

Figure CN121746966B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power inspection technology, specifically relating to a method, system, equipment, and medium for intelligent identification of defects in power distribution drone inspections. Background Technology
[0002] As the "last mile" of the power system, the power distribution network's equipment is exposed to the outdoor environment for extended periods, making it prone to typical defects such as bird nests and foreign objects (leading to short circuits), damaged insulators (causing flashover faults), broken conductor strands (causing power outages), and equipment corrosion (reducing insulation performance). According to the "DL / T1575-2016 Technical Guidelines for Unmanned Aerial Vehicle Inspection of Power Distribution Lines," power outages caused by distribution defects account for more than 62% of all distribution network accidents. Unmanned aerial vehicle (UAV) inspection technology has become the mainstream method for power distribution inspection due to its high efficiency (increasing the average daily inspection mileage per person to 20km) and strong safety (avoiding the risks of high-altitude operations). However, existing defect identification technologies still have three major pain points: 1. High cost and low efficiency of sample labeling. The distribution defect scenarios are highly diverse (e.g., differences in bird nest shapes in different seasons, and graded damage levels of insulators). Traditional manual annotation requires professionals to annotate defect areas pixel by pixel, with a single defect annotation taking over 30 minutes and accounting for more than 65% of the total algorithm development cost. Existing weakly supervised annotation algorithms are mostly designed for general scenarios (e.g., ImageNet) and are not adapted to the semantic features of distribution equipment (e.g., specific structures such as "insulator skirts" and "conductor clamps"), resulting in a defect region segmentation IoU (Intersection over Union) of only 68%~72%, which cannot meet the sample quality requirements. 2. Defect inference is uninterpretable and has poor reliability. Existing models often rely on single visual features (such as visible light image texture) and ignore the collaboration of multimodal data such as infrared thermal imaging (abnormal equipment temperature) and lidar (spatial geometric information), resulting in a misclassification rate of ≥20% for bird nests and tree branches, and a missed detection rate of ≥18% for minor insulator damage (area <5%). Furthermore, the model's reasoning process is a "black box," unable to output the logical basis for defect judgment (such as "why this area is judged to be a bird nest rather than a tree branch"), leading to low trust in the results among inspection personnel and requiring manual verification, thus increasing maintenance costs. 3. Insufficient performance and poor real-time capabilities of edge deployment. Multimodal large models (such as ViT-B / 16+GPT-4V) have hundreds of millions to tens of billions of parameters, which require cloud computing power support. However, power distribution scenarios are mostly located in suburban and mountainous areas, where 4G / 5G network coverage is low (only 65%) and fluctuates greatly, resulting in inference latency of over 500ms, which cannot meet the real-time inspection requirements of drones (edge latency requirement ≤200ms). Existing lightweight models mostly only optimize visual features and do not consider the preservation of text inference capabilities. After distillation, the defect identification accuracy drops by 10%~15%, which cannot balance "lightweight" and "high precision". Summary of the Invention
[0003] The purpose of this invention is to address the problems in the prior art by providing a method, system, device, and medium for intelligent identification of defects in power distribution drone inspections, reducing sample labeling costs, improving the interpretability and reliability of reasoning, and enabling real-time deployment at the edge.
[0004] To achieve the above objectives, the present invention provides the following technical solution: Firstly, a method for intelligent defect identification during power distribution drone inspections is provided, including: The multimodal data collected by the power distribution drone is input into a pre-constructed hierarchical label tree for defect semantics, and the defect region is located based on multimodal weak supervised contrastive learning to generate structured samples of the defect region. Multimodal reasoning is performed on structured samples of defective regions to generate interpretable reasoning chains, and the defect type and attributes are determined based on the interpretable reasoning chains. Cross-modal lightweight distillation is performed on the defect type and attribute determination results. The knowledge is then transferred to the lightweight model and deployed to the UAV edge inference engine to achieve real-time defect identification and inference.
[0005] As a preferred embodiment, the defect semantic hierarchical tag tree is a three-level semantic tag tree of "equipment category - defect type - attribute parameter". The first-level tag is the equipment category, which corresponds to the main objects inspected by the power distribution drone; the second-level tag is the defect type, which corresponds to the typical defects associated with each equipment category; and the third-level tag is the attribute parameter, which corresponds to the quantitative attributes of each defect type. Tree structure model This represents a three-level semantic tag tree, where, V Represents a set of nodes. , To represent the first-level node of the equipment category, For secondary nodes representing defect types, A third-level node representing attribute parameters; E Let be the set of edges. Represents a node and The father-son relationship; L For the set of label weights, Representing an edge The semantic association weights are set based on the experience of power distribution experts and the frequency of defect occurrence. The semantic similarity of the three-level semantic tag tree is calculated using cosine similarity. For any two tag nodes... , Semantic similarity S ( Calculate according to the following formula:
[0006] In the formula, For nodes v semantic embedding vector, for and The shortest path, d For path length, Let be the weight of the k-th edge of the path.
[0007] As a preferred embodiment, the step of locating defect regions and generating structured samples of defect regions based on multimodal weakly supervised contrastive learning includes: Multimodal feature extraction from visual, textual, and point cloud data collected by power distribution drones; Through cross-modal contrastive learning, semantic alignment of "visual-text-point cloud" features is achieved, forcing the features of defective regions to be similar to the text description features, and the features of non-defective regions to be significantly different from the text description features; Based on the semantically aligned "visual-text-point cloud" features, defect regions are segmented, and defect regions are labeled using a hierarchical semantic label tree to generate structured samples of defect regions.
[0008] As a preferred embodiment, when extracting multimodal features (visual, textual, and point cloud) from the multimodal data collected by the power distribution drone, the multimodal data collected by the power distribution drone includes visible light images, infrared thermal images, lidar point clouds, and corresponding textual descriptions of defects. For multimodal feature extraction, a lightweight convolutional neural network (CNN) is used to extract multi-scale features from the visible light images, outputting visual feature maps. Infrared images are segmented using a temperature threshold, and infrared feature maps are extracted and fused into joint visual features. For point cloud feature extraction, point clouds are first downsampled using a voxel grid, and spatial geometric features are extracted using a point net (PointNet). Then, region of interest (RoI) pooling is used to map the point cloud features to the image coordinate system. For textual feature extraction, a hierarchical label tree based on defect semantics is used to parse the text description T into a label sequence, and DistilBERT encoding is used to generate textual features.
[0009] As a preferred approach, in the step of achieving semantic alignment of "visual-text-point cloud" features through cross-modal contrastive learning, and forcing defect region features to be similar to text description features while ensuring that non-defect region features are significantly different from text description features, the expression for the loss function is defined as follows:
[0010] In the formula, Visual-text contrast loss is calculated using the following formula to compute joint visual features. Text features Comparative loss:
[0011] In the formula, B Batch size; C This refers to the number of channels in the feature map. Let be the visual feature vector of the b-th sample and the c-th channel; For the first Text features of each sample Temperature coefficient; For point cloud-text contrast loss, the point cloud mapping features are calculated using the following formula. Text features Comparative loss:
[0012] For the visual-point cloud contrast loss, the visual joint features are calculated using the following formula. Point cloud mapping features Comparative loss:
[0013] In the formula, It is an L2 norm.
[0014] As a preferred approach, in the step of segmenting defect regions based on semantically aligned "visual-text-point cloud" features and labeling defect regions using a hierarchical semantic label tree to generate structured samples of defect regions, the visual joint features are calculated. Text features Similarity map; and, calculating point cloud mapping features. Text features The similarity graphs are fused together, and the defect region mask is obtained by Otsu threshold segmentation. Based on the obtained defect region mask, the corresponding labels are labeled for the defect region in combination with the defect semantic hierarchical label tree, and a structured sample library is generated.
[0015] As a preferred embodiment, in the step of performing multimodal reasoning on structured samples of defective regions to generate interpretable reasoning chains, and determining the defect type and attributes based on these interpretable reasoning chains, an attention fusion mechanism is used to analyze visible light images. Feature-level fusion with infrared images is expressed as follows: Among them, attention weight pass Calculations show that the visual-infrared fusion features are... Point cloud features Text label features The embedding matrix of the cross-attention module and the defect semantic label tree is used. , obtain fusion features The calculation expression is: The inference chain probabilistic model is represented as ,in, The initial state is "to be reasoned". to These correspond to equipment identification, defect type determination, and attribute parameter quantification status, respectively. This represents the inference probability from state k-1 to k; in the device identification stage, it is based on fused features. With the first-level node V1 of the tag tree, through the expression Calculate the probability of device category, where, For equipment identification, a weight matrix is used to select the equipment category with the highest probability as the result. In defect type determination, secondary nodes associated with the label tree are filtered based on the equipment identification results, using an expression... Calculate the probability of defect type, where, Weight matrix for determining defect type, For semantic weights; in the attribute parameter quantification stage, based on the defect type results, related third-level nodes are filtered, and attribute parameters are quantified through regression. The severity is based on the proportion of defect areas in visual features. The determination is based on direct quantization of size or temperature using point cloud volume or infrared temperature; the inference chain output integrates the three-step inference results into a natural language form.
[0016] As a preferred approach, the step of performing cross-modal lightweight distillation on the defect type and attribute determination results, transferring knowledge to a lightweight model, and then deploying it to the UAV edge-side inference engine to achieve real-time defect identification and inference involves knowledge distillation enabling the student model to acquire key information from the teacher model. The distillation loss function is a weighted sum of multi-task losses, including feature alignment loss, semantic consistency loss, and task loss. The feature alignment loss uses MSE loss to ensure that the multimodal feature distribution of the student model is consistent with that of the teacher model, allowing the student model to possess the feature extraction capabilities of the teacher model. The semantic consistency loss is based on cosine similarity loss to ensure that the semantics of the inference chain of the student model are consistent with those of the teacher model, ensuring that their inference logic is the same. The task loss is composed of a weighted sum of classification loss and regression loss. The classification loss uses cross-entropy loss to calculate the classification error of the device and defect type, and the regression loss uses L1 loss to calculate the regression error of position, size, and temperature, ensuring the accuracy of the student model in the defect identification task. The distillation weights were determined by validating the power distribution defect dataset. The distilled student model is optimized, including model quantization, which uses INT8 quantization technology to convert the model weights from 32-bit floating-point numbers to 8-bit integers; and operator optimization, which uses the TensorRT framework to optimize the core operators and reduce memory access latency; the optimized model is then converted to ONNX format and deployed to the UAV edge inference engine.
[0017] Secondly, a power distribution drone inspection defect intelligent identification system is provided, including: The defect region localization module is used to input the multimodal data collected by the power distribution drone into a pre-constructed hierarchical label tree for defect semantics, and to locate the defect region based on multimodal weak supervised contrastive learning, generating a structured sample of the defect region. The thought chain reasoning module is used to perform multimodal thought chain reasoning on the structured samples of the defect area, generate an interpretable reasoning chain, and determine the defect type and attributes based on the interpretable reasoning chain. The lightweight distillation module is used to perform cross-modal lightweight distillation on the defect type and attribute determination results. After the knowledge is transferred to the lightweight model, it is deployed to the UAV edge inference engine to realize real-time defect identification and inference.
[0018] As a preferred embodiment, when the defect area localization module inputs the multimodal data collected by the power distribution drone into a pre-constructed hierarchical semantic labeling tree, the hierarchical semantic labeling tree is a three-level semantic labeling tree of "equipment category - defect type - attribute parameter". The first-level label is the equipment category, which corresponds to the main objects inspected by the power distribution drone; the second-level label is the defect type, which corresponds to the typical defects associated with each equipment category; and the third-level label is the attribute parameter, which corresponds to the quantitative attributes of each defect type. Tree structure model This represents a three-level semantic tag tree, where, V Represents a set of nodes. , To represent the first-level node of the equipment category, For secondary nodes representing defect types, A third-level node representing attribute parameters; E Let be the set of edges. Represents a node and The father-son relationship; L For the set of label weights, Representing an edge The semantic association weights are set based on the experience of power distribution experts and the frequency of defect occurrence. The semantic similarity of the three-level semantic tag tree is calculated using cosine similarity. For any two tag nodes... , Semantic similarity S ( Calculate according to the following formula:
[0019] In the formula, For nodes v semantic embedding vector, for and The shortest path, d For path length, Let be the weight of the k-th edge of the path.
[0020] As a preferred embodiment, the defect region localization module locates the defect region based on multimodal weakly supervised contrastive learning. When generating structured samples of the defect region, it extracts multimodal features of vision, text, and point cloud from the multimodal data collected by the power distribution drone. Through cross-modal contrastive learning, semantic alignment of "visual-text-point cloud" features is achieved, forcing the features of defective regions to be similar to the text description features, and the features of non-defective regions to be significantly different from the text description features; Based on the semantically aligned "visual-text-point cloud" features, defect regions are segmented, and defect regions are labeled using a hierarchical semantic label tree to generate structured samples of defect regions.
[0021] As a preferred embodiment, when the defect area localization module extracts multimodal features (visual, textual, and point cloud) from the multimodal data collected by the power distribution drone, the multimodal data includes visible light images, infrared thermal images, lidar point clouds, and corresponding textual descriptions of defects. During multimodal feature extraction, a lightweight convolutional neural network (CNN) is used to extract multi-scale features from the visible light images, outputting a visual feature map. Infrared images are segmented using a temperature threshold, and infrared feature maps are extracted and fused into joint visual features. When extracting point cloud features, the point cloud is first downsampled using a voxel grid, and spatial geometric features are extracted using a point net (PointNet). Then, region-of-interest pooling (RoIPooling) maps the point cloud features to the image coordinate system. Text feature extraction is based on a hierarchical semantic label tree for defects, parsing the text description T into a label sequence and generating text features using DistilBERT encoding.
[0022] As a preferred embodiment, the defect region localization module achieves semantic alignment of "visual-text-point cloud" features through cross-modal contrastive learning, forcing the defect region features to be similar to the text description features. When the non-defect region features differ significantly from the text description features, the expression for the loss function is defined as follows:
[0023] In the formula, Visual-text contrast loss is calculated using the following formula to compute joint visual features. Text features Comparative loss:
[0024] In the formula, B Batch size; C This refers to the number of channels in the feature map. Let be the visual feature vector of the b-th sample and the c-th channel; For the first Text features of each sample Temperature coefficient; For point cloud-text contrast loss, the point cloud mapping features are calculated using the following formula. Text features Comparative loss:
[0025] For the visual-point cloud contrast loss, the visual joint features are calculated using the following formula. Point cloud mapping features Comparative loss:
[0026] In the formula, It is an L2 norm.
[0027] As a preferred embodiment, the defect region localization module performs defect region segmentation based on semantically aligned "visual-text-point cloud" features, and combines a hierarchical semantic label tree to annotate the defect regions. When generating structured samples of defect regions, it calculates visual joint features. Text features Similarity map; and, calculating point cloud mapping features. Text features The similarity graphs are fused together, and the defect region mask is obtained by Otsu threshold segmentation. Based on the obtained defect region mask, the corresponding labels are labeled for the defect region in combination with the defect semantic hierarchical label tree, and a structured sample library is generated.
[0028] As a preferred embodiment, the thought chain reasoning module utilizes an attention fusion mechanism to analyze visible light images. Feature-level fusion with infrared images is expressed as follows: Among them, attention weight pass Calculations show that the visual-infrared fusion features are... Point cloud features Text label features The embedding matrix of the cross-attention module and the defect semantic label tree is used. , obtain fusion features The calculation expression is: The inference chain probabilistic model is represented as ,in, The initial state is "to be reasoned". to These correspond to equipment identification, defect type determination, and attribute parameter quantification status, respectively. This represents the inference probability from state k-1 to k; in the device identification stage, it is based on fused features. With the first-level node V1 of the tag tree, through the expression Calculate the probability of device category, where, For equipment identification, a weight matrix is used to select the equipment category with the highest probability as the result. In defect type determination, secondary nodes associated with the label tree are filtered based on the equipment identification results, using an expression... Calculate the probability of defect type, where, Weight matrix for determining defect type, For semantic weights; in the attribute parameter quantification stage, based on the defect type results, related third-level nodes are filtered, and attribute parameters are quantified through regression. The severity is based on the proportion of defect areas in visual features. The determination is based on direct quantization of size or temperature using point cloud volume or infrared temperature; the inference chain output integrates the three-step inference results into a natural language form.
[0029] As a preferred embodiment, the lightweight distillation module enables the student model to acquire key information from the teacher model through knowledge distillation. The distillation loss function is a weighted sum of multi-task losses, including feature alignment loss, semantic consistency loss, and task loss. The feature alignment loss uses MSE loss to ensure that the multimodal feature distribution of the student model is consistent with that of the teacher model, enabling the student model to have the feature extraction capabilities of the teacher model. The semantic consistency loss is based on cosine similarity loss to ensure that the semantics of the inference chain of the student model are consistent with those of the teacher model, ensuring that the inference logic of both is the same. The task loss is composed of a weighted sum of classification loss and regression loss. The classification loss uses cross-entropy loss to calculate the classification error of the device and defect type, and the regression loss uses L1 loss to calculate the regression error of position, size, and temperature, ensuring the accuracy of the student model in the defect identification task. The distillation weights were determined by validating the power distribution defect dataset. The distilled student model is optimized, including model quantization, which uses INT8 quantization technology to convert the model weights from 32-bit floating-point numbers to 8-bit integers; and operator optimization, which uses the TensorRT framework to optimize the core operators and reduce memory access latency; the optimized model is then converted to ONNX format and deployed to the UAV edge inference engine.
[0030] Thirdly, an electronic device is provided, including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the intelligent identification method for defects in power distribution drone inspections.
[0031] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one instruction, which, when executed by a processor, implements the intelligent identification method for defects in power distribution drone inspections.
[0032] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects: Addressing the core pain points of high annotation costs, uninterpretable reasoning, and poor real-time performance in defect identification during power distribution drone inspections, this invention combines multimodal weakly supervised contrastive learning with a hierarchical semantic labeling tree for defects. This achieves automatic defect region segmentation with minimal text guidance, employs multimodal feature alignment for precise defect region localization, and generates structured samples of defect regions. Multimodal reasoning is then applied to these structured samples to generate interpretable reasoning chains, enabling accurate determination of defect types and attributes, improving the interpretability and reliability of the reasoning process, and ensuring its traceability and verifiability. Through cross-modal lightweight distillation, the defect types and attributes determined by the multimodal reasoning chain are transferred to obtain a lightweight model, which is then deployed to the drone's edge-side inference engine for real-time defect identification reasoning, meeting the real-time inspection needs of power distribution drones. Verification has shown that, in sample annotation, this invention, through multimodal weakly supervised contrastive learning combined with a hierarchical semantic labeling tree for defects, improves annotation efficiency by 4.3 times, reduces the cost of annotation per defect from 30 yuan to 6.9 yuan, shortens the annotation cycle for 100,000 frames of data from 100 days to 23 days, and achieves an IoU of 87.2% for defect region segmentation, a 17.2 percentage point improvement over traditional weakly supervised algorithms. In defect reasoning, multimodal thinking chain reasoning enables visualized output of "step-by-step reasoning + semantic basis," increasing inspectors' trust in the results from 65% to 92%, reducing the manual review rate from 100% to 10%, achieving a defect type accuracy of 95.2%, and reducing the misclassification rate for bird nests and tree branches to 3%. The efficiency was reduced by 1%, a decrease of 16.9 percentage points compared to the traditional single-modal model. In terms of edge deployment performance, the cross-modal lightweight distillation model parameter size reached 7.8M, a 99.94% reduction compared to the teacher model. Edge storage usage decreased from 50GB to 30MB, edge inference latency was 112ms, and total latency was ≤200ms, a 57.6% improvement compared to cloud inference. In inspection applications, the drone inspection efficiency reached 22km / day, 4.4 times that of manual inspection. The inspection time for 100 poles was shortened from 20 days to 5 days, realizing "machine inspection replacing human inspection". The defect missed rate was reduced from 15% to 4.8%, effectively solving the core pain point of defect identification in power distribution drone inspection.
[0033] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0034] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 Flowchart of the intelligent defect identification method for power distribution drone inspection according to an embodiment of the present invention; Figure 2 Principle architecture diagram of the intelligent defect identification method for power distribution drone inspection according to an embodiment of the present invention; Figure 3 A structural block diagram of the intelligent identification system for power distribution drone inspection defects according to an embodiment of the present invention. Detailed Implementation
[0036] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail. Flowcharts are used in the embodiments of this application to illustrate the operations performed by the apparatus according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more steps may be removed from these processes.
[0037] Please see Figure 1This invention proposes an intelligent defect identification method for power distribution drone inspections. By integrating Multimodal Weakly Supervised Comparative Labeling (MW-SCL), Multimodal Thinking Chain Inference (M-CoT), Cross-Modal Lightweight Distillation (Light-MMD), and a hierarchical defect semantic label tree construction algorithm, a full-link technical framework of "efficient sample construction - interpretable defect identification - real-time deployment at the edge" is constructed. This solves the problems of high labeling costs, uninterpretable reasoning, and poor real-time performance at the edge in typical defects in power distribution scenarios such as foreign objects in bird nests, insulator damage, broken conductor strands, and equipment corrosion. It is suitable for tasks such as autonomous inspection of power distribution drones, intelligent defect classification, and operation and maintenance decision support, and also accommodates extended applications such as collaborative inspection of live-line working robots in power distribution networks and AR defect visualization. The intelligent defect identification method for power distribution drone inspections in this invention aims to overcome the shortcomings of existing power distribution drone defect identification technologies, including high labeling costs, uninterpretable reasoning, and poor real-time performance at the edge. Specifically, its objectives include: 1. Reducing sample labeling costs. By combining the MW-SCL annotation algorithm with a hierarchical semantic labeling tree for defects, weakly supervised annotation with "minimal text guidance + multimodal feature alignment" is achieved, improving annotation efficiency by more than 4 times and reducing the cost of single defect annotation by 75%, while ensuring that the IoU of defect region segmentation is ≥85%; 2. Improve the interpretability and reliability of inference. By integrating visible light, infrared, and lidar multimodal features through the M-CoT inference algorithm, an inference chain conforming to power distribution professional knowledge is generated, reducing the defect misjudgment rate to below 5%, and the inference process is traceable and verifiable; 3. Achieve real-time deployment on the edge. By transferring knowledge from a large model to a lightweight model through the Light-MMD distillation algorithm, the number of parameters is compressed to below 8M, the edge inference latency is ≤150ms, and the accuracy decrease is ≤3%, meeting the real-time inspection requirements of power distribution drones. The intelligent defect identification method for power distribution drone inspection in this embodiment of the invention mainly includes the following steps: S1. Input the multimodal data collected by the power distribution drone into the pre-constructed hierarchical label tree of defect semantics, and locate the defect region based on multimodal weak supervised contrastive learning to generate structured samples of the defect region. S2. Perform multimodal reasoning on the structured samples of the defect area to generate an interpretable reasoning chain, and determine the defect type and attributes based on the interpretable reasoning chain. S3. Perform cross-modal lightweight distillation on the defect type and attribute determination results, transfer the knowledge to the lightweight model and then deploy it to the UAV edge inference engine to achieve real-time defect identification and inference.
[0038] In one possible implementation, step S1 constructs a three-level semantic tag tree of "equipment category - defect type - attribute parameter" to address the characteristics of diverse power distribution defect types and complex semantic relationships, providing structured semantic constraints for subsequent annotation and reasoning. In terms of the tag tree structure design, the first-level tags (equipment categories) cover the core equipment for power distribution inspection, including five categories: "insulators," "conductors," "towers," "transformers," and "ring main units," corresponding to the main objects inspected by power distribution drones. The second-level tags (defect types) are based on the "Power Distribution Line Operation Regulations," with each equipment category associated with typical defects. For example, "insulators" are associated with "damage," "spontaneous explosion," and "contamination," while "conductors" are associated with "broken strands," "wear," and "bird's nest foreign objects," totaling 12 core defect categories. In the third-level tags (attribute parameters), each defect type includes quantitative attributes such as "location," "severity," and "size / temperature." For example, for "bird's nest foreign objects," the location (conductor #XX pole segment, tower crossarm), severity (mild: coverage section <10%, moderate: 10%~30%, severe: >30%), and size (volume / m³) are included, and for "equipment corrosion," the temperature (temperature difference from the environment / ℃) is included.
[0039] In the mathematical representation of the tag tree, a tree structure model is adopted. Represents a label tree, where, V Represents a set of nodes. , This is a primary node (equipment category, 5 in total). It is a secondary node (12 defect types). It is a three-level node (with 36 attribute parameters); E Let be the set of edges. Represents a node and Parent-child relationships (e.g., "insulator" - "damaged" - "location"); L For the set of label weights, Representing an edge The semantic association weights are set based on the experience of power distribution experts and the frequency of defect occurrence. For example, the weight of "conductor-bird's nest foreign object" is l=1.2 (high frequency of occurrence), and the weight of "power distribution-oil leakage" is l=0.8 (low frequency of occurrence).
[0040] The semantic similarity of the tag tree is calculated using cosine similarity, for any two tag nodes. , Its semantic similarity S ( Calculate according to the following formula:
[0041] in, For nodes vThe semantic embedding vector (generated by a BERT model pre-trained on a power distribution defect corpus, dimension 128). for and The shortest path, d For path length, For path number k The weights of the edges. This formula ensures that semantically related labels (such as "insulator damage" and "insulator self-explosion") have higher similarity, providing constraints for subsequent multimodal alignment.
[0042] In one possible implementation, step S1, based on multimodal weakly supervised contrastive learning (MW-SCL) and combined with a hierarchical semantic labeling tree for defects, achieves weakly supervised annotation of "a small amount of text description → automatic segmentation of defect regions." The core is to accurately locate defect regions through visual-text-LiDAR multimodal feature alignment. The steps in step S1, which involve locating defect regions based on multimodal weakly supervised contrastive learning and generating structured samples of the defect regions, include: (1) Multimodal feature extraction. First, input the multimodal data collected by the power distribution UAV. This includes visible light images with a resolution of H=1080 and W=1920. Infrared thermal imaging with a temperature measurement range of -20℃ to 150℃ Point cloud quantity N≥ LiDAR point cloud The text describes the defects, including the corresponding text description T (e.g., "Bird's nest above pole #087 of the 10kV East Line, covering 30% of the cross-section"). Next, feature extraction is performed. For visual feature extraction, a lightweight CNN (MobileNetV3) is used to extract multi-scale features from the visible light image, outputting a visual feature map. Infrared images are processed by temperature threshold segmentation (e.g., the temperature of the rusted area is 5°C higher than the ambient temperature) to extract infrared feature maps. and through (Conv is a 1x1 convolution dimensionality reduction) fused into joint visual features; during point cloud feature extraction, the point cloud is first downsampled using VoxelGrid (voxel size is 0.05 cubic meters), and then spatial geometric features are extracted using PointNet. (M is the number of voxels), and then RoI Pooling (Region of Interest Pooling) maps the point cloud features to the image coordinate system to obtain... Text feature extraction is based on a defect semantic hierarchical label tree, which parses the text description T into a label sequence (e.g., "wire-bird's nest foreign object-location": #087 pole-severity-major), and generates text features through DistilBERT encoding. .
[0043] (2) Through cross-modal contrastive learning, semantic alignment of "visual-text-point cloud" features is achieved, forcing the features of defective regions to be similar to the text description features, and the features of non-defective regions to be significantly different from the text description features. The loss function is defined as:
[0044] in, For visual-text contrast loss, compute joint visual features. Text features The contrast loss is calculated using the following formula:
[0045] in, B For batch size, C The number of feature map channels. For the first b The sample, the first c Visual feature vectors of each channel, For the first Text features of each sample (negative sample). For the temperature coefficient (optimized through cross-validation in power distribution scenarios), this loss ensures that the visual features of the defective area are highly similar to the corresponding text features. For point cloud-text contrast loss, similar to Calculate point cloud mapping features Text features The contrast loss is calculated using the following formula:
[0046] This loss function utilizes the spatial geometric information of point clouds (such as the volume and location of bird nests) to enhance text semantic alignment and avoid mislabeling of visually similar objects (such as bird nests and tree branches). For visual-point cloud contrast loss, compute visual joint features Point cloud mapping features The contrast loss is used to ensure consistency within multimodal features, and the formula is:
[0047] in, Using the L2 norm, this loss minimizes the difference between visual and point cloud features, improving annotation robustness.
[0048] (3) Defect region segmentation and annotation. Based on the feature extraction network trained with MW-SCL loss, defect region segmentation is performed on the input multimodal data. Specifically, the visual joint features are calculated first. Text features Similarity graph , Secondly, calculate the point cloud mapping features. Text features Similarity graph , Then, the similarity maps are merged. (Visual features have higher weights), and the defect region mask is obtained through Otsu thresholding. (1 represents the defect area); finally, combined with the hierarchical label tree of defect semantics, the defect area is automatically labeled with three-level labels (such as "wire-bird's nest foreign object-location: #087 pole-severity: severe"), generating a structured sample library.
[0049] In one possible implementation, step S2 integrates multimodal features and a defect semantic label tree through a three-step reasoning logic of "feature fusion - step-by-step reasoning - result output" to generate an interpretable reasoning chain, achieving accurate determination of defect type and attributes. In the multimodal feature fusion stage, the structured samples output from multimodal weakly supervised contrastive learning are used as input for further processing. First, the visible light image... Feature-level fusion with infrared images, utilizing an attention fusion mechanism. Attention weights pass Calculations show that this mechanism can highlight defect areas, such as the high-temperature infrared characteristics of corroded areas. Next, the visual-infrared fusion features are then analyzed. Point cloud features Text label features The embedding matrix of the cross-attention module and the defect semantic label tree is used. (by all tag nodes) The components are fused together to obtain fusion characteristics. The formula is The CrossAttn cross-attention mechanism ensures that the fused features are closely linked to the power distribution semantics.
[0050] The step-by-step inference logic, as the core of M-CoT, breaks down the inference process into three steps: "equipment identification → defect type determination → attribute parameter quantification." Each step of the inference is based on fused features and a semantic label tree. The probabilistic model of the inference chain is represented as follows: ,in The initial state is "to be reasoned". to These correspond to equipment identification, defect type determination, and attribute parameter quantification status, respectively. This represents the inference probability from state k-1 to k. During the device identification stage, based on fused features... With the first-level node V1 of the tag tree, through the formula Calculate the probability of device category, where A weight matrix is used to identify equipment, and the equipment category with the highest probability is selected as the result. In defect type determination, secondary nodes associated with the label tree are filtered based on the equipment identification results, using a formula... Calculate the probability of defect type, where Weight matrix for determining defect type, For semantic weights. In the attribute parameter quantization stage, based on the defect type results, associated third-level nodes are selected, and attribute parameters are quantified through regression, such as location using LiDAR point cloud spatial coordinates. The severity is calculated based on the proportion of defective areas according to visual features. The determination of size or temperature is based on direct quantification of point cloud volume or infrared temperature. Finally, the inference chain output integrates the three-step inference results into natural language form, such as "Step 1: The fusion feature has the highest similarity to the semantic embedding of 'wire' (0.92) → the device is identified as 'wire'; Step 2: The similarity of the fusion feature to the semantic embedding of 'bird's nest foreign object' × weight = 0.88 × 1.2 = 1.056, which is higher than 'broken strand' (0.75 × 1.0 = 0.75) → the defect type is 'bird's nest foreign object'; Step 3: The lidar point cloud volume is 0.12m³ → the size is 0.12m³; the defect area covers 32% of the wire cross section → the severity is 'severe'; the point cloud spatial coordinates correspond to pole #087 of the line → the location is 'pole #087'; conclusion: there is a severe bird's nest foreign object in the pole segment of wire #087, with a volume of 0.12m³", thus achieving accurate and interpretable determination of defect type and attribute.
[0051] In one possible implementation, step S3 is to achieve real-time deployment on the UAV end side. Through cross-modal lightweight distillation (Light-MMD), the knowledge of the above-mentioned multimodal thinking chain reasoning large model (teacher model, 13B parameters) is transferred to the lightweight model (student model, 8M parameters). The core is the triple constraint of "feature alignment - semantic consistency - task loss" to ensure that the performance loss is minimized after lightweighting. (1) Teacher and student model architecture. The teacher model adopts the combination of "ViT-B / 16 (vision) + PointNet++ (point cloud) + BERT (text) + Transformer (reasoning)", with a parameter count of 13B. Although the accuracy is high, the model size is huge. The student model uses "MobileViT-XXS (vision) + simplified PointNet (point cloud) + DistilBERT (text) + lightweight Transformer (reasoning)", with a parameter count of only 8M. It has the advantages of small size and fast reasoning speed. It aims to obtain key information from the teacher model through knowledge distillation; (2) Distillation loss function. The distillation loss function is a weighted sum of multi-task losses. Its core function is to maintain the performance of the teacher model as close as possible while ensuring the lightweight nature of the student model. It mainly consists of three parts: First, feature alignment loss. Using MSE loss, we ensure that the multimodal feature distribution of the student model is consistent with that of the teacher model, so that the student model can learn the feature extraction ability of the teacher model; Second, semantic consistency loss. Based on cosine similarity loss, we ensure that the semantics of the inference chain of the student model are consistent with those of the teacher model, so that the inference logic of the two is the same; Finally, task loss. It consists of a weighted sum of classification loss and regression loss. The former uses cross-entropy loss to calculate the classification error of the device and the defect type, and the latter uses L1 loss to calculate the regression error of the position, size and temperature, which aims to ensure the accuracy of the student model in the defect identification task. Finally, the distillation weights are determined by the power distribution defect dataset to keep the accuracy of the student model within 3%; (3) End-side deployment optimization. In order to adapt to the end-side hardware of the UAV (such as NVIDIA Jetson OrinNX), the distilled student model is optimized in many aspects. The first is model quantization. By employing INT8 quantization technology, the model weights were converted from 32-bit floating-point numbers to 8-bit integers, significantly compressing the model size by 75% and doubling the inference speed. Secondly, operator optimization was implemented. Using the TensorRT framework, core operators such as convolution and attention were optimized to reduce memory access latency. Finally, the inference engine was deployed. The optimized model was converted to ONNX format and deployed to the UAV edge inference engine to achieve real-time defect identification inference.
[0052] Please see Figure 2 The present invention proposes an intelligent defect identification method for power distribution drone inspections: 1. Defect semantic hierarchical tag tree construction method To address the semantic association between power distribution equipment and defects, a three-level tag tree of "equipment category - defect type - attribute parameters" is constructed. Semantic association is calculated using hierarchical cosine similarity, providing structured constraints for multimodal labeling and reasoning. Unlike existing flat-structured tags, this invention's tag tree achieves precise semantic constraints, improving multimodal alignment accuracy.
[0053] 2. MW-SCL Weakly Supervised Labeling Algorithm By integrating multimodal features from visible light, infrared, and lidar, a triple contrast loss function of "visual-text-point cloud" was designed, achieving "automatic segmentation of defect areas + labeling with minimal text guidance," thus solving the problems of low accuracy and high cost of traditional weakly supervised labeling. Compared to existing technologies that only use visual-text features, this invention introduces point cloud contrast loss, improving the labeling accuracy in power distribution scenarios.
[0054] 3. M-CoT Multimodal Inference Algorithm This algorithm breaks down defect identification into a three-step reasoning process: "equipment identification - defect type determination - attribute parameter quantification." Based on multimodal fusion features and semantic constraints of the tag tree, it generates an interpretable natural language reasoning chain, breaking the "black box" limitation of traditional models. Unlike existing reasoning methods that directly output defect results, the reasoning chain of this invention is traceable, verifiable, and integrates multimodal features, resulting in higher reliability.
[0055] 4. Light-MMD cross-modal distillation algorithm The algorithm employs a triple distillation loss function of "feature alignment - semantic consistency - task loss" to transfer knowledge from a large multimodal model to a lightweight model, ensuring a balance between "lightweight" and "high accuracy" during edge deployment. Unlike existing distillation techniques that only optimize visual features, this invention introduces semantic consistency loss, preserving the ability to generate inference chains, with an accuracy decrease of ≤3%.
[0056] 5. Multi-module collaborative defect identification framework The framework integrates four major modules: "Label Tree - MW - SCL - M - CoT - Light - MMD," forming a full-link technical framework of "sample annotation - model inference - edge deployment," suitable for power distribution drone inspection scenarios. Unlike existing methods that optimize a single module, this invention's four modules support each other, with the label tree constraining annotation and inference, and distillation preserving inference capabilities, achieving optimal overall performance.
[0057] Please see Figure 3 Another embodiment of the present invention also proposes an intelligent defect identification system for power distribution drone inspection, comprising: The defect region localization module 301 is used to input the multimodal data collected by the power distribution drone into a pre-constructed hierarchical label tree for defect semantics, and to locate the defect region based on multimodal weak supervised contrastive learning, and generate a structured sample of the defect region. The thought chain reasoning module 302 is used to perform multimodal thought chain reasoning on the structured sample of the defect area, generate an interpretable reasoning chain, and determine the defect type and attributes based on the interpretable reasoning chain. The lightweight distillation module 303 is used to perform cross-modal lightweight distillation on the defect type and attribute determination results. After transferring the knowledge to the lightweight model, it is deployed to the UAV end-side inference engine to realize real-time defect identification and inference.
[0058] In one possible implementation, when the defect area localization module 301 inputs the multimodal data collected by the power distribution drone into a pre-constructed hierarchical semantic labeling tree, the hierarchical semantic labeling tree is a three-level semantic labeling tree of "equipment category - defect type - attribute parameter". The first-level label is the equipment category, which corresponds to the main objects inspected by the power distribution drone; the second-level label is the defect type, which corresponds to the typical defects associated with each equipment category; and the third-level label is the attribute parameter, which corresponds to the quantitative attributes of each defect type. Tree structure model This represents a three-level semantic tag tree, where, V Represents a set of nodes. , To represent the first-level node of the equipment category, For secondary nodes representing defect types, A third-level node representing attribute parameters; E Let be the set of edges. Represents a node and The father-son relationship; L For the set of label weights, Representing an edge The semantic association weights are set based on the experience of power distribution experts and the frequency of defect occurrence. The semantic similarity of the three-level semantic tag tree is calculated using cosine similarity. For any two tag nodes... , Semantic similarity S ( Calculate according to the following formula:
[0059] In the formula, For nodes v semantic embedding vector, for and The shortest path, d For path length, Let be the weight of the k-th edge of the path.
[0060] In one possible implementation, the defect region localization module 301 locates the defect region based on multimodal weak supervised contrastive learning. When generating a structured sample of the defect region, it extracts multimodal features of vision, text, and point cloud from the multimodal data collected by the power distribution drone. Through cross-modal contrastive learning, semantic alignment of "visual-text-point cloud" features is achieved, forcing the features of defective regions to be similar to the text description features, and the features of non-defective regions to be significantly different from the text description features; Based on the semantically aligned "visual-text-point cloud" features, defect regions are segmented, and defect regions are labeled using a hierarchical semantic label tree to generate structured samples of defect regions.
[0061] In one possible implementation, when the defect region localization module 301 extracts multimodal features (visual, textual, and point cloud) from the multimodal data collected by the power distribution drone, the multimodal data includes visible light images, infrared thermal images, lidar point clouds, and corresponding textual descriptions of defects. During multimodal feature extraction, a lightweight convolutional neural network (CNN) is used to extract multi-scale features from the visible light images, outputting a visual feature map. The infrared images are processed by temperature threshold segmentation, and infrared feature maps are extracted and fused into joint visual features. When extracting point cloud features, the point cloud is first downsampled using a voxel grid, and spatial geometric features are extracted using a point net (PointNet). Then, the point cloud features are mapped to the image coordinate system using region of interest pooling (RoIPooling). Text feature extraction is based on a hierarchical semantic label tree for defects, parsing the text description T into a label sequence and generating text features using DistilBERT encoding.
[0062] In one possible implementation, the defect region localization module 301 achieves semantic alignment of "visual-text-point cloud" features through cross-modal contrastive learning, forcing the defect region features to be similar to the text description features. When the non-defect region features differ significantly from the text description features, the expression for the loss function is defined as follows:
[0063] In the formula, Visual-text contrast loss is calculated using the following formula to compute joint visual features. Text features Comparative loss:
[0064] In the formula, B Batch size; C This refers to the number of channels in the feature map. Let be the visual feature vector of the b-th sample and the c-th channel; For the first Text features of each sample Temperature coefficient; For point cloud-text contrast loss, the point cloud mapping features are calculated using the following formula. Text features Comparative loss:
[0065] For the visual-point cloud contrast loss, the visual joint features are calculated using the following formula. Point cloud mapping features Comparative loss:
[0066] In the formula, It is an L2 norm.
[0067] In one possible implementation, the defect region localization module 301 performs defect region segmentation based on semantically aligned "visual-textual-point cloud" features, and combines a hierarchical defect semantic label tree to annotate the defect region. When generating structured samples of the defect region, it calculates the visual joint features. Text features Similarity map; and, calculating point cloud mapping features. Text features The similarity graphs are fused together, and the defect region mask is obtained by Otsu threshold segmentation. Based on the obtained defect region mask, the corresponding labels are labeled for the defect region in combination with the defect semantic hierarchical label tree, and a structured sample library is generated.
[0068] In one possible implementation, the thought chain reasoning module 302 utilizes an attention fusion mechanism to analyze visible light images. Feature-level fusion with infrared images is expressed as follows: Among them, attention weight pass Calculations show that the visual-infrared fusion features are... Point cloud features Text label features The embedding matrix of the cross-attention module and the defect semantic label tree is used. , obtain fusion features The calculation expression is: The inference chain probabilistic model is represented as ,in, The initial state is "to be reasoned". to These correspond to equipment identification, defect type determination, and attribute parameter quantification status, respectively. This represents the inference probability from state k-1 to k; in the device identification stage, it is based on fused features. With the first-level node V1 of the tag tree, through the expression Calculate the probability of device category, where, For equipment identification, a weight matrix is used to select the equipment category with the highest probability as the result. In defect type determination, secondary nodes associated with the label tree are filtered based on the equipment identification results, using an expression... Calculate the probability of defect type, where, Weight matrix for determining defect type, For semantic weights; in the attribute parameter quantification stage, based on the defect type results, related third-level nodes are filtered, and attribute parameters are quantified through regression. The severity is based on the proportion of defect areas in visual features. The determination is based on direct quantization of size or temperature using point cloud volume or infrared temperature; the inference chain output integrates the three-step inference results into a natural language form.
[0069] In one possible implementation, the lightweight distillation module 303 enables the student model to obtain key information from the teacher model through knowledge distillation; the distillation loss function is a weighted sum of multi-task losses, including feature alignment loss, semantic consistency loss, and task loss; the feature alignment loss uses MSE loss to ensure that the multimodal feature distribution of the student model is consistent with that of the teacher model, allowing the student model to have the feature extraction capability of the teacher model; the semantic consistency loss is based on cosine similarity loss to ensure that the semantics of the inference chain of the student model are consistent with those of the teacher model, ensuring that the inference logic of the two is the same; the task loss is composed of a weighted sum of classification loss and regression loss, the classification loss uses cross-entropy loss to calculate the classification error of the device and defect type, and the regression loss uses L1 loss to calculate the regression error of position, size, and temperature, ensuring the accuracy of the student model in the defect identification task; The distillation weights were determined by validating the power distribution defect dataset. The distilled student model is optimized, including model quantization, which uses INT8 quantization technology to convert the model weights from 32-bit floating-point numbers to 8-bit integers; and operator optimization, which uses the TensorRT framework to optimize the core operators and reduce memory access latency; the optimized model is then converted to ONNX format and deployed to the UAV edge inference engine.
[0070] Another embodiment of the present invention also proposes an electronic device, including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the intelligent identification method for defects in power distribution drone inspections.
[0071] Another embodiment of the present invention also proposes a computer-readable storage medium storing at least one instruction, which, when executed by a processor, implements the intelligent identification method for defects in power distribution drone inspections.
[0072] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. For ease of explanation, the above content only shows the parts related to the embodiments of the present invention; for specific technical details not disclosed, please refer to the method section of the embodiments of the present invention. This computer-readable storage medium is non-transitory and can be stored in storage devices formed by various electronic devices, enabling the execution process described in the method of the embodiments of the present invention.
[0073] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0074] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0075] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0076] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0077] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.
Claims
1. A method for intelligent identification of defects in power distribution drone inspections, characterized in that, include: The multimodal data collected by the power distribution drone is input into a pre-constructed hierarchical label tree for defect semantics, and the defect region is located based on multimodal weak supervised contrastive learning to generate structured samples of the defect region. The defect semantic hierarchical tag tree is a three-level semantic tag tree of "equipment category - defect type - attribute parameter". The first-level tag is the equipment category, which corresponds to the main objects of the power distribution drone inspection; the second-level tag is the defect type, which corresponds to the typical defects associated with each equipment category; and the third-level tag is the attribute parameter, which corresponds to the quantitative attributes of each defect type. The steps for locating defect regions and generating structured samples of defect regions based on multimodal weakly supervised contrastive learning include: Multimodal feature extraction from visual, textual, and point cloud data collected by power distribution drones; Through cross-modal contrastive learning, semantic alignment of "visual-text-point cloud" features is achieved, forcing the features of defective regions to be similar to the text description features, and the features of non-defective regions to be significantly different from the text description features. Based on the semantically aligned "visual-text-point cloud" features, defect regions are segmented, and defect regions are labeled using a hierarchical semantic label tree to generate structured samples of defect regions. Multimodal reasoning is performed on structured samples of defective regions to generate interpretable reasoning chains, and the defect type and attributes are determined based on these interpretable reasoning chains. In the step of performing multimodal reasoning on structured samples of defective regions to generate interpretable reasoning chains and determining the defect type and attributes based on these interpretable reasoning chains, an attention fusion mechanism is used to analyze visible light images. Feature-level fusion with infrared images is expressed as follows: Among them, attention weight pass Calculations show that the visual-infrared fusion features are... Point cloud features Text label features The embedding matrix of the cross-attention module and the defect semantic label tree is used. By performing fusion, fusion characteristics are obtained. The calculation expression is: The inference chain probabilistic model is represented as ,in, The initial state is "to be reasoned". to These correspond to equipment identification, defect type determination, and attribute parameter quantification status, respectively. This represents the inference probability from state k-1 to k; in the device identification stage, it is based on fused features. With the first-level node V1 of the tag tree, through the expression Calculate the probability of device category, where, For equipment identification, a weight matrix is used to select the equipment category with the highest probability as the result. In defect type determination, secondary nodes associated with the label tree are filtered based on the equipment identification results, using an expression... Calculate the probability of defect type, where, Weight matrix for determining defect type, For semantic weights; in the attribute parameter quantification stage, based on the defect type results, related third-level nodes are filtered, and attribute parameters are quantified through regression. The severity is based on the proportion of defect areas in visual features. The determination is based on direct quantization of size or temperature using point cloud volume or infrared temperature; the inference chain output integrates the three-step inference results into natural language form. A cross-modal lightweight distillation process is performed on the defect type and attribute determination results. This knowledge is then transferred to a lightweight model and deployed to the UAV edge-side inference engine to achieve real-time defect identification and inference. In this step, knowledge distillation enables the student model to acquire key information from the teacher model. The distillation loss function is a weighted sum of multi-task losses, including feature alignment loss, semantic consistency loss, and task loss. The feature alignment loss uses MSE loss to ensure that the multimodal feature distribution of the student model is consistent with that of the teacher model, allowing the student model to possess the feature extraction capabilities of the teacher model. The semantic consistency loss is based on cosine similarity loss, ensuring that the semantics of the inference chain of the student model are consistent with those of the teacher model, ensuring that their inference logic is the same. The task loss consists of a weighted sum of classification loss and regression loss. The classification loss uses cross-entropy loss to calculate the classification error of the device and defect type, while the regression loss uses L1 loss to calculate the regression error of position, size, and temperature, ensuring the accuracy of the student model in the defect identification task. The distillation weights were determined by validating the power distribution defect dataset. The distilled student model is optimized, including model quantization, which uses INT8 quantization technology to convert the model weights from 32-bit floating-point numbers to 8-bit integers; and operator optimization, which uses the TensorRT framework to optimize the core operators and reduce memory access latency; the optimized model is then converted to ONNX format and deployed to the UAV edge inference engine.
2. The intelligent defect identification method for power distribution drone inspection according to claim 1, characterized in that, Tree structure model This represents a three-level semantic tag tree, where, V Represents a set of nodes. , To represent the first-level node of the equipment category, For secondary nodes representing defect types, A third-level node representing attribute parameters; E Let be the set of edges. Represents a node and The father-son relationship; L For the set of label weights, Representing an edge The semantic association weights are set based on the experience of power distribution experts and the frequency of defect occurrence. The semantic similarity of the three-level semantic tag tree is calculated using cosine similarity. For any two tag nodes... , Semantic similarity S ( Calculate according to the following formula: In the formula, For nodes v semantic embedding vector, for and The shortest path, d For path length, For path number z The weight of the edge.
3. The intelligent defect identification method for power distribution drone inspection according to claim 1, characterized in that, When extracting multimodal features (visual, textual, and point cloud) from the multimodal data collected by the power distribution drone, the multimodal data collected by the power distribution drone includes visible light images, infrared thermal images, lidar point clouds, and corresponding textual descriptions of defects. For multimodal feature extraction, a lightweight convolutional neural network (CNN) is used to extract multi-scale features from the visible light images, outputting visual feature maps. Infrared images are segmented using a temperature threshold, and infrared feature maps are extracted and fused into joint visual features. For point cloud feature extraction, point clouds are first downsampled using a voxel grid, and spatial geometric features are extracted using a point net (PointNet). Then, region-of-interest (RoI) pooling is used to map the point cloud features to the image coordinate system. For textual feature extraction, a hierarchical label tree based on defect semantics is used to parse the text description T into a label sequence, and DistilBERT encoding is used to generate textual features.
4. The intelligent defect identification method for power distribution drone inspection according to claim 1, characterized in that, In the step of achieving semantic alignment of "visual-text-point cloud" features through cross-modal contrastive learning, and forcing defect region features to be similar to text description features while ensuring that non-defect region features are significantly different from text description features, the expression for the loss function is defined as follows: In the formula, Visual-text contrast loss is calculated using the following formula to compute joint visual features. Text features Comparative loss: In the formula, B Batch size; C This refers to the number of channels in the feature map. Let be the visual feature vector of the b-th sample and the c-th channel; For the first Text features of each sample Temperature coefficient; For point cloud-text contrast loss, the point cloud mapping features are calculated using the following formula. Text features Comparative loss: For the visual-point cloud contrast loss, the visual joint features are calculated using the following formula. Point cloud mapping features Comparative loss: In the formula, It is an L2 norm.
5. The intelligent defect identification method for power distribution drone inspection according to claim 4, characterized in that, In the step of segmenting defect regions based on semantically aligned "visual-text-point cloud" features and labeling defect regions using a hierarchical semantic label tree to generate structured samples of defect regions, the visual joint features are calculated. Text features Similarity map; and, calculating point cloud mapping features. Text features The similarity graphs are fused together, and the defect region mask is obtained by Otsu threshold segmentation. Based on the obtained defect region mask, the corresponding labels are labeled for the defect region in combination with the defect semantic hierarchical label tree, and a structured sample library is generated.
6. A power distribution drone inspection defect intelligent identification system, characterized in that, include: The defect region localization module is used to input the multimodal data collected by the power distribution drone into a pre-constructed hierarchical label tree for defect semantics, and to locate the defect region based on multimodal weak supervised contrastive learning, generating a structured sample of the defect region. The defect semantic hierarchical tag tree is a three-level semantic tag tree of "equipment category - defect type - attribute parameter". The first-level tag is the equipment category, which corresponds to the main objects of the power distribution drone inspection; the second-level tag is the defect type, which corresponds to the typical defects associated with each equipment category; and the third-level tag is the attribute parameter, which corresponds to the quantitative attributes of each defect type. The defect region localization module locates defect regions based on multimodal weak supervised contrastive learning. When generating structured samples of defect regions, it extracts multimodal features of vision, text, and point cloud from the multimodal data collected by the power distribution drone. Through cross-modal contrastive learning, semantic alignment of "visual-text-point cloud" features is achieved, forcing the features of defective regions to be similar to the text description features, and the features of non-defective regions to be significantly different from the text description features. Based on the semantically aligned "visual-text-point cloud" features, defect regions are segmented, and defect regions are labeled using a hierarchical semantic label tree to generate structured samples of defect regions. The thought chain reasoning module is used to perform multimodal thought chain reasoning on structured samples of defect regions, generating interpretable reasoning chains, and determining the defect type and attributes based on these interpretable reasoning chains. The thought chain reasoning module utilizes an attention fusion mechanism on visible light images. Feature-level fusion with infrared images is expressed as follows: Among them, attention weight pass Calculations show that the visual-infrared fusion features are... Point cloud features Text label features The embedding matrix of the cross-attention module and the defect semantic label tree is used. By performing fusion, fusion characteristics are obtained. The calculation expression is: The inference chain probabilistic model is represented as ,in, The initial state is "to be reasoned". to These correspond to equipment identification, defect type determination, and attribute parameter quantification status, respectively. This represents the inference probability from state k-1 to k; in the device identification stage, it is based on fused features. With the first-level node V1 of the tag tree, through the expression Calculate the probability of device category, where, For equipment identification, a weight matrix is used to select the equipment category with the highest probability as the result. In defect type determination, secondary nodes associated with the label tree are filtered based on the equipment identification results, using an expression... Calculate the probability of defect type, where, Weight matrix for determining defect type, For semantic weights; in the attribute parameter quantification stage, based on the defect type results, related third-level nodes are filtered, and attribute parameters are quantified through regression. The severity is based on the proportion of defect areas in visual features. The determination is based on direct quantization of size or temperature using point cloud volume or infrared temperature; the inference chain output integrates the three-step inference results into natural language form. The lightweight distillation module performs cross-modal lightweight distillation on defect type and attribute determination results, transferring knowledge to a lightweight model and deploying it to the UAV edge inference engine for real-time defect identification and inference. This module enables the student model to acquire key information from the teacher model through knowledge distillation. The distillation loss function is a weighted sum of multi-task losses, including feature alignment loss, semantic consistency loss, and task loss. The feature alignment loss uses MSE loss to ensure that the multimodal feature distribution of the student model is consistent with that of the teacher model, allowing the student model to possess the feature extraction capabilities of the teacher model. The semantic consistency loss is based on cosine similarity loss, ensuring that the semantics of the inference chain of the student model are consistent with those of the teacher model, ensuring that their inference logic is the same. The task loss is a weighted sum of classification loss and regression loss. The classification loss uses cross-entropy loss to calculate the classification error of the device and defect type, while the regression loss uses L1 loss to calculate the regression error of location, size, and temperature, ensuring the accuracy of the student model in the defect identification task. The distillation weights were determined by validating the power distribution defect dataset. The distilled student model is optimized, including model quantization, which uses INT8 quantization technology to convert the model weights from 32-bit floating-point numbers to 8-bit integers; and operator optimization, which uses the TensorRT framework to optimize the core operators and reduce memory access latency; the optimized model is then converted to ONNX format and deployed to the UAV edge inference engine.
7. The intelligent defect identification system for power distribution drone inspection according to claim 6, characterized in that, When the defect region localization module inputs the multimodal data collected by the power distribution drone into the pre-constructed hierarchical defect semantic labeling tree, it adopts a tree structure model. This represents a three-level semantic tag tree, where, V Represents a set of nodes. , To represent the first-level node of the equipment category, For secondary nodes representing defect types, A third-level node representing attribute parameters; E Let be the set of edges. Represents a node and The father-son relationship; L For the set of label weights, Representing an edge The semantic association weights are set based on the experience of power distribution experts and the frequency of defect occurrence. The semantic similarity of the three-level semantic tag tree is calculated using cosine similarity. For any two tag nodes... , Semantic similarity S ( Calculate according to the following formula: In the formula, For nodes v semantic embedding vector, for and The shortest path, d For path length, For path number z The weight of the edge.
8. The intelligent defect identification system for power distribution drone inspection according to claim 6, characterized in that, When the defect area localization module extracts multimodal features (visual, textual, and point cloud) from the multimodal data collected by the power distribution drone, the multimodal data includes visible light images, infrared thermal images, lidar point clouds, and corresponding textual descriptions of defects. For multimodal feature extraction, a lightweight convolutional neural network (CNN) is used to extract multi-scale features from the visible light images, outputting a visual feature map. Infrared images are segmented using a temperature threshold, and infrared feature maps are extracted and fused into joint visual features. For point cloud feature extraction, point clouds are first downsampled using a voxel grid, and spatial geometric features are extracted using a point net (PointNet). Then, region of interest (RoI) pooling is used to map the point cloud features to the image coordinate system. For textual feature extraction, a hierarchical label tree based on defect semantics is used to parse the text description T into a label sequence, and DistilBERT encoding is used to generate textual features.
9. The intelligent defect identification system for power distribution drone inspection according to claim 6, characterized in that, The defect region localization module achieves semantic alignment of "visual-text-point cloud" features through cross-modal contrastive learning, forcing defect region features to be similar to text description features. When non-defect region features differ significantly from text description features, the expression for the loss function is defined as follows: In the formula, Visual-text contrast loss is calculated using the following formula to compute joint visual features. Text features Comparative loss: In the formula, B Batch size; C This refers to the number of channels in the feature map. Let be the visual feature vector of the b-th sample and the c-th channel; For the first Text features of each sample Temperature coefficient; For point cloud-text contrast loss, the point cloud mapping features are calculated using the following formula. Text features Comparative loss: For the visual-point cloud contrast loss, the visual joint features are calculated using the following formula. Point cloud mapping features Comparative loss: In the formula, It is an L2 norm.
10. The intelligent defect identification system for power distribution drone inspection according to claim 9, characterized in that, The defect region localization module performs defect region segmentation based on semantically aligned "visual-text-point cloud" features, and annotates the defect regions using a hierarchical semantic label tree. When generating structured samples of defect regions, it calculates joint visual features. Text features Similarity map; and, calculating point cloud mapping features. Text features The similarity graphs are fused together, and the defect region mask is obtained by Otsu threshold segmentation. Based on the obtained defect region mask, the corresponding labels are labeled for the defect region in combination with the defect semantic hierarchical label tree, and a structured sample library is generated.
11. An electronic device, characterized in that, It includes a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the intelligent identification method for power distribution drone inspection defects as described in any one of claims 1 to 5.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which, when executed by a processor, implements the intelligent identification method for power distribution drone inspection defects as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Intelligent industrial defect detection method and system based on fusion model
CN117853492A
Multi-modal information tagging method, apparatus and device, and storage medium and product
WO2025148651A1