Power transmission line defect detection method and device and electronic equipment

By combining multimodal data with a target architecture model for preliminary identification and correction, the problem of inaccurate transmission line defect detection was solved, achieving higher detection accuracy and power supply reliability.

CN121582258BActive Publication Date: 2026-05-08STATE GRID BEIJING ELECTRIC POWER CO +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STATE GRID BEIJING ELECTRIC POWER CO
Filing Date
2026-01-27
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The inaccurate detection of defects in transmission lines in existing technologies leads to insufficient fault prevention and power supply reliability.

Method used

Defect detection is performed by combining multimodal data with the target architecture model. The first part initially identifies defects, and the second part corrects the results to ensure that the model output distribution is aligned and improves the detection accuracy.

Benefits of technology

This improved the accuracy and reliability of power transmission line defect detection, ensuring timely fault detection and handling, and reducing power outage time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121582258B_ABST
    Figure CN121582258B_ABST
Patent Text Reader

Abstract

The application discloses a power transmission line defect detection method and device and electronic equipment. The method comprises the following steps: acquiring multi-modal data corresponding to a target power transmission line; calling a target architecture model; inputting a power transmission line image into a first part of the target architecture model to obtain an initial defect detection result corresponding to the target power transmission line; inputting the initial defect detection result and the multi-modal data into a second part of the target architecture model to correct the initial defect detection result and obtain a target defect detection result corresponding to the target power transmission line. The application solves the technical problem of inaccurate defect detection when a power transmission line is detected in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of defect detection, and more specifically, to a method, apparatus, and electronic equipment for detecting defects in power transmission lines. Background Technology

[0002] With the continuous growth of electricity demand and the increasing scale of the power grid, the operation and maintenance management of transmission lines has become increasingly important. Timely detection and handling of defects in transmission lines, such as insulator damage, broken conductor strands, and tower corrosion, can effectively prevent faults, reduce power outage time, and improve power supply reliability. However, in related technologies, there are technical problems with inaccurate defect detection when performing defect detection on transmission lines.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This invention provides a method, apparatus, and electronic device for detecting defects in power transmission lines, which at least solves the technical problem of inaccurate defect detection when performing defect detection on power transmission lines in related technologies.

[0005] According to one aspect of the present invention, a method for detecting defects in a power transmission line is provided, comprising: acquiring multimodal data corresponding to a target power transmission line, wherein the multimodal data includes a power transmission line image; retrieving a target architecture model, wherein the target architecture model includes a first part and a second part, the target architecture model being trained based on a joint training function and a training dataset, the joint training function being used to determine the similarity between the output of the first part and the output of the second part, the first part being used for defect identification, and the second part being used to correct the output result of the first part; inputting the power transmission line image into the first part of the target architecture model to obtain an initial defect detection result corresponding to the target power transmission line; inputting the initial defect detection result and the multimodal data into the second part of the target architecture model to correct the initial defect detection result to obtain a target defect detection result corresponding to the target power transmission line.

[0006] Optionally, before retrieving the target architecture model, the process includes: determining training target parameters when the training data set includes multiple sample images, wherein the training target parameters include a predetermined similarity threshold; inputting any one of the multiple sample images into a first part of the target architecture model to obtain a training output corresponding to the first part under the given sample image; determining a training output corresponding to the second part under the given sample image based on the given sample image and the training output corresponding to the first part under the given sample image; determining the output similarity between the training output corresponding to the first part and the training output corresponding to the second part under the given sample image based on the joint training function, and determining whether the output similarity is greater than or equal to the predetermined similarity threshold; and stopping training the target architecture model when the determination result is that the output similarity is greater than or equal to the predetermined similarity threshold.

[0007] Optionally, before retrieving the target architecture model, the process includes: if the training data set includes positive sample data and negative sample data, inputting the positive sample data into the second part to obtain a positive sample output corresponding to the positive sample data, wherein the positive sample data is sample data with a defect severity greater than a defect threshold, and the negative sample data is sample data with a defect severity less than or equal to the defect threshold; inputting the negative sample data into the second part to obtain a negative sample output corresponding to the negative sample data; determining a comparison loss function corresponding to the second part, wherein the comparison loss function is used to determine the accuracy of defect detection in the second part; determining the defect detection accuracy corresponding to the second part based on the positive sample output, the negative sample output, and the comparison loss function, until the defect detection accuracy is greater than or equal to the defect detection accuracy threshold, and then stopping training on the second part.

[0008] Optionally, the step of inputting the transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line includes: determining a defect feature map corresponding to the transmission line image; dividing the defect feature map to obtain multiple sub-maps; determining sub-defect features corresponding to each of the multiple sub-maps, and multiple association features corresponding to each of the multiple sub-maps, wherein the multiple association features represent the defect association relationship between the corresponding sub-map and other sub-maps; and obtaining the initial defect detection result corresponding to the target transmission line based on the sub-defect features and multiple association features corresponding to each of the multiple sub-maps.

[0009] Optionally, obtaining the initial defect detection result corresponding to the target transmission line based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs includes: determining the enhanced feature map corresponding to the defect feature map based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs; and determining the initial defect detection result corresponding to the target transmission line based on the enhanced feature map.

[0010] Optionally, when the second part includes a reference defect detection set, the initial defect detection result and the multimodal data are input into the second part of the target architecture model to correct the initial defect detection result and obtain a target defect detection result corresponding to the target transmission line. This includes: determining multimodal defect features corresponding to the target transmission line based on the multimodal data; determining similar transmission lines corresponding to the target transmission line from multiple reference transmission lines based on the multimodal defect features, wherein the similar transmission lines are reference transmission lines whose difference index between the reference defect features and the multimodal defect features is less than a difference threshold; wherein the reference defect detection set includes multiple reference transmission lines, and reference defect features and reference defect detection results corresponding to the multiple reference transmission lines respectively; and correcting the initial defect detection result based on the multimodal defect features and the reference defect detection results corresponding to the similar transmission lines to obtain a target defect detection result corresponding to the target transmission line.

[0011] Optionally, after inputting the transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line, the method further includes: if the target transmission line is a predetermined transmission line and the initial defect detection result includes defect region parameters, determining the target scene parameters corresponding to the target transmission line, wherein the predetermined transmission line is a transmission line under a predetermined scene; determining a fused image based on the target scene parameters, the defect region parameters, and the transmission line image; and performing incremental training on the first part based on the fused image to obtain the incrementally trained first part.

[0012] According to one aspect of the present invention, a transmission line defect detection device is provided, comprising: an acquisition module for acquiring multimodal data corresponding to a target transmission line, wherein the multimodal data includes an image of the transmission line; a retrieval module for retrieving a target architecture model, wherein the target architecture model includes a first part and a second part, the target architecture model being trained based on a joint training function and a training dataset, the joint training function being used to determine the similarity between the output of the first part and the output of the second part, the first part being used for defect identification, and the second part being used to correct the output result of the first part; a first determination module for inputting the transmission line image into the first part of the target architecture model to obtain an initial defect detection result corresponding to the target transmission line; and a second determination module for inputting the initial defect detection result and the multimodal data into the second part of the target architecture model to correct the initial defect detection result and obtain a target defect detection result corresponding to the target transmission line.

[0013] According to one aspect of the present invention, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the transmission line defect detection method described in any of the preceding claims.

[0014] According to one aspect of the present invention, a computer-readable storage medium is provided, comprising: when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to perform the transmission line defect detection method described in any of the preceding claims.

[0015] In this embodiment of the invention, multimodal defect features corresponding to the target transmission line are determined. Based on the multimodal defect features, similar transmission lines corresponding to the target transmission line are determined from multiple reference transmission lines. The similar transmission lines are those whose difference index between the reference defect features and the multimodal defect features is less than a difference threshold. The reference defect detection set includes multiple reference transmission lines, and reference defect features and reference defect detection results corresponding to each of the multiple reference transmission lines. Based on the multimodal defect features and the reference defect detection results corresponding to the similar transmission lines, the initial defect detection results are corrected to obtain the target defect detection result corresponding to the target transmission line. Preliminary defect identification is performed through the first part of the target architecture model, and then the preliminary results are corrected by the second part. Since the target architecture model is trained based on a joint training function and a training data set, and the joint training function is used to ensure the alignment of the output distributions of the first and second parts, the collaboration capability between models can be maximized according to the target architecture model, thereby improving the accuracy of transmission line defect detection. This solves the technical problem of inaccurate defect detection when performing defect detection on transmission lines in related technologies. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0017] Figure 1 This is a flowchart of a transmission line defect detection method according to an embodiment of the present invention;

[0018] Figure 2 This is a flowchart of a transmission line defect detection method in an optional embodiment of the present invention;

[0019] Figure 3 This is a structural block diagram of a power transmission line defect detection device according to an embodiment of the present invention. Detailed Implementation

[0020] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0021] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0022] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0023] YOLOv5: YOLOv5 is an object detection algorithm belonging to the YOLO (You Only Look Once) series. It completes the localization and classification of objects through a single forward propagation and has the ability to detect objects quickly.

[0024] Faster R-CNN: Faster R-CNN is a two-stage object detection algorithm consisting of a Region Proposal Network (RPN) and Fast R-CNN. It detects objects by generating candidate regions, classifying them, and regressing bounding boxes.

[0025] FLOPs: FLOPs (Floating Point Operations per Second) refers to the number of floating-point operations performed per second, used to measure the computational load of a computing task.

[0026] YOLOv12: YOLOv12 is one of the YOLO series of object detection models. It is a real-time object detection framework with attention mechanism as its core.

[0027] OverLoCK: OverLoCK is a technique for feature enhancement and refined modeling, used to improve the detection capabilities of models in complex scenes.

[0028] mAP@0.5: mAP@0.5 (mean Average Precision at Intersection over Union threshold of 0.5) refers to the average precision when the IoU (Intersection over Union) threshold is 0.5.

[0029] LLaVA-MoD: A lightweight multimodal large model that integrates a sparse expert hybrid (MoE) architecture and knowledge distillation techniques.

[0030] UNIGEN: UNIGEN is a technique for generative adversarial networks (GANs) and data augmentation that generates synthetic data and performs pseudo-re-labeling to improve the generalization ability of the model.

[0031] Backbone: Backbone refers to the feature extraction part of the YOLOv12 model, which is used to extract feature maps from the input image.

[0032] Base-Net: Base-Net is the foundational network component of the OverLock architecture, used to extract basic visual features.

[0033] Overview-Net: Overview-Net is the global semantic overview network part of the OverLoCK architecture, used to generate a global semantic overview.

[0034] Focus-Net: Focus-Net is the fine-grained modeling network part of the OverLoCK architecture, used for fine-grained modeling of regions of interest (ROI).

[0035] Sigmoid: Sigmoid is an activation function whose output ranges from 0 to 1 and has an S-shaped curve.

[0036] KL divergence: KL divergence is a statistic that measures the difference between two probability distributions.

[0037] FAISS Vector Database: FAISS (Facebook AI Similarity Search) is a high-performance similarity search library for quickly retrieving and comparing vectors.

[0038] Transformer Decoder: The Transformer decoder is the decoding part of the Transformer architecture, used to generate sequence output.

[0039] OTA technology: OTA technology is a technology that provides software updates and firmware upgrades to electronic devices via wireless networks. It allows devices to receive and install update packages via wireless networks without physical contact, thereby enabling feature optimization, bug fixes, or new feature expansions.

[0040] Context-Mixing Dynamic Convolution: Context-Mixing dynamic convolution (ContMix) is a type of dynamic convolution that allows convolutional layers to dynamically adjust their weights based on contextual information while maintaining strong local inductive bias.

[0041] Example 1

[0042] According to an embodiment of the present invention, an embodiment of a method for detecting defects in transmission lines is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0043] Figure 1 This is a flowchart of a transmission line defect detection method according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0044] S102, acquire multimodal data corresponding to the target transmission line, wherein the multimodal data includes images of the transmission line.

[0045] In step S102 of this application, multimodal data corresponding to the target transmission line is obtained.

[0046] This involves the target transmission line, which is the transmission line that needs to be detected and analyzed in the transmission line defect detection task.

[0047] This involves multimodal data, which is a dataset containing various types of data (such as transmission line images, meteorological data, and equipment ledger information). In the scenario of transmission line defect detection, it covers various types of information related to transmission lines, which describe the status and environment of the transmission lines from different perspectives, helping to conduct more comprehensive defect detection and analysis.

[0048] This includes images of power transmission lines, which are obtained by taking pictures of power transmission lines and their ancillary facilities using imaging devices (such as cameras mounted on drones or cameras on line inspection robots). These images provide visual information for the inspection of power transmission lines.

[0049] Acquire multimodal data corresponding to the target transmission line to provide a comprehensive and multidimensional data foundation for transmission line defect detection.

[0050] S104, retrieve the target architecture model, which includes a first part and a second part. The target architecture model is trained based on a joint training function and a training data set. The joint training function is used to determine the similarity between the output of the first part and the output of the second part. The first part is used for defect identification, and the second part is used to correct the output of the first part.

[0051] In step S104 of this application, the target architecture model is retrieved.

[0052] This involves a target architecture model, which is used for transmission line defect detection. It consists of two main parts: a first part and a second part. This model is trained using a joint training function and a training dataset.

[0053] The first part involves the module responsible for preliminary defect identification in the target architecture model, including YOLOv12 and OverLoCK, which are mainly used to detect potential defects from transmission line images, such as insulator damage and broken conductor strands.

[0054] This involves the second part, which is a module in the target architecture model responsible for correcting the output results of the first part, including LLaVA-MoD. It uses multimodal data (such as meteorological data, equipment ledger information, etc.) to analyze and adjust the preliminary detection results to improve the accuracy and reliability of the detection results.

[0055] This involves a joint training function, which is used to train the target architecture model. By calculating the similarity between the outputs of the first and second parts, it ensures that the output distributions of the two parts are aligned, thereby optimizing the overall performance of the model. The joint training function ensures that the two modules can work collaboratively by minimizing the difference between the outputs of the first and second parts, maximizing the model's collaborative ability and improving the accuracy and stability of defect detection.

[0056] This involves a training dataset, which serves as the dataset for training the target architecture model. It contains various types of sample data, such as labeled images of transmission lines, meteorological data, and equipment ledger information. The training dataset provides the foundation for the model's learning; by learning the features and patterns in this data, the model can better identify and correct defects in transmission lines.

[0057] This involves defect identification, which is the process of detecting potential defects from images of transmission lines. This includes analyzing various features in the images to determine if there are any anomalies, such as broken insulators or broken strands in conductors.

[0058] This includes the output results, which are the detection results generated after processing the transmission line image in the first part, including information such as the location of defects.

[0059] The target architecture model consists of a first part (responsible for initial defect identification) and a second part (responsible for correcting the initial detection results). It is trained by a joint training function and training dataset to ensure that the two parts work together and maximize the collaborative ability of the first and second parts. The first part quickly identifies potential defects in the image, while the second part uses multimodal data to analyze and adjust the initial results, which helps to improve the accuracy and reliability of the detection results.

[0060] S106, input the transmission line image into the first part of the target architecture model to obtain the initial defect detection results corresponding to the target transmission line.

[0061] In step S106 provided in this application, the image of the transmission line is input into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line.

[0062] This includes initial defect detection results, which are the preliminary detection results generated by the first part of the target architecture model after processing transmission line images. These results are preliminary judgments based on image data to identify and locate possible defects in the transmission line.

[0063] By inputting the image of the transmission line into the first part of the target architecture model, the initial defect detection results corresponding to the target transmission line are obtained, which can quickly identify possible defects in the image and provide basic data for subsequent defect analysis and correction.

[0064] S108, input the initial defect detection results and multimodal data into the second part of the target architecture model to correct the initial defect detection results and obtain the target defect detection results corresponding to the target transmission line.

[0065] In step S108 provided in this application, the initial defect detection results and multimodal data are input into the second part of the target architecture model to obtain the target defect detection results corresponding to the target transmission line.

[0066] This involves the target defect detection results, which are the final detection results after being corrected by the second part of the target architecture model. This result integrates the initial defect detection results and multimodal data (such as meteorological data, equipment ledger information, etc.), and can provide more accurate and reliable defect information.

[0067] By inputting the initial defect detection results and multimodal data into the second part of the target architecture model, the defect location, category, and confidence level in the initial detection results can be precisely adjusted through the supplementation and verification of multimodal data, thereby obtaining more accurate and reliable final target defect detection results.

[0068] Through the above steps S102-S108, based on the multimodal data, the multimodal defect features corresponding to the target transmission line are determined; based on the multimodal defect features, similar transmission lines corresponding to the target transmission line are determined from multiple reference transmission lines, wherein the similar transmission lines are reference transmission lines whose difference index between the reference defect features and the multimodal defect features is less than the difference threshold; the reference defect detection set includes multiple reference transmission lines, as well as reference defect features and reference defect detection results corresponding to each of the multiple reference transmission lines; based on the multimodal defect features and the reference defect detection results corresponding to the similar transmission lines, the initial defect detection results are corrected to obtain the target defect detection results corresponding to the target transmission line. The first part of the target architecture model is used for preliminary defect identification, and the second part is used to correct the preliminary results. Since the target architecture model is trained based on the joint training function and the training data set, and the joint training function is used to ensure the alignment of the output distribution of the first part and the second part, the collaboration ability between the models can be maximized according to the target architecture model, thereby improving the accuracy of transmission line defect detection. This solves the technical problem of inaccurate defect detection when detecting defects in transmission lines in related technologies.

[0069] As an optional embodiment, before retrieving the target architecture model, the process includes: determining training target parameters, wherein the training target parameters include a predetermined similarity threshold, when the training dataset includes multiple sample images; inputting any one of the multiple sample images into a first part of the target architecture model to obtain a training output corresponding to the first part for any sample image; determining a training output corresponding to a second part for any sample image based on any sample image and the training output corresponding to the first part for any sample image; determining the output similarity between the training output corresponding to the first part and the training output corresponding to the second part for any sample image based on a joint training function, and determining whether the output similarity is greater than or equal to the predetermined similarity threshold; and stopping training the target architecture model when the determination result is that the output similarity is greater than or equal to the predetermined similarity threshold.

[0070] This embodiment describes the specific steps before retrieving the target architecture model.

[0071] This involves multiple sample images, which form the image dataset used to train the target architecture model. These images include transmission line images from various scenarios, used to train the model to identify and detect defects. These sample images provide the model with rich data, helping it learn the characteristics of different types of defects and improving its generalization ability and detection accuracy. They cover various possible defect types (such as insulator damage, conductor strand breakage, etc.) and different environmental conditions (such as different lighting conditions, weather, etc.), ensuring that the model can accurately identify defects in practical applications.

[0072] This involves a predetermined similarity threshold, a pre-set value used to evaluate whether the similarity between the outputs of the first and second parts of the target architecture model reaches an acceptable level. By setting this threshold, it can be determined whether the model has been trained well enough, i.e., whether the outputs of the first and second parts have reached the expected standard in terms of similarity. When the output similarity is greater than or equal to this threshold, the model is considered to have converged, and training can be stopped to avoid overfitting and wasting computational resources.

[0073] This involves using any sample image, randomly selected from the training dataset, to evaluate the model's performance during training. By inputting different sample images into the model, its performance under various conditions can be comprehensively evaluated, ensuring the model has good recognition capabilities for various types of defects and environmental conditions. The input and output results of each sample image are used to adjust the model's parameters to optimize its overall performance.

[0074] This involves the training output, which is the first part of the target architecture model, and is the result of training by processing any sample image. These results are the initial judgments made by the first part to identify and locate potential defects in any sample image.

[0075] This involves output similarity, which is the degree of similarity between the outputs of the first and second parts of the target architecture model. It is typically calculated using a joint training function. Output similarity is used to evaluate the collaborative effect of the two parts, ensuring they work together to improve the model's detection accuracy. By minimizing the difference between the first and second part outputs, the overall performance of the model can be optimized, improving the accuracy and stability of defect detection.

[0076] By inputting sample images into the first part to obtain preliminary detection results, and then determining the corrected output of the second part based on these results, the similarity between the outputs of the two parts is calculated using a joint training function and compared with a predetermined threshold. Training stops when the similarity reaches or exceeds the threshold. This helps to prevent model overfitting, improve the model's generalization ability and detection performance in practical applications, and ensure that the first and second parts of the target architecture model can work together to achieve the expected detection accuracy and stability.

[0077] As an optional embodiment, before retrieving the target architecture model, the process includes: if the training dataset includes positive sample data and negative sample data, inputting the positive sample data into the second part to obtain positive sample outputs corresponding to the positive sample data, wherein the positive sample data are sample data with a defect severity greater than a defect threshold, and the negative sample data are sample data with a defect severity less than or equal to the defect threshold; inputting the negative sample data into the second part to obtain negative sample outputs corresponding to the negative sample data; determining the alignment loss function corresponding to the second part, wherein the alignment loss function is used to determine the accuracy of defect detection in the second part; and determining the defect detection accuracy corresponding to the second part based on the positive sample output, the negative sample output, and the alignment loss function, until the defect detection accuracy is greater than or equal to the defect detection accuracy threshold, at which point training of the second part is stopped.

[0078] In this embodiment, the specific steps preceding the second part of the target architecture are described.

[0079] This includes positive sample data, which consists of samples in the training dataset with defect severity exceeding a defect threshold. These samples are explicitly labeled as containing significant defects and are used to train the second part to identify and detect obvious defects in transmission lines. Positive sample data provides the second part with known and obvious defect features, helping it learn how to accurately identify and locate these defects. This data typically includes various types of severe defects, such as severe insulator damage and obvious conductor strand breaks, ensuring that the second part can identify multiple types of severe defects.

[0080] This involves negative sample data, which consists of samples in the training dataset with defect severity less than or equal to a defect threshold. These samples are explicitly labeled as containing no significant defects or with minor defects, and are used to train the second part to identify normal or slightly defective states. The negative sample data provides the second part with features indicating normal or slightly defective states, helping it learn how to distinguish between them. This data typically includes transmission line images with no obvious defects or only minor defects, ensuring the second part can correctly identify normal states and reducing false positives.

[0081] This involves positive sample output, which is the output result generated by the second part of the target architecture model when processing positive sample data.

[0082] This involves negative sample output, which is the output of the second part of the target architecture model when processing negative sample data.

[0083] This involves a comparison loss function, which is used to evaluate the accuracy of defect detection in the second part.

[0084] This includes the accuracy of defect detection, which refers to the accuracy of the second part in detecting defects and reflects the second part's ability to identify positive and negative samples.

[0085] This involves a defect detection accuracy threshold, a pre-set value used to determine whether the defect detection accuracy of the second part has reached an acceptable level. By setting this threshold, it can be determined whether the second part has been trained well enough. When the defect detection accuracy is greater than or equal to this threshold, the second part is considered to have converged, and training can be stopped to avoid overfitting and wasting computational resources.

[0086] Before retrieving the target architecture model, positive sample data (defect severity greater than the defect threshold) and negative sample data (defect severity less than or equal to the defect threshold) are input into the second part to obtain corresponding positive and negative sample outputs. The accuracy of the second part in defect detection is evaluated using a comparison loss function. Based on the comparative learning of positive and negative samples, the parameters of the second part can be optimized. When the defect detection accuracy reaches or exceeds the preset defect detection accuracy threshold, training is stopped to ensure that the second part has high accuracy and stability in identifying significant defects and normal or minor defect states, thereby improving the generalization ability and reliability of the second part in practical applications.

[0087] As an optional embodiment, the transmission line image is input into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line, including: determining the defect feature map corresponding to the transmission line image; dividing the defect feature map to obtain multiple sub-maps; determining the sub-defect features corresponding to each of the multiple sub-maps, and multiple associated features corresponding to each of the multiple sub-maps, wherein the multiple associated features represent the defect association relationship between the corresponding sub-map and other sub-maps; and obtaining the initial defect detection result corresponding to the target transmission line based on the sub-defect features and multiple associated features corresponding to each of the multiple sub-maps.

[0088] This embodiment describes the specific steps of inputting a transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line.

[0089] This involves a defect feature map, which is an image representation extracted from transmission line images and contains potential defect features, used to highlight potentially defective parts in the image.

[0090] This involves multiple sub-images, which divide the defect feature map into smaller regions or blocks, each representing a local area of ​​the defect feature map. Dividing the defect feature map into multiple sub-images allows for more detailed analysis of defect features in the image, helping the first part capture local details and contextual information. This division method improves the first part's ability to detect small target defects and also helps handle complex scenes and occlusion issues in the image.

[0091] This involves sub-defect features, which are features extracted from each sub-image and related to potential defects. These features can include information such as texture, shape, and color, used to describe the defect characteristics in the sub-image. Sub-defect features reflect the potential defects in each sub-image identified in the first part, providing basic information for subsequent defect detection and classification.

[0092] This involves multiple association features, which describe the relationships between subgraphs. These features represent the defect relationships between subgraphs, such as whether defects in adjacent subgraphs belong to the same defect object, or whether there is some spatial association. Association features help the first part understand the contextual relationships between subgraphs, thus more accurately determining the integrity and continuity of defects. By considering the association features between subgraphs, the first part can better handle occlusion and fragmentation problems in complex scenes, improving the accuracy and robustness of defect detection.

[0093] By identifying the defect feature map corresponding to the transmission line image and dividing it into multiple sub-images, the first part extracts sub-defect features and related features from each sub-image. This allows the first part to capture local details and contextual information in the image more meticulously, thereby more accurately identifying and locating potential defects. Sub-defect features provide defect information for each local region, while related features help the first part understand the relationships between these local regions, ensuring accurate judgments on the integrity and continuity of defects, thus improving the first part's ability to detect small target defects.

[0094] As an optional embodiment, an initial defect detection result corresponding to the target transmission line is obtained based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs, including: determining an enhanced feature map corresponding to the defect feature map based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs; and determining the initial defect detection result corresponding to the target transmission line based on the enhanced feature map.

[0095] This embodiment describes the specific steps for obtaining the initial defect detection result corresponding to the target transmission line based on the sub-defect features and multiple associated features corresponding to multiple sub-graphs.

[0096] This involves enhanced feature maps, which are feature maps obtained by fusing and enhancing the sub-defect features and associated features of multiple sub-maps, in order to more prominently and accurately represent the features of potential defects.

[0097] By determining the enhanced feature map based on the sub-defect features corresponding to multiple sub-maps and multiple associated features, the features of potential defects can be represented more prominently and accurately. The initial defect detection result is then determined based on the enhanced feature map, which improves the detection accuracy of small target defects in complex scenarios and provides more accurate and reliable initial detection results for subsequent defect analysis and processing.

[0098] As an optional embodiment, when the second part includes a reference defect detection set, the initial defect detection results and multimodal data are input into the second part of the target architecture model to correct the initial defect detection results and obtain the target defect detection results corresponding to the target transmission line. This includes: determining multimodal defect features corresponding to the target transmission line based on the multimodal data; determining similar transmission lines corresponding to the target transmission line from multiple reference transmission lines based on the multimodal defect features, wherein similar transmission lines are reference transmission lines whose difference index between the reference defect features and the multimodal defect features is less than a difference threshold; wherein the reference defect detection set includes multiple reference transmission lines, and reference defect features and reference defect detection results corresponding to the multiple reference transmission lines respectively; and correcting the initial defect detection results based on the multimodal defect features and the reference defect detection results corresponding to the similar transmission lines to obtain the target defect detection results corresponding to the target transmission line.

[0099] In this embodiment, specific steps are described to correct the initial defect detection results based on multimodal data and the second part of the target architecture model when the second part includes a reference defect detection set, so as to obtain the target defect detection results corresponding to the target transmission line.

[0100] This involves multimodal defect features, which are comprehensive features extracted by integrating various types of data (such as images, meteorological data, equipment ledger information, etc.) to more comprehensively describe the nature and status of transmission line defects.

[0101] This involves multiple reference transmission lines, which are included in the reference defect detection set and contain transmission lines with known defect characteristics and detection results, for the learning and comparison in the second part. These multiple reference transmission lines can provide transmission line characteristics and defect patterns under various different scenarios, helping the second part to better understand and identify various possible defect situations.

[0102] This involves similar transmission lines, which are those in the reference defect detection set that have relatively small differences in multimodal defect characteristics compared to the target transmission line. Specifically, the difference index between the reference defect characteristics of the similar transmission line and the multimodal defect characteristics of the target transmission line is less than a set difference threshold. The reference defect detection results of the similar transmission line can serve as a basis for correcting the initial defect detection results of the target transmission line, helping to improve the accuracy and reliability of the detection.

[0103] This involves a difference index, a quantitative indicator used to measure the difference between the multimodal defect characteristics of a target transmission line and the reference defect characteristics of a reference transmission line. The difference index calculates the distance or similarity between features to determine which reference transmission lines are most similar to the target transmission line, thus selecting appropriate reference defect detection results for correction.

[0104] This involves a difference threshold, a pre-set value used to determine whether the difference between the target transmission line and a reference transmission line is small enough to be considered similar. The difference threshold controls the selection criteria for similar transmission lines. Only when the difference index is less than the difference threshold will the reference transmission line be selected as a similar transmission line, and its corresponding reference defect detection results will be used as a reference to correct the initial defect detection results of the target transmission line.

[0105] This involves reference defect features, which are multimodal defect features determined based on multimodal data from a reference transmission line. The reference defect features provide known defect patterns and characteristics for comparison with the multimodal defect features of the target transmission line, thereby aiding in the identification and correction of defects in the target transmission line.

[0106] This involves a reference defect detection set, which is a dataset containing multiple reference transmission lines and their corresponding reference defect features and detection results. The reference defect detection set provides rich information on known defects. By comparing the features of the target transmission line with those in the reference defect detection set, the second part learns how to more accurately correct the initial defect detection results.

[0107] This includes reference defect detection results, which are the defect detection results of a reference transmission line.

[0108] Through the above steps, multimodal defect features provide a more comprehensive description of defects, while reference results from similar transmission lines provide a reliable basis for correction, thereby ensuring that the final target defect detection results are closer to the actual situation, reducing false detections and missed detections, and improving the overall performance of transmission line defect detection.

[0109] As an optional embodiment, after inputting the transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line, the method further includes: if the target transmission line is a predetermined transmission line and the initial defect detection result includes defect region parameters, determining the target scene parameters corresponding to the target transmission line, wherein the predetermined transmission line is the transmission line under the predetermined scene; determining the fused image based on the target scene parameters, defect region parameters, and transmission line image; and performing incremental training on the first part based on the fused image to obtain the incrementally trained first part.

[0110] This embodiment describes the specific steps after inputting the image of the transmission line into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line.

[0111] This involves predetermined transmission lines, which are transmission lines operating under predetermined scenarios or conditions. These lines typically have predetermined characteristics or requirements, necessitating specialized testing and analysis.

[0112] This includes defect region parameters, which are used to describe the location, extent, and confidence level of defects.

[0113] This involves target scene parameters, which are parameters in a specific scenario related to the predetermined transmission line. These parameters describe the environment and conditions in which the transmission line is located, such as lighting conditions, weather conditions, and terrain features (such as altitude).

[0114] This involves predefined scenarios, which are predefined scenarios or conditions that typically have specific characteristics or requirements and require specialized detection and analysis.

[0115] This involves fused images, which are generated by combining target scene parameters, defect area parameters, and transmission line images. These images integrate multiple pieces of information, providing a more comprehensive description of the defect area.

[0116] This involves incremental training, which is the process of further training the first part to gradually optimize it and improve its performance. Incremental training allows the first part to be adjusted and optimized based on new data or information, improving its adaptability to specific scenarios or circuits and ensuring that the first part can continuously improve and enhance detection accuracy in practical applications.

[0117] Based on the above embodiments and optional embodiments, an optional implementation method is provided, which is described in detail below.

[0118] In related technologies, with the continuous growth of electricity demand and the increasing scale of power grids, the operation and maintenance management of transmission lines has become increasingly important. Timely detection and handling of defects in transmission lines, such as insulator damage, broken conductor strands, and tower corrosion, can effectively prevent faults, reduce power outage time, and improve power supply reliability. However, in related technologies, there are technical problems with inaccurate defect detection when performing defect detection on transmission lines.

[0119] Specifically, in transmission line defect detection, most transmission line defect detection still relies on a combination of traditional manual inspection and single-model detection. Manual inspection requires maintenance personnel to check the line section by section, which is extremely inefficient. In complex terrain areas such as mountainous areas and cross-river areas, manual inspection is not only time-consuming but also poses significant safety risks. Taking a 500-kilometer-long transmission line as an example, it would take about 30 days to complete a full manual inspection, during which time factors such as personnel fatigue and environmental interference can easily lead to missed detections. In automated detection technology, existing single-model detection methods have many shortcomings. Common target detection models, such as YOLOv5, can achieve a certain degree of automated detection, but their detection accuracy is severely insufficient when dealing with small target defects (such as micro-cracks in insulators and minor strand breaks in conductors). Because the network structure of such models has limited ability to extract features from small targets, in practical applications, the missed detection rate for small target defects with a pixel ratio of less than 0.5% is as high as 35% or more. Furthermore, when faced with complex occlusion scenarios (such as pole components obscured by tree branches or billboards), the detection performance of a single model drops significantly, making it difficult to accurately identify targets, with a false detection rate exceeding 25%.

[0120] Furthermore, existing technologies lack the ability to effectively integrate multimodal data and perform knowledge reasoning. Transmission line inspection data includes multimodal information such as images, meteorological data, and equipment records, but existing methods often process each modality independently, failing to uncover potential correlations between the data. For example, they cannot combine humidity and wind speed information from meteorological data with equipment images to analyze the impact of environmental factors on transmission equipment. Simultaneously, in the defect analysis phase, existing technologies rely on manual experience and lack the ability to automatically retrieve and apply relevant domain knowledge, making it difficult to accurately analyze the causes of defects and assess risks, thus failing to provide comprehensive and accurate basis for operation and maintenance decisions.

[0121] In terms of model deployment, existing deep learning models are computationally intensive and have many parameters, making it difficult to run efficiently on edge devices. For example, Faster R-CNN has a computational cost of over 2000 GFLOPs, and when running on edge computing devices, the single-frame processing latency exceeds 800ms, which cannot meet the requirements for real-time analysis of drone inspection video streams and limits the practical application and promotion of automated inspection technology.

[0122] Furthermore, existing technologies have poor adaptability to different environments and equipment models. Transmission lines are widely distributed with significant environmental variations; lighting conditions and equipment models differ across regions. Existing models, when trained, rely on a single data source and lack coverage of diverse scenarios, resulting in weak model generalization ability. When faced with new environments and equipment models, model performance significantly degrades, failing to accurately detect defects and severely impacting inspection effectiveness and the safe operation of the power grid.

[0123] There is currently no effective solution to the above problems.

[0124] In view of this, the optional embodiments of the present invention provide a method for detecting defects in transmission lines, with specific objectives including: high-precision defect detection, knowledge-enhanced semantic reasoning, improved cross-domain generalization ability, and lightweight edge deployment. It can effectively solve the technical problem of inaccurate defect detection when performing defect detection on transmission lines in related technologies.

[0125] in:

[0126] For high-precision defect detection: By combining YOLOv12 and OverLoCK dynamic convolution, the detection accuracy of small targets and low-contrast defects is improved (mAP@0.5 is improved to over 92%).

[0127] For knowledge-enhanced semantic reasoning: by utilizing the sparse expert network of LLaVA-MoD, inspection images and industry knowledge are integrated to achieve automatic classification of defect types (such as "glass insulator damage → spontaneous explosion risk level III") and generation of handling suggestions;

[0128] Regarding the improvement of cross-domain generalization ability: Through UNIGEN's unlabeled data generation and pseudo-relabeling technology, the model's generalization ability to different lighting (backlight / strong light) and equipment models (such as towers of different voltage levels) is enhanced, and the domain transfer error is reduced (such as the false detection rate caused by the difference in inspection data in different areas is reduced by 25%).

[0129] For lightweight edge deployment: Through model compression and shared backbone design, the computational load is reduced by more than 50%, enabling real-time inference on edge computing devices (single frame processing latency <150ms).

[0130] Figure 2 This is a flowchart of a transmission line defect detection method in an optional embodiment of the present invention, such as... Figure 2 As shown, by integrating YOLOv12, OverLoCK, LLaVA-MoD and UNIGEN, the collaboration of large and small models for power transmission line inspection is achieved, thereby enabling the detection of defects in power transmission lines. Specifically, this is achieved through four parts: a lightweight visual inspection module, a knowledge-enhanced reasoning module, a cross-domain data generation module, and a joint optimization module, which are described in detail below.

[0131] S1. Acquire multimodal data corresponding to the target transmission line, wherein the multimodal data includes images of the transmission line;

[0132] S2. Retrieve the target architecture model, which includes a first part and a second part. The target architecture model is trained based on a joint training function and a training data set. The joint training function is used to determine the similarity between the output of the first part and the output of the second part. The first part is used for defect identification, and the second part is used to correct the output of the first part.

[0133] The target architecture model includes a lightweight visual detection module (i.e., the first part) and a knowledge-enhanced reasoning module (i.e., the second part). The lightweight visual detection module includes YOLOv12 and OverLoCK, while the knowledge-enhanced reasoning module includes LLaVA-MoD.

[0134] The lightweight visual inspection module is designed as follows:

[0135] (1) YOLOv12 basic detection architecture.

[0136] The YOLOv12 basic detection architecture includes a region attention mechanism and efficient residual aggregation.

[0137] For region attention mechanisms:

[0138] To address the issue of insufficient feature extraction for small targets (such as broken conductor strands and loose bolts) in power transmission inspection images (i.e., power transmission line images), a region attention mechanism is introduced into the Backbone of YOLOv12, specifically implemented as follows:

[0139] The input feature map (same as the defect feature map above) is divided into... There are three non-overlapping regions, and sparse attention is used to calculate feature associations within each region, as follows:

[0140]

[0141] in:

[0142] This is a functional representation of the region attention mechanism, used to calculate the feature associations of the input feature map in different regions, thereby enhancing the feature representation of small targets;

[0143] For the feature dimension in the attention mechanism;

[0144] Input feature map;

[0145] The first Query, key, and value vectors for each region;

[0146] Key vector The transpose of .

[0147] By partitioning regions, the computational complexity is reduced from that of global attention. Down to This significantly improves the feature response of small targets (pixel percentage <0.5%). Experiments show that this module increases the feature activation value for microcrack detection by 40%. Let represent the computational complexity of the global attention mechanism, indicating that the computational complexity is proportional to the square of the number of pixels in the feature map. The computational complexity is calculated after partitioning the region.

[0148] For efficient residual aggregation:

[0149] A residual efficient layer aggregation network (R-ELAN) is designed in the detection head to enhance the fusion of shallow detailed features and deep semantic features through cross-layer residual connections, as shown in the following formula:

[0150]

[0151] in:

[0152] For output features;

[0153] These are feature maps at different levels (e.g., C3, C4, C5).

[0154] Input features;

[0155] This is the residual scaling factor.

[0156] Specifically, for layers C3, C4, and C5, the following is indicated:

[0157] C3 layer: A shallower feature map containing more detailed information, suitable for capturing details of small targets.

[0158] C4 layer: The feature map of the middle layer contains both detailed information and semantic information, and is suitable for capturing medium-sized targets.

[0159] C5 layer: A deeper feature map containing more semantic information, suitable for capturing the semantic information of large targets.

[0160] This structure effectively alleviates the gradient vanishing problem in deep networks, improving the mAP@0.5 of insulator string detection from 82.3% to 87.1%.

[0161] (2) OverLoCK level feature enhancement.

[0162] Based on the three-branch architecture of OverLoCK, a hierarchical feature decomposition network is constructed, including: Base-Net, Overview-Net, and Focus-Net.

[0163] For Base-Net: it consists of 3 layers of 3×3 convolutions with a stride of 2, extracting basic visual features (such as tower structure and insulator outline), and the output feature size is [missing information]. ,in, The height of the input image. The width of the input image;

[0164] For Overview-Net: a coarse-grained semantic overview (such as "suspected discharge area") is generated through global average pooling, and then... Convolution compression to As a global guidance signal;

[0165] For Focus-Net: Guided by overview features, the Region of Interest (ROI) is refined and modeled using Context-Mixing dynamic convolution (ContMix), as follows:

[0166]

[0167] in:

[0168] This is the result of the convolution;

[0169] r represents 9 region centers (using a 3x3 grid for sampling);

[0170] For pixels Affinity matrix with region center r (calculated by coordinate distance and feature similarity);

[0171] Use the Sigmoid activation function;

[0172] The kernel is a dynamically generated 3x3 convolution kernel.

[0173] This module reduced the missed detection rate of obscured tower components from 32% to 13%.

[0174] Prior to S2, the specific design of the knowledge-enhanced reasoning module (LLaVA-MoD) also includes a sparse expert distillation mechanism, which specifically includes imitation distillation and preference distillation.

[0175] Specifically, for imitation distillation, the process includes: determining training target parameters, wherein the training target parameters include a predetermined similarity threshold, when the training dataset includes multiple sample images; inputting any one of the multiple sample images into the first part of the target architecture model to obtain the training output corresponding to the first part under any sample image; determining the training output corresponding to the second part under any sample image based on any sample image and the training output corresponding to the first part under any sample image; determining the output similarity between the training output corresponding to the first part and the training output corresponding to the second part under any sample image based on the joint training function, and determining whether the output similarity is greater than or equal to the predetermined similarity threshold; and stopping training the target architecture model when the determination result is that the output similarity is greater than or equal to the predetermined similarity threshold.

[0176] To address the knowledge transfer problem between the large model (LLaVA-MoD) and the small model (YOLOv12), a method is designed to mimic the distillation loss to align the output distributions of both models. The joint training function is as follows:

[0177]

[0178] in:

[0179] This is the loss function in Knowledge Distillation, used to measure the difference in output distribution between the first part of YOLOv12 and the second part of LLaVA-MoD (similar to the output similarity mentioned above).

[0180] Let represent the expectation (or average) of all samples (x,y) in dataset D;

[0181] D is a labeled dataset containing inspection images (same as the multiple sample images mentioned above).

[0182] This represents the detection probability distribution for the small model.

[0183] This represents the semantic reasoning distribution of a large model.

[0184] Knowledge distillation is achieved by minimizing KL divergence.

[0185] Furthermore, preference distillation includes: when the training dataset includes positive and negative sample data, inputting positive sample data into the second part to obtain positive sample outputs corresponding to the positive sample data; inputting negative sample data into the second part to obtain negative sample outputs corresponding to the negative sample data; determining the alignment loss function corresponding to the second part, wherein the alignment loss function is used to determine the accuracy of defect detection in the second part; determining the defect detection accuracy corresponding to the second part based on the positive sample outputs, negative sample outputs, and alignment loss function, until the defect detection accuracy is greater than or equal to the defect detection accuracy threshold, at which point training of the second part is stopped.

[0186] For example, for defect level discrimination (such as Level I minor defect, Level III severe defect), a preference contrast loss is introduced, and the contrast loss function is as follows:

[0187]

[0188] in:

[0189] The preference contrast loss function (same as the contrast loss function above) is used to measure the degree of preference of the model for samples of different categories;

[0190] For activation functions;

[0191] It is a balancing factor used to adjust the weights between positive and negative samples;

[0192] This represents the model's predicted probability distribution for positive samples.

[0193] This represents the model's predicted probability distribution for negative samples.

[0194] This represents the predicted probability distribution for positive samples in the second part.

[0195] This represents the predicted probability distribution for negative samples in the second part.

[0196] Positive samples (same as the positive sample data above, such as severe defects);

[0197] For negative samples (same as the negative sample data above, such as minor defects).

[0198] Comparative learning improved the accuracy of the large model in judging defect levels by 19%.

[0199] S3. Input the transmission line image into the first part of the target architecture model to obtain the initial defect detection results corresponding to the target transmission line;

[0200] Specifically, S3 also includes: determining the defect feature map corresponding to the transmission line image, and providing a region attention mechanism introduced in the YOLOv12 Backbone to divide the defect feature map into multiple sub-maps, and determining the sub-defect features corresponding to each of the multiple sub-maps, as well as multiple associated features corresponding to each of the multiple sub-maps. Then, based on the sub-defect features and multiple associated features corresponding to each of the multiple sub-maps, through the aforementioned OverLoCK hierarchical feature enhancement, an enhanced feature map corresponding to the defect feature map is determined; based on the enhanced feature map, the initial defect detection result corresponding to the target transmission line is determined.

[0201] In addition, after S3, the first part is incrementally trained by the cross-domain data generation module (UNIGEN). That is, when the target transmission line is a predetermined transmission line and the initial defect detection result includes defect area parameters, the target scene parameters corresponding to the target transmission line are determined, where the predetermined transmission line is the transmission line under the predetermined scene; based on the target scene parameters, defect area parameters, and transmission line image, a fused image is determined; based on the fused image, the first part is incrementally trained to obtain the first part after incremental training.

[0202] The cross-domain data generation module (UNIGEN) is designed as follows:

[0203] (1) Generation of unlabeled data and pseudo-relabeling.

[0204] To address the issue of sample scarcity in specific scenarios, including new equipment (such as ultra-high voltage power transmission towers) or extreme environments (such as heavy rain or backlighting), UNIGEN's text-image generation capabilities are used to synthesize cross-domain data, as follows:

[0205]

[0206] in:

[0207] The synthesized image generated by UNIGEN (same as the fused image above) is used to supplement scenes with scarce samples;

[0208] This is a text-image generation model that generates corresponding images based on input text prompts.

[0209] The text prompts input into UNIGEN, also known as defect region parameters, describe the content of the target image.

[0210] After generating the image, soft labels are generated using pseudo-re-labeling technology, as follows:

[0211]

[0212] in:

[0213] These are soft tags generated using pseudo-re-labeling technology.

[0214] The Softmax function converts the input vector into a probability distribution.

[0215] To synthesize images The predicted probability distribution;

[0216] For the first A composite image;

[0217] These are general annotation prompts;

[0218] This refers to the temperature parameter.

[0219] (2) Domain transfer enhancement training.

[0220] Generative Adversarial Networks (GANs) are employed to further enhance cross-domain generalization capabilities through domain conditional vectors. Control the attributes of the generated samples, such as lighting and device model, as follows:

[0221]

[0222] in:

[0223] Synthetic images generated by Generative Adversarial Networks (GANs);

[0224] For generator functions;

[0225] It is Gaussian noise;

[0226] For target domain distribution (such as the distribution of equipment models in different regions).

[0227] Introducing domain adversarial loss during training This forces the model to ignore inter-domain differences and focus on defect features, as follows:

[0228]

[0229] in:

[0230] For the expected (or average) value;

[0231] For domain discriminator;

[0232] These are the domain labels predicted by the model.

[0233] Experiments show that this mechanism reduces the cross-domain false positive rate from 41% to 23%.

[0234] S4. Input the initial defect detection results and multimodal data into the second part of the target architecture model to correct the initial defect detection results and obtain the target defect detection results corresponding to the target transmission line.

[0235] Specifically, S4 further includes: determining multimodal defect features corresponding to the target transmission line based on multimodal data; determining similar transmission lines corresponding to the target transmission line from multiple reference transmission lines based on the multimodal defect features, wherein similar transmission lines are reference transmission lines whose difference index between the reference defect features and the multimodal defect features is less than a difference threshold, wherein the reference defect detection set includes multiple reference transmission lines, and reference defect features and reference defect detection results corresponding to each of the multiple reference transmission lines; and correcting the initial defect detection results based on the multimodal defect features and the reference defect detection results corresponding to the similar transmission lines to obtain the target defect detection result corresponding to the target transmission line. This can be implemented through domain knowledge injection and reasoning, as specifically designed below:

[0236] A knowledge graph (KG, similar to the aforementioned reference defect detection set) is constructed for the transmission line domain. This graph includes entities and their relationships (e.g., "ice thickness > 10mm → insulator flashover risk ↑") such as equipment type (e.g., insulators, towers), defect modes (e.g., damage, discharge), and environmental risks (e.g., icing, lightning strikes). The knowledge embeddings are indexed using the FAISS vector database to enable real-time retrieval, as follows:

[0237]

[0238] in:

[0239] This is the knowledge embedding vector obtained by retrieving the knowledge graph (KG) in the transmission line domain;

[0240] This is a retrieval function used to retrieve the knowledge entries most relevant to the input feature vector from the knowledge graph (KG) of the transmission line domain;

[0241] This is a fusion vector of visual features and meteorological data;

[0242] To retrieve the number of nearest neighbor knowledge entries.

[0243] The search results are concatenated with the detection features and then input into the Transformer decoder to generate a natural language explanation (such as "Currently, glass insulator damage has been detected. According to the transmission line status evaluation standard, the risk level is determined to be Level III. Emergency handling within 24 hours is recommended").

[0244] In addition, the initial defect detection results and multimodal data are input into the second part of the target architecture model to correct the initial defect detection results. After obtaining the target defect detection results corresponding to the target transmission line, a joint optimization module is also set up, the specific design of which is as follows:

[0245] (1) Shared feature backbone design.

[0246] The detection branch and the inference branch share the parameters of the first 12 convolutional layers (i.e., the first 12 layers of the backbone), and the output feature map sizes are as follows: Detection branch (Used for defect localization); reasoning branch (For semantic feature extraction), through parameter sharing, computational redundancy is reduced by 40%, FLOPs are reduced from 1800G to 1080G, while maintaining a detection accuracy loss of <1.5%;

[0247] (2) Multi-task multi-joint loss function.

[0248] Design a multi-task loss function that includes detection, distillation, and classification. ,as follows:

[0249]

[0250] in:

[0251] Includes classification loss (Focal Loss, i.e.) ) and regression loss (DIoU Loss, i.e. );

[0252] The knowledge distillation loss function;

[0253] For the preference contrast loss function;

[0254] The cross-entropy loss function;

[0255] , , These are the weighting coefficients.

[0256] A balance between detection and inference tasks was achieved through hyperparameter tuning. After joint optimization, the detection mAP@0.5 improved from 89.2% to 90.4%, and the pose estimation error (ADD-S) decreased from 78.5 mm to 74.3 mm.

[0257] The following description, with specific examples, will further illustrate this point.

[0258] S1. Data Acquisition and Preprocessing.

[0259] Specifically, S1 includes:

[0260] S11. Synchronous acquisition of multi-source data (same as the multimodal data mentioned above).

[0261] A drone equipped with a 4K camera and LiDAR was used to fly along the power transmission line at a speed of 8 meters per second, maintaining a flight altitude of 10-15 meters above the ground wire. While acquiring images, meteorological data (including temperature, humidity, wind speed, and light intensity) and equipment log information (such as tower type and commissioning time) were recorded simultaneously. The various data were aligned using timestamps to form a multimodal dataset (similar to the multimodal data mentioned above).

[0262] S12, Image Enhancement and Calibration.

[0263] For the acquired images (same as the transmission line images mentioned above), adaptive median filtering (5×5 window) is first used to remove salt-and-pepper noise (i.e., impulse noise); then homomorphic filtering (cutoff frequency 0.3) is used to eliminate uneven illumination and improve image quality. Camera intrinsic parameters are obtained through camera calibration, and extrinsic parameters are obtained through hand-eye calibration, achieving centimeter-level coordinate calibration to ensure accurate image positioning.

[0264] S13. Region of Interest (ROI) Extraction.

[0265] The image is cropped, retaining only the ground line and a 2-meter radius around it, compressing the data volume to 30% of the original image. This significantly reduces the subsequent computational load while preserving the effective information.

[0266] S2, edge deployment and real-time detection.

[0267] S21, Lightweight Model Deployment.

[0268] Deploy the YOLOv12+OverLoCK model on edge devices. Scale the input image (same as the transmission line image mentioned above) to 640×640 and use 16-bit floating-point (FP16) precision to accelerate inference, keeping the processing time per frame within 120ms to meet real-time detection requirements.

[0269] S22, Multi-scale feature fusion detection.

[0270] The YOLOv12 backbone network (CSPDarknet) outputs feature maps at different scales in layers P3, P4, and P5 (similar to the sub-defect features corresponding to multiple sub-maps, and multiple associated features corresponding to multiple sub-maps). OverLoCK's three-branch network performs deep decomposition and refined modeling of these features to obtain enhanced feature maps. The detection head ultimately outputs the initial defect detection results, including the defect location, category (e.g., insulator damage, conductor strand breakage), and confidence information. P3, P4, and P5 are feature maps at different scales output by the YOLOv12 backbone network (CSPDarknet). These feature maps represent feature extraction results at different scales and are used for subsequent multi-scale target detection.

[0271] S3, Cloud-based Knowledge Reasoning and Analysis.

[0272] S31. Feature extraction and transmission.

[0273] The 1024-dimensional feature vector of the detection area is extracted and transmitted to the cloud server via a 5G network. The transmission time is less than 50ms, ensuring fast data upload.

[0274] S32. Knowledge Graph Retrieval and Risk Assessment.

[0275] The LLaVA-MoD model in the cloud uses FAISS vector retrieval to match the knowledge graph of the power transmission field. Combining meteorological data and equipment ledgers, it analyzes the causes of defects and generates risk levels (such as Level III) and handling suggestions (such as "arrange replacement within 24 hours") based on the power transmission line status evaluation standards.

[0276] S4, Cross-domain data generation and model optimization.

[0277] S41, Synthetic Data Generation.

[0278] For special scenarios such as high altitude and severe weather, UNIGEN is used to generate synthetic images. After inputting prompts (such as "altitude 4000 meters, 500kV line, insulator icing, light intensity 300 lux"), 10,000 synthetic images with pseudo-labels are generated. Samples with a confidence level below 70% are filtered out, ultimately retaining 9,200 valid images.

[0279] S42, Model Iterative Update.

[0280] Incremental training is performed every two weeks, mixing hard samples (confidence levels between 50% and 70%) from edge devices with synthetic data. The batch size is set to 64, the learning rate to 0.0001, and training is conducted for 50 epochs. After training, the updated model is pushed to edge devices via OTA (Over-The-Air) technology for continuous model optimization.

[0281] The above optional implementation methods can achieve at least the following beneficial effects:

[0282] (1) Compared with related technologies, the present invention performs preliminary defect identification through the first part of the target architecture model, and then corrects the preliminary results through the second part. Since the target architecture model is trained based on the joint training function and the training data set, and the joint training function is used to ensure that the output distribution of the first part and the second part are aligned, the collaboration ability between the models can be maximized according to the target architecture model, thereby improving the accuracy of defect detection of transmission lines. This solves the technical problem of inaccurate defect detection when performing defect detection on transmission lines in related technologies.

[0283] (2) Compared with related technologies, the present invention obtains preliminary detection results by inputting sample images into the first part, and then determines the corrected output of the second part based on these results. The similarity between the outputs of the two parts is calculated using a joint training function and compared with a predetermined threshold. Training is stopped when the similarity reaches or exceeds the threshold. This helps to prevent model overfitting, improve the generalization ability of the model and the detection performance in practical applications, so as to ensure that the first part and the second part of the target architecture model can work together to achieve the expected detection accuracy and stability.

[0284] (3) Compared with related technologies, the present invention determines the defect feature map corresponding to the transmission line image and divides it into multiple sub-images. It extracts the sub-defect features and related features of each sub-image, enabling the first part to capture local details and contextual information in the image more meticulously, thereby more accurately identifying and locating potential defects. The sub-defect features provide defect information for each local area, while the related features help the first part understand the relationship between these local areas, ensuring accurate judgment of the integrity and continuity of the defect, thereby helping to improve the first part's ability to detect small target defects.

[0285] (4) Compared with related technologies, the present invention provides a more comprehensive description of defects through multimodal defect features, while the reference results of similar transmission lines provide a reliable basis for correction, thereby ensuring that the final target defect detection results are closer to the actual situation, reducing false detections and missed detections, and improving the overall performance of transmission line defect detection.

[0286] (5) Compared with related technologies, this invention, through innovative integration of YOLOv12, OverLoCK, LLaVA-MoD and UNIGEN, demonstrates significant advantages in the field of power transmission inspection, as follows:

[0287] By combining YOLOv12 with OverLoCK, detection performance in small targets and complex occlusion scenarios is significantly improved. The region attention mechanism (A2 module) divides the input feature map into multiple regions for sparse attention computation, improving the feature response of small targets by 40%. On the COCO-TransmissionLine dataset, the mAP@0.5 for small target detection (area <32² pixels) is improved from 68.5% to 77.8%. OverLoCK's depth decomposition strategy and Context-Mixing dynamic convolution can accurately capture tower components occluded by vegetation, improving the detection accuracy in complex occlusion scenarios from 72.1% to 89.2%, greatly reducing false negatives and missed detections.

[0288] By introducing a sparse expert distillation mechanism and domain knowledge graph through the LLaVA-MoD module, a leap from simple image detection to intelligent analysis and decision-making is achieved. Imitation distillation and preference distillation enable a defect classification accuracy of 91.5%, far exceeding the 67.3% of traditional methods. Combined with the knowledge graph, the system can automatically associate with transmission line condition evaluation standards, assess the risk level of detected defects, and generate handling suggestions. The risk interpretation generation compliance rate reaches 95%, providing a reliable basis for operation and maintenance management.

[0289] By utilizing the UNIGEN module, based on generative adversarial networks and pseudo-relabeling techniques, the challenges of data scarcity and cross-domain generalization are effectively addressed. Diverse cross-domain samples can be synthesized to address differences in lighting conditions and device models, enabling the model to maintain stable performance across various scenarios. In different lighting scenarios, the false detection rate is reduced from 41% (backlight) and 37% (strong light) of traditional methods to 23% and 19%, respectively; data migration error between devices in different regions is reduced by 27%, significantly enhancing the model's adaptability to complex environments.

[0290] By adopting a shared feature backbone design and model compression strategy, a balance between computational efficiency and accuracy is achieved. The total number of model parameters is reduced from 78M to 42M, the computational load is reduced by 46%, and FLOPs are reduced from 1800G to 1080G. On edge devices, the single-frame processing latency is <150ms, achieving an inference speed of 25FPS, meeting the real-time analysis requirements of UAV inspection video streams, significantly improving inspection efficiency, and reducing operation and maintenance costs.

[0291] By constructing a complete closed loop of "detection-inference-data augmentation-model optimization", collecting difficult samples from the edge and cross-domain data generated by UNIGEN, and regularly performing joint optimization training to continuously update model parameters, this continuous evolution mechanism enables the model to adapt to the upgrading of transmission line equipment, environmental changes and other situations, maintain high performance in the long term, and provide a lasting guarantee for the safe and stable operation of transmission lines.

[0292] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0293] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0294] Example 2

[0295] According to embodiments of the present invention, an apparatus for implementing the above-described method for detecting defects in transmission lines is also provided. Figure 3 This is a structural block diagram of a power transmission line defect detection device according to an embodiment of the present invention, such as... Figure 3 As shown, the device includes: an acquisition module 302, a retrieval module 304, a first determination module 306, and a second determination module 308. The device will be described in detail below.

[0296] The acquisition module 302 is used to acquire multimodal data corresponding to the target transmission line, wherein the multimodal data includes transmission line images; the retrieval module 304, connected to the acquisition module 302, is used to retrieve the target architecture model, wherein the target architecture model includes a first part and a second part, the target architecture model is trained based on a joint training function and a training dataset, the joint training function is used to determine the similarity between the output of the first part and the output of the second part, the first part is used for defect identification, and the second part is used to correct the output of the first part; the first determination module 306, connected to the retrieval module 304, is used to input the transmission line image into the first part of the target architecture model to obtain an initial defect detection result corresponding to the target transmission line; the second determination module 308, connected to the first determination module 306, is used to input the initial defect detection result and the multimodal data into the second part of the target architecture model to correct the initial defect detection result and obtain a target defect detection result corresponding to the target transmission line.

[0297] It should be noted here that the above-mentioned acquisition module 302, retrieval module 304, first determination module 306 and second determination module 308 correspond to steps S102 to S108 in the method for detecting defects in transmission lines. The multiple modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in the above embodiment 1.

[0298] Example 3

[0299] According to another aspect of the present invention, an electronic device is also provided, comprising: a processor; and a memory for storing processor-executable instructions, wherein the processor is configured to execute instructions to implement the transmission line defect detection method of any of the above embodiments.

[0300] Example 4

[0301] According to another aspect of the present invention, a computer-readable storage medium is also provided, which, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform the transmission line defect detection method described above.

[0302] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0303] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0304] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0305] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0306] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0307] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0308] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for detecting defects in power transmission lines, characterized in that, include: Acquire multimodal data corresponding to the target transmission line, wherein the multimodal data includes images of the transmission line; The target architecture model is retrieved, wherein the target architecture model includes a first part and a second part. The target architecture model is trained based on a joint training function and a training dataset. The joint training function is used to determine the similarity between the output of the first part and the output of the second part, and to minimize the difference between the outputs of the first part and the second part. The first part is used for defect identification, and the second part is used to correct the output of the first part. The first part includes YOLOv12 and OverLoCK, and the second part includes LLaVA-MoD. The image of the transmission line is input into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line; The initial defect detection results and the multimodal data are input into the second part of the target architecture model to correct the initial defect detection results and obtain the target defect detection results corresponding to the target transmission line.

2. The method according to claim 1, characterized in that, Before retrieving the target architecture model, the following steps are included: When the training dataset includes multiple sample images, training target parameters are determined, wherein the training target parameters include a predetermined similarity threshold; Input any one of the multiple sample images into the first part of the target architecture model to obtain the training output corresponding to the first part under any one sample image; Based on any sample image and the training output corresponding to the first part under any sample image, determine the training output corresponding to the second part under any sample image; Based on the joint training function, determine the output similarity between the training output corresponding to the first part and the training output corresponding to the second part under any sample image, and determine whether the output similarity is greater than or equal to the determination result of the predetermined similarity threshold; Training of the target architecture model will stop when the determined result is that the output similarity is greater than or equal to the predetermined similarity threshold.

3. The method according to claim 1, characterized in that, Before retrieving the target architecture model, the following steps are included: When the training data set includes positive sample data and negative sample data, the positive sample data is input into the second part to obtain the positive sample output corresponding to the positive sample data. The positive sample data is sample data with a defect degree greater than the defect threshold, and the negative sample data is sample data with a defect degree less than or equal to the defect threshold. The negative sample data is input into the second part to obtain the negative sample output corresponding to the negative sample data; Determine the comparison loss function corresponding to the second part, wherein the comparison loss function is used to determine the accuracy of defect detection in the second part; Based on the positive sample output, the negative sample output, and the comparison loss function, the defect detection accuracy corresponding to the second part is determined until the defect detection accuracy is greater than or equal to the defect detection accuracy threshold, at which point training of the second part is stopped.

4. The method according to claim 1, characterized in that, The step of inputting the transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line includes: Determine the defect feature map corresponding to the transmission line image; The defect feature map is divided into multiple sub-maps; Determine the sub-defect features corresponding to the plurality of subgraphs, and the plurality of associated features corresponding to the plurality of subgraphs, wherein the plurality of associated features represent the defect association relationships between the corresponding subgraphs and other subgraphs; Based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs, the initial defect detection results corresponding to the target transmission line are obtained.

5. The method according to claim 4, characterized in that, The initial defect detection result corresponding to the target transmission line is obtained based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs, including: Based on the sub-defect features and multiple associated features corresponding to the multiple sub-graphs, an enhanced feature map corresponding to the defect feature map is determined; Based on the enhanced feature map, the initial defect detection result corresponding to the target transmission line is determined.

6. The method according to claim 1, characterized in that, In the case where the second part includes a reference defect detection set, the initial defect detection results and the multimodal data are input into the second part of the target architecture model to correct the initial defect detection results and obtain target defect detection results corresponding to the target transmission line, including: Based on the multimodal data, determine the multimodal defect characteristics corresponding to the target transmission line; Based on the multimodal defect features, similar transmission lines corresponding to the target transmission line are determined from multiple reference transmission lines. The similar transmission lines are reference transmission lines whose difference index between the reference defect features and the multimodal defect features is less than a difference threshold. The reference defect detection set includes multiple reference transmission lines, as well as reference defect features and reference defect detection results corresponding to the multiple reference transmission lines respectively. Based on the multimodal defect characteristics and the reference defect detection results corresponding to the similar transmission lines, the initial defect detection results are corrected to obtain the target defect detection results corresponding to the target transmission line.

7. The method according to any one of claims 1 to 6, characterized in that, After inputting the transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line, the method further includes: If the target transmission line is a predetermined transmission line and the initial defect detection result includes defect area parameters, the target scene parameters corresponding to the target transmission line are determined, wherein the predetermined transmission line is the transmission line under the predetermined scene. Based on the target scene parameters, the defect area parameters, and the transmission line image, a fused image is determined; Based on the fused image, the first part is incrementally trained to obtain the incrementally trained first part.

8. A transmission line defect detection device, characterized in that, include: An acquisition module is used to acquire multimodal data corresponding to a target transmission line, wherein the multimodal data includes images of the transmission line; The retrieval module is used to retrieve the target architecture model, wherein the target architecture model includes a first part and a second part. The target architecture model is trained based on a joint training function and a training dataset. The joint training function is used to determine the similarity between the output of the first part and the output of the second part, and to minimize the difference between the outputs of the first part and the second part. The first part is used for defect identification, and the second part is used to correct the output of the first part. The first part includes YOLOv12 and OverLoCK, and the second part includes LLaVA-MoD. The first determining module is used to input the transmission line image into the first part of the target architecture model to obtain the initial defect detection result corresponding to the target transmission line; The second determining module is used to input the initial defect detection result and the multimodal data into the second part of the target architecture model to correct the initial defect detection result and obtain the target defect detection result corresponding to the target transmission line.

9. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the transmission line defect detection method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the transmission line defect detection method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Domain knowledge-driven multi-mode power inspection method, device and system

    CN119048842A

  • Power transmission line key component defect identification method based on cloud edge cooperation

    CN120470463A