Power transmission line defect detection method and system
Through a multimodal large language model and a global-local image representation learning algorithm, the problem of insufficient vision-language alignment in transmission line defect detection is solved, the detection accuracy is improved, and more efficient transmission line defect detection is achieved.
Patent Information
- Application Number
- CN202510862536.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Existing drone inspection technology has problems in transmission line defect detection, such as the small size of defects in the image and the complex background, resulting in insufficient visual information, insufficient alignment between expert description text and image, and inability to effectively utilize the visual-linguistic information in the inspection image.
A multimodal large language model is used to introduce knowledge description texts from different sources, and the detector is pre-trained in combination with an attention-based global-local image representation learning algorithm to construct a joint knowledge transmission line image-text dataset and improve the vision-language alignment capability.
It improves the accuracy of transmission line defect detection, effectively alleviates the problem of insufficient visual information, and achieves higher-precision defect detection.
Smart Images

Figure CN120765573A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power transmission line detection, and in particular relates to a method and system for detecting defects in a power transmission line. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] The stable operation of transmission lines is crucial for power system safety and power supply quality. Transmission lines, exposed to harsh outdoor environments for extended periods, are prone to damage and defects (such as rust, deformation, detachment, and displacement). Furthermore, due to human and animal activity, non-component defects (such as bird nests and foreign objects) caused by external interference are also prone to occur in transmission lines. These defects pose a serious threat to power system security. With the construction of power systems and the increasing distances between transmission lines, traditional manual inspection methods are no longer sufficient for the tedious inspection workload. The rapid development of drone technology has made drone-based transmission line inspection a new mainstream inspection method. Using drone-mounted cameras to capture images of transmission line inspections, deep learning-based defect detectors are used to process and analyze these images and identify defects, effectively enhancing the intelligence of the inspection process and reducing the workload. However, due to imaging distance and angle, defects captured in drone images are often small, partially occluded, and have complex backgrounds, resulting in insufficient visual information related to the defects. This makes it difficult for defect detectors to learn sufficient knowledge for accurate defect identification.
[0004] During power inspections, inspectors use descriptive text (expert descriptions) as the filename for each defect-bearing inspection image to help subsequent data annotators accurately label the defect locations. This text contains rich semantic information about the defects, which can be used to alleviate the problem of insufficient visual information related to the defects in the images. However, no research has yet used this text as a supervisory signal to train defect detectors. In recent years, vision-language pre-training technology has provided ideas for fully utilizing this text. However, incorporating this textual information into defect detectors through vision-language pre-training still presents the following three issues: (1) The proportion of defects in the image is too small, and a large amount of redundant background information seriously interferes with the visual-linguistic alignment of the expert description text and the original inspection image. Therefore, it is impossible to directly use the expert description text and the original inspection image for visual-linguistic pre-training.
[0005] (2) In addition to defects, there are a large number of normal components in the inspection image that are not mentioned in the expert description text, and these components can provide additional visual clues for the model to capture the key features of the defects. Therefore, learning only the visual-linguistic alignment of the expert description text and the defect site is less efficient in utilizing the original inspection image.
[0006] (3) The existing mainstream visual-linguistic pre-training technology and open-source pre-training model are mainly based on image-level supervision, lacking region-word-level alignment and understanding ability. The expert description text often involves the defect site and its surrounding components, so the existing visual-linguistic pre-training technology can only get a suboptimal representation when processing the expert description text, which is not conducive to learning fine-grained visual-linguistic alignment. SUMMARY
[0007] To solve the above problems, the present application provides a power transmission line defect detection method and system, which introduces description text containing knowledge of different sources using a multi-modal large language model to provide additional semantic information for the detector, and introduces a global-local image representation learning algorithm based on attention to pre-train the backbone of the detector to better learn the global and local visual-linguistic alignment, thereby improving the accuracy of power transmission line defect detection.
[0008] According to some embodiments, the first aspect of the present application provides a power transmission line defect detection method, which adopts the following technical solution: A power transmission line defect detection method, comprising: obtaining a power transmission line inspection image and a power transmission line instruction following data set; obtaining generated description text based on the obtained power transmission line instruction following data set, and combining the obtained power transmission line inspection image to form an image-text pair containing multi-modal large language knowledge; obtaining expert description text based on the obtained power transmission line inspection image, matching the generated description text and the expert description text, and obtaining an image-text pair containing field expert knowledge; constructing a joint knowledge power transmission line image-text data set according to the image-text pair containing multi-modal large language knowledge and the image-text pair containing field expert knowledge; performing defect detection of the power transmission line according to the constructed joint knowledge power transmission line image-text data set and a defect detector.
[0009] As a further technical limitation, based on the constructed joint knowledge power transmission line image-text data set, a global-local image representation learning algorithm based on attention is introduced to pre-train the visual base model, and the knowledge in the image-text data set is introduced into the visual base model to complete the pre-training of the visual base model.
[0010] Furthermore, a defect detector is used to detect defects in transmission lines. That is, the text encoder is removed from the pre-trained visual base model, and the image encoder is retained as the backbone network of the defect detector. A feature pyramid network, a region proposal network, and a bounding box head are added to the output end of the backbone network to construct a defect detector. The feature pyramid network provides multi-level features for defect detection, and the region proposal network and the bounding box head are used to analyze multi-level features and locate and classify defects in the input image. The defect detector is fine-tuned using the original inspection image with bounding box annotations, and the knowledge obtained from pre-training is transferred to the defect detection task to obtain the final defect detector.
[0011] As a further technical limitation, high-quality images are screened from the transmission line instruction following dataset, and the screened images are inferred using the tuned InternVL to obtain generated description texts; based on the acquired transmission line inspection images and their corresponding generated description texts, image-text pairs containing multimodal large language model knowledge are obtained.
[0012] Furthermore, ChatGLM is used to semantically match the generated description texts corresponding to all screened high-quality images containing defects with the expert description texts provided by the inspectors for the original inspection images to which the defects belong. Based on the defect images and the expert description texts matched with the generated description texts corresponding to the defect images, image-text pairs containing domain expert knowledge are obtained.
[0013] As a further technical limitation, in the process of constructing a joint knowledge transmission line image-text dataset based on image-text pairs containing multimodal large language knowledge and image-text pairs containing domain expert knowledge, in order to enhance the supervision signal, the category label of each transmission line image is filled into a predefined text template to obtain a description text based on the predefined template, and an image-text pair based on the predefined template is constructed based on the obtained description text and the transmission line image; each image in the image portion of the pretraining dataset is traversed to check the type of image-text pair it participates in. If the image participates in the image-text pair containing domain expert knowledge, the image-text pair containing domain expert knowledge is used as a sample of the pretraining dataset; otherwise, one of the image-text pairs containing multimodal large language model knowledge and the image-text pair based on the predefined template in which the image participates is randomly selected as a sample of the pretraining dataset, and a pretraining dataset that integrates knowledge from the multimodal large language model and domain expert knowledge is constructed, that is, a transmission line image-text dataset with joint knowledge is obtained.
[0014] According to some embodiments, a second solution of the present invention provides a power transmission line defect detection system, which adopts the following technical solution: A transmission line defect detection system, comprising: an acquisition module configured to acquire a transmission line inspection image and a transmission line instruction following dataset; A construction module is configured to generate description text based on the acquired transmission line instruction following dataset, and combine it with the acquired transmission line inspection images to form image-text pairs containing multimodal large language knowledge; obtain expert description text based on the acquired transmission line inspection images, match the generated description text with the expert description text to obtain image-text pairs containing domain expert knowledge; and construct a transmission line image-text dataset with joint knowledge based on the image-text pairs containing multimodal large language knowledge and the image-text pairs containing domain expert knowledge; A detection module is configured to perform transmission line defect detection based on the constructed joint knowledge transmission line image-text dataset and the defect detector.
[0015] According to some embodiments, a third solution of the present invention provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium stores a program, which, when executed by a processor, implements the steps of the power transmission line defect detection method according to the first embodiment of the present invention.
[0016] According to some embodiments, a fourth solution of the present invention provides an electronic device, which adopts the following technical solution: An electronic device comprises a memory, a processor, and a program stored in the memory and running on the processor, wherein when the processor executes the program, the steps of the power transmission line defect detection method according to the first embodiment of the present invention are implemented.
[0017] According to some embodiments, a fifth solution of the present invention provides a computer program product, which adopts the following technical solution: A computer program product includes software codes, wherein the program in the software codes executes the steps of the power transmission line defect detection method according to the first embodiment of the present invention.
[0018] Compared with the prior art, the present invention has the following beneficial effects: This paper uses a multimodal large language model and vision-language pre-training technology to introduce text containing knowledge descriptions from different sources to provide additional semantic information for the detector. It also introduces an attention-based global-local image representation learning algorithm to pre-train the detector backbone to better learn global and local vision-language alignment, thereby improving the accuracy of transmission line defect detection. This invention combines knowledge from two sources, combined with an attention-based global-local image representation learning algorithm, to train a high-performance image encoder. Using the pre-trained image encoder as the backbone network for a power line defect detector effectively alleviates the problem of insufficient visual information related to defects in the original inspection images. Fine-tuning results in a more accurate transmission line defect detector, which has important research implications for intelligent transmission line inspection. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings constituting a part of the specification of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions of this embodiment are used to explain this embodiment and do not constitute an improper limitation on this embodiment.
[0020] Figure 1 This is a flow chart of a method for detecting defects in a power transmission line according to the first embodiment of the present invention; Figure 2 Detailed steps of the power transmission line defect detection method according to the first embodiment of the present invention are shown in FIG. Figure 3 Schematic diagram of the overall training process in Example 1 of the present invention; Figure 4 A schematic diagram of a method for introducing and generating description text and expert description text in the first embodiment of the present invention; Figure 5 Schematic diagram of the attention-based global-local image representation learning algorithm in Example 1 of the present invention; Figure 6 This is a diagram showing the effect of power transmission line defect detection in the first embodiment of the present invention, wherein Figure 6 (b) is the effect diagram of transmission line defect detection. Figure 6 (a) in the Figure 6 A partial schematic diagram of (b); Figure 7 This is a structural block diagram of a power transmission line defect detection system in Embodiment 2 of the present invention. DETAILED DESCRIPTION
[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0022] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0023] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0024] In the present invention, terms such as "upper", "lower", "left", "right", "front", "back", "vertical", "horizontal", "side", "bottom", etc. indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are relational words determined only for the convenience of describing the structural relationships of the various parts or elements of the present invention, and do not specifically refer to any part or element in the present invention, and should not be understood as limiting the present invention.
[0025] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0026] Example 1 Embodiment 1 of the present invention introduces a method for detecting defects in a power transmission line.
[0027] like Figure 1 A method for detecting defects in a transmission line is shown, comprising: Acquire transmission line inspection images and transmission line instruction following datasets; Generate description text based on the acquired transmission line instruction following dataset, and combine it with the acquired transmission line inspection image to form an image-text pair containing multimodal large language knowledge; Obtain expert description text based on the acquired transmission line inspection image, match the generated description text with the expert description text, and obtain an image-text pair containing domain expert knowledge; Based on image-text pairs containing multimodal large language knowledge and image-text pairs containing domain expert knowledge, a joint knowledge transmission line image-text dataset is constructed; Transmission line defect detection is performed based on the constructed joint knowledge transmission line image-text dataset and defect detector.
[0028] This embodiment combines a multimodal large model and joint knowledge to detect defects in transmission lines. The specific steps are as follows: Figure 2 As shown, specifically: Step S1: Label and crop the original transmission line inspection images and use the large language model ChatGLM to batch generate question-answer pairs to construct a transmission line command following dataset. This dataset contains multimodal conversation data related to various transmission line components and defects, covering various concepts related to transmission line defect detection. Step S2: Using the QLoRA technology on the power line instruction following dataset, we fine-tune the high-performance multimodal large language model InternVL to learn visual-language alignment of various concepts related to power line defect detection. Step S3: Filter high-quality images from the power line instruction following dataset and use the tuned InternVL to infer all the filtered images to obtain generated description text. Each image and the corresponding generated description text constitute an image-text pair that contains knowledge of the multimodal large language model. Step S4: Use ChatGLM to semantically match the generated description texts corresponding to all images containing defects with the expert description texts provided by the inspectors for the original inspection images containing the defects. The defect image and the expert description text that correctly matches its generated description text constitute an image-text pair that contains domain expert knowledge. Step S5: constructing a joint knowledge power transmission line image-text dataset using the image-text pairs containing the multimodal large language model knowledge and the image-text pairs containing the domain expert knowledge; Step S6: On the joint knowledge power line image-text dataset, an attention-based global-local image representation learning algorithm is introduced to pre-train the visual basic model Chinese-CLIP, introducing the rich knowledge in the dataset into Chinese-CLIP; Step S7: Use the pre-trained image encoder in Chinese-CLIP as the backbone network to build a defect detector, and use a small amount of labeled defect detection data to fine-tune the final defect detector.
[0029] This embodiment overcomes the problem of insufficient visual information related to defects in transmission line inspection images from the perspective of multimodal learning. The overall training process is as follows: Figure 3 As shown in the figure. In raw transmission line inspection data, the expert descriptions provided by inspectors only describe a small portion of the defective area within the high-resolution inspection image. The image often contains numerous components and other complex backgrounds that are irrelevant to the expert descriptions. Therefore, image-text constructed directly from raw inspection data is not suitable for learning visual-language alignment. Otherwise, the image-text alignment would be adversely affected by the large amount of interfering background information in the image, resulting in poor learning results and low utilization of image information.
[0030] To better learn visual-language alignment and fully utilize the information contained in inspection images, we first need to crop the defective areas and other normal parts in the image. This produces single-instance images (for all defects and parts) that are more conducive to visual-language alignment, as well as context-enhanced instance images (for defects only, as expert descriptions often cover the defect and its surrounding parts). Then, considering that only a small portion of the cropped images can be matched with expert descriptions, the remaining images (images of normal parts and images of defects not described by experts) lack matching descriptions. Therefore, a multimodal large language model is needed to generate matching descriptions for them.
[0031] While existing large language models and multimodal large language models contain knowledge related to transmission line inspection (derived from open-source documentation and data on the internet), they either lack image input or, due to data barriers, lack access to a large number of images of transmission line components or defects during training. This results in large language models and multimodal large language models providing good responses when fed only queries related to transmission line inspection. However, when fed both queries and images related to transmission line inspection, multimodal large language models often give incorrect answers. This indicates that existing multimodal large language models are unable to effectively align visual-linguistic information in transmission line scenarios. Therefore, with the help of large language models, a dedicated instruction-following dataset for transmission line scenarios can be constructed using the aforementioned single-instance images and context-enhanced instance images. This can be used to fine-tune the multimodal large language model to help it learn visual-linguistic alignment in transmission line scenarios, thereby generating higher-quality descriptions for images.
[0032] In step S1, the original transmission line inspection images are labeled and cropped, and question-answer pairs are generated in batches using the large language model ChatGLM to construct a transmission line command following dataset. The dataset contains multimodal conversation data related to various transmission line components and defects, covering various concepts related to transmission line defect detection, including: First, the components and defects in the original transmission line inspection images are annotated and the annotated boxes are cropped to obtain a large number of single-instance images with category labels (the images mainly focus on a single component or defect) and context-enhanced instance images with category labels and defect location coordinates (the images contain contextual information such as the defect and the component where it is located). These two types of images constitute the image portion of the transmission line command following dataset.
[0033] Secondly, for single-instance images, the large language model ChatGLM is first prompted to generate a set of question-answer pairs about the category based on the category label, such as: "Please generate 30 sets of single-round questions and answers about the appearance, color, location, or function of the shielding ring in the transmission line." From the generated question-answer pairs, the parts that are helpful for learning the relevant concepts of transmission line defect detection are selected as seed question-answer pairs. Then, the context learning ability of the large language model is used to generate coarse-grained paired question-answer pairs for each single-instance image belonging to the category based on the seed question-answer pairs.
[0034] Then, for each context-enhanced instance image, a basic description of the image is obtained based on a manually designed description template, using the category label and defect location coordinates. For example, "The image shows a damaged grading ring on a transmission line, located in the lower left corner of the image. The area of this area is moderate in size." This description is then used to prompt ChatGLM to generate a coarse-grained question-answer pair corresponding to the image. For example, "Please rewrite this sentence in a different way while maintaining the semantic integrity of the original sentence. Then, using this sentence as the answer, a matching question is generated." Finally, the question-answer pairs that are coarsely matched with single-instance images and the question-answer pairs that are coarsely matched with context-enhanced instance images constitute the text part of the transmission line instruction following dataset. This part, together with the image part of the transmission line instruction following dataset, constitutes the transmission line instruction following dataset. This dataset covers visual and language instances of various concepts related to transmission line defect detection, and can therefore help multimodal large language models learn visual-language alignment in transmission line scenarios.
[0035] In this embodiment, considering that the expert description texts are all in Chinese and that existing equipment conditions and data scale make it difficult to support end-to-end tuning of a large multimodal language model, it is necessary to adopt a large multimodal language model that supports Chinese input and a model tuning technique that is more efficient than end-to-end fine-tuning. Therefore, in step S2, the high-performance large multimodal language model InternVL is tuned using the QLoRA technology on the power line instruction following dataset to enable it to learn visual-language alignment of multiple concepts related to transmission line defect detection, specifically including: First, we selected InternVL, a high-performance multimodal large language model that has demonstrated excellent performance in Chinese vision-language tasks, for optimization for power transmission line scenarios. Taking into account the memory constraints of the available devices, we selected the open-source InternVL2-2B weights to initialize the model parameters.
[0036] Then, on the constructed transmission line instruction following dataset, QLoRA technology is used to efficiently tune the InternVL model, effectively avoiding the problem of traditional end-to-end tuning requiring a large amount of computing power and a lot of time under the condition of limited computing resources.
[0037] In this embodiment, although the cropping in step S1 yields a large number of component and defect images, many of these images are low-quality and there is a certain degree of imbalance in the distribution of component and defect categories. To improve the quality of text generation and subsequent visual-language pre-training, data screening is necessary. Using a tuned large multimodal language model for inference on the high-quality images obtained through screening can incorporate knowledge related to the observed components or defects from the large multimodal language model into the generated text.
[0038] Therefore, in step S3, high-quality images are screened from the power line instruction following dataset, and the tuned InternVL is used to infer all the screened images to obtain generated description texts. Each image and the corresponding generated description text constitute an image-text pair containing the knowledge of the multimodal large language model, specifically including: First, high-quality defect and component images are screened from the image portion of the transmission line instruction following dataset, which constitutes the image portion of the pre-training dataset required for vision-language pre-training.
[0039] Then, to ensure the diversity of the generated results, a variety of manually designed instructions, such as "Describe the component in the image and determine whether it has defects", are used to prompt the tuned InternVL to reason about all the screened images and generate a fine-grained matching description text for each image. This text contains the power domain knowledge provided by the tuned InternVL (including the power domain knowledge provided by ChatGLM and contained in InternVL itself). Each image and the corresponding generated description text constitute an image-text pair containing the knowledge of the multimodal large language model. The process is as follows: Figure 4 shown.
[0040] In this example, the original expert description text contains some information irrelevant to the defect image. For example, the voltage level, line name, and tower number described in the sentence "500kV Mengji I Line #0091 Tower, Right Conductor Small Side, Left String, Tower End, Third Insulator Self-Explosion" don't directly correspond to the visual information in the image, thus interfering with visual-language pre-training. Furthermore, an image may contain multiple defects, while the expert description text only mentions one. Therefore, it's not possible to directly pair the cropped defect image with the expert description text corresponding to the original image.
[0041] In step S4, ChatGLM is used to perform semantic matching between the generated description texts corresponding to all images screened out to contain defects and the expert description texts provided by the inspectors for the original inspection images to which the defects belong. The defect images and the expert description texts that correctly match the generated description texts constitute image-text pairs containing domain expert knowledge, specifically including: First, the expert description text provided by the inspectors for the original inspection images is preprocessed to delete redundant content, such as voltage level, line name and tower number, to prevent these contents from interfering with the subsequent pre-training effect.
[0042] Then, for each image containing a defect, ChatGLM is instructed to judge the degree of semantic matching between the generated description text of the image and the expert description text of the original inspection image of the defect through the instruction "Please judge the degree of matching between the following two texts and give a matching score between 0-10". A score between 0-10 is given. When the score is greater than or equal to the threshold score of 6, the image and the expert description text are considered to be correctly matched. The two constitute an image-text pair containing domain expert knowledge. The process is as follows: Figure 4 shown.
[0043] This embodiment uses the two image-text pairs containing knowledge from different sources obtained in steps S3 and S4 to construct a joint knowledge multimodal dataset for visual-language pre-training. At the same time, considering that the supervisory signal strength provided by the two types of text is limited, the category labels retained when cropping the image can be used to enhance the supervisory signal during the pre-training process. Therefore, in step S5, the image-text pairs containing multimodal large language model knowledge and the image-text pairs containing domain expert knowledge are used to construct a joint knowledge transmission line image-text dataset, specifically including: First, in order to enhance the supervision signal, the category label of each image is filled into a predefined text template to obtain a descriptive text based on the predefined template, such as: "There is a normal equalizing ring in the picture". This text and the image together constitute an image-text pair based on the predefined template.
[0044] Then, each image in the image part of the pre-training dataset is traversed to check the type of image-text pair constituted by the image. If the image participates in the formation of an image-text pair containing domain expert knowledge, then this group of image-text pairs containing domain expert knowledge is used as a sample of the pre-training dataset. Otherwise, one item is randomly selected from the image-text pairs containing multimodal large language model knowledge and the image-text pairs based on the predefined template as a sample of the pre-training dataset. The constructed pre-training dataset integrates the knowledge from the multimodal large language model and domain experts, and is called the joint knowledge transmission line image-text dataset.
[0045] In this embodiment, considering that visual-language pre-training from scratch requires billions of data points and consumes a significant amount of computing power, to avoid training the model from scratch, this embodiment employs transfer learning to migrate an existing open-source pre-trained model to the power transmission line image-text dataset constructed in step S5 for pre-training on the power transmission line scenario. On the other hand, existing mainstream visual-language pre-training techniques and open-source pre-training models primarily rely on image-level supervision. While they possess global image-text alignment capabilities, they lack alignment and understanding capabilities at the region-to-word level. Expert descriptions often refer to the defect location and surrounding components. For example, in the sentence "500kV Linhui I Line #0146 Tower Conductor Small Side String Bowl Head Hanging Plate and Connecting Plate Connecting Bolt Missing Pin," in addition to mentioning the "bolt missing pin" defect, it also mentions the connecting hardware on both sides, the "bowl head hanging plate" and "connecting plate." Therefore, to better learn fine-grained region-to-word alignment, it is necessary to improve existing visual-language pre-training algorithms to help the model learn local representations while simultaneously learning global image and text representations.
[0046] In step S6, an attention-based global-local image representation learning algorithm is introduced on the joint knowledge power line image-text dataset to pre-train the visual basic model Chinese-CLIP, thereby introducing the rich knowledge in the dataset into Chinese-CLIP. Specifically, the algorithm includes: First, we use the open-source Chinese-CLIP ViT-B-16 pre-trained weight initialization model to improve pre-training efficiency, so as to avoid the problem of consuming a lot of computing power to pre-train the model from scratch. On this basis, we use the attention-based global-local image representation learning algorithm to pre-train the Chinese-CLIP model, such as Figure 5 For each image-text pair, the image encoder in Chinese-CLIP encodes the image into a global image representation and a local image representation, as follows: ; (1) in, represents the image encoder of Chinese-CLIP, and They represent the global image representation learning function and the local image representation learning function that convert the image encoder output into the multimodal feature space, respectively. is the input image, the global representation of the image For one d dimensional vector, local representation of the image Include m indivuald dimensional vectors (corresponding to the image segmentation m Image blocks), similarly, the text encoder in Chinese-CLIP encodes the text in the image-text pair into global text representation and local text representation, as follows: ; (2) in, represents the text encoder of Chinese-CLIP, and They represent the global text representation learning function and the local text representation learning function that convert the text encoder output into the multimodal feature space, respectively. For input text, global representation of text For one d dimensional vector, local representation of text Include w indivual d dimensional vectors (corresponding to the text segmentation w participles).
[0047] Secondly, in order to enable Chinese-CLIP to learn a global image representation that is fully aligned with the text on the power line image-text dataset with joint knowledge while avoiding damage to the original pre-trained representation, Chinese-CLIP is pre-trained using the image-text contrastive learning objective. The loss calculation formula is: (3) (4) (5) in, and They are the first i Hedi j The global image representation of the image, and They are the first i Hedi j The global representation of the text, b is the batch size, τ 1 is the temperature coefficient, and are the image-text global contrastive losses from image to text and from text to image, respectively. is the global contrast loss.
[0048] Thirdly, in order to enable Chinese-CLIP to learn local image representations that are fully aligned with the text on the power line image-text dataset with joint knowledge, Chinese-CLIP is pre-trained using an attention-based region-segmentation comparison learning objective. Specifically, the dot product similarity between the feature vector of each image block and the feature vector of the text segmentation is calculated, as follows: (6) in, represents the dot product operation, is the similarity matrix, for Middle i Rank j The elements of the column represent the similarity score between the i-th image block feature vector and the j-th text word segmentation feature vector. The attention matrix can be obtained by normalizing each column in , where the attention weight The calculation formula is: (7) in, τ 2 is the temperature coefficient, for Middle i Rank k The elements of the column are right The context-aware image representation can be obtained by weighted summing of the image block feature vectors in , and the calculation formula is: (8) in, for The j image block feature vectors, Based on the i The context-aware image representation of each word segmentation is then compared with the local text representation to help Chinese-CLIP capture a finer-grained visual-language alignment. The loss is calculated as: (9) (10) (11) in, Based on the j Context-aware image representation of word segments, and Separate i Hedi ja feature vector corresponding to each segmented word, and are image-text local contrastive losses from image to text and from text to image, respectively, is a local contrastive loss.
[0049] Finally, the total pre-training loss is a weighted sum of the global contrastive loss and the local contrastive loss, and the calculation formula is: (12) wherein, is the total pre-training loss, and are the weights of the global contrastive loss and the local contrastive loss, respectively.
[0050] The method uses the Chinese-CLIP model pre-trained in the above step S6 to construct and fine-tune the final defect detector, so as to achieve the effect of transferring the knowledge in the multi-modal large language model and the expert description text to the defect detection. In the step S7, the image encoder in the pre-trained Chinese-CLIP is used as the backbone network to construct the defect detector, and a small amount of labeled defect detection data is used to fine-tune the final defect detector, which specifically includes: First, the text encoder is removed from the pre-trained Chinese-CLIP model, and the image encoder is retained as the backbone network of the defect detector, and a simple feature pyramid network, a region proposal network and a bounding box head are added at the output end to construct a complete defect detector, wherein the simple feature pyramid network provides multi-level features for defect detection, and the region proposal network and the bounding box head are used to analyze the multi-level features and locate and classify the defects in the input image.
[0051] Then, a small-scale defect detection data, i.e. an original inspection image with bounding box annotation, is used to fine-tune the entire defect detector, so as to finally transfer the knowledge obtained by pre-training to the defect detection task and obtain the final defect detector.
[0052] The detection effect of the method of the application in the defect detection task of the power transmission line is as follows: Figure 6As shown in (a) and (b) in the figure. The present invention uses a multimodal large language model to fully utilize the visual and language information contained in the original transmission line inspection data, introduces the generated description text provided by the multimodal large language model and the expert description text provided by the inspection personnel to construct a joint knowledge transmission line image-text dataset, which is used to train a detector backbone network containing rich knowledge, thereby effectively alleviating the problem of insufficient defect-related visual information in the image; in addition, in view of the problem that the existing visual-language pre-training method is difficult to learn fine-grained visual-language alignment in the transmission line scene, this paper introduces an attention-based global-local image representation learning algorithm to pre-train the Chinese-CLIP model for the transmission line scene, helping to achieve finer-grained visual-language alignment to better capture the knowledge contained in the text, and having a better migration effect in downstream detection tasks. It can be seen that the method of the present invention can jointly utilize the two modalities of vision and language to obtain richer information, helping the transmission line defect detection model to overcome the problem of insufficient defect-related visual information in the original inspection image, thereby effectively improving the detection accuracy of transmission line defects.
[0053] This embodiment utilizes a multimodal large language model and vision-language pre-training technology to introduce text containing knowledge descriptions from different sources to provide additional semantic information for the detector. It also introduces an attention-based global-local image representation learning algorithm to pre-train the detector backbone to better learn global and local vision-language alignment, thereby improving the accuracy of power transmission line defect detection. Combined with the attention-based global-local image representation learning algorithm, the goal of training a high-performance image encoder by combining knowledge from both sources is achieved. Using the pre-trained image encoder as the backbone network of the power transmission line defect detector can effectively alleviate the problem of insufficient visual information related to defects in the original inspection images. Fine-tuning can yield a more accurate power transmission line defect detector, which has important research significance for intelligent power transmission line inspection.
[0054] Example 2 The second embodiment of the present invention introduces a method and system for detecting defects in a power transmission line.
[0055] like Figure 7 A transmission line defect detection system shown includes: an acquisition module configured to acquire a transmission line inspection image and a transmission line instruction following dataset; A construction module is configured to generate description text based on the acquired transmission line instruction following dataset, and combine it with the acquired transmission line inspection images to form image-text pairs containing multimodal large language knowledge; obtain expert description text based on the acquired transmission line inspection images, match the generated description text with the expert description text to obtain image-text pairs containing domain expert knowledge; and construct a transmission line image-text dataset with joint knowledge based on the image-text pairs containing multimodal large language knowledge and the image-text pairs containing domain expert knowledge; A detection module is configured to perform transmission line defect detection based on the constructed joint knowledge transmission line image-text dataset and the defect detector.
[0056] The detailed steps are the same as those of the power transmission line defect detection method provided in Example 1 and will not be repeated here.
[0057] Example 3 A third embodiment of the present invention provides a computer-readable storage medium.
[0058] A computer-readable storage medium stores a program thereon, which, when executed by a processor, implements the steps of the power transmission line defect detection method as described in the first embodiment of the present invention.
[0059] The detailed steps are the same as those of the power transmission line defect detection method provided in Example 1 and will not be repeated here.
[0060] Example 4 A fourth embodiment of the present invention provides an electronic device.
[0061] An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, the steps of the power transmission line defect detection method as described in the first embodiment of the present invention are implemented.
[0062] The detailed steps are the same as those of the power transmission line defect detection method provided in Example 1 and will not be repeated here.
[0063] Example 5 A fifth embodiment of the present invention provides a computer program product.
[0064] A computer program product includes software code, wherein the program in the software code executes the steps of the power transmission line defect detection method according to the first embodiment of the present invention.
[0065] The detailed steps are the same as those of the power transmission line defect detection method provided in Example 1 and will not be repeated here.
[0066] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk drives, CD-ROMs, optical storage devices, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0067] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0068] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0069] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0070] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0071] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
[0072] The above description is merely a preferred embodiment of this embodiment and is not intended to limit this embodiment. Those skilled in the art will readily appreciate that this embodiment may be modified and varied in various ways. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this embodiment shall be within the scope of protection of this embodiment.
Claims
1. A method for detecting defects in a transmission line, characterized in that: include: Acquire transmission line inspection images and transmission line instruction following datasets; Generate description text based on the acquired transmission line instruction following dataset, and combine it with the acquired transmission line inspection image to form an image-text pair containing multimodal large language knowledge; Obtain expert description text based on the acquired transmission line inspection image, match the generated description text with the expert description text, and obtain an image-text pair containing domain expert knowledge; Based on image-text pairs containing multimodal large language knowledge and image-text pairs containing domain expert knowledge, a joint knowledge transmission line image-text dataset is constructed; Transmission line defect detection is performed based on the constructed joint knowledge transmission line image-text dataset and defect detector.
2. A method for detecting defects in a power transmission line as claimed in claim 1, characterized in that: Based on the constructed joint knowledge transmission line image-text dataset, an attention-based global-local image representation learning algorithm is introduced to pre-train the visual base model, and the knowledge in the image-text dataset is introduced into the visual base model to complete the pre-training of the visual base model.
3. A method for detecting defects in a power transmission line as claimed in claim 2, characterized in that: A defect detector is used to detect defects in transmission lines. Specifically, the text encoder is removed from the pre-trained visual base model, and the image encoder is retained as the backbone network of the defect detector. A feature pyramid network, a region proposal network, and a bounding box head are added to the output of the backbone network to construct the defect detector. The feature pyramid network provides multi-level features for defect detection, and the region proposal network and bounding box head are used to analyze the multi-level features and locate and classify defects in the input image. The defect detector is fine-tuned using original inspection images with bounding box annotations, and the knowledge obtained from pre-training is transferred to the defect detection task to obtain the final defect detector.
4. A method for detecting defects in a power transmission line as claimed in claim 1, characterized in that: High-quality images are screened from the transmission line instruction following dataset, and the screened images are inferred using the tuned InternVL to obtain generated description texts. Based on the acquired transmission line inspection images and their corresponding generated description texts, image-text pairs containing multimodal large language model knowledge are obtained.
5. A method for detecting defects in a power transmission line as claimed in claim 4, characterized in that: ChatGLM is used to semantically match the generated description texts corresponding to all screened high-quality images containing defects with the expert description texts provided by inspectors for the original inspection images containing the defects. Based on the defect images and the expert description texts matched with the generated description texts corresponding to the defect images, image-text pairs containing domain expert knowledge are obtained.
6. A method for detecting defects in a power transmission line as claimed in claim 1, characterized in that: In the process of constructing a joint knowledge transmission line image-text dataset based on image-text pairs containing multimodal large language knowledge and image-text pairs containing domain expert knowledge, in order to enhance the supervision signal, the category label of each transmission line image is filled into a predefined text template to obtain a description text based on the predefined template. The obtained description text and the transmission line image constitute an image-text pair based on the predefined template; each image in the image portion of the pretraining dataset is traversed to check the type of image-text pair it participates in. If the image participates in the image-text pair containing domain expert knowledge, the image-text pair containing domain expert knowledge is used as a sample of the pretraining dataset. Otherwise, one of the image-text pairs containing multimodal large language model knowledge and the image-text pair based on the predefined template that the image participates in is randomly selected as a sample of the pretraining dataset, and a pretraining dataset that integrates the knowledge from the multimodal large language model and the domain expert knowledge is constructed, i.e., a joint knowledge transmission line image-text dataset is obtained.
7. A power transmission line defect detection system, characterized in that: include: an acquisition module configured to acquire a transmission line inspection image and a transmission line instruction following dataset; A construction module is configured to generate description text based on the acquired transmission line instruction following dataset, and combine it with the acquired transmission line inspection images to form image-text pairs containing multimodal large language knowledge; obtain expert description text based on the acquired transmission line inspection images, match the generated description text with the expert description text to obtain image-text pairs containing domain expert knowledge; and construct a transmission line image-text dataset with joint knowledge based on the image-text pairs containing multimodal large language knowledge and the image-text pairs containing domain expert knowledge; The detection module is configured to perform transmission line defect detection based on the constructed joint knowledge transmission line image-text dataset and the defect detector.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the power transmission line defect detection method according to any one of claims 1 to 6 are implemented.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of the power transmission line defect detection method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising software code, characterized in that The program in the software code executes the steps of the power transmission line defect detection method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Electric power defect image detection method based on image-text question-answer multi-modal model
CN117763107A
Bolt missing detection method based on cross-media and knowledge reasoning
CN119514680A
Method for detecting defects of power equipment and related products
CN119557716A
Power transmission line defect identification method based on multi-modal comparative learning
CN119741253A
Defect detection method and system for unmanned aerial vehicle inspection equipment in open domain
CN119942378A