Defect detection method and system for unmanned aerial vehicle inspection equipment in open domain
By constructing a multimodal defect recognition network structure, combining BERT and Swin Transformer models, and fusing image and text features, the problem of insufficient defect recognition capabilities of power equipment in open scenarios is solved, and more efficient and accurate defect detection effects are achieved.
Patent Information
- Application Number
- CN202510036238.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The existing UAV inspection equipment defect identification methods lack effective identification capabilities in open scenarios and cannot adapt to the diversity of power equipment defects and the complexity of scenarios.
The multimodal defect identification network structure is adopted, and the category and location information of equipment defects are identified by building input modules, feature extraction modules, graphic and text feature fusion modules, feature decoders and output modules, combined with BERT and Swin Transformer models.
It improves the accuracy and efficiency of power equipment defect detection in open scenarios, enhances the model's semantic understanding and spatial perception of the scene, adapts to changeable on-site conditions, and supports more complex defect recognition tasks.
Smart Images

Figure CN119942378A_ABST
Abstract
Description
Background Art
[0002] As an advanced aviation technology, drone technology has made significant progress in recent years. From its initial military use to its current widespread application in agriculture, medical care, environmental monitoring and other fields, drone technology has become an indispensable part of contemporary society.
[0003] The application of power UAV technology in the field of power transmission and distribution has received widespread attention. Equipped with high-resolution cameras and sensors, it can achieve efficient monitoring of power lines and detect various potential defects. In recent years, the rise of computer vision technology has made the automated application of defect recognition algorithms gradually become mainstream.
[0004] The locations where drones are photographed usually have varied terrain and complex environments, and there are a large number of different types of power equipment on the lines, and the equipment defects are diverse. However, current recognition models mostly use recognition modes with fixed output categories, which cannot effectively identify equipment defects in open scenarios and do not meet the current characteristics of diverse equipment defects and complex and changing scenarios.
[0005] Prior art 1 related to the present invention: single-modal defect recognition method based on drone inspection images, such as Faster RCNN, YOLO, CenterNet, DETR, etc.; Existing defect recognition methods mostly use single-mode (image) recognition mode, such as Faster RCNN, YOLO, DETR, etc. The recognition mode of this method is as follows Figure 1 As shown in the figure, the process is to input an image, then use a feature extraction network (such as a convolutional neural network) to extract the equipment defect features, and finally output the recognition results through post-processing and other means. This type of single-modality-based method does not have the ability to input text or utilize multi-modal features, and does not support defect recognition in open scenarios, lacking versatility.
[0006] Compared with multimodal recognition methods, single-modal recognition methods have the following disadvantages: 1. The unimodal recognition method lacks the strong semantics of text features and relies on the features of a single image data, resulting in poor generalization ability of the model under different data sets.
[0007] 2. The single-modal recognition method does not support text input and lacks interactive capabilities. The model can only output fixed defect categories, which does not conform to the current characteristics of diversified equipment defects and complex and changeable scenarios, resulting in the lack of universality of the model. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide an open domain UAV inspection equipment defect detection method and system in view of the deficiencies in the above-mentioned prior art, so as to solve the technical problem of poor defect recognition ability of power equipment in existing open scenarios.
[0009] The present invention adopts the following technical solutions: A method for detecting defects of open domain drone inspection equipment comprises the following steps: Construct a multimodal defect recognition network structure; The multimodal defect recognition network structure constructed by training the training data set; The drone images and texts to be identified are input into the trained multimodal defect recognition network structure, and the category and location information of the equipment defects are output.
[0010] Preferably, the multimodal defect recognition network structure includes: An input module inputs drone images and texts to be identified. The drone images include drone inspection images and depth maps to be identified. The texts include text descriptions of equipment defects input by users according to actual scenarios. The text descriptions include phrases or single words. Feature extraction module, using BERT to extract text features and Swin Transformer to extract corresponding image features; The image-text feature fusion module fuses the image and text features, obtains the features of the two image-text fusions, and inputs them into the self-attention module respectively, and obtains the fused image features and text features through the FFN feedforward neural network; Feature decoder, obtains the position feature representation of the candidate box; The output module compares and matches the input text features with the image features corresponding to the candidate boxes. By setting the score threshold hyperparameter, the model outputs the candidate box category and position coordinates whose similarity with the text feature matching is higher than the threshold, and draws specific defect information in the image.
[0011] Preferably, the BERT model adopts the Transformer architecture and learns deep language representation through large-scale unsupervised pre-training tasks; Swin Transformer uses the Transformer structure and cross-layer connections, and adopts the local attention mechanism and windowed self-attention mechanism. For the input depth image, the parameter-shared Swin Transformer is used to extract features.
[0012] Preferably, the feature extraction using parameter-sharing Swin Transformer is as follows:
[0013] in, represents the deep image features, Represents the visible light image features.
[0014] Preferably, feature fusion is specifically: For image features, feature concatenation and flattening are performed on the feature maps of different resolutions of Swin Transformer; The Q, K, and V matrices of image features and text features are obtained respectively. The Q matrix of image features and the K and V matrices of text features are used as the input of cross-modal self-attention from image to text features. The Q matrix of text features and the K and V matrices of image features are used as inputs for cross-modal self-attention from text to image features; After obtaining the features of the two image-text fusions, they are input into the self-attention module respectively, and then the fused image features and text features are obtained through the feedforward neural network.
[0015] Preferably, the text features and image features are fused as follows:
[0016] in, The Q matrix represents the image features, The Q matrix represents the text features, The K matrix and V matrix represent the text features respectively; The K matrix and V matrix represent the image features respectively.
[0017] Preferably, and The calculation is as follows:
[0018] in, To fuse the image features with the text features, It is the text feature fused with image feature.
[0019] Preferably, the position feature representation of the candidate box is obtained as follows: Through the Embedding layer, we get a set of initial candidate box Queries, get the Q, K, V matrices, and input them into the self-attention module, whose output is used as the Q matrix. At the same time, we get the K and V matrices of the image and text features output by the previous image-text fusion module respectively. The Q matrix is fused with the K and V matrices of text features for text cross-modal self-attention; The output Q matrix is fused with the K and V matrices of the image features through image cross-modal self-attention; After passing through the feedforward neural network, the refined candidate box feature vector is output.
[0020] Preferably, the training data set includes: Defect images, by collecting and summarizing drone inspection images of transmission lines and distribution lines and carrying out data annotation; Power equipment images: by collecting and summarizing drone inspection images of transmission lines and distribution lines and carrying out data annotation of power equipment, common power equipment is annotated; General defect images, defects in other scenes except drone inspection images, partial defects; Open world images, introduce open source datasets into the training set for mixed training in a set ratio.
[0021] Preferably, the annotation methods of defect images are divided into equipment defect description and scene description. The equipment defect description is to annotate all defects in the picture, and the annotation content includes the defect name and equipment name; the scene description is an overall description of the defects in the whole picture; the annotation method of power equipment images is the English description of the equipment; the annotation method of general defect images adopts general defect word annotation.
[0022] In a second aspect, an embodiment of the present invention provides an open domain drone inspection equipment defect detection system, comprising: Network module, building a multimodal defect recognition network structure; A training module uses a training data set to train the constructed multimodal defect recognition network structure; The detection module inputs the drone image and text to be identified into the trained multimodal defect recognition network structure and outputs the category and location information of the equipment defect.
[0023] In a third aspect, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned open domain drone inspection equipment defect detection method when executing the computer program.
[0024] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned open domain drone inspection equipment defect detection method.
[0025] In a fifth aspect, a chip comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned open domain UAV inspection equipment defect detection method when executing the computer program.
[0026] In a sixth aspect, an embodiment of the present invention provides an electronic device, including a computer program, which, when executed by the electronic device, implements the steps of the above-mentioned open domain drone inspection equipment defect detection method.
[0027] Compared with the prior art, the present invention has at least the following beneficial effects: An open domain UAV inspection equipment defect detection method aims to improve the accuracy and efficiency of power equipment defect detection by building a multimodal defect recognition network structure, using a training data set for network training, inputting UAV images and texts to be identified, and outputting the category and location information of equipment defects. This method enhances the model's semantic understanding and spatial perception of the scene by fusing multimodal information of images and texts, enabling defect detection to adapt to changing field conditions and providing technical support for the safe operation of power equipment.
[0028] Furthermore, the setting of the multimodal defect recognition network structure aims to enhance the model's ability to identify power equipment defects by integrating image and text modal data. This structure uses BERT and Swin Transformer to extract text and image features respectively, and then integrates these features through feature fusion techniques such as self-attention mechanism to obtain richer representations, which can improve the model's ability to understand equipment defects.
[0029] Furthermore, the purpose of introducing the BERT model is to utilize its deep language representation capabilities. The rich semantic information learned by the model through large-scale unsupervised pre-training tasks can enhance the model's ability to extract text features. The BERT model uses the Transformer architecture, which can capture the bidirectional dependencies in the text, thereby providing more accurate text features for power equipment defect detection.
[0030] Furthermore, using parameter-sharing Swin Transformer to extract features from depth images and visible light images can reduce the complexity and computational complexity of the model while ensuring that the features extracted from the two different types of images are consistent in the representation space. The parameter sharing strategy enables the model to maintain the consistency of the learned feature representation when processing the two types of images, thereby more effectively combining information from different modalities in the subsequent feature fusion and defect recognition process.
[0031] Furthermore, the purpose of the feature fusion setting is to integrate feature information from different modalities to obtain a more comprehensive feature representation and enhance the model's ability to identify defects. Through feature fusion, the model can more accurately locate and classify power equipment defects, making defect detection results more reliable.
[0032] Furthermore, the purpose of setting the fusion mode of text features and image features is to achieve cross-modal information integration. This fusion mode combines the semantic information of text features with the visual information of image features, so that the model can understand the defect conditions of power equipment from multiple dimensions. In principle, by calculating the interaction between the Q matrix of image features and the K and V matrices of text features, as well as the interaction between the Q matrix of text features and the K and V matrices of image features, a cross-modal self-attention mechanism from image to text and from text to image is realized, achieving more accurate defect detection in complex open domain scenarios.
[0033] Furthermore, the setting of the position feature representation of the candidate box is aimed at accurately locating and identifying the device defects in the image. This process initializes the candidate box Queries and combines the image features and text features, so that the model can optimize the position features of the candidate box according to the semantic information described in the text. The final output candidate box feature vector can more accurately reflect the location of the defect.
[0034] Furthermore, the training dataset is designed to provide rich and diverse samples for the multimodal defect recognition network to enhance the learning and generalization capabilities of the model. By collecting and annotating defect images, power equipment images, general defect images, and open-world images from different scenes, the training dataset can cover a variety of possible defect types and environmental conditions, so that the model can learn more comprehensive data features during the training process.
[0035] It can be understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.
[0036] In summary, the present invention breaks through the limitations of traditional single-mode recognition mode, can effectively identify various equipment problems, and helps to enhance the ability to identify defects in power equipment in open scenarios, thereby reducing the operating risks of the entire power system; in the future, it can play an important role in the fields of smart grid maintenance, line equipment operation and maintenance, etc.
[0037] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 It is a schematic diagram of an existing single-mode defect recognition method; Figure 2 The overall flow chart of the method of the present invention is as follows; Figure 3 This is a schematic diagram of the input module; Figure 4 It is a schematic diagram of the image-text feature fusion module; Figure 5 Schematic diagram of feature decoder; Figure 6 This is a schematic diagram of the output module; Figure 7 This is an example of a defect image of a drone inspection device; Figure 8 This is an example diagram of power equipment images; Fig. 9 Schematic diagram of wall cracks; Fig.10 For open world image examples; Fig.11 This is a schematic diagram of the visible defect recognition results; Fig.12 This is a schematic diagram of the invisible class recognition results; Fig.13 A schematic diagram of a computer device provided by an embodiment of the present invention; Fig.14 The present invention is a block diagram of an electronic device provided according to an embodiment. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] In the description of the present invention, it should be understood that the terms “include” and “comprises” indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0042] It should also be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0043] It should be further understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects are in an "or" relationship.
[0044] It should be understood that, although the terms first, second, third, etc. may be used to describe preset ranges, etc. in the embodiments of the present invention, these preset ranges should not be limited to these terms. These terms are only used to distinguish preset ranges from each other. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.
[0045] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0046] Various structural schematic diagrams of the embodiments disclosed in the present invention are shown in the accompanying drawings. These figures are not drawn to scale, and some details are magnified and some details may be omitted for the purpose of clear expression. The shapes of various regions and layers shown in the figures and the relative sizes and positional relationships therebetween are only exemplary, and may deviate in practice due to manufacturing tolerances or technical limitations, and those skilled in the art may additionally design regions / layers with different shapes, sizes, and relative positions according to actual needs.
[0047] The present invention provides an open domain drone inspection equipment defect detection method, which adopts a multimodal image-text fusion model based on drone inspection images, supports defect recognition in open scenarios, utilizes the strong semantic characteristics of multimodal text features, and combines image-text feature fusion, self-attention mechanism, and contrastive learning methods to improve the defect recognition capability of power equipment in open scenarios. Compared with other defect recognition methods based on image feature extraction, the method of the present invention has text input capability and can detect defects of any equipment based on text prompts entered by the user.
[0048] The present invention provides an open domain unmanned aerial vehicle inspection equipment defect detection method, comprising the following steps: S1. Constructing a multimodal defect recognition network structure See also Figure 2,The multimodal defect recognition network structure includes input module, feature extraction module, image and text feature fusion module, feature decoder and output module, as follows: Input Module See also Figure 3 ,The input module includes two parts: image input and text input.
[0049] Image input refers to the drone inspection images and depth maps to be identified; text input is the user inputting text descriptions of equipment defects based on actual scenarios. These descriptions are phrases or single words.
[0050] For text modality, the advantage of text modality input is that it can express rich semantic information and abstract concepts. Text can accurately describe specific objects, relationships, emotions, and reasoning, making up for the shortcomings of other modalities in semantic understanding. In addition, text input is highly flexible and can adapt to different contexts, providing more precise instructions and descriptions, and enhancing the system's semantic understanding and reasoning capabilities.
[0051] For image modalities, the advantage of fusing visible light images with depth maps is that it can significantly enhance the spatial perception of the system. The depth map provides distance information of objects in three-dimensional space, which can help the model understand the scene structure and the relative position between objects more accurately. After being fused with visible light image features, it can make up for the shortcomings of a single image modality in terms of position and perspective changes, and enhance the robustness of the target. In addition, the depth map has a strong ability to handle occluded objects, and can more effectively identify partially occluded targets, improve the accuracy and reliability of recognition, and is suitable for environments with complex backgrounds.
[0052] Feature extraction module In the feature extraction module, BERT is used to extract text features. The BERT model adopts the Transformer architecture and learns deep language representation through large-scale unsupervised pre-training tasks, so that the model can better associate the feature information between images and texts, and has a deeper semantic understanding, which helps to achieve better transfer learning effects and facilitates subsequent image and text feature fusion.
[0053] For the input image, Swin Transformer is used to extract the corresponding features. Through the Transformer structure and cross-layer connections, it can better capture long-distance dependencies and improve the ability to understand the global information of the image. At the same time, it adopts the local attention mechanism and windowed self-attention mechanism to effectively reduce the computational complexity and maintain high efficiency when processing large-size images.
[0054] For the input depth image, the parameter-sharing Swin Transformer is used to extract features, and the fusion method is as follows: (1) in, represents the deep image features, Represents visible light image features, Flatten represents feature flattening operation, Concat represents feature concatenation operation, , , ; Indicates the output feature map of 8, 16, and 32 times downsampling in Swin Transformer.
[0055] Image and text feature fusion module See also Figure 4 ,After the features of the image and text are extracted, the features of the two modalities need to be fused.
[0056] For image features, the feature maps of different resolutions of Swin Transformer are first concatenated and flattened to align with the dimensions of text features for easy feature fusion. Afterwards, the Q, K, and V matrices of the image features and text features are obtained respectively. The Q matrix of the image features and the K and V matrices of the text features are used as the input of the image-text cross-modal self-attention. The Q matrix of text features and the K and V matrices of image features serve as the input of the text-image cross-modal self-attention. After obtaining the features of the two image-text fusions, they are input into the self-attention module respectively, and then the fused image features and text features are obtained through the FFN feedforward neural network.
[0057] The fusion method of text features and image features is as follows: (2) in, The Q matrix represents the image features, The Q matrix represents the text features, The K matrix and V matrix represent the text features respectively; The K matrix and V matrix represent the image features respectively.
[0058] Feature Decoder See also Figure 5 , the feature decoder is mainly used to obtain the position feature representation of the candidate box, as follows: First, a set of initial candidate box Queries is obtained through the Embedding layer, and the Q, K, and V matrices are obtained and input into the self-attention module, whose output is used as the Q matrix. At the same time, the K and V matrices of the image and text features output by the previous image-text fusion module are obtained respectively. Then, the Q matrix is fused with the K and V matrices of the text features using the Text CrossAttention method. Then, the output Q matrix is fused with the K and V matrices of the image features using ImageCross Attention. After passing through FFN (feed-forward neural network), the refined candidate box feature vector is output.
[0059] Self-Attention and Cross-Attention are calculated as follows: (3) Among them, Q represents the Q matrix of Queries, K and V represent the K matrix and V matrix of text or image features.
[0060] Output Module See also Figure 6 The output module is used to output the category and location information of equipment defects and supports defect recognition in open scenarios. Therefore, the model outputs specific defect categories based on the user's input prompt words. The input text features are compared and matched with the image features corresponding to the candidate boxes. Finally, by setting the score threshold hyperparameter, the model outputs the category and location coordinates of the candidate boxes whose similarity with the text feature matching is higher than the threshold, and draws specific defect information in the picture.
[0061] S2. Construct a multimodal model training data set, and input the multimodal model training data set into the multimodal defect recognition network structure constructed in step S2 for training; after the training is completed, input the visible light image, depth image and text prompt word respectively, and the model outputs the defect recognition result after forward propagation, including category and coordinate information.
[0062] Dataset composition See also Figure 7, Defective images: By collecting and summarizing the drone inspection images of transmission lines and distribution lines of various provincial companies across the country and carrying out data annotation, the annotation method is divided into two parts: equipment defect description and scene description. The equipment defect description is to annotate all the defects in the image. The annotation content is generally composed of the defect name and the equipment name. For example, the damaged insulator is marked as damaged insulator. The scene description is an overall description of the defects in the entire image. There are damaged insulators and cracked walls on the pole, which are marked as: There are defects such as damaged insulators and cracked walls on the pole.
[0063] See also Figure 8 , Power equipment images: By collecting and summarizing the drone inspection images of transmission lines and distribution lines of provincial companies across the country and carrying out data annotation of power equipment, we mainly annotate common power equipment. The annotation method is divided into English descriptions of equipment, such as disc insulators are marked as disc insulators, glass insulators are marked as glassinsulators, and transformers are marked as Transformer.
[0064] See also Fig. 9 ,General defect images: In addition to drone inspection images, there are also a large number of defects in other scenes. Some defects, such as wall cracks and surface dirt, have characteristics that are highly similar to some equipment defects (such as wall cracks on towers). Adding these images with similar characteristics to the training data will help improve the model's ability to identify equipment defect characteristics. The annotation method uses general defect word annotation, and wall cracks are annotated as Crack.
[0065] See also Fig.10 ,Open world images: In order to improve the versatility of the model in open scenarios and support diverse input texts, this method introduces open source datasets into the training set in a certain proportion for mixed training, such as the general target detection dataset Object 365, the image description dataset Flickr30k, etc. Image examples are as follows Fig.10 As shown, the image is labeled Several people are milling about outside a large off-white colored house with a terra cotta roof while someone in the distance takes their picture.
[0066] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as systems, methods or program products. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "platforms".
[0067] In yet another embodiment of the present invention, a system for defect detection of open domain UAV inspection equipment is provided. The system can be used to implement the above-mentioned method for defect detection of open domain UAV inspection equipment. Specifically, the system for defect detection of open domain UAV inspection equipment includes a network module, a training module and a detection module.
[0068] Among them, the network module builds a multimodal defect recognition network structure; A training module uses a training data set to train the constructed multimodal defect recognition network structure; The detection module inputs the drone image and text to be identified into the trained multimodal defect recognition network structure and outputs the category and location information of the equipment defect.
[0069] In another embodiment of the present invention, a terminal device is provided, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, graphics processors (GPU), tensor processors (TPU), digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present invention can be used for the operation of the defect detection method of the open domain unmanned aerial vehicle inspection equipment, including: Construct a multimodal defect recognition network structure; use the training data set to train the constructed multimodal defect recognition network structure; input the drone image and text to be identified into the trained multimodal defect recognition network structure, and output the category and location information of the equipment defect.
[0070] See also Fig.13 , the terminal device is a computer device. The computer device 60 of this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the computer program 63 is executed by the processor 61, the open domain drone inspection equipment defect detection method in the embodiment is implemented. To avoid repetition, it is not described here one by one. Alternatively, when the computer program 63 is executed by the processor 61, the functions of each model / unit in the open domain drone inspection equipment defect detection system in the embodiment are implemented. To avoid repetition, it is not described here one by one.
[0071] The computer device 60 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will appreciate that Fig.13 This is only an example of the computer device 60 and does not constitute a limitation of the computer device 60. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, etc.
[0072] The processor 61 may be a central processing unit (CPU), or other general-purpose processors, graphics processing units (GPU), tensor processing units (TPU), digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc.
[0073] The memory 62 may be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 may also be an external storage device of the computer device 60, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the computer device 60.
[0074] Furthermore, the memory 62 may include both an internal storage unit of the computer device 60 and an external storage device. The memory 62 is used to store computer programs and other programs and data required by the computer device. The memory 62 may also be used to temporarily store data that has been output or is to be output.
[0075] See also Fig.14 The terminal device is an electronic device 600, which is in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0076] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps according to various exemplary embodiments of the present invention described in the above method section of this specification. For example, the processing unit 610 can perform the following steps: Figure 2 Follow the steps shown in .
[0077] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0078] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0079] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0080] The electronic device 600 may also communicate with one or more external devices 700 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 660. The network adapter 660 may communicate with other modules of the electronic device 600 via a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms, etc.
[0081] In another embodiment of the present invention, the present invention also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device, and can be any tangible medium containing or storing a program, which can be used by an instruction execution system, device or device or used in combination with it. The computer-readable storage medium provides a storage space, which stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that more specific examples (non-exhaustive list) of the computer-readable storage medium here include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0082] Computer readable storage media also include data signals propagated in baseband or as part of a carrier wave, which carry readable program codes. Such propagated data signals can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable storage medium can also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or device. The program code contained on the readable storage medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.
[0083] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0084] The processor may load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the open domain drone inspection equipment defect detection method in the above embodiment; the processor may load and execute the following steps: Construct a multimodal defect recognition network structure; use the training data set to train the constructed multimodal defect recognition network structure; input the drone image and text to be identified into the trained multimodal defect recognition network structure, and output the category and location information of the equipment defect.
[0085] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present invention described and shown in the drawings here can usually be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0086] See also Fig.11 , showing the recognition results based on the web page visualization interface. The input prompt words are the fixed input categories preset by the model, and obvious recognition effects can be seen.
[0087] See also Fig.12 , showing the recognition effect under open vocabulary input. The category corresponding to the input vocabulary is not in the training set. It can be seen that the model has the ability to generalize to new categories.
[0088] The present invention uses the following evaluation indicators to evaluate the effectiveness of the model: defect discovery rate, defect false detection ratio, and defect accuracy rate.
[0089] Defect Detection Rate :
[0090] Among them, FR represents the defect detection rate, It represents the number of correct detection boxes (IoU>0.5) output by the model, and B represents the total number of annotated boxes. The FR value range is 0~1. The closer it is to 1, the better the defect detection effect of the model.
[0091] Defect False Detection Ratio :
[0092] Among them, ER represents the defect false detection ratio, is the total number of error boxes (IoU < 0.5) output by the model, and B is the total number of labeled boxes. The ER value range is The closer the value is to 0, the better the model's ability to classify defects is.
[0093] Defect recognition accuracy: Since the discovery rate and false positive rate are two different evaluation methods, it is impossible to evaluate the effectiveness at the same level. In order to evaluate the discovery rate and false positive rate at the same time and make it possible to evaluate the effectiveness of the model more accurately, this paper proposes a new comprehensive accuracy calculation index:
[0094] Among them, k is the coefficient, FR represents the defect detection rate, and ER represents the defect false detection ratio.
[0095] The value range of the comprehensive accuracy calculation index is 0~1. The closer the result is to 1, the better the model effect is. In order to balance the false detection rate and the discovery rate and better meet the needs of business application scenarios, the present invention takes a value of 16.
[0096] Table 1 shows the defect recognition comparison results of this method after fine-tuning and training on power distribution data and the current mainstream recognition model.
[0097] It can be seen from Table 1 that the method of the present invention has achieved the best results in the detection rate, false positive rate and accuracy rate.
[0098] In summary, the present invention provides an open domain UAV inspection equipment defect detection method and system, which can improve the accuracy of power equipment defect detection in open scenarios and improve the level of practicality. The image features and text features are fully integrated through the cross-modal self-attention module Images-Text Cross Attention Module (ITCAM), thereby improving the model's ability to utilize multimodal features; the Depth Images Fusion Module (DIFM), a visible light image and depth image feature fusion method, is adopted to integrate the advantages of the two image modal features and utilize the difference information between the modalities to improve the model's feature expression ability; a new evaluation index is proposed based on the discovery rate and false detection ratio to solve the problem that the indicators under the two dimensions of discovery rate and false detection ratio cannot effectively evaluate the effectiveness of the model.
[0099] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0100] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0101] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0102] In the embodiments provided by the present invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0103] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0105] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0106] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0107] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0109] The above contents are only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A method for detecting defects in open domain drone inspection equipment, characterized in that: The following steps are involved: Construct a multimodal defect recognition network structure; The multimodal defect recognition network structure constructed by training the training data set; The drone images and texts to be identified are input into the trained multimodal defect recognition network structure, and the category and location information of the equipment defects are output.
2. The open domain drone inspection equipment defect detection method according to claim 1 is characterized in that: The multimodal defect recognition network structure includes: An input module inputs drone images and texts to be identified. The drone images include drone inspection images and depth maps to be identified. The texts include text descriptions of equipment defects input by users according to actual scenarios. The text descriptions include phrases or single words. Feature extraction module, using BERT to extract text features and Swin Transformer to extract corresponding image features; The image-text feature fusion module fuses the image and text features, obtains the features of the two image-text fusions, and inputs them into the self-attention module respectively, and obtains the fused image features and text features through the FFN feedforward neural network; Feature decoder, obtains the position feature representation of the candidate box; The output module compares and matches the input text features with the image features corresponding to the candidate boxes. By setting the score threshold hyperparameter, the model outputs the candidate box category and position coordinates whose similarity with the text feature matching is higher than the threshold, and draws specific defect information in the image.
3. The open domain drone inspection equipment defect detection method according to claim 2 is characterized in that: The BERT model uses the Transformer architecture and learns deep language representation through large-scale unsupervised pre-training tasks; SwinTransformer uses the Transformer structure and cross-layer connections, and adopts local attention mechanism and windowed self-attention mechanism. For the input deep image, the parameter-shared Swin Transformer is used to extract features.
4. The open domain drone inspection equipment defect detection method according to claim 3 is characterized in that: The specific features extracted by Swin Transformer using parameter sharing are: in, represents the deep image features, Represents the visible light image features.
5. The open domain drone inspection equipment defect detection method according to claim 2 is characterized in that: The specific feature fusion is: For image features, feature concatenation and flattening are performed on the feature maps of different resolutions of Swin Transformer; The Q, K, and V matrices of image features and text features are obtained respectively. The Q matrix of image features and the K and V matrices of text features are used as the input of cross-modal self-attention from image to text features. The Q matrix of text features and the K and V matrices of image features are used as inputs for cross-modal self-attention from text to image features; After obtaining the features of the two image-text fusions, they are input into the self-attention module respectively, and then the fused image features and text features are obtained through the feedforward neural network.
6. The open domain drone inspection equipment defect detection method according to claim 5 is characterized in that: The fusion method of text features and image features is as follows: in, The Q matrix represents the image features, The Q matrix represents the text features, The K matrix and V matrix represent the text features respectively; The K matrix and V matrix represent the image features respectively.
7. The open domain drone inspection equipment defect detection method according to claim 6 is characterized in that: and The calculation is as follows: in, To fuse the image features with text information, It is the text feature that integrates image information.
8. The open domain drone inspection equipment defect detection method according to claim 2 is characterized in that: The position feature representation of the candidate box is as follows: Through the Embedding layer, we get a set of initial candidate box Queries, get the Q, K, V matrices, and input them into the self-attention module, whose output is used as the Q matrix. At the same time, we get the K and V matrices of the image and text features output by the previous image-text fusion module respectively. The Q matrix is fused with the K and V matrices of text features for text cross-modal self-attention; The output Q matrix is fused with the K and V matrices of the image features through image cross-modal self-attention; After passing through the feedforward neural network, the refined candidate box feature vector is output.
9. The open domain drone inspection equipment defect detection method according to claim 1, characterized in that: The training dataset includes: Defect images, by collecting and summarizing drone inspection images of transmission lines and distribution lines and carrying out data annotation; Power equipment images: by collecting and summarizing drone inspection images of transmission lines and distribution lines and carrying out data annotation of power equipment, common power equipment is annotated; General defect images, defects in other scenes except drone inspection images, partial defects; Open world images, introduce open source datasets into the training set for mixed training in a set ratio.
10. The open domain drone inspection equipment defect detection method according to claim 9, characterized in that: The annotation methods of defect images are divided into equipment defect description and scene description. Equipment defect description is to annotate all defects in the image, and the annotation content includes defect name and equipment name; scene description is an overall description of the defects in the entire image; the annotation method of power equipment images is the English description of the equipment; the annotation method of general defect images is to use general defect words.
11. An open domain drone inspection equipment defect detection system, characterized in that: include: Network module, building a multimodal defect recognition network structure; A training module uses a training data set to train the constructed multimodal defect recognition network structure; The detection module inputs the drone image and text to be identified into the trained multimodal defect recognition network structure and outputs the category and location information of the equipment defect.
12. A computer-readable storage medium storing one or more programs, characterized in that: The one or more programs include instructions, which, when executed by a computing device, cause the computing device to perform the method of any one of claims 1 to 10.
13. A computing device, characterized in that: include: One or more processors, a memory and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include steps for executing the method according to any one of claims 1 to 10.
Citation Information
Cited By
Open type question automatic scoring method and system based on multi-modal large model and target detection
CN120635927A
Industrial defect image-text joint detection method and related equipment
CN120707521A
Power transmission line defect detection method and system
CN120765573A