High-altitude insulator detection method and system based on visual image
By combining the multi-level feature fusion technology of CNN convolutional neural network and self-attention Transformers module, the problem of insufficient robustness of traditional detection methods in complex environments is solved, efficient insulator defect detection is achieved, and detection accuracy and reliability are improved.
Patent Information
- Application Number
- CN202510351978.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional insulator defect detection methods are not robust enough in complex environments, making it difficult to effectively detect insulator defects with small target features, limiting the application of drone detection.
The high-altitude insulator detection method based on visual images is adopted, combined with the CNN convolutional neural network and the self-attention Transformers module, and the accuracy and robustness of the detection are improved through multi-level feature fusion and dynamic offset subnetwork.
It significantly improves the accuracy of insulator defect detection under complex background and light changes, can effectively classify and locate multiple defects, and improves the reliability of detection and fault handling efficiency.
Smart Images

Figure CN120298852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for detecting high-altitude insulators based on visual images, belonging to the technical field of power equipment defect detection. Background Art
[0002] With the continuous steady expansion of the scale of transmission lines. As a key component of transmission lines, insulators are often exposed to harsh natural environments, prone to aging, flashover or damage, posing a significant threat to the stable operation of the power grid. Once an insulator is damaged, it will lead to the interruption of power transmission and may cause short circuits and large-scale outages of power lines. In addition, the damage of insulators may cause current to flow from both poles to the ground, increasing the risk of electric shock and fire hazards. Therefore, insulator defect detection has become an important part of the routine inspection of transmission lines, with important engineering significance and research value.
[0003] Traditional insulator defect detection methods usually rely on feature engineering, extracting features such as color and edges, and then using algorithms such as edge detection, HOG or SIFT for identification. Although these methods can achieve defect detection to a certain extent, their performance highly depends on high-quality images and shooting angles, which limits their robustness in complex environments. With the rapid development of drone technology, drone detection, known for its flexibility and efficiency, is gradually replacing traditional manual detection methods. However, the limitations of traditional detection algorithms restrict their wide application in drone inspections.
[0004] In recent years, with the rapid development of deep learning technology and computing hardware, more and more researchers have adopted deep learning models for insulator defect detection. Deep learning-based methods overcome the limitations of traditional feature engineering and demonstrate powerful feature extraction capabilities in complex scenarios. However, there are still challenges in insulator defect detection mainly focusing on small target features. Summary of the Invention
[0005] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a method and system for detecting high-altitude insulators based on visual images, aiming to improve the accuracy and robustness of insulator defect detection.
[0006] The technical solution of the present invention is as follows:
[0007] On the one hand, the present invention provides a method for detecting high-altitude insulators based on visual images, including the following steps:
[0008] Collect the original images containing high-altitude insulators and perform preprocessing, mark the defect types and position data of the high-altitude insulators in the original images, and construct a training sample set through the preprocessed data;
[0009] Construct a visual detection model, including a CNN convolutional neural network module and a self-attention Transformers module. The CNN convolutional neural network module is used to extract hierarchical features of the input data, and the self-attention Transformers module outputs prediction data through the self-attention mechanism and the input hierarchical features;
[0010] Iteratively train the visual detection model with a training sample set to obtain a trained high-altitude insulator defect detection model;
[0011] Input the high-altitude insulator image obtained from the inspection into the high-altitude insulator defect detection model, output the defect type and position data of the target high-altitude insulator, and output the target image of the high-altitude insulator according to the position data.
[0012] As a preferred embodiment, the CNN convolutional neural network module uses a multi-level convolutional neural network to extract multi-level features of the input data, and uses a cross-stage fusion method to fuse features of different levels into an integrated feature, and projects the integrated feature to a fixed dimension through linear projection as the input of the self-attention Transformers module.
[0013] As a preferred embodiment, in the self-attention Transformers module:
[0014] Use learnable feature matrices W Q 、W K and W V to map the input integrated feature into a query vector, a key vector, and a value vector, specifically as follows:
[0015] Q = XW Q , K = XW K , V = XW V ;
[0016] where Q, K, and V are the query vector, the key vector, and the value vector respectively, and X is the integrated feature;
[0017] The calculation formula of the self-attention Transformer module is as follows:
[0018]
[0019] where SAtransformer(·) represents the self-attention Transformer calculation function, and Softmax(·) represents the Softmax normalization function.
[0020] As a preferred embodiment, a dynamic offset sub-network is also introduced in the self-attention Transformers module to adjust the input feature mapping, and the specific steps are as follows:
[0021] Calculate the offset value through the dynamic offset sub-network:
[0022] Δp = θ offset (Q);
[0023] where θ offset () represents the dynamic offset sub-network, and Δp represents the offset value calculated through the query vector;
[0024] Update the input features according to the offset value:
[0025] X' = n(X → p + Δp);
[0026] where X' is the updated integrated feature, and n(·) represents the bilinear interpolation function;
[0027] Update the key vector and value vector through the updated integrated feature:
[0028]
[0029] where, is the updated key vector, is the updated value vector.
[0030] On the other hand, the present invention also provides a high-altitude insulator detection system based on visual images, including:
[0031] A dataset construction module for collecting original images containing high-altitude insulators and performing preprocessing, marking the defect types and position data of high-altitude insulators in the original images, and constructing a training sample set through the preprocessed data;
[0032] A model construction module for constructing a visual detection model, including a CNN convolutional neural network module and a self-attention Transformers module. The CNN convolutional neural network module is used to extract hierarchical features of the input data, and the self-attention Transformers module outputs prediction data through the self-attention mechanism and the input hierarchical features;
[0033] A model training module for iteratively training the visual detection model through the training sample set to obtain a trained high-altitude insulator defect detection model;
[0034] An insulator detection module for inputting the high-altitude insulator images obtained from the inspection into the high-altitude insulator defect detection model, outputting the defect types and position data of the target high-altitude insulators, and outputting the target images of high-altitude insulators according to the position data.
[0035] As a preferred embodiment, the CNN convolutional neural network module extracts multi-level features of the input data using a multi-level convolutional neural network, and fuses the features at different levels into integrated features in a cross-stage fusion manner, and projects the integrated features through linear projection into a fixed dimension as the input of the self-attention Transformers module.
[0036] As a preferred embodiment, in the self-attention Transformers module:
[0037] Use learnable feature matrices W Q , W K and W V to map the input integrated features into query vectors, key vectors, and value vectors, specifically as follows:
[0038] Q = XW Q , K = XW K , V = XW V ;
[0039] where Q, K, and V are the query vector, key vector, and value vector respectively, and X is the integrated feature;
[0040] The calculation formula of the self-attention Transformer module is as follows:
[0041]
[0042] where SAtransformer(·) represents the self-attention Transformer calculation function, and Softmax(·) represents the Softmax normalization function.
[0043] As a preferred embodiment, a dynamic offset sub-network is also introduced in the self-attention Transformers module to adjust the input feature mapping, and the specific steps are as follows:
[0044] Calculate the offset value through the dynamic offset sub-network:
[0045] Δp = θ offset (Q);
[0046] where θ offset () represents the dynamic offset sub-network, and Δp represents the offset value calculated through the query vector;
[0047] Update the input features according to the offset value:
[0048] X' = n(X → p + Δp);
[0049] where X' is the updated integrated feature, and n(·) represents the bilinear interpolation function;
[0050] Update the key vector and value vector with the updated integrated features:
[0051]
[0052] Wherein, is the updated key vector, is the updated value vector.
[0053] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for detecting high-altitude insulators based on visual images as described in any embodiment of the present invention.
[0054] On the other hand, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for detecting high-altitude insulators based on visual images as described in any embodiment of the present invention.
[0055] The beneficial effects of the present invention are as follows:
[0056] The present invention utilizes the multi-scale feature fusion technology, combines the efficient feature extraction ability of the convolutional network of deep learning with the global context capture ability of self-attention Transformers, can effectively process various insulator defects (such as cracks, aging, fouling, etc.) in special scenarios such as blur, occlusion, or small targets, and can classify and locate each defect. Especially in the case of complex backgrounds or changing lighting conditions, the robustness and reliability of detection are significantly improved. It provides more detailed diagnostic information for maintenance personnel and further improves the fault handling efficiency.
[0057] The additional aspects and advantages of the present invention will be clarified in the following description, and some of them will be obvious from the description, or can be understood by practicing the present invention. In addition, the various aspects and advantages of the present invention can be realized and obtained through the method steps and combinations specifically pointed out in the appended claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic flowchart of the method according to Embodiment 1 of the present invention;
[0059] Figure 2 is a schematic diagram of the model structure of the visual detection model in the embodiment of the present invention;
[0060] Figure 3 is a schematic diagram of the algorithm principle of the spatial fast pyramid in the embodiment of the present invention;
[0061] Figure 4 is a schematic diagram of the principle of multi-scale feature fusion in the embodiment of the present invention. Specific Embodiments
[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0063] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0064] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0065] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0066] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0067] Embodiment 1:
[0068] Refer to Figure 1 , this embodiment provides a method for detecting high-altitude insulators based on visual images, including the following steps:
[0069] S100. Use the visual sensor carried by the unmanned aerial vehicle platform to collect the original images containing high-altitude insulators and perform preprocessing, mark the defect types and position data of the high-altitude insulators in the original images, and construct a training sample set through the preprocessed data;
[0070] S200. Construct a visual detection model. The schematic diagram of the model structure of the visual detection model constructed in this embodiment is specifically as Figure 2 shown, including a CNN convolutional neural network module and a self-attention Transformers module. The CNN convolutional neural network module is used to extract the hierarchical features of the input data, and the self-attention Transformers module outputs the prediction data through the self-attention mechanism and the input hierarchical features.
[0071] After obtaining the hierarchical features, global feature modeling is performed through a Transformer, and learnable object queries are introduced to enable the network to directly predict the object category and bounding box for detection on the fused features. The implementation details are as follows:
[0072] On the parsed hierarchical features, a detection head is added to predict: object category, object location, and object confidence. Here, the detection head uses a Transformer for global feature modeling and directly predicts the center point, size, and offset of the target without using a pre-anchoring point. This process uses the Transformer to directly predict the target box and assigns the target through Hungarian matching.
[0073] The specific process is as follows: First, hierarchical features are extracted through convolution. A feature map of size H×W×C is obtained, where C is the number of channels, and H and W are the spatial dimensions of the feature map. The feature map is flattened into a series of serialized embedding vectors, and position encoding is added, then input into the Transformer for encoding. The self-attention mechanism is added. After variable Transformer decoding, each query predicts the category and bounding box of a target. Then, the Hungarian algorithm is used and optimized based on classification loss, L1 loss, and CIoU loss to ensure that each predicted box uniquely corresponds to a real target, thus avoiding redundant object detection. Finally, object detection is completed, making full use of the global modeling ability of the Transformer and combining multi-scale feature fusion to achieve accurate recognition and positioning of the target.
[0074] This embodiment adopts a multi-scale feature fusion technology, which combines the efficient feature extraction ability of the convolutional network in deep learning with the global context capture ability of self-attention Transformers to improve the small object detection performance.
[0075] S300. Iteratively train the visual detection model with the training sample set to obtain a trained high-altitude insulator defect detection model;
[0076] S400. Input the high-altitude insulator image obtained from the inspection into the high-altitude insulator defect detection model, output the defect type and position data of the target high-altitude insulator, and output the high-altitude insulator target image according to the position data; after detecting the insulator, generate a detection report, including the insulator defect type, insulator position, and image schematic diagram.
[0077] For specific reference, see Figure 4, in this embodiment, the CNN convolutional neural network module uses a multi-level convolutional neural network to extract multi-level features of the input data, and uses cross-stage feature fusion to fuse features at different levels into integrated features, enhancing the representation of key regions. Partial connections are used to improve the model's ability to capture key features to ensure a robust target region expression. And the integrated features are mapped to a fixed dimension through linear projection as the input of the self-attention Transformers module. Thereby enhancing global context feature learning. In addition, a fast spatial pyramid structure is introduced to extract detailed information at a unified scale, combining different receptive fields to improve the description of small objects and the representation of local details.
[0078] Specifically refer to Figure 3 , the empty fast spatial pyramid adds batch normalization and the ReLU activation function to the depth convolutional layer of the input CNN convolutional neural network module. It is used to perform convolutional operations on the input feature map to extract features and batch normalization to speed up the training process and improve the stability of the model, while the ReLU activation function introduces non-linearity to enhance the model's expressive ability. Through multiple max-pooling operations and concatenation operations, feature fusion at different scales is achieved, and finally spatial expansion convolutions with different expansion rates are performed to further extract features. These convolutional layers can cover a larger receptive field and capture a wider range of context information without increasing the computational cost. By operating along the horizontal and vertical directions respectively, they allow for more detailed processing of image features and enhance the model's understanding of the spatial relationships in the image.
[0079] As a preferred implementation of this embodiment, in the self-attention Transformers module:
[0080] Use the learnable feature matrices W Q , W K and W V to map the input integrated features into query vectors, key vectors, and value vectors, specifically as follows:
[0081] Q = XW Q , K = XW K , V = XW V ;
[0082] where Q, K, and V are the query vector, key vector, and value vector respectively, and X is the integrated feature; W Q , W K and W V are the query vector feature matrix, key vector feature matrix, and value vector feature matrix respectively.
[0083] Among them, by performing a linear transformation on the integrated features, a query vector (Q) is obtained, which represents the information queried by each hierarchical feature marker. The key vector (K) is also obtained through the linear transformation of the input features and represents the key information of all hierarchical features. The value vector (V) is also obtained from the integrated features and is used to characterize the actual feature values for updating the integrated features.
[0084] The calculation formula of the self-attention Transformer module is as follows:
[0085]
[0086] Among them, SAtransformer(·) represents the self-attention Transformer calculation function, and Softmax(·) represents the Softmax normalization function.
[0087] Here, the similarity between the query vector and all key vectors is calculated through the dot product between Q and K to form a correlation matrix. The shape of this matrix represents the relationship between all feature points in the sequence length. Then, the correlation matrix is scaled by the square root of the feature dimension d k to prevent unstable gradients caused by overly large dot product values. Next, we perform a Softmax operation on the scaled correlation matrix to generate a probability distribution representing the attention weights of each feature marker. Finally, the value vector V is weighted using the normalized attention weights to obtain the attention feature representation.
[0088] As a preferred implementation manner of this embodiment, a dynamic offset sub-network is further introduced in the self-attention Transformers module to adjust the input feature mapping, which can flexibly focus on relevant regions, capture more information features, enable the network to dynamically adjust the attention focus according to the input feature map instead of being fixed on a grid or a fixed position, and significantly improve the detection of small target objects. The specific steps are as follows:
[0089] First, through the introduced dynamic offset sub-network, the offset value is calculated:
[0090] Δp = θ offset (Q);
[0091] where θ offset () represents the dynamic offset sub-network, Δp represents the offset value calculated through the query vector, and the offset value is dynamically generated and constrained by the tanh function to avoid excessive offset.
[0092] Next, the input features are updated according to the offset value:
[0093] X' = n(X → p + Δp);
[0094] Among them, X' is the updated integrated feature, and n(·) represents a bilinear interpolation function used to calculate the feature difference at non-integer positions to ensure the microscopic nature of the operation.
[0095] The updated integrated feature is mapped into the key vector and value vector to form a deformable Transformer module:
[0096]
[0097] Among them, is the updated key vector, is the updated value vector.
[0098] Embodiment 2:
[0099] This embodiment provides an overhead insulator detection system based on visual images, including:
[0100] A dataset construction module for collecting and preprocessing the original images containing overhead insulators, marking the defect types and position data of the overhead insulators in the original images, and constructing a training sample set through the preprocessed data; this module is used to implement the function of step S100 in Embodiment 1 and will not be elaborated here;
[0101] A model construction module for constructing a visual detection model, including a CNN convolutional neural network module and a self-attention Transformers module. The CNN convolutional neural network module is used to extract the hierarchical features of the input data, and the self-attention Transformers module outputs prediction data through the self-attention mechanism and the input hierarchical features; this module is used to implement the function of step S200 in Embodiment 1 and will not be elaborated here;
[0102] A model training module for iteratively training the visual detection model through the training sample set to obtain a trained overhead insulator defect detection model; this module is used to implement the function of step S300 in Embodiment 1 and will not be elaborated here;
[0103] An insulator detection module for inputting the overhead insulator images obtained from the inspection into the overhead insulator defect detection model, outputting the defect types and position data of the target overhead insulators, and outputting the overhead insulator target images according to the position data; this module is used to implement the function of step S400 in Embodiment 1 and will not be elaborated here.
[0104] As a preferred implementation manner of this embodiment, the CNN convolutional neural network module extracts multi-level features of the input data by using a multi-level convolutional neural network, and fuses the features of different levels into integrated features by using a cross-stage fusion method, and maps the integrated features to a fixed dimension through linear projection as the input of the self-attention Transformers module.
[0105] As a preferred implementation manner of this embodiment, in the self-attention Transformers module:
[0106] Use the learnable feature matrices W Q 、W K and W V to map the input integrated features into query vectors, key vectors, and value vectors, specifically as follows:
[0107] Q = XW Q , K = XW K , V = XW V ;
[0108] where Q, K, and V are the query vector, key vector, and value vector respectively, and X is the integrated feature;
[0109] The calculation formula of the self-attention Transformer module is as follows:
[0110]
[0111] where SAtransformer(·) represents the self-attention Transformer calculation function, and Softmax(·) represents the Softmax normalization function.
[0112] As a preferred implementation manner of this embodiment, a dynamic offset sub-network is further introduced in the self-attention Transformers module to adjust the input feature mapping, and the specific steps are as follows:
[0113] Calculate the offset value through the dynamic offset sub-network:
[0114] Δp = θ offset (Q);
[0115] where θ offset () represents the dynamic offset sub-network, and Δp represents the offset value calculated through the query vector;
[0116] Update the input features according to the offset value:
[0117] X' = n(X → p + Δp);
[0118] Among them, X' is the updated integrated feature, and n(·) represents the bilinear interpolation function;
[0119] Update the key vector and value vector with the updated integrated feature:
[0120]
[0121] Among them, is the updated key vector, is the updated value vector.
[0122] Embodiment Three:
[0123] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for detecting high-altitude insulators based on visual images as described in any embodiment of the present invention.
[0124] Embodiment Four:
[0125] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for detecting high-altitude insulators based on visual images as described in any embodiment of the present invention.
[0126] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent the situation where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c may be single or multiple.
[0127] Those of ordinary skill in the art can realize that the units and algorithm steps described in the embodiments disclosed herein can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0128] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0129] In several embodiments provided in the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (hereinafter referred to as ROM), random access memory (hereinafter referred to as RAM), magnetic disks, or optical discs that can store program codes.
[0130] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A method for detecting high-altitude insulators based on visual images, characterized in that, It includes the following steps: Collect the original images containing high-altitude insulators and perform preprocessing, mark the defect types and location data of the high-altitude insulators in the original images, and construct a training sample set through the preprocessed data; Construct a visual detection model, including a CNN convolutional neural network module and a self-attention Transformers module. The CNN convolutional neural network module is used to extract hierarchical features of the input data, and the self-attention Transformers module outputs prediction data through the self-attention mechanism and the input hierarchical features; Iteratively train the visual detection model with the training sample set to obtain a trained high-altitude insulator defect detection model; Input the high-altitude insulator images obtained from the inspection into the high-altitude insulator defect detection model, output the defect types and location data of the target high-altitude insulators, and output the target images of the high-altitude insulators according to the location data.
2. The method for detecting high-altitude insulators based on visual images according to claim 1, wherein, The CNN convolutional neural network module uses a multi-level convolutional neural network to extract multi-level features of the input data, and uses a cross-stage fusion method to fuse features of different levels into integrated features, and projects the integrated features to a fixed dimension through linear projection as the input of the self-attention Transformers module.
3. The method for detecting high-altitude insulators based on visual images according to claim 2, wherein, In the self-attention Transformers module: Use the learnable feature matrix W Q , W K and W V to map the input integrated features into query vectors, key vectors, and value vectors as follows: Q = XW Q , K = XW K , V = XW V ; Among them, Q, K, and V are the query vector, key vector, and value vector respectively, and X is the integrated feature; The calculation formula of the self-attention Transformer module is as follows: Among them, SAtransformer(·) represents the self-attention Transformer calculation function, and Softmax(·) represents the Softmax normalization function.
4. The method for detecting high-altitude insulators based on visual images according to claim 3, characterized in that A dynamic offset sub-network is also introduced in the self-attention Transformers module to adjust the input feature mapping. The specific steps are as follows: Calculate the offset value through the dynamic offset sub-network: Δp = θ offset (Q); Among them, θ offset () represents the dynamic offset sub-network, and Δp represents the offset value calculated by the query vector; Update the input features according to the offset value: X' = n(X → p + Δp); Among them, X' is the updated integrated feature, and n(·) represents the bilinear interpolation function; Update the key vector and value vector through the updated integrated feature: Among them, is the updated key vector, is the updated value vector.
5. An overhead insulator detection system based on visual images, characterized in that It includes: A dataset construction module, which is used to collect the original images containing high-altitude insulators and perform preprocessing, mark the defect types and location data of the high-altitude insulators in the original images, and construct a training sample set through the preprocessed data; A model construction module, which is used to construct a visual detection model, including a CNN convolutional neural network module and a self-attention Transformers module. The CNN convolutional neural network module is used to extract hierarchical features of the input data, and the self-attention Transformers module outputs prediction data through the self-attention mechanism and the input hierarchical features; A model training module, which is used to iteratively train the visual detection model with the training sample set to obtain a trained high-altitude insulator defect detection model; An insulator detection module, which is used to input the high-altitude insulator images obtained from the inspection into the high-altitude insulator defect detection model, output the defect types and location data of the target high-altitude insulators, and output the target images of the high-altitude insulators according to the location data.
6. The high-altitude insulator detection system based on visual images according to claim 5, characterized in that, The CNN convolutional neural network module extracts multi-level features of the input data using a multi-level convolutional neural network, and fuses features at different levels into integrated features in a cross-stage fusion manner, and maps the integrated features to a fixed dimension through linear projection as the input of the self-attention Transformers module.
7. An aerial insulator detection system based on visual images according to claim 6, characterized in that In the self-attention Transformers module: Use the learnable feature matrix W Q , W K and W V Map the input integrated features to query vectors, key vectors, and value vectors as follows: Q = XW Q , K = XW K , V = XW V ; where Q, K, and V are the query vector, key vector, and value vector respectively, and X is the integrated feature; The calculation formula of the self-attention Transformer module is as follows: where SAtransformer(·) represents the self-attention Transformer calculation function, and Softmax(·) represents the Softmax normalization function.
8. The high-altitude insulator detection system based on visual images according to claim 7, wherein, A dynamic offset sub-network is also introduced in the self-attention Transformers module to adjust the input feature mapping, and the specific steps are as follows: Calculate the offset value through the dynamic offset sub-network: Δp = θ offset (Q); where θ offset () represents the dynamic offset sub-network, and Δp represents the offset value calculated through the query vector; Update the input features according to the offset value: X' = n(X → p + Δp); where X' is the updated integrated feature, and n(·) represents the bilinear interpolation function; Update the key vector and value vector through the updated integrated feature: Among them, is the updated key vector, is the updated value vector.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for detecting high-altitude insulators based on visual images according to any one of claims 1 to 4.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method for detecting high-altitude insulators based on visual images according to any one of claims 1 to 4.