Substation foreign object intrusion detection method and system based on improved YOLOv11 model

By improving the YOLOv11 model, the feature extraction and multi-scale detection capabilities are enhanced, and the problems of missed detection, misjudgment and slow recognition rates of small and medium-sized targets in substation equipment and line detection are solved, and efficient and accurate detection of foreign object intrusion is achieved.

CN119722662BActive Publication Date: 2025-07-08ANHUI DERUN ELECTRIC & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510213834.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-08
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art has problems such as small target miss detection, misjudgment, low recognition accuracy and slow recognition rate in substation equipment and line detection, especially in the case of occlusion, which is low recognition rate and does not adapt to multiple target sizes.

Method used

Using the improved YOLOv11 model, by adding the C2PSA-DHSA module and the C3k2-Dual module after the SPPF module, combining the non-local attention mechanism module, the feature extraction and multi-scale detection capabilities are enhanced, and a small object detection layer is added to improve the real-time performance and adaptability of the model.

Benefits of technology

It improves the accuracy and efficiency of substation equipment and line detection, can effectively identify small targets and occlusion targets, adapt to targets of different sizes, and improves the real-time performance and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119722662B_ABST
    Figure CN119722662B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and specifically discloses a substation foreign object intrusion detection method and system based on an improved YOLOv11 model. Aiming at the problem of substation foreign object intrusion detection, the method proposes a substation foreign object intrusion detection model based on an improved YOLOv11 model to achieve automatic image recognition of intrusion foreign objects. The model is improved by adding a C2PSA-DHSA module after the SPPF module, using a C3k2-Dual module, adding a new small target detection layer, and adding a non-local attention mechanism module in front of the SPPF module, etc., enabling the model to more effectively identify those targets with smaller areas. While reducing the computational cost and the number of parameters of the model, it also improves the problems of low recognition rate and misrecognition caused by occlusion of the recognition device, improves the accuracy of the model, and enhances the adaptability of the model to targets of different sizes. The present invention well solves the problems of missed detection, misjudgment, low recognition accuracy and slow rate existing in the detection process of traditional models for substation equipment and lines.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of image processing and foreign object detection in substations, and particularly relates to a method and system for detecting foreign object intrusion in substations based on an improved YOLOv11 model. Background Art

[0002] Substations are an important part of the power grid, and the safe operation of substations is crucial for the safety and stability of the entire power grid. On the equipment and lines of outdoor substations, it is common to have birds building nests and windborne objects such as plastic bags, cardboard, woven bags, and kites. These invading foreign objects can easily cause accidents such as short circuits, flashovers, and tripping of equipment and lines, seriously endangering the safe and reliable operation of the power grid. Compared with traditional methods for preventing foreign object intrusion such as manual inspection, bird spikes, and ultrasonic waves, automatic image recognition and manual removal of invading foreign objects is a very effective and economical method. Although there have been some deep learning object detection models for foreign object detection in substations in the prior art, such as YOLOv4, YOLOv5, etc. However, these models have problems such as missed detection and misjudgment of small targets during the detection process of substation equipment and lines. In addition, problems such as low recognition accuracy will also occur due to occlusion and other reasons. Moreover, there are also problems such as slow recognition speed when deploying the above models on inspection equipment for real-time detection. Summary of the Invention

[0003] The purpose of the present invention is to propose a method for detecting foreign object intrusion in substations based on an improved YOLOv11 model, so as to achieve automatic image recognition of foreign objects invading substations, avoid missed detection and misjudgment of small targets, facilitate improving the accuracy and detection efficiency of foreign object intrusion detection, and thus ensure the safe and reliable operation of substations.

[0004] In order to achieve the above purpose, the present invention adopts the following technical solutions:

[0005] A method for detecting foreign object intrusion in substations based on an improved YOLOv11 model includes the following steps:

[0006] Step 1. Obtain multiple images of substation equipment and lines, and perform image preprocessing and annotation on the obtained images of substation equipment and lines to obtain a training data set for the following foreign object intrusion detection model in substations;

[0007] Step 2. Build a foreign object intrusion detection model in substations based on the YOLOv11 model architecture. This foreign object intrusion detection model in substations includes a backbone network, a neck network, and a detection head network;

[0008] Among them, the backbone network includes a convolutional block, four groups of convolutional blocks and a C3k2-Dual module, a non-local attention mechanism module, an SPPF module, and a C2PSA-DHSA module;

[0009] Define four groups of convolutional blocks and C3k2-Dual modules as the first group, the second group, the third group, and the fourth group of convolutional blocks and C3k2-Dual modules respectively, and each group includes a convolutional block and a C3k2-Dual module;

[0010] The C3k2-Dual module is based on the traditional C3k2 module. By replacing the Conv convolution in the traditional C3k2 module with Dual Conv convolution, the model can process the same input feature map channels simultaneously and efficiently arrange the convolutional filters using group convolution technology;

[0011] The C2PSA-DHSA module is obtained by inserting the DHSA module into the C2PSA module;

[0012] The DHSA module enables the model to capture the global and local features of the power grid image simultaneously through the parallel processing of BHR and FHR, and then fuses the global and local features through the self-attention mechanism to finally achieve high-quality image restoration;

[0013] The image features input to the YOLOv11 model architecture first enter the backbone network and are processed as follows:

[0014] The image features first undergo feature extraction through a convolutional block, and then sequentially pass through the first group, the second group, the third group, and the fourth group of convolutional blocks and C3k2-Dual modules for feature extraction, and then pass through a non-local attention mechanism module, enabling the network to not only capture the local details of the equipment and line images of the input substation but also obtain global context information;

[0015] The output features of the non-local attention mechanism module then pass through the SPPF module and finally output through the C2PSA-DHSA module;

[0016] The neck network is used to achieve the deep fusion of features at different levels, and the detection head network contains four detection heads;

[0017] Step 3. Train the built substation foreign object intrusion detection model based on the training dataset in Step 1 to obtain a trained substation foreign object intrusion detection model, and use the trained model to detect the intrusion of foreign objects on the transformer.

[0018] In addition, based on the above substation foreign object intrusion detection method based on the improved YOLOv11 model, the present invention also proposes a substation foreign object intrusion detection system based on the improved YOLOv11 model, which adopts the following technical solutions:

[0019] The substation foreign object intrusion detection system based on the improved YOLOv11 model includes a camera and computer equipment. Among them, the computer equipment includes a memory and at least one processor; an executable code is stored in the memory; when the processor executes the executable code, it is used to implement the substation foreign object intrusion detection method based on the improved YOLOv11 model as described above.

[0020] The present invention has the following advantages:

[0021] As described above, the present invention relates to a substation foreign object intrusion detection method and system based on an improved YOLOv11 model. The method of the present invention constructs a substation foreign object intrusion detection model based on the improved YOLOv11 model architecture to enhance the multi-scale detection ability and feature extraction ability, improve the real-time performance of the model, improve the adaptability of the model, and at the same time improve the detection performance of small targets and occluded targets. Specifically, in terms of enhancing the multi-scale detection ability and feature extraction ability, the present invention adds a C3k2-Dual module to extract features at a deep level, greatly improving the model's extraction of substation equipment and line features. Then, before the SPPF module, a non-local attention mechanism module is added to extract features of the input feature map of substation equipment and lines in the spatial dimension and channel dimension, realizing the acquisition of global information. This method can maintain a high detection performance on targets of different scales while realizing deep feature extraction. In terms of improving the real-time performance of the model, the special structure of the C3k2-Dual module enables the model to process multiple input channels simultaneously and efficiently arrange convolution filters using grouped convolution technology, effectively reducing the computational cost and the number of parameters of the recognition model, while also improving the accuracy of the model, providing technical support for the rapid and accurate inspection of unmanned substation equipment and line inspection robots, and at the same time providing a solution to the problem of slow recognition rate of current unmanned inspection equipment. In terms of improving the adaptability of the model, the present invention adds a C2PSA-DHSA module after the SPPF module. The C2PSA-DHSA module is obtained by introducing the DHSA module on the basis of the structure of the traditional C2PSA module, which can selectively focus on the spatial features of the dynamic range of substation lines and equipment. For long-distance images of substation equipment and lines, similar degraded pixels in the feature map can be processed in parallel, making the model perform excellently in common outdoor harsh weather image restoration tasks such as image de-snowing, de-rain and fog, and de-raining. And this structure helps to maintain the correct relative position and structural relationship at the global level, enabling the model to no longer rely on high-pixel input images, but can be more widely applied in low-performance recognition devices. In terms of improving the detection performance of small targets and occluded targets, after adding an additional small target detection layer, the model can separately identify and extract features of images with the median ratio of the bounding box area to the image area between 0.08% and 0.58%, solving the problems of missed detection and misjudgment of small targets in the detection process of substation equipment and lines; combined with the non-local attention mechanism module, the model can transfer the recognition attention from large-area occluded areas to ordinary areas through non-local attention, enabling the model to not only detect small targets, but also improve problems such as low recognition rate and misrecognition caused by occlusion, enhancing the adaptability of the model to various target sizes. Description of the Drawings

[0022] Figure 1 This is the flowchart of the substation foreign object intrusion detection method based on the improved YOLOv11 model in the embodiments of the present invention;

[0023] Figure 2 This is the structural diagram of the substation foreign object intrusion detection model built in the embodiments of the present invention;

[0024] Figure 3 This is the network structure diagram of the C2PSA-DHSA module in the embodiments of the present invention;

[0025] Figure 4 This is the network structure diagram of the PSAB-DHSA module in the embodiments of the present invention;

[0026] Figure 5 This is the network structure diagram of the DHSA module in the embodiments of the present invention;

[0027] Figure 6 This is the network structure diagram of the C3k2-Dual module in the embodiments of the present invention;

[0028] Figure 7 This is the network structure diagram of the Dual Conv convolution in the embodiments of the present invention;

[0029] Figure 8 This is the network structure diagram of the non-local attention mechanism module in the embodiments of the present invention. Detailed implementation manners

[0030] The present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners:

[0031] Embodiment 1

[0032] This Embodiment 1 describes a substation foreign object intrusion detection method based on an improved YOLOv11 model to solve problems such as missed detection, misjudgment, low recognition accuracy, and slow rate in the detection process of substation equipment and lines by traditional models.

[0033] Specifically, based on the YOLOv11 (You Only Look Once version 11) algorithm, the present invention proposes a substation foreign object intrusion detection model based on an improved YOLOv11 model to achieve automatic image recognition of foreign objects invading the substation. The substation foreign object intrusion detection model adds a C2PSA-DHSA module after the SPPF module. The C2PSA-DHSA module is obtained by inserting a DHSA module (Dynamic-range Histogram Self-Attention) into the C2PSA module (Cross Channel Pyramid Saliency Attention). At the same time, the C3k2-Dual module is used in cooperation, adding a new small target detection layer and adding a non-local attention mechanism module in front of the SPPF module and other improvements, enabling the model to more effectively identify those targets with smaller areas. While reducing the computational cost and the number of parameters of the model, it also improves the problems of low recognition rate and misrecognition caused by occlusion of the recognition device, improves the accuracy of the model, and enhances the adaptability of the model to targets of different sizes.

[0034] As Figure 1 shown, the substation foreign object intrusion detection method based on the improved YOLOv11 model includes the following steps:

[0035] Step 1. Obtain multiple images of substation equipment and lines, and perform image preprocessing and annotation on the obtained images of substation equipment and lines to obtain a training data set for the following substation foreign object intrusion detection model.

[0036] The training data set adopted by the model of the present invention is from a self-collected database, and the characteristics of this data set are as follows:

[0037] Large amount of data and diverse data. The database contains more than 2,500 images, from 12 sites respectively, with a time span from September 2020 to June 2024, including images of substation equipment and lines in multiple seasons and multiple weather conditions.

[0038] The data annotation of this data set is accurate. The images in the database have all been manually annotated with high annotation quality. The annotation categories are, for example: a is a bird's nest, b is a beehive, c is plastic and cloth foreign objects, d is a kite, e is a balloon, f is a branch.

[0039] Before using the substation foreign object intrusion detection model to detect foreign objects in the substation, first perform preprocessing operations on the obtained images of substation equipment and lines (i.e., the above data set) for model training.

[0040] The preprocessing operations mainly include processes such as reading the device and line image files of the substation into memory, scaling the images to the 640×640 size supported by the model, and performing normalization using the Z-score, etc.

[0041] Using Z-score normalization, the data is converted into a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0042] Z-score normalization converts the original data into a new dataset with zero mean and unit variance by calculating the deviation of each data point from the mean of the dataset and dividing it by the standard deviation of the dataset; the calculation formula is as follows:

[0043] ; Z is the normalized value, x is the original data point, which refers to a certain pixel value in the device and line image data of the substation or a certain line feature value in the feature map (which may be normal or have foreign objects), μ is the mean of all image pixel values or feature values, is the standard deviation of the original dataset.

[0044] Z-score normalization does not change the distribution shape of the data; if the original data is approximately normally distributed, then the normalized data will have a mean of 0 and a standard deviation of 1, but still maintain its original distribution shape, maintaining the data distribution and scale invariance; the data is centered at a mean of 0, which is conducive to the improved YOLOv11 model for recognition.

[0045] Then, perform color channel conversion to convert the image from RGB format to CHW format, that is, channel-height-width format.

[0046] Step 2. Build a substation foreign object intrusion detection model based on the YOLOv11 model architecture, and its network structure is as Figure 2 shown. This substation foreign object intrusion detection model includes a backbone network, a neck network, and a detection head network.

[0047] Among them, the backbone network includes a convolutional block, four groups of convolutional blocks and C3k2-Dual modules, a non-local attention mechanism module (NonLocalBlockND module), an SPPF module, and a C2PSA-DHSA module.

[0048] Define the four groups of convolutional blocks and C3k2-Dual modules as the first group, the second group, the third group, and the fourth group of convolutional blocks and C3k2-Dual modules respectively, and each group includes a convolutional block and a C3k2-Dual module.

[0049] The C3k2-Dual module is based on the traditional C3k2 module. It replaces the Conv convolution in the traditional C3k2 module with Dual Conv convolution, enabling the model to process the same input feature map channels simultaneously. By using the group convolution technique to efficiently arrange the convolution filters, it helps to improve the model's accuracy while reducing the computational cost and the number of parameters.

[0050] The C2PSA-DHSA module is a newly introduced module after the SPPF module. It is obtained by inserting the DHSA module into the C2PSA module. The C2PSA-DHSA module combines the convolution blocks of the DHSA block to enhance the feature extraction ability. The C2PSA-DHSA module enhances the feature extraction ability through the multi-head attention mechanism and the feed-forward neural network, ultimately achieving high-quality image restoration.

[0051] Through the parallel processing of BHR and FHR, the DHSA module enables the model to capture the global and local features of the power grid image simultaneously. Then, it fuses the global and local features through the self-attention mechanism, ultimately achieving high-quality image restoration.

[0052] The image features input into the YOLOv11 model architecture first enter the backbone network and are processed as follows:

[0053] The image features first undergo a convolution block for feature extraction, and then successively pass through the first group, second group, third group, and fourth group of convolution blocks and the C3k2-Dual module for feature extraction. Then, they pass through a non-local attention mechanism module, enabling the network to not only capture the local details of the equipment and line images of the input substation but also obtain global context information. The output features of the non-local attention mechanism module then pass through the SPPF module and finally are output by the C2PSA-DHSA module.

[0054] The neck network is used to achieve the deep fusion of features at different levels, and the detection head network contains four detection heads.

[0055] The neck network includes the C3k2-Dual module, convolution blocks, upsampling modules, and combination layers.

[0056] It is defined that there are six C3k2-Dual modules in the neck network, namely the first, second, third, fourth, fifth, and sixth C3k2-Dual modules; it is defined that there are two convolution blocks in the neck network, namely the first and second convolution blocks.

[0057] It is defined that there are six combination layers in the neck network, namely the first, second, third, fourth, fifth, and sixth combination layers, and it is defined that there are three upsampling modules in the neck network, namely the first, second, and third upsampling modules.

[0058] The output features of the C2PSA-DHSA module first pass through the first upsampling module and enter the first combination layer, where they are fused with the output features of the fourth convolutional block group and the C3k2-Dual module in the C3k2-Dual module.

[0059] The output features of the first combination layer sequentially pass through the first C3k2-Dual module and the second upsampling module.

[0060] The output features of the second upsampling module enter the second combination layer, where they are fused with the output features of the third convolutional block group and the C3k2-Dual module in the C3k2-Dual module.

[0061] The output features of the second combination layer sequentially pass through the second C3k2-Dual module and the third upsampling module.

[0062] The output features of the third upsampling module enter the third combination layer, where they are fused with the output features of the second convolutional block group and the C3k2-Dual module in the C3k2-Dual module.

[0063] The output features of the third combination layer sequentially pass through the third C3k2-Dual module and the first convolutional block; the output features of the first convolutional block enter the fourth combination layer, where they are fused with the output features of the second combination layer.

[0064] The output features of the fourth combination layer sequentially pass through the fourth C3k2-Dual module and the second convolutional block; the output features of the second convolutional block enter the fifth combination layer, where they are fused with the output features of the first combination layer.

[0065] The output features of the fifth combination layer pass through the fifth C3k2-Dual module and then enter the sixth combination layer, where they are fused with the output features of the C2PSA-DHSA module, and the output of the sixth combination layer enters the sixth C3k2-Dual module.

[0066] The output ends of the third, fourth, fifth, and sixth C3k2-Dual modules are respectively connected to a detection head.

[0067] The C2PSA-DHSA module inserts the DHSA module into the C2PSA module to replace the traditional attention. The DHSA module enables the detection model to simultaneously capture the global and local features of the power grid image through the parallel processing of BHR and FHR, and then fuses these features through the self-attention mechanism, ultimately achieving high-quality image restoration. The improved model can focus on the regions with similar degradation patterns caused by rain, fog, or snow, meeting the requirement of stable operation under extreme weather conditions.

[0068] As Figure 3 shown, the network structure of the C2PSA-DHSA module is presented. The C2PSA-DHSA module includes a data splitting layer, n PSAB-DHSA modules, and two convolutional layers. As Figure 4 shown, the PSAB-DHSA module consists of a DHSA module and a feedforward neural network composed of two convolutional layers. The image features input to the C2PSA-DHSA module first undergo feature extraction through a convolutional layer, and then pass through a data splitting layer to obtain two features, which are respectively input into one branch; one of the branch features is processed by n PSAB-DHSA modules, while the other branch feature remains unprocessed; the output features of the above two branches are concatenated, and the concatenated features are output after passing through a convolutional layer.

[0069] As Figure 5 shown, the processing flow of the DHSA module is as follows: Define the device and line image feature tensor F ∈ R C×H×W of a substation input to the DHSA module; where C is the number of channels, H is the height, and W is the width. The input feature map undergoes dynamic-range convolution, splitting the feature F into two branch features, namely feature F1 and F2. Feature F1 is sorted horizontally and vertically to obtain the sorted feature, and the sorted feature is concatenated with feature F2 to form a new feature F'. Feature F' is processed through 1×1 point convolution and 3×3 depth convolution to obtain the feature of dynamic-range convolution. Next, histogram reshaping operations, namely BHR and FHR operations, are performed on the feature of dynamic-range convolution, and the output of dynamic-range convolution is divided into value feature V and query-key pair F Q,K , that is:

[0070] ;

[0071] ;

[0072] ;

[0073] Among them, represents the query feature, K represents the key feature, V represents the value feature, represents the query feature matrix, represents the key feature matrix, represents the value feature matrix, and X represents the input image feature tensor F.

[0074] Sort V and rearrange F according to the sorting index Q,K , and the sorted feature V is in the form of C×B× .

[0075] Where B is the number of histogram bins, and the BHR operation makes each bin contain more device and line image pixels of the substation, which is used to capture the global features of the entire line; the FHR operation makes each bin contain fewer pixels, which is used to capture the local features of the line. Then, the self-attention maps A B and A F are calculated for the features of the BHR and FHR paths respectively, that is:

[0076] ;

[0077] .

[0078] Where RB and RF represent the reshaping operations of BHR and FHR respectively, and k is the scaling factor.

[0079] and represent Query vectors of different channels, and represent Key vectors of different channels. The subscript 1 is used to calculate self-attention in the global range, and the subscript 2 is used to calculate self-attention in the local range.

[0080] , , represent the reshaping operations on Q1, K1, and V in the BHR channel, making the number and frequency of elements contained in each histogram bin fixed, and reorganizing the Value features into a shape capable of self-attention calculation.

[0081] , , represent the reshaping operations on Q2, K2, and V in the FHR channel, making the number and frequency of elements contained in each histogram bin fixed, and reorganizing the Value features into a shape capable of self-attention calculation.

[0082] Finally, the two self-attention maps A B and A F are multiplied element-wise to obtain the output feature A E, that is A E= A B ⊙A F .

[0083] Where the horizontally and vertically sorted feature data outputs are element-projected after channel splitting and recombined with A E to obtain the final self-attention output feature map of the device and line A. Finally, the sorted features are re-projected back to the original spatial position through 1×1 point convolution to maintain spatial consistency.

[0084] As Figure 6 shown in the network structure diagram of the C3k2-Dual module.

[0085] Figure 6 In (a), C3k is False, and at this time C3k degrades to C2f for data processing; (b) is the case where C3k is True, and at this time C3k operates normally; (c) is the diagram of the stacking situation of C3k in the C3k2-Dual module.

[0086] Compared with the traditional C3k2 module, the C3k2-Dual module replaces the traditional Conv convolution with Dual Conv convolution, enabling the model to process the same input feature map channels simultaneously, and using group convolution technology to efficiently arrange convolution filters, while reducing the computational cost and the number of parameters and improving the accuracy of the model. As Figure 7 shown in the network structure diagram of Dual Conv convolution. Dual Conv convolution divides N convolutional kernels into G groups, and controls the ratio of convolutional kernels by adjusting the number of G groups to reduce the number of floating-point operations (FLOPs). Each group processes the device and line feature maps of all input substations.

[0087] Define that the line feature map of M / G input retains all the information of the input feature map through parallel 3×3 and 1×1 convolutional kernels to help the deep convolutional layer extract information more effectively. The remaining (M - M / G) input channels are only processed by 1×1 convolutional kernels to reduce the number of parameters of the model; finally, the results of the 3×3 and 1×1 convolutional kernels are summed to obtain the output feature map.

[0088] In addition, the present invention also adds a non-local attention mechanism module, namely the NonLocalBlockND module, before the SPPF module. Compared with CNN, the NonLocalBlockND module is no longer limited to the features of the local neighborhood. By adding the NonLocalBlockND module, the network can not only capture the local details of the device and line images of the input substation, but also obtain global context information, so that the recognition ability of small objects (small birds, small flying objects, bird nests) and large occlusions (kites, large plastic products and other large flying objects, trees, lens occlusions) in line and device detection of the model will be significantly enhanced.

[0089] The non-local attention mechanism module captures the mutual relationship of all positions in the feature map globally to improve the model performance, that is: .

[0090] Where is to obtain the response of the output position i of the feature map, is the input signal, is a pairwise function for calculating the relationship between positions i and j, is a univariate function for calculating the feature representation of position j, is the normalization factor; represents the input signal at position i, represents the input signal at position j.

[0091] As Figure 8 shown, the non-local attention mechanism module generates a feature representation with reduced dimensions through three branches θ, , and g, which is used to calculate the similarity between the transformer line and equipment features and perform weighting.

[0092] Among them, the θ branch is used to reduce the input features to 512 channels for generating the query vector query; The branch is also reduced to 512 channels for generating the key vector key, which is used to calculate the similarity with the query vector; the g branch is reduced to 512 channels to generate the value vector value of the input features for weighted aggregation, that is:

[0093] ;

[0094] ;

[0095] .

[0096] Then, the attention weight matrix is generated through the similarity calculation of θ and , and it is multiplied by the feature of the g branch to obtain the globally weighted feature representation , and the formula is as follows:

[0097] ;

[0098] ;

[0099] Among them, , q is the query vector, k is the key vector, is the dimension of the key vector.

[0100] Finally, the globally weighted feature representation restores the number of channels through 1×1 convolution, and adopts a residual structure to retain the original features and the globally weighted features, combining local and global information to improve the perception ability of the model.

[0101] In addition, the present invention also enhances the detection ability of the model for small targets (the median of the ratio of the bounding box area to the image area is between 0.08% and 0.58%) by adding an additional detection layer. The network with the new detection layer can identify small targets separately during the recognition process, and to a certain extent, improves the multi-scale detection ability and generalization ability of the model.

[0102] Step 3. Train the built substation foreign object intrusion detection model based on the training dataset in Step 1 to obtain a trained substation foreign object intrusion detection model, and use the trained model to detect foreign objects in the transformer.

[0103] When training the built substation foreign object intrusion detection model, set the learning rate to 0.08; the model optimization algorithm uses the AdamW optimization algorithm; the batch size is 32, and the number of iterations of the model is 400.

[0104] The learning rate is an important hyperparameter in deep learning, which controls the update speed of the model during training. Given that this model is a small model applied to image classification and recognition, this model uses a training method with a warm-up learning rate. By constructing a learning rate scheduler, the learning rate adjustment rate for each iteration is set to 0.001, and the cycle period is specified as 80 iterations. After 80 iterations, the base learning rate is used. Through multiple experiments, it is found that when the learning rate is higher than 0.1, the model often has the problem of overfitting; when the learning rate is lower than 0.001, the model training takes a long time. Therefore, the learning rate of this model is set to 0.08.

[0105] The AdamW optimization algorithm is a variant of the Adam optimizer, which can directly introduce weight decay into the optimization step to improve the generalization performance of the model. And AdamW combines the advantages of adaptive learning rate and momentum. Different from the original Adam optimizer, it can directly apply weight decay to the parameter update step, achieving better regularization effect and generalization performance.

[0106] The AdamW optimizer, after initialization, defines the initial values of the first-order matrix estimate m t and the second-order matrix estimate v t as m0 = 0 and v0 = 0 respectively. By continuously updating the first-order matrix estimate m t and the second-order matrix estimate v t to calculate the bias-corrected first-order matrix estimate m t and the second-order matrix estimate v t , and finally update the optimized model parameters θ; its calculation formula is as follows:

[0107] .

[0108] Where Refers to the calculation of the gradient in the t-th iteration, and Refers to the first-order moment estimate and the second-order moment estimate in the t-th iteration, Refers to the decay rate of the first-order moment estimate, with a value of 0.9, Refers to the decay rate of the second-order moment estimate, with a value of 0.999, is a constant used to prevent the denominator from being zero; is the weight decay coefficient used for regularization to prevent overfitting; Refers to the learning rate that controls the size of each update step; and respectively represent the model parameters θ in the t-th and (t - 1)-th iterations.

[0109] After multiple debuggings, the number of iterations of the model is set to 400. At this time, the model performance is better and overfitting does not occur.

[0110] After the model training is completed, for the input images of the equipment and lines of the substation to be detected, first perform image preprocessing operations according to step 1 in the above method, and then input the preprocessed image features into the substation foreign object intrusion detection model for prediction. The substation foreign object intrusion detection model finally outputs the category, location, and confidence score of the target.

[0111] Through the above method steps, the method of the present invention can more effectively identify those targets with smaller areas. While reducing the computational cost and the number of parameters of the model, it also improves the problems of low recognition rate and misrecognition caused by occlusion of the recognition device, improves the accuracy of the model, enhances the adaptability of the model to targets of different sizes, and significantly improves the real-time performance.

[0112] Embodiment 2

[0113] This Embodiment 2 describes a substation foreign object intrusion detection system based on an improved YOLOv11 model. The substation foreign object intrusion detection system includes a camera and a computer device. The camera is respectively used to collect images of the equipment and lines of the substation, and upload the image pair composed of the collected images of the equipment and lines of the substation to the computer device.

[0114] The computer device includes a memory and at least one processor. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement the substation foreign object intrusion detection method based on the improved YOLOv11 model in the above Embodiment 1.

[0115] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to the enumerated embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.

Claims

1. A substation foreign object intrusion detection method based on an improved YOLOv11 model, characterized in that, It includes the following steps: Step 1. Obtain multiple images of the equipment and lines in the substation, and perform image preprocessing and annotation on the obtained images of the equipment and lines in the substation to obtain a training dataset for the following substation foreign object intrusion detection model; Step 2. Build a substation foreign object intrusion detection model based on the YOLOv11 model architecture. This substation foreign object intrusion detection model includes a backbone network, a neck network, and a detection head network; Among them, the backbone network includes a convolutional block, four groups of convolutional blocks and a C3k2-Dual module, a non-local attention mechanism module, an SPPF module, and a C2PSA-DHSA module; Define the four groups of convolutional blocks and the C3k2-Dual module as the first group, the second group, the third group, and the fourth group of convolutional blocks and the C3k2-Dual module respectively, and each group includes a convolutional block and a C3k2-Dual module; The C3k2-Dual module is based on the traditional C3k2 module. It replaces the Conv convolution in the traditional C3k2 module with Dual Conv convolution, enabling the model to process the same input feature map channels simultaneously and efficiently arranging the convolutional filters using group convolution technology; the Dual Conv convolution divides N convolutional kernels into G groups and controls the ratio of convolutional kernels by adjusting the number of G groups to reduce the number of floating-point operations. Each group processes all the input feature maps of the equipment and lines in the substation; Define that the line feature map channels of M / G inputs retain all the information of the input feature map through parallel 3×3 and 1×1 convolutional kernels to help the deep convolutional layer extract information more effectively, and the remaining (M - M / G) input channels are only processed by 1×1 convolutional kernels to reduce the number of parameters of the model; finally, the results of the 3×3 and 1×1 convolutional kernels are summed to obtain the output feature map; The C2PSA-DHSA module is obtained by inserting the DHSA module into the traditional C2PSA module; The DHSA module enables the model to capture the global and local features of the power grid image simultaneously through the parallel processing of BHR and FHR, and then fuses the global and local features through the self-attention mechanism to finally achieve high-quality restoration of the image; The image features input into the YOLOv11 model architecture first enter the backbone network and are processed as follows: The image features first undergo feature extraction through a convolutional block, and then sequentially undergo feature extraction through the first group, the second group, the third group, and the fourth group of convolutional blocks and the C3k2-Dual module, and then pass through a non-local attention mechanism module, enabling the network to not only capture the local details of the input images of the equipment and lines in the substation but also obtain global context information; the output features of the non-local attention mechanism module pass through the SPPF module and finally are output through the C2PSA-DHSA module; The neck network is used to achieve deep fusion of features at different levels, and the detection head network contains four detection heads; Step 3. Train the built substation foreign object intrusion detection model based on the training dataset in Step 1 to obtain a trained substation foreign object intrusion detection model, and use the trained model to detect foreign object intrusion in the transformer.

2. The method for detecting foreign object intrusion in a substation based on the improved YOLOv11 model according to claim 1, characterized in that, The neck network includes a C3k2-Dual module, a convolutional block, an upsampling module, and a combination layer; It is defined that there are six C3k2-Dual modules in the neck network, which are the first, second, third, fourth, fifth, and sixth C3k2-Dual modules respectively; it is defined that there are two convolutional blocks in the neck network, which are the first and second convolutional blocks respectively; It is defined that there are six combination layers in the neck network, which are the first, second, third, fourth, fifth, and sixth combination layers respectively, and it is defined that there are three upsampling modules in the neck network, which are the first, second, and third upsampling modules respectively; The output features of the C2PSA-DHSA module first pass through the first upsampling module and enter the first combination layer, where they are fused with the output features of the fourth convolutional block and the C3k2-Dual module in the C3k2-Dual module; The output features of the first combination layer sequentially pass through the first C3k2-Dual module and the second upsampling module; The output features of the second upsampling module enter the second combination layer, where they are fused with the output features of the third convolutional block and the C3k2-Dual module in the C3k2-Dual module; The output features of the second combination layer sequentially pass through the second C3k2-Dual module and the third upsampling module; The output features of the third upsampling module enter the third combination layer, where they are fused with the output features of the second convolutional block and the C3k2-Dual module in the C3k2-Dual module; The output features of the third combination layer sequentially pass through the third C3k2-Dual module and the first convolutional block; the output features of the first convolutional block enter the fourth combination layer, where they are fused with the output features of the second combination layer; The output features of the fourth combination layer sequentially pass through the fourth C3k2-Dual module and the second convolutional block; the output features of the second convolutional block enter the fifth combination layer, where they are fused with the output features of the first combination layer; The output features of the fifth combination layer pass through the fifth C3k2-Dual module and then enter the sixth combination layer, where they are fused with the output features of the C2PSA-DHSA module, and the output of the sixth combination layer enters the sixth C3k2-Dual module; Among them, the output ends of the third, fourth, fifth, and sixth C3k2-Dual modules are respectively connected to one of the detection heads.

3. The substation foreign object intrusion detection method based on the improved YOLOv11 model according to claim 1, wherein, The C2PSA-DHSA module includes a data segmentation layer, n PSAB-DHSA modules, and two convolutional layers; Among them, the processing flow of the C2PSA-DHSA module is as follows: The image features input into the C2PSA-DHSA module first go through a convolutional layer for feature extraction, and then go through a data segmentation layer to obtain two features and are respectively input into a branch; One of the branch features is processed through n PSAB-DHSA modules, and the other branch feature is not processed; The output features of the above two branches are concatenated, and the concatenated features are output after passing through a convolutional layer; The PSAB-DHSA module includes a DHSA module and a feedforward neural network composed of two convolutional layers.

4. The substation foreign object intrusion detection method based on the improved YOLOv11 model according to claim 3, characterized in that The processing flow of the DHSA module is as follows: Define the device and line image feature tensor F ∈ R of a substation input into the DHSA module C×H×W ; Where C is the number of channels, H is the height, and W is the width; the input feature map undergoes dynamic range convolution, which divides F into two branch features, namely feature F1 and F2; the feature F1 is sorted horizontally and vertically to obtain the sorted feature, and the sorted feature is concatenated with the feature F2 to form a new feature F'; the feature F' is processed by 1×1 point convolution and 3×3 depth convolution to obtain the feature of dynamic range convolution; then, a histogram reshaping operation is performed on the feature of dynamic range convolution, namely the BHR and FHR operations, which divide the output of dynamic range convolution into value feature V and query-key pair F Q,K , that is: Q = X·W Q ; K = X · W K ; V = X·W V ; Among them, Q represents the query feature, K represents the key feature, V represents the value feature, and W Q represents the query feature matrix, and W K represents the key feature matrix, and W V represents the value feature matrix, and X represents the input image feature tensor F; Sort V and rearrange F according to the sorting index Q,K , where the sorted feature V is in the form of; Where B is the number of histogram bins, the BHR operation is used to capture the global features of the entire line; The FHR operation is used to capture the local features of the line; Then, calculate the self-attention maps A B and A F for the features of the BHR and FHR paths respectively, i.e.: B and A F , i.e.: Where RB and RF respectively represent the reshaping operations of BHR and FHR, and k is the scaling factor; Q1 and Q2 represent Query vectors of different channels, and K1 and K2 represent Key vectors of different channels; The subscript 1 indicates that it is used to calculate self-attention within the global range, and the subscript 2 indicates that it is used to calculate self-attention within the local range; R B (Q1), R B (K1), R F (V) represents reshaping operations on Q1, K1, and V in the BHR channel to fix the number and frequency of elements in each histogram bin and reorganize the Value feature into a shape capable of performing self-attention calculations; R F (Q2), R F (K2), R F (V) represents reshaping operations on Q2, K2, and V in the FHR channel to make the number and frequency of elements in each histogram bin fixed, and reorganizing the Value feature into a shape capable of performing self-attention calculation; Finally, multiply the two self-attention maps A B and A F element-wise to obtain the output feature A E, That is, A E = A B ⊙ A F ; Among them, the feature data outputs after horizontal sorting and vertical sorting are subjected to element projection after channel splitting, and are recombined with the output feature A E to obtain the final feature map A of the self-attention output device and the circuit; finally, the sorted features are re-projected back to the original spatial position through 1×1 point convolution to maintain spatial consistency.

5. The substation foreign object intrusion detection method based on the improved YOLOv11 model according to claim 1, characterized in that, The non-local attention mechanism module captures the mutual relationships of all positions in the feature map within the global range, that is: where y i is the response for obtaining the output position i of the feature map, x is the input signal, f() is a pairwise function for calculating the relationship between positions i and j, g() is a univariate function for calculating the feature representation of position j, and C(x) is a normalization factor; x i represents the input signal at position i, and x j represents the input signal at position j; The non-local attention mechanism module generates a reduced-dimensional feature representation through three branches θ, φ, and g, which is used to calculate the similarity between the transformer line and equipment features and perform weighting; Among them, the θ branch is used to reduce the input features to 512 channels, which is used to generate the query vector query; The φ branch is also reduced to 512 channels, which is used to generate the key vector key, which is used to calculate the similarity with the query vector; The g branch is reduced to 512 channels, generating the value vector value of the input features, which is used for weighted aggregation; Then, the attention weight matrix attentionweights is generated through the similarity calculation of θ and φ, and it is multiplied by the feature v of the g branch to obtain the globally weighted feature representation y. The formula is expressed as follows: y = attentionweights·v; where similarity = q T k, q are query vectors, k is the key vector, d k is the dimension of the key vector; Finally, the globally weighted feature representation y is restored to the number of channels through a 1×1 convolution, and the residual structure is used to retain the original features and the globally weighted features, combining local and global information to improve the perception ability of the model.

6. The substation foreign object intrusion detection method based on the improved YOLOv11 model according to claim 1, characterized in that, In step 1, before using the substation foreign object intrusion detection model to detect foreign objects in the substation, first obtain the device and line images of the substation for model training, and perform preprocessing operations, including reading the device and line image files of the substation into memory, scaling the images to the 640×640 size supported by the model, and using Z-score normalization.

7. The method for detecting foreign object intrusion in a substation based on the improved YOLOv11 model according to claim 6, characterized in that, In step 1, Z-score normalization converts the original data into a new data set with zero mean and unit variance by calculating the deviation of each data point from the mean of the data set and dividing it by the standard deviation of the data set; The calculation formula is as follows: Among them, Z is the normalized value; x is the original data point, that is, a certain pixel value in the equipment and line image data of the substation or a certain line feature value in the feature map, μ is the mean of all image pixel values or feature values, and σ is the standard deviation of the original data set.

8. The substation foreign object intrusion detection method based on the improved YOLOv11 model according to claim 1, wherein In step 3, when training the built substation foreign object intrusion detection model, set the learning rate to 0.08; the model optimization algorithm uses the AdamW optimization algorithm; the batch size is 32, and the number of iterations of the model is 400; In step 3, after initialization, the AdamW optimizer defines the first-order matrix estimate m t and the second-order matrix estimate v t with initial values of m0 = 0 and v0 = 0 respectively. By continuously updating the first-order matrix estimate m t and the second-order matrix estimate v t to calculate the bias-corrected first-order matrix estimate m t and the second-order matrix estimate v t , and finally update the optimized model parameters θ. The calculation formula is as follows: where g t denotes the computed gradient in the t-th iteration, m t , v t denote the first and second moment estimates in the t-th iteration, β1 denotes the decay rate of the first moment estimate with a value of 0.9, β2 denotes the decay rate of the second moment estimate with a value of 0.999, ε is a constant used to prevent the denominator from being zero; λ is the weight decay coefficient for regularization to prevent overfitting; η denotes the learning rate to control the size of each update step; θ t , θ t-1 represent the model parameters θ in the t-th and (t - 1)-th iterations respectively.

9. The substation foreign object intrusion detection system based on the improved YOLOv11 model includes a camera and computer equipment, where, The computer device includes a memory and at least one processor; characterized in that The memory stores executable code; when the processor executes the executable code, it is used to implement the substation foreign object intrusion detection method based on the improved YOLOv11 model as described in any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • Insect behavior identification method based on YOLO V7 and 2D convolutional network

    CN117275084A

  • Transformer fault detection method based on improved YOLOv8 model

    CN119323565A