Substation insulator segmentation method and system based on BOX supervision

Through a BOX supervision-based method, combined with the attention mechanism of SAM and GrabCut models and the U-Net network, the complex background and diversity problems of insulator identification in the substation are solved, and high-precision insulator segmentation is achieved, ensuring the safe and stable operation of the equipment.

CN120259647APending Publication Date: 2025-07-04STATE GRID SHANDONG ELECTRIC POWER CO LIAOCHENG POWER SUPPLY CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237037.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-01
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

When identifying insulators in substations, the prior art faces complex background environment, insulator morphology diversity and algorithm limitations, resulting in low recognition accuracy and affecting the safe and stable operation of the equipment.

Method used

Using a BOX supervision-based method, the SAM segmentation model and the GrabCut segmentation model are used to obtain the segmentation mask, and the U-Net network combined with channel attention and spatial attention mechanism is trained to achieve accurate segmentation of insulators.

Benefits of technology

It improves the accuracy and efficiency of insulator identification, and ensures the safety and stability of substation equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259647A_ABST
    Figure CN120259647A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of power image processing. According to the transformer substation insulator segmentation method and system based on BOX supervision, target BOX labels of training images are used for obtaining corresponding segmentation masks through an SAM segmentation algorithm and a GrabCut segmentation algorithm, then mask refinement is carried out through the images and operation, and final segmentation mask labels are obtained. Based on the training image data and the corresponding segmentation mask label data, U-Net network training based on a channel attention mechanism and a space attention mechanism is carried out so as to realize accurate segmentation of the transformer substation insulator, the feasibility and the accuracy are high, the insulator part in the transformer substation can be well identified, and the identification efficiency of the transformer substation insulator is improved. And the equipment safety in the transformer substation is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power image processing, and particularly to a substation insulator segmentation method and system based on BOX supervision. Background Technique

[0002] The statements in this part only provide background techniques related to the present invention and do not necessarily constitute prior art.

[0003] Insulators are essential key devices in substations. Their main functions are to ensure the insulation performance and mechanical support ability of electrical equipment, which is particularly important in high-voltage power systems. It is not only a key component for electrical isolation but also an important support for ensuring the safe and stable operation of equipment. Its performance and reliability are directly related to the overall operation quality of the power system and are an important link that cannot be ignored in the construction and operation of substations. Identifying insulators in substation scenarios is a highly challenging task, and the accuracy of identification is crucial for intelligent analysis and processing in substations.

[0004] Although there are currently various insulator recognition algorithms developed and applied in practical scenarios, there are still significant problems in the recognition accuracy of these algorithms. The main reasons include the following aspects: (1) Complex and variable background environment. Insulators are usually installed outdoors, and their background environment is complex and variable, including various elements such as trees, buildings, and the sky. The interference of these background elements in the image may cause the algorithm to be difficult to accurately extract the features of insulators, thus affecting the recognition accuracy; (2) Diversity of insulator morphology and types. Insulators have diverse morphologies and types, including different sizes, shapes, and colors, etc. This diversity increases the difficulty of the algorithm in feature extraction and classification, making it difficult for the algorithm to accurately identify all types of insulators; (3) Limitations of the algorithm itself. Existing insulator recognition algorithms may have limitations in feature extraction, classification, and recognition, etc. For example, some algorithms may rely too much on specific features or conditions, resulting in poor performance in complex or changing environments. In addition, the computational efficiency and robustness of the algorithm may also affect its practicality and recognition accuracy. Summary of the Invention

[0005] To address the deficiencies of the prior art, the present invention provides a method and system for substation insulator segmentation based on BOX supervision. By using the target BOX annotation of training images, the SAM segmentation model and the GrabCut segmentation model are respectively used to obtain their corresponding segmentation masks. Then, image operations are used to refine the masks to obtain the final segmentation mask labels. Based on the training images and the corresponding segmentation mask label data, a U-Net network based on channel attention mechanism and spatial attention mechanism is trained to achieve accurate segmentation of substation insulators, which has high feasibility and accuracy, can better identify the insulator parts in the substation, and ensure the safety of the equipment in the substation.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a method for substation insulator segmentation based on BOX supervision.

[0007] A method for substation insulator segmentation based on BOX supervision includes the following processes: Obtain the substation image to be segmented; According to the substation image and the pre-trained U-Net network based on spatial attention mechanism and channel attention mechanism, obtain the segmentation result of the insulator in the substation image; Among them, the training of the U-Net network includes: annotating the insulator part in the training set images with BOX, inputting each BOX-annotated training set image into the SAM segmentation model and the GrabCut segmentation model respectively, taking the intersection of the two segmentation results to obtain the training set label corresponding to each training set image, and training the U-Net network according to the training set images and the training set labels.

[0008] As a further limitation of the first aspect of the present invention, taking the intersection of the two segmentation results includes: for any training set image, obtaining the first segmentation result output by the SAM segmentation model and the second segmentation result output by the GrabCut segmentation model, and performing a pixel-by-pixel AND operation on the first segmentation result and the second segmentation result to obtain the training set label corresponding to this training set image.

[0009] As a further limitation of the first aspect of the present invention, the U-Net network based on spatial attention mechanism and channel attention mechanism includes: an encoder, a decoder, channel attention processing, and spatial attention processing; The encoder includes multiple stages of convolution operations. After each stage of convolution operation, max pooling operation is used to downsample the features. The decoder includes multiple stages of convolution operations. After each stage of convolution operation, the features are upsampled. The number of stages of convolution operations in the encoder is the same as the number of stages of convolution operations in the decoder; After the convolutional operation of each stage of the encoder is completed, the features are sequentially subjected to channel attention processing and spatial attention processing. The result of the spatial attention processing is output to the convolutional operation stage of the corresponding decoder for splicing operation.

[0010] As a further limitation of the first aspect of the present invention, the channel attention processing includes: Performing global max pooling and global average pooling on a spatial dimension of an input insulator feature map F with a size of H×W×C to obtain two insulator feature maps of 1×1×C; Feeding the two insulator feature maps of 1×1×C into a shared multi-layer perceptron for learning to obtain two processed insulator feature maps of 1×1×C; Performing an addition operation on the results output from the learning in the multi-layer perceptron, and then performing mapping processing through a Sigmoid activation function to finally obtain a channel attention weight matrix, and obtaining a feature map after channel attention processing according to the channel attention weight matrix.

[0011] As a further limitation of the first aspect of the present invention, the spatial attention processing includes: Performing global max pooling and global average pooling on a channel dimension of an input feature map F with a size of H×W×C to obtain two insulator feature maps of H×W×1; Splicing the two insulator feature maps of H×W×1 according to the channels to obtain a feature map with a size of H×W×2. Finally, performing a 7×7 convolutional operation on the spliced result to obtain a feature map with a size of H×W×1, and then passing through a Sigmoid activation function to obtain a spatial attention weight matrix, and obtaining a feature map after spatial attention processing according to the spatial attention weight matrix.

[0012] In a second aspect, the present invention provides a substation insulator segmentation system based on BOX supervision.

[0013] A substation insulator segmentation system based on BOX supervision includes: An image acquisition unit configured to acquire a substation image to be segmented; An insulator segmentation unit configured to obtain an insulator segmentation result in the substation image according to the substation image and a pre-trained U-Net network based on a spatial attention mechanism and a channel attention mechanism; Among them, the training of the U-Net network includes: annotating the insulator part in the training set images with BOX, inputting each training set image after BOX annotation into the SAM segmentation model and the GrabCut segmentation model respectively, taking the intersection of the two segmentation results to obtain the training set labels corresponding to each training set image, and training the U-Net network according to the training set images and the training set labels.

[0014] As a further limitation of the second aspect of the present invention, in the insulator segmentation unit, the U-Net network based on the spatial attention mechanism and the channel attention mechanism includes: an encoder, a decoder, channel attention processing, and spatial attention processing; The encoder includes convolutional operations in multiple stages. After each stage of convolutional operation, max pooling operation is used to downsample the features. The decoder includes convolutional operations in multiple stages. After each stage of convolutional operation, the features are upsampled. The number of stages of convolutional operations in the encoder is the same as the number of stages of convolutional operations in the decoder; After each stage of convolutional operation in the encoder, channel attention processing and spatial attention processing are sequentially performed on the features, and the result of the spatial attention processing is output to the corresponding convolutional operation stage of the decoder for splicing operation.

[0015] In a third aspect, the present invention provides a computer device, which is characterized in that it includes: a processor and a computer-readable storage medium; The processor is adapted to execute a computer program; The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the BOX-supervised substation insulator segmentation method as described in the first aspect of the present invention.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the BOX-supervised substation insulator segmentation method as described in the first aspect of the present invention.

[0017] In a fifth aspect, the present invention provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the BOX-supervised substation insulator segmentation method as described in the first aspect of the present invention.

[0018] Compared with the prior art, the beneficial effects of the present invention are: The present invention innovatively proposes a method for segmenting substation insulators based on BOX supervision. The target BOX annotations of the training images are used to obtain the corresponding segmentation masks by using the SAM segmentation algorithm and the GrabCut segmentation algorithm respectively. Then, image operations are used to refine the masks to obtain the final segmentation mask labels. Based on the training image data and the corresponding segmentation mask label data, a U-Net network based on channel attention mechanism and spatial attention mechanism is trained to achieve accurate segmentation of substation insulators, which has high feasibility and accuracy, can better identify the insulator parts in the substation, and ensure the safety of the equipment in the substation.

[0019] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings, which form a part of this specification, are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0021] Figure 1 Schematic diagram of the U-Net network training method based on spatial attention mechanism and channel attention mechanism provided in Embodiment 1 of the present invention; Figure 2 Schematic diagram of the promptable segmentation provided in Embodiment 1 of the present invention; Figure 3 Schematic diagram of the U-Net network based on channel attention mechanism and spatial attention mechanism provided in Embodiment 1 of the present invention; Figure 4 Schematic diagram of the encoder part structure provided in Embodiment 1 of the present invention; Figure 5 Schematic diagram of the decoder part structure provided in Embodiment 1 of the present invention; Figure 6 Schematic diagram of the bottleneck part structure provided in Embodiment 1 of the present invention; Figure 7 Flowchart of the channel attention mechanism provided in Embodiment 1 of the present invention; Figure 8 Flowchart of the spatial attention mechanism provided in Embodiment 1 of the present invention; Figure 9 Flowchart of the hybrid attention mechanism provided in Embodiment 1 of the present invention; Figure 10 Schematic diagram of a system for segmenting substation insulators based on BOX supervision provided in Embodiment 2 of the present invention; Figure 11A schematic diagram of the computer device provided in Embodiment 3 of the present invention. Detailed implementation manners

[0022] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0023] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0024] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0025] Embodiment 1: This implementation manner proposes a method for segmenting substation insulators based on BOX supervision, including the following processes: S1: Obtain the substation image to be segmented; S2: According to the substation image and the pre-trained U-Net network based on the spatial attention mechanism and the channel attention mechanism, obtain the segmentation result of the insulators in the substation image.

[0026] The specific training process is as Figure 1 shown. First, it is the training set data acquisition part. The training set images are the photos with insulators taken by the camera in the substation. After that, the insulator parts in each image are marked with BOX. The method for obtaining the training set labels is to use the SAM image segmentation model and the GrabCut image segmentation model to process each image respectively, and segment the insulators in it. Then, take the intersection result of the two output results as a training set image. By performing the above operations on a large number of images, the training set labels can be obtained. Secondly, use the U-Net based on the spatial attention mechanism and the channel attention mechanism to segment the training set images, segment the insulator parts, and then compare them pixel by pixel with the training set labels. After calculating the loss function, adjust the network parameters to reduce the loss function and improve the segmentation accuracy of the model. Finally, the target task is achieved. This method can achieve good results for insulators of different types and different shape features, and has the advantages of fast detection speed and high robustness.

[0027] Training the model for the task of segmenting insulators requires training set images and training set labels. The method for obtaining the training labels in the present invention is to use the BOX annotation of the area with insulators in the image as a hint, and use the SAM image segmentation model and the GrabCut image segmentation model to mark the insulators in the image in the form of a mask as the training labels.

[0028] The SAM image segmentation model can segment objects in images, including objects and visual fields that have not been seen before. This means that in the present invention, the SAM image segmentation model can be used for various insulator image segmentation tasks, improving the flexibility and adaptability of the work. SAM can perform various segmentation tasks by framing a box including the insulator on the image, and its output result is to return a valid segmentation mask according to the segmentation prompt (see Figure 2 ), and the SAM image segmentation model can be used to label the insulators in the image to obtain segmentation labels.

[0029] Since SAM is an interactive large model, users need to give prompts for SAM to output masks at the corresponding positions. Insulators are everywhere in substations and on transmission lines, so there will be many insulator images obtained by camera photography. If workers use the SAM image segmentation model to segment the insulators in each image, then corresponding prompts need to be made for each image, and workers may miss labeling insulators due to poor shooting angles or inattentiveness during labeling. This work is very cumbersome. Therefore, it is not appropriate to simply use the SAM image segmentation model for segmentation tasks in this scenario. However, considering the excellent segmentation effect of SAM, it is used in the process of obtaining segmentation labels.

[0030] Although the SAM image segmentation model has obvious advantages in insulator image segmentation tasks, SAM also has certain disadvantages. For example, for insulators made of glass, due to their transparent state, the insulators are easily affected by the background image, resulting in poor segmentation effects and unable to obtain correct segmentation labels. Therefore, the present invention proposes to use the GrabCut algorithm to process images with insulators. GrabCut is an efficient image segmentation algorithm. It is based on the graph cut technology and segments the image through the rough information of the foreground (object) and background provided by the user. The user only needs to draw a rectangular box around the insulator in the image to specify the approximate position of the insulator, and the algorithm can automatically separate the insulator from the background of the image. This method is different from algorithms such as KMeans and MeanShift because it not only considers the color information of pixels but also considers the relationship between pixels, making the result of segmenting insulators more accurate. The core idea of GrabCut is to transform the image segmentation problem into a graph cut problem and separate the insulator and the background by minimizing an energy function. The following is the principle of the GrabCut algorithm: The energy function consists of two parts: (1); Data term Represents the cost of classifying a pixel as foreground or background, based on the pixel's color distribution (represented by GMM), and the smoothness term Indicates whether adjacent pixels should have the same label, used to preserve the smoothness of the segmentation. L represents the label (foreground or background). Represents the GMM parameters, and Z is the pixel color value.

[0031] Graph model construction: The graph model includes the nodes and edges of the graph. The nodes of the graph represent the pixels of the image, and the edges of the graph include the edges between pixels and the foreground / background source nodes (data term weights) and the edges between pixels (smoothness term weights).

[0032] Gaussian Mixture Model (GMM): GrabCut assumes that the color distributions of the foreground and background respectively follow independent GMMs. The parameters of the GMM (means, covariances, and mixture weights) are learned through the Expectation-Maximization (EM) algorithm and updated according to the initialization information provided by the user.

[0033] Graph Cut: Use the Min-Cut algorithm to cut the graph and find an optimal solution to segment the foreground and background, minimizing the energy function.

[0034] The advantage of the GrabCut algorithm lies in its high accuracy. It is an algorithm based on iterative optimization, so it can continuously iterate to ensure the accuracy of segmenting insulator images. In addition, the GrabCut algorithm can handle complex images, including segmentation in the case of complex backgrounds. In substations and on transmission lines, the backgrounds in the images of insulators captured by cameras are necessarily diverse. Therefore, it is reasonable to use the GrabCut algorithm to segment insulator images with BOXes. However, there are still inconveniences in simply using the GrabCut algorithm to complete the task: GrabCut has high requirements for hardware. The GrabCut algorithm requires multiple iterations, and each iteration needs to calculate parameters such as Gaussian mixture model parameters and joint probability distributions. Therefore, the computational complexity of the algorithm is large, and a powerful computer is required to perform real-time segmentation. In addition, like the SAM image segmentation model, GrabCut is interactive. Users need to manually mark the foreground and background to guide the algorithm to perform segmentation in order to improve the accuracy of the algorithm. This cumbersome and redundant work is obviously inefficient and meaningless. However, considering its high segmentation accuracy, the present invention uses it in the process of obtaining segmentation labels.

[0035] In terms of the performance of SAM and GrabCut, they are both very good at segmenting the parts to be segmented under the prompt of the BOX. However, they both have two common drawbacks. One is the large amount of computation and the relatively high requirements for hardware. More importantly, they both complete the segmentation task interactively. In this task, the insulator in the image needs to be masked without any prompt. Obviously, neither of these two methods can achieve this. Therefore, another model needs to be built in the present invention. After inputting the feature map with the insulator, the neural network can output the feature map of the insulator masked in the image. The training of the model requires training set images and training set labels. The training image is an original image without any processing, and the training set label is the image with the insulator masked. The training set label is the result of performing a per-pixel AND operation on each image processed by the SAM image segmentation model and GrabCut respectively. The mathematical expression is as follows: (2); Among them, Result is the result image, Image1 is the image processed by the SAM image segmentation model, and Image2 is the image processed by GrabCut.

[0036] The present invention requires inputting the image with the insulator into the model, and the model represents the insulator part in the image in the form of a mask. The entire process does not require any manual annotation of the image. To meet the above requirements and ensure the accuracy of the detection results, the present invention proposes a U-Net model based on the channel attention mechanism and the spatial attention mechanism. After each convolutional layer stage in the U-Net encoder ends and the feature map is processed pixel by pixel using the ReLU function, a copy enters the CBAM for processing, and finally, after central cropping, it enters the decoder part for processing. The improved partial structure diagram is as Figure 3 shown. Taking the insulator image with BOX annotation as the input and adding it to the SAM image segmentation model and the GrabCut algorithm respectively, they can both represent the insulator in the BOX in the form of a mask. Take the intersection of the two as the training set label, and the original image as the training set image. Use this training set to train the attention mechanism-based U-Net network proposed in the present invention, and test the trained model with the test set images, obtaining good results.

[0037] U-Net is a convolutional neural network. It has performed excellently in many image segmentation tasks due to its unique architecture and performance, and has become a classic network model in the segmentation field. It has the advantages of full-resolution prediction ability, efficient parameter utilization, accurate boundary segmentation, simple implementation, and stable training. This exactly meets the core requirements for the task of segmenting insulator images. Therefore, the present invention selects U-Net for insulator segmentation.

[0038] The U-Net model consists of an encoder and a decoder. The encoder is responsible for extracting features from the input insulator images, while the decoder is responsible for upsampling the intermediate features and generating the final output. Moreover, the encoder and decoder are symmetric and connected by paths. The feature maps are passed through the encoder composed of repeated convolutional layers and max pooling layers. These pooling layers extract intermediate features, and then these extracted features are upsampled through the corresponding decoder. The insulator feature copies in the encoder are connected to the insulator features in the decoder through the connection paths. The last layer generates the output, which is a mask here. At this time, the loss function value corresponding to a real mask can be calculated, and the gradient is backpropagated through the network to improve the prediction ability of the model.

[0039] The encoder consists of 3×3 convolutional layers repeated in each stage. After each convolutional layer, the ReLu activation function is applied element-wise to each feature map. Between each stage, a 2×2 max pooling operation downsamples the features. This is an operation with a stride of 2, equivalent to rolling a non-overlapping window over the image and selecting the maximum value. This reduces the spatial dimension of the features. To compensate for this, the number of channels is doubled after each downsampling operation. This is the process of encoding the features of the insulator, and part of its working process is as Figure 4 shown.

[0040] The decoder is, in many ways, the reverse of the encoder. It also consists of a series of 3×3 convolutional layers, each followed by the ReLu activation function. Different from using max pooling for downsampling, the decoder upsamples the current feature set and then applies a 2×2 convolutional layer to halve the number of channels. The upsampling operation is used to restore the spatial resolution of the features lost in the encoding stage. After multiple stages, the feature map with the insulator is represented by a mask, and part of its working process is as Figure 5 shown.

[0041] There are two types of connections between the encoder and the decoder, which are called bottlenecks and connection paths. The connection paths are essentially copying and cropping. They simply copy the features of the symmetric part of the encoder and connect them to the corresponding stages in the decoder. This means that the convolutional layers in the decoder can operate on the features of both the decoder and the encoder at the same time. Simply put, for the size of the output of each stage, it needs to be copied and centrally cropped to facilitate splicing with the size generated by subsequent upsampling. The decoded features can contain more semantic information, while the encoded features contain more spatial information.

[0042] The other connection method is the bottleneck, which is where the encoder is transformed into the decoder. First, the features are downsampled, then passed through recognizable convolutional layers, and finally upsampled again to the corresponding resolution before the bottleneck. The process is as Figure 6 shown.

[0043] The attention mechanism is a data processing method in machine learning and is widely used in various types of machine learning tasks such as image segmentation. The degree of attention (importance) to different information is reflected by weights. In this task, the present invention only focuses on the insulator part in the image and does not care about other parts of the image. Therefore, the attention mechanism is introduced in this task to enable a better network to segment the insulator in the image. According to the attention focus domain, it can be divided into the spatial domain and the channel domain.

[0044] The channel attention mechanism (CAM) obtains the channel attention mechanism through the relationship between features inside the features. Each channel of the feature map is regarded as a feature detection, and its structure is as Figure 7 shown.

[0045] The idea and process of the present invention using the channel attention mechanism are as follows: First, perform global max pooling and global average pooling on the input insulator feature map F with a size of H×W×C in the spatial dimension to obtain two insulator feature maps of 1×1×C; (pooling in the spatial dimension compresses the spatial size to facilitate learning the features of the channels later) Then, send the results of global max pooling and global average pooling into a shared multi-layer perceptron (MLP) for learning respectively to obtain two insulator feature maps of 1×1×C. The number of neurons in the first layer of the MLP is C / r, the activation function is Relu, and the number of neurons in the second layer is C; Finally, perform the Add operation on the results output by the MLP, and then perform mapping processing through the Sigmoid activation function to finally obtain the channel attention weight matrix M C 。

[0046] The channel attention weight matrix M c can be expressed as: (3); To reduce the calculation parameters, a dimensionality reduction coefficient r is adopted in the MLP 。

[0047] In summary, the calculation formula of the channel attention is as follows: (4); In the above formula, and respectively represent the global average pooling feature and the max pooling feature.

[0048] Generate the spatial attention feature map through the relationship inside the feature map space, and the spatial attention process is as Figure 8As shown in the figure, first, global max pooling and global average pooling are performed on an input feature map F with dimensions H×W×C in the channel dimension to obtain two insulator feature maps of H×W×1; (pooling in the channel dimension compresses the channel size to facilitate learning spatial features later), then, the results of global max pooling and global average pooling are concatenated along the channel dimension (concat) to obtain a feature map with dimensions H×W×2. Finally, a 7×7 convolution operation is performed on the concatenated result to obtain a feature map with dimensions H×W×1, and then through the Sigmoid activation function, a spatial attention weight matrix M is obtained S 。

[0049] Spatial attention weight matrix , which can be expressed as: (5); Similarly, two pooling methods are used in the channel dimension to generate 2D feature maps: (6); (7); In summary, the calculation formula for spatial attention is as follows: (8); The Convolutional Block Attention Module (CBAM) is a representative model of the hybrid attention mechanism. It includes a channel attention module and a spatial attention module. The model structure of CBAM is as follows. In this model, for the input insulator feature map, it first undergoes channel attention module processing; the obtained result then undergoes spatial attention module processing, and finally, the adjusted insulator feature map is obtained, as Figure 9 shown

[0050] Example 2: As Figure 10 shown, this implementation provides a substation insulator segmentation system based on BOX supervision, including: An image acquisition unit, configured to: acquire a substation image to be segmented; An insulator segmentation unit, configured to: obtain the insulator segmentation result in the substation image according to the substation image and a pre-trained U-Net network based on spatial attention mechanism and channel attention mechanism; Among them, the training of the U-Net network includes: annotating the insulator part in the training set images with BOX, inputting each training set image after BOX annotation into the SAM segmentation model and the GrabCut segmentation model respectively, taking the intersection of the two segmentation results to obtain the training set labels corresponding to each training set image, and training the U-Net network according to the training set images and the training set labels.

[0051] For the specific training process and network structure, please refer to the introduction in Embodiment 1, which will not be elaborated here.

[0052] It can be understood that the above-mentioned units can be separately or wholly combined into one or several other units to form, or some of them can be further split into multiple smaller units with functional division to form, which can achieve the same operation without affecting the implementation of the technical effects of the embodiments of this application. The above units are divided based on logical functions. In practical applications, the function of one unit can be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of this application, the system may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.

[0053] According to another embodiment of this application, the system described in this embodiment can be constructed and the method of Embodiment 1 of this application can be implemented by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). The computer program can be recorded on a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.

[0054] Embodiment 3: As Figure 11 shown, this implementation provides an electronic device, which includes a processor 1001, a communication interface 1002, and a computer-readable storage medium 1003. Among them, the processor 1001, the communication interface 1002, and the computer-readable storage medium 1003 can be connected through a bus or other means.

[0055] Among them, the communication interface 1002 is used to receive and send data. The computer-readable storage medium 1003 can be stored in the memory of the electronic device. The computer-readable storage medium 1003 is used to store computer programs, and the computer programs include program instructions. The processor 1001 is used to execute the program instructions stored in the computer-readable storage medium 1003.

[0056] The processor 1001 (or CPU (Central Processing Unit, central processing unit)) is the computing core and control core of the electronic device, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0057] The processor 1001 is configured to execute the following process: Obtain the substation image to be segmented; According to the substation image and the pre-trained U-Net network based on the spatial attention mechanism and the channel attention mechanism, obtain the insulator segmentation result in the substation image; For the specific training process and network structure, see the introduction in Embodiment 1, which will not be elaborated here.

[0058] Embodiment 4: This implementation provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the electronic device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.

[0059] Moreover, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0060] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the following process: Obtain the substation image to be segmented; According to the substation image and the pre-trained U-Net network based on the spatial attention mechanism and the channel attention mechanism, obtain the insulator segmentation result in the substation image; For the specific training process and network structure, please refer to the description in Embodiment 1, which will not be elaborated here.

[0061] Embodiment 5: This implementation provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to perform the following process: Obtain a substation image to be segmented; According to the substation image and the pre-trained U-Net network based on the spatial attention mechanism and the channel attention mechanism, obtain the insulator segmentation result in the substation image; For the specific training process and network structure, please refer to the description in Embodiment 1, which will not be elaborated here.

[0062] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0063] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0064] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A substation insulator segmentation method based on BOX supervision, characterized in that, Including the following processes: Obtain the substation image to be segmented; According to the substation image and the pre-trained U-Net network based on the spatial attention mechanism and the channel attention mechanism, obtain the insulator segmentation result in the substation image; Among them, the training of the U-Net network includes: annotating the insulator part in the training set image with a BOX, inputting each training set image after BOX annotation into the SAM segmentation model and the GrabCut segmentation model respectively, taking the intersection of the two segmentation results to obtain the training set label corresponding to each training set image, and training the U-Net network according to the training set image and the training set label.

2. The substation insulator segmentation method based on BOX supervision according to claim 1, wherein Taking the intersection of the two segmentation results includes: for any training set image, obtaining the first segmentation result output by the SAM segmentation model and the second segmentation result output by the GrabCut segmentation model, performing an AND operation on the first segmentation result and the second segmentation result pixel by pixel to obtain the training set label corresponding to this training set image.

3. The substation insulator segmentation method based on BOX supervision according to claim 1, wherein The U-Net network based on the spatial attention mechanism and the channel attention mechanism includes: an encoder, a decoder, channel attention processing, and spatial attention processing; The encoder includes convolutional operations in multiple stages. After each stage of convolutional operation, max pooling operation is used to downsample the features. The decoder includes convolutional operations in multiple stages. After each stage of convolutional operation, the features are upsampled. The number of stages of convolutional operations in the encoder is the same as the number of stages of convolutional operations in the decoder; After each stage of convolutional operation in the encoder, channel attention processing and spatial attention processing are sequentially performed on the features. The result of spatial attention processing is output to the corresponding convolutional operation stage of the decoder for splicing operation.

4. The substation insulator segmentation method based on BOX supervision according to claim 3, wherein Channel attention processing includes: Performing global max pooling and global average pooling on the spatial dimension of an input insulator feature map F with a size of H×W×C to obtain two insulator feature maps of 1×1×C; Sending the two insulator feature maps of 1×1×C into a shared multi-layer perceptron for learning to obtain two processed insulator feature maps of 1×1×C; Performing an addition operation on the results output by the multi-layer perceptron during learning, and then performing mapping processing through the Sigmoid activation function to finally obtain the channel attention weight matrix, and obtaining the feature map after channel attention processing according to the channel attention weight matrix.

5. The substation insulator segmentation method based on BOX supervision according to claim 3, wherein Spatial attention processing includes: Performing global max pooling and global average pooling on the channel dimension of an input feature map F with a size of H×W×C to obtain two insulator feature maps of H×W×1; Two insulator feature maps of H×W×1 are concatenated along the channels to obtain a feature map with a size of H×W×2. Finally, a 7×7 convolution operation is performed on the concatenated result to obtain a feature map with a size of H×W×1. Then, through the Sigmoid activation function, a spatial attention weight matrix is obtained, and the feature map after spatial attention processing is obtained according to the spatial attention weight matrix.

6. A substation insulator segmentation system based on BOX supervision, characterized in that, It includes: An image acquisition unit configured to: acquire a substation image to be segmented; An insulator segmentation unit configured to: obtain the insulator segmentation result in the substation image according to the substation image and a pre-trained U-Net network based on a spatial attention mechanism and a channel attention mechanism; Among them, the training of the U-Net network includes: annotating the insulator part in the training set image with a BOX, respectively inputting each training set image annotated with a BOX into the SAM segmentation model and the GrabCut segmentation model, taking the intersection of the two segmentation results to obtain the training set label corresponding to each training set image, and training the U-Net network according to the training set image and the training set label.

7. The substation insulator segmentation system based on BOX supervision according to claim 6, wherein In the insulator segmentation unit, the U-Net network based on a spatial attention mechanism and a channel attention mechanism includes: an encoder, a decoder, channel attention processing, and spatial attention processing; The encoder includes convolutional operations in multiple stages. After each stage of convolutional operation, max pooling operation is used to downsample the features. The decoder includes convolutional operations in multiple stages. After each stage of convolutional operation, the features are upsampled. The number of stages of convolutional operations in the encoder is the same as the number of stages of convolutional operations in the decoder; After each stage of convolutional operation in the encoder, channel attention processing and spatial attention processing are sequentially performed on the features, and the result of spatial attention processing is output to the corresponding convolutional operation stage of the decoder for concatenation operation.

8. A computer device, characterized in that, It includes: A processor and a computer-readable storage medium; The processor is adapted to execute a computer program; The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the substation insulator segmentation method based on BOX supervision according to any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the substation insulator segmentation method based on BOX supervision according to any one of claims 1 to 5.

10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the substation insulator segmentation method based on BOX supervision according to any one of claims 1 to 5.