Insulator instance segmentation method and system based on improved YoloV8 network

By improving the YOLOv8 network structure, including replacing the Bottleneck module in the neck network with a star-shaped structure, the problem of insufficient detection accuracy of the insulator instance segmentation algorithm in the prior art is solved, and high-precision insulator detection on resource-constrained devices is realized.

CN120088478APending Publication Date: 2025-06-03NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510157144.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing insulator instance segmentation algorithm based on lightweight YOLOv8 network cannot maintain sufficient detection accuracy while ensuring the detection speed. Especially in the case of complex background interference, the instance segmentation effect of small insulator targets is poor.

Method used

By improving the YOLOv8 network structure, including replacing the Bottleneck module in the neck network with a star-shaped structure, and corresponding improvements are made in the backbone network and the segmentation head to build a lightweight insulator instance segmentation model.

Benefits of technology

It realizes that while reducing the amount of network parameters, maintain or improve detection accuracy, so that the model can better adapt to resource-constrained equipment and provide high-precision insulator detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088478A_ABST
    Figure CN120088478A_ABST
Patent Text Reader

Abstract

According to the lightweight insulator instance segmentation method and system based on the improved YloV8 network, the current YloV8 network structure is improved, the parameter quantity of the YloV8 network is reduced, and meanwhile the high precision of the YloV8 network is guaranteed. The improved YoloV8 network is utilized to construct an insulator instance segmentation model, and better detection precision can be achieved with less parameter quantity, so that the model can well adapt to the condition that resources of mobile equipment are limited, and the insulator detection capability with higher precision is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly relates to an insulator instance segmentation method and system based on an improved YoloV8 network. Background Art

[0002] In the power system, the safe and stable operation of transmission lines is crucial. As a key insulating component in transmission lines, the state of insulators directly affects the reliability of the power grid. Currently, drones equipped with insulator instance segmentation algorithms are usually used to achieve automatic inspection of transmission lines and automatic detection of insulators. Considering the resource limitations of drone terminals, an insulator instance segmentation algorithm based on the lightweight YOLOv8 network is widely used at present. However, although the lightweight YOLOv8 network runs faster on resource-constrained devices, they sacrifice a certain network accuracy. In the application of transmission line instance segmentation, in the face of complex background interference, for the instance segmentation of small insulator targets, the network needs to have more accurate feature extraction and identification capabilities, which puts higher requirements on the detection accuracy of the network. The existing lightweight YOLOv8 networks cannot maintain the detection accuracy while ensuring the detection speed, which has become an urgent problem to be solved in the industry. Summary of the Invention

[0003] One or more embodiments of this specification describe an insulator instance segmentation method and system based on an improved YoloV8 network, which can achieve better detection accuracy with fewer parameters.

[0004] In a first aspect, a lightweight insulator instance segmentation method based on an improved YoloV8 network is provided. The YoloV8 network includes a backbone network, a neck network, and a segmentation head. The method includes:

[0005] Replace the bottleneck module in the C2f module in the neck network with a star-shaped Star_Block module to obtain an improved C2f_Star module. The Star_Block module is used to send the input feature image into the first 7×7 depthwise separable convolution kernel for convolution dimensionality reduction, then send the dimensionality reduction results into two 1×1 convolution kernels for convolution dimensionality increase respectively. Then, multiply the two parts of the dimensionality increase results element-wise, send the product into a 1×1 convolution kernel for convolution dimensionality reduction again, and then send it into the second 7×7 depthwise separable convolution kernel for convolution processing. The residual formed by this convolution result and the feature image is used as the output result of the Star_Block module.

[0006] Train the YoloV8 network with pre-collected sample images to obtain an insulator instance segmentation model.

[0007] Use the insulator instance segmentation model to detect the target image to be detected, and obtain the insulator instance segmentation result in the target image.

[0008] As an optional implementation of the method described in the first aspect, the backbone network includes the first to fifth layer SubNet sub-networks and a spatial pyramid pooling module, where the output features of the third layer SubNet sub-network, the output features of the fourth layer SubNet sub-network, and the output features of the spatial pyramid pooling module are used as the three-way output features of the backbone network.

[0009] Specifically, the first to fifth layer SubNet sub-networks are used to perform layer-by-layer feature extraction on the input image and achieve feature decoupling during the layer-by-layer feature extraction process; each layer of the SubNet sub-network contains N sub-modules, and the N sub-modules are used to perform feature extraction from shallow to deep on the input features of this layer. Moreover, in adjacent layers of the SubNet sub-networks, there is a reversible connection between the shallow sub-modules of the higher-level SubNet sub-network and the deep sub-modules of the lower-level SubNet sub-network.

[0010] As an optional implementation of the method described in the first aspect, the segmentation head includes 3 1×1 convolutional kernels and two 3×3 convolutional kernels, and the weights of the two 3×3 convolutional kernels are shared; the 3 1×1 convolutional kernels respectively perform convolutional processing on the three-way output features output by the neck network. The three processed results are subjected to shared convolution through the first 3×3 convolutional kernel and the second 3×3 convolutional kernel, and the results of the shared convolution are respectively sent into the classification head and the regression head to perform classification tasks and regression tasks.

[0011] Specifically, the segmentation head further includes a scaling layer, and the scaling layer is used to perform feature scaling on the output result of the regression head.

[0012] In the second aspect, a lightweight insulator instance segmentation system based on the improved YoloV8 network is provided. The system includes:

[0013] A data acquisition module configured to acquire the target image to be detected;

[0014] A training module, configured to train an improved YoloV8 network using pre-collected sample images to obtain an insulator instance segmentation model; the improved YoloV8 network includes a backbone network, a neck network, and a segmentation head. Among them, the bottleneck module in the C2f module in the neck network is replaced by a star-shaped Star_Block module to obtain an improved C2f_Star module; the Star_Block module is used to send the input feature image into a first 7×7 depthwise separable convolution kernel for convolution dimensionality reduction, and then send the dimensionality reduction results into two 1×1 convolution kernels for convolution dimensionality increase. Then, the two upsampled results are multiplied element-wise, and the product is sent into a 1×1 convolution kernel for convolution dimensionality reduction again and then sent into a second 7×7 depthwise separable convolution kernel for convolution processing. The residual formed by the convolution result and the feature image is used as the output result of the Star_Block module.

[0015] A detection module, configured to use the insulator instance segmentation model to detect a target image to be detected and obtain the insulator instance segmentation result in the target image.

[0016] As an optional implementation manner of the system described in the second aspect, the backbone network includes a first-layer to fifth-layer SubNet sub-network and a spatial pyramid pooling module. Among them, the output features of the third-layer SubNet sub-network, the output features of the fourth-layer SubNet sub-network, and the output features of the spatial pyramid pooling module are used as the three-way output features of the backbone network.

[0017] Specifically, the first-layer to fifth-layer SubNet sub-networks are used to perform layer-by-layer feature extraction on the input image and achieve feature decoupling during the layer-by-layer feature extraction process; each layer of the SubNet sub-network contains N sub-modules, and the N sub-modules are used to perform feature extraction from shallow to deep on the input features of this layer. And in adjacent two-layer SubNet sub-networks, there is a reversible connection between the shallow sub-module of the high-layer SubNet sub-network and the deep sub-module of the low-layer SubNet sub-network.

[0018] As an optional implementation manner of the system described in the second aspect, the segmentation head includes 3 1×1 convolution kernels and two 3×3 convolution kernels, and the weights of the two 3×3 convolution kernels are shared; the 3 1×1 convolution kernels respectively perform convolution processing on the three-way output features output by the neck network. The three processed results are subjected to shared convolution through the first 3×3 convolution kernel and the second 3×3 convolution kernel, and the results of the shared convolution are respectively sent into a classification head and a regression head to perform classification tasks and regression tasks.

[0019] Specifically, the segmentation head further includes a scale layer, which is used to perform feature scaling on the output result of the regression head.

[0020] Beneficial effects: By improving the current YoloV8 network structure, the present invention reduces the number of parameters of the YoloV8 network while ensuring a relatively high accuracy of the YoloV8 network. Using the improved YoloV8 network to construct an insulator instance segmentation model can achieve better detection accuracy with fewer parameters, enabling the model to well adapt to the situation of resource constraints of mobile devices and providing a relatively high-precision insulator detection ability. Description of the Drawings

[0021] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a schematic flowchart of an insulator instance segmentation method based on an improved YoloV8 network involved in the embodiments of this specification.

[0023] Figure 2 It is a schematic structural diagram of the backbone network of the existing Yolov8 network.

[0024] Figure 3 It is a schematic structural diagram of the backbone network involved in the embodiments of this specification.

[0025] Figure 4 It is a schematic structural diagram of the SubNet sub-network involved in the embodiments of this specification.

[0026] Figure 5 It is a schematic structural diagram of the sub-module Level involved in the embodiments of this specification.

[0027] Figure 6 It is a schematic structural diagram of the Fusion module involved in the embodiments of this specification.

[0028] Figure 7 It is a schematic structural diagram of the neck network of the existing Yolov8 network.

[0029] Figure 8 It is a schematic structural diagram of the Bottleneck module in the C2f module of the existing Yolov8 network.

[0030] Figure 9It is a schematic structural diagram of the C2f_Star module involved in the embodiments of this specification.

[0031] Figure 10 It is a schematic structural diagram of the Star_Block module involved in the embodiments of the specification.

[0032] Figure 11 It is a schematic structural diagram of the LSC-Segment segmentation head involved in the embodiments of the specification.

[0033] Figure 12 It is a schematic structural diagram of an improved YoloV8 network involved in the embodiments of the specification.

[0034] Figure 13 It is involved in the embodiments of the specification using Figure 12 The schematic diagram of the segmentation result after the insulator instance segmentation of the target image to be detected by the insulator instance segmentation model trained by the improved YoloV8 network shown.

[0035] Figure 14 It is a schematic structural diagram of an insulator instance segmentation system based on an improved YoloV8 network involved in the embodiments of the specification. Detailed implementation manners

[0036] First of all, it should be noted that the terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0037] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of the embodiments. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0038] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, the single steps described in this specification may be decomposed into multiple steps for description in other embodiments; and the multiple steps described in this specification may also be combined into a single step for description in other embodiments.

[0039] In the process of using drones for automatic inspection to achieve instance segmentation of insulators in transmission lines, due to the insufficient detection accuracy of the existing instance segmentation algorithm based on the YoloV8 network, resulting in poor accuracy of detection results, and the high network complexity of the existing instance segmentation algorithm based on the YoloV8 network, resulting in slow detection speed, one or more embodiments of this specification propose an insulator instance segmentation method and system based on an improved YoloV8 network to at least partially solve the above technical problems.

[0040] The following will further elaborate on the insulator instance segmentation method and system based on the improved YoloV8 network described in one or more embodiments of this specification in combination with the accompanying drawings of the specification and specific embodiments, but this detailed description does not constitute a limitation on the embodiments of this specification.

[0041] Please refer to Figure 1 , Figure 1 which schematically shows a flowchart of an insulator instance segmentation method based on an improved YoloV8 network. The method includes steps S100 to S104:

[0042] S100: Improve the YoloV8 network to obtain an improved YoloV8 network.

[0043] The YoloV8 network mainly includes three parts, namely the backbone network, the neck network, and the segmentation head.

[0044] In some embodiments, the backbone network can be improved. The existing backbone network of Yolov8, as Figure 2 shown, adopts a 10-layer DarkNet structure, mainly composed of Conv modules, C2f modules, and a spatial pyramid pooling module SPPF module. Although this network structure can effectively extract the features of insulator targets, the number of layers in the backbone network is too large, greatly increasing the complexity of the model, and some information of insulator targets will be lost during the feature extraction process.

[0045] To maintain the integrity of information during the feature extraction process of the backbone network and avoid information loss and compression, this embodiment improves the 10-layer Darken backbone network in the Yolov8 network to a 6-layer backbone network as Figure 3 shown. As Figure 3As shown in the figure, the improved backbone network includes the first to fifth layer SubNet sub-networks and a spatial pyramid pooling module. These five SubNet sub-networks are denoted as SubNet0, SubNet1, SubNet2, SubNet3, and SubNet4 from top to bottom. The structure of each SubNet sub-network is as Figure 4 shown, and it contains N sub-modules Level. The N sub-modules Level in the same layer sequentially perform feature extraction on the feature image input to this layer to obtain features from shallow to deep. In addition, in adjacent SubNet sub-networks, the deep sub-module Level of the upper layer is reversibly connected to the shallow sub-module Level of the lower layer. Based on Figure 4 the network structure shown in the figure, on the one hand, the input image undergoes layer-by-layer feature extraction along the column direction through SubNet0, SubNet1, SubNet2, SubNet3, and SubNet4, and feature decoupling is achieved during this process. On the other hand, the feature information of the input image is transmitted between parallel columns through reversible connections. Through this serial-parallel cross-level connection between Levels, it can be ensured that during the forward propagation process, each column can share the information of other columns, so that the shallow to deep features output by SubNet4 contain target features at different levels, which can not only avoid the problem of information loss but also help improve the overall performance of the model, thereby enhancing the model's ability to express and understand complex data.

[0046] The module structure of the sub-module Level in the SubNet sub-network is as Figure 5 shown, and it is composed of a Fusion Block module and a C2f module connected in series. Among them, the Fusion module is as Figure 6 shown. It performs a Conv convolution operation on the low-resolution features (Low Resolution Features), and then inputs them into the upsampling (Upsample) module. At the same time, it performs a Conv2d convolution operation on the high-resolution features (High Resolution Features), and then inputs them into the batch normalization (BatchNorm2d) module, and adds the results of the two to obtain the final output. It can be seen that the Level module extracts features at different levels, recognizes complex objects and scenes, and greatly improves the performance and accuracy of the model.

[0047] In this embodiment, a 6-layer backbone network design is adopted, and its purpose is to align the data channels with the neck network. Specifically, it extracts the output features of the SubNet2, SubNet3, and SPPF modules in the backbone network, and then uses the features of these three different scales as the input of the neck network. By improving the 10-layer Darken backbone network in the Yolov8 network to be asFigure 3 The 6-layer backbone network shown can reduce the number of network layers and parameters of the improved Yolov8 network while ensuring that the network accuracy remains unchanged.

[0048] In some embodiments, the Neck network can be improved. The structure of the Neck network in the existing Yolov8 is as Figure 7 shown. After the input data is convolved through the first convolutional layer, the convolutional result is divided into two parts. One part is processed through n Bottleneck modules and then concatenated with the other part in the channel dimension. The concatenated feature data is convolved through the second convolutional layer to obtain the final output. The structure of the Bottleneck module is as Figure 8 shown. First, a 3x3 convolution is used to reduce the dimension and eliminate redundant information, then a 3×3 convolution is used to increase the dimension, and then added to another result. By reducing the spatial dimension of the image and continuously increasing the number of channels, hierarchical features in the image are obtained. However, although this structure can obtain rich gradient flow information during the feature extraction process, at the same time, multiple feature maps carry the same information, resulting in a certain amount of information redundancy and increasing the model complexity, leading to more memory occupation.

[0049] In this embodiment, the Bottleneck module in the C2f module in the Neck network is replaced with a Star_Block module with a star structure to obtain an improved C2f_Star module. The structure of the improved C2f_Star module is as Figure 9 shown. The Star_Block module with a star structure is as Figure 10 shown. It includes two 7×7 depthwise separable convolution kernels, three 1×1 convolution kernels, and a multiplication module. The processing flow of the Star_Block module for the input feature map is as follows: First, the first 7×7 depthwise separable convolution kernel is used to perform depthwise separable convolution on the input data, reducing the parameters of the same-level ordinary convolution and streamlining the calculation model. Then, two 1×1 convolution kernels are used to perform convolution to increase the dimension on the results of the depthwise separable convolution respectively, and then the multiplication module is used to perform element-wise multiplication on the two upsampled results. The result of the element-wise multiplication is then convolved and downsampled through a 1×1 convolution kernel. Thus, the important upsampling and downsampling operations on the input feature map are completed. Finally, the second 7×7 depthwise separable convolution kernel is used to perform depthwise separable convolution on the result of the convolution downsampling, increasing the network depth. The finally obtained feature map forms a residual network with the input feature map, improving the expression ability and training stability of the model.

[0050] Since there are many C2f modules in the neck network of the YOLOv8 network, and the C2f module uses relatively inefficient addition operations, in this embodiment, the traditional addition operation in C2f is improved to a faster multiplication operation, making the overall C2f module lighter and reducing redundant calculations. At the same time, replacing the C2f module with a C2f_Star module with a simpler star connection structure can effectively reduce the model complexity and the memory footprint of the entire network.

[0051] In some embodiments, the segmentation head of the YoloV8 network can also be improved. The original YoloV8 network model goes through segmentation heads corresponding to output layers with three different depths, which will increase a large number of parameters and computational amounts. Based on this, this embodiment proposes a brand-new segmentation head (hereinafter referred to as the LSC-Segment segmentation head). The structure of the LSC-Segment segmentation head is as Figure 11 shown, including three 1×1 convolutional kernels and two 3×3 convolutional kernels, and the weights of the two 3×3 convolutional kernels are shared. The LSC-Segment segmentation head receives input features at three different levels (P3 to P5). First, three 1×1 convolutional kernels are used to perform 1×1 convolutional processing on these three input features one by one to increase the information exchange in the channel dimension. Then, shared convolutions of two 3×3 convolutional kernels are used for information aggregation, reducing redundant information and increasing the learning probability of adjacent feature information. Finally, the information extracted by the shared convolution is input into the classification head and the regression head, and the output features of the regression head can also be feature-scaled through a Scale layer to enhance the ability to retain multi-scale features.

[0052] In this embodiment, the mechanism of sharing convolutional weight parameters is used to reconstruct the segmentation head structure, introducing a brand-new lightweight shared convolutional segmentation head (LSCD-Segment) to further reduce the number of model parameters and improve the model detection performance, realizing the improvement of the model's multi-scale perception ability, and making the improved YOLOv8 network more suitable for resource-constrained devices.

[0053] S102: Train the improved YoloV8 network using pre-collected sample images to obtain an insulator instance segmentation model.

[0054] The above-mentioned sample images refer to images containing insulators, and these images can be from drone aerial images. After obtaining these sample images, it is necessary to mark the insulator category and the insulator position box in the sample images as training labels.

[0055] During training, the above sample images are input into the improved YoloV8 network to obtain insulator detection results, which include the class prediction results of the detection targets and the detection box prediction results. Based on the insulator detection results and the corresponding training labels, a loss function is constructed to update the parameters of the YoloV8 network, thereby obtaining an insulator instance segmentation model that meets the requirements.

[0056] S104: Use the insulator instance segmentation model to detect the target image to be detected, and obtain the insulator instance segmentation result in the target image.

[0057] In the application scenario of unmanned aerial vehicle (UAV) automatic inspection and detection of transmission line insulators, the aerial images of the UAV can be obtained as the above target images to be detected, and then the aerial images are input into the insulator instance segmentation model to obtain the insulator instance segmentation result in the target image.

[0058] To prove the technical effect of the insulator instance segmentation method based on the improved YoloV8 network described in this embodiment, the technical effect of this method will be verified below in combination with specific experimental data.

[0059] Light reference Figure 12 , Figure 12 shows a schematic structural diagram of an improved YoloV8 network. The above improvements to the backbone network, neck network, and detection head in step S100 are respectively made to this improved YoloV8 network. That is, for the backbone network part, the 10-layer DarkNet backbone network in the original Yolov8 network is improved to a 6-layer backbone network SubNet. For the neck network part, the C2f module in the original Yolov8 network is improved to a C2f_Star module with a star structure. For the detection head part, the original detection head is improved to a lightweight shared convolutional segmentation head (LSCD-Segment). By improving the three parts of the backbone network, neck network, and detection head, the Figure 12 shown improved YoloV8 network is obtained.

[0060] In this embodiment, 3590 UAV aerial images containing insulators are selected as the Figure 12 dataset of the shown improved YoloV8 network. This dataset is divided into a training set and a validation set. The training set contains 3231 UAV aerial images, and the validation set contains 359 UAV aerial images. Use the above training set to train the Figure 12 shown improved YoloV8 network to obtain an insulator instance segmentation model.

[0061] In this embodiment, to ensure the reliability and effectiveness of the comparative experiment, the algorithm is run on the same device, and the parameters selected for the same network are kept consistent. The hardware, software environment, and training parameter settings are shown in Table 1.

[0062] Table 1 Experimental environment and training parameters

[0063]

[0064] In this experiment, the mean average precision (mAP), parameters, Gflops, weights, and FPS are used as parameter indicators. Among them, mAP demonstrates the performance of the model, and its calculation formula is as follows:

[0065]

[0066] Among them, TP is the number of correctly detected targets; FP is the number of incorrectly detected targets; FN is the number of missed detections and false negatives; the mean average precision mAP is the average of the APs for all classes. However, in this experiment, there is only one class, namely the insulator, so mAP is equal to AP. Parameters refer to the total number of parameters in the model to judge the complexity of the model. Gflops is the number of floating-point operations that the model can perform per second to judge the computing power of the model. The higher the Gflops, the higher the computing requirements for the edge computing hardware. Weights refer to the size of the trained network model file, with the unit of M (megabyte). FPS represents the number of video frames that can be processed per second. The higher the FPS, the smoother the video.

[0067] To verify the performance of the backbone network proposed in this embodiment, a comparative experiment is conducted by comparing the backbone network SubNet with other lightweight backbone network models. SubNet is compared with Yolov8-n, efficientViT, fasternet, MobileNetV3, and MobileNetV4 respectively. The experimental results are shown in Table 2.

[0068] Table 2 Comparison of experimental data of different backbone networks

[0069]

[0070] As can be seen from the data in Table 2, compared with the Yolov8 network itself, the backbone network SubNet has improved in both mAP and FPS, with improvements of 0.1% and 9.6% respectively. At the same time, SubNet has a more obvious improvement in the lightweight index, with a decrease of 22.4%, 15.8% and 20.6% in parameters, Gflops and weights respectively. When the mAP index increases, the parameters, Gflops and weight indexes decrease significantly, and the FPS index also increases, proving that the backbone network SubNet is a superior lightweight model, reducing the complexity of the model and the computational burden.

[0071] While the mAP of efficientViT and FasterNet increases, it is accompanied by an increase in the number of parameters, Gflops and weights, increasing the computational complexity. Although the MobileNetV3 model has a significant decrease in parameters, Gflops and weights, the mAP parameter decreases severely and the model accuracy is severely lacking. MobileNetV4 not only increases the number of parameters but also loses accuracy and frames. Through comprehensive analysis, it is concluded that the backbone network SubNet proposed in this embodiment has better multi-scale feature fusion ability, reduces redundant calculations, retains more key information, and realizes the lightweight of the model.

[0072] To verify the lightweight performance of the C2f_star module proposed in the present invention, a comparative experiment was conducted by comparing the improvement with other lightweight models. C2f_star was compared with Yolov8-n, C2f_faster, c2 and c3 respectively. The experimental results are shown in Table 3.

[0073] Table 3 Comparison of experimental data of different C2f modules

[0074]

[0075]

[0076] As can be seen from the data presented in Table 3, C2f_Star has a significant improvement on the map. Compared with the original Yolov8, the map has increased by 1.0%, and there are decreases in parameters, Gflops, and weights, which are 6.1%, 3.3%, and 5.9% respectively. This proves that C2f_Star is a lightweight model, reducing the computational complexity while improving the model's accuracy. The lightweight metrics of C2f_faster have decreased, but the accuracy cannot be maintained normally. C2 has a small decrease in parameters, Gflops, and weights, and a small increase in the map. Compared with C2f_Star, the effect of C2 is worse and it cannot achieve obvious lightweighting. When the parameters of C3 have decreased by 7.8%, the map has only increased by 0.1%. Compared with C2f_Star, the effect is not significant. Through comprehensive analysis, it is concluded that C2f_Star proposed in this embodiment reduces information loss, retains more key information, greatly increases the accuracy of the module, and at the same time achieves the effect of lightweighting.

[0077] In order to more intuitively and specifically understand the gain effect of each proposed module on the network structure, the experimental results of multiple ablation experiments are also provided in this embodiment, as shown in Table 4.

[0078] Experimental results of ablation experiments in Table 4

[0079]

[0080] As can be seen from the ablation experiment results in Table 4, all the various improvement methods proposed in this embodiment can have a positive impact. When using the SubNet module alone, the parameters are reduced by 22.4%. At the same time, the map and FPS are improved, increasing by 0.1% and 9.6% respectively, increasing the accuracy and frame rate of the model and achieving the lightweight of the model. When using the c2f_star model alone, there is a slight reduction in parameters, but there is a significant increase in the map, with a 1.0% increase. c2f_star reduces the loss of data information and increases the accuracy of the module. When using the LSC-Segment module alone, the parameters decrease significantly, by 23.2%. The map can be stably maintained, proving that the lightweight shared convolution can not only reduce the computational amount of the model but also effectively maintain the model accuracy. When using SubNet and c2f_star simultaneously, while achieving a decrease in parameters, the map also increases. The parameters decrease by 28.5% and the map increases by 0.6%. When SubNet and LSC-Segment act together, the parameters are greatly reduced, by 45.6%, and the FPS is greatly improved, by 16.1%. When c2f_star and LSC-Segment act together, the parameters decrease, by 29.4%, and the map is stably maintained without a decrease. Finally, when Revcol, c2f_star, and LSC-Segment act simultaneously, the parameters decrease significantly, by 51.8%, and the map and FPS also increase slightly, by 0.4% and 9.6% respectively.

[0081] In summary, the insulator instance segmentation method based on the improved YoloV8 network proposed in this embodiment improves in three parts: the backbone network, the neck network, and the detection head, successfully lightweighting the model. When the model parameters decrease by 50%, a relatively high map is maintained, the accuracy is stable, and better detection accuracy is achieved with fewer parameters, making it suitable for running on mobile devices.

[0082] Please refer to Figure 13 , Figure 13 which shows Figure 12 the segmentation result after performing insulator instance segmentation on the target image to be detected using the insulator instance segmentation model trained with the improved YoloV8 network shown. It can be seen that the insulator instance segmentation method based on the improved YoloV8 network described in this embodiment can well segment the insulators.

[0083] Corresponding to the above insulator instance segmentation algorithm based on the improved YoloV8 network, this embodiment also provides an insulator instance segmentation system based on the improved YoloV8 network. As Figure 14 shown, the system includes:

[0084] The data acquisition module is configured to acquire the target image to be detected.

[0085] The training module is configured to train the improved YoloV8 network using pre-collected sample images to obtain an insulator instance segmentation model.

[0086] The detection module is configured to detect the target image to be detected using the insulator instance segmentation model to obtain the insulator instance segmentation result in the target image.

[0087] Specifically, the above-mentioned improved YoloV8 network can be improved by using the relevant steps in step S100 in the above-mentioned insulator instance segmentation method based on the improved YoloV8 network, and this embodiment will not be repeated here.

[0088] It is to be understood that the structure illustrated in the embodiments of this specification does not constitute a specific limitation on the system of the embodiments of this specification. In other embodiments of the specification, the above system may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0089] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0090] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0091] It should be noted that the above examples are only specific embodiments of the present invention, and the present invention is obviously not limited to the above examples, and there are many similar variations. All variations directly derived or associated from the contents disclosed by the technicians in this field should fall within the protection scope of the present invention.

Claims

1. A lightweight insulator instance segmentation method based on an improved YoloV8 network, wherein the YoloV8 network includes a backbone network, a neck network and a segmentation head, characterized in that: The method comprises: The bottleneck module in the C2f module in the neck network is replaced by a Star_Block module with a star structure to obtain an improved C2f_Star module; the Star_Block module is used to send the input feature image into the first 7×7 depth-separable convolution kernel for convolution dimensionality reduction, and then send the dimension reduction results into two 1×1 convolution kernels for convolution dimensionality increase, and then perform element-wise multiplication on the two parts of the dimension increase results, and the product is sent to a 1×1 convolution kernel for convolution dimensionality reduction again, and then sent to the second 7×7 depth-separable convolution kernel for convolution processing, and the residual formed by the convolution result and the feature image is used as the output result of the Star_Block module; The YoloV8 network is trained using the pre-collected sample images to obtain an insulator instance segmentation model; The insulator instance segmentation model is used to detect the target image to be detected, and an insulator instance segmentation result in the target image is obtained.

2. The method according to claim 1, characterized in that The backbone network includes the first to fifth SubNet subnetworks and a spatial pyramid pooling module, wherein the output features of the third SubNet subnetwork, the output features of the fourth SubNet subnetwork and the output features of the spatial pyramid pooling module serve as three-way output features of the backbone network.

3. The method according to claim 2, characterized in that The first to fifth SubNet subnetworks are used to extract features of the input image layer by layer and realize feature decoupling in the process of layer by layer feature extraction; each layer of SubNet subnetwork includes N submodules, and the N submodules are used to extract features of the input features of the layer from shallow to deep layers, and in two adjacent layers of SubNet subnetworks, the shallow submodules of the high-level SubNet subnetwork and the deep submodules of the low-level SubNet subnetwork are reversibly connected.

4. The method according to claim 1, characterized in that: The segmentation head includes three 1×1 convolution kernels and two 3×3 convolution kernels, and the weights of the two 3×3 convolution kernels are shared; the three 1×1 convolution kernels perform convolution processing on the three-way output features output by the neck network one by one, and the three-way results after convolution processing are shared convolution through the first 3×3 convolution kernel and the second 3×3 convolution kernel. The results of the shared convolution are respectively sent to the classification head and the regression head to perform classification tasks and regression tasks.

5. The method according to claim 4, characterized in that The segmentation head also includes a scaling layer, and the scaling layer is used to perform feature scaling on the output result of the regression head.

6. A lightweight insulator instance segmentation system based on an improved YoloV8 network, characterized in that: include: A data acquisition module, configured to acquire a target image to be detected; A training module is configured to train an improved YoloV8 network using pre-collected sample images to obtain an insulator instance segmentation model; the improved YoloV8 network includes a backbone network, a neck network and a segmentation head, wherein the bottleneck module in the C2f module in the neck network is replaced by a Star_Block module of a star structure to obtain an improved C2f_Star module; the Star_Block module is used to send the input feature image to a first 7×7 depth-separable convolution kernel for convolution dimensionality reduction, and then send the dimension reduction results to two 1×1 convolution kernels for convolution dimensionality increase, and then perform element-wise multiplication on the two parts of the dimension increase results, and the product is sent to a 1×1 convolution kernel for convolution dimensionality reduction again, and then sent to a second 7×7 depth-separable convolution kernel for convolution processing, and the residual formed by the convolution result and the feature image is used as the output result of the Star_Block module; The detection module is configured to detect the target image to be detected by using the insulator instance segmentation model to obtain the insulator instance segmentation result in the target image.

7. The system according to claim 6, characterized in that The backbone network includes the first to fifth SubNet subnetworks and a spatial pyramid pooling module, wherein the output features of the third SubNet subnetwork, the output features of the fourth SubNet subnetwork and the output features of the spatial pyramid pooling module serve as three-way output features of the backbone network.

8. The system according to claim 7, characterized in that The first to fifth SubNet subnetworks are used to extract features of the input image layer by layer and realize feature decoupling in the process of layer by layer feature extraction; each layer of SubNet subnetwork includes N submodules, and the N submodules are used to extract features of the input features of the layer from shallow to deep layers, and in two adjacent layers of SubNet subnetworks, the shallow submodules of the high-level SubNet subnetwork and the deep submodules of the low-level SubNet subnetwork are reversibly connected.

9. The system according to claim 6, characterized in that The segmentation head includes three 1×1 convolution kernels and two 3×3 convolution kernels, and the weights of the two 3×3 convolution kernels are shared; the three 1×1 convolution kernels perform convolution processing on the three-way output features output by the neck network one by one, and the three-way results after convolution processing are shared convolution through the first 3×3 convolution kernel and the second 3×3 convolution kernel. The results of the shared convolution are respectively sent to the classification head and the regression head to perform classification tasks and regression tasks.

10. The system according to claim 9, characterized in that The segmentation head also includes a scaling layer, and the scaling layer is used to perform feature scaling on the output result of the regression head.