A target detection method for power infrastructure scenarios improved based on Yolov8

By introducing ConvFormer, CGLU network module and FSConv module in the Yolov8 model, the problem of insufficient target detection accuracy and robustness in power infrastructure scenarios is solved, and more efficient calculations and higher detection accuracy are achieved.

CN119785183BActive Publication Date: 2025-05-27ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510277003.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-05-27
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

The existing object detection algorithms have low detection accuracy in power infrastructure scenarios, poor robustness, and high computing resources consumption, making it difficult to deploy in embedded devices.

Method used

Design a module that combines ConvFormer and CGLU network, replaces the C2f network module in the Yolov8 model framework, and replaces the SPPF network module through the shared convolution module FSConv to reduce model calculation parameters and improve detection accuracy.

Benefits of technology

On the premise of ensuring the detection accuracy, the calculation parameters of the model are effectively reduced, the detection accuracy of the model in complex scenarios is improved, and the hardware requirements are reduced, making the model easier to deploy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785183B_ABST
    Figure CN119785183B_ABST
Patent Text Reader

Abstract

The present invention provides an object detection method for power infrastructure scenarios improved based on Yolov8, belonging to the technical fields of artificial intelligence and object detection. The present invention integrates artificial intelligence technology into power operation scenarios, which can prevent potential safety hazards during the operation process; by constructing a module combining ConvFormer and Convolutional Gated Linear Unit network to replace the C2f network module in the Yolov8 model framework, it can effectively reduce the calculation parameters of the model while ensuring the detection accuracy of the algorithm; at the same time, by constructing a Feature Shared Conv module to replace the SPPF network module in the Yolov8 model framework, it can capture more fine-grained features in the image and improve the detection accuracy of the model when facing complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and target detection, and specifically relates to a target detection method for power infrastructure scenarios improved based on Yolov8. Background Art

[0002] With the expansion of the scale of infrastructure projects, the construction of power at all levels such as UHV and EHV has entered a new peak. The new requirements for power construction and safety management have brought many challenges to the infrastructure safety management work. The construction site is more complex, the safety operation risks and the number of participants have increased significantly, and the supervision is insufficient, etc., making the safety problems at the infrastructure site increasingly prominent.

[0003] During the infrastructure operation process, in order to timely discover and prevent potential safety risks, employees are usually given professional safety training to improve their operation skills and safety awareness, and on-site or video monitoring is carried out manually to ensure operation safety. However, this manual method not only consumes a large amount of human resources, but also is prone to omissions or misdetections during the monitoring process. Therefore, in order to strengthen the safety control of the infrastructure site and ensure the quality and safety of construction personnel and projects, it is necessary to introduce an efficient and intelligent detection technology. Accurately identifying targets such as construction site workers, safety belts, and operation tools is a crucial task in the research.

[0004] With the development of deep learning technology, its characteristics of learning the internal structure in image data and extracting deeper features of targets by training neural networks enable deep learning to be well applied in the detection field. Therefore, using deep learning methods to perform target detection on complex infrastructure scenarios has important application value and practical significance. Currently, the most commonly used target detection algorithms, such as the Yolo series, Faster-RCNN, SSD, etc., have achieved good performance and speed in target detection tasks. However, these methods have high requirements for device computing power and are difficult to be deployed on embedded devices. At the same time, their detection accuracy is not high and the robustness is poor in complex scenarios such as illumination, occlusion, fast movement, and background interference. For some lightweight models such as the ShuffleNet series, these models reduce the dependence on hardware, but their model accuracy is greatly reduced. Therefore, in practical applications, how to reduce the model calculation parameters and improve the calculation speed without sacrificing the detection performance of the model is an urgent problem to be solved currently.

[0005] Deeply integrating artificial intelligence technology into the power operation scenario aims to improve the professional qualities of operators, standardize the work behaviors of operators, and at the same time prevent potential safety hazards during the operation process. The application of this technology not only helps to build a more standardized and standardized power operation environment, but also can promote the development of the power operation scenario towards intelligent and refined management. Summary of the Invention

[0006] In view of the above problems existing in the prior art, the purpose of the present invention is to provide an object detection method for the power infrastructure scenario improved based on Yolov8, which can capture more fine-grained features in the image, improve the detection accuracy of the model, and at the same time reduce the calculation parameters of the model by optimizing the network model to solve the problems proposed in the above background technology; specifically, it includes the following steps:

[0007] Step 1: Determine the detection targets and collect the data sets required for the research. In the real power infrastructure site, collect image data through surveillance videos. The collection can be carried out under different conditions such as light, weather, time periods, etc. to increase the diversity of the data;

[0008] Step 2: Process the collected data such as annotation and format conversion to obtain the processed data set, and divide it into an experimental training set, a test set and a validation set according to a preset ratio;

[0009] Step 3: Design a module combining ConvFormer and Convolutional Gated Linear Unit (CGLU) network to replace the C2f network module in the Yolov8 model framework, which effectively reduces the calculation parameters of the model while ensuring the detection accuracy of the algorithm;

[0010] Step 4: Design a Feature Shared Conv (FSConv) module to replace the SPPF network module in the algorithm framework. This module uses convolutional operations for feature extraction, which can capture more fine-grained features in the image and improve the detection accuracy of the model when facing complex scenarios;

[0011] Step 5: Use the data set to train the improved model, obtain the optimal object detection model based on the performance evaluation index, detect the images in the test set, and complete the object detection task.

[0012] Further, the specific process of the above Step 1 is as follows:

[0013] For the complex scenarios of power infrastructure, determine the targets to be detected, such as workers, safety helmets, safety belts, excavators, fire extinguishers, roadblocks, cranes, etc. Obtain pictures by extracting frames from the surveillance video. To ensure the effectiveness of the pictures, reduce the frame extraction interval during normal working hours.

[0014] Further, the specific process of the above Step 2 is as follows:

[0015] Remove low-quality images such as blurred and overexposed ones. Determine the detection target according to the requirements of the application scenario, use the Lableme tool to label the target of the image data, and convert the labeled data into the Yolo format required for model training. Divide the processed dataset into a training set, a test set, and a validation set according to the ratio of 8:1:1.

[0016] Further, the specific process of step 3 is as follows:

[0017] Based on the ConvFormer and CGLU network models, design a separable convolution module C2f_CFGLU to replace all C2f modules in the Yolov8 model. The core feature of ConvFormer is to use separable convolution to mix feature vector tokens, and this convolution method consists of two steps: depth convolution and pointwise convolution. Depth convolution is responsible for capturing the spatial features of the image, while pointwise convolution is responsible for capturing the relationships between channels. CGLU is a channel mixer that combines channel attention and local feature extraction, which can effectively capture local features, thereby enhancing the model's understanding and expression ability of image details. Based on the ConvFormer structure (i.e., input embedding, feature mixer Token Mixer, and multi-layer perceptron MLP, etc.), use CGLU to replace the multi-layer perceptron MPL in the ConvFormer network structure to obtain the separable convolution module C2f_CFGLU. This module can effectively enhance the robustness of the model, making the model better adapt to different application scenarios, and the number of parameters and computational complexity of this module are lower, making it easier to train and deploy.

[0018] Further, the specific process of step 4 is as follows:

[0019] Design a shared convolution module FSConv. This module constructs a feature pyramid. First, use a 1x1 convolution layer to process the input feature map and extract the main features. Then, according to different dilation rates, use shared convolution layers to process feature maps of different scales to capture richer multi-scale information. Finally, connect and output all feature maps. Use this module to replace the SPPF module in the Yolov8 model. Compared with the pooling operation of SPPF, the convolution operation has higher flexibility and expression ability in feature extraction, and can better capture the details in the image. Further improve the model's multi-scale feature representation ability and enhance the model's recognition ability for targets in complex scenarios.

[0020] Further, the specific process of step 5 is as follows:

[0021] The performance evaluation metrics specifically include Precision, Recall, mean average precision mAP50, and mean average precision mAP50-95.

[0022] By using performance evaluation metrics, it is convenient to screen out the optimal object detection model for completing the object detection task.

[0023] Furthermore, the expression of Precision is:

[0024] ;

[0025] It refers to the ratio of the positive samples (TP) correctly detected by the classifier to all the samples marked as positive by the classifier (TP + FP).

[0026] The expression of Recall is:

[0027] ;

[0028] It refers to the ratio of the positive samples (TP) correctly detected by the classifier to all the true positive samples (TP + FN).

[0029] The expressions of mAP50 and mAP50-95 are:

[0030] ;

[0031] ;

[0032] In the formula, IoU is the overlapping degree between the predicted bounding box and the true bounding box. AP is the average value of calculating the precision at different recall levels, usually the area under the PR curve. N is the total number of categories, and M is the number of IoU thresholds. mAP50 refers to the mean average precision when the IoU threshold is 0.5, and mAP50-95 calculates the mean average precision with the IoU threshold ranging from 0.5 to 0.95 (usually increasing in steps of 0.05).

[0033] The performance evaluation metrics of the detection model can be calculated through the above formulas.

[0034] By adopting the above technology, compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] By constructing a module combining ConvFormer and CGLU networks to replace the C2f network module in the Yolov8 model framework, the present invention can effectively reduce the calculation parameters of the model while ensuring the detection accuracy of the algorithm; at the same time, by constructing an FSConv network module to replace the SPPF network module in the Yolov8 model framework, it can capture more fine-grained features in the image and improve the detection accuracy of the model when facing complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1It is the method flowchart of the present invention;

[0037] Figure 2 It is the improved Yolov8 network structure diagram in the embodiment of the present invention;

[0038] Figure 3 It is the module diagram of the separable convolution module C2F_CFGLU in the embodiment of the present invention;

[0039] Figure 4 It is the MLP module diagram in the separable convolution module C2F_CFGLU in the embodiment of the present invention;

[0040] Figure 5 It is the FSConv module diagram in the embodiment of the present invention. Detailed implementation manners

[0041] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the specification drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, but not to limit the present invention.

[0042] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope of the present invention defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.

[0043] As Figure 1 shown, a target detection method for the power infrastructure scenario based on Yolov8 specifically includes the following steps:

[0044] Step 1: Shoot the real power infrastructure scenario through a camera, and obtain picture data by frame extraction from the captured video. In order to enhance the effectiveness of the data, reduce the frame extraction interval during normal working hours.

[0045] Step 2: Process the picture data such as screening, annotation, data format conversion, etc. to generate a standard data set that conforms to the Yolo format; divide the standard data set into a training set, a test set and a validation set according to the ratio of 8:1:1 for subsequent experiments.

[0046] Step 3: As Figure 2As shown, based on the ConvFormer and CGLU networks, the C2F_CFGLU network is designed to replace all C2f modules in the Yolov8 model. The backbone network of the improved Yolov8 model includes a convolutional layer 01 (Conv01) and a convolutional layer 02 (Conv02) connected in sequence, a separable convolutional module 01 (C2F_CFGLU01), a convolutional layer 03 (Conv03), a separable convolutional module 02 (C2F_CFGLU02), a convolutional layer 04 (Conv04), a separable convolutional module 03 (C2F_CFGLU03), a convolutional layer 05 (Conv05), and a separable convolutional module 04 (C2F_CFGLU04). The neck network includes an upsampling layer 01 (Upsample01), a connection module 01 (Concat01), a separable convolutional module 05 (C2F_CFGLU05), an upsampling layer 02 (Upsample02), a connection module 02 (Concat02), a separable convolutional module 06 (C2F_CFGLU06), a convolutional layer 06 (Conv06), a connection module 03 (Concat03), a separable convolutional module 07 (C2F_CFGLU07), a convolutional layer 07 (Conv07), a connection module 04 (Concat04), and a separable convolutional module 08 (C2F_CFGLU08).

[0047] Specifically, the separable convolutional module C2F_CFGLU designed in step 3 has the following specific structure Figure 3 As shown, the specific feature extraction process of the separable convolutional module C2F_CFGLU is as follows

[0048] First, the input is normalized through a normalization layer to standardize the feature values. The normalized feature map passes through a SepConv mixer layer to fuse different features. This layer uses separable convolutions to achieve feature mixing, decomposing the convolution operation into depthwise convolution and pointwise convolution, thereby reducing the computational complexity and the number of parameters of the model. The output of the SepConv mixer is then processed through a scaling layer Scale and a regularization layer Dropout to prevent overfitting, and added to the residual path. The output after residual connection is processed through a multi-layer perceptron MLP. Among them, the module diagram of the multi-layer perceptron MLP is as Figure 4 shown. It processes the input through CGLU, which combines convolution and attention mechanisms to enhance the model's expressive power. Convolution can effectively capture local features, while the attention mechanism helps the model focus on the most important parts of the input features. Finally, it passes through another regularization layer Dropout and a scaling layer Scale as the final feature map output. Compared with the original C2f module, the optimized module reduces the number of parameters and the computational amount, and is more easily deployable to hardware devices.

[0049] Step 4: As Figure 2 shown, design a shared convolution module FSConv to replace the SPPF module in the Yolov8 model, that is, connect this module behind the separable convolution module C2F_CFGLU04 in the backbone network.

[0050] Furthermore, for the FSConv module designed in Step 4, the specific structure is as Figure 5 shown, and the specific operation process of the FSConv module is as follows:

[0051] First, pass the input features through a 1x1 convolutional layer to reduce the number of channels and computational amount of the input features. Then define a convolutional kernel and the dilation rates of different dilated convolutions to be used (default is 1, 3, 5). By looping through different dilation rates, use the same convolutional kernel to perform convolutional operations on the input features. Although multiple convolutional operations are performed in this way, all operations use the same convolutional kernel weights, which helps to reduce the number of model parameters while maintaining the unity of convolutional operations, enabling the network to process features of different scales in the same way. Finally, after all convolutional operations, fuse the resulting feature maps and further adjust the number of channels to generate the final output features. Compared with the pooling operation in the original SPPF module, this module can reduce computational resources while providing finer image features for the subsequent detection process, improving the detection accuracy of the model.

[0052] Step 5: Use the prepared standard datasets, including the training set (8000 images), test set (1000 images), and validation set (1000 images), to train the improved Yolov8 model. The experimental software and hardware configurations are shown in Table 1. Use the Pytorch framework, adopt the SGD optimizer, with an initial learning rate of 0.001, a momentum factor of 0.973, 300 iterations, and scale the input image size to 640×640. Based on the above experimental environment, select the optimal object detection model according to the performance evaluation metrics.

[0053] Table 1 Software and Hardware Configuration Table

[0054]

[0055] In Step 5, the performance evaluation metrics specifically include Precision, Recall, mAP50, and mAP50-95.

[0056] With the above settings, it is convenient to screen out the optimal object detection model through the performance evaluation metrics to complete the object detection task.

[0057] Specifically, the expression of Precision is:

[0058] ;

[0059] It refers to the ratio of the positive samples (TP) correctly detected by the classifier to all the samples marked as positive by the classifier (TP + FP).

[0060] The expression of Recall is:

[0061] ;

[0062] It refers to the ratio of the positive samples (TP) correctly detected by the classifier to all the true positive samples (TP + FN).

[0063] The expressions of mAP50 and mAP50-95 are:

[0064] ;

[0065] ;

[0066] In the formula, IoU is the overlapping degree between the predicted bounding box and the true bounding box. AP is the average value of the precision calculated at different recall levels, usually the area under the PR curve. N is the total number of categories, and M is the number of IoU thresholds. mAP50 refers to the mean average precision when the IoU threshold is 0.5, and mAP50-95 calculates the mean average precision with the IoU threshold ranging from 0.5 to 0.95 (usually increasing in steps of 0.05).

[0067] In order to further verify the effects of the improved Yolov8 of the present invention and the existing Yolov8 algorithm, ablation experiments were conducted for comparison between the two algorithms. The experimental results are shown in Table 2:

[0068] Table 2 Comparison table of ablation experiment results

[0069]

[0070] Experiments show that: in the object detection for complex infrastructure scenarios, the improved Yolov8 algorithm of this embodiment has better performance in terms of accuracy Precision, recall Recall, and mAP value compared with the Yolov8 algorithm. In the comparison of the mAP index, the index of Yolov8 is 32.39%, while the algorithm of this embodiment The highest value is 35.07%, which is 2.6 percentage points higher than Yolov8. In the comparison of the Precision value, the Precision value of the algorithm in this embodiment reaches 87.15%, which is 3.3 percentage points higher than the Precision value of Yolov8. Therefore, the comprehensive performance of the algorithm in this embodiment is better.

[0071] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A target detection method for power infrastructure scenes based on Yolov8, characterized in that: The steps include: Step 1: Determine the detection target and collect the required data set: In the power infrastructure scenario, collect the corresponding image data through monitoring video; Step 2: Screen, annotate, and convert the collected image data to obtain a standard data set that conforms to the Yolo format, and divide it into a training set, a test set, and a validation set according to the set ratio; Step 3: Combine ConvFormer with the convolutional gated linear unit network to build a separable convolution module C2f_CFGLU, replacing the C2f network module in the Yolov8 model framework, so as to reduce the calculation parameters of the model while ensuring the accuracy of algorithm detection; the specific process of step 3 is as follows: Based on the ConvFormer and Convolutional Gated Linear Unit network models, on the basis of the ConvFormer network structure, the Convolutional Gated Linear Unit is used to replace the multi-layer perceptron MPL in the ConvFormer network structure to obtain the separable convolution module C2f_CFGLU, and the separable convolution module C2f_CFGLU is used to replace all C2f modules in the Yolov8 model; Step 4, construct a shared convolution module Feature Shared Conv to replace the SPPF network module in the Yolov8 model framework; the shared convolution module Feature Shared Conv uses convolution operation to extract features; Step 5: Use the training set to train the improved Yolov8 model, obtain the optimal target detection model based on the performance evaluation index, detect the images in the test set, and complete the target detection task.

2. According to claim 1, a target detection method for electric power infrastructure scene based on Yolov8 improvement is characterized in that: The specific process of step 1 is as follows: For power infrastructure scenarios, determine the targets that need to be detected; obtain images by extracting frames from surveillance videos. To ensure the validity of the images, reduce the interval between frame extractions.

3. According to claim 2, a target detection method for electric power infrastructure scene based on Yolov8 improvement is characterized in that: The specific process of step 2 is as follows: Screen the collected images and remove images with defects; determine the detection target according to the needs of the application scenario, use the Labelme tool to annotate the image data, and convert the annotated data into the Yolo format required for model training; divide the processed data set into training set, test set, and validation set according to the set ratio.

4. According to claim 3, a target detection method based on Yolov8 improved power infrastructure scene is characterized in that: In step 4, a shared convolution module Feature SharedConv is constructed by constructing a feature pyramid; the specific working process of the shared convolution module Feature SharedConv is as follows: 1) Use a 1x1 convolutional layer to process the input feature map and extract the required features; 2) According to different expansion rates, shared convolutional layers are used to process feature maps of different scales to capture multi-scale information; 3) All feature maps are concatenated and output.

5. According to claim 4, a target detection method based on Yolov8 improved power infrastructure scene is characterized in that: In step 5, The performance evaluation indicators include precision, recall, mAP50 and mAP50-95; the expression of the precision is: It represents the ratio of the positive samples TP correctly detected by the classifier to all the positive samples TP+FP marked by the classifier; the expression of recall rate Recall is: It represents the ratio of the positive sample TP correctly detected by the classifier to all true positive samples TP+FN; The expressions of mAP50 and mAP50-95 are: Where IoU is the overlap between the predicted bounding box and the true bounding box, AP is the average precision calculated at different recall levels, N is the total number of categories, M is the number of IoU thresholds, mAP50 refers to the mean average precision when the IoU threshold is 0.5, and mAP50-95 calculates the mean average precision when the IoU threshold ranges from 0.5 to 0.95.

Citation Information

Patent Citations

  • Low-illumination target detection method based on convolutional neural network

    CN117576540A

  • Power equipment infrared image identification method and device and medium

    CN117746212A