Small target detection method and system for electric power high-altitude operation

By using the small object detection model based on YOLOv8 in power altitude operations, the problem of low detection accuracy of small object is solved, higher detection performance and accuracy are achieved, and missed detection and false detection are reduced.

CN120070843APending Publication Date: 2025-05-30STATE GRID HUNAN ELECTRIC POWER COMPANY LIMITED +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411409314.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In power altitude operations, small targets account for a small proportion of pixels in the image and are easily blocked. The existing models lack the capture of target features in different scenarios, resulting in low detection accuracy of small targets and problems of missed detection.

Method used

A small object detection model based on YOLOv8 is adopted, which includes a feature extraction module, an aggregation and diffusion pyramid module and a task alignment dynamic detection head module. Through multiple feature aggregation and diffusion, the model's positioning and classification capabilities of small objects are enhanced.

Benefits of technology

The model's detection performance and accuracy of small targets is improved, missed and missed detection problems are reduced, and the model's positioning and classification capabilities are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070843A_ABST
    Figure CN120070843A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power high-altitude operation small target detection method comprising the following steps: obtaining electric power high-altitude operation small target historical image data, and obtaining a training data set; based on YOLOv8, constructing an initial small target detection model; using the obtained training data set to train the initial small target detection model to obtain a small target detection model; and carrying out actual high-altitude power operation small target detection by using the small target detection model. The invention also discloses a system for realizing the electric power high-altitude operation small target detection method. According to the method, the model is constructed based on YOLOv8n, the model combines the aggregation diffusion pyramid network and the task alignment dynamic detection head module, the calculation amount and the parameter amount during model operation are reduced, the detection performance of the model is effectively improved, and the positioning capacity and the classification capacity of the model for small targets are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of high-altitude operation of power lines, and specifically relates to a method and system for detecting small targets in high-altitude power operation. Background Art

[0002] High-altitude power operation refers to the operation carried out at high altitudes of power facilities such as transmission lines, substations, and poles. These tasks usually require operators to work at a very high altitude from the ground, which brings great risks to the safety of the operators. During the power production and system maintenance process, power enterprises generally face the problem of operators falling from high altitudes and need to carry out targeted risk prevention and control.

[0003] However, since the size of the operators in the high-altitude power operation scenario is constantly changing, when the operators become smaller in the scenario, due to the small pixel proportion of small targets in the image, there is a situation of being blocked, and the model captures insufficient target features in different scenarios, which is prone to false detection and missed detection. Eventually, there are many challenges and difficulties in the accuracy of small target detection. Summary of the Invention

[0004] One of the purposes of the present invention is to provide a method for detecting small targets in high-altitude power operation, which improves the detection performance and accuracy of the model, and at the same time improves the positioning ability and classification ability of small targets.

[0005] Another purpose of the present invention is to provide a system for implementing the method for detecting small targets in high-altitude power operation.

[0006] The present invention provides a method for detecting small targets in high-altitude power operation, including the following steps:

[0007] S1. Obtain historical image data of high-altitude power operation to obtain a training data set;

[0008] S2. Based on YOLOv8, construct an initial small target detection model; the initial small target detection model includes a feature extraction module, an aggregation and diffusion pyramid module, and a task alignment dynamic detection head module connected in series in sequence;

[0009] The feature extraction module extracts features and then inputs them into the aggregation and diffusion pyramid module;

[0010] The aggregation and diffusion pyramid module performs several times of feature aggregation and feature diffusion on the input features to obtain context semantic information features of different scales, and then inputs them into the task alignment dynamic detection head module;

[0011] The task alignment dynamic detection head module extracts interaction features from the input context semantic information features of different scales, completes target bounding box regression and target category prediction, and outputs target boxes and categories;

[0012] S3. Use the training dataset obtained in step S1 to train the initial small target detection model obtained in step S2 to obtain a small target detection model;

[0013] S4. Use the small target detection model obtained in step S3 to perform actual small target detection for high-altitude power operations.

[0014] The feature extraction module uses the backbone network in YOLOv8n to extract features; the feature extraction module processes the input image to obtain three different-scale features P1, P2, and P3; among them, the output of the second c2f module in the backbone network is P3, the output of the third c2f module in the backbone network is P2, and the output of the SPPF module in the backbone network is P1.

[0015] The aggregation and diffusion pyramid module includes a first feature aggregation module and a second feature aggregation module; the first feature aggregation module uses features P1, P2, and P3 as inputs; the output of the first feature aggregation module is respectively upsampled and downsampled, the upsampled result is concatenated with P1 in channels to obtain P1', the downsampled result is concatenated with P3 in channels to obtain P3', and the obtained results P1' and P3' together with the output P2' of the first feature aggregation module are used as the inputs of the second feature aggregation module;

[0016] The output of the second feature aggregation module is respectively upsampled and downsampled, the upsampled result is concatenated with P1' in channels to obtain P1”, the downsampled result is concatenated with P3' in channels to obtain P3”, and thus three features P1”, P3” and the output P2” of the second feature aggregation module are used as the output of the aggregation and diffusion pyramid module.

[0017] The first feature aggregation module and the second feature aggregation module have the same structure; the feature aggregation module includes a downsampling module, a first 1×1 convolutional layer, an upsampling module, a second 1×1 convolutional layer, a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a third depthwise separable convolutional layer, a fourth depthwise separable convolutional layer, and a third 1×1 convolutional layer;

[0018] The feature aggregation module uses features P1, P2, and P3 as inputs; feature P1 is input into the downsampling module for downsampling to change the feature size to obtain feature P1'; feature P2 is input into the first 1×1 convolutional layer for feature extraction processing to obtain feature P2'; feature P3 is successively processed by the upsampling module and the second 1×1 convolutional layer to obtain feature P3'; then features P1', P2', and P3' are concatenated in channels to obtain feature P; feature P is respectively input into the first to fourth depthwise separable convolutional layers for diffusion processing to obtain four different-scale features P1 , P 2 , P 3 , P 4 , and then for feature P 1 , P 2 , P 3 , P 4 , perform feature fusion with feature P using a residual connection to obtain feature P'; input feature P' into the third 1×1 convolutional layer to adjust the number of channels, and then fuse the result with feature P to obtain feature P" as the output of the feature aggregation module;

[0019] The task alignment dynamic detection head module includes a feature extractor module, a localization branch, and a classification branch; the feature extractor module extracts interaction features from the input data; the localization branch and the classification branch use the interaction features output by the feature extractor module as input; the outputs of the localization branch and the classification branch are used as the outputs of the model;

[0020] The feature extractor module includes a first 3×3 convolutional layer and a second 3×3 convolutional layer connected in series; concatenate the outputs of the first 3×3 convolutional layer and the second 3×3 convolutional layer in channels to obtain interaction features as the output of the feature extraction module, which is represented by the following formula:

[0021] X inter = Concat[Conv 3×3 (Conv 3×3 (X P )), Conv 3×3 (X P )]

[0022] where X inter is the interaction feature output by the feature extractor module; Conv 3×3 is a 3×3 convolutional layer; X P is the feature output by the aggregation diffusion pyramid network; Cconcat is channel concatenation;

[0023] The localization branch uses deformable convolution to process the input interaction features to obtain the mask and offset of the deformable convolution, and the output of the classification branch is used as the category of the target, which is represented by the following formula:

[0024] X Reg = Concat(Dcnv2(TAP(X inter )), Mask&Offset(X inter ))

[0025] where X Reg is the output of the localization branch; Dcnv2 is deformable convolution; TAP is layer attention mechanism; Mask&Offset is the mask and offset of Dcnv2;

[0026] The classification branch includes a layer attention mechanism, a first 1×1 convolutional layer, a third 3×3 convolutional layer, and a second 1×1 convolutional layer; the layer attention mechanism uses the interaction features as the input; after the interaction features are sequentially processed by the first 1×1 convolutional layer, the third 3×3 convolutional layer, and the second 1×1 convolutional layer, they are concatenated with the output of the layer attention mechanism in channels, and the obtained result is used as the result of the classification branch. The output of the localization branch is the target bounding box of the target, which is represented by the following formula:

[0027] X Cls =Concat(TAP(X inter ),Conv 1×1 (Conv 3×3 (Conv 1×1 (X inter ))))

[0028] Wherein, X Cls is the output of the classification branch; Conv 1×1 is the processing of the 1×1 convolutional layer.

[0029] The present invention also provides a system for implementing the above-mentioned small target detection method for high-altitude electric power operation, including a data acquisition module, a model construction module, a model training module, and a target detection module connected in series in sequence;

[0030] The data acquisition module acquires the historical image data of small targets for high-altitude electric power operation, obtains the training data set, and uploads the data to the model construction module;

[0031] The model construction module constructs an initial small target detection model based on YOLOv8n according to the received data, and uploads the data to the model training module;

[0032] The model training module trains the initial small target detection model using the training data set according to the received data, obtains the small target detection model, and uploads the data to the target detection module;

[0033] The target detection module performs actual small target detection for high-altitude electric power operation according to the received data.

[0034] The present invention discloses a small target detection method and system for high-altitude electric power operation. A model is constructed based on YOLOv8n. The model combines an aggregation diffusion pyramid network and a task alignment dynamic detection head module, which not only reduces the computational amount and the number of parameters during the operation of the model, but also effectively improves the detection performance of the model, and improves the localization ability and classification ability of the model for small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic flow chart of the method of the present invention;

[0036] Figure 2 It is a schematic structural diagram of the system of the present invention;

[0037] Figure 3 It is a comparison chart of the small target recognition results of the method of the present invention and the mainstream model under different scenarios in the embodiment, where Figure 3 a is the night scene; Figure 3 b is the day scene; Figure 3 c is the dense scene;

[0038] Figure 4 It is an actual application diagram of the small target recognition of the method of the present invention for high-altitude power operation in the embodiment. Specific implementation manners

[0039] The present invention provides a method for detecting small targets in high-altitude power operation, and its process schematic diagram is as Figure 1 shown, including the following steps:

[0040] S1. Obtain historical image data of high-altitude power operation to obtain a training data set;

[0041] S2. Based on YOLOv8, construct an initial small target detection model; the initial small target detection model includes a feature extraction module, an aggregation and diffusion pyramid module, and a task alignment dynamic detection head module connected in series in sequence;

[0042] The feature extraction module extracts features and then inputs them into the aggregation and diffusion pyramid module;

[0043] The aggregation and diffusion pyramid module performs several times of feature aggregation and feature diffusion on the input features to obtain context semantic information features of different scales, and then inputs them into the task alignment dynamic detection head module;

[0044] The task alignment dynamic detection head module extracts interaction features from the context semantic information features of different scales of the input, completes target bounding box regression and target category prediction, and outputs target boxes and categories;

[0045] The feature extraction module uses the backbone network backbone in YOLOv8n network to extract features; the feature extraction module processes the input image to obtain three different scales of features P1, P2, and P3; among them, the output of the second c2f module in the backbone network backbone is P3, the output of the third c2f module in the backbone network backbone is P2, and the output of the SPPF module in the backbone network backbone is P1.

[0046] The aggregation and diffusion pyramid module includes a first feature aggregation module and a second feature aggregation module; the first feature aggregation module uses features P1, P2, and P3 as inputs; the output of the first feature aggregation module is respectively upsampled and downsampled. The result of upsampling is concatenated with P1 along the channel dimension to obtain P1', and the result of downsampling is concatenated with P3 along the channel dimension to obtain P3'. The obtained results P1' and P3', together with the output P2' of the first feature aggregation module, are used as the inputs of the second feature aggregation module;

[0047] The output of the second feature aggregation module is respectively upsampled and downsampled. The result of upsampling is concatenated with P1' along the channel dimension to obtain P1'', and the result of downsampling is concatenated with P3' along the channel dimension to obtain P3''. In this way, P1'', P3'', and the output P2'' of the second feature aggregation module are used as the outputs of the aggregation and diffusion pyramid module.

[0048] The first feature aggregation module has the same structure as the second feature aggregation module; the feature aggregation module includes a downsampling module, a first 1×1 convolutional layer, an upsampling module, a second 1×1 convolutional layer, a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a third depthwise separable convolutional layer, a fourth depthwise separable convolutional layer, and a third 1×1 convolutional layer;

[0049] The feature aggregation module uses features P1, P2, and P3 as inputs; feature P1 is input into the downsampling module for downsampling to change the feature size and obtain feature P1'; feature P2 is input into the first 1×1 convolutional layer for feature extraction processing to obtain feature P2'; feature P3 passes through the upsampling module and the second 1×1 convolutional layer in sequence to obtain feature P3'; then features P1', P2', and P3' are concatenated along the channel dimension to obtain feature P; feature P is respectively input into the first to fourth depthwise separable convolutional layers for diffusion processing to obtain four different-scale features P 1 、P 2 、P 3 、P 4 , and then for feature P 1 、P 2 、P 3 、P 4 and feature P, residual connection is used for feature fusion to obtain feature P'; feature P' is input into the third 1×1 convolutional layer to adjust the number of channels, and then the result is fused with feature P to obtain feature P'' as the output of the feature aggregation module;

[0050] The task alignment dynamic detection head module includes a feature extractor module, a localization branch, and a classification branch; the feature extractor module extracts interaction features from the input data; the localization branch and the classification branch use the interaction features output by the feature extractor module as inputs; the outputs of the localization branch and the classification branch are used as the outputs of the model;

[0051] The feature extractor module includes a first 3×3 convolutional layer and a second 3×3 convolutional layer connected in series; the outputs of the first 3×3 convolutional layer and the second 3×3 convolutional layer are concatenated in channels to obtain the interaction feature as the output of the feature extraction module, which is represented by the following formula:

[0052] X inter =Concat[Conv 3×3 (Conv 3×3 (X P )),Conv 3×3 (X P )]

[0053] Where X inter is the interaction feature output by the feature extractor module; Conv 3×3 is a 3×3 convolutional layer; X P is the feature output by the aggregation diffusion pyramid network; Cconcat is channel concatenation;

[0054] The localization branch uses deformable convolution to process the input interaction feature to obtain the mask and offset of the deformable convolution. The output of the localization branch is the target bounding box, which is represented by the following formula:

[0055] X Reg =Concat(Dcnv2(TAP(X inter )),Mask&Offset(X inter ))

[0056] Where X Reg is the output of the localization branch; Dcnv2 is deformable convolution; TAP is the layer attention mechanism; Mask&Offset is the mask and offset of Dcnv2;

[0057] The classification branch includes a layer attention mechanism, a first 1×1 convolutional layer, a third 3×3 convolutional layer, and a second 1×1 convolutional layer; the layer attention mechanism uses the interaction feature as the input; after the interaction feature is processed by the first 1×1 convolutional layer, the third 3×3 convolutional layer, and the second 1×1 convolutional layer in sequence, it is concatenated with the output of the layer attention mechanism in channels, and the result obtained is used as the result of the classification branch. The output of the classification branch is used as the category of the target, which is represented by the following formula:

[0058] X Cls =Concat(TAP(X inter ),Conv 1×1 (Conv 3×3 (Conv 1×1 (X inter ))))

[0059] Where XCls Output for classification branch; Conv 1×1 Processed by a 1×1 convolutional layer.

[0060] S3. Use the training dataset obtained in step S1 to train the initial small target detection model obtained in step S2 to obtain a small target detection model;

[0061] S4. Use the small target detection model obtained in step S3 to perform actual small target detection for high-altitude power operations.

[0062] The present invention also provides a system for implementing the method for detecting small targets in high-altitude power operations, including a data acquisition module, a model construction module, a model training module, and a target detection module connected in series in sequence;

[0063] The data acquisition module acquires historical image data of small targets in high-altitude power operations to obtain a training dataset and uploads the data to the model construction module;

[0064] The model construction module constructs an initial small target detection model based on YOLOv8n according to the received data and uploads the data to the model training module;

[0065] The model training module trains the initial small target detection model using the training dataset according to the received data to obtain a small target detection model and uploads the data to the target detection module;

[0066] The target detection module performs actual small target detection for high-altitude power operations according to the received data.

[0067] The following describes the content of the present invention in conjunction with an embodiment:

[0068] For different scenarios, small target detection is performed using the mainstream model and the method of the present invention, and the obtained results are as Figure 3 shown.

[0069] Compared with YOLOv5s, YOLOv5m, and YOLOv7-tiny, the model of the method of the present invention has the smallest computational amount and the highest accuracy, and both the number of parameters and the weights have decreased. This is due to the model adopting an aggregation diffusion pyramid network and a task-aligned dynamic detection head module, which fuses cross-scale features multiple times, greatly enhancing the richness and effectiveness of feature semantics. It not only reduces the computational amount and the number of parameters during model operation, but also effectively improves the detection performance of the model, reduces the problems of missed detection and false detection, and improves the localization ability and classification ability of the model for small targets.

[0070] The actual application diagram of the method of the present invention for identifying small targets in high-altitude power is as Figure 4 shown.

Claims

1. A method for detecting small targets in power aerial work, characterized in that: The following steps are involved: S1. Obtain historical image data of power aerial work to obtain a training data set; S2. Based on YOLOv8, an initial small target detection model is constructed; the initial small target detection model includes a feature extraction module, an aggregate diffusion pyramid module, and a task alignment dynamic detection head module connected in series; The feature extraction module extracts features and then inputs them into the aggregate diffusion pyramid module; The aggregation diffusion pyramid module aggregates and diffuses the input features several times to obtain contextual semantic information features of different scales, which are then input into the task alignment dynamic detection head module. The task-aligned dynamic detection head module extracts interactive features from the contextual semantic information features of different scales of the input, completes the target bounding box regression and target category prediction, and outputs the target box and category; S3. Using the training data set obtained in step S1, the initial small target detection model obtained in step S2 is trained to obtain a small target detection model; S4. Use the small target detection model obtained in step S3 to perform actual small target detection in high-altitude power operations.

2. The method for detecting small targets during power aerial work according to claim 1, characterized in that: The feature extraction module uses the backbone network backbone in the YOLOv8n network for feature extraction; The extraction module processes the input image to obtain three features of different scales, P1, P2, and P3; among them, the output of the second c2f module in the backbone network is P3, the output of the third c2f module in the backbone network is P2, and the output of the SPPF module in the backbone network is P1.

3. The method for detecting small targets during power aerial work according to claim 2, characterized in that: The aggregate diffusion pyramid module includes a first feature aggregation module and a second feature aggregation module; The first feature aggregation module uses features P1, P2, and P3 as inputs; the output of the first feature aggregation module is upsampled and downsampled respectively, the upsampled result is channel-joined with P1 to obtain P1', the downsampled result is channel-joined with P3 to obtain P3', and the obtained results P1' and P3' and the output P2' of the first feature aggregation module are used as the input of the second feature aggregation module; The output of the second feature aggregation module is upsampled and downsampled respectively, and the upsampled result is channel-joined with P1' to obtain P1", and the downsampled result is channel-joined with P3' to obtain P3", so that the three features P1", P3" and the output of the second feature aggregation module P2" are obtained as the output of the aggregated diffusion pyramid module.

4. The method for detecting small targets during high-altitude power operations according to claim 4 is characterized in that: The first feature aggregation module has the same structure as the second feature aggregation module; the feature aggregation module includes a downsampling module, a first 1×1 convolutional layer, an upsampling module, a second 1×1 convolutional layer, a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a third depthwise separable convolutional layer, a fourth depthwise separable convolutional layer, and a third 1×1 convolutional layer; The feature aggregation module uses features P1, P2, and P3 as input; feature P1 is input to the downsampling module for downsampling to change the feature size to obtain feature P1'; feature P2 is input to the first 1×1 convolutional layer for feature extraction processing to obtain feature P2'; Feature P3 is processed by the upsampling module and the second 1×1 convolutional layer in turn to obtain feature P3'; Then, features P1', P2', and P3' are channel-joined to obtain feature P; The feature P is input into the first to fourth depth-separable convolutional layers for diffusion processing, and four different scale feature Ps are obtained. 1 , P 2 , P 3 , P 4 , and then for feature P 1 , P 2 , P 3 , P 4 The feature is fused with feature P using residual connection to obtain feature P'; The feature P' is input into the third 1×1 convolutional layer to adjust the number of channels, and then the result is fused with the feature P to obtain the feature P" as the output of the feature aggregation module.

5. The method for detecting small targets during power aerial work according to claim 2, characterized in that: The task-aligned dynamic detection head module includes a feature extractor module, a positioning branch and a classification branch; the feature extractor module extracts interactive features from the input data; the positioning branch and the classification branch use the interactive features output by the feature extractor module as input; and the outputs of the positioning branch and the classification branch serve as the output of the model.

6. The method for detecting small targets during power aerial work according to claim 6, characterized in that: The feature extractor module includes a first 3×3 convolutional layer and a second 3×3 convolutional layer connected in series. The output of the first 3×3 convolutional layer and the output of the second 3×3 convolutional layer are concatenated to obtain the interactive feature as the output of the feature extraction module, which is expressed by the following formula: X inter =Concat[Conv 3×3 (Conv 3×3 (X P )),Conv 3×3 (X P )] Among them, X inter Interaction features output by the feature extractor module; Conv 3×3 is a 3×3 convolutional layer; X P is the feature output by the aggregated diffusion pyramid network; Cconcat is channel concatenation; The localization branch uses deformable convolution to process the input interactive features to obtain the mask and offset of the deformable convolution. The output of the localization branch is the target box of the target, which is expressed by the following formula: X Reg =Concat(Dcnv2(TAP(X inter )),Mask&Offset(X inter )) Among them, X Reg is the output of the positioning branch; Dcnv2 is deformable convolution; TAP is the layer attention mechanism; Mask&Offset is the mask and offset of Dcnv2; the output of the positioning branch is the target box of the target; The classification branch includes the layer attention mechanism, the first 1×1 convolution layer, the third 3×3 convolution layer and the second 1×1 convolution layer; the layer attention mechanism uses the interaction feature as input; the interaction feature is processed by the first 1×1 convolution layer, the third 3×3 convolution layer and the second 1×1 convolution layer in sequence, and then the channel is spliced ​​with the output of the layer attention mechanism. The result is used as the result of the classification branch. The output of the classification branch is used as the category of the target, which is expressed by the following formula: X Cls =Concat(TAP(X inter ),Conv 1×1 (Conv 3×3 (Conv 1×1 (X inter )))) Among them, X Cls is the output of the classification branch; Conv 1×1 It is processed by 1×1 convolution layer.

7. A system for implementing the method for detecting small targets in power aerial work according to any one of claims 1 to 7, characterized in that: It includes a data acquisition module, a model building module, a model training module and a target detection module connected in series in sequence; The data acquisition module acquires historical image data of small targets of power aerial work, obtains a training data set, and uploads the data to the model construction module; The model building module builds an initial small target detection model based on YOLOv8n according to the received data, and uploads the data to the model training module; The model training module uses the training data set to train the initial small target detection model based on the received data, obtains the small target detection model, and uploads the data to the target detection module; The target detection module performs actual small target detection for high-altitude power operations based on the received data.