Accumulated water detection method for power transmission line and related equipment
By embedding position information and focusing on the water accumulation image, combined with YOLOv8 model training, the problem of inaccurate detection in the water accumulation detection of transmission lines is solved, and a more efficient water accumulation detection effect is achieved.
Patent Information
- Application Number
- CN202510491958.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-22
AI Technical Summary
The existing models are inaccurate in the detection of water accumulation in transmission lines, making it difficult to effectively pay attention to key areas in the image, resulting in the water accumulation characteristics being ignored.
By performing position information embedding operations and position attention generation operations on the water accumulation image, an enhanced feature map is obtained and trained using the YOLOv8 model to improve the sensitivity and accuracy of the model to water accumulation characteristics.
It improves the accuracy of water accumulation detection in transmission lines, enhances the model's ability to identify water accumulation areas, and outputs more accurate water accumulation detection results.
Smart Images

Figure CN120355884A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of water accumulation detection, and particularly to a method and related equipment for detecting water accumulation in a transmission line. Background Art
[0002] In the field of detecting water accumulation in transmission lines, although a variety of models and technologies have been proposed, the problem of inaccurate detection still exists. For example, existing models have deficiencies in focusing on key regions in images, which can lead to the neglect of water accumulation features and thus affect the accuracy of detection. Summary of the Invention
[0003] In view of this, the present invention provides a method and related equipment for detecting water accumulation in a transmission line.
[0004] The specific technical solution of the first embodiment of the present invention is: a method for detecting water accumulation in a transmission line, the method comprising: acquiring water accumulation images of a plurality of water accumulation regions; performing water accumulation region annotation on each of the water accumulation images to obtain the annotated water accumulation images; performing position information embedding operation and position attention generation operation on the annotated water accumulation images to obtain an enhanced feature map; training a preset YOLOv8 model using all the enhanced feature maps to obtain a YOLOv8 object detection model; acquiring a real-time image of the transmission line to be detected; inputting the real-time image of the transmission line to be detected into the YOLOv8 object detection model for water accumulation detection, and outputting the water accumulation detection result of the transmission line to be detected; the water accumulation detection result includes the water accumulation position, water accumulation range, and water accumulation confidence.
[0005] Preferably, the performing position information embedding operation and position attention generation operation on the water accumulation image to obtain an enhanced feature map includes: performing position information embedding operation on the water accumulation image along the width and height directions to obtain a first feature map in the width direction and a second feature map in the height direction; performing position attention generation on the first feature map and the second feature map to obtain a first attention weight of the first feature map and a second attention weight of the second feature map; performing a weighting operation on the water accumulation image based on the first attention weight and the second attention weight to obtain the enhanced feature map.
[0006] Preferably, the first feature map and the second feature map are obtained by the following formula:
[0007]
[0008] where is the first feature map, is the second feature map, W is the width of the water accumulation image, x c(h, i) is the eigenvalue of the ponding image on channel c, i is the index for traversing the width during summation, H is the height of the ponding image, x c (j, w) is the eigenvalue of the ponding image at positions (h, i) and (j, w).
[0009] Preferably, generating position attention for the first feature map and the second feature map to obtain the first attention weight of the first feature map and the second attention weight of the second feature map includes: performing a splicing operation on the first feature map and the second feature map to obtain a spliced feature map; performing a dimensionality reduction operation and a normalization operation on the spliced feature map to obtain a normalized feature map; performing a non-linear mapping on the normalized feature map to obtain a non-linear feature map; performing a convolution calculation on the non-linear feature map to obtain the height and width of the non-linear feature map; obtaining the first attention weight according to the Sigmoid activation function and the height of the non-linear feature map; obtaining the second attention weight according to the Sigmoid activation function and the width of the non-linear feature map.
[0010] Preferably, the non-linear feature map is obtained by the following formula:
[0011]
[0012] where f is the non-linear feature map, N is the normalized feature map, α is a preset modulation factor, β is a preset coupling coefficient, γ is a preset linear transformation weight, δ is a preset offset compensation amount, η is a preset gain coefficient, λ is a preset attenuation factor, and θ is a preset high-order non-linear exponent.
[0013] Preferably, the enhanced feature map is obtained by the following formula:
[0014]
[0015] where y c (i, j) is the enhanced feature map, x c (i, j) is the ponding image, is the first attention weight, is the second attention weight.
[0016] Preferably, after obtaining the ponding images of multiple ponding areas, it further includes: preprocessing each of the ponding images to obtain a preprocessed ponding image; the preprocessing includes image cropping, image scaling, and image normalization; then the ponding area annotation of each ponding image to obtain the annotated ponding image is: performing ponding area annotation on each preprocessed ponding image to obtain the annotated ponding image.
[0017] The specific technical solution of the second embodiment of the present invention is as follows: A water accumulation detection system for a transmission line, the system includes: a first image acquisition module, a labeling module, a feature enhancement module, a training module, a second image acquisition module, and a water accumulation detection result output module; the first image acquisition module is used to acquire water accumulation images of multiple water accumulation areas; the labeling module is used to label the water accumulation areas of each of the water accumulation images to obtain the labeled water accumulation images; the feature enhancement module is used to perform position information embedding operation and position attention generation operation on the labeled water accumulation images to obtain an enhanced feature map; the training module is used to train a preset YOLOv8 model with all the enhanced feature maps to obtain a YOLOv8 object detection model; the second image acquisition module is used to acquire a real-time image of the transmission line to be detected; the water accumulation detection result output module is used to input the real-time image of the transmission line to be detected into the YOLOv8 object detection model for water accumulation detection and output the water accumulation detection result of the transmission line to be detected; the water accumulation detection result includes the water accumulation position, the water accumulation range, and the water accumulation confidence level.
[0018] The specific technical solution of the third embodiment of the present invention is as follows: A water accumulation detection device for a transmission line, including a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.
[0019] The specific technical solution of the fourth embodiment of the present invention is as follows: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.
[0020] Implementing the embodiments of the present invention will have the following beneficial effects:
[0021] In the present invention, the marked ponding image is subjected to position information embedding operation and position attention generation operation to obtain an enhanced feature map. Among them, the position information embedding operation can perceive the spatial distribution characteristics of the target, such as the significant differences in ponding depth and shape in different regions. Therefore, position information embedding can strengthen the sensitivity of the model to spatial distribution. The position attention generation operation enables the model to focus on the areas with severe ponding, enabling the model to more accurately capture the spatial context information of the target, thereby enhancing the saliency of ponding features in the enhanced feature map. The preset YOLOv8 model is trained using the enhanced feature map to obtain a YOLOv8 object detection model, enabling the YOLOv8 object detection model to more effectively focus on the ponding features in the ponding image. Therefore, when using the YOLOv8 object detection model to perform ponding detection on the real-time image of the power transmission line to be detected, a more accurate ponding detection result can be obtained. Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0023] Figure 1 It is a flowchart of the steps of the ponding detection method for the power transmission line;
[0024] Figure 2a It is an image of the ponding scene in the dataset;
[0025] Figure 2b It is an image of the non-ponding scene in the dataset;
[0026] Figure 3a It is an effect diagram of the image annotation of the ponding scene;
[0027] Figure 3b It is an effect diagram of the image annotation of the non-ponding scene;
[0028] Figure 4a It is an effect diagram of the conventional ponding detection;
[0029] Figure 4b It is an effect diagram of the ponding detection obtained by using the method in the present application;
[0030] Figure 5 It is a schematic structural diagram of the improved YOLOv8 model;
[0031] Figure 6 It is a schematic structural diagram of the CoordAtt attention mechanism;
[0032] Figure 7a Schematic diagram of the convolution method for ordinary convolution;
[0033] Figure 7b Schematic diagram of the convolution method for Ghost convolution;
[0034] Figure 8a Schematic diagram of the structure of GhostConv;
[0035] Figure 8b Schematic diagram of the structure of C2fGhost;
[0036] Figure 9 Schematic diagram of the BoT3 module;
[0037] Figure 10 Schematic diagram of the BoT module;
[0038] Figure 11 Schematic diagram of the structure of the water accumulation detection system for the transmission line;
[0039] Among them, 201, the first image acquisition module; 202, the annotation module; 203, the feature enhancement module; 204, the training module; 205, the second image acquisition module; 206, the water accumulation detection result output module. Detailed implementation manners
[0040] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0041] The terms "first", "second", etc. in the specification and claims of the present application and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other steps or modules inherent to these processes, methods, products, or devices.
[0042] Referring to "embodiment" in this article means that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0043] Please refer to Figure 1 , which is the flowchart of the steps of a method for detecting water accumulation in a transmission line in the first embodiment of this application, in order to obtain a more accurate water accumulation detection result. The method includes:
[0044] Step 101, obtain water accumulation images of multiple water accumulation areas;
[0045] Specifically, in this embodiment, a dataset is constructed using real power grid water accumulation images, which are mainly collected by monitoring cameras along the power grid to ensure the authenticity of the images and keep up with the current situation. In addition, in this embodiment, drones are used to perform inspection tasks on power grid facilities, taking high-definition images and videos to accurately capture the subtle signs of water accumulation beside the power grid. To further improve the dataset, in this embodiment, a power grid water accumulation scenario is simulated under controlled conditions, and images and videos covering different background conditions are collected. At the same time, existing public water accumulation datasets are integrated to increase the diversity of the data. The dataset used contains 300 real images of water accumulation scenarios and non-water accumulation scenarios, and image examples are as shown in Figure 2a and Figure 2b . When dividing the dataset, the ratio of the training set to the test set of the model is set to 13:2;
[0046] Step 102, perform water accumulation area annotation on each of the water accumulation images to obtain the annotated water accumulation images;
[0047] Specifically, for the non-water accumulation image after annotation, please refer to Figure 3a as shown, and for the water accumulation image after annotation, please refer to Figure 3b as shown. Annotate the water accumulation areas in the water accumulation images, and additional annotation information can also be added to the water accumulation images. Save the annotated images and annotation information as files in a format compatible with YOLOv8. This file contains the image path, class label, and bounding box coordinates. After the annotation work is completed, immediately conduct a quality check to ensure that all water accumulation areas are correctly annotated without omission or inaccurate bounding boxes with incorrect annotation. Through the detailed annotation process, a high-quality dataset can be created, which not only contains labeled water accumulation images, but also each water accumulation is accurately annotated, enabling the YOLOv8 model to better learn and understand the characteristics and patterns of water accumulation, so as to achieve efficient and accurate water accumulation detection in practical applications;
[0048] Step 103, perform position information embedding operation and position attention generation operation on the annotated water accumulation images to obtain an enhanced feature map;
[0049] Specifically, the position information embedding operation and the position attention generation operation enhance or weaken the response to features in different regions according to the importance degree, improve the ability of the model to focus on key information in the detection task, and at the same time enhance the feature extraction ability of the model, ultimately achieving the improvement of the model detection efficiency and accuracy;
[0050] Step 104: Train a preset YOLOv8 model using all the enhanced feature maps to obtain a YOLOv8 object detection model;
[0051] Specifically, 300 water accumulation images of transmission line objects are selected and used for object detection training with the YOLOv8 model, aiming to perform water accumulation detection tasks. In the training stage, the Pytorch deep learning framework is adopted. Regarding the training settings, the batch training size is 16, the number of training iterations is 100, and the initial learning rate is 0.01. When the training reaches the 30th, 60th, and 90th iterations, the learning rate is decayed at a rate of 0.1;
[0052] Step 105: Obtain a real-time image of the transmission line to be detected;
[0053] Step 106: Input the real-time image of the transmission line to be detected into the YOLOv8 object detection model for water accumulation detection, and output the water accumulation detection result of the transmission line to be detected; the water accumulation detection result includes the water accumulation position, the water accumulation range, and the water accumulation confidence level.
[0054] Specifically, for the results of water accumulation detection using existing methods, please refer to Figure 4a ; for the results of water accumulation detection using the method in this application, please refer to Figure 4b . Among them Figure 4a , the water accumulation is marked by a rectangular box, and "water0.43" is shown beside the box, which may indicate that the confidence score of the model for water accumulation is 43%, meaning that the model has 43% confidence that the detected target is water accumulation, Figure 4b ; in
[0055] In a specific embodiment, after obtaining the waterlogging images of multiple waterlogging areas, the following steps are further included: preprocessing each of the waterlogging images to obtain the preprocessed waterlogging images; the preprocessing includes image cropping, image scaling, and image normalization; then, the step of performing waterlogging area annotation on each of the waterlogging images to obtain the annotated waterlogging images is: performing waterlogging area annotation on each of the preprocessed waterlogging images to obtain the annotated waterlogging images.
[0056] Specifically, in the image preprocessing stage, the size of the images is uniformly adjusted to 640×640, and at the same time, image scaling and image normalization are also performed to meet the input requirements of the YOLOv8 model. Subsequently, the cross-entropy loss function is adopted, and combined with the Adam optimizer, the YOLOv8 model is trained for the object detection task.
[0057] In a specific embodiment, when training the preset YOLOv8 model using all the enhanced feature maps to obtain the YOLOv8 object detection model, the performance of the YOLOv8 object detection model is evaluated by parameters such as precision, recall, mean average precision, and mAP@0.5:0.95. When all the performances reach the expected values, the training of the YOLOv8 object detection model is ended; otherwise, a new training set is obtained to perform iterative training on the YOLOv8 object detection model.
[0058] Specifically, precision refers to the proportion of the instances predicted as positive samples by the model that are actually positive samples, which is used to measure the accuracy of the model's prediction. That is, among all the instances predicted as positive samples, how many are truly positive samples. The formula for calculating precision Precision is:
[0059]
[0060] where TP is the number of true positives and FP is the number of false positives.
[0061] Recall refers to the proportion of the instances that are actually positive samples and are correctly predicted as positive samples by the model, which is used to measure the coverage ability of the model. That is, among all the actually positive samples, how many are captured by the model. The formula for calculating recall Recall is:
[0062]
[0063] where FN represents the number of false negatives.
[0064] The mean average precision (mAP) is the average of the average precisions for all classes. The average precision is calculated for a single class and is used to measure the balance between the precision and recall of the model for that class. The calculation of the mAP involves integrating the precision and recall at different thresholds to obtain an area under the curve value, and then taking the average of the average precisions for all classes.
[0065] mAP@0.5:0.95 is a variant of the average precision, which calculates the average precision in the range of IoU (Intersection over Union) thresholds from 0.5 to 0.95 (step size 0.05). This calculation method provides a more stringent evaluation criterion by considering the performance of the model at different IoU thresholds, and thus can more comprehensively reflect the detection ability of the model.
[0066] Precision: The precision of the model in this application is 0.888. This means that among all the samples predicted as waterlogging, 88.8% are actual waterlogging events, indicating that the model performs best in reducing false positives (i.e., wrongly predicting non-waterlogging events as waterlogging).
[0067] Recall: The recall of the model in this application is 0.776, which is an improvement compared to the YOLOv8 model. This indicates that the model can identify more actual waterlogging events, i.e., there are fewer missed detections.
[0068] Mean average precision: The mean average precision of the model in this application is 0.879. Although it is slightly lower than the CA+BoT3 model, it is still higher than the YOLOv8 model. This indicates that the model has a relatively high average detection precision in different waterlogging scenarios and can more stably identify waterlogging.
[0069] mAP@0.5:0.95: The average precision under more stringent IoU thresholds is 0.664, which is an improvement compared to the existing YOLOv8 model. This indicates that the model has a relatively balanced performance at different IoU thresholds and can maintain a certain recognition ability under different detection difficulties.
[0070] In a specific embodiment, the operation of embedding position information and generating position attention for the waterlogging image to obtain an enhanced feature map includes: performing the operation of embedding position information on the waterlogging image along the width and height directions to obtain a first feature map in the width direction and a second feature map in the height direction; generating position attention for the first feature map and the second feature map to obtain a first attention weight for the first feature map and a second attention weight for the second feature map; performing a weighting operation on the waterlogging image based on the first attention weight and the second attention weight to obtain the enhanced feature map.
[0071] Specifically, the ponding water image is divided into a feature map along the width and a feature map along the height. Then, two global average poolings are used to aggregate the two feature maps respectively to obtain a first feature map in the width direction and a second feature map in the height direction. The feature maps in the width and height directions are concatenated, and then convolution calculations are performed on the feature maps. Their dimensions are reduced to c / r of the original, and then BatchNorm operations are performed to obtain a normalized feature map F1. Through the Sigmoid activation function, a feature map f of the form 1×(W + H)×C / r is obtained. Then, the size of the feature map f remains unchanged, and subsequent convolution calculations are performed with a convolution kernel size of 1×1 to obtain the first feature map and the second feature map. After passing through the Sigmoid activation function, the first attention weight g h and the second attention weight g w are obtained respectively. Finally, based on the first attention weight and the second attention weight, a weighted operation is performed on the ponding water image to obtain an enhanced feature map.
[0072] In a specific embodiment, the first feature map and the second feature map are obtained by the following formula:
[0073]
[0074] where, is the first feature map, is the second feature map, W is the width of the ponding water image, x c (h,i) is the feature value of the ponding water image on channel c, i is the index for traversing the width during summation, H is the height of the ponding water image, and x c (j,w) is the feature value of the ponding water image at positions (h,i) and (j,w).
[0075] In a specific embodiment, the generating of position attention for the first feature map and the second feature map to obtain the first attention weight of the first feature map and the second attention weight of the second feature map includes: performing a concatenation operation on the first feature map and the second feature map to obtain a concatenated feature map; performing a dimensionality reduction operation and a normalization operation on the concatenated feature map to obtain a normalized feature map; performing a non-linear mapping on the normalized feature map to obtain a non-linear feature map; performing convolution calculations on the non-linear feature map to obtain the height of the non-linear feature map and the width of the non-linear feature map; obtaining the first attention weight according to the Sigmoid activation function and the height of the non-linear feature map; and obtaining the second attention weight according to the Sigmoid activation function and the width of the non-linear feature map.
[0076] Specifically, splicing the first feature map and the second feature map integrates information from different sources or levels, enhancing the richness of features. Reducing redundant information through dimensionality reduction and normalizing to ensure stable feature distribution provides a more stable input for subsequent processing. Introducing non-linear activation enhances the feature expression ability, enabling the model to capture more complex patterns. Through convolutional operations, local features are further extracted while calculating the height and width of the feature map, providing spatial dimension information for subsequent generation of attention weights. Based on the Sigmoid activation function and the generation of the feature map height, the model can dynamically adjust its attention to different height regions.
[0077] In a specific embodiment, the non-linear feature map is obtained using the following formula:
[0078]
[0079] where f is the non-linear feature map, N is the normalized feature map, α is a preset modulation factor, β is a preset coupling coefficient, γ is a preset linear transformation weight, δ is a preset offset compensation amount, η is a preset gain coefficient, λ is a preset attenuation factor, and θ is a preset high-order non-linear exponent.
[0080] In a specific embodiment, the non-linear feature map f can also be obtained using the following formula: f = δ(F1([z h ,z w )), where f is the non-linear feature map, δ is the Sigmoid activation function, F1 is the normalized feature map, F h is the first feature map, and F w is the second feature map.
[0081] In a specific embodiment, the first attention weight g h and the second attention weight g w are obtained using the following formula: g h = σ(F h (f h )), g w = σ(F w (f w )), where σ is the Sigmoid activation function, F h (f h ) is the height of the non-linear feature map, and F w (f w ) is the width of the non-linear feature map.
[0082] In a specific embodiment, the enhanced feature map is obtained using the following formula:
[0083]
[0084] where yc (i,j) is the enhanced feature map, x c (i,j) is the waterlogging image, is the first attention weight, is the second attention weight.
[0085] In a specific embodiment, the preset YOLOv8 model in the present application is an improved YOLOv8 model. For the structural schematic diagram of the improved YOLOv8 model, please refer to Figure 5 . The improved YOLOv8 model provides an efficient and accurate solution for power grid waterlogging detection by integrating advanced feature extraction and fusion technologies. The Backbone part of the model performs efficient feature extraction through the Focus module and GhostConv, ensuring that the key features of waterlogging can be quickly captured even in high-resolution images. The introduction of the CoordAtt attention mechanism and the BoT3 module enables the model to pay more attention to the areas related to waterlogging in the image, thereby improving the detection accuracy. The Neck part enhances the model's ability to identify waterlogging targets of different sizes through multi-scale feature fusion, which is particularly important for the diverse waterlogging scenarios that may occur in the power grid, ranging from small-area waterlogging to large-area waterlogging. The Head part is responsible for precise object detection at different scales and outputs the specific location and confidence level of the waterlogging.
[0086] Among them, the CoordAtt attention mechanism realizes the position information embedding operation and the position attention generation operation. For the structural schematic diagram of the CoordAtt attention mechanism, please refer to Figure 6 .
[0087] Ghost convolution is an innovative module proposed in GhostNet. Compared with ordinary convolution, Ghost convolution can generate more feature maps while achieving fewer parameters. Figure 7a and Figure 7bThe schematic diagram showing the differences between Ghost convolution and ordinary convolution is given. Ghost convolution divides the input feature map into "main" channels and "cheap" channels. The convolution operation is performed on the "main" channels, while the "cheap" channels use residual connections to splice the results of grouped convolution, thereby reducing the computational complexity. For Ghost convolution, assuming the input feature map size is h×w×c, the output feature map size is h′×w′×c′, where w and h are the width and height of the input feature map, w′ and h′ are the width and height of the output feature map, and the size of the convolution kernel is k×k. In the process of ordinary convolution, the number of convolution kernels n and the number of channels c are usually very large, and the FLOPs required for calculation are n×h′×w′×n×k×k. The FLOPs required for Ghost convolution calculation are n / s·h′·w′·c·k·k+(s - 1)n / s·h′·w′·d·d. In addition, the relationship between the computational amounts of ordinary convolution and Ghost convolution is as follows:
[0088]
[0089] Obviously, the computational amount of ordinary convolution is about s times that of Ghost convolution. That is, Ghost convolution has more advantages in terms of computational amount compared to ordinary convolution.
[0090] Ghost convolution is divided into three parts: conventional convolution, Ghost generation, and feature map splicing. The steps are as follows:
[0091] (1) Use conventional convolution to calculate the intrinsic feature maps Y w′×h′×m ;
[0092] (2) Then, perform grouped convolution operations on the feature maps V of each channel i to generate Ghost feature maps y ij , and the specific formula is:
[0093] (3) Splice the Ghost feature maps and the intrinsic feature maps to obtain the output result.
[0094] Both the GhostConv and C2fGhost modules are constructed based on the principle of Ghost convolution. Their detailed structures are as shown in Figure 8a and Figure 8bAs shown, in this embodiment, the GhostConv module decomposes the Conv module in YOLOv8 into a Conv operation and an "economical" residual connection. During this process, the number of channels of the Conv module is reduced to a part of the original. On the other hand, the GhostBottleneck module and the GhostC2f module are improved based on the Bottleneck and C2f modules of YOLOv8. By using the GhostConv module to replace the original Conv module, the number of parameters and the computational load are reduced without significantly increasing the computational amount. This method will improve the parameter efficiency of the improved YOLOv8 model, enabling the model to learn richer feature representations under limited computational resources, thereby enhancing the accuracy of power grid waterlogging detection.
[0095] The BoT3 module can more effectively comprehensively understand semantic data. This successfully addresses the concerns related to alleviating false positives and false negatives in distinguishing between objects and backgrounds within the model. BoT3 optimizes the components of the encoder and decoder models. The maintainability and scalability and modular functionality of the model are enhanced by adopting adaptive technologies. The BoT3 module adopted in this embodiment mainly consists of a convolutional block and a Bottleneck Transformer (BoT). The composition of the BoT3 module is as Figure 9 shown.
[0096] BoT stands for Bottleneck Transformer, which is a neural network module for image classification and object detection tasks. It represents an improved and optimized version of the Transformer network architecture. BoT includes three basic components: expansion, multi-head self-attention (MHSA), and contraction, as Figure 10 shown. Among them, MHSA is the key component of the BoT model, which features an enhanced channel attention mechanism derived from the self-attention mechanism. To sum up, this embodiment will adopt the BoT3 module. By leveraging the self-attention mechanism in the Transformer architecture, the model's ability to model long-range dependencies is enhanced to help the model identify global features related to waterlogging in complex power grid environments and improve the robustness of detection.
[0097] The Focus module is a convolutional neural network module for efficient downsampling. It was initially introduced in YOLOv5 and has been widely used in subsequent YOLO series models. Its core idea is to reorganize the spatial information of the input image into the channel dimension through a special slicing and splicing operation, thereby achieving lossless downsampling.
[0098] The working principle of the Focus module is to split the input image into multiple sub-images according to certain rules and then concatenate these sub-images along the channel dimension. The specific steps are as follows: The input image is split into four sub-images according to a 2×2 grid, corresponding to the upper left, upper right, lower left, and lower right corners respectively. These four sub-images are concatenated along the channel dimension to form a new feature map. In this way, the width and height of the input image are halved, while the number of channels is quadrupled.
[0099] The advantage of the Focus module is that it realizes efficient downsampling by reorganizing the spatial information of the input image into the channel dimension, while significantly reducing the computational amount. Different from traditional pooling or convolutional downsampling methods, the Focus module can better retain the detailed information of the image and avoid information loss caused by downsampling. In addition, by increasing the number of channels, the Focus module provides richer feature representations for subsequent convolutional layers, thus enhancing the model's ability to capture target details. These characteristics make the Focus module perform excellently in processing small object detection and multi-scale object detection tasks, significantly improving the overall performance and efficiency of the model.
[0100] Therefore, in this embodiment, the Focus module will be introduced into the improved YOLOv8 model to reduce the spatial dimension of the input image while retaining the key spatial information. Thus, the purpose of reducing the computational burden of the model and accelerating the processing speed is achieved, and at the same time, high-quality feature maps are provided for subsequent feature extraction.
[0101] Traditional YOLO series models mainly rely on a single attention mechanism (such as SE, CBAM) or local feature extraction modules, and it is difficult to balance global semantics and fine-grained features. This improved model innovatively integrates the CoordAtt coordinate attention mechanism and the global self-attention of the BoT3 module. Through CoordAtt, the spatial position sensitivity is enhanced to capture the edge features of ponding water, and at the same time, the long-range dependence modeling ability of BoT3 is used to extract the global information of the ponding water diffusion trend. The synergistic effect of the two enables the model to accurately locate ponding water under complex background interference, with the detection accuracy (mAP) increased by 12.7% and the false alarm rate reduced by 35%.
[0102] In the field of object detection, in order to improve the accuracy of the model, it is usually necessary to adopt dense convolutional operations or complex network structures, which often leads to a significant increase in computational resources. However, this increased computational burden may limit the application of the model in resource-constrained environments, such as mobile devices or edge computing devices. To address this challenge, the improved YOLOv8 model in this embodiment introduces the Ghost convolution module and the Focus module, aiming to achieve lightweight and efficient feature reconstruction.
[0103] In a specific embodiment, please refer to Figure 11, which is a schematic structural diagram of a water accumulation detection system for a transmission line in the second embodiment of the present application. The system includes: a first image acquisition module 201, an annotation module 202, a feature enhancement module 203, a training module 204, a second image acquisition module 205, and a water accumulation detection result output module 206; the first image acquisition module 201 is used to acquire water accumulation images of multiple water accumulation areas; the annotation module 202 is used to annotate the water accumulation areas of each water accumulation image to obtain the annotated water accumulation image; the feature enhancement module 203 is used to perform position information embedding operation and position attention generation operation on the annotated water accumulation image to obtain an enhanced feature map; the training module 204 is used to train a preset YOLOv8 model using all the enhanced feature maps to obtain a YOLOv8 object detection model; the second image acquisition module 205 is used to acquire a real-time image of the transmission line to be detected; the water accumulation detection result output module 206 is used to input the real-time image of the transmission line to be detected into the YOLOv8 object detection model for water accumulation detection and output the water accumulation detection result of the transmission line to be detected; the water accumulation detection result includes the water accumulation position, the water accumulation range, and the water accumulation confidence level.
[0104] In a specific embodiment, the third embodiment of the present application provides a water accumulation detection device for a transmission line, including a memory and a processor. The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.
[0105] In a specific embodiment, the fourth embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of the first embodiments of the present application.
[0106] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
[0107] The above is only a preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as the technical content of the present invention is not departed from, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still belong to the protection scope of the technical solution of the present invention.
Claims
1. A method for detecting water accumulation in a transmission line, characterized in that The method includes: Obtaining waterlogging images of multiple waterlogging areas; Performing waterlogging area annotation on each of the waterlogging images to obtain the annotated waterlogging images; Performing position information embedding operation and position attention generation operation on the annotated waterlogging images to obtain enhanced feature maps; Training a preset YOLOv8 model using all the enhanced feature maps to obtain a YOLOv8 object detection model; Obtaining a real-time image of the power transmission line to be detected; Inputting the real-time image of the power transmission line to be detected into the YOLOv8 object detection model for waterlogging detection, and outputting the waterlogging detection result of the power transmission line to be detected; the waterlogging detection result includes the waterlogging position, waterlogging range, and waterlogging confidence.
2. The water accumulation detection method for a transmission line according to claim 1, wherein The performing the position information embedding operation and position attention generation operation on the waterlogging images to obtain enhanced feature maps includes: Performing position information embedding operation on the waterlogging images along the width and height directions to obtain a first feature map in the width direction and a second feature map in the height direction; Performing position attention generation on the first feature map and the second feature map to obtain a first attention weight of the first feature map and a second attention weight of the second feature map; Performing a weighting operation on the waterlogging images based on the first attention weight and the second attention weight to obtain the enhanced feature maps.
3. The water accumulation detection method for a transmission line according to claim 2, characterized in that The first feature map and the second feature map are obtained by the following formula: Among them, is the first feature map, is the second feature map, W is the width of the ponding image, and x c (h, i) is the eigenvalue of the ponding image on channel c, i is the index for traversing the width during summation, H is the height of the ponding image, and x c (j, w) is the eigenvalue of the ponding image at positions (h, i) and (j, w).
4. The water accumulation detection method for a transmission line according to claim 2, characterized in that, The performing the position attention generation on the first feature map and the second feature map to obtain the first attention weight of the first feature map and the second attention weight of the second feature map includes: Performing a splicing operation on the first feature map and the second feature map to obtain a spliced feature map; Performing a dimensionality reduction operation and a normalization operation on the spliced feature map to obtain a normalized feature map; Performing a non-linear mapping on the normalized feature map to obtain a non-linear feature map; Performing convolution calculation on the non-linear feature map to obtain the height of the non-linear feature map and the width of the non-linear feature map; Obtaining the first attention weight according to the Sigmoid activation function and the height of the non-linear feature map; Obtaining the second attention weight according to the Sigmoid activation function and the width of the non-linear feature map.
5. The water accumulation detection method for a transmission line according to claim 4, wherein The non-linear feature map is obtained by the following formula: Where f is the non-linear feature map, N is the normalized feature map, α is a preset modulation factor, β is a preset coupling coefficient, γ is a preset linear transformation weight, δ is a preset offset compensation amount, η is a preset gain coefficient, λ is a preset attenuation factor, and θ is a preset high-order non-linear exponent.
6. The water accumulation detection method for a transmission line according to claim 2, wherein The enhanced feature maps are obtained by the following formula: Among them, y c (i, j) is the enhanced feature map, x c (i, j) is the ponding image, is the first attention weight, is the second attention weight.
7. The water accumulation detection method for a transmission line according to claim 1, characterized in that After obtaining the waterlogging images of multiple waterlogging areas, it further includes: Performing preprocessing on each of the waterlogging images to obtain preprocessed waterlogging images; the preprocessing includes image cropping, image scaling, and image normalization; Then the performing waterlogging area annotation on each of the waterlogging images to obtain the annotated waterlogging images is: Performing waterlogging area annotation on each of the preprocessed waterlogging images to obtain the annotated waterlogging images.
8. A water accumulation detection system for a transmission line, characterized in that The system includes: a first image acquisition module, an annotation module, a feature enhancement module, a training module, a second image acquisition module, and a water accumulation detection result output module; The first image acquisition module is used to acquire water accumulation images of multiple water accumulation areas; The annotation module is used to annotate the water accumulation areas of each of the water accumulation images to obtain the annotated water accumulation images; The feature enhancement module is used to perform position information embedding operation and position attention generation operation on the annotated water accumulation images to obtain enhanced feature maps; The training module is used to train a preset YOLOv8 model using all the enhanced feature maps to obtain a YOLOv8 object detection model; The second image acquisition module is used to acquire real-time images of the transmission line to be detected; The water accumulation detection result output module is used to input the real-time images of the transmission line to be detected into the YOLOv8 object detection model for water accumulation detection and output the water accumulation detection results of the transmission line to be detected; the water accumulation detection results include the water accumulation position, the water accumulation range, and the water accumulation confidence level.
9. An accumulated water detection device for a transmission line, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 7.