Abnormality detection method and device of power transmission tower, electronic equipment and storage medium

By using an automated detection method based on a dual-branch semantic segmentation network model, combined with a bidirectional attention fusion module for semantic and spatial branches, the problems of low efficiency and high cost in transmission tower detection are solved, achieving efficient and accurate anomaly detection.

CN120182673BActive Publication Date: 2026-01-06SHAOGUAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510231506.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2026-01-06
Estimated Expiration
2045-02-27

AI Technical Summary

Technical Problem

Existing technologies for detecting defects in power transmission towers are inefficient, costly, and pose safety risks, and manual inspection methods are insufficient to meet the requirements.

Method used

An anomaly detection method based on a bi-branch semantic segmentation network model is adopted. Images are acquired by UAVs and automatically detected by a pre-trained detection model. The method combines a bi-directional attention fusion module with semantic and spatial branches to achieve complementarity and enhancement of image semantic and spatial features.

Benefits of technology

It has automated the detection of anomalies in power transmission towers, reducing detection costs and risks, and improving detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182673B_ABST
    Figure CN120182673B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, electronic device, and storage medium for anomaly detection of transmission towers. The method first acquires an image of the transmission tower to be detected, then inputs the image into a pre-trained detection model to obtain a target semantic segmentation result. The detection model is trained on a bi-branch semantic segmentation network model with semantic and spatial branches and a bidirectional attention fusion module based on training data. This technical solution achieves automated anomaly detection of transmission towers by inputting the acquired image of the transmission tower to be detected into a pre-trained detection model to obtain the target semantic segmentation result. The bidirectional attention fusion module complements and enhances the semantic and spatial features of the transmission tower image, thereby improving the efficiency and accuracy of anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power transmission technology, and in particular to a method, device, electronic equipment and storage medium for detecting anomalies in power transmission towers. Background Technology

[0002] In the field of power transmission, transmission towers are key structures that support power lines and ensure power transmission; their integrity is crucial to the safety and reliability of the power system.

[0003] In the existing technology, manual inspection is used to detect defects in transmission towers. However, manual inspection has technical problems such as low detection efficiency, high cost and safety risks. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for detecting anomalies in transmission towers, in order to solve the technical problems of low efficiency, high cost, and safety risks in the prior art for detecting defects in transmission towers.

[0005] In a first aspect, embodiments of this application provide a method for detecting anomalies in transmission towers, including:

[0006] Acquire the image of the transmission tower to be inspected;

[0007] The image to be detected is input into a pre-trained detection model to obtain the target semantic segmentation result. The target semantic segmentation result is the damage information of each component in the transmission tower to be detected. The detection model is obtained by training a bi-branch semantic segmentation network model based on training data. The training data includes: multiple first images corresponding to the transmission tower and the annotation information corresponding to the multiple first images. The bi-branch semantic segmentation network model includes: multiple first basic modules of the semantic branch and multiple second basic modules of the spatial branch. A bidirectional attention fusion module is set between each two adjacent first basic modules and each two adjacent second basic modules.

[0008] In one possible implementation, before inputting the image to be detected into a pre-trained detection model to obtain the target semantic segmentation result, the method further includes:

[0009] Obtain multiple first images corresponding to the power transmission towers and the annotation information corresponding to the multiple first images;

[0010] The detection model is obtained by training the dual-branch semantic segmentation network model based on the plurality of first images and the corresponding annotation information of the plurality of first images.

[0011] In one possible implementation, training the dual-branch semantic segmentation network model based on the plurality of first images and the corresponding annotation information of the plurality of first images to obtain the detection model includes:

[0012] For each first image, the first image is input into the backbone network of the dual-branch semantic segmentation network model to obtain a second image after downsampling.

[0013] The second image is input into the semantic branch and the spatial branch respectively, and the semantic features and spatial features are fused through the bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch respectively.

[0014] The third image is input into the pyramid pooling module of the dual-branch semantic segmentation network model to obtain the first semantic feature containing multi-scale information.

[0015] The fourth image and the first semantic feature are input into the feature fusion module in the dual-branch semantic segmentation network model to obtain the first fusion information;

[0016] Based on the semantic segmentation results corresponding to the first fusion information of each first image and the annotation information corresponding to each first image, the parameters in the dual-branch semantic segmentation network model are adjusted, and the adjusted dual-branch semantic segmentation network model is used as the detection model.

[0017] In one possible implementation, the step of inputting the second image into the semantic branch and the spatial branch respectively, and performing semantic and spatial feature fusion processing through the bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch respectively, includes:

[0018] Based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module, determine the second semantic feature and the first spatial feature corresponding to the second image;

[0019] The third semantic feature and the second spatial feature are determined based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module.

[0020] Based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module, the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch are determined.

[0021] In one possible implementation, determining the second semantic feature and the first spatial feature corresponding to the second image based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module includes:

[0022] The second image is input into the first basic module of the semantic branch to obtain the fifth image after convolution and downsampling processing;

[0023] The second image is input into the first second basic module of the spatial branch to obtain the sixth image after convolution processing;

[0024] The fifth image and the sixth image are input into the first bidirectional attention fusion module to obtain the second semantic feature corresponding to the fused fifth image and the first spatial feature corresponding to the sixth image.

[0025] In one possible implementation, determining the third semantic feature and the second spatial feature based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module includes:

[0026] The second semantic feature and the fifth image are input into the second first basic module of the semantic branch to obtain the seventh image after convolution and downsampling processing;

[0027] The first spatial feature and the sixth image are input into the second basic module of the spatial branch to obtain the eighth image after convolution processing;

[0028] The seventh and eighth images are input into the second bidirectional attention fusion module to obtain the third semantic feature corresponding to the fused seventh image and the second spatial feature corresponding to the eighth image.

[0029] In one possible implementation, determining the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module includes:

[0030] The third semantic feature and the seventh image are input into the third first basic module of the semantic branch to obtain the third image after convolution and downsampling processing;

[0031] The second spatial feature and the eighth image are input into the third second basic module of the spatial branch to obtain the fourth image after convolution processing.

[0032] Secondly, embodiments of this application provide an anomaly detection device for transmission towers, comprising:

[0033] The acquisition module is used to acquire the image of the transmission tower to be inspected.

[0034] The determination module is used to input the image to be detected into a pre-trained detection model to obtain the target semantic segmentation result. The target semantic segmentation result is the damage information of each component in the transmission tower to be detected. The detection model is obtained by training a bi-branch semantic segmentation network model based on training data. The training data includes: multiple first images corresponding to the transmission tower and the annotation information corresponding to the multiple first images. The bi-branch semantic segmentation network model includes: multiple first basic modules of the semantic branch and multiple second basic modules of the spatial branch. A bidirectional attention fusion module is set between each two adjacent first basic modules and each two adjacent second basic modules.

[0035] In one possible implementation, before inputting the image to be detected into a pre-trained detection model to obtain the target semantic segmentation result, the determining module is further configured to:

[0036] Obtain multiple first images corresponding to the power transmission towers and the annotation information corresponding to the multiple first images;

[0037] The detection model is obtained by training the dual-branch semantic segmentation network model based on the plurality of first images and the corresponding annotation information of the plurality of first images.

[0038] In one possible implementation, the step of training the dual-branch semantic segmentation network model based on the plurality of first images and the corresponding annotation information of the plurality of first images to obtain the detection model, wherein the determination module is specifically used for:

[0039] For each first image, the first image is input into the backbone network of the dual-branch semantic segmentation network model to obtain a second image after downsampling.

[0040] The second image is input into the semantic branch and the spatial branch respectively, and the semantic features and spatial features are fused through the bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch respectively.

[0041] The third image is input into the pyramid pooling module of the dual-branch semantic segmentation network model to obtain the first semantic feature containing multi-scale information.

[0042] The fourth image and the first semantic feature are input into the feature fusion module in the dual-branch semantic segmentation network model to obtain the first fusion information;

[0043] Based on the semantic segmentation results corresponding to the first fusion information of each first image and the annotation information corresponding to each first image, the parameters in the dual-branch semantic segmentation network model are adjusted, and the adjusted dual-branch semantic segmentation network model is used as the detection model.

[0044] In one possible implementation, the second image is input to the semantic branch and the spatial branch respectively, and the semantic and spatial features are fused through the bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch respectively. The determination module is specifically used for:

[0045] Based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module, determine the second semantic feature and the first spatial feature corresponding to the second image;

[0046] The third semantic feature and the second spatial feature are determined based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module.

[0047] Based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module, the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch are determined.

[0048] In one possible implementation, the step of determining the second semantic feature and the first spatial feature corresponding to the second image based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module, wherein the determining module is specifically used for:

[0049] The second image is input into the first basic module of the semantic branch to obtain the fifth image after convolution and downsampling processing;

[0050] The second image is input into the first second basic module of the spatial branch to obtain the sixth image after convolution processing;

[0051] The fifth image and the sixth image are input into the first bidirectional attention fusion module to obtain the second semantic feature corresponding to the fused fifth image and the first spatial feature corresponding to the sixth image.

[0052] In one possible implementation, the step of determining the third semantic feature and the second spatial feature based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module, wherein the determining module is specifically used for:

[0053] The second semantic feature and the fifth image are input into the second first basic module of the semantic branch to obtain the seventh image after convolution and downsampling processing;

[0054] The first spatial feature and the sixth image are input into the second basic module of the spatial branch to obtain the eighth image after convolution processing;

[0055] The seventh and eighth images are input into the second bidirectional attention fusion module to obtain the third semantic feature corresponding to the fused seventh image and the second spatial feature corresponding to the eighth image.

[0056] In one possible implementation, the step of determining the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module, wherein the determining module is specifically used for:

[0057] The third semantic feature and the seventh image are input into the third first basic module of the semantic branch to obtain the third image after convolution and downsampling processing;

[0058] The second spatial feature and the eighth image are input into the third second basic module of the spatial branch to obtain the fourth image after convolution processing.

[0059] Thirdly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0060] The memory stores computer-executed instructions;

[0061] The processor executes computer execution instructions stored in the memory to implement the method as described in the first aspect or any of the above methods.

[0062] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method described in the first aspect or any of the above-described methods.

[0063] Fifthly, embodiments of this application provide a computer program, the computer program product including a computer program stored in a computer-readable storage medium, at least one processor can read the computer program from the computer-readable storage medium, and the at least one processor can implement the method described in the first aspect or any of the above methods when executing the computer program.

[0064] The present application provides an anomaly detection method, device, electronic device, and storage medium for transmission towers. The method first acquires the image to be detected corresponding to the transmission tower to be detected, and then inputs the image to be detected into a pre-trained detection model to obtain the target semantic segmentation result. The target semantic segmentation result is the damage information of each component in the transmission tower to be detected. The detection model is obtained by training a bi-branch semantic segmentation network model based on training data. The training data includes: multiple first images corresponding to the transmission tower and the annotation information corresponding to the multiple first images. The bi-branch semantic segmentation network model includes: multiple first basic modules of the semantic branch and multiple second basic modules of the spatial branch. A bidirectional attention fusion module is set between each two adjacent first basic modules and each two adjacent second basic modules. This technical solution inputs the acquired images of the transmission towers to be detected into a pre-trained detection model with semantic and spatial branches and a bidirectional attention fusion module to obtain the target semantic segmentation result. This achieves automated detection of transmission tower anomalies, reduces the cost and danger of transmission tower anomaly detection, and utilizes the bidirectional attention fusion module to achieve complementarity and enhancement of semantic and spatial features of transmission tower images, thereby improving the efficiency and accuracy of transmission tower anomaly detection. Attached Figure Description

[0065] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0066] Figure 1 A schematic diagram of the framework of the dual-branch semantic segmentation network model provided in the embodiments of this application;

[0067] Figure 2 A flowchart illustrating the anomaly detection method for transmission towers provided in this application embodiment. Figure 1 ;

[0068] Figure 3 A flowchart illustrating the anomaly detection method for transmission towers provided in this application embodiment. Figure 2 ;

[0069] Figure 4 A flowchart illustrating the anomaly detection method for transmission towers provided in this application embodiment. Figure 3 ;

[0070] Figure 5 This is a schematic diagram of the bidirectional attention fusion module in the detection model provided in the embodiments of this application;

[0071] Figure 6 This is a schematic diagram of the pyramid pooling module in the detection model provided in the embodiments of this application;

[0072] Figure 7 This is a schematic diagram of the feature fusion module in the detection model provided in the embodiments of this application;

[0073] Figure 8 A schematic diagram of the structure of the anomaly detection device for transmission towers provided in the embodiments of this application;

[0074] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.

[0075] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0076] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0077] First, the terms used in the embodiments of this application will be explained:

[0078] Bilateral Attention Fusion Module (BAFM): This is a module used in deep learning networks that aims to improve model performance by combining bidirectional (or dual-channel) attention mechanisms.

[0079] The Pyramid Pooling Module (PPM) is a module used in computer vision, particularly for image segmentation tasks. Its main purpose is to capture contextual information at different scales within an image through multi-scale pooling operations, thereby helping the model better understand the global structure and local details of the image.

[0080] Seg Head: The module responsible for generating the final segmentation result, usually located in the last layer of the neural network or near the output.

[0081] Basic Block: The basic modules or units that make up a neural network. Each basic block usually contains one or more operations such as convolution, batch normalization, and activation functions. These operations work together to achieve a specific function and can be stacked repeatedly to build more complex networks.

[0082] Before introducing the embodiments of this application, the application background of the embodiments of this application will be explained first:

[0083] In the field of power transmission, transmission towers are key structures that support power lines and ensure power transmission; their integrity is crucial to the safety and reliability of the power system.

[0084] With the growth in scale and complexity of power grids, existing technologies that rely on manual inspections for defect detection of transmission towers suffer from low detection efficiency, high costs, and safety risks.

[0085] In recent years, deep learning technology has made significant breakthroughs in computer vision, particularly in the subfield of semantic segmentation. Semantic segmentation methods can not only accurately classify each pixel in an image but also identify the contours and components of objects in complex environments. Deep neural networks, through end-to-end training, have demonstrated powerful performance in image recognition and object detection tasks. Compared to earlier techniques, deep learning models are more advanced in feature extraction and better able to handle challenges such as varying lighting conditions and background noise.

[0086] In the application of drones for inspecting power transmission towers, the use of deep learning technology has opened up new avenues for solving existing problems. Using semantic segmentation models, numerous structures and components on the towers can be accurately identified, including insulators, crossarms, bolts, and other devices, even detecting minute defects and damage under complex environmental conditions. Furthermore, deep learning models can automatically learn from massive amounts of inspection data and perform self-optimization, improving the accuracy and efficiency of anomaly detection on power transmission towers.

[0087] To address the technical problems existing in the prior art, the inventors of this application propose the following solution: Due to the low efficiency, high cost, and safety risks associated with manual inspections, a drone is used to acquire images of the transmission towers to be inspected. These images are then input into a pre-trained detection model for processing, yielding target semantic segmentation results. This allows for the identification of anomalies in the transmission towers within the inspected images. The aforementioned detection model is based on a bi-branch semantic segmentation network model, which includes semantic and spatial branches, as well as a bidirectional attention fusion module positioned between the basic modules of the semantic and spatial branches. The use of semantic and spatial branches, along with the bidirectional attention fusion module, enhances information exchange between the two branches, achieving complementarity and enhancement of image semantic and spatial features. This allows for more efficient acquisition of the category and anomaly features of each component of the transmission tower in the inspected image, effectively improving the detection efficiency of transmission towers and reducing the cost and risks associated with anomaly detection.

[0088] First, a brief explanation of the two-branch semantic segmentation network model will be given. Figure 1 A schematic diagram of the framework of the dual-branch semantic segmentation network model provided in the embodiments of this application is shown below. Figure 1 As shown, the dual-branch semantic segmentation network model includes: a backbone network, a semantic branch module, a spatial branch module, BAFM, PPM, a feature fusion module, and a segmentation head for outputting results.

[0089] The backbone network forms the foundation of the entire segmentation model, responsible for extracting multi-level features from the input image. The semantic branch module primarily extracts and focuses on the semantic information of the image. This module typically focuses on global semantic context features to capture overall semantic category information, such as the component categories and damage information of power transmission towers. The spatial branch module focuses on capturing spatial structural information in the image. It processes spatial details of the image, such as the edges, shapes, and positions of objects. BAFM is used to fuse information between the semantic and spatial branches. By introducing a bidirectional attention mechanism, BAFM can adaptively perform weighted fusion between semantic and spatial information, effectively improving the network's feature representation ability. This allows information from different sources (such as spatial and semantic) to complement each other, optimizing the final segmentation result. PPM captures multi-scale contextual information by performing pooling operations on feature maps at different scales, helping the network understand objects of different sizes and adaptively adjust segmentation boundaries in the image. The feature fusion module fuses features from different modules (such as the semantic branch, spatial branch, BAFM, and PPM) to obtain a more comprehensive feature description, thereby improving segmentation accuracy. The segmentation head is the last layer of the network, used to map the fused feature map onto the class label of each pixel to generate the final segmentation result.

[0090] The dual-branch semantic segmentation network model provided in this application embodiment, through the collaborative work of the above modules, can simultaneously capture semantic and spatial information, perform multi-scale processing, and fuse image semantic features and spatial features to ultimately output accurate segmentation results efficiently.

[0091] The technical solution of this application will now be described in detail through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0092] The executing entity of this application may be a drone, or an electronic device such as a backend server or terminal equipment.

[0093] Figure 2 A flowchart illustrating the anomaly detection method for transmission towers provided in this application embodiment. Figure 1 ,like Figure 2 As shown, the method may include the following steps:

[0094] Step 21: Obtain the image of the transmission tower to be inspected.

[0095] In this step, sensor devices are used to acquire images of the transmission towers to be inspected.

[0096] The sensor device can be a high-definition camera or a drone equipped with a high-definition camera.

[0097] In one possible implementation, the image to be detected may include transmission tower image information acquired in real time or in advance under different angles and lighting conditions.

[0098] Optionally, after acquiring the image to be detected, preprocessing (e.g., noise reduction, brightness adjustment, coordinate transformation, and normalization) can be performed on the image to be detected to improve the image quality and provide better image input for subsequent anomaly detection of transmission towers.

[0099] Step 22: Input the image to be detected into the pre-trained detection model to obtain the target semantic segmentation result.

[0100] The target semantic segmentation result is the damage information of each component in the transmission tower to be detected. The detection model is obtained by training a bi-branch semantic segmentation network model based on training data. The training data includes: multiple first images corresponding to the transmission tower and the annotation information corresponding to the multiple first images. The bi-branch semantic segmentation network model includes: multiple first basic modules of the semantic branch and multiple second basic modules of the spatial branch. A bidirectional attention fusion module is set between each two adjacent first basic modules and each two adjacent second basic modules.

[0101] In this step, the preprocessed image to be detected is input into a pre-trained detection model. After a series of convolution and feature fusion operations, the target semantic segmentation result of the image to be detected is output.

[0102] The target semantic segmentation result is the category label of each pixel in the image to be detected and the corresponding damage information. The category label of each pixel is obtained by assigning each pixel in the input image to a specific category (such as the tower body, tower base, conductor suspension device, crossbeam and other components in the transmission tower) in the semantic segmentation task. The damage status of each component (such as the tower body, tower base and other components) on the transmission tower in the image can be identified through the target semantic segmentation result.

[0103] The anomaly detection method for transmission towers provided in this application first acquires the image to be detected corresponding to the transmission tower to be detected, and then inputs the image to be detected into a pre-trained detection model to obtain the target semantic segmentation result. The target semantic segmentation result is the damage information of each component in the transmission tower to be detected. The detection model is obtained by training a bi-branch semantic segmentation network model based on training data. The training data includes: multiple first images corresponding to the transmission tower and the annotation information corresponding to the multiple first images. The bi-branch semantic segmentation network model includes: multiple first basic modules of the semantic branch and multiple second basic modules of the spatial branch. A bidirectional attention fusion module is set between each two adjacent first basic modules and each two adjacent second basic modules. This technical solution inputs the acquired images of the transmission towers to be detected into a pre-trained detection model with semantic and spatial branches and a bidirectional attention fusion module to obtain the target semantic segmentation result. This achieves automated detection of anomalies in transmission towers, reduces the cost and danger of anomaly detection, and utilizes the bidirectional attention fusion module to achieve complementarity and enhancement of semantic and spatial features of the transmission tower images. This allows for more efficient acquisition of the category features and anomaly features of each component of the transmission tower in the image to be detected, thereby improving the efficiency and accuracy of anomaly detection.

[0104] Based on the above embodiments, Figure 3A flowchart illustrating the anomaly detection method for transmission towers provided in this application embodiment. Figure 2 ,like Figure 3 As shown, prior to step 22, the method further includes the following steps:

[0105] Step 31: Obtain multiple first images corresponding to the transmission towers and the annotation information corresponding to the multiple first images;

[0106] In this step, a large number of images of power transmission towers are collected using sensor devices to obtain the corresponding images of the power transmission towers (i.e., the first images) and the annotation information corresponding to the images of the power transmission towers.

[0107] The first image consists of a large number of pre-acquired images of power transmission towers in different environments and from various perspectives.

[0108] In addition, for each image of a power transmission tower, it is also necessary to obtain the corresponding annotation information, namely the category of each component in each image of the power transmission tower and whether each component is damaged (such as broken, aged, etc.). The annotation information can be obtained manually or through some auxiliary tools to ensure that the category annotation of each pixel is accurate.

[0109] After obtaining the images corresponding to the transmission towers, multiple first images are preprocessed (e.g., noise reduction, brightness adjustment, coordinate transformation, and normalization) to obtain higher quality first images.

[0110] Step 32: Train the dual-branch semantic segmentation network model based on multiple first images and their corresponding annotation information to obtain the detection model.

[0111] In this step, multiple first images and their corresponding annotations are input into a two-branch semantic segmentation model for training. The two-branch semantic model learns the features and semantic information of the first images, adjusts parameters based on the pixel features of the first images and the category information of the annotations, and optimizes the semantic segmentation performance so that it can effectively identify different elements in the image and determine whether they are abnormal. After training, the model will have the ability to automatically analyze the image to be detected and output the semantic segmentation results, that is, the detection model is obtained.

[0112] Among them, the multiple first images can be multiple images after preprocessing.

[0113] The anomaly detection method for transmission towers provided in this application first acquires multiple first images corresponding to the transmission tower and their corresponding annotation information. Then, based on the multiple first images and their corresponding annotation information, a bi-branch semantic segmentation network model is trained to obtain a detection model. This technical solution uses multiple first images corresponding to the transmission tower and their corresponding annotation information to provide comprehensive image information and image category features, improving the generalization ability of the detection model. The multiple first images and their corresponding annotation information are then input into the bi-branch semantic segmentation model for training, fusing the semantic and spatial features of the transmission tower in each image. This allows the model to accurately learn the component category information and damage status of the transmission tower in each image. The trained detection model can achieve fast and efficient anomaly detection for transmission towers.

[0114] Based on the above embodiments, Figure 4 A flowchart illustrating the anomaly detection method for transmission towers provided in this application embodiment. Figure 3 ,like Figure 4 As shown, step 32 above may include the following steps:

[0115] Step 41: For each first image, input the first image into the backbone network of the dual-branch semantic segmentation network model to obtain the second image after downsampling.

[0116] In this step, each first image is input into the backbone network of the dual-branch semantic segmentation network, and each first image is downsampled to obtain a second image.

[0117] In one possible implementation, such as Figure 1 As shown, the backbone network is an optimization based on the ResNet-18 network model. The initial module (stem block) of the original ResNet-18 is improved to obtain the backbone network. Specifically, the 7x7 convolutional layer is replaced with two 3x3 convolutional layers with a stride of 2. This can downsample the first image to 1 / 8 of the resolution of the first image. The above improvement not only reduces the number of parameters, but also helps to extract the features of the first image and enhances the nonlinearity of the backbone network, thereby improving the processing efficiency of image features.

[0118] Step 42: Input the second image into the semantic branch and the spatial branch respectively, and perform semantic feature and spatial feature fusion processing through the bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch respectively;

[0119] In this step, the second image is input to the semantic branch and the spatial branch respectively. The semantic branch is mainly responsible for capturing the semantic features of the second image, while the spatial branch focuses on capturing the spatial layout or structural feature information in the second image. Then, the semantic features and spatial features are fused through the bidirectional attention fusion module to further optimize the complementarity of the semantic features and spatial features, and the fused third image corresponding to the semantic branch and the fused fourth image corresponding to the spatial branch are obtained respectively.

[0120] Optionally, step 42 includes the following implementation:

[0121] Step 1: Based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module, determine the second semantic features and the first spatial features corresponding to the second image;

[0122] In this step, the second image is input into the first basic module and the second basic module of the semantic branch and the spatial branch, respectively. Then, the processed second image feature information is input into the first bidirectional attention fusion module to determine the second semantic feature and the first spatial feature corresponding to the second image.

[0123] In one possible implementation, such as Figure 1 As shown, the resolution of the second image after processing by the basic modules in the backbone network is 1 / 8 of the resolution of the first image. On the one hand, the basic modules in the backbone network input the second image into the first basic module in the semantic branch for processing, resulting in an image with a resolution of 1 / 16 of the first image. On the other hand, the basic modules in the backbone network input the second image into the first basic module in the spatial branch for processing, resulting in an image with a resolution of 1 / 8 of the first image. Then, the images processed by the semantic and spatial branches are input into BAFM for feature fusion, resulting in the fused second semantic features and first spatial features corresponding to the second image.

[0124] The semantic branch processes low-resolution images, while the spatial branch processes high-resolution images. Information fusion is achieved through BAFM, which helps to complement and enhance the second image features.

[0125] Optionally, step 1 includes the following implementation:

[0126] S1. Input the second image into the first basic module of the semantic branch to obtain the fifth image after convolution and downsampling processing;

[0127] In this implementation, the second image is input into the first basic module of the semantic branch, where convolution and downsampling operations are performed to obtain deep semantic information, and the fifth image is obtained after processing.

[0128] In one possible implementation, such as Figure 1 As shown, the resolution of the fifth image obtained after the second image is processed by the first basic module of the semantic branch is 1 / 16 of that of the first image.

[0129] S2. Input the second image into the first second basic module of the spatial branch to obtain the sixth image after convolution processing;

[0130] In this implementation, the second image is input into the first second basic module of the spatial branch and convolution is performed to obtain the processed sixth image.

[0131] In one possible implementation, the first second basic module's convolution is a 3×3 convolution with a stride of 1, which preserves the feature map of the first image at 1 / 8 resolution to supplement the spatial features of the semantic branches, while removing redundant convolution operations, thereby reducing the computational cost and parameter count of the network. The resulting sixth image has a resolution of 1 / 8 of the first image.

[0132] S3. Input the fifth and sixth images into the first bidirectional attention fusion module to obtain the second semantic feature corresponding to the fifth image and the first spatial feature corresponding to the sixth image after fusion.

[0133] In this implementation, both the fifth and sixth images contain rich feature information. Since there are differences in the image features of the semantic branch and the spatial branch, the fifth and sixth images are input into the first bidirectional attention fusion module for feature fusion processing to obtain the second semantic feature corresponding to the fifth image and the first spatial feature corresponding to the sixth image.

[0134] Among them, the bidirectional attention fusion module is designed based on the attention mechanism. This module aims to enhance the information exchange between the two branches, realize the complementarity and enhancement of semantic features and spatial features, and help improve the accuracy and robustness of the detection model.

[0135] In one possible implementation, Figure 5 This is a schematic diagram of the bidirectional attention fusion module in the detection model provided in the embodiments of this application, as shown below. Figure 5 As shown, BAFM contains two main paths: one is a low-resolution-high-resolution path from the semantic branch to the spatial branch, and the other is a high-resolution-low-resolution path from the spatial branch to the semantic branch.

[0136] In high-resolution-low-resolution paths (such as...) Figure 5The dashed lines in the diagram represent the spatial branch's feature map, which contains rich spatial details. The number of channels in the spatial branch's feature map is expanded to match that of the semantic branch through downsampling, and then the spatial and semantic features are directly added together for fusion.

[0137] In low-resolution-high-resolution paths (such as...) Figure 5 (The solid lines in the diagram represent the parts). First, the low-resolution feature map of the semantic branch is upsampled using bilinear interpolation. Then, the upsampled feature map is added to the convolutional feature map of the spatial branch and input into the attention mechanism module to obtain the fused feature. Next, the fused feature is input into the sigmoid function to obtain the relative attention weight α of the semantic branch, while the relative attention weight of the spatial branch is 1-α. Based on these weights, the feature maps of the two branches are multiplied point-by-point and the results are added together. Finally, this sum is added to the original image features of the spatial branch to obtain the fused result of the spatial branch.

[0138] The final fusion output of the spatial branch and the semantic branch can be expressed by the formula:

[0139] F′ H =F H +(F H ·(1-α)+F L ·α)

[0140] F′ L =Downsample(F H )+F L

[0141] Among them, F H and F L F′ represents the spatial branch input features (e.g., the features corresponding to the sixth image) and the semantic branch input features (e.g., the features corresponding to the fifth image), respectively. H and F′ L These are the spatial branch output features (e.g., the first spatial feature) and the semantic branch output features (e.g., the second semantic feature), respectively. Downsample(*) represents the downsampling operation.

[0142] Step 2: Determine the third semantic feature and the second spatial feature based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module.

[0143] In this step, the second semantic feature and the first spatial feature are respectively input into the second first basic module and the second second basic module of the semantic branch for processing, and the processed image features are input into the second bidirectional attention fusion module for feature fusion, thereby determining the third semantic feature and the second spatial feature.

[0144] Optionally, step 2 may include the following implementation:

[0145] S1. Input the second semantic feature and the fifth image into the second basic module of the semantic branch to obtain the seventh image after convolution and downsampling processing;

[0146] In this implementation, the second semantic features and the fifth image are input into the second first basic module of the semantic branch, and convolution and downsampling operations are performed in the second first basic module to obtain the seventh image.

[0147] In one possible implementation, such as Figure 1 As shown, the resolution of the seventh image obtained after processing by the second basic module is 1 / 32 of that of the first image.

[0148] S2. Input the first spatial features and the sixth image into the second basic module of the spatial branch to obtain the eighth image after convolution processing;

[0149] In this implementation, the first spatial feature and the sixth image are input into the second basic module of the spatial branch, and a convolution operation is performed in the second basic module to obtain the eighth image.

[0150] In one possible implementation, such as Figure 1 As shown, the convolution in the second basic module is a 3×3 convolution with a stride of 1, and the resolution of the eighth image obtained after processing is 1 / 8 of that of the first image.

[0151] S3. Input the seventh and eighth images into the second bidirectional attention fusion module to obtain the third semantic feature corresponding to the seventh image and the second spatial feature corresponding to the eighth image after fusion.

[0152] In this implementation, the seventh and eighth images are input into the second bidirectional attention fusion module. Feature fusion processing is performed according to the feature fusion method of the low-resolution-high-resolution path from the semantic branch to the spatial branch and the high-resolution-low-resolution path from the spatial branch to the semantic branch, so as to obtain the third semantic feature corresponding to the seventh image and the second spatial feature corresponding to the eighth image after feature fusion.

[0153] Step 3: Based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module, determine the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch.

[0154] In this step, the third semantic feature is input into the third first basic module for processing to obtain the third image, and the second spatial feature is input into the third second basic module for processing to obtain the fourth image.

[0155] In one possible implementation, the resolution of the third image is 1 / 64 of the resolution of the first image, and the resolution of the fourth image is 1 / 8 of the resolution of the first image.

[0156] Optionally, step 3 can be implemented as follows:

[0157] S1. Input the third semantic feature and the seventh image into the third basic module of the semantic branch to obtain the third image after convolution and downsampling.

[0158] In this implementation, the third semantic feature and the seventh image are input into the third basic module of the semantic branch, and convolution and downsampling are performed in the first basic module to obtain the third image.

[0159] The third semantic feature is obtained after processing by the second bidirectional attention fusion module, and the seventh image is obtained after processing by the second first basic module of the semantic branch.

[0160] S2. Input the second spatial features and the eighth image into the third second basic module of the spatial branch to obtain the fourth image after convolution processing.

[0161] In this implementation, the second spatial features and the eighth image are input into the third second basic module of the spatial branch, and convolution processing is performed in the second basic module to obtain the fourth image.

[0162] Among them, the second spatial feature is obtained after processing by the second bidirectional attention fusion module, and the eighth image is obtained after processing by the second basic module of the spatial branch.

[0163] Step 43: Input the third image into the pyramid pooling module in the dual-branch semantic segmentation network model to obtain the first semantic feature containing multi-scale information;

[0164] In this step, the third image is input into the pyramid pooling module for processing, and feature information at different scales in the third image is extracted to obtain the first semantic feature containing multi-scale information.

[0165] Among them, the multi-level pooling operation of the pyramid pooling module can extract semantic information of the image from different scales. Through this multi-scale feature extraction, the dual-branch semantic segmentation network model can capture contextual information at different scales, which helps to improve the detection capability of the detection model in complex power transmission tower image scenes.

[0166] In one possible implementation, Figure 6 This is a schematic diagram of the pyramid pooling module in the detection model provided in the embodiments of this application, combined with... Figure 1 and Figure 6 Using a feature map of 1 / 64 resolution from the first image as input, channel compression is first performed through a 1×1 convolution layer, followed by mean pooling with strides of 2, 4, and 8 (corresponding to...). Figure 6 The AvgPooling1, AvgPooling2, and AvgPooling3 modules generate feature maps of 1 / 128, 1 / 256, and 1 / 512 resolution, respectively. The feature map of the third image (i.e., the original image input to the pyramid pooling module) is then concatenated with the global mean pooling feature map to expand the size of the third image's feature map. If the input feature is x... in Then each scale pooling feature It can be expressed as follows:

[0167]

[0168] Where i represents different scales, C 1×1 For a 1×1 convolution operation, P k,s (*) represents a pooling operation with a kernel size of k and a stride of s, P global (*) indicates a global pooling operation.

[0169] Among them, multi-scale feature images are obtained through pooling operations. Next, bilinear interpolation was used for upsampling to maintain consistent image size. Then, 3×3 depthwise convolution was used for feature extraction, and contextual information at different scales was continuously incorporated using a bottom-up, stepwise residual approach. The feature map of the third image is not upsampled; it is directly compared with the upsampled feature map. The features are then fused together, followed by a 3×3 convolution operation. The output of this convolution operation is then fused with the features from the next scale. All output results... It can be expressed as follows:

[0170]

[0171] Among them, C 3×3 This is a 3×3 convolution operation, and U(*) is a bilinear interpolation upsampling operation. This indicates additive fusion. To capture the relationships between different scales and reduce the dimensionality of the channels, the individual output features are... Perform a concatenation operation, followed by channel compression using a 1×1 convolution. Then, concatenate all features, and finally fuse them with the output features of the third image to obtain the final output feature x. outIt can be expressed by the formula:

[0172]

[0173] Among them, C 1×1 This is a 1×1 convolution, and concat(*) is used for concatenation. This indicates a feature addition and fusion operation.

[0174] The final output feature x mentioned above out This is the first semantic feature that contains multi-scale information.

[0175] Step 44: Input the fourth image and the first semantic features into the feature fusion module in the dual-branch semantic segmentation network model to obtain the first fused information;

[0176] In this step, the fourth image and the first semantic feature are input into the feature fusion module, which effectively combines the fourth image and the first semantic feature to obtain the first fusion information of the image corresponding to the fourth image and the first semantic feature, thereby improving the accuracy of the detection model.

[0177] In one possible implementation, Figure 7 This is a schematic diagram of the feature fusion module in the detection model provided in the embodiments of this application, as shown below. Figure 7 As shown, firstly, the semantic branch is used to perform convolution and upsampling operations on the lower-resolution image, and the spatial branch is used to perform two consecutive convolution operations with a stride of 2 on the higher-resolution image to achieve downsampling of the spatial branch. This spatial branch is then multiplied with the convolved semantic branch to achieve the interaction of semantic and spatial information. Next, the result of the first multiplication operation is multiplied a second time with the image features after the convolution operation of the semantic branch, followed by downsampling to obtain the downsampled result of the semantic branch. Then, the result of the first multiplication operation is multiplied a third time with the image features after the convolution operation of the spatial branch, followed by a fourth multiplication operation with the downsampled result of the semantic branch. Finally, the image is processed through a convolution layer, batch normalization, and a Rectified Linear Unit (ReLU) activation function to obtain the final fused first information.

[0178] Among them, based on the feature fusion module, feature fusion of spatial and semantic branches is realized, which realizes the fusion of rich spatial information in high-resolution feature maps with rich semantic information in low-resolution maps, and realizes mutual complementarity and enhancement of image features. By using such a bidirectional feature complementarity mechanism, a comprehensive understanding of image information can be effectively achieved, improving the accuracy and robustness of semantic segmentation tasks.

[0179] Step 45: Based on the semantic segmentation results corresponding to the first fusion information of each first image and the annotation information corresponding to each first image, adjust the parameters in the bi-branch semantic segmentation network model, and use the adjusted bi-branch semantic segmentation network model as the detection model.

[0180] In this step, the first fusion information corresponding to each first image is input into the segmentation head (e.g., Figure 1 The first fusion information is processed to obtain the semantic segmentation result. Based on the semantic segmentation result and the annotation information corresponding to each first image, the parameters in the dual-branch semantic segmentation network model are adjusted, and the adjusted dual-branch semantic segmentation network model is used as the detection model.

[0181] The segmentation head mainly consists of a 3×3 convolution, a batch normalization (BN) layer, ReLU activation, and finally a 1×1 convolution for dimensionality compression. The main branch loss handles the majority of the network's learning, while the auxiliary loss is responsible for optimization. Different weights are assigned to each, and the final loss is calculated as follows:

[0182] L f =L m +αL a

[0183] Where L f L represents the final loss. m L represents the loss of the main branch. a This indicates an auxiliary loss.

[0184] In one possible implementation, the parameters may include the learning rate, the number of iterations, the loss, etc.

[0185] In one possible implementation, the bi-branch semantic segmentation model is trained using multiple first images and their corresponding annotations. The model parameters are adjusted, and the model training employs stochastic gradient descent with an initial learning rate of 0.01 and a polynomial decay rate of 0.9. The learning rate formula is as follows:

[0186]

[0187] Among them, lr init The initial learning rate is given by `iter`, and the current iteration number is given by `iter`. max This represents the total number of iterations.

[0188] Furthermore, using auxiliary losses at appropriate locations in the bi-branch semantic segmentation network model can effectively address the propagation problem in shallow layers of the network during backpropagation. Therefore, drawing on this approach, this application adds a segmentation head after each stage in the proposed bi-branch model to calculate the auxiliary loss. Considering the trade-off between computational load and efficiency, the bi-branch semantic segmentation network model used in this application only adds an auxiliary loss function (e.g., the Online Hard Example Mining (OHEM) loss function) after the first interaction fusion (i.e., the first bidirectional attention fusion module) of the spatial branch to optimize training, thereby obtaining the various parameters in the bi-branch semantic segmentation network model.

[0189] The anomaly detection method for transmission towers provided in this application firstly inputs each first image into the backbone network of a two-branch semantic segmentation network model to obtain a downsampled second image. The second image is then input into both the semantic and spatial branches, and a bidirectional attention fusion module is used to fuse semantic and spatial features, resulting in a third image corresponding to the semantic branch and a fourth image corresponding to the spatial branch. The third image is then input into the pyramid pooling module of the two-branch semantic segmentation network model to obtain first semantic features containing multi-scale information. The fourth image and the first semantic features are then input into the feature fusion module of the two-branch semantic segmentation network model to obtain first fusion information. Finally, based on the semantic segmentation results corresponding to the first fusion information for each first image and the annotation information corresponding to each first image, the parameters in the two-branch semantic segmentation network model are adjusted, and the adjusted two-branch semantic segmentation network model is used as the detection model. This technical solution effectively improves the accuracy and robustness of semantic segmentation through multi-stage, multi-level feature extraction and fusion, combined with a two-branch structure and a bidirectional attention mechanism. Through precise parameter adjustment, the detection model also exhibits good adaptability and flexibility in the application of anomaly detection for transmission towers.

[0190] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0191] Figure 8 This is a schematic diagram of the anomaly detection device for power transmission towers provided in an embodiment of this application. Figure 8 As shown, the device includes:

[0192] The acquisition module 81 is used to acquire the image of the transmission tower to be detected.

[0193] The determination module 82 is used to input the image to be detected into the pre-trained detection model to obtain the target semantic segmentation result. The target semantic segmentation result is the damage information of each component in the transmission tower to be detected. The detection model is obtained by training a bi-branch semantic segmentation network model based on training data. The training data includes: multiple first images corresponding to the transmission tower and the annotation information corresponding to the multiple first images. The bi-branch semantic segmentation network model includes: multiple first basic modules of the semantic branch and multiple second basic modules of the spatial branch. A bidirectional attention fusion module is set between each two adjacent first basic modules and each two adjacent second basic modules.

[0194] In one possible implementation, before inputting the image to be detected into a pre-trained detection model to obtain the target semantic segmentation result, the determination module 82 is further configured to:

[0195] Obtain multiple first images corresponding to the power transmission towers and the annotation information corresponding to the multiple first images;

[0196] Based on multiple first images and their corresponding annotation information, a dual-branch semantic segmentation network model is trained to obtain a detection model.

[0197] In one possible implementation, a dual-branch semantic segmentation network model is trained based on multiple first images and their corresponding annotation information to obtain a detection model. The determination module is specifically used for:

[0198] For each first image, the first image is input into the backbone network of the dual-branch semantic segmentation network model to obtain a second image after downsampling.

[0199] The second image is input into the semantic branch and the spatial branch respectively, and the semantic features and spatial features are fused through the bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch respectively.

[0200] The third image is input into the pyramid pooling module in the dual-branch semantic segmentation network model to obtain the first semantic feature containing multi-scale information.

[0201] The fourth image and the first semantic features are input into the feature fusion module in the dual-branch semantic segmentation network model to obtain the first fused information.

[0202] Based on the semantic segmentation results corresponding to the first fusion information of each first image and the annotation information corresponding to each first image, the parameters in the bi-branch semantic segmentation network model are adjusted, and the adjusted bi-branch semantic segmentation network model is used as the detection model.

[0203] In one possible implementation, the second image is input to both the semantic branch and the spatial branch, and the semantic and spatial features are fused using a bidirectional attention fusion module to obtain the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch, respectively. The determination module 82 is specifically used for:

[0204] Based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module, determine the second semantic features and the first spatial features corresponding to the second image;

[0205] The third semantic feature and the second spatial feature are determined based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module.

[0206] Based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module, determine the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch.

[0207] In one possible implementation, based on the second image, the first basic module of the semantic branch, the first second basic module of the spatial branch, and the first bidirectional attention fusion module, the second semantic feature and the first spatial feature corresponding to the second image are determined. The determining module 82 is specifically used for:

[0208] The second image is input into the first basic module of the semantic branch to obtain the fifth image after convolution and downsampling processing;

[0209] The second image is input into the first second basic module of the spatial branch to obtain the sixth image after convolution processing;

[0210] The fifth and sixth images are input into the first bidirectional attention fusion module to obtain the second semantic feature corresponding to the fifth image and the first spatial feature corresponding to the sixth image after fusion.

[0211] In one possible implementation, the third semantic feature and the second spatial feature are determined based on the second semantic feature, the second first basic module of the semantic branch, the first spatial feature, the second second basic module of the spatial branch, and the second bidirectional attention fusion module. The determining module 82 is specifically used for:

[0212] The second semantic feature and the fifth image are input into the second first basic module of the semantic branch to obtain the seventh image after convolution and downsampling processing;

[0213] The first spatial feature and the sixth image are input into the second basic module of the spatial branch to obtain the eighth image after convolution processing;

[0214] The seventh and eighth images are input into the second bidirectional attention fusion module to obtain the third semantic feature corresponding to the seventh image and the second spatial feature corresponding to the eighth image after fusion.

[0215] In one possible implementation, based on the third semantic feature, the third first basic module, the second spatial feature, and the third second basic module, the third image corresponding to the semantic branch and the fourth image corresponding to the spatial branch are determined. The determining module 82 is specifically used for:

[0216] The third semantic feature and the seventh image are input into the third basic module of the semantic branch to obtain the third image after convolution and downsampling.

[0217] The second spatial features and the eighth image are input into the third second basic module of the spatial branch to obtain the fourth image after convolution processing.

[0218] The apparatus provided in this application embodiment can be used to execute the determination method in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0219] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. Additionally, these modules can be fully or partially integrated together, or implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through the integrated logic circuits in the hardware of the processor element or through software instructions.

[0220] Figure 9 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 9 As shown, the electronic device may include: a processor 91, a memory 92, and computer program instructions stored in the memory 92 and executable on the processor 91. When the processor 91 executes the computer program instructions, it implements the method provided in any of the foregoing embodiments.

[0221] Optionally, the various components of the electronic device can be connected via a system bus.

[0222] The memory 92 can be a separate memory unit or a memory unit integrated into the processor 91. The number of processors 91 can be one or more.

[0223] It should be understood that the processor 91 can be a Central Processing Unit (CPU), or other general-purpose processors 91, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor 91 can be a microprocessor 91, or any conventional processor 91. The steps of the method disclosed in this application can be directly manifested as being executed by the hardware processor 91, or being executed by a combination of hardware and software modules within the processor 91.

[0224] The system bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus. Memory 92 may include random access memory (RAM) 92, and may also include non-volatile memory (NVM) 92, such as at least one disk storage device 92.

[0225] All or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory 92. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory 92 (storage medium) includes: read-only memory 92 (ROM), RAM, flash memory 92, hard disk, solid-state hard disk, magnetic tape, floppy disk, optical disk, and any combination thereof.

[0226] The electronic device provided in this application embodiment can be used to execute the method provided in any of the above method embodiments. Its implementation principle and technical effect are similar, and will not be repeated here.

[0227] This application provides a computer-readable storage medium storing computer instructions that, when executed on a computer, cause the computer to perform the above-described method.

[0228] The aforementioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0229] Optionally, a readable storage medium can be coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Alternatively, the readable storage medium can be an integral part of the processor. Both the processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components within the device.

[0230] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium, and the at least one processor can implement the above-described method when executing the computer program.

[0231] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method of anomaly detection of a power transmission tower, characterized by, The method comprises: acquiring a to-be-detected image corresponding to a to-be-detected power transmission tower; inputting the to-be-detected image into a pre-trained detection model to obtain a target semantic segmentation result, the target semantic segmentation result being damage information of each element in the to-be-detected power transmission tower, the detection model being obtained by training a double-branch semantic segmentation network model based on training data, the training data comprising a plurality of first images corresponding to a power transmission tower and label information corresponding to the plurality of first images, the double-branch semantic segmentation network model comprising a plurality of first basic modules of a semantic branch and a plurality of second basic modules of a spatial branch, and a bidirectional attention fusion module being arranged between each two adjacent first basic modules and each two adjacent second basic modules; before the inputting, the method further comprises: for each first image, inputting the first image into a backbone network in the double-branch semantic segmentation network model to obtain a second image subjected to down-sampling processing; inputting the second image into a first first basic module of the semantic branch to obtain a fifth image subjected to convolution and down-sampling processing, inputting the second image into a first second basic module of the spatial branch to obtain a sixth image subjected to convolution processing, and inputting the fifth image and the sixth image into a first bidirectional attention fusion module to obtain second semantic features corresponding to the fifth image after fusion and first spatial features corresponding to the sixth image; inputting the second semantic features and the fifth image into a second first basic module of the semantic branch to obtain a seventh image subjected to convolution and down-sampling processing, inputting the first spatial features and the sixth image into a second second basic module of the spatial branch to obtain an eighth image subjected to convolution processing, and inputting the seventh image and the eighth image into a second bidirectional attention fusion module to obtain third semantic features corresponding to the seventh image after fusion and second spatial features corresponding to the eighth image; inputting the third semantic features and the seventh image into a third first basic module of the semantic branch to obtain a third image subjected to convolution and down-sampling processing, and inputting the second spatial features and the eighth image into a third second basic module of the spatial branch to obtain a fourth image subjected to convolution processing; inputting the third image into a pyramid pooling module in the double-branch semantic segmentation network model to obtain first semantic features containing multi-scale information; inputting the fourth image and the first semantic features into a feature fusion module in the double-branch semantic segmentation network model to obtain first fusion information; adjusting parameters in the double-branch semantic segmentation network model according to semantic segmentation results corresponding to the first fusion information corresponding to each first image and label information corresponding to each first image, and taking the double-branch semantic segmentation network model after the adjustment as the detection model.

2. The method of claim 1, wherein, Before inputting the first image into a backbone network in the double-branch semantic segmentation network model to obtain a second image subjected to down-sampling processing, the method further comprises: Obtaining the plurality of first images corresponding to the power transmission tower and the annotation information corresponding to the plurality of first images.

3. An abnormality detection device for a power transmission tower, characterized by comprising: The device comprises: An acquisition module configured to acquire a to-be-detected image corresponding to a to-be-detected power transmission tower; A determination module configured to input the to-be-detected image into a pre-trained detection model to obtain a target semantic segmentation result, the target semantic segmentation result being damage information of each element in the to-be-detected power transmission tower, the detection model being obtained by training a double-branch semantic segmentation network model based on training data, the training data comprising a plurality of first images corresponding to a power transmission tower and annotation information corresponding to the plurality of first images, the double-branch semantic segmentation network model comprising a plurality of first basic modules of a semantic branch and a plurality of second basic modules of a spatial branch, and a bidirectional attention fusion module being arranged between each adjacent two first basic modules and each adjacent two second basic modules; Before inputting the to-be-detected image into the pre-trained detection model to obtain the target semantic segmentation result, the determination module is further configured to: Input the first image into the backbone network in the double-branch semantic segmentation network model to obtain the second image subjected to down-sampling processing; Input the second image into a first first basic module of the semantic branch to obtain a fifth image subjected to convolution and down-sampling processing, input the second image into a first second basic module of the spatial branch to obtain a sixth image subjected to convolution processing, and input the fifth image and the sixth image into a first bidirectional attention fusion module to obtain a second semantic feature corresponding to the fifth image after fusion and a first spatial feature corresponding to the sixth image; Input the second semantic feature and the fifth image into a second first basic module of the semantic branch to obtain a seventh image subjected to convolution and down-sampling processing, input the first spatial feature and the sixth image into a second second basic module of the spatial branch to obtain an eighth image subjected to convolution processing, and input the seventh image and the eighth image into a second bidirectional attention fusion module to obtain a third semantic feature corresponding to the seventh image after fusion and a second spatial feature corresponding to the eighth image; Input the third semantic feature and the seventh image into a third first basic module of the semantic branch to obtain a third image subjected to convolution and down-sampling processing, and input the second spatial feature and the eighth image into a third second basic module of the spatial branch to obtain a fourth image subjected to convolution processing; Input the third image into a pyramid pooling module in the double-branch semantic segmentation network model to obtain a first semantic feature containing multi-scale information; Input the fourth image and the first semantic feature into a feature fusion module in the double-branch semantic segmentation network model to obtain first fusion information; According to the semantic segmentation result corresponding to the first fusion information corresponding to each first image and the annotation information corresponding to each first image, parameters in the double-branch semantic segmentation network model are adjusted, and the double-branch semantic segmentation network model after adjustment is taken as the detection model.

4. An electronic device, comprising: Comprise: A processor, and a memory connected with the processor in communication; The memory stores computer-executed instructions; The processor executes the computer-executed instructions stored in the memory to implement the method of any one of claims 1 to 2.

5. A computer readable storage medium, characterized in that, The computer-readable storage medium stores computer-executed instructions, and the computer-executed instructions are executed by the processor to implement the method of any one of claims 1 to 2.

Citation Information

Patent Citations

  • Double-resolution real-time semantic segmentation method based on detail enhancement

    CN117409412A

  • High-precision semantic segmentation method for automatic driving road scene

    CN117649526A

  • Foreign matter attachment identification method and device for power transmission line, and computer equipment

    CN117953426A