Method and device for detecting small animals in transformer substation and electronic equipment
By using detection models in the substation to extract multi-scale features and perform alternate attention treatments, the problem of low detection efficiency and accuracy of small animals is solved, and more efficient and accurate detection of small animals is achieved.
Patent Information
- Application Number
- CN202411794068.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-06
AI Technical Summary
Small and medium-sized animals in substations enter the live equipment area and cause short circuits and tripping accidents, and the existing detection methods are less efficient and accurate.
A substation small animal detection method is adopted. By obtaining the substation target image and inputting it to the trained detection model, multi-scale features are extracted using the pyramid transformer module, the boundary generation module performs feature fusion, the feature enhancement module performs feature enhancement, the channel reduction module performs channel dimensionality reduction, and finally, the multi-guided cross-aggregation module performs alternate attention processing to predict the target area where the small animal is located.
It improves the efficiency and accuracy of small animals detection in the substation, and can more accurately locate the target area where the small animals are located.
Smart Images

Figure CN119942585A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of electric power technology, and in particular, relates to a method, device and electronic equipment for detecting small animals in a substation. Background Art
[0002] Substations are an important part of the power system. Their main function is to convert the voltage of the electricity generated by the power plant for easy transmission and distribution to end users. Substations usually include a variety of electrical equipment, such as transformers, switchgear, and protection devices, which are used to increase or decrease the voltage, control the current, and protect the power system. Because substations occupy a large area, they are often located in remote areas. In substations, small animals such as birds, cats, mice, and geckos are happy to live in substations. If small animals enter the live equipment area in the substation, short circuits and tripping accidents will occur, which will destroy the operating environment of the power equipment, threaten the continued safe and stable operation of the power equipment, and increase the maintenance pressure on operators.
[0003] In some scenarios, small animal baffles, mouse traps, and mouse traps are often used to detect small animals in substations. However, this method is relatively passive and requires regular inspections by maintenance personnel. Therefore, when using this method to detect small animals in substations, its detection efficiency and accuracy are relatively low. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide a method, device and electronic equipment for detecting small animals in a substation, which can solve the problem of low detection efficiency and accuracy when detecting small animals in a substation.
[0005] In a first aspect, a method for detecting small animals in a substation is provided, which is executed by a terminal, and the method includes: acquiring a target image of the substation; inputting the target image into a trained detection model; extracting multi-scale features of different resolutions of the target image through a pyramid transformer module of the detection model, performing feature fusion on the multi-scale features through a boundary generation module of the detection model to obtain boundary guidance of the target image, performing feature enhancement on the highest-level features in the multi-scale features through a feature enhancement module of the detection model to obtain global guidance of the target image, performing channel dimension reduction on the multi-scale features through a channel reduction module of the detection model to obtain trunk guidance; inputting the boundary guidance, the global guidance and the trunk guidance into a multi-guide cross-aggregation module of the detection model for alternating attention processing, and predicting a target area where small animals are located in the target image.
[0006] In a second aspect, a small animal detection device in a substation is provided, comprising: an acquisition module for acquiring a target image of the substation; an input module for inputting the target image into a trained detection model; a detection module for extracting multi-scale features of different resolutions of the target image through a pyramid transformer module of the detection model, performing feature fusion on the multi-scale features through a boundary generation module of the detection model to obtain boundary guidance of the target image, performing feature enhancement on the highest-level features in the multi-scale features through a feature enhancement module of the detection model to obtain global guidance of the target image, and performing channel dimension reduction on the multi-scale features through a channel reduction module of the detection model to obtain trunk guidance; and a processing module for inputting the boundary guidance, the global guidance and the trunk guidance into a multi-guide cross-aggregation module of the detection model for alternating attention processing to predict a target area where small animals are located in the target image.
[0007] In a third aspect, an embodiment of the present application provides an electronic device comprising: a memory, a processor, and computer executable instructions stored in the memory and executable on the processor, wherein the computer executable instructions implement the steps of executing the method of the first aspect when executed by the processor.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which is used to store computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the method of the first aspect are implemented.
[0009] In an embodiment of the present application, a target image of a substation is first acquired; then the target image is input into a trained detection model; multi-scale features of different resolutions of the target image are extracted through a pyramid transformer module of the detection model, and feature fusion is performed on the multi-scale features through a boundary generation module of the detection model to obtain boundary guidance of the target image, and feature enhancement is performed on the highest-level features of the multi-scale features through a feature enhancement module of the detection model to obtain global guidance of the target image, and channel dimension reduction is performed on the multi-scale features through a channel reduction module of the detection model to obtain trunk guidance; finally, the boundary guidance, the global guidance and the trunk guidance are input into a multi-guide cross-aggregation module of the detection model for alternating attention processing to predict the target area where the small animals in the target image are located.
[0010] In this way, the embodiment of the present application can input the collected target image of the substation into a pre-trained detection model to obtain multi-scale features of different resolutions, perform feature fusion on the multi-scale features, perform feature enhancement on the highest-level features in the multi-scale features, and perform channel dimension reduction on the multi-scale features to obtain boundary guidance, global guidance, and trunk guidance. Finally, the boundary guidance, the global guidance, and the trunk guidance are alternately processed through a multi-guide cross-aggregation module to obtain the target area where the small animals are located in the target image, and gradually refine it in a cascade manner to obtain the target area where the small animals are located with higher accuracy and integrity. When the method provided in the embodiment of the present application is used to detect small animals in the substation, the detection efficiency and accuracy are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0012] Figure 1 A schematic flow chart of a method for detecting small animals in a substation provided in an embodiment of the present application.
[0013] Figure 2 A schematic diagram of the structure of a detection model provided in an embodiment of the present application.
[0014] Figure 3 A schematic diagram of the structure of a boundary generation module in a detection model provided in an embodiment of the present application.
[0015] Figure 4 A schematic diagram of the structure of a feature enhancement module and a channel reduction module in a detection model provided in an embodiment of the present application.
[0016] Figure 5 A schematic diagram of the structure of a multi-guide cross-aggregation module in a detection model provided in an embodiment of the present application.
[0017] Figure 6 A schematic diagram of the structure of a small animal detection device in a substation provided in an embodiment of the present application.
[0018] Figure 7 This is one of the structural schematic diagrams of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.
[0020] The technical solution provided by the embodiments of the present application is described in detail below through some embodiments and their application scenarios in combination with the accompanying drawings.
[0021] Figure 1 The flowchart of a method for detecting small animals in a substation provided by an embodiment of the present application is shown. The method can be executed by an electronic device, such as a terminal device or a server device. In other words, the method can be executed by software or hardware installed in the terminal device or the server device. Figure 1 As shown, the method may include the following steps.
[0022] Step S101, acquiring a target image of a substation.
[0023] Specifically, the target image of the substation can be taken by using a drone equipped with a camera to shoot various areas of the substation. The target image includes but is not limited to the equipment, installation area, and various small animals in the substation.
[0024] Step S102, inputting the target image into the trained detection model.
[0025] Specifically, the detection model can be a trained detection model, such as Figure 2 As shown, Figure 2 A schematic diagram of the structure of a detection model provided in an embodiment of the present application. Figure 2 In , the detection model includes a pyramid transformer module, a boundary generation module, a feature enhancement module, a channel reduction module, and a multi-guided cross aggregation module.
[0026] Further, such as Figures 3 to 5 As shown, Figure 3 A schematic diagram of the structure of a boundary generation module in a detection model provided in an embodiment of the present application. Figure 4 A schematic diagram of the structure of a feature enhancement module and a channel reduction module in a detection model provided in an embodiment of the present application. Figure 5 A schematic diagram of the structure of a multi-guide cross-aggregation module in a detection model provided in an embodiment of the present application. Figure 3 In , the boundary generation module consists of multiple 3×3 convolutional layers and one 1×1 convolutional layer. Figure 4In the proposed method, the feature enhancement module consists of an asymmetric convolutional layer, a hole convolutional layer, a normal convolutional layer, and a 1×1 convolutional layer. The channel reduction module consists of two 3×3 convolutional layers. Figure 5 In , the multi-guided cross-aggregation module consists of a forward flow branch, a reverse flow branch, and a 1×1 convolutional layer.
[0027] Step S103, extracting multi-scale features of different resolutions of the target image through the pyramid transformer module of the detection model, performing feature fusion on the multi-scale features through the boundary generation module of the detection model to obtain boundary guidance of the target image, performing feature enhancement on the highest-level features in the multi-scale features through the feature enhancement module of the detection model to obtain global guidance of the target image, and performing channel dimension reduction on the multi-scale features through the channel reduction module of the detection model to obtain trunk guidance.
[0028] Specifically, the collected target image is input into the pre-trained pyramid transformer module to obtain four levels of multi-scale features with different resolutions {F i ,i=1,2,3,4}, where i represents the level of the feature. The higher the level of the feature, the higher the resolution of the image. Then the multi-scale feature {F i ,i=1,2,3,4} is input into the boundary generation module, the boundary generation module includes multiple first convolution layers and one second convolution layer, the first convolution layer can be a 3×3 convolution layer, and the second convolution layer can be a 1×1 convolution layer.
[0029] Further, as an optional embodiment of the present application, the multi-scale features are subjected to feature fusion through the boundary generation module of the detection model to obtain the boundary guidance of the target image, including: inputting the multi-scale features into multiple first convolutional layers of the boundary generation module, and fusing the multi-scale features from bottom to top adjacent convolutional layers to obtain intermediate features of the multi-scale features; inputting the intermediate features into multiple first convolutional layers of the boundary generation module again, and fusing the intermediate features from bottom to top adjacent convolutional layers to obtain a fourth fused feature that fuses the multi-scale features; inputting the fourth fused feature into the second convolutional layer for channel dimensionality reduction to obtain the boundary guidance.
[0030] Specifically, if Figure 3 As shown, the multi-scale features {F i , i=1,2,3,4} are input to the 3×3 convolution layer of the boundary generation module respectively, and then the output of the 3×3 convolution layer is fused from bottom to top to obtain four intermediate features {f i e , i=1,2,3,4}. In the embodiment of the present application, the following formula is used to represent the intermediate characteristics:
[0031]
[0032] In the above formula, Conv represents a 3×3 convolutional layer, up represents upsampling, and * represents an element-wise multiplication operation.
[0033] Further, the intermediate feature {f i e , i=1,2,3,4} is input again into multiple 3×3 convolutional layers of the boundary generation module, and the intermediate features are fused from bottom to top in adjacent convolutional layers to obtain the fourth fusion feature f1 that fuses the multi-scale features d The embodiment of the present application uses the following formula to represent the fourth fusion feature:
[0034]
[0035] Finally, the fourth fusion feature f1 d Input to the second convolutional layer for channel dimension reduction to obtain the boundary guidance G edge .
[0036] Furthermore, when the highest-level features in the multi-scale features are enhanced by the feature enhancement module, as an optional embodiment of the present application, the highest-level features are input into the hole convolution branch of the feature enhancement module to expand the receptive field to obtain the first processing feature, the highest-level features are input into the asymmetric convolution branch of the feature enhancement module to increase the inference speed to obtain the second processing feature, and the highest-level features are input into the normal convolution branch of the feature enhancement module to improve the diversity of multi-level features to obtain the third processing feature; the first processing feature, the second processing feature and the third processing feature are connected and then channel dimension reduction is performed to obtain the fourth processing feature; the highest-level features are input into the residual branch of the feature enhancement module to perform channel dimension reduction to obtain the fifth processing feature; the fourth processing feature and the fifth processing feature are added and activated by an activation function to obtain the global guidance.
[0037] Specifically, if Figure 4As shown, the highest-level feature F4 is input into the hole convolution branch, and a 3×3 hole convolution is used to expand the receptive field, reduce the loss of details caused by the large hole rate, and set the hole rate to 3. The highest-level feature F4 is input into the asymmetric convolution branch of the feature enhancement module, and the inference speed is increased through a pair of asymmetric convolutions of 1×3 convolution layers and 3×1 convolution layers. The highest-level feature F4 is input into the normal convolution branch of the feature enhancement module, and a 3×3 convolution is used to increase the diversity of the extracted multi-level features. The outputs of the above-mentioned convolution branches are connected, and channel dimensionality reduction is performed through a 3×3 convolution layer to obtain the fourth processing feature. The embodiment of the present application uses the following formula to represent the fourth processing feature:
[0038] F cat =Conv(Cat[C asy (F4),C dil (F4),C ori (F4)])
[0039] In the above formula, Cat represents the connection operation. Conv represents the 3×3 convolutional layer. asy It is an asymmetric convolution. C dil is a dilated convolution. C ori is a normal convolution. F cat Indicates that the fourth processing feature F is obtained by performing channel dimensionality reduction through a 3×3 convolution layer cat .
[0040] Furthermore, the highest-level feature F4 is input into the residual branch of the feature enhancement module, and channel dimension reduction is performed through a 1×1 convolution layer to retain a certain amount of original detail information. cat The output of the residual branch is added and activated by the GeLU activation function to obtain the global guide G glo .
[0041] Furthermore, when the multi-scale features are subjected to channel dimensionality reduction through the channel reduction module of the detection model, as an optional embodiment of the present application, the multi-scale features are input into the third convolutional layer of the channel reduction module for channel dimensionality reduction to obtain a first feature; the first feature is input into the fourth convolutional layer of the channel reduction module to remove noise interference to obtain the trunk guidance.
[0042] Specifically, the third convolutional layer and the fourth convolutional layer can be 3×3 convolutional layers, and the obtained multi-scale features {F i , i=1,2,3,4} is first input into a 3×3 convolution to reduce the channel dimension and obtain the first feature. Then the first feature is input into the 3×3 convolution layer of the channel reduction module to reduce noise interference and obtain the backbone guidance In the present application, the following formula is used to represent
[0043]
[0044] In the above formula, Conv represents a 3×3 convolutional layer. Indicates trunk boot.
[0045] Step S104, inputting the boundary guidance, global guidance and trunk guidance into the multi-guide cross-aggregation module of the detection model for alternating attention processing, and predicting the target area where the small animal is located in the target image.
[0046] Specifically, when predicting the target area where the small animal is located in the target image, as an optional embodiment of the present application, the global guidance is used as the first fusion feature; the boundary guidance, the trunk guidance and the first fusion feature are input into the multi-guided cross-aggregation module of the third layer of the detection model, and the multi-guided cross-aggregation module of the third layer alternately focuses on the target area containing the small animal and the background area not containing the small animal in the target image to obtain the three-level features and the three-level prediction map of the target area; the three-level prediction map is added to the global guidance to obtain the second fusion feature; the boundary guidance, the trunk guidance, the second fusion feature and the three The multi-guide cross-aggregation module of the second layer inputs the level features, alternately focusing on the target area containing small animals and the background area not containing small animals in the target image, to obtain the secondary features and the secondary prediction map of the target area; the secondary prediction map is added to the global guidance to obtain the third fusion feature; the boundary guidance, the trunk guidance, the third fusion feature and the secondary feature are input to the multi-guide cross-aggregation module of the first layer, alternately focusing on the target area containing small animals and the background area not containing small animals in the target image, to obtain the primary features and the primary prediction map of the target area, the primary prediction map includes the finally predicted image of the target area containing the small animal.
[0047] Specifically, the global guide is used as the first fusion feature The boundary guide G edge , the trunk guide and the first fusion feature The multi-guided cross-aggregation module input to the third layer of the detection model alternately focuses on the target area containing small animals and the background area that does not contain small animals through a dual-branch structure to obtain three-level features and the three-level prediction map Then the three-level prediction map With global guidance G glo Add together to get the second fusion feature Further guide the boundary G edge, the trunk guide The second fusion feature And the three-level features The input is sent to the multi-guided cross-aggregation module of the second layer, which continues to alternately focus on the target area and the background area through a dual-branch structure to obtain the secondary features. and the secondary prediction graph Furthermore, the obtained secondary prediction graph and global guide G glo Add together to get the third fusion feature And guide the boundary to G edge , the trunk guide The third fusion feature And the secondary features The multi-guided cross-aggregation module input to the first layer alternately focuses on the target area and the background area through a dual-branch structure to obtain the first-level features and the first-level prediction graph
[0048] Furthermore, when the target area containing small animals and the background area not containing small animals in the target image are alternately focused on through the multi-guided cross-aggregation module of the third layer to obtain the three-level features and the three-level prediction map of the target area, as an optional embodiment of the present application, the first fusion feature is upsampled and grouped and fused with the trunk guidance, and then multiplied by a first predetermined parameter to obtain a forward flow branch result; the first fusion feature is upsampled and reverse edge integrated, and a pooling operation is used to focus on the boundary of the background area in the target image to obtain a reverse map, and the reverse map is grouped and fused with the trunk guidance and multiplied by a second predetermined parameter to obtain a reverse flow branch result; the first fusion feature is subtracted from the forward flow branch result and then added to the reverse flow branch result to obtain a fourth fusion feature of the fused two branches; the boundary guidance is supplemented to the fourth fusion feature to obtain the three-level feature; the three-level feature is input into the convolution layer of the multi-guided cross-aggregation module of the third layer for channel compression to obtain the three-level prediction map.
[0049] Specifically, the first fusion feature After upsampling and the backbone guidance Perform group fusion, and then multiply by the first predetermined parameter α to obtain the forward flow branch result. The embodiment of the present application uses the following formula to represent the forward flow branch result
[0050]
[0051] In the above formula, represents the forward flow branch result. α represents the first predetermined parameter. GF represents group fusion. up represents upsampling. Indicates a backbone guide. * indicates an element-wise multiplication operation. Represents the first fused feature.
[0052] Furthermore, the first fusion feature After upsampling, reverse edge integration is performed, and the pooling operation is used to focus on the boundaries in the background area to obtain the reverse graph R, which reduces the amount of calculation while mining the potential information of the features. In the embodiment of the present application, the following formula is used to calculate the reverse graph R:
[0053]
[0054] In the above formula, MaxPool represents the maximum pooling operation with a kernel of 5. σ represents the sigmoid function. Ε represents the unit 1 matrix. up represents upsampling. Represents the first fused feature.
[0055] Further, the reverse graph R is connected to the trunk guide After grouping and merging, multiply it by the second predetermined parameter β to obtain the reverse flow branch result The present application embodiment adopts the following formula for calculation:
[0056]
[0057] In the above formula, represents the reverse flow branch result. β represents the second predetermined parameter. R represents the reverse graph. = represents backbone guidance. * represents element-wise multiplication. GF represents group fusion.
[0058] Furthermore, the first fusion feature Subtract the forward flow branch result Then branch with the reverse flow result Add together to get the fourth fusion feature F of the fused two branches nr , the fourth fusion feature F is represented by the following formula in the present embodiment: nr :
[0059]
[0060] In the above formula, F nr Represents the fourth fusion feature. Indicates the result of the forward flow branch. Represents the first fused feature. Represents the result of a reverse flow branch.
[0061] Furthermore, the conditional batch normalization method is used to guide the boundary G edge The rich detail information in the fourth fusion feature F nr In the above example, we get the third-level features Then the third level features After a 1×1 convolution for channel compression, a three-level prediction map is obtained
[0062] Furthermore, after obtaining the target area where the small animal is located in the target image, the first-level prediction image is interpolated. The size of is adjusted to the size of the target image, and the detection image P of the small animal is obtained. final .
[0063] Furthermore, in the embodiment of the present application, a joint supervision strategy is adopted to train the detection model, and a hybrid loss function is used to stimulate the network to learn segmentation features at the pixel level, the object level, and the pixel level. Therefore, the loss function in the collaborative supervision strategy can be a hybrid loss function consisting of a two-level cross entropy loss function, an enhanced alignment loss function, and a weighted IOU loss function.
[0064] Specifically, the embodiment of the present application uses the following formula to represent the loss function of the detection model:
[0065]
[0066] In the above formula, L total represents the loss function of the detection model. G represents the true value map. L hybrid represents the mixed loss function. G edge Indicates boundary guidance. Represents a prediction graph.
[0067] Furthermore, the present embodiment uses the following formula to calculate L hybrid :
[0068]
[0069] In the above formula, is the weighted binary cross entropy loss. λ1 represents The coefficient of L ξ represents the enhanced alignment loss. λ2 represents L ξ The coefficient of . is the weighted IOU loss. λ3 represents In order to balance the numerical values, we set λ1=λ2=λ3=1.
[0070] The embodiment of the present application can input the collected target image of the substation into a pre-trained detection model to obtain multi-scale features of different resolutions, perform feature fusion on the multi-scale features, perform feature enhancement on the highest-level features in the multi-scale features, and perform channel dimension reduction on the multi-scale features to obtain boundary guidance, global guidance, and trunk guidance. Finally, the boundary guidance, the global guidance, and the trunk guidance are alternately processed through a multi-guide cross-aggregation module to obtain the target area where the small animals are located in the target image, and gradually refine it in a cascade manner to obtain the target area where the small animals are located with higher accuracy and integrity. When the method provided in the embodiment of the present application is used to detect small animals in the substation, the detection efficiency and accuracy are improved.
[0071] Furthermore, the embodiments of the present application utilize pyramid transformers to extract image features, model global information, and are suitable for pixel-level detection tasks to generate feature maps of more scales. The boundary generation module uses bottom-up and top-down fusion methods to fuse adjacent layers, reducing the noise interference caused by cross-layer fusion. The feature enhancement module uses parallel normal convolution, hole convolution, and asymmetric convolution to expand the receptive field, capture information at different levels, and restore fine-grained information. The channel reduction module performs channel dimensionality reduction on the multi-scale features extracted by the backbone network to reduce computational costs. The multi-guided cross-aggregation module maximizes the use of previously learned features, using forward flow branches and reverse flow branches to alternately focus on target areas and background areas, refine target boundaries, explore potential features, and improve the detection accuracy of hidden small targets.
[0072] Figure 6 It is a structural schematic diagram of a small animal detection device in a substation provided by an exemplary embodiment of the present application. The device 600 includes: an acquisition module 601, which is used to acquire a target image of a substation; an input module 602, which is used to input the target image into a trained detection model; a detection module 603, which is used to extract multi-scale features of different resolutions of the target image through the pyramid transformer module of the detection model, perform feature fusion on the multi-scale features through the boundary generation module of the detection model to obtain the boundary guidance of the target image, perform feature enhancement on the highest level features in the multi-scale features through the feature enhancement module of the detection model to obtain the global guidance of the target image, and perform channel dimension reduction on the multi-scale features through the channel reduction module of the detection model to obtain the trunk guidance; a processing module 604, which is used to input the boundary guidance, the global guidance and the trunk guidance into the multi-guide cross aggregation module of the detection model for alternating attention processing, and predict the target area where the small animal is located in the target image.
[0073] The embodiment of the present application can input the collected target image of the substation into a pre-trained detection model to obtain multi-scale features of different resolutions, perform feature fusion on the multi-scale features, perform feature enhancement on the highest-level features in the multi-scale features, and perform channel dimension reduction on the multi-scale features to obtain boundary guidance, global guidance, and trunk guidance. Finally, the boundary guidance, the global guidance, and the trunk guidance are alternately processed through a multi-guide cross-aggregation module to obtain the target area where the small animals are located in the target image, and gradually refine it in a cascade manner to obtain the target area where the small animals are located with higher accuracy and integrity. When the method provided in the embodiment of the present application is used to detect small animals in the substation, the detection efficiency and accuracy are improved.
[0074] Optionally, the processing module 604 is further used to use the global guidance as a first fusion feature; input the boundary guidance, the trunk guidance and the first fusion feature into the multi-guided cross-aggregation module of the third layer of the detection model, and use the multi-guided cross-aggregation module of the third layer to alternately focus on the target area containing small animals and the background area not containing small animals in the target image to obtain the three-level features and the three-level prediction map of the target area; add the three-level prediction map to the global guidance to obtain the second fusion feature; input the boundary guidance, the trunk guidance, the second fusion feature and the three-level feature into the multi-guided cross-aggregation module of the second layer The cross aggregation module alternately focuses on the target area containing small animals and the background area not containing small animals in the target image to obtain the secondary features and the secondary prediction map of the target area; the secondary prediction map is added to the global guidance to obtain the third fusion feature; the boundary guidance, the trunk guidance, the third fusion feature and the secondary feature are input to the multi-guide cross aggregation module of the first layer to alternately focus on the target area containing small animals and the background area not containing small animals in the target image to obtain the primary features and the primary prediction map of the target area, and the primary prediction map includes the finally predicted image of the target area containing the small animal.
[0075] Optionally, the processing module 604 is further used to adjust the size of the primary prediction image to the size of the target image to obtain the detection image of the small animal.
[0076] Optionally, the processing module 604 is also used to upsample the first fusion feature, group and fuse it with the trunk guide, and then multiply it with a first predetermined parameter to obtain a forward flow branch result; upsample the first fusion feature, perform reverse edge integration, use pooling operation to focus on the boundary of the background area in the target image, obtain a reverse image, group and fuse the reverse image with the trunk guide, and then multiply it with a second predetermined parameter to obtain a reverse flow branch result; subtract the first fusion feature from the forward flow branch result, and then add it to the reverse flow branch result to obtain a fourth fusion feature of the fused two branches; add the boundary guidance to the fourth fusion feature to obtain the third-level feature; input the third-level feature into the convolution layer of the multi-guided cross aggregation module of the third layer for channel compression to obtain the third-level prediction map.
[0077] Optionally, the processing module 604 is also used to input the multi-scale features into multiple first convolutional layers of the boundary generation module, and perform bottom-up adjacent convolutional layer fusion on the multi-scale features to obtain intermediate features of the multi-scale features; input the intermediate features into multiple first convolutional layers of the boundary generation module again, and perform bottom-up adjacent convolutional layer fusion on the intermediate features to obtain fourth fused features that fuse the multi-scale features; input the fourth fused features into the second convolutional layer for channel dimensionality reduction to obtain the boundary guidance.
[0078] Optionally, the processing module 604 is also used to input the highest-level feature into the hole convolution branch of the feature enhancement module to expand the receptive field to obtain a first processing feature, input the highest-level feature into the asymmetric convolution branch of the feature enhancement module to increase the inference speed to obtain a second processing feature, input the highest-level feature into the normal convolution branch of the feature enhancement module to improve the diversity of multi-level features to obtain a third processing feature; connect the first processing feature, the second processing feature and the third processing feature and perform channel dimensionality reduction to obtain a fourth processing feature; input the highest-level feature into the residual branch of the feature enhancement module to perform channel dimensionality reduction to obtain a fifth processing feature; add the fourth processing feature and the fifth processing feature and activate them through an activation function to obtain the global guidance.
[0079] Optionally, the processing module 604 is also used to input the multi-scale feature into the third convolutional layer of the channel reduction module to perform channel dimension reduction to obtain a first feature; input the first feature into the fourth convolutional layer of the channel reduction module to remove noise interference to obtain the trunk guidance.
[0080] Optionally, the loss function of the detection model is a hybrid loss function consisting of a two-level cross entropy loss function, an enhanced alignment loss function, and a weighted IOU loss function.
[0081] The device 600 provided in the embodiment of the present application can execute each method in the foregoing method embodiment and realize the functions and beneficial effects of each method in the foregoing method embodiment, which will not be repeated here.
[0082] Figure 7 It is one of the structural diagrams of an electronic device provided by an exemplary embodiment of the present application. Referring to the figure, at the hardware level, the electronic device in the substation includes a processor, and optionally, an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage, etc. Of course, the electronic device may also include hardware required for other services.
[0083] The processor, network interface and memory can be interconnected through an internal bus, which can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0084] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0085] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a device for locating the specified user at the logical level. The processor executes the program stored in the memory and is specifically used to execute: Figure 1 The method disclosed in the illustrated embodiment implements the functions and beneficial effects of each method in the previous method embodiments, which will not be described in detail here.
[0086] The above application Figure 1The method disclosed in the illustrated embodiment can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0087] The small animal detection device in the substation can also execute the methods in the foregoing method embodiments and realize the functions and beneficial effects of the methods in the foregoing method embodiments, which will not be described in detail here.
[0088] Of course, in addition to software implementation methods, the small animal detection device in the substation of the present application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0089] The present application also provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple application programs, the electronic device executes Figure 1 The method disclosed in the illustrated embodiment implements the functions and beneficial effects of each method in the previous method embodiments, which will not be described in detail here.
[0090] The computer-readable storage medium includes a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0091] Furthermore, an embodiment of the present application also provides a computer program product, the computer program product includes a computer program stored on a non-transitory computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the following process is implemented: Figure 1 The method disclosed in the illustrated embodiment implements the functions and beneficial effects of each method in the previous method embodiments, which will not be described in detail here.
[0092] In short, the above are only preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
[0093] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0094] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0095] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0096] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
Claims
1. A method for detecting small animals in a substation, characterized in that: include: Acquire target images of substations; Inputting the target image into a trained detection model; Extracting multi-scale features of different resolutions of the target image through the pyramid transformer module of the detection model, performing feature fusion on the multi-scale features through the boundary generation module of the detection model to obtain boundary guidance of the target image, performing feature enhancement on the highest-level features in the multi-scale features through the feature enhancement module of the detection model to obtain global guidance of the target image, and performing channel dimension reduction on the multi-scale features through the channel reduction module of the detection model to obtain trunk guidance; The boundary guidance, the global guidance and the trunk guidance are input into the multi-guide cross-aggregation module of the detection model for alternating attention processing, so as to predict the target area where the small animal in the target image is located.
2. The method for detecting small animals in a substation according to claim 1, characterized in that: The boundary guidance, the global guidance and the trunk guidance are input into the multi-guide cross-aggregation module of the detection model for alternate attention processing, and the target area where the small animal in the target image is predicted to be located includes: Using the global guide as a first fusion feature; Inputting the boundary guide, the trunk guide and the first fusion feature into a multi-guide cross-aggregation module of the third layer of the detection model, and alternately focusing on a target area containing small animals and a background area not containing small animals in the target image through the multi-guide cross-aggregation module of the third layer, to obtain a three-level feature and a three-level prediction map of the target area; Adding the three-level prediction map to the global guide to obtain a second fusion feature; Inputting the boundary guidance, the trunk guidance, the second fusion feature and the third-level feature into a multi-guide cross-aggregation module of the second layer to alternately focus on the target area containing the small animal and the background area not containing the small animal in the target image, so as to obtain the secondary feature and the secondary prediction map of the target area; Adding the secondary prediction map to the global guide to obtain a third fusion feature; The boundary guidance, the trunk guidance, the third fusion feature and the secondary feature are input into a multi-guided cross-aggregation module of the first layer, which alternately focuses on the target area containing the small animal and the background area not containing the small animal in the target image, and obtains the primary feature and the primary prediction map of the target area, wherein the primary prediction map includes the finally predicted image of the target area containing the small animal.
3. The method for detecting small animals in a substation according to claim 2, characterized in that: After inputting the boundary guidance, the global guidance and the trunk guidance into the multi-guide cross-aggregation module of the detection model for alternate attention processing and predicting the target area where the small animal is located in the target image, the method further includes: The size of the primary prediction image is adjusted to the size of the target image to obtain a detection image of the small animal.
4. The method for detecting small animals in a substation according to claim 2, characterized in that: The boundary guide, the trunk guide and the first fusion feature are input into the multi-guide cross aggregation module of the third layer of the detection model, and the target area containing small animals and the background area not containing small animals in the target image are alternately focused on by the multi-guide cross aggregation module of the third layer, so as to obtain the three-level features and the three-level prediction map of the target area, including: The first fusion feature is upsampled, grouped and fused with the trunk guide, and then multiplied by a first predetermined parameter to obtain a forward flow branch result; After upsampling the first fusion feature, reverse edge integration is performed, and a pooling operation is used to focus on the boundary of the background area in the target image to obtain a reverse map, and the reverse map is grouped and fused with the trunk guide and then multiplied with a second predetermined parameter to obtain a reverse flow branch result; Subtracting the first fusion feature from the forward flow branch result and adding the result to the reverse flow branch result to obtain a fourth fusion feature of the fused double branches; Adding the boundary guidance to the fourth fusion feature to obtain the third-level feature; The three-level features are input into the convolution layer of the multi-guided cross aggregation module of the third layer for channel compression to obtain the three-level prediction map.
5. The method for detecting small animals in a substation according to claim 1, characterized in that: The step of fusing the multi-scale features through the boundary generation module of the detection model to obtain the boundary guidance of the target image includes: Inputting the multi-scale features into a plurality of first convolutional layers of a boundary generation module, and fusing the multi-scale features from bottom to top in adjacent convolutional layers to obtain intermediate features of the multi-scale features; Inputting the intermediate features into the plurality of first convolutional layers of the boundary generation module again, and fusing the intermediate features from bottom to top in adjacent convolutional layers to obtain fourth fused features that fuse the multi-scale features; The fourth fusion feature is input into the second convolutional layer to perform channel dimension reduction to obtain the boundary guidance.
6. The method for detecting small animals in a substation according to claim 1, characterized in that: The step of performing feature enhancement on the highest-level features in the multi-scale features through the feature enhancement module of the detection model to obtain the global guidance of the target image includes: Input the highest level feature to the hole convolution branch of the feature enhancement module to expand the receptive field, obtain the first processing feature, input the highest level feature to the asymmetric convolution branch of the feature enhancement module to increase the inference speed, obtain the second processing feature, input the highest level feature to the normal convolution branch of the feature enhancement module to improve the diversity of multi-level features, and obtain the third processing feature; The first processing feature, the second processing feature and the third processing feature are connected and then channel dimension reduction is performed to obtain a fourth processing feature; Inputting the highest level feature into the residual branch of the feature enhancement module to perform channel dimension reduction to obtain a fifth processing feature; The fourth processing feature and the fifth processing feature are added together and then activated by an activation function to obtain the global guidance.
7. The method for detecting small animals in a substation according to claim 1, characterized in that: The step of performing channel dimension reduction on the multi-scale features through the channel reduction module of the detection model to obtain a trunk guide comprises: Inputting the multi-scale features into the third convolutional layer of the channel reduction module to perform channel dimension reduction to obtain a first feature; The first feature is input into the fourth convolutional layer of the channel reduction module to remove noise interference and obtain the trunk guidance.
8. The method for detecting small animals in a substation according to claim 1, characterized in that: The loss function of the detection model is a hybrid loss function consisting of a two-level cross entropy loss function, an enhanced alignment loss function, and a weighted IOU loss function.
9. A small animal detection device in a substation, characterized in that: include: An acquisition module, used for acquiring a target image of a substation; An input module, used for inputting the target image into a trained detection model; A detection module, used to extract multi-scale features of different resolutions of the target image through the pyramid transformer module of the detection model, perform feature fusion on the multi-scale features through the boundary generation module of the detection model to obtain boundary guidance of the target image, perform feature enhancement on the highest-level features in the multi-scale features through the feature enhancement module of the detection model to obtain global guidance of the target image, and perform channel dimension reduction on the multi-scale features through the channel reduction module of the detection model to obtain trunk guidance; A processing module is used to input the boundary guidance, the global guidance and the trunk guidance into the multi-guide cross-aggregation module of the detection model for alternating attention processing, so as to predict the target area where the small animal is located in the target image.
10. An electronic device, characterized in that: The electronic device comprises: a processor and a memory; wherein the memory is used to store a computer program that can be run on the processor; the processor is used to execute the program stored in the memory to implement the steps of the small animal detection method in the substation as described in any one of claims 1-8.