Detection model training method and detection method for easily floating objects along power transmission channel
By using a combination method of deep feature extraction and feature attention modules in the remote sensing image of the transmission channel, the problem of low accuracy in detection of easy floating objects in the prior art is solved, and higher detection accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510218538.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-26
AI Technical Summary
The prior art is difficult to accurately match the target boundaries of floating objects in the remote sensing image of the transmission channel, resulting in low accuracy of detection results.
A training method for easy-to-float detection model along the transmission channel was designed. By acquiring historical remote sensing images, the deep feature extraction module of easy-to-float feature and the attention module of easy-to-float feature are used to extract the features of easy-to-float feature of different scales, and the feature representations of the easy-to-float feature are enhanced through pooling operations and channel weighting. The detection model parameters are finally iteratively updated to improve the detection accuracy.
Through this method, the accuracy and efficiency of floating objects detection are significantly improved, the possibility of missed detection is reduced, and the model's adaptability to complex scenarios is enhanced.
Smart Images

Figure CN120032115A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, computer equipment, computer-readable storage medium and computer program product for training a model for detecting floating objects along a power transmission channel, and a method, device, computer equipment, computer-readable storage medium and computer program product for detecting floating objects along a power transmission channel. Background Art
[0002] Floating objects along the transmission channels, such as plastic greenhouses, ground films and dust-proof green nets, are easily blown up and attached to transmission lines under severe weather conditions such as strong winds and heavy rains, causing transmission line short circuits, tripping and other faults, seriously affecting the stable operation of the power system and creating major safety hazards in the transmission channels.
[0003] With the rapid development of remote sensing technology, deep learning algorithms are used in existing technologies to analyze and process remote sensing images of transmission channels to achieve automatic identification and detection of floating objects.
[0004] However, since floating objects such as plastic greenhouses and ground films may present complex shapes and textures in remote sensing images, it is difficult to accurately match the target boundaries in the detection of floating objects in remote sensing images of power transmission channels in the existing technology. Therefore, the detection results of floating objects are not accurate. Summary of the invention
[0005] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product for training a floating object detection model along a transmission channel that can improve the accuracy of floating object detection in response to the above-mentioned technical problems.
[0006] In a first aspect, the present application provides a method for training a model for detecting floating objects along a power transmission channel, comprising:
[0007] Acquire remote sensing images along historical power transmission channels, wherein the remote sensing images along historical power transmission channels contain category labels and location information of real frames of floating objects;
[0008] The remote sensing images along the historical power transmission channel are input into the constructed initial floating object detection model. The initial floating object detection model extracts the floating object features of different scales from the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates a transmission channel floating object feature map based on the floating object features of different scales.
[0009] The easy-to-float feature map of the transmission channel is subjected to a pooling operation and a channel weighting operation through a easy-to-float feature attention module, so as to capture the easy-to-float feature and enhance the easy-to-float feature representation in the easy-to-float feature map of the transmission channel, and obtain an easy-to-float feature enhanced map of the transmission channel;
[0010] Based on the features of the floatable objects at different scales and the enhanced map of the features of the floatable objects along the transmission channel, determining the category, location information and confidence of the prediction box, wherein the prediction box is used to identify the floatable objects in the remote sensing images along the historical transmission channel, and the initial floatable object detection model includes a deep feature extraction module for the floatable objects and an attention module for the features of the floatable objects;
[0011] Determine the error between the real frame and the predicted frame according to the category label and position information of the real frame, and the category, position information and confidence of the predicted frame;
[0012] Based on the error, the parameters of the initial easy-to-float object detection model are iteratively updated until a preset training end condition is reached to obtain a trained easy-to-float object detection model.
[0013] In the second aspect, the present application also provides a method for detecting floating objects along a power transmission channel, comprising:
[0014] Acquire remote sensing images along the transmission corridor;
[0015] Taking the remote sensing image along the power transmission channel as input, calling the trained floating object detection model to obtain the floating object detection result, wherein the floating object detection model is trained based on the above-mentioned floating object detection model training method along the power transmission channel;
[0016] An early warning is issued based on the detection result of the easily floating objects.
[0017] In a third aspect, the present application also provides a model training device for detecting floating objects along a power transmission channel, comprising:
[0018] An image acquisition module is used to acquire remote sensing images along the historical power transmission channel, wherein the remote sensing images along the historical power transmission channel contain category labels and location information of real frames of floating objects;
[0019] A model training module is used to input the remote sensing images along the historical power transmission channel into the constructed initial floating object detection model, wherein the initial floating object detection model extracts floating object features of different scales from the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates a power transmission channel floating object feature map based on the floating object features of different scales; the floating object feature attention module performs pooling operations and channel weighting operations on the floating object feature map of the transmission channel, captures the floating object features and enhances the floating object feature representation in the floating object feature map of the transmission channel, and obtains the floating object feature enhancement map of the transmission channel; based on the floating object features of different scales and the floating object feature enhancement map of the transmission channel, determines the category, location information and confidence of the prediction box, wherein the prediction box is used to identify the floating objects in the remote sensing images along the historical power transmission channel, and the initial floating object detection model includes the floating object deep feature extraction module and the floating object feature attention module;
[0020] An error determination module, configured to determine an error between the real frame and the predicted frame according to the category label and position information of the real frame, and the category, position information and confidence of the predicted frame;
[0021] The parameter updating module is used to iteratively update the parameters of the initial floating object detection model until a preset training end condition is reached to obtain a trained floating object detection model.
[0022] In a fourth aspect, the present application also provides a device for detecting floating objects along a power transmission channel, comprising:
[0023] A remote sensing image acquisition module is used to acquire remote sensing images along the power transmission channel;
[0024] A floating object detection module, used to use the remote sensing image along the power transmission channel as input, call a trained floating object detection model, and obtain a floating object detection result, wherein the floating object detection model is trained based on the above-mentioned floating object detection model training method along the power transmission channel;
[0025] The early warning module is used to issue an early warning based on the detection result of the floating object.
[0026] In a fifth aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps in any one of the above-mentioned steps in the training method for detecting floating objects along a transmission channel and the steps in the method for detecting floating objects along a transmission channel.
[0027] In a sixth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in any one of the above-mentioned steps in an embodiment of a method for training a model for detecting floating objects along a transmission channel, and an embodiment of a method for detecting floating objects along a transmission channel.
[0028] In a seventh aspect, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in any one of the above-mentioned steps in the training method for detecting floating objects along a transmission channel and the steps in the method for detecting floating objects along a transmission channel.
[0029] The above-mentioned method, device, computer equipment, computer-readable storage medium and computer program product for the detection model of floating objects along the transmission channel take into account the influence of the complex morphology and texture of floating objects along the transmission channel in remote sensing images on the detection accuracy, design a floating object deep feature extraction module, extract the deep features of floating objects of different scales in remote sensing images, generate a floating object feature map of the transmission channel based on the floating object features of different scales, which is conducive to enhancing the robustness and detection accuracy of the floating object detection model. At the same time, considering that floating objects may appear as small-sized or low-contrast targets in remote sensing images, such as thin materials such as ground film and dust-proof green net, and are easily blocked by trees and buildings, or confused with the background, design a floating object feature attention module, and perform pooling operations and channel weighting operations on the floating object feature map of the transmission channel through the floating object feature attention module, capture the floating object features in the floating object feature map of the transmission channel and enhance the floating object feature representation, and obtain the floating object feature enhancement map of the transmission channel, which is conducive to improving the detection accuracy of the model and reducing the possibility of missed detection. Thus, according to the features of floating objects at different scales and the enhanced map of floating objects in the transmission channel, the category, location information and confidence of the prediction box are determined, and then according to the category label and location information of the real box, as well as the category, location information and confidence of the prediction box, the error between the real box and the prediction box is determined, and the model parameters are iteratively updated according to the error until the preset training end condition is reached, and the trained floating object detection model is obtained. Detection using the trained floating object detection model is conducive to improving the accuracy of floating object detection in different detection scenarios along the transmission channel, as well as the model's adaptability to high-precision floating object detection in different complex scenarios.
[0030] The above-mentioned method, device, computer equipment, computer calibration storage medium and computer program product for detecting floating objects along the transmission channel obtain remote sensing images along the transmission channel, input the images into the trained floating object detection model obtained by the above-mentioned training method for detecting floating objects along the transmission channel, obtain floating object detection results, and issue early warnings based on the floating object detection results. On the one hand, the accuracy and efficiency of floating detection are improved through model detection. On the other hand, early warnings are issued based on the rapid and accurate detection results of floating objects along the transmission channel based on the model, which is conducive to timely discovery of potential hidden dangers of the power system, so as to take corresponding measures to reduce the possibility of accidents and improve the stability and safety of power system operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 A diagram showing an application environment of a method for training a model for detecting floating objects along a power transmission channel in one embodiment;
[0033] Figure 2 A schematic diagram of a flow chart of a method for training a model for detecting floating objects along a power transmission channel in one embodiment;
[0034] Figure 3 It is a structural block diagram of a model for detecting floating objects along a power transmission channel in one embodiment;
[0035] Figure 4 It is a structural block diagram of a deep feature extraction module for easily floating objects in one embodiment;
[0036] Figure 5 is a schematic diagram of the structure of a floating object feature attention module in one embodiment;
[0037] Figure 6 A schematic diagram of a method for detecting floating objects along a power transmission channel;
[0038] Figure 7 It is a structural block diagram of a model training device for detecting floating objects along a power transmission channel in one embodiment;
[0039] Figure 8 It is a structural block diagram of a device for detecting floating objects along a power transmission channel in one embodiment;
[0040] Fig. 9FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0042] The method for training a model for detecting floating objects along a power transmission channel provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers.
[0043] Specifically, the operator may upload the collected remote sensing images along the historical power transmission channel to the server 104 through the terminal 102, and then send the model training message to the server 104 through the terminal 102, and the server 104 obtains the remote sensing images along the power transmission channel. Secondly, the remote sensing images along the historical power transmission channel are input into the constructed initial floating object detection model, and the initial floating object detection model extracts the floating object features of different scales of the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates a transmission channel floating object feature map based on the floating object features of different scales. Then, the floating object feature attention module performs pooling operations and channel weighting operations on the transmission channel floating object feature map to capture the floating object features in the transmission channel floating object feature map. The easy-to-float object features and enhanced easy-to-float object feature representations are obtained to obtain an easy-to-float object feature enhancement map of the transmission channel. Again, based on the easy-to-float object features of different scales and the easy-to-float object feature enhancement map of the transmission channel, the category, position information and confidence of the prediction frame are determined. The prediction frame is used to identify the easy-to-float objects in the remote sensing images along the historical transmission channel. The initial easy-to-float object detection model includes an easy-to-float object deep feature extraction module and an easy-to-float object feature attention module. Finally, according to the category label and position information of the real frame, and the category, position information and confidence of the prediction frame, the error between the real frame and the prediction frame is determined. Based on the error, the parameters of the initial easy-to-float object detection model are iteratively updated until a preset training end condition is reached to obtain a trained easy-to-float object detection model.
[0044] The terminal 102 may be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart car devices, projection devices, etc. Portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. Head-mounted devices may be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 104 may be an independent physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.
[0045] In an exemplary embodiment, Figure 2 As shown in the figure, a method for training a model for detecting floating objects along a power transmission channel is provided. Figure 1 The server 104 in the example is used for explanation, and includes the following S100 to S600. Among them:
[0046] S100, obtaining remote sensing images along a historical power transmission channel, wherein the remote sensing images along the historical power transmission channel contain category labels and location information of real frames of floating objects.
[0047] Among them, the remote sensing images along the historical transmission channel can be images obtained by collecting images of the area along the transmission channel during the historical time period. Floating objects along the transmission channel may refer to objects that are easily moved or floated under the action of wind or other natural forces and may pose a threat to the transmission line. Floating objects include but are not limited to plastic greenhouses, ground films, and dust-proof green nets. When floating objects come into contact with power lines, they may cause power failures such as short circuits and tripping, affecting the stable operation and safety of the power system. The category labels of the real frame may include floating objects and background. The location information may include the upper left corner coordinates and the lower right corner coordinates of the real frame.
[0048] In practical applications, it is possible to collect images of areas along the transmission channel through UAV remote sensing. Acquiring historical transmission channel remote sensing images can be to obtain multiple historical transmission channel remote sensing images containing floating objects collected by UAV remote sensing during a historical time period. In order to facilitate model processing, the size of these remote sensing images can be standardized. Specifically, the remote sensing images along the historical transmission channel are cropped and resized to make the image size uniform. For these historical transmission channel remote sensing images, use the annotation tool to select the location of the floating objects, that is, determine the coordinates of the upper left corner vertex and the lower right corner vertex of the box, and annotate them with category labels.
[0049] In other embodiments, in order to enhance the performance of the model, data enhancement processing is performed on the remote sensing images along the historical power transmission channel. The data enhancement processing includes but is not limited to rotation processing, flipping processing, scaling and cropping processing, brightness adjustment processing and noise addition processing. After the data enhancement processing, the remote sensing images along the historical power transmission channel after data enhancement processing can also be divided into a training set, a verification set and a test set in a ratio of 8:1:1. The remote sensing images along the historical power transmission channel in the training set are used to train the initial floating object detection model.
[0050] S200, inputting the remote sensing images along the historical power transmission channel into the constructed initial floating object detection model, the initial floating object detection model extracts the floating object features of different scales from the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates a power transmission channel floating object feature map based on the floating object features of different scales.
[0051] The initial floating object detection model is improved based on the pre-training model. The deep feature extraction module of floating objects includes multiple feature processing branches for extracting deep features of floating objects.
[0052] In practical applications, in order to improve the accuracy of capturing the features of floating objects in remote sensing images, this application designs a deep feature extraction module for floating objects that includes multiple feature processing branches, which is used to extract features of different dimensions of floating objects in remote sensing images. Through the designed deep feature extraction module for floating objects, the architecture of the pre-trained model (such as the YOLO model) is improved to obtain an initial floating object detection model. Taking the remote sensing images along the historical power transmission channel as input, multi-dimensional feature extraction and feature fusion operations are performed through multiple feature processing branches to extract the features of floating objects at different scales in the remote sensing images along the historical power transmission channel, and feature enhancement, further feature extraction and feature fusion processing are performed on the features of floating objects at different scales to generate a feature map of floating objects in the transmission channel.
[0053] S300, performing a pooling operation and a channel weighting operation on the easy-to-float feature map of the transmission channel through a easy-to-float feature attention module, capturing the easy-to-float features and enhancing the easy-to-float feature representation in the easy-to-float feature map of the transmission channel, and obtaining an easy-to-float feature enhanced map of the transmission channel.
[0054] Among them, the easy-to-float feature attention module is used to enhance the easy-to-float feature representation.
[0055] In practical applications, considering that floating objects may appear as small or low-contrast targets in remote sensing images, such as thin materials such as ground film and dust-proof green nets, and are easily blocked by trees and buildings, or confused with the background, a floating object feature attention module is designed. The architecture of the pre-trained model (such as the YOLO model) is improved through the floating object deep feature extraction module and the floating object feature attention module to obtain the initial floating object detection model. Figure 3 As shown. Initialize all neural network parameters of the initial easy-to-float object detection model, and set the hyperparameters related to the initial easy-to-float object detection model, such as: training rounds, batch size, optimizer selection, learning rate, etc. Perform pooling operations (such as global pooling and maximum pooling) on the easy-to-float object feature map of the transmission channel through the easy-to-float feature attention module to capture the easy-to-float feature of the input feature map, enhance the easy-to-float feature representation through channel weighted operations, and obtain the easy-to-float feature enhancement map of the transmission channel.
[0056] S400, based on the features of the floatable objects at different scales and the enhanced map of the features of the floatable objects in the transmission channel, determining the category, location information and confidence of a prediction box, wherein the prediction box is used to identify the floatable objects in the remote sensing images along the historical transmission channel, and the initial floatable object detection model includes a floatable object deep feature extraction module and a floatable object feature attention module.
[0057] In practical applications, by performing feature extraction and feature fusion on the features of floating objects at different scales and the enhanced map of floating objects in the transmission channel, a prediction box is generated to identify the position of floating objects in the remote sensing image, and the category, location information and confidence of the prediction box are output.
[0058] S500, determining an error between the real frame and the predicted frame according to the category label and position information of the real frame, and the category, position information and confidence of the predicted frame.
[0059] In practical applications, the intersection-and-union ratio between the predicted box and the real box can be determined based on the position information of the predicted box and the position information of the real box, and the difference between the position information of the predicted box and the real box, the difference between the category of the predicted box and the real category, the confidence of the predicted box and the intersection-and-union ratio can be calculated through the loss function to determine the error between the predicted box and the real box. The loss function includes but is not limited to the mean square error loss function, the cross entropy loss function and the smooth L1 loss function.
[0060] S600, iteratively updating the parameters of the initial floating object detection model based on the error until a preset training end condition is reached to obtain a trained floating object detection model.
[0061] In practical applications, after obtaining the error between the real frame and the predicted frame through the loss function, back propagation is started from the output layer of the initial floating object detection model according to the loss function value (error), the parameter gradient value of each layer is calculated, and the parameters of the initial floating object detection model are updated to minimize the loss function value. The preset training end condition can be that the loss function value is continuously less than the preset loss threshold within the preset number of iterations, the training of the initial floating object detection model is stopped, and the trained floating object detection model is obtained.
[0062] In other embodiments, after completing a round of training on the images in the training set, the model performance is verified based on the historical transmission channel remote sensing images in the verification set, and the training strategy is adjusted based on the verification results, such as terminating the training early or adjusting the learning rate.
[0063] In the above-mentioned method for training the detection model of floating objects along the transmission channel, considering the influence of the complex morphology and texture of floating objects along the transmission channel in the remote sensing image on the detection accuracy, a deep feature extraction module of floating objects is designed to extract the deep features of floating objects of different scales in the remote sensing image, and generate the floating object feature map of the transmission channel according to the floating object features of different scales, which is conducive to enhancing the robustness and detection accuracy of the floating object detection model. At the same time, considering that floating objects may appear as small-sized or low-contrast targets in remote sensing images, such as thin materials such as ground film and dust-proof green net, and are easily blocked by trees and buildings, or confused with the background, a floating object feature attention module is designed. Through the floating object feature attention module, the floating object feature map of the transmission channel is pooled and channel weighted, the floating object features in the floating object feature map of the transmission channel are captured and the floating object feature representation is enhanced, and the floating object feature enhancement map of the transmission channel is obtained, which is conducive to improving the detection accuracy of the model and reducing the possibility of missed detection. Thus, according to the features of floating objects at different scales and the enhanced map of floating objects in the transmission channel, the category, location information and confidence of the prediction box are determined, and then according to the category label and location information of the real box, as well as the category, location information and confidence of the prediction box, the error between the real box and the prediction box is determined, and the model parameters are iteratively updated according to the error until the preset training end condition is reached, and the trained floating object detection model is obtained. Detection using the trained floating object detection model is conducive to improving the accuracy of floating object detection in different detection scenarios along the transmission channel, as well as the model's adaptability to high-precision floating object detection in different complex scenarios.
[0064] In an exemplary embodiment, the initial floating object detection model includes a feature extraction module, such as Figure 3 As shown, S200 includes S220 to S260. Among them:
[0065] S220, extracting features of the remote sensing images along the historical power transmission channel through a feature extraction module to obtain a first feature map.
[0066] Among them, the first feature map is a feature map obtained by the feature extraction module performing feature extraction operations on remote sensing images along the historical power transmission channel.
[0067] In this embodiment, the feature extraction module is a Conv_BN_ReLu module (i.e. Figure 3 In practical applications, the Conv_BN_ReLu module extracts features from the input remote sensing image along the historical power transmission channel through a sliding convolution kernel to generate the first feature map.
[0068] S240, performing convolution operation, channel weighting operation and pooling operation on the first feature map respectively through the feature dimension reduction branch, the multi-scale feature fusion branch and the feature pooling branch in the deep feature extraction module of the easy-to-float object, to obtain a second feature map.
[0069] The feature dimension reduction branch may include multiple convolution modules for performing feature dimension reduction on the input feature map. The multi-scale feature fusion branch may include a pooling layer and a feature fusion module for fusing multi-scale features. The feature pooling branch may include different pooling layers (such as a maximum pooling layer and an average pooling layer). The second feature map is a feature map obtained by extracting features from the first feature map by the deep feature extraction module for floating objects.
[0070] In practical applications, the first feature map may be input into modules in a feature dimension reduction branch, a multi-scale feature fusion branch, and a feature pooling branch respectively: a convolution operation is performed on the first feature map through multiple convolution modules in the feature dimension reduction branch, a pooling operation is performed on the first feature map through a pooling layer in a multi-scale feature fusion branch, and feature fusion is performed on the feature map obtained after the pooling operation and the feature map obtained through the convolution operation of the feature dimension reduction branch, and a pooling operation is performed on the first feature map through different pooling layers in the feature pooling branch, and feature fusion is performed on the feature maps obtained from different feature processing branches to obtain a second feature map.
[0071] S260, based on the feature extraction module and the easy-to-float deep feature extraction module, extract the features of the second feature map at different scales, and obtain the easy-to-float feature map of the transmission channel based on the features at different scales.
[0072] In practical applications, the feature extraction module (Conv_BN_ReLu module) and the easy-to-float deep feature extraction module may be stacked in sequence for multiple times. Through multiple convolution operations of the stacked feature extraction modules and the easy-to-float deep feature extraction module, the features of the second feature map at different scales are extracted, and convolution operations and feature fusion operations are performed on the features of different scales to obtain the easy-to-float feature map of the transmission channel.
[0073] In this embodiment, the input feature map is processed by the feature dimension reduction branch, the multi-scale feature fusion branch and the feature pooling branch, which is beneficial to enhancing the robustness of the model and the feature representation capability of floating objects, thereby facilitating improving the detection accuracy of floating objects.
[0074] In an exemplary embodiment, S240 includes S241 to S245. Among them:
[0075] S241, performing a dilated convolution on the first feature map through a dilated convolution layer in a deep feature extraction module for floating objects to obtain a first deep feature map.
[0076] The first deep feature map is a feature map obtained by performing a dilated convolution operation on the first deep feature map by the dilated convolution layer in the deep feature extraction module for floating objects.
[0077] In practical applications, this application designs an Easily Floating Objects Deep-feature Extraction module EFODE (Easily Floating Objects Deep-feature Extraction module). The structure of the Easily Floating Objects Deep-feature Extraction module is as follows: Figure 4 As shown in the figure, the first feature map Q1 is input into a dilated convolution layer with a convolution kernel size of 3×3 and a dilation rate of 7. The dilated convolution layer performs a dilated convolution operation on the first feature map to expand the receptive field and capture a wider range of contextual information, thereby obtaining the first deep feature map Q2.
[0078] S242, performing convolution operation, batch normalization operation, nonlinear processing and convolution operation on the first deep feature map in sequence through the feature dimension reduction branch to obtain a second deep feature map.
[0079] Among them, the feature dimension reduction branch includes a series of convolutions and a convolution layer of size 1×1, a batch normalization layer ( Figure 4 The convolutional layers are constructed with BN batch normalization in ), ReLu activation function, and a convolutional kernel size of 1×1.
[0080] In practical applications, through the series-connected convolutional layer, batch normalization layer, ReLu activation function and convolutional layer in the feature dimension reduction branch, the first deep feature map Q2 is sequentially subjected to convolution operations, batch normalization operations, nonlinear processing and convolution operations to achieve feature dimension reduction and nonlinear enhancement, and obtain the second deep feature map Q3.
[0081] S243, performing a pooling operation on the first feature map through a multi-scale feature fusion branch to obtain a third deep feature map, performing a channel weighted operation on the third deep feature map and the second deep feature map through a convolution layer to obtain a fourth deep feature map, and performing a convolution operation on the fourth deep feature map to obtain a fifth deep feature map.
[0082] Among them, the multi-scale feature fusion branch includes a maximum pooling layer, a channel weighted module, and a convolution layer with a convolution kernel size of 1×1.
[0083] In practical applications, the input first feature map Q2 is processed by the multi-scale feature fusion branch. First, the first feature map Q2 is pooled through the maximum pooling layer to obtain the third deep feature map Q4; the second deep feature map Q3 and the third feature map Q4 are channel-weighted to obtain the fourth deep feature map Q5, and then the fourth deep feature map is convolved with a convolution kernel size of 1×1 to obtain the fifth deep feature map Q6.
[0084] S244, performing maximum pooling operation and average pooling operation on the first feature map respectively through the feature pooling branch, performing feature fusion on the results of the maximum pooling operation and the average pooling operation to obtain a sixth deep feature map, performing channel weighting operation on the sixth deep feature map to obtain a seventh deep feature map.
[0085] Among them, the feature pooling branch includes a series of maximum pooling layers and Relu activation functions, a series of average pooling layers and Relu activation functions, a feature fusion module, and a channel weighting module.
[0086] In practical applications, the first feature map Q2 is pooled and processed nonlinearly by the maximum pooling and Relu activation function connected in series in the feature pooling branch, and the first feature map Q2 is averaged and processed nonlinearly by the average pooling layer and Relu activation function connected in series. The results of the two processes are subjected to channel dimension feature fusion to obtain the sixth deep feature map Q7. In this way, the feature representation is enriched through multiple pooling methods, and the fifth deep feature map Q6 and the sixth deep feature map Q7 are subjected to channel weighted operation to obtain the seventh deep feature map Q8.
[0087] S245, performing feature fusion and convolution operations on the second deep feature map and the seventh deep feature map to obtain a second feature map.
[0088] In practical applications, the second deep feature map and the seventh deep feature map are input into the series-connected feature fusion module and the convolution module with a convolution kernel size of 1×1 for feature fusion and convolution operations to obtain the second feature map Q9 (i.e. Figure 3 feature map X4 in .
[0089] In this embodiment, a feature fusion and weighting method is designed and applied to the deep feature extraction module of easy-to-float objects. The receptive field is effectively expanded through the dilated convolution to capture a wide range of contextual information. The feature dimension reduction branch uses convolution to perform feature dimension reduction and nonlinear enhancement; the multi-scale feature fusion branch fuses the feature maps of easy-to-float objects at different scales to enhance the robustness of the model; the feature pooling branch introduces multiple pooling methods to enrich the feature representation. In addition, the channel weighting and fusion operations in the deep feature extraction module of easy-to-float objects further enhance the feature extraction capability of easy-to-float objects.
[0090] In an exemplary embodiment, the number of the deep feature extraction module and the feature extraction module is multiple, and S260 includes S261 to S265. Among them:
[0091] S261, respectively connecting the output end of each feature extraction module with the input end of each deep feature extraction module of the floating object to obtain a first feature extraction combination module, a second feature extraction combination module and a third feature extraction combination module.
[0092] S262, extracting features of the third feature map through the first feature extraction and combination module to obtain a first scale feature map, extracting features of the first scale feature map through the second feature extraction and combination module to obtain a second scale feature map, and extracting features of the second scale feature map through the third feature extraction and combination module to obtain a third scale feature map.
[0093] S263, extracting features from the third scale feature map through a deep feature extraction module for floating objects, and performing a pooling operation on the extracted features to obtain a fourth scale feature map.
[0094] S264, by upsampling the fourth scale feature map, performing feature fusion on the upsampled feature map and the second scale feature map to obtain a feature fusion map, and performing feature extraction on the feature map through a deep feature extraction module for floating objects to obtain a fifth scale feature map.
[0095] S265, performing an upsampling operation on the fifth-size feature map, and performing feature fusion on the feature map after the upsampling operation and the first-scale feature map to obtain a transmission channel easy-to-float feature map.
[0096] In practical applications, such as Figure 3As shown, in the initial floating object detection model, the output end of the feature extraction module is connected to the input end of the floating object feature extraction module, and the connected feature extraction module and the floating object feature extraction module are stacked three times to obtain a first feature extraction combination module, a second feature extraction combination module and a third feature extraction combination module respectively.
[0097] The input is the second feature map ( Figure 3 The feature map X4 in the first feature extraction module is extracted by the feature extraction module in the first feature extraction module, and the feature map X5 is output. The feature map X5 is extracted by the deep feature extraction module of the easy floating object in the first feature extraction module to obtain the first scale feature map X6; the feature map X6 is extracted by the feature extraction module in the second feature extraction module, and the feature map X7 is output. The feature X7 is input to the deep feature extraction module of the easy floating object for feature extraction, and the spatial dimension is reduced to obtain the second scale feature map X8; the feature map X8 is extracted by the feature extraction module in the third feature extraction module, and the feature map X9 is output. The feature map X9 is extracted by the deep feature extraction module of the easy floating object in the third feature extraction module, and the spatial dimension is reduced to obtain the third scale feature map X10. The step of extracting the feature map by the deep feature extraction module of the easy floating object in this embodiment refers to the step of extracting the feature of the first feature map by the deep feature extraction module of the easy floating object in the above embodiment to obtain the second feature map in advance, and it will not be repeated here.
[0098] The third-scale feature map X10 is input into the deep feature extraction module of the floating object for feature extraction to obtain the feature map X11. The feature map X11 is input into the SPPF pooling layer (Spatial Pyramid Pooling - Fast) for fast pooling operation to obtain the fourth-scale feature map X12.
[0099] The fourth-scale feature map X12 is upsampled to obtain a feature map X13, and the feature map X13 is fused with the second-scale feature map X8 to obtain a feature fusion map X14. The feature fusion map X14 is input into the deep feature extraction module for floating objects to extract features, and a fifth-scale feature map X15 is obtained.
[0100] The fifth scale feature map X15 is upsampled to obtain a feature map X16, and the feature map X16 is fused with the first scale feature map X6 in the channel dimension to obtain a transmission channel feature map ( Figure 3 feature map X17 in .
[0101] In this embodiment, the features of the second feature map at different scales are extracted by the feature extraction combination module, which is beneficial to improving the detection accuracy and robustness of the module.
[0102] In an exemplary embodiment, the floating object feature attention module includes a maximum pooling layer and an average pooling layer, and S300 includes S310 to S350. Among them:
[0103] S310, performing pooling operations on the floating object feature map of the power transmission channel through a maximum pooling layer and an average pooling layer, respectively, to obtain a first attention feature map and a second attention feature map.
[0104] S320, performing hole convolution operations of different scales on the easy-to-float feature map of the transmission channel through the easy-to-float feature attention module to obtain a third attention feature map, a fourth attention feature map, and a fifth attention feature map.
[0105] S330, performing a channel weighting operation on the first attention feature map and the third attention feature map to obtain a sixth attention feature map.
[0106] S340, performing a channel weighting operation on the second attention feature map and the fifth attention feature map to obtain a seventh attention feature map.
[0107] S350, performing feature fusion and feature extraction on the fourth attention feature map, the sixth attention feature map and the seventh attention feature map to obtain an eighth attention feature map.
[0108] In order to reduce background interference and enhance the features of easily floating objects, this application designs an easily floating object feature attention module EFO Attention (Easily Floating Objects Feature Attention module). Its structure is as follows Figure 5 As shown. Figure 5 In the example, the input is the characteristic map of floating objects in the transmission channel S1 ( Figure 3 feature map X17 in .
[0109] First, the transmission channel easy-to-float feature map S1 is pooled through the global maximum pooling layer to obtain the first attention feature map S2, and the transmission channel easy-to-float feature map S1 is pooled through the global average pooling layer to obtain the second attention feature map S3.
[0110] Secondly, the transmission channel easy floating object feature map S1 is subjected to a dilated convolution operation with a dilated convolution kernel size of 3×3 and a dilation rate of 3, batch normalization and ReLu activation operations to obtain the third attention feature map S4; the transmission channel easy floating object feature map S1 is subjected to a dilated convolution operation with a dilated convolution kernel size of 5×5 and a dilation rate of 5, batch normalization and ReLu activation operations to obtain the third attention feature map S6; the transmission channel easy floating object feature map S1 is subjected to a dilated convolution operation with a dilated convolution kernel size of 9×9 and a dilation rate of 9, batch normalization and ReLu activation operations to obtain the third attention feature map S7.
[0111] Again, channel weighting operation is performed on the first attention feature map S2 and the third attention feature map S4 to obtain the sixth attention feature map S5.
[0112] Afterwards, a channel weighting operation is performed on the second attention feature map S3 and the fifth attention feature map S7 to obtain the seventh attention feature map S8.
[0113] Finally, the fourth attention feature map S6, the sixth attention feature map S5 and the seventh attention feature map S8 are fused in the channel dimension, and then a convolution operation with a convolution kernel of size 1×1 is performed to obtain the transmission channel easy floating object feature enhancement map S9.
[0114] In this embodiment, the most significant and average features of the input feature map are captured by global maximum pooling and global average pooling; multi-scale dilated convolution is used, combined with convolution kernels of different sizes and dilation rates, to expand the receptive field and fuse the contextual information of easy-to-float objects; through channel weighted operation, global pooling features and dilated convolution features are combined to generate deep feature maps, and enhance the sensitivity of the model to the key features of easy-to-float objects; finally, the multi-scale features are fused in the channel dimension and integrated through convolution to obtain the final enhanced feature map of easy-to-float objects in the transmission channel, which is conducive to reducing background interference and highlighting target features, thereby reducing the possibility of missed detection and false detection of the model.
[0115] In an exemplary embodiment, based on the floating object features of different scales and the enhanced map of floating object features of the transmission channel, determining the category, location information and confidence of the prediction box includes S420 to S460.
[0116] S420, extracting features of the easy-to-float object feature enhancement map of the power transmission channel through the feature extraction module in the initial easy-to-float object detection model to obtain a first target feature map, performing feature fusion on the first target feature map and the fifth-scale feature map to obtain a second target feature map, enhancing the feature representation of the second target feature map through the easy-to-float object feature attention module to obtain a first-scale target feature fusion map.
[0117] In practical applications, such as Figure 3As shown, the transmission channel easy-to-float feature enhancement map is input into the feature extraction module for feature extraction, and the first target feature map X19 is output. The first target feature map X19 and the fifth-size feature map X15 are feature fused to obtain the second target feature map X20. The second target feature map X20 is input into the easy-to-float feature attention module for feature enhancement processing to obtain the first-scale target feature fusion map X21. In this embodiment, the step of performing feature enhancement processing on the input feature map by the easy-to-float feature attention module refers to the step of performing pooling operation and channel weighting operation on the transmission channel easy-to-float feature map by the easy-to-float feature attention module, capturing the easy-to-float features in the transmission channel easy-to-float feature map and enhancing the easy-to-float feature representation, and obtaining the transmission channel easy-to-float feature enhancement map, which will not be repeated here.
[0118] S440, extracting features of the first-scale target feature fusion map through the feature fusion module in the initial easy-to-float object detection model to obtain a third target feature map, performing feature fusion on the third target feature map and the fourth-scale feature map to obtain a fourth target feature map, enhancing feature representation of the fourth target feature map through the easy-to-float object feature attention module, and obtaining a second-scale target feature fusion map.
[0119] In practical applications, the first-scale target feature fusion map X21 is input into the feature fusion module for feature extraction to obtain the third target feature map X22, the third target feature map X22 and the fourth-scale feature map X12 are subjected to channel dimension feature fusion to obtain the fourth target feature map X23, and the enhanced fourth target feature map X23 is input into the floating object feature attention module for feature enhancement processing to obtain the second-scale target feature fusion map X24.
[0120] S460, input the first scale target feature fusion map, the second scale target feature fusion map and the transmission channel easy floating object feature enhancement map into the target detection head of the initial easy floating object detection model, generate multiple prediction boxes, and output the category, location information and confidence of each prediction box.
[0121] In practical applications, the first-scale target feature fusion map, the second-scale target feature fusion map and the transmission channel floating object feature enhancement map are respectively input into the target prediction head of the initial floating object detection model. The target prediction head generates multiple prediction boxes for identifying floating objects, predicts the category of each prediction box, and outputs the category, location information and confidence of the prediction box.
[0122] In this embodiment, the first-scale target feature fusion map, the second-scale target feature fusion map and the transmission channel easy-to-float feature enhancement map of different scales are extracted through the easy-to-float feature attention module and the feature extraction module, so that the category, position information and confidence of the predicted box are output by the detection head, thereby improving the detection accuracy.
[0123] In order to measure the prediction quality of the model, in an exemplary embodiment, S500 includes S510 to S550. Among them:
[0124] S510, determining an intersection-over-union ratio between the real frame and the predicted frame based on the position information of the real frame and the position information of the predicted frame.
[0125] In practical applications, the intersection-and-union ratio of the real box and the predicted box is calculated based on the position information of the real box and the predicted box to measure the degree of overlap between them. The intersection-and-union ratio is the ratio of the intersection area of the real box and the predicted box to the union area. The higher the intersection-and-union ratio, the higher the degree of overlap between the two.
[0126] S520, determining a prediction box loss value between the real box and the prediction box based on the intersection-over-union ratio.
[0127] Among them, the predicted box loss value represents the position difference between the predicted box and the true box.
[0128] In practical applications, the predicted box loss value between the real box and the predicted box can be obtained according to a preset loss function and intersection-over-union ratio. The preset loss function includes but is not limited to a mean square error loss function, a cross entropy loss function, and a smooth L1 loss function.
[0129] S530, determining a category loss value between the real box and the predicted box according to the category label of the real box, the category of the predicted box and a preset loss function.
[0130] Among them, the category loss value is used to measure the difference between the predicted box category distribution and the actual category.
[0131] In practical applications, the category label of the true box and the category probability distribution of the predicted box can be obtained by combining sigmoid with the binary cross entropy loss function to determine the category loss value between the true box and the predicted box.
[0132] S540, determining a confidence loss value between the real box and the predicted box according to the intersection-over-union ratio between the real box and the predicted box and the confidence of the predicted box.
[0133] Among them, the confidence loss value is used to measure the prediction error of the model on whether there are floating objects in the prediction box.
[0134] In practical applications, the intersection-over-union ratio of the true box and the predicted box can be the true confidence. The confidence loss value between the true box and the predicted box is determined according to the cross entropy loss function, the true confidence and the confidence of the predicted box.
[0135] S550, determining the error between the predicted box and the true box according to the predicted box loss value, the category loss value and the confidence loss value.
[0136] In practical applications, weights can be assigned to the prediction box loss value, category loss value, and confidence loss value respectively, and the loss value (error) between the prediction box and the true box can be obtained by weighted summation.
[0137] In this embodiment, the error between the prediction result of the model and the true label is determined by the prediction box loss value, category loss value and confidence loss value between the true box and the prediction box, which is conducive to guiding the adjustment of model parameters according to the error to minimize the error.
[0138] In an exemplary embodiment, S520 includes S522 to S524. Among them:
[0139] S522, determine the minimum bounding box between the real box and the predicted box.
[0140] In practical applications, the minimum bounding box is determined to be the minimum enclosing rectangle containing the real box and the predicted box.
[0141] S524, determining a weight factor based on the size of the minimum bounding box, the size of the real box, and the size of the predicted box, and determining a predicted box loss value between the real box and the predicted box according to the intersection-over-union ratio and the weight factor.
[0142] Considering that in the acquired remote sensing images, the proportion of floating objects in the image is small and the interference of background pixels is large, this application specially designs an improved intersection-over-union loss function: The loss function introduces weight factors that are closely related to the size of the true box and the predicted box. These factors comprehensively consider the minimum bounding box and the width and height of the true box and the predicted box, thereby significantly improving the detection accuracy of floating objects.
[0143] The loss function is defined as follows:
[0144]
[0145] In the formula, It represents the intersection-over-union ratio of the predicted box and the true box, that is, , and Represent the predicted box and the true box respectively; and Represent the width and height of the real box respectively; and Respectively represent the width and height of the prediction box; and Represent the width and height of the smallest bounding box consisting of the predicted box and the true box, respectively.
[0146] In this embodiment, the calculation method of the intersection-over-union loss function is improved by introducing a weight factor that is closely related to the size of the real box and the predicted box, thereby improving the accuracy of calculating the loss value of the predicted box and thus improving the detection accuracy of the model.
[0147] In order to make a clearer description of the method for training a model for detecting floating objects along a power transmission channel provided by the present application, the following is a specific embodiment and the attached drawings. Figure 6 To illustrate, this specific embodiment includes the following steps:
[0148] S1, obtain remote sensing images along the historical transmission channel, which contain category labels and location information of the real boxes of floating objects.
[0149] S2, inputting the remote sensing images along the historical power transmission channel into the constructed initial floating object detection model, extracting the features of the remote sensing images along the historical power transmission channel through a feature extraction module, and obtaining a first feature map.
[0150] S3, through the hole convolution layer in the deep feature extraction module of easy floating objects, the first feature map is subjected to hole convolution to obtain the first deep feature map, the first deep feature map is subjected to convolution operation, batch normalization operation, nonlinear processing and convolution operation in sequence through the feature dimension reduction branch to obtain the second deep feature map, the first feature map is subjected to pooling operation through the multi-scale feature fusion branch to obtain the third deep feature map, the third deep feature map and the second deep feature map are subjected to channel weighting operation through the convolution layer to obtain the fourth deep feature map, the first feature map is subjected to maximum pooling operation and average pooling operation respectively through the feature pooling branch, the results of the maximum pooling operation and the average pooling operation are subjected to feature fusion to obtain the sixth deep feature map, the sixth deep feature map is subjected to channel weighting operation to obtain the seventh deep feature map, the second deep feature map and the seventh deep feature map are subjected to feature fusion and convolution operation to obtain the second feature map.
[0151] S4, respectively connecting the output end of each feature extraction module with the input end of each easy-to-float deep feature extraction module to obtain a first feature extraction combination module, a second feature extraction combination module and a third feature extraction combination module, extracting features of the third feature map through the first feature extraction combination module to obtain a first scale feature map, extracting features of the first scale feature map through the second feature extraction combination module to obtain a second scale feature map, extracting features of the second scale feature map through the third feature extraction combination module to obtain a third scale feature map, performing feature extraction on the third scale feature map through the easy-to-float deep feature extraction module, performing a pooling operation on the extracted features to obtain a fourth scale feature map, performing an upsampling operation on the fourth scale feature map, performing feature fusion on the upsampled feature map with the second scale feature map to obtain a feature fusion map, performing feature extraction on the easy-to-float deep feature extraction module to obtain a fifth scale feature map, performing an upsampling operation on the fifth scale feature map, performing feature fusion on the upsampled feature map with the first scale feature map to obtain a transmission channel easy-to-float feature map.
[0152] S5, performing pooling operations on the easy-to-float feature map of the transmission channel through the maximum pooling layer and the average pooling layer respectively, to obtain a first attention feature map and a second attention feature map, performing hole convolution operations of different scales on the easy-to-float feature map of the transmission channel through the easy-to-float feature attention module respectively, to obtain a third attention feature map, a fourth attention feature map and a fifth attention feature map, performing channel weighting operations on the first attention feature map and the third attention feature map to obtain a sixth attention feature map, performing channel weighting operations on the second attention feature map and the fifth attention feature map to obtain a seventh attention feature map, performing feature fusion and feature extraction on the sixth attention feature map and the seventh attention feature map to obtain an eighth attention feature map.
[0153] S6, extracting features of the enhanced map of easy-to-float objects in the transmission channel through the feature extraction module in the initial easy-to-float object detection model to obtain a first target feature map, performing feature fusion on the first target feature map and the fifth-scale feature map to obtain a second target feature map, enhancing the feature representation of the second target feature map through the easy-to-float object feature attention module to obtain a first-scale target feature fusion map, extracting features of the first-scale target feature fusion map through the feature fusion module in the initial easy-to-float object detection model to obtain a third target feature map, performing feature fusion on the third target feature map and the fourth-scale feature map to obtain a fourth target feature map, enhancing the feature representation of the fourth target feature map through the easy-to-float object feature attention module to obtain a second-scale target feature fusion map, inputting the first-scale target feature fusion map, the second-scale target feature fusion map and the enhanced map of easy-to-float objects in the transmission channel into the target detection head of the initial easy-to-float object detection model, generating multiple prediction frames, and outputting the category, location information and confidence of each prediction frame.
[0154] S7, based on the position information of the real frame and the position information of the predicted frame, determine the intersection-and-union ratio of the real frame and the predicted frame, determine the minimum bounding box of the real frame and the predicted frame, determine the weight factor based on the size of the minimum bounding box, the size of the real frame and the size of the predicted frame, determine the prediction box loss value between the real frame and the predicted frame according to the intersection-and-union ratio and the weight factor, determine the category loss value between the real frame and the predicted frame according to the category label of the real frame, the category of the predicted frame and a preset loss function, determine the confidence loss value between the real frame and the predicted frame according to the intersection-and-union ratio of the real frame and the predicted frame, and the confidence of the predicted frame, and determine the error between the predicted frame and the real frame according to the prediction box loss value, the category loss value and the confidence loss value.
[0155] In practical applications, the size of the feature map is set to H×W×C, where H is the height of the feature map, W is the width of the feature map, and C is the number of channels of the feature map.
[0156] like Figure 3 As shown in the figure, the remote sensing image X1 of floating objects along the transmission channel is used as the input of the remote sensing detection model of floating objects along the transmission channel, and the size of X1 is 512×512×3. First, the Conv_BN_ReLu feature extraction module is applied to process X1 to obtain the feature map X2, and the size of X2 is 512×512×64; then, the Conv_BN_ReLu feature extraction module is applied to X2 for convolution, batch normalization and activation operations to obtain the first feature map X3, and the size of the first feature map X3 is 128×128×128; then, the EFODE module is applied to process the first feature map X3 to obtain the second feature map X4, and the size of the second feature map X4 is 128×128×256; secondly, the Conv_BN_ReLu extraction module is applied to process the second feature map X4 to obtain the feature map X5, and the size of X5 is 64×64×512, and X5 is input to the EFODE module. The block is processed to obtain the first-scale feature map X6, and the size of X6 is 64×64×512; then, the Conv_BN_ReLu feature extraction module is applied to X6 for convolution, batch normalization and activation operations to obtain the feature map X7, the size of X7 is 32×32×512, and X7 is input to the EFODE module for processing to obtain the second-scale feature map X8, the size of X8 is 32×32×512; finally, the Conv_BN_ReLu module is applied to X8 for convolution, batch normalization and activation operations to obtain the feature map X9, the size of X9 is 16×16×512, and X9 is input to the EFODE module for processing to obtain the third-scale feature map X10, the size of X10 is 16×16×512.
[0157] The EFODE module is applied to process the third-scale feature map X10 to obtain the feature map X11, and the size of X11 is 16×16×512; secondly, the SPPF is applied to process X11 to obtain the fourth-scale feature map X12, and the size of X12 is 16×16×512. After upsampling, the feature map X13 is obtained, and the size of X13 is 32×32×512; then, the second-scale feature map X8 is fused with X13 in the channel dimension to obtain the thirteenth floating object feature map X14, and the size of X14 is The size of the feature fusion map X14 is 32×32×1024. Then, the sixth EFODE module is applied to process the feature fusion map X14 to obtain the fifth-scale feature map X15. The size of X15 is 32×32×1024. X15 is upsampled to obtain the feature map X16. The size of X16 is 64×64×512. Finally, the first-scale feature map X6 and the fifth-scale feature map X15 are fused in the channel dimension to obtain the transmission channel floating object feature map X17. The size of X17 is 64×64×1024.
[0158] The transmission channel easy floating object feature map X17 is input into the EFO Attention module to obtain the transmission channel easy floating object feature enhancement map X18, and the size of X18 is 64×64×512; secondly, the Conv_BN_ReLu module is applied to X18 for convolution, batch normalization and activation operations to obtain the first target feature map X19, and the size of X19 is 32×32×512, and X19 is fused with the fifth size feature map X15 in the channel dimension to obtain the second target feature map X20, and the size of X20 is 32×32×1536; then, EFO is applied The Attention module processes X20 to obtain the first-scale target feature fusion map X21, the size of which is 32×32×512, and the Conv_BN_ReLu module is used to perform convolution, batch normalization and activation operations on X21 to obtain the third target feature map X22, the size of which is 16×16×512; then, the fourth-scale feature map X12 is fused with the third target feature map X22 in the channel dimension to obtain the fourth target feature map X23, the size of which is 16×16×1024; finally, the EFO Attention module is used to process X23 to obtain the second-scale target feature fusion map X24, the size of which is 16×16×512. Finally, X18, X21, and X24 are respectively input into the target prediction head Head, and a tensor containing the prediction information is output. Each row of the tensor corresponds to a detection result, including information such as the predicted bounding box coordinates, defect category label, and confidence score. This structure achieves accurate detection and classification of floating objects along power transmission channels in remote sensing images through layer-by-layer extraction, splicing and fusion.
[0159] S8, iteratively updating the parameters of the initial floating object detection model based on the error until a preset training end condition is reached to obtain a trained floating object detection model.
[0160] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0161] The method for detecting tree obstacle hazards in a power transmission channel provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 through a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed on the cloud or other network servers.
[0162] Specifically, the operator may send a message about the detection of floating objects in the power transmission channel to the server 104 through the terminal. The server receives the message about the detection of floating objects in the power transmission channel. Next, the remote sensing images along the power transmission channel are obtained. The remote sensing images along the power transmission channel are used as input, and a trained floating object detection model is called to obtain the floating object detection results. An early warning is issued based on the floating object detection results.
[0163] In an exemplary embodiment, the present application also provides a method for detecting floating objects along a power transmission channel, such as Figure 6 As shown, it includes the following S700 to S900. Among them:
[0164] S700: Acquire remote sensing images along the power transmission channel.
[0165] The remote sensing images along the power transmission channel may be images acquired by collecting images along the power transmission channel.
[0166] In practical applications, during power grid operations, regular inspections can be carried out along the transmission channel, and remote sensing images along the transmission channel can be obtained through drone remote sensing and satellite remote sensing.
[0167] S800, using the remote sensing image along the power transmission channel as input, calling the trained floating object detection model to obtain the floating object detection result, the floating object detection model is trained based on the above-mentioned floating object detection model training method along the power transmission channel.
[0168] The floating object detection result may include a detection frame and a category label of the area where the floating object is located.
[0169] In practical applications, during the inspection and collection of remote sensing images along the power transmission channel, the acquired remote sensing images along the power transmission channel are input into the trained model obtained by the above-mentioned training method for the model for detecting floating objects along the power transmission channel, and the area where the floating objects are located in the remote sensing images along the power transmission channel is detected to obtain the detection result of the floating objects. The floating object detection model is obtained by training using the steps in any of the above-mentioned training methods for detecting floating objects along the power transmission channel. The specific model training process will not be repeated here.
[0170] S900, provides early warning based on the detection results of floating objects.
[0171] In practical applications, when the detection of easily floating objects is indicated in the easily floating object detection results, the location information of the easily floating objects in the image can be obtained based on the remote sensing image along the transmission channel where the easily floating objects are detected, and the specific location information of the easily floating objects in the transmission channel can be obtained based on the source of the remote sensing image along the transmission channel. The specific location information of the easily floating objects and the time when the easily floating objects are detected are integrated to obtain the warning information and push the warning information. The method of pushing the warning information can include pushing the alarm information in the notification bar, displaying the alarm information in the form of a full-screen or half-screen pop-up window, giving an alarm prompt by emitting a specific sound or vibration mode, giving a visual alarm prompt by flashing an indicator light, etc. It can be understood that the method of pushing the alarm information can be any one of the aforementioned methods or a combination of any multiple methods, which is not limited here.
[0172] In this embodiment, by acquiring remote sensing images along the transmission channel and inputting them into the trained floating object detection model obtained by the above-mentioned floating object detection model training method along the transmission channel, a floating object detection result is obtained, and an early warning is given based on the floating object detection result. On the one hand, the accuracy and efficiency of floating detection are improved through model detection. On the other hand, early warning is given based on the rapid and accurate detection results of floating objects along the transmission channel based on the model, which is conducive to timely discovery of potential hidden dangers of the power system, so as to take corresponding measures to reduce the possibility of accidents and improve the stability and safety of power system operation.
[0173] In an exemplary embodiment, Figure 7As shown, a model training device 600 for detecting floating objects along a power transmission channel is provided, comprising: an image acquisition module 610, a model training module 620, an error determination module 630 and a parameter updating module 640, wherein:
[0174] The image acquisition module 610 is used to acquire remote sensing images along the historical power transmission channel, and the remote sensing images along the historical power transmission channel contain category labels and location information of the real frames of the floating objects.
[0175] The model training module 620 is used to input the remote sensing images along the historical power transmission channel into the constructed initial floating object detection model. The initial floating object detection model extracts the floating object features of different scales in the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates the power transmission channel floating object feature map based on the floating object features of different scales; the floating object feature attention module performs pooling operations and channel weighting operations on the power transmission channel floating object feature map to capture the floating object features and enhance the floating object feature representation in the power transmission channel floating object feature map to obtain the power transmission channel floating object feature enhancement map; based on the floating object features of different scales and the power transmission channel floating object feature enhancement map, the category, location information and confidence of the prediction box are determined. The prediction box is used to identify the floating objects in the remote sensing images along the historical power transmission channel. The initial floating object detection model includes the floating object deep feature extraction module and the floating object feature attention module.
[0176] The error determination module 630 is used to determine the error between the real box and the predicted box according to the category label and position information of the real box, and the category, position information and confidence of the predicted box.
[0177] The parameter updating module 640 is used to iteratively update the parameters of the initial floating object detection model until a preset training end condition is reached to obtain a trained floating object detection model.
[0178] In an exemplary embodiment, the model training module 620 is also used to extract features of remote sensing images along the historical power transmission channel through the feature extraction module to obtain a first feature map; perform convolution operations, channel weighting operations and pooling operations on the first feature map through the feature dimension reduction branch, multi-scale feature fusion branch and feature pooling branch in the deep feature extraction module for floating objects to obtain a second feature map; based on the feature extraction module and the deep feature extraction module for floating objects, extract features of the second feature map at different scales, and obtain a floating object feature map of the power transmission channel based on the features at different scales.
[0179] In an exemplary embodiment, the model training module 620 is further used to perform a dilated convolution on the first feature map through the dilated convolution layer in the easy-to-float deep feature extraction module to obtain a first deep feature map; perform a convolution operation, a batch normalization operation, a nonlinear processing and a convolution operation on the first deep feature map in sequence through the feature dimension reduction branch to obtain a second deep feature map; perform a pooling operation on the first feature map through the multi-scale feature fusion branch to obtain a third deep feature map, perform a channel weighting operation on the third deep feature map and the second deep feature map through the convolution layer to obtain a fourth deep feature map, and perform a convolution operation on the fourth deep feature map to obtain a fifth deep feature map; perform a maximum pooling operation and an average pooling operation on the first feature map through the feature pooling branch, perform feature fusion on the results of the maximum pooling operation and the average pooling operation to obtain a sixth deep feature map, perform a channel weighting operation on the fifth deep feature map and the sixth deep feature map to obtain a seventh deep feature map; perform feature fusion and convolution operations on the second deep feature map and the seventh deep feature map to obtain a second feature map.
[0180] In an exemplary embodiment, the model training module 620 is further used to connect the output end of each feature extraction module with the input end of each easy-floating object deep feature extraction module to obtain a first feature extraction combination module, a second feature extraction combination module and a third feature extraction combination module;
[0181] The features of the second feature map are extracted by the first feature extraction combination module to obtain a first scale feature map, the features of the first scale feature map are extracted by the second feature extraction combination module to obtain a second scale feature map, and the features of the second scale feature map are extracted by the third feature extraction combination module to obtain a third scale feature map; the third scale feature map is extracted by the deep feature extraction module for easy floating objects, and the extracted features are pooled to obtain a fourth scale feature map; the fourth scale feature map is upsampled, and the upsampled feature map is feature-fused with the second scale feature map to obtain a feature fusion map, and the deep feature extraction module for easy floating objects is used to extract features to obtain a fifth scale feature map; the fifth scale feature map is upsampled, and the upsampled feature map is feature-fused with the first scale feature map to obtain a transmission channel easy floating object feature map.
[0182] In an exemplary embodiment, the model training module 620 is also used to perform pooling operations on the transmission channel easy-to-float feature map through a maximum pooling layer and an average pooling layer, respectively, to obtain a first attention feature map and a second attention feature map; perform hole convolution operations of different scales on the transmission channel easy-to-float feature map through the easy-to-float feature attention module, respectively, to obtain a third attention feature map, a fourth attention feature map and a fifth attention feature map; perform channel weighting operations on the first attention feature map and the third attention feature map to obtain a sixth attention feature map; perform channel weighting operations on the second attention feature map and the fifth attention feature map to obtain a seventh attention feature map; perform feature fusion and feature extraction on the fourth attention feature map, the sixth attention feature map and the seventh attention feature map to obtain a transmission channel easy-to-float feature enhancement map.
[0183] In an exemplary embodiment, the model training module 620 is further used to extract features of the enhanced map of easy-to-float objects in the transmission channel through the feature extraction module in the initial easy-to-float object detection model to obtain a first target feature map, perform feature fusion on the first target feature map and the fifth-size feature map to obtain a second target feature map, enhance the feature representation of the second target feature map through the easy-to-float object feature attention module to obtain a first-scale target feature fusion map; extract features of the first-scale target feature fusion map through the feature fusion module in the initial easy-to-float object detection model to obtain a third target feature map, perform feature fusion on the third target feature map and the fourth-scale feature map to obtain a fourth target feature map, enhance the feature representation of the fourth target feature map through the easy-to-float object feature attention module to obtain a second-scale target feature fusion map; input the first-scale target feature fusion map, the second-scale target feature fusion map and the enhanced map of easy-to-float objects in the transmission channel into the target detection head of the initial easy-to-float object detection model, generate multiple prediction frames, and output the category, location information and confidence of each prediction frame.
[0184] In an exemplary embodiment, the error determination module 630 is further used to determine the intersection-over-union ratio of the real box and the predicted box based on the position information of the real box and the position information of the predicted box; determine the prediction box loss value between the real box and the predicted box based on the intersection-over-union ratio; determine the category loss value between the real box and the predicted box according to the category label of the real box, the category of the predicted box and a preset loss function; determine the confidence loss value between the real box and the predicted box according to the intersection-over-union ratio of the real box and the predicted box, and the confidence of the predicted box; determine the error between the predicted box and the real box according to the prediction box loss value, the category loss value and the confidence loss value.
[0185] In an exemplary embodiment, the error determination module 640 is further used to determine the minimum bounding box between the real box and the predicted box; determine a weight factor based on the size of the minimum bounding box, the size of the real box, and the size of the predicted box; and determine the prediction box loss value between the real box and the predicted box according to the intersection-over-union ratio and the weight factor.
[0186] In an exemplary embodiment, Figure 8 As shown, a floating object detection device 700 along a power transmission channel is provided, comprising: a remote sensing image acquisition module 710, a floating object detection module 720 and an early warning module 730, wherein:
[0187] A remote sensing image acquisition module 710 is used to acquire remote sensing images along the transmission channel;
[0188] The tree obstacle hidden danger detection module 720 is used to use the remote sensing image along the transmission channel as input, call the trained floating object detection model, and obtain the floating object detection result, wherein the floating object detection model is trained based on the above-mentioned floating object detection model training method along the transmission channel;
[0189] The early warning module 730 is used to issue an early warning based on the detection result of the floating object.
[0190] Each module in the above-mentioned floating object detection model training device 600 and floating object detection device 700 can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each of the above modules.
[0191] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig. 9As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for training a model for detecting floating objects along a power transmission channel is implemented.
[0192] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0193] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps in any one of the above-mentioned embodiments of the method for training a model for detecting floating objects along a power transmission channel are implemented.
[0194] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in any one of the above-mentioned embodiments of the method for training a model for detecting floating objects along a power transmission channel are implemented.
[0195] In one embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the steps in any of the above-mentioned embodiments of the method for training a model for detecting floating objects along a power transmission channel.
[0196] It should be noted that the data involved in this application (including but not limited to data used for analysis, storage, display, etc.) are all data fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0197] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but is not limited thereto.
[0198] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0199] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A method for training a model for detecting floating objects along a power transmission channel, characterized in that: The method comprises: Acquire remote sensing images along historical power transmission channels, wherein the remote sensing images along historical power transmission channels contain category labels and location information of real frames of floating objects; The remote sensing images along the historical power transmission channel are input into the constructed initial floating object detection model. The initial floating object detection model extracts the floating object features of different scales from the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates a transmission channel floating object feature map based on the floating object features of different scales. The easy-to-float feature map of the transmission channel is subjected to a pooling operation and a channel weighting operation through a easy-to-float feature attention module, so as to capture the easy-to-float feature and enhance the easy-to-float feature representation in the easy-to-float feature map of the transmission channel, and obtain an easy-to-float feature enhanced map of the transmission channel; Based on the features of the floatable objects at different scales and the enhanced map of the features of the floatable objects along the transmission channel, determining the category, location information and confidence of the prediction box, wherein the prediction box is used to identify the floatable objects in the remote sensing images along the historical transmission channel, and the initial floatable object detection model includes a deep feature extraction module for the floatable objects and an attention module for the features of the floatable objects; Determine the error between the real frame and the predicted frame according to the category label and position information of the real frame, and the category, position information and confidence of the predicted frame; Based on the error, the parameters of the initial easy-to-float object detection model are iteratively updated until a preset training end condition is reached to obtain a trained easy-to-float object detection model.
2. The method according to claim 1, characterized in that The initial floating object detection model includes a feature extraction module; The initial floating object detection model extracts floating object features of different scales from the remote sensing images along the historical power transmission channel through a floating object deep feature extraction module, and generates a floating object feature map of the power transmission channel based on the floating object features of different scales, including: Extracting features of the remote sensing image along the historical power transmission channel by the feature extraction module to obtain a first feature map; Through the feature dimension reduction branch, the multi-scale feature fusion branch and the feature pooling branch in the deep feature extraction module of the easy-to-float object, the first feature map is subjected to convolution operation, channel weighting operation and pooling operation respectively, to obtain a second feature map; Based on the feature extraction module and the deep feature extraction module for easily floating objects, the features of the second feature map at different scales are extracted, and based on the features at different scales, a feature map of easily floating objects in a power transmission channel is obtained.
3. The method according to claim 2, characterized in that The feature dimension reduction branch, the multi-scale feature fusion branch and the feature pooling branch in the deep feature extraction module of the floating object are used to perform convolution operation, channel weighting operation and pooling operation on the first feature map respectively, so as to obtain the second feature map including: Performing a dilated convolution on the first feature map through the dilated convolution layer in the deep feature extraction module for floating objects to obtain a first deep feature map; Performing convolution operation, batch normalization operation, nonlinear processing and convolution operation on the first deep feature map in sequence through the feature dimension reduction branch to obtain a second deep feature map; Performing a pooling operation on the first feature map through the multi-scale feature fusion branch to obtain a third deep feature map, performing a channel weighting operation on the third deep feature map and the second deep feature map through a convolution layer to obtain a fourth deep feature map, and performing a convolution operation on the fourth deep feature map to obtain a fifth deep feature map; Performing a maximum pooling operation and an average pooling operation on the first feature map through a feature pooling branch, performing feature fusion on the results of the maximum pooling operation and the average pooling operation to obtain a sixth deep feature map, and performing a channel weighted operation on the fifth deep feature map and the sixth deep feature map to obtain a seventh deep feature map; Perform feature fusion and convolution operations on the second deep feature map and the seventh deep feature map to obtain a second feature map.
4. The method according to claim 3, characterized in that The number of the deep feature extraction module for easy floating objects and the number of the feature extraction modules are multiple; based on the feature extraction module and the deep feature extraction module for easy floating objects, the features of the second feature map at different scales are extracted to obtain the feature map of easy floating objects in the transmission channel, including: Respectively connecting the output end of each of the feature extraction modules with the input end of each of the deep-layer feature extraction modules of the easily floating objects to obtain a first feature extraction combination module, a second feature extraction combination module and a third feature extraction combination module; Extracting features of the second feature map through the first feature extraction and combination module to obtain a first scale feature map, extracting features of the first scale feature map through the second feature extraction and combination module to obtain a second scale feature map, and extracting features of the second scale feature map through the third feature extraction and combination module to obtain a third scale feature map; Performing feature extraction on the third scale feature map by using the deep feature extraction module for easy floating objects, and performing a pooling operation on the extracted features to obtain a fourth scale feature map; By performing an upsampling operation on the fourth scale feature map, performing feature fusion on the upsampled feature map and the second scale feature map to obtain a feature fusion map, and performing feature extraction on the feature map through the easy-floating object deep feature extraction module to obtain a fifth scale feature map; An upsampling operation is performed on the fifth-size feature map, and feature fusion is performed on the feature map after the upsampling operation and the first-scale feature map to obtain a transmission channel easy-to-float feature map.
5. The method according to claim 4, characterized in that The easy-to-float feature attention module includes a maximum pooling layer and an average pooling layer; the easy-to-float feature attention module performs a pooling operation and a channel weighting operation on the easy-to-float feature map of the transmission channel, captures the easy-to-float feature and enhances the easy-to-float feature representation in the easy-to-float feature map of the transmission channel, and obtains the easy-to-float feature enhanced map of the transmission channel, including: Performing a pooling operation on the power transmission channel easy-to-float feature map through the maximum pooling layer and the average pooling layer respectively to obtain a first attention feature map and a second attention feature map; By using the easy-to-float feature attention module, the easy-to-float feature map of the transmission channel is subjected to hole convolution operations of different scales to obtain a third attention feature map, a fourth attention feature map and a fifth attention feature map; Performing a channel weighted operation on the first attention feature map and the third attention feature map to obtain a sixth attention feature map; Performing a channel weighted operation on the second attention feature map and the fifth attention feature map to obtain a seventh attention feature map; The fourth attention feature map, the sixth attention feature map and the seventh attention feature map are subjected to feature fusion and feature extraction to obtain a transmission channel easy-to-float feature enhancement map.
6. The method according to claim 4, characterized in that The determining of the category, location information and confidence of the prediction box based on the features of different scales and the enhanced map of the features of the power transmission channel that are prone to floating objects includes: Extracting features of the enhanced map of the easy-to-float object feature of the power transmission channel by the feature extraction module in the initial easy-to-float object detection model to obtain a first target feature map, performing feature fusion on the first target feature map and the fifth-size feature map to obtain a second target feature map, and enhancing feature representation of the second target feature map by the easy-to-float object feature attention module to obtain a first-scale target feature fusion map; Extracting features of the first-scale target feature fusion map through the feature fusion module in the initial easy-to-float object detection model to obtain a third target feature map, performing feature fusion on the third target feature map and the fourth-scale feature map to obtain a fourth target feature map, and enhancing feature representation of the fourth target feature map through the easy-to-float object feature attention module to obtain a second-scale target feature fusion map; The first-scale target feature fusion map, the second-scale target feature fusion map and the power transmission channel easy-to-float object feature enhancement map are input into the target detection head of the initial easy-to-float object detection model to generate multiple prediction frames, and the category, location information and confidence of each prediction frame are output.
7. The method according to claim 6, characterized in that The determining, according to the category label and position information of the real frame, and the category, position information, and confidence of the predicted frame, an error between the real frame and the predicted frame includes: Determine an intersection-over-union ratio between the real frame and the predicted frame based on the position information of the real frame and the position information of the predicted frame; Based on the intersection-over-union ratio, determining a prediction box loss value between the real box and the prediction box; Determine a category loss value between the real frame and the predicted frame according to the category label of the real frame, the category of the predicted frame and a preset loss function; Determine a confidence loss value between the real frame and the predicted frame according to the intersection-over-union ratio of the real frame and the predicted frame, and the confidence of the predicted frame; The error between the predicted box and the true box is determined according to the predicted box loss value, the category loss value and the confidence loss value.
8. The method according to claim 7, characterized in that The determining, based on the intersection-over-union ratio, a prediction frame loss value between the real frame and the prediction frame comprises: Determine the minimum bounding box of the real box and the predicted box; A weight factor is determined based on the size of the minimum bounding box, the size of the real box, and the size of the predicted box, and a prediction box loss value between the real box and the predicted box is determined according to the intersection-over-union ratio and the weight factor.
9. A method for detecting floating objects along a power transmission channel, characterized in that: The method comprises: Acquire remote sensing images along the transmission corridor; Taking the remote sensing image along the transmission channel as input, calling the trained floating object detection model to obtain the floating object detection result, wherein the floating object detection model is trained based on the floating object detection model training method along the transmission channel according to any one of claims 1 to 8; An early warning is issued based on the detection result of the easily floating objects.
10. A model training device for detecting floating objects along a power transmission channel, characterized in that: The device comprises: An image acquisition module is used to acquire remote sensing images along the historical power transmission channel, wherein the remote sensing images along the historical power transmission channel contain category labels and location information of real frames of floating objects; A model training module is used to input the remote sensing images along the historical power transmission channel into the constructed initial floating object detection model, wherein the initial floating object detection model extracts floating object features of different scales from the remote sensing images along the historical power transmission channel through the floating object deep feature extraction module, and generates a power transmission channel floating object feature map based on the floating object features of different scales; the floating object feature attention module performs pooling operations and channel weighting operations on the floating object feature map of the transmission channel, captures the floating object features and enhances the floating object feature representation in the floating object feature map of the transmission channel, and obtains the floating object feature enhancement map of the transmission channel; based on the floating object features of different scales and the floating object feature enhancement map of the transmission channel, determines the category, location information and confidence of the prediction box, wherein the prediction box is used to identify the floating objects in the remote sensing images along the historical power transmission channel, and the initial floating object detection model includes the floating object deep feature extraction module and the floating object feature attention module; An error determination module, configured to determine an error between the real frame and the predicted frame according to the category label and position information of the real frame, and the category, position information and confidence of the predicted frame; The parameter updating module is used to iteratively update the parameters of the initial easy-to-float object detection model until a preset training end condition is reached to obtain a trained easy-to-float object detection model.
Citation Information
Patent Citations
Small target floating garbage detection method based on improved YOLOv7 model
CN117292313A
Power transmission channel floater identification monitoring system and method based on satellite remote sensing
CN118411627A
Method and system for detecting easily floating objects in power transmission channel based on optical and radar fusion
CN118566901A
Classifying and detecting foreign objects using a power amplifier controller integrated circuit in wireless power transmission systems
US20210091606A1