A Remote Sensing Rotating Aircraft Detection Algorithm Based on Vector Regression and Attention Mechanism

By adopting aircraft object detection algorithm based on vector regression and attention mechanisms in remote sensing aerial image processing, the shortcomings of traditional algorithms in rotation, tilt, non-rectangular object detection and scale change processing are solved, and higher detection accuracy and robustness are achieved.

CN117830876BActive Publication Date: 2025-06-27CHINA THREE GORGES UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311556266.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-06-27
Estimated Expiration
2043-11-21

AI Technical Summary

Technical Problem

Traditional object detection algorithms have problems such as rotation, tilt, non-rectangular object detection difficulties, poor processing of scale changes, and background interference affecting accuracy when processing remote sensing aerial images.

Method used

The aircraft object detection algorithm based on vector regression and attention mechanism is adopted, and the accuracy and robustness of rotating aircraft detection are improved through the Cross-latitude Attention model and the ClA-DLANet model, combined with five-parameter annotation and Authentic SmoothL1 Loss loss function.

Benefits of technology

The accuracy and robustness of rotary aircraft detection were significantly improved, and compared with networks without attention mechanisms, it improved 1.22%, 0.76%, and 0.97% in Precision, Recall, and F1 respectively, and performed better than other loss functions in the verification of angle loss validity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117830876B_ABST
    Figure CN117830876B_ABST
Patent Text Reader

Abstract

A remote sensing rotating aircraft detection algorithm based on vector regression and attention mechanism, which comprises the following steps: Step 1: Select the UAV aerial photography dataset as the dataset; Step 2: Uniformly convert the four-parameter and eight-parameter annotation methods of the obtained dataset into five-parameter annotation, and use a rotation box with angle specificity to improve the detection accuracy; Step 3: Divide the dataset into a training set, a validation set, and a test set according to a certain ratio; Step 4: Construct a Cross-latitude Attention model and a ClA-DLANet model; Step 5: Test on the test dataset through the trained model, evaluate the model through evaluation indicators, and adjust the parameters; Detect the rotating aircraft through the above steps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and relates to an aircraft target detection algorithm based on vector regression and attention mechanism. Background Art

[0002] With the development of modern technology, remote sensing target detection is widely used in urban planning and detection, water resource management, and farmland detection. However, some traditional target detection algorithms have some defects:

[0003] (1) Traditional target detection algorithms usually perform target detection based on rectangular bounding boxes. However, in remote sensing aerial images, the targets may be rotated, tilted, or non-rectangular. This results in limitations in processing images by traditional algorithms and inability to accurately locate and identify targets.

[0004] (2) The targets in remote sensing aerial images have different scales, ranging from small to large. Traditional algorithms often use fixed-scale windows or sliding window methods for target detection, which leads to the inability to effectively process targets with scale changes and may result in missed detections or false detections.

[0005] (3) The background of aerial images is usually complex, with a large number of textures, buildings, etc. The algorithm is easily interfered by the background during processing, affecting the accuracy of target detection. Summary of the Invention

[0006] The present invention is a remote sensing rotating aircraft detection algorithm based on vector regression and attention mechanism, which is proposed to solve the problems of insufficient semantic information, low accuracy, and high missed detection rate in the existing technology during the detection process of remote sensing aircraft.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0008] A remote sensing rotating aircraft algorithm based on vector regression and attention mechanism, which includes the following steps:

[0009] Step 1: Select the UAV aerial photography dataset as the dataset;

[0010] Step 2: Uniformly convert the four-parameter and eight-parameter annotation methods of the obtained dataset into five-parameter annotation, and use a rotation box with angle specificity to improve the detection accuracy;

[0011] Step 3: Divide the dataset into a training set, a validation set, and a test set according to a certain ratio;

[0012] Step 4: Construct a Cross-latitude Attention model and a ClA-DLANet model

[0013] Step 5: Test on the test dataset using the trained model, evaluate the model with evaluation metrics, and adjust the parameters;

[0014] Detect the rotating aircraft through the above steps.

[0015] In Step 2, the specific process of converting the dataset annotation method to five-parameter annotation is as follows:

[0016] Step 2-1: Obtain the coordinates of 4 points from the annotated image, which are (x1, y1), (x2, y2), (x3, y3), and (x4, y4);

[0017] Step 2-2: Transform the 8 parameters of the obtained 4-point coordinates into 5 parameters through the following formula. The specific formula is as follows:

[0018] x = (x1 + x2 + x3 + x4) / 4

[0019] y = (y1 + y2 + y3 + y4) / 4

[0020]

[0021]

[0022] θ = atan(y2 - y1, x2 - x1)

[0023] Where x is the average abscissa mean, y is the average ordinate mean, W is the width of the calculated rotated bounding box, H is the height of the calculated rotated bounding box, a is the proportionality coefficient, and θ is the angle between the rotated bounding box and the horizontal direction. Through the conversion of the dataset marking method, the rotated bounding box can be made more accurate.

[0024] In Step 4, the designed Cross-latitude Attention model is as follows:

[0025] The first output is connected to the input of the first MaxPooling layer 1. The output of the first MaxPooling layer 1 is connected to the input of the first Average Pooling layer 1. The output of the first Average Pooling layer 1 is connected to the input of the first Conv layer 1. The first output and the output of the first Conv layer 1 are multiplied to obtain the first result;

[0026] The second output is connected to the input of the second MaxPooling layer 1, the output of the second MaxPooling layer 1 is connected to the input of the second Average Pooling layer 1, the output of the second Average Pooling layer 1 is connected to the input of the second Conv layer 1, and the second output is multiplied by the output of the second Conv layer 1 to obtain the result of the second time;

[0027] The third output is connected to the input of the third MaxPooling layer 1, the output of the third MaxPooling layer 1 is connected to the input of the third Average Pooling layer 1, the output of the third Average Pooling layer 1 is connected to the input of the third Conv layer 1, and the third output is multiplied by the output of the third Conv layer 1 to obtain the result of the third time;

[0028] The results of the first time, the second time, and the third time are added together to obtain the final output.

[0029] In step 4, the constructed ClA-DLANet model is as follows:

[0030] The input is connected to the input of the Convolution layer 1, the output of the Convolution layer 1 is connected to the input of the Convolution layer 2, the output of the Convolution layer 2 is connected to the input of the Cross-latitude Attention layer 1, the output of the Cross-latitude Attention layer 1 is connected to the input of the HDA layer 1, the output of the HDA layer 1 is connected to the input of the Cross-latitude Attention layer 2, the output of the Cross-latitude Attention layer 2 is connected to the input of the HDA layer 2, the output of the HDA layer 2 is connected to the input of the HDA layer 3, the output of the HDA layer 3 is connected to the input of the HDA layer 4, the output of the HDA layer 4 is connected to the input of the Cross-latitude Attention layer 3, and the Cross-latitude Attention layer 3 is output.

[0031] In step 4, the Authentic SmoothL1 Loss function designed for the ClA-DLANet model is as follows:

[0032] The calculation formula of the Authentic SmoothL1 Loss function (hereinafter referred to as ASL1) is as follows:

[0033]

[0034] Where d is the difference value between the predicted box and the ground truth box, and μ is a parameter that, combined with the angle, can effectively reduce the sensitivity to outliers.

[0035] Compared with the prior art, the present invention has the following technical effects:

[0036] 1) The present invention adopts a carefully designed Cross-latitude Attention. This attention mechanism module is divided into three branches. The first branch interacts with the channel C and spatial H dimensions to obtain better feature expression ability; the second branch interacts with the channel C and spatial W dimensions to enhance the feature representation ability; the third branch introduces deformable convolution DCNv3 to simulate the geometric changes in the spatial dimension, continuously refining and denoising the feature map. Through the fusion of the three branches, the angle feature learning becomes more obvious, improving the accuracy of object detection.

[0037] 2) The present invention proposes an Authentic SmoothL1 Loss function, which is more conducive to the regression of the predicted angle.

[0038] 3) In the experiment on verifying the effectiveness of the attention mechanism of the present invention, compared with the network without the attention mechanism, the Precision, Recall, and F1 are increased by 1.22%, 0.76%, and 0.97% respectively. Compared with other mainstream attention mechanisms, there are also improvements to varying degrees.

[0039] 4) In the experiment on verifying the effectiveness of the angle loss of the present invention, there is also a small improvement compared with other loss functions. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The present invention will be further described below in conjunction with the drawings and embodiments:

[0041] Figure 1 is the algorithm design flow chart of the present invention;

[0042] Figure 2 is the Cross-latitude Attention model diagram of the present invention;

[0043] Figure 3 is the diagram of converting the annotation method of the Drone Vehicle dataset into a five-parameter marker;

[0044] Figure 4 is the DLANet model diagram;

[0045] Figure 5 is the ClA-DLANet model diagram;

[0046] Figure 6 is the DLANet detection effect diagram;

[0047] Figure 7 It is the detection effect diagram of the ClA-DLANet model. Specific implementation manner

[0048] A remote sensing rotating aircraft detection algorithm based on vector regression and attention mechanism. Through the designed Cross-latitude Attention, this method can better focus on the local features of the target, the background features around the target, and the global features of the entire image, thereby improving the effect of target detection. At the same time, the designed loss function has a smaller deviation between the predicted angle and the true value during the rotation box regression, and the gradient curve of the loss function is smoother, which is more conducive to the regression of the predicted angle.

[0049] It includes the following steps:

[0050] Step 1: Select the DroneVehicle large-scale UAV aerial photography dataset as the dataset for this experiment;

[0051] Step 2: Uniformly convert the four-parameter and eight-parameter annotation methods used in the DroneVehicle dataset into five-parameter annotations, and use rotation boxes with angle specificity to improve the detection accuracy;

[0052] Step 3: Divide the DroneVehicle dataset into a training set, a validation set, and a test set in the ratio of 8:1:1;

[0053] Step 4: Construct the Cross-latitude Attention model and the ClA-DLANet model

[0054] Step 5: Test on the test dataset through the trained model, evaluate the model through evaluation indicators, and adjust the parameters;

[0055] A rotating aircraft detection algorithm is proposed through the above steps.

[0056] As Figure 2 shown, in Step 4, the constructed Cross-latitude Attention model is as follows:

[0057] The first output is connected to the input of the first MaxPooling layer 1, the output of the first MaxPooling layer 1 is connected to the input of the first Average Pooling layer 1, the output of the first Average Pooling layer 1 is connected to the input of the first Conv layer 1, and the first output is multiplied by the output of the first Conv layer 1 to obtain the first result;

[0058] The second output is connected to the input of the second MaxPooling layer 1. The output of the second MaxPooling layer 1 is connected to the input of the second Average Pooling layer 1. The output of the second Average Pooling layer 1 is connected to the input of the second Conv layer 1. The second output is multiplied by the output of the second Conv layer 1 to obtain the second result;

[0059] The third output is connected to the input of the third MaxPooling layer 1. The output of the third MaxPooling layer 1 is connected to the input of the third Average Pooling layer 1. The output of the third Average Pooling layer 1 is connected to the input of the third Conv layer 1. The third output is multiplied by the output of the third Conv layer 1 to obtain the third result;

[0060] The first result, the second result, and the third result are added together to obtain the final output.

[0061] Its structure is as follows:

[0062] Input Tensor1 → First Max Pooling; First Max Pooling → First AveragePooling; First Average Pooling → First Conv; First Conv, Input Tensor1 → OutputTensor1;

[0063] Input Tensor2 → Second Max Pooling; Second Max Pooling → Second AveragePooling; Second Average Pooling → Second Conv; Second Conv, Input Tensor2 → OutputTensor2;

[0064] Input Tensor3 → Third Max Pooling; Third Max Pooling → Third AveragePooling; Third Average Pooling → DCNv3; Third Conv, Input Tensor3 → OutputTensor3;

[0065] Output Tensor1, Output Tensor2, Output Tensor3 → Output Tensor.

[0066] The extraction of remote sensing aircraft feature information is enhanced through the Cross-latitude Attention model to improve the detection accuracy

[0067] As Figure 5 shown, in step 4, the constructed ClA-DLANet model is as follows:

[0068] The input is connected to the input of the first Convolution layer. The output of the first Convolution layer is connected to the input of the second Convolution layer. The output of the second Convolution layer is connected to the input of the first Cross-latitude Attention layer. The output of the first Cross-latitude Attention layer is connected to the input of the first HDA layer. The output of the first HDA layer is connected to the input of the second Cross-latitude Attention layer. The output of the second Cross-latitude Attention layer is connected to the input of the second HDA layer. The output of the second HDA layer is connected to the input of the third HDA layer. The output of the third HDA layer is connected to the input of the fourth HDA layer. The output of the fourth HDA layer is connected to the input of the third Cross-latitude Attention layer, and the output of the third Cross-latitude Attention layer is outputted.

[0069] Its structure is:

[0070] Input → First Convolution; First Convolution → Second Convolution; Second Convolution → First Cross-latitude Attention;

[0071] First Cross-latitude Attention → First HDA; First HDA → Second Cross-latitude Attention; Second Cross-latitude Attention → Second HDA; Second HDA → Third HDA; Third HDA → Third Cross-latitude Attention;

[0072] Third Cross-latitude Attention → Output.

[0073] Through the ClA-DLANet model, a network for optimizing remote sensing aircraft detection is built from beginning to end, retaining more target features of aircraft and improving the detection accuracy of the model.

[0074] This method captures the interaction information in the channel and spatial dimensions through the designed Cross-latitude Attention mechanism module, enabling the channel dimension to perceive the H and W spatial dimensions respectively. The deformable convolution DCNv3 is used to simulate the geometric changes in the spatial dimension, so as to better focus on the local features of the target, the background features around the target, and the global features of the entire image, thereby improving the effect of target detection. At the same time, the designed loss function has a smaller deviation between the predicted angle and the true value during the rotation box regression, and the gradient curve of the loss function is smoother, which is more conducive to the regression of the predicted angle.

[0075] The designed Authentic SmoothL1 Loss function in step 4 is as follows:

[0076] The calculation formula of the Authentic SmoothL1 Loss function (hereinafter referred to as ASL1) is as follows:

[0077]

[0078] Where d is the difference value between the predicted box and the true box, and μ is a parameter that can effectively reduce the sensitivity to outliers when combined with the angle.

[0079] Complete the network construction, start training, and save the trained model;

[0080] Use the trained model to test on the test set, and evaluate it using the Precision, Recall, and F1 metric parameters. Among them, Precision represents the proportion of true positives among the samples classified as positive examples, Recall represents the proportion of true positives among the samples classified as positive examples, and F1 is a metric that comprehensively considers Precision and Recall. It is the harmonic mean of Precision and Recall. The value range of the F1 value is from 0 to 1, and the closer it is to 1, the better the performance of the model. The F1 value combines Precision and Recall, considering both the accuracy of positive examples and the integrity of positive examples;

[0081] The calculation formula of Precision is as follows:

[0082] Precision = TP / (TP + FP)

[0083] Where TP is a parameter representing the number of true positives (the number of samples correctly predicted as positive examples by the model), and FP is also a parameter representing the number of false positives (the number of samples incorrectly predicted as positive examples by the model).

[0084] The calculation formula of Recall is as follows:

[0085] Recall = TP / (TP + FN)

[0086] Where FN is a parameter representing false negatives (the number of samples that the model wrongly predicts as negatives).

[0087] The formula for F1 is as follows:

[0088] F1 = 2 * Precision * Recall / (Precision + Recall).

[0089] Example:

[0090] The code of the present invention is implemented based on the Pytorch framework, and the model is trained using an NVIDIA 3070 GPU under Ubuntu. The dataset uses the large-scale drone aerial photography dataset of Drone Vehicle published by Tianjin University, mainly for vehicle detection and vehicle counting tasks. The dataset contains five types of labels: car, truck, bus, van, and freight car. The annotation methods use the eight-parameter method [(x1, y1), (x2, y2), (x3, y3), (x4, y4)] and the four-parameter method (x, y, w, h), and the dataset is uniformly transformed into five-parameter markings (x, y, W, H, θ) after processing. The dataset contains RGB images and infrared images, which can reflect the real distribution of urban vehicles in China from the perspective of drone aerial photography to a certain extent, and has relatively high requirements for the generalization of the detection algorithm. In this paper, 3362 RGB images are selected from the Drone Vehicle dataset, including 161642 annotation instances. The image environment is mainly daytime, with occlusions and scale changes of targets in the real environment. To ensure better experimental results, during the training phase, we perform data augmentation on the data, including random flipping, adding noise, etc., to enhance the robustness of the detector. The dataset is divided into a training set, a validation set, and a test set according to 8:1:1. The backbone feature extraction network selects DLANet to verify the effectiveness of the attention mechanism and the effectiveness of the angle loss. Among them, for the verification of the effectiveness of the attention mechanism, two groups of experiments are carried out. In Experiment 1, Cross-latitude Attention is added to DLANet and compared with the network without it; in Experiment 2, under the same experimental conditions, Cross-latitude Attention and other mainstream attention mechanisms are added to the DLANet network to compare the detection effects of the models; for the verification of the effectiveness of the angle loss, the backbone feature extraction network selects DLANet, and the experimental group introduces L1 Loss, SmoothL1Loss, and Authentic SmoothL1 Loss as the angle loss respectively, and compares the detection effects of the models under the same experimental conditions. The results of Experiment 1 are shown in Table 1.

[0091] Table 1 Comparison of Cross-latitude Attention Performance

[0092] Attention mechanism Precision Recall F1 FLOPS None 95.86 86.56 90.97 7.22GB Cross-latitude Attention 97.08 87.32 91.94 7.22GB

[0093] Based on the DLANet network, a Cross-latitude Attention module is newly added. As can be seen from Table 1, after adding the Cross-latitude Attention module, the floating-point operation amount does not increase significantly, and the accuracy, recall rate, and F1 have been improved to a certain extent, which are 1.22%, 0.76%, and 0.97% higher than the baseline respectively.

[0094] The results of Experiment 2 are shown in Table 2.

[0095] Table 2 Performance Comparison of Different Attention Mechanisms

[0096] Attention mechanism Precision Recall F1 CBAM 98.32 56.89 72.08 Triplet Attention 98.1 39.15 55.97 SENet 96.89 85.16 90.65 SKNet 96.67 64.69 77.51 Shuffle Attention 96.6 69.46 80.81 Cross-latitude Attention 97.08 87.32 91.94

[0097] From the performance comparison results of different attention mechanisms (all added at the same position as Cross-latitude Attention) in Table 2, it can be seen that compared with DLANet, the networks integrated with attention mechanisms have different improvements in the accuracy performance of the detection effect, but the recall rate has not been improved, and even has decreased to varying degrees. In the previous chapter, we discussed the impact of the "angle" instance on the detection performance. As an instance in space, the angle is mainly affected by the feature space. Among them, attention mechanisms such as CBAM and Triplet Attention involve spatial attention, calculate the mean value for different channels in the same pixel point, and then obtain the Spitial Attention Mask through operations such as convolution and upsampling, giving different weights to the pixel points of the spatial features. Channel attention mechanisms such as SENet assign weights to each channel C and do not involve the spatial level, so the impact on the "angle" instance is not obvious. Compared with other attention mechanisms, Cross-latitude Attention introduces pixel offsets in the spatial dimension, increases the angle learning in the spatial dimension, and improves the overall performance of the network.

[0098] The experimental results of the angle effectiveness verification are shown in Table 3.

[0099] Table 3 Comparison of Different Angle Losses

[0100] Angle loss Precision Recall F1 L1 Loss 97.51 63.39 76.84 Smooth L1 Loss 95.86 86.56 90.97 Authentic Smooth L1 Loss 97.1 89.08 92.91

[0101] The feature extraction network for angle loss effectiveness comparison is DLANet. Based on DLANet, the experimental group sets the angle loss function based on L1, Smooth L1, and Authentic Smooth L1 respectively. The regression effect of L2 loss for angles is too poor to be described in the comparison table. As can be seen from Table 3, the algorithms based on the DLANet feature extraction network have an accuracy of 95.86% for the angle loss based on L1 Loss and SmoothL1 Loss respectively. The overall accuracy of the angle loss model based on Authentic SmoothL1 Loss in this paper is 97.1%, which is 1.24% higher than that of SmoothL1 Loss. Moreover, the recall rate and F1 index have been significantly improved, fully demonstrating the effectiveness of the angle loss based on Authentic SmoothL1 Loss proposed in this paper in the rotating object detection task. Specifically, by adding the long-side offset of the angle to the bounding box loss, the angle loss corrects the errors generated during training in the rotating bounding box, which helps the model better understand the rotating bounding box and its relationship with the target, thereby improving the detection accuracy and further reducing the redundant information repetition of the background.

[0102] From the results of the above three groups of experiments, it can be seen that the Cross-latitude Attention attention mechanism module and the angle loss function based on Authentic SmoothL1 Loss designed in the present invention have achieved remarkable results on the Drone Vehicle large-scale drone aerial photography dataset. There is a certain improvement in terms of Precision, Recall, and F1 compared to the baseline network, improving the accuracy of object detection and having good application prospects.

[0103] In summary, the network precision (Precision), recall rate (Recall), and F1 of the ClA-DLANet model incorporating Cross-latitude Attention are 97.1, 89.08, and 92.91 respectively. Compared with the original DLANet network, they have increased by 1.24%, 2.52%, and 1.94%. The increase in precision (Precision) indicates that the constructed model can detect more airplanes, and the increase in recall rate (Recall) indicates that the constructed model has fewer missed detections of airplanes.

[0104] Figure 6 It is the visualization detection effect diagram of the original network DLANet. Figure 7This is the visualization detection effect diagram of ClA-DLANet. From the perspective of actual application scenarios, in the face of high-altitude and long-distance aircraft scenarios, the ClA-DLANet model can detect more aircraft, with high accuracy and few missed detections, and has good application value.

Claims

1. A remote sensing rotating aircraft detection method based on vector regression and attention mechanism, characterized in that It includes the following steps: Step 1: Select the UAV aerial photography dataset as the dataset; Step 2: Uniformly convert the four-parameter and eight-parameter annotation methods of the obtained dataset into five-parameter annotation, and use a rotation box with angle specificity to improve the detection accuracy; Step 3: Divide the dataset into a training set, a validation set, and a test set according to a certain ratio; Step 4: Construct a Cross-latitude Attention model and a ClA-DLANet model; Step 5: Test on the test dataset through the trained model, evaluate the model through evaluation indicators, and adjust the parameters; Detect the rotating aircraft through the above steps; In Step 4, the constructed ClA-DLANet model is as follows: The input is connected to the input of the first layer of Convolution, the output of the first layer of Convolution is connected to the input of the second layer of Convolution, the output of the second layer of Convolution is connected to the input of the first layer of Cross-latitude Attention, the output of the first layer of Cross-latitude Attention is connected to the input of the first layer of HDA, the output of the first layer of HDA is connected to the input of the second layer of Cross-latitude Attention, the output of the second layer of Cross-latitude Attention is connected to the input of the second layer of HDA, the output of the second layer of HDA is connected to the input of the third layer of HDA, the output of the third layer of HDA is connected to the input of the fourth layer of HDA, the output of the fourth layer of HDA is connected to the input of the third layer of Cross-latitude Attention, and the third layer of Cross-latitude Attention outputs.

2. The method according to claim 1, wherein In Step 2, the specific process of converting the dataset annotation method into five-parameter annotation is as follows: Step 2-1: Obtain the coordinates of 4 points from the annotated image, which are (x1, y1), (x2, y2), (x3, y3), and (x4, y4) respectively; Step 2-2: Transform the 8 parameters of the obtained 4-point coordinates into 5 parameters through the following formula. The specific formula is as follows: where x is the average abscissa mean, y is the average ordinate mean, W is the width of the calculated rotated bounding box, H is the height of the calculated rotated bounding box, is the proportionality coefficient, and θ is the angle between the rotated bounding box and the horizontal direction.

3. The method according to claim 1, characterized in that In Step 4, the designed Cross-latitudeAttention model is as follows: The first output is connected to the input of the first layer of the first MaxPooling, the output of the first layer of the first MaxPooling is connected to the input of the first layer of the first Average Pooling, the output of the first layer of the first Average Pooling is connected to the input of the first layer of the first Conv, and the first output is multiplied by the output of the first layer of the first Conv to obtain the first result; The second output is connected to the input of the second MaxPooling layer 1, the output of the second MaxPooling layer 1 is connected to the input of the second Average Pooling layer 1, the output of the second Average Pooling layer 1 is connected to the input of the second Conv layer 1, and the second output is multiplied by the output of the second Conv layer 1 to obtain the result of the second time; The third output is connected to the input of the third MaxPooling layer 1, the output of the third MaxPooling layer 1 is connected to the input of the third Average Pooling layer 1, the output of the third Average Pooling layer 1 is connected to the input of the third Conv layer 1, and the third output is multiplied by the output of the third Conv layer 1 to obtain the result of the third time; The results of the first time, the second time, and the third time are added together to obtain the final output.

Citation Information

Patent Citations

  • Lightweight target detection method for unmanned aerial vehicle to inspect insulator

    CN116385911A