Aerial target detection method based on cross weighted pixel reconstruction

By using the cross-weighted pixel reconstruction and feature selection path aggregation network module in drone detection, the problem of low stability and accuracy of target detection is solved, more efficient feature extraction and fusion is achieved, and detection performance is improved.

CN119992373AActive Publication Date: 2025-05-13CHINA SATELLITE MARITIME MEASUREMENT & CONTROL DEPT
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202411971610.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The prior art has problems in the detection of drones with low target detection stability and low detection accuracy, especially in complex backgrounds and low resolution images.

Method used

The aerial object detection method based on cross-weighted pixel reconstruction is adopted, and the spatial information and channel information of the feature map are mined by constructing the CPR module and the FS-PANet network module, and the quality of the low-level feature map is improved through the guiding feature selection path aggregation network module.

Benefits of technology

It improves the stability and accuracy of object detection, can more accurately locate and identify the target of interest, reduces interference from background noise, and improves feature fusion efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992373A_ABST
    Figure CN119992373A_ABST
Patent Text Reader

Abstract

The invention discloses an aerial target detection method based on cross weighted pixel reconstruction, which comprises the following steps: acquiring an image data set containing an aerial target, and making a training sample; the method comprises the following steps: constructing an air target detection network model comprising a CPR module and an FS-PANet network module by taking a Yolov10 network as a basic network; inputting a training sample into the air target detection network model for training, and storing parameters of the trained network model; and inputting a to-be-detected image into the trained aerial target detection network model for aerial target detection. The method has the remarkable effects that direction perception and position perception information can be captured by mining information of the height, width and channel direction of the input feature map, so that the model can more accurately position and identify an interested target, and noise contained in the low-resolution feature map can be well removed through average pooling operation, so that the accuracy of the target recognition is improved. Interference enhancement target features are reduced; and the feature fusion efficiency is improved, so that the target detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of aerial target detection, and in particular to an aerial target detection method based on cross-weighted pixel reconstruction. Background Art

[0002] In recent years, drones have been increasingly used in urban management, precision agriculture, environmental monitoring, traffic control, military operations, and disaster relief due to their flexibility, multiple functions, low cost, and ease of operation. However, the rapid popularization of drones has also caused many problems, such as the use of drones for illegal attacks, interference, or surveillance, which seriously affects social security and public privacy. Therefore, effective detection of drones is crucial and has become a hot topic in current research. The current mainstream drone detection methods mainly include radar detection, audio detection, optoelectronic visual detection, etc. Compared with other detection technologies, optoelectronic visual detection technology has the characteristics of more intuitive detection results, higher accuracy, and wider application scenarios. It is the mainstream method for detecting drones.

[0003] However, drones are developing in the direction of miniaturization. Especially when drones are far away from sensors, the imaging scale of drones in optical sensors is significantly reduced, and there are problems such as less shape texture and degradation of apparent features. In addition, drones have weak contrast with the background, low signal-to-clutter ratio, and are easily interfered by cloud clutter and flying birds. The above unfavorable factors have posed serious challenges to image detection. Traditional target detection methods usually rely on manually designed features, which are difficult to accurately represent the complex and changeable target morphology and appearance. Traditional target detection methods usually use fixed feature extractors and classifiers. When faced with complex scenes, they often show poor detection performance and are difficult to be effectively applied in drone detection.

[0004] In recent years, with the rapid progress of GPU parallel computing technology, deep learning algorithms have developed rapidly and have been widely used in the field of target detection. At present, the mainstream target detection algorithms based on deep learning are mainly divided into two categories. One is the two-stage detection algorithm, such as R-CNN (Convolutional Neural Network, CNN), Fast R-CNN, Faster R-CNN, Mask R-CNN, etc. These algorithms first generate region candidate boxes on the input image, and then extract and classify the image features. They have high detection accuracy, but the model parameter scale is large and the detection speed is slow, which makes it difficult to meet the requirements of high real-time detection. The other is the single-stage algorithm, which mainly includes SSD (Single-Shot MultiBoxDetector), RetinaNet, YOLO (You Only Look Once), etc. The single-stage algorithm directly converts the target frame positioning into a regression problem for solution. It does not need to generate candidate boxes, and the detection speed is greatly improved, but the detection accuracy is not as good as the two-stage detection algorithm. Therefore, in order to maintain a high detection speed and detection accuracy for drones, it is necessary to improve the detection algorithm based on deep learning.

[0005] To this end, a research on anti-UAV target detection algorithm based on YOLOv5s-AntiUAV[J]. (Tan Liang, Zhao Liangjun, Zheng Liping, Xiao Bo, Electro-Optics and Control, 2024, 31(5):40-45.) discloses a detection method for drones. The method first introduces the Slim Neck paradigm combined with deep hyperparameter convolution to enhance the algorithm's feature extraction and representation learning capabilities; secondly, the SPD-Conv module is introduced in the backbone and neck networks respectively to avoid information loss caused by strided convolution and pooling, thereby improving the algorithm's detection performance of small targets in low-resolution images; finally, the loss function is optimized, and Alpha-CIoULoss is used to replace CIOU Loss in YOLOv5s to enhance the algorithm's versatility.

[0006] However, this method still has the following disadvantages:

[0007] 1. The feature extraction capability is not strong, which affects the stability of target detection. This technical solution assigns weights to different channels by calculating the feature correlation in the channel dimension. However, it only learns the degree of correlation between channels and cannot capture the spatial information of the image, which makes it difficult to effectively extract targets in complex backgrounds, greatly affecting the stability of target detection.

[0008] 2. The feature fusion capability is not strong, which affects the target detection accuracy. The feature fusion method of this technical solution can effectively utilize feature maps of different levels, which is conducive to multi-scale target detection. However, this feature fusion method will also dilute the semantic information of the feature to a certain extent, thereby losing some important information inside the feature map, resulting in a decrease in target detection accuracy. Summary of the invention

[0009] In view of the shortcomings of the prior art, the purpose of the present invention is to provide an aerial target detection method based on cross-weighted pixel reconstruction, which can effectively solve the problems of low target detection stability and low target detection accuracy existing in the background technology by constructing a cross-weighted pixel reconstruction module to fully mine the spatial information and channel information of the feature map, and by designing an instructive feature selection path aggregation network module to improve the quality of the low-level feature map.

[0010] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0011] A method for detecting aerial targets based on cross-weighted pixel reconstruction includes the following steps:

[0012] Step 1: Obtain an image dataset containing aerial targets and create training samples;

[0013] Step 2: Using the Yolov10 network as the basic network, build an aerial target detection network model including a CPR module and a FS-PANet network module;

[0014] The CPR module is used to capture different dimensional information along the height, width and channel directions of the feature map after two convolutions in the backbone of the Yolov10 network, and generate cross weights to reconstruct the pixels of the feature map;

[0015] The FS-PANet network module is used to filter lower-level features from top to bottom using adjacent higher-level features for feature maps of different scales output by the main part of the Yolov10 network, and fuse the filtered lower-level features with the higher-level features. The fused features are integrated from bottom to top into features of different scales and output to the head part of the Yolov10 network;

[0016] Step 3: input the training samples into the aerial target detection network model for training, and save the trained network model parameters;

[0017] Step 4: Input the image to be detected into the trained aerial target detection network model for aerial target detection.

[0018] Furthermore, the preparation process of the training samples is as follows:

[0019] Step 1.1, select a data source and obtain a sample set from the data source;

[0020] Step 1.2: Label the images in the sample set;

[0021] Step 1.3, cut the labeled image to the required size;

[0022] Step 1.4: Convert the label format to the required format to obtain training samples.

[0023] Furthermore, the training samples are composed of a training set and a test set, the training set is used to train the constructed aerial target detection network model, and the test set is used to evaluate the performance of the aerial target detection network model after training.

[0024] Furthermore, the CPR module processes the input feature map as follows:

[0025] Calculate the channel-wise feature weights of the input feature map;

[0026] For the input feature map, features are aggregated in the height and width directions to obtain feature maps in the height and width directions respectively;

[0027] Apply the channel-wise weights to the feature maps in the height and width directions, and calculate the feature weights in the height and width directions respectively;

[0028] The feature weights in the height and width directions are crossed to obtain cross weights in multiple dimensions, and the cross weights are dot-multiplied with the input feature map to recalibrate the input feature map to obtain the output feature map.

[0029] Furthermore, the CPR module includes a global average pooling unit, a height average pooling unit, a width average pooling unit, three BConv units, three Sigmoid function units, and four Hadamard product units. The input feature map is respectively input into the input ends of the global average pooling unit, the height average pooling unit, and the width average pooling unit. The output end of the global average pooling unit is connected to the input end of the first Sigmoid function unit via the first BConv unit, the output end of the height average pooling unit is connected to the first input end of the first Hadamard product unit via the second BConv unit, and the output end of the width average pooling unit is connected to the first input end of the second Hadamard product unit via the third BConv unit. The first input end, the two output ends of the first Sigmoid function unit are respectively connected to the second input end of the first Hadamard product unit and the second input end of the second Hadamard product unit, the output end of the first Hadamard product unit is connected to the first input end of the third Hadamard product unit via the second Sigmoid function unit, the output end of the second Hadamard product unit is connected to the second input end of the third Hadamard product unit via the third Sigmoid function unit, the output end of the third Hadamard product unit is connected to the first input end of the fourth Hadamard product unit, the second input end of the fourth Hadamard product unit inputs the input feature map, and the output end of the fourth Hadamard product unit outputs the output feature map, wherein:

[0030] The global average pooling unit is used to aggregate channel-wise features in the input feature map;

[0031] The height average pooling unit is used to aggregate the features in the height direction of the input feature map;

[0032] The width average pooling unit is used to aggregate the features in the width direction of the input feature map;

[0033] Three BConv units are used to learn the parameters;

[0034] The three Sigmoid function units are used to calculate the feature weights in the channel, height, and width directions respectively;

[0035] The first Hadamard product unit is used to apply the feature weight of the channel direction to the feature map in the height direction;

[0036] The second Hadamard product unit is used to apply the feature weights in the channel direction to the feature map in the width direction;

[0037] The third Hadamard product unit is used to cross the feature weights in the height and width directions to obtain cross weights in multiple dimensions;

[0038] The fourth Hadamard product unit is used to apply the cross weights to the input feature map to obtain the output feature map.

[0039] Furthermore, the calculation formula of the feature weight in the channel direction is:

[0040] ω ch =σ(conv 1×1 (ReLU(conv 1×1 (GAP(F)))))

[0041] Among them, ω ch is the weight in the channel direction, GAP(·) represents the global average pooling with a pooling size of 1×H×W, and conv 1×1 (·) represents 1×1 convolution, ReLU(·) represents activation function, and σ(·) represents sigmoid activation function;

[0042] The calculation formula of the feature weight in the height direction is:

[0043] ω height =σ(F height ⊙expand(ω ch ))

[0044] The calculation formula of the feature weight in the width direction is:

[0045] ω width =σ(F width ⊙expand(ω ch ))

[0046] Where expand(·) represents the dimension expansion operation, and ⊙ represents the Hadamard product.

[0047] Furthermore, the FS-PANet network module includes two AFS modules, three C2f modules, a convolution module, two Concat modules, a SCDown downsampling module and a C2fCIB module, the two input ends of the first AFS module are respectively input with the high-level feature map and the middle-level feature map extracted from the main part of the Yolov10 network, the output end of the first AFS module is connected to the input end of the first C2f module, the two output ends of the first C2f module are respectively connected to an input end of the second AFS module and an input end of the first Concat module, the other input end of the second AFS module is input with the low-level feature map extracted from the main part of the Yolov10 network, and the The output ends of the two AFS modules are connected to the input end of the second C2f module, one output end of the second C2f module is connected to the other input end of the first Concat module via the convolution module, the output end of the first Concat module is connected to one input end of the second Concat module via the third C2f module and the SCDown downsampling module, the other input end of the second Concat module inputs the high-level feature map, the output end of the second Concat module is connected to the input end of the C2fCIB module, and the other output end of the second C2f module, the other output end of the third C2f module, and the output end of the C2fCIB module are all connected to the head part of the Yolov10 network.

[0048] Furthermore, the AFS module includes a transposed convolution unit, a CA attention unit, a multiplication unit, and a feature fusion unit. The transposed convolution unit is used to perform a transposed convolution on the input higher-level feature map to obtain an intermediate feature map of the same size as the lower-level feature map. The CA attention unit is used to calculate the weight of each feature channel. The multiplication unit is used to multiply the channel weight of the obtained higher-level feature map with the lower-level feature map to enhance the useful features in the lower-level feature map. The feature fusion unit is used to perform feature fusion on the higher-level feature map and the enhanced lower-level feature map and output it.

[0049] Furthermore, the AFS module calculates the weight of each feature channel using the following formula:

[0050] ω=σ(BConv(GAP 1×2H×2W (C' high ))+BConv(GMP 1×2H×2W (C' high )))

[0051] Among them, ω represents the weight of each feature channel, σ(·) represents the sigmoid activation function, GAP 1×2H×2W (·) GMP 1×2H×2W(·) represent global average pooling and global maximum pooling with pooling size of 1×2H×2W, respectively.

[0052] Furthermore, the loss function of the aerial target detection network model is:

[0053]

[0054] Among them, L inner-MPDIoU is the loss function, IoU inner is the inner intersection and union ratio, d1 is the distance between the upper left corner coordinates of the true bounding box and the predicted bounding box, d2 is the distance between the lower right corner coordinates of the true bounding box and the predicted bounding box, and w and h are the width and height of the input feature map respectively.

[0055] The remarkable effects of the present invention are:

[0056] 1. Improved the stability of target detection. The present invention mines information on the height, width, and channel direction of the input feature map. On the one hand, it can capture direction perception and position perception information, allowing the model to more accurately locate and identify targets of interest. On the other hand, the average pooling operation can effectively remove the noise contained in the low-resolution feature map, reducing interference and strengthening the target features.

[0057] 2. Improved target detection accuracy. The present invention fuses adjacent feature maps, uses higher-level feature maps to filter important features of adjacent lower-level feature maps, and then fuses the filtered lower-level features with higher-level features, further integrating them into features of different scales, thereby improving the feature fusion efficiency and thus improving the target detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a flow chart of the method of the present invention;

[0059] Figure 2 It is the overall structure diagram of the aerial target detection network model of the present invention;

[0060] Figure 3 It is the structural diagram of the CPR module;

[0061] Figure 4 This is the structural diagram of the FS-PANet network module. DETAILED DESCRIPTION

[0062] The specific implementation manner and working principle of the present invention are further described in detail below with reference to the accompanying drawings.

[0063] In this embodiment, in order to solve the problem that drones are small in size and easily disturbed by similar-shaped targets such as flying birds in drone detection tasks, a Cross-weighted Pixel Reconstructionand Feature Selection Network CFS-YOLO is proposed based on Yolov10n as the basic network. First, through the CPR module added in the backbone part, it is possible to fully mine the information of different dimensions of the feature map, generate cross weights to reconstruct the pixels of the feature map, so as to reduce background noise interference and improve the feature expression ability of the network; then, a selective feature fusion mechanism is designed to replace the neck part in the traditional Yolov10n, by using the rich semantic information of the high-level feature map to guide the low-level feature map to select important features, and fully fuse the selected features in a top-down and bottom-up bidirectional manner, effectively combining the detailed information of the feature maps of different layers to improve the performance of the model when dealing with small targets, similar targets and occluded targets. The specific embodiments are described as follows:

[0064] This embodiment provides an aerial target detection method based on cross-weighted pixel reconstruction, such as Figure 1 As shown, the specific steps are as follows:

[0065] Step 1: Obtain an image dataset containing aerial targets and create training samples;

[0066] Step 1.1, select a data source and obtain a sample set from the data source;

[0067] Step 1.2: Label the images in the sample set;

[0068] Step 1.3, cut the labeled image to the required size;

[0069] Step 1.4: Convert the label format to the required format to obtain training samples.

[0070] In this patent, the training samples are composed of a training set and a test set, and the ratio of the training set to the test set is 8:2. The training set is used to train the constructed aerial target detection network model, and the test set is used to evaluate the performance of the aerial target detection network model after training.

[0071] Step 2: Using the Yolov10n network as the basic network, a network model for aerial target detection including a CPR module and a FS-PANet network module is constructed;

[0072] In this example, the aerial target detection network model is constructed as follows Figure 2As shown, the CPR module is added to the main part of the Yolov10n network to mine information about the height, width, and channel direction of the input feature map, and the FS-PANet network module is used to replace the neck part of the Yolov10n network to fuse adjacent feature maps. Specifically:

[0073] The CPR (Cross-weighted Pixel Reconstruction) module is used to capture different dimensional information along the height, width and channel directions of the feature map after two convolutions in the main part of the Yolov10 network, and use the average pooling operation to reduce information redundancy and suppress background noise along the height and width directions, thereby highlighting the features of the target of interest; use the global average pooling operation to aggregate channel information, assign weights to channels according to the importance of different channels, and guide the model to use limited resources to learn important channel features. Further, for the different dimensional information after the pooling operation, the BConv module composed of pointwise convolution and ReLU activation function is used to learn the parameters. The attention weights of different dimensions are generated by the Sigmoid function, and the different weights are crossed and applied to the original feature map, and each pixel of the original feature map is reconstructed to reduce background interference while improving the model's recognition ability for the target of interest.

[0074] like Figure 3 As shown, the CPR module includes a global average pooling unit, a height average pooling unit, a width average pooling unit, three BConv units, three Sigmoid function units, and four Hadamard product units. The input feature map is respectively input into the input ends of the global average pooling unit, the height average pooling unit, and the width average pooling unit. The output end of the global average pooling unit is connected to the input end of the first Sigmoid function unit via the first BConv unit, the output end of the height average pooling unit is connected to the first input end of the first Hadamard product unit via the second BConv unit, and the output end of the width average pooling unit is connected to the first input end of the second Hadamard product unit via the third BConv unit. The first input end of the first Hadamard product unit is connected to the second input end of the first Hadamard product unit and the second input end of the second Hadamard product unit respectively. The output end of the first Hadamard product unit is connected to the first input end of the third Hadamard product unit via the second Sigmoid function unit. The output end of the second Hadamard product unit is connected to the second input end of the third Hadamard product unit via the third Sigmoid function unit. The output end of the third Hadamard product unit is connected to the first input end of the fourth Hadamard product unit. The second input end of the fourth Hadamard product unit inputs the input feature map, and the output end of the fourth Hadamard product unit outputs the output feature map, wherein:

[0075] The global average pooling unit is used to aggregate channel-wise features in the input feature map;

[0076] The height average pooling unit is used to aggregate the features in the height direction of the input feature map;

[0077] The width average pooling unit is used to aggregate the features in the width direction of the input feature map;

[0078] Three BConv units are used to learn the parameters;

[0079] The three Sigmoid function units are used to calculate the feature weights in the channel, height, and width directions respectively;

[0080] The first Hadamard product unit is used to apply the feature weight of the channel direction to the feature map in the height direction;

[0081] The second Hadamard product unit is used to apply the feature weights in the channel direction to the feature map in the width direction;

[0082] The third Hadamard product unit is used to cross the feature weights in the height and width directions to obtain cross weights in multiple dimensions;

[0083] The fourth Hadamard product unit is used to apply the cross weights to the input feature map to obtain the output feature map.

[0084] In this example, the CPR module is processing the input feature map The processing steps are as follows:

[0085] Step A1: The input image is convolved twice on the backbone to obtain the input feature map Calculate input feature map The feature weight of the channel direction is calculated as follows:

[0086] ω ch =σ(conv 1×1 (ReLU(conv 1×1 (GAP(F))))) (1)

[0087] Among them, ω ch is the weight in the channel direction, GAP(·) represents the global average pooling with a pooling size of 1×H×W, and conv 1×1 (·) represents 1×1 convolution, ReLU(·) represents activation function, and σ(·) represents sigmoid activation function.

[0088] Step A2: Use the average pooling operation and BConv unit to aggregate the features of the input feature map in the height and width directions to obtain feature maps in the height and width directions respectively. The calculation process is as follows:

[0089] F height =conv 1×1 (ReLU(conv 1×1 (AvgPool 1×H×1 (F)))) (2)

[0090] F width =conv 1×1 (ReLU(conv 1×1 (AvgPool 1×1×W (F)))) (3)

[0091] Among them, F height 、F width They are feature maps in height and width directions respectively, AvgPool 1×H×1 (·) represents the average pooling with a pooling size of 1×H×1, AvgPool 1×1×W (·) is 1×1×W average pooling.

[0092] Step A3: Apply the channel direction weight to the feature maps in the height and width directions, and calculate the feature weights in the height and width directions respectively;

[0093] Among them, the calculation formula of the feature weight in the height direction is:

[0094] ω height =σ(F height ⊙expand(ω ch )) (4)

[0095] Among them, the calculation formula of the feature weight in the width direction is:

[0096] ω width =σ(F width ⊙expand(ω ch )) (5)

[0097] In the formula, expand(·) represents the dimension expansion operation, and v represents the Hadamard product.

[0098] Step A4: Cross the feature weights in the height and width directions to obtain cross weights in multiple dimensions, perform dot multiplication of the cross weights with the input feature map, and recalibrate the input feature map to obtain the output feature map;

[0099] The calculation formula for the output feature map is as follows:

[0100] F'=Fv(expand(ω width )vexpand(ω height )) (6)

[0101] In this example, in order to improve the efficiency of feature fusion at different layers, a FS-PANet (Feature Selection Path Aggregation Network) network module is proposed to replace the neck part of the traditional Yolov10n network. First, an Adjacent Feature Selection (AFS) module is constructed to filter the low-level features from top to bottom using adjacent high-level features, and then fuse the filtered low-level features with the high-level features. The fused features further integrate features of different scales from bottom to top and output them to the head part of the Yolov10n network.

[0102] Specifically, the structure of the FS-PANet network module is as follows: Figure 4 As shown, it includes two AFS modules, three C2f modules, a convolution module, two Concat modules, a SCDown downsampling module and a C2fCIB module. The two input ends of the first AFS module are respectively input into the high-level feature map and the middle-level feature map extracted from the backbone of the Yolov10 network. The output end of the first AFS module is connected to the input end of the first C2f module. The two output ends of the first C2f module are respectively connected to an input end of the second AFS module and an input end of the first Concat module. The other input end of the second AFS module is input into the low-level feature map extracted from the backbone of the Yolov10 network. The output end is connected to the input end of the second C2f module, one output end of the second C2f module is connected to the other input end of the first Concat module via the convolution module, the output end of the first Concat module is connected to an input end of the second Concat module via the third C2f module and the SCDown downsampling module, the other input end of the second Concat module inputs the high-level feature map, the output end of the second Concat module is connected to the input end of the C2fCIB module, and the other output end of the second C2f module, the other output end of the third C2f module, and the output end of the C2fCIB module are all connected to the head part of the Yolov10 network.

[0103] exist Figure 4In the figure, C2, C3, C4, and C5 are feature maps extracted from the backbone. As the scale of the feature map decreases, the receptive field of the feature map on the input image gradually increases. Therefore, the high-level feature map contains more semantic information than the low-level feature map. Using the adjacent high-level feature map to filter the low-level feature map can improve the quality of the low-level features, which is conducive to distinguishing similar target features and improving the fusion efficiency of features at different layers. The main part of the FS-PANet network module that implements adjacent feature filtering and then fusion is the AFS module. The AFS module takes the feature maps of two adjacent layers as input, and the high-level feature map is adjusted to the same scale as the low-level feature map through a 3×3 transposed convolution, that is, Figure 4 The feature maps M3, M4, and M5 in the image are used to measure the importance of each feature channel using the CA (Channel attention) attention module, and the channel weight is output. The high-level channel weight is multiplied by the low-level feature map to enhance the useful features in the low-level feature map and suppress the useless features. The feature maps P3, P4, and P5 after feature fusion are obtained, thereby improving the efficiency of feature fusion, making full use of semantic information, and reducing information redundancy in feature fusion.

[0104] from Figure 4 It can also be seen that the AFS module includes a transposed convolution unit, a CA attention unit, a multiplication unit, and a feature fusion unit. The transposed convolution unit is used to perform a transposed convolution on the input higher-level feature map to obtain an intermediate feature map of the same size as the lower-level feature map. The CA attention unit is used to calculate the weight of each feature channel. The multiplication unit is used to multiply the channel weight of the obtained higher-level feature map with the lower-level feature map to enhance the useful features in the lower-level feature map. The feature fusion unit is used to perform feature fusion on the higher-level feature map and the enhanced lower-level feature map and output it.

[0105] Based on the above structure, the main processing process of the AFS module is as follows:

[0106] Step B1: Take two adjacent feature maps As input, for the high-level feature map C high Perform transposed convolution to obtain the intermediate feature map C' high , C' high And the low-level feature map C low The dimensions are the same, and the calculation formula is:

[0107] C′ high =TransConv(C high ) (7)

[0108] Among them, TransConv(·) is a 3×3 transposed convolution with a stride of 2.

[0109] Step B2: Intermediate feature map C' high After the CA attention unit performs global average pooling (GAP) and global maximum pooling (GMP) to extract weight features, the calculation formula is as follows:

[0110] ω=σ(BConv(GAP 1×2H×2W (C' high ))+BConv(GMP 1×2H×2W (C' high ))) (8)

[0111] Among them, ω represents the weight of each feature channel, σ(·) represents the sigmoid activation function, GAP 1×2H×2W (·) GMP 1×2H×2W (·) represent global average pooling and global maximum pooling with pooling size of 1×2H×2W, respectively.

[0112] The extracted weight features are then added after an excitation operation, and then converted into a weight of 0 to 1 through a sigmoid function.

[0113] Step B3, the multiplication unit multiplies the obtained channel weight of the higher-level feature map by the lower-level feature map to enhance the useful features in the lower-level feature map;

[0114] Step B4: The feature fusion unit performs feature fusion on the higher-level feature map and the enhanced lower-level feature map and outputs the result.

[0115] In this example, MPDIoU loss is used in combination with Inner-IoU based on auxiliary bounding boxes to calculate the regression loss of the bounding box. Compared with MPDIoU, Inner-MPDIoU uses auxiliary bounding boxes of different scales to calculate the loss, pays more attention to the core part of the bounding box, can make more accurate judgments on overlapping areas, and effectively accelerates the bounding box regression process. The calculation process of the Inner-MPDIoU loss function is as follows:

[0116]

[0117]

[0118] Among them, L inner-MPDIoU is the loss function, IoU inner is the inner intersection over union (Inner-IoU), d1 is the true bounding box Β gt and the predicted bounding box Β prd The distance between the upper left corner coordinates, d2 is the real bounding box Βgt and the predicted bounding box Β prd The distance between the lower right corner coordinates of the predicted bounding box Β prd The width and height are w prd 、h prd , the coordinates of its upper left and lower right corners are The coordinates of its center point are The true bounding box Β gt The width and height are w gt 、h gt , the coordinates of its upper left and lower right corners are The coordinates of its center point are w and h are the width and height of the input feature map respectively.

[0119] In this embodiment, the calculation formula of the inner intersection over union (Inner-IoU) is as follows:

[0120]

[0121] uinon=(w gt *h gt )*ratio 2 +(w prd *h prd )*ratio 2 -inter (12)

[0122]

[0123]

[0124] Here, ratio is the scale factor, and its value range is [0.5, 1.5].

[0125] Step 3: input the training samples into the aerial target detection network model for training, and save the trained network model parameters;

[0126] This example uses Windows 10 system, 16G memory, Intel(R) Core(TM) i7-13700KF CPU, and the experimental environment is python 3.11.9, pytorch 2.0.1, and cuda 11.8. All models are trained and tested on NVIDIA RTX 4080 GPU.

[0127] Training and evaluation process:

[0128] ①Use the conda create-n yolov10 python=3.11.9 command to create a virtual environment named yolov10 and activate the environment;

[0129] ②Then install PyTorch and related dependent libraries;

[0130] ③ According to the specific situation of the data set, modify the corresponding yolov10.yaml configuration file to specify the number of categories, the paths of the training set and the test set, and other information;

[0131] ④ Use the training command to start training. After the training is completed, the model is evaluated on the test set and the evaluation indicators such as accuracy, recall rate, mAP and other test results are output.

[0132] Step 4: Input the image to be detected into the trained aerial target detection network model for aerial target detection.

[0133] To sum up, in order to solve the problems of low target detection stability and low target detection accuracy in the background technology, the present invention, firstly, mines information on the height, width, and channel direction of the input feature map through the CPR module. On the one hand, it can capture direction perception and position perception information, so that the model can more accurately locate and identify the target of interest. On the other hand, the average pooling operation can well remove the noise contained in the low-resolution feature map, reduce interference and enhance the target features; secondly, the FS-PANet network module is used to fuse adjacent feature maps, and the higher-level feature maps are used to filter the important features of the adjacent lower-level feature maps. Then, the filtered lower-level features are fused with the higher-level features, and further integrated into features of different scales, thereby improving the feature fusion efficiency and thus improving the detection accuracy of the target.

[0134] The technical solution provided by the present invention is described in detail above. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A method for aerial target detection based on cross-weighted pixel reconstruction, characterized in that: The steps include: Step 1: Obtain an image dataset containing aerial targets and create training samples; Step 2: Using the Yolov10 network as the basic network, build an aerial target detection network model including a CPR module and a FS-PANet network module; The CPR module is used to capture different dimensional information along the height, width and channel directions of the feature map after two convolutions in the backbone of the Yolov10 network, and generate cross weights to reconstruct the pixels of the feature map; The FS-PANet network module is used to filter lower-level features from top to bottom using adjacent higher-level features for feature maps of different scales output by the main part of the Yolov10 network, and fuse the filtered lower-level features with the higher-level features. The fused features are integrated from bottom to top into features of different scales and output to the head part of the Yolov10 network; Step 3: input the training samples into the aerial target detection network model for training, and save the trained network model parameters; Step 4: Input the image to be detected into the trained aerial target detection network model for aerial target detection.

2. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 1, characterized in that: The process of making the training samples is as follows: Step 1.1, select a data source and obtain a sample set from the data source; Step 1.2: Label the images in the sample set; Step 1.3, cut the marked image to the required size; Step 1.4: Convert the label format to the required format to obtain training samples.

3. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 2, characterized in that: The training samples are composed of a training set and a test set, wherein the training set is used to train the constructed aerial target detection network model, and the test set is used to evaluate the performance of the aerial target detection network model after training.

4. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 1, characterized in that: The CPR module processes the input feature map as follows: Calculate the channel-wise feature weights of the input feature map; For the input feature map, features are aggregated in the height and width directions to obtain feature maps in the height and width directions respectively; Apply the channel-wise weights to the feature maps in the height and width directions, and calculate the feature weights in the height and width directions respectively; The feature weights in the height and width directions are crossed to obtain cross weights in multiple dimensions, and the cross weights are dot-multiplied with the input feature map to recalibrate the input feature map to obtain the output feature map.

5. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 4 is characterized in that: The CPR module includes a global average pooling unit, a height average pooling unit, a width average pooling unit, three BConv units, three Sigmoid function units, and four Hadamard product units. The input feature map is respectively input into the input ends of the global average pooling unit, the height average pooling unit, and the width average pooling unit. The output end of the global average pooling unit is connected to the input end of the first Sigmoid function unit via the first BConv unit, the output end of the height average pooling unit is connected to the first input end of the first Hadamard product unit via the second BConv unit, and the output end of the width average pooling unit is connected to the first input end of the second Hadamard product unit via the third BConv unit. , the two output ends of the first Sigmoid function unit are respectively connected to the second input end of the first Hadamard product unit and the second input end of the second Hadamard product unit, the output end of the first Hadamard product unit is connected to the first input end of the third Hadamard product unit via the second Sigmoid function unit, the output end of the second Hadamard product unit is connected to the second input end of the third Hadamard product unit via the third Sigmoid function unit, the output end of the third Hadamard product unit is connected to the first input end of the fourth Hadamard product unit, the second input end of the fourth Hadamard product unit inputs the input feature map, and the output end of the fourth Hadamard product unit outputs the output feature map, wherein: The global average pooling unit is used to aggregate channel-wise features in the input feature map; The height average pooling unit is used to aggregate the features in the height direction of the input feature map; The width average pooling unit is used to aggregate the features in the width direction of the input feature map; Three BConv units are used to learn the parameters; The three Sigmoid function units are used to calculate the feature weights in the channel, height, and width directions respectively; The first Hadamard product unit is used to apply the feature weight of the channel direction to the feature map in the height direction; The second Hadamard product unit is used to apply the feature weights in the channel direction to the feature map in the width direction; The third Hadamard product unit is used to cross the feature weights in the height and width directions to obtain cross weights in multiple dimensions; The fourth Hadamard product unit is used to apply the cross weights to the input feature map to obtain the output feature map.

6. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 5, characterized in that: The calculation formula of the feature weight in the channel direction is: ω ch =σ(conv 1×1 (ReLU(conv 1×1 (GAP(F))))); The calculation formula of the feature weight in the height direction is: oh height =σ(F height ⊙expand(ω ch )) The calculation formula of the feature weight in the width direction is: oh width =σ(F width ⊙expand(ω ch )) Where expand(·) represents the dimension expansion operation, and ⊙ represents the Hadamard product.

7. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 1, characterized in that: The FS-PANet network module includes two AFS modules, three C2f modules, a convolution module, two Concat modules, a SCDown downsampling module and a C2fCIB module. The two input ends of the first AFS module are respectively input with the high-level feature map and the middle-level feature map extracted from the main part of the Yolov10 network. The output end of the first AFS module is connected to the input end of the first C2f module. The two output ends of the first C2f module are respectively connected to an input end of the second AFS module and an input end of the first Concat module. The other input end of the second AFS module is input with the low-level feature map extracted from the main part of the Yolov10 network. The output end of the FS module is connected to the input end of the second C2f module, one output end of the second C2f module is connected to the other input end of the first Concat module via the convolution module, the output end of the first Concat module is connected to one input end of the second Concat module via the third C2f module and the SCDown downsampling module, the other input end of the second Concat module inputs the high-level feature map, the output end of the second Concat module is connected to the input end of the C2fCIB module, and the other output end of the second C2f module, the other output end of the third C2f module, and the output end of the C2fCIB module are all connected to the head part of the Yolov10 network.

8. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 7, characterized in that: The AFS module includes a transposed convolution unit, a CA attention unit, a multiplication unit, and a feature fusion unit. The transposed convolution unit is used to perform a transposed convolution on the input higher-level feature map to obtain an intermediate feature map of the same size as the lower-level feature map. The CA attention unit is used to calculate the weight of each feature channel. The multiplication unit is used to multiply the channel weight of the obtained higher-level feature map with the lower-level feature map to enhance the useful features in the lower-level feature map. The feature fusion unit is used to perform feature fusion on the higher-level feature map and the enhanced lower-level feature map and output them.

9. The aerial target detection method based on cross-weighted pixel reconstruction according to claim 8, characterized in that: The AFS module calculates the weight of each feature channel using the following formula: ω=σ(BConv(GAP 1×2H×2W (C' high ))+BConv(GMP 1×2H×2W (C' high ))) Among them, ω represents the weight of each feature channel, σ(·) represents the sigmoid activation function, GAP 1×2H×2W (·) GMP 1 ×2H×2W (·) represent global average pooling and global maximum pooling with pooling size of 1×2H×2W, respectively.

10. The aerial target detection method based on cross-weighted pixel reconstruction according to any one of claims 1 to 9, characterized in that: The loss function of the aerial target detection network model is: Among them, L inner-MPDIoU is the loss function, IoU inner is the inner intersection and union ratio, d1 is the distance between the upper left corner coordinates of the true bounding box and the predicted bounding box, d2 is the distance between the lower right corner coordinates of the true bounding box and the predicted bounding box, and w and h are the width and height of the input feature map respectively.

Citation Information

Patent Citations

  • Low, small and slow target detection method based on adjacent scale weight distribution feature fusion

    CN114926718A

  • YOLOV4 remote sensing target detection method fusing feature transfer and attention mechanism

    CN115497005A

  • SAR small target detection based on coordinate awareness attention and spatial semantic context

    CN116385873A

  • Joint segmentation tracking depth estimation model training method and use method

    CN117115786A

  • Road surface crack detection method based on edge reconstruction network

    CN118096672A