Unmanned aerial vehicle detection method based on polarization information enhancement and polarization attention characteristics
By performing multi-channel splitting and demosaic processing on infrared polarized mosaic images, combined with YOLOv8 backbone network and attention module, the problem of failure to effectively utilize polarization information in the existing technology is solved, and accurate detection of drone targets is achieved.
Patent Information
- Application Number
- CN202510381996.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art fails to effectively utilize the polarization information in infrared polarization mosaic images, resulting in low detection performance of drone targets, and noise affects the detection process after demosaic and polarization information are solved.
Through polarization decoding student network, multi-channel splitting and demosaic processing of infrared polarization mosaic images is generated to generate multi-dimensional polarization infographics, and preliminary features are extracted using YOLOv8's backbone network, and target positioning features are generated through global spatial, multi-dimensional polarization and local spatial attention modules, and finally input the detection head for drone detection.
The contrast between the drone target and the background is improved, the precise detection of drone targets is achieved, and the impact of noise on the detection process is reduced.
Smart Images

Figure CN120388306A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a drone detection method, device, medium, and equipment based on polarization information enhancement and polarization attention features. Background Art
[0002] Infrared polarization imaging technology can greatly improve the contrast between drone targets and the background, reducing the difficulty of detecting and identifying drone targets. However, traditional infrared polarization imaging methods such as time-sharing imaging methods, amplitude-splitting imaging methods, and aperture-splitting imaging methods have defects such as inability to image in real time, excessive volume, and mutual interference between channels, making it difficult to be effectively applied to the drone target detection scenario. In recent years, with the development of micro-nano processing technology, the focal-plane infrared polarization imaging method with advantages such as small volume, low weight, and strong real-time performance has gradually received extensive attention in the field of target detection and other fields.
[0003] The polarization mosaic image obtained by the focal-plane infrared polarization imaging method needs to be demosaicked to restore different polarization direction images of the same size as the original image to calculate the image polarization information. Existing target detection methods do not consider the polarization information contained in the polarization mosaic image. If the infrared polarization mosaic image is directly applied to the drone target detection process, it is difficult to use the polarization information in the image to improve the target detection performance; while if the image is demosaicked and the polarization information is calculated before detection, a large amount of noise contained in the calculated polarization information will affect the detection process of the algorithm, reducing the detection performance of the algorithm. Summary of the Invention
[0004] The main purpose of this application is to provide a drone detection method, device, medium, and equipment based on polarization information enhancement and polarization attention features, aiming to solve the technical problem that traditional methods do not consider the polarization information contained in the mosaic image and cannot effectively utilize the infrared polarization information containing the physical and chemical characteristics of the target surface, resulting in low detection performance of drone targets.
[0005] To achieve the above purpose, this application provides a drone detection method based on polarization information enhancement and polarization attention features, including: performing multi-channel splitting and demosaicking on the infrared polarization mosaic image through a polarization decoding student network to generate a multi-dimensional polarization information map; extracting features from the multi-dimensional polarization information map based on the backbone network of YOLOv8, and outputting preliminary features including semantic and texture levels; sequentially inputting the preliminary features into a global spatial attention module, a multi-dimensional polarization attention module, and a local spatial attention module to generate target localization features, and inputting the target localization features into a detection head to obtain the drone detection result.
[0006] Optionally, the infrared polarization mosaic image is subjected to multi-channel splitting and demosaicing by the polarization decoding student network to generate a multi-dimensional polarization information map, including: reconstructing the infrared polarization mosaic image into sub-images of four polarization channels through the first layer of the polarization decoding student network; after splitting the sub-images of the four polarization channels through the second layer of the polarization decoding student network, constructing denoised and enhanced information for each channel through the third layer of the polarization decoding student network; constructing four-channel feature information by combining the denoised and enhanced information of each channel with the sub-images through the second layer of the polarization decoding student network; reconstructing a mosaic multi-dimensional polarization information map by combining the four-channel feature information with the infrared polarization mosaic image through the first layer of the polarization decoding student network.
[0007] Optionally, the training process of the polarization decoding student network includes: obtaining a 4-channel demosaicing ground truth map captured by an infrared camera, calculating the weighted sum of the 4-channel demosaicing ground truth map to obtain an S0 information map; calculating the first Stokes vector and the second Stokes vector based on the 4-channel demosaicing ground truth map, and obtaining a Dolp information map based on the ratio of the root of the sum of the squares of the first Stokes vector and the second Stokes vector to the S0 information map; obtaining an Aolp information map based on the weighted value of the arccotangent value of the ratio of the second Stokes vector to the first Stokes vector; obtaining a ground truth information map based on the S0 information map, the Dolp information map, and the Aolp information map, inputting the infrared polarization mosaic image into the polarization decoding student network to obtain a first multi-dimensional polarization information map; inputting the infrared polarization mosaic image into the teacher network to obtain a second multi-dimensional polarization information map; calculating the first L1 loss between the first multi-dimensional polarization information map and the ground truth information map; calculating the second L1 loss between the first multi-dimensional polarization information map and the second multi-dimensional polarization information map; using the sum of the first L1 loss and the second L1 loss as the objective function to train the polarization decoding student network.
[0008] Optionally, the preliminary features are sequentially input into the global spatial attention module, the multi-dimensional polarization attention module, and the local spatial attention module to generate target localization features, and the target localization features are input into the detection head to obtain the UAV detection result, including: extracting the global spatial attention features in the preliminary features based on the global spatial attention module; extracting the multi-dimensional polarization attention features in the global spatial attention features based on the multi-dimensional polarization attention module; extracting the local spatial attention features in the global spatial attention features based on the local spatial attention module; inputting the global spatial attention features, the multi-dimensional polarization attention features, and the local spatial attention features into the detection head and outputting the UAV detection result.
[0009] Optionally, extracting the global spatial attention feature in the preliminary feature by the global spatial attention module includes: performing three groups of independent convolutions on the preliminary feature to obtain a query feature Q containing approximate position information of the target, a key feature K containing different hierarchical information, and a convolutional preliminary value V containing preliminary attention information of the target; performing matrix multiplication on the query feature Q and the key feature K and then normalizing by Softmax to obtain a normalized feature; performing weighted summation on the normalized feature and the convolutional preliminary value V to obtain the global spatial attention feature.
[0010] Optionally, extracting the multi-dimensional polarization attention feature in the global spatial attention feature by the multi-dimensional polarization attention module includes: performing average pooling and max pooling on each channel of the global spatial attention feature respectively to obtain pooling results; accumulating the pooling results to obtain the multi-dimensional polarization attention feature.
[0011] Optionally, extracting the local spatial attention feature in the global spatial attention feature by the local spatial attention module includes: performing dilated convolutions with a convolutional kernel size of 1×1, a convolutional kernel size of 3×3, and a convolutional kernel size of 5×5 on the global spatial attention feature in parallel to obtain three-way output features; concatenating the three-way output features to obtain the local spatial attention feature.
[0012] In addition, to achieve the above object, a drone detection device based on polarization information enhancement and polarization attention features includes: a polarization information generation module for performing multi-channel splitting and demosaicing processing on an infrared polarization mosaic image through a polarization decoding student network to generate a multi-dimensional polarization information map; a preliminary feature generation module for extracting features from the multi-dimensional polarization information map based on the backbone network of YOLOv8 and outputting preliminary features including semantic and texture levels; a detection module for sequentially inputting the preliminary features into a global spatial attention module, a multi-dimensional polarization attention module, and a local spatial attention module to generate target localization features, and inputting the target localization features into a detection head to obtain a drone detection result.
[0013] To achieve the above object, the present application also provides a computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the drone detection method based on polarization information enhancement and polarization attention features provided in the above embodiments.
[0014] To achieve the above object, the present application also provides an electronic device, which includes: at least one processor, a memory, and an input / output unit; wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the drone detection method based on polarization information enhancement and polarization attention features provided in any of the foregoing embodiments.
[0015] A drone detection method, device, medium, and equipment based on polarization information enhancement and polarization attention features proposed in an embodiment of the present application perform multi-channel splitting and demosaicing on an infrared polarization mosaic image through a polarization decoding student network to generate a multi-dimensional polarization information map; extract features from the multi-dimensional polarization information map based on the backbone network of YOLOv8 and output preliminary features including semantic and texture levels; sequentially input the preliminary features into a global spatial attention module, a multi-dimensional polarization attention module, and a local spatial attention module to generate target localization features, and input the target localization features into a detection head to obtain drone detection results. The present application can obtain multi-dimensional polarization information from a polarization mosaic image through a polarization decoding student network, use different polarization information to enhance the contrast between the target and the background, obtain preliminary features through the backbone network of YOLOv8 for the multi-dimensional polarization information, and process the preliminary features through a polarization attention feature network constructed by a global spatial attention module, a multi-dimensional polarization attention module, and a local spatial attention module to achieve precise detection of drone targets on the multi-dimensional polarization information with significantly enhanced contrast. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 FIG. 6 is a schematic flowchart provided for an embodiment of the drone detection method based on polarization information enhancement and polarization attention features of the present application;
[0017] Figure 2 FIG. 10 is a general flowchart of the method provided for an embodiment of the drone detection method based on polarization information enhancement and polarization attention features of the present application;
[0018] Figure 3 FIG. 14 is a schematic diagram of the polarization decoding student network provided for an embodiment of the drone detection method based on polarization information enhancement and polarization attention features of the present application;
[0019] Figure 4 FIG. 18 is a schematic diagram of the polarization attention feature network provided for an embodiment of the drone detection method based on polarization information enhancement and polarization attention features of the present application;
[0020] Figure 5 FIG. 22 is a diagram of the drone detection result provided for an embodiment of the drone detection method based on polarization information enhancement and polarization attention features of the present application;
[0021] Figure 6 FIG. 26 is a block diagram of the functional module structure provided for an embodiment of the drone detection device based on polarization information enhancement and polarization attention features in the present application.
[0022] The implementation, functional features, and advantages of the objectives of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] It should be understood that the specific embodiments described herein are merely used to explain the present application and are not intended to limit the present application.
[0024] Since the target detection method based on the split focal plane infrared polarization imaging method in the prior art does not consider the polarization information contained in the polarization mosaic image, if the infrared polarization mosaic image is directly applied to the UAV target detection process, it is difficult to utilize the polarization information in the image to improve the target detection performance; and if the image is demosaicked and the polarization information is calculated before detection, a large amount of noise contained in the calculated polarization information will affect the detection process of the algorithm and reduce the detection performance of the algorithm. In view of the problems existing in the traditional method, a UAV detection method based on polarization information enhancement and polarization attention features is designed. This method can obtain multi-dimensional polarization information from the polarization mosaic image, use different polarization information to enhance the contrast between the target and the background, and achieve accurate detection of UAV targets by constructing a polarization attention feature network.
[0025] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Figure 2 is the overall flowchart of the method of the embodiment of the present invention. As Figure 2 shown, it is the overall flowchart of the method of the embodiment of the present invention. First, use an infrared polarization camera to capture a UAV image, and use the obtained infrared polarization mosaic image as the input of the network; secondly, use a polarization decoding student network to obtain a demosaicked multi-dimensional polarization information image from the input infrared polarization mosaic image I M to enhance the saliency of the UAV target; then send the enhanced multi-dimensional information image I D into the detection backbone network YOLOv8 to obtain three layers of preliminary features F Y ; finally, use a polarization attention module to construct an attention feature map F A by using the attention mechanism, and obtain the final detection result I R .
[0027] Refer to Figure 1 , Figure 1 is a schematic flowchart provided by an embodiment of the UAV detection method based on polarization information enhancement and polarization attention features of the present application. The present application provides a UAV detection method based on polarization information enhancement and polarization attention features, and this method can be executed by a processor of a server or a terminal. The UAV detection method based on polarization information enhancement and polarization attention features may include:
[0028] S10. Perform multi-channel splitting and demosaicking processing on the infrared polarization mosaic image through a polarization decoding student network to generate a multi-dimensional polarization information map;
[0029] Among them, the processor first establishes a polarization decoding student network, obtains an infrared polarization mosaic image based on the infrared polarization camera shooting the drone target image, and uses the obtained infrared polarization mosaic image as the input of the network; then uses the established polarization decoding student network to obtain the demosaicked multi-dimensional polarization information image from the input infrared polarization mosaic image I M to enhance the saliency of the drone target.
[0030] In an embodiment of the present application, step S10 may include the following execution process:
[0031] S101. Reconstruct the infrared polarization mosaic image into sub-images of four polarization channels through the first layer of the polarization decoding student network;
[0032] S102. After splitting the sub-images of the four polarization channels through the second layer of the polarization decoding student network, construct the denoising and enhancement information of each channel through the third layer of the polarization decoding student network;
[0033] S103. Construct four-channel feature information by combining the denoising and enhancement information of each channel with the sub-images through the second layer of the polarization decoding student network;
[0034] S104. Reconstruct the mosaic multi-dimensional polarization information map by combining the four-channel feature information with the infrared polarization mosaic image through the first layer of the polarization decoding student network.
[0035] Among them, the polarization decoding student network is divided into three layers. The input infrared polarization mosaic image I M is split into 4 images of different polarization channels by the reconstruction module and sent to the second layer of the network; these 4 images of different infrared polarization channels are further split and sent to the third layer of the network to construct the denoising and enhancement information F Channel of each channel, and sent to the second layer of the network together with the images of different channels to construct the four-channel feature information F E , and this feature information will be sent to the backbone network of the first layer together with the input polarization mosaic image and used as auxiliary information to reconstruct the demosaicked multi-dimensional polarization information map I Channel . The sub-focal plane demosaicking dataset is used for training, and the PJDNDMnet is used as the teacher network to train the multi-dimensional polarization information map output by the polarization decoding student network. For the multi-dimensional information map I D output by the polarization decoding student network, I s0 , I Dolp , I Aolp and the multi-dimensional information map obtained from the ground truth image calculate the first L1 loss. For the multi-dimensional information map I s0 , I Dolp , I Aolp and the multi-dimensional information map obtained from the teacher network Calculate the second L1 loss. Based on the sum of the L1 loss and the L2 loss, jointly calculate the loss value L PDMSN .
[0036] In the specific execution process, Figure 3 This is the schematic diagram of the polarization decoding student network according to the embodiment of the present invention. As Figure 3 shown, the convolution module is composed of 1 convolution layer and 1 residual layer, which is used as the basic structure of the network to obtain convolution features; the residual module is composed of M convolution layers and 1 residual layer, and a shortcut link structure is introduced to avoid the disappearance of features caused by excessive convolution; the splitting step will split an image of H*W*1 into an image of H / 2*W / 2*4
[0037] The polarization decoding student network is divided into three layers. The input infrared polarization mosaic image I M is split into 4 images of different polarization channels by the reconstruction module I Channel , and is sent to the second layer of the network:
[0038] I Channel = Shuffle(I M ) (1)
[0039] where I Channel = [I0, I 45 , I 90 , I 135 , which are infrared polarization images of 0°, 45°, 90°, and 135° channels respectively, and the length and width dimensions are both 1 / 2 of the input polarization mosaic image. The second layer of the network further splits these 4 infrared polarization channel images into 16 images, and each channel is split into 4 images with the length and width being 1 / 2 of the original in the same way as the previous step
[0040] It is sent to the third layer of the network to construct the denoising and enhancement information F E of 0°, 45°, 90°, and 135° channels, and is sent to the second layer of the network together with I Channel = [I0, I 45 , I 90 , I 135 to construct the 4-channel feature information F Channel :
[0041] F Channel = SecondLayer([F E , I Channel ) (2)
[0042] where F Channel = [F0, F 45 , F 90 , F 135It represents the 4-channel feature information obtained by the second-layer network from the 4-channel image. This feature information will be used as auxiliary information and input into the first-layer backbone network composed of N convolutional modules together with the input polarized mosaic image to construct the mosaic multi-dimensional polarization information map I. D :
[0043] I D = FirstLayer([F Channel ,I M ) (3)
[0044] In an embodiment of the present application, the training process of the polarization decoding student network includes:
[0045] Obtain the 4-channel demosaicked ground truth map captured by the infrared camera, calculate the weighted sum of the 4-channel demosaicked ground truth map to obtain the S0 information map;
[0046] Calculate the first Stokes vector and the second Stokes vector based on the 4-channel demosaicked ground truth map, and obtain the Dolp information map based on the ratio of the root of the sum of the squares of the first Stokes vector and the second Stokes vector to the S0 information map;
[0047] Obtain the Aolp information map based on the weighted value of the arccotangent value of the ratio of the second Stokes vector to the first Stokes vector;
[0048] Obtain the ground truth information map based on the S0 information map, the Dolp information map, and the Aolp information map;
[0049] Input the infrared polarized mosaic image into the polarization decoding student network to obtain the first multi-dimensional polarization information map;
[0050] Input the infrared polarized mosaic image into the teacher network to obtain the second multi-dimensional polarization information map;
[0051] Calculate the first L1 loss between the first multi-dimensional polarization information map and the ground truth information map;
[0052] Calculate the second L1 loss between the first multi-dimensional polarization information map and the second multi-dimensional polarization information map;
[0053] Take the sum of the first L1 loss and the second L1 loss as the objective function to train the polarization decoding student network.
[0054] It should be noted that in order to ensure the overall real-time performance of the algorithm, the polarization decoding student network is designed with a simple structure and low computational complexity. However, this design will reduce the algorithm performance while improving the real-time performance. Therefore, the multi-dimensional information map obtained by the student network using the ground truth image While training, it is also necessary to conduct joint training with an already trained teacher network with high computational complexity and high accuracy, so that the student network can learn the output of the teacher network, thereby making the student network highly real-time and highly accurate.
[0055] In the specific implementation process, the multi-dimensional information graph obtained from the true value image is represented as follows:
[0056]
[0057] in This represents the four-channel demosaiced true values obtained by the infrared camera after rotating the external infrared polarizer placed in front of the ordinary infrared camera lens to 0°, 45°, 90°, and 135° respectively. Represents the Stokes vectors s1 and s2 calculated from the 4-channel true value image; Represents the S0, Dolp, and Aolp images obtained from the demosaiced true value image.
[0058] This application uses PJDNDMnet as a teacher network to train the multi-dimensional polarization information graph output by the polarization decoding student network, and uses L1 loss to train the polarization decoding student network with polarization mosaic image I M Multidimensional information diagram I obtained for input D =[I s0 ,I Dolp ,I Aolp ], the teacher network is used to polarize the mosaic image I M Multidimensional information graph obtained for input Multidimensional information graph obtained by calculating the true value image Calculate the loss value L together PDMSN :
[0059]
[0060] L PDMSGN =L GT +L Teacher (9)
[0061] in Teacher Network with I M The S0, Dolp, and Aolp images are obtained as inputs, λ1, λ2, ..., λ6 are weights, and the sum is 1; ||·||1 represents the L1 loss, L GT With L Teacher are the loss values between the multi-dimensional information graph and the ground-truth image and the output of the teacher network, respectively. As a result, the polarization decoding student network can extract polarization information with relatively less noise from the polarization mosaic image, thereby improving the saliency of the polarization information.
[0062] S20. Extract features from the multi-dimensional polarization information map based on the backbone network of YOLOv8, and output preliminary features including semantic and texture levels.
[0063] Figure 4 Schematic diagram of the polarization attention feature network according to an embodiment of the present invention. As Figure 4 shown, in this step, the multi-dimensional polarization information map output by the polarization decoding student network will be subjected to preliminary feature extraction through the backbone network based on Yolov8 to obtain different-level preliminary features F = [F1, F2, F3] including semantic information, texture information, etc.
[0064] S30. Input the preliminary features into the global spatial attention module, the multi-dimensional polarization attention module, and the local spatial attention module in sequence to generate target location features, and input the target location features into the detection head to obtain the UAV detection result.
[0065] In this step, the processor establishes a polarization attention feature network, uses the preliminary feature F as the input of this network to obtain the attention feature of the image, and quickly determines the target position to weaken the influence of the noise in the polarization information on the detection result.
[0066] Among them, the polarization attention feature network is divided into 3 sub-networks: a sub-network for obtaining global spatial attention information, a sub-network for obtaining multi-dimensional polarization attention information, and a sub-network for obtaining local spatial attention information. The preliminary features input into the polarization attention feature network will first obtain the global attention information A of the image through the global spatial attention network G to determine the approximate position of the target. Then, the global attention features will be sent to the multi-dimensional polarization attention network and the local spatial attention network respectively to obtain the multi-dimensional polarization attention A P and the local spatial attention A L , and use the polarization information to improve the detection accuracy. The architecture of the global spatial attention network is improved based on the Self-attention structure of the Transformer network. Through three groups of different convolution operations, the global spatial attention network will obtain the query feature Q containing the approximate position information of the target, the key feature K containing different-level information such as the texture, shape, and semantics of the target, and the convolutional preliminary value V containing the preliminary attention information of the target from the input preliminary feature F. After fusing the query feature and the key feature and performing softmax processing, and then weighted fusion with the convolutional preliminary value, the global attention information A of the image can be obtained G . The multi-dimensional polarization attention network adopts a channel attention architecture design. By respectively performing average pooling and max pooling on the different-channel features of the input global attention information, and adding the pooling results to obtain the multi-dimensional polarization attention feature A PAfter obtaining the approximate position of the target by global spatial attention, the local spatial attention module will further refine the target position using the attention mechanism. The local spatial attention module adopts an atrous convolution pyramid and uses 1×1 convolution, 3×3 atrous convolution, and 5×5 atrous convolution to obtain the local spatial attention feature A from the global attention feature. L Finally, the detection result of the UAV target is obtained through the detection head.
[0067] In an embodiment of the present application, step S30 may include the following execution process:
[0068] S301. Extract the global spatial attention feature in the preliminary feature based on the global spatial attention module;
[0069] S302. Extract the multi-dimensional polarization attention feature in the global spatial attention feature based on the multi-dimensional polarization attention module;
[0070] S303. Extract the local spatial attention feature in the global spatial attention feature based on the local spatial attention module;
[0071] S304. Input the global spatial attention feature, multi-dimensional polarization attention feature, and local spatial attention feature into the detection head, and output the UAV detection result.
[0072] Specifically, step S301 may include the following execution process:
[0073] S3011. Perform three groups of independent convolutions on the preliminary feature to obtain a query feature Q containing the approximate position information of the target, a key feature K containing different hierarchical information, and a convolutional preliminary value V containing the preliminary attention information of the target;
[0074] S3012. Perform matrix multiplication on the query feature Q and the key feature K and then perform Softmax normalization to obtain a normalized feature;
[0075] S3013. Perform weighted summation on the normalized feature and the convolutional preliminary value V to obtain the global spatial attention feature.
[0076] In the specific execution process, the architecture of the global spatial attention network is improved based on the Self-attention structure of the Transformer network. Through three groups of different convolutional operations, the global spatial attention network will obtain a query feature Q containing the approximate position information of the target, a key feature K containing different hierarchical information such as the texture, shape, and semantics of the target, and a convolutional preliminary value V containing the preliminary attention information of the target from the input preliminary feature F:
[0077] K = convblocks2(F); (10)
[0078] After fusing the query feature with the key feature, performing softmax processing, and then fusing it with the initial convolution value through weighting, the global attention information A of the image can be obtained. G :
[0079] A t = λ1softmax(Q, K) + λ2V (11)
[0080] where λ1 and λ2 represent weights.
[0081] Specifically, step S302 may include the following execution process:
[0082] S3021. Perform average pooling and max pooling on each channel of the global spatial attention feature respectively to obtain the pooling results;
[0083] S3022. Accumulate the pooling results to obtain the multi-dimensional polarization attention feature.
[0084] In the specific execution process, since the polarization decoding student network saves the multi-dimensional polarization information into different channels of the image, the multi-dimensional polarization attention network adopts a channel attention architecture design. By performing average pooling and max pooling on the feature of different channels of the input global attention information respectively, and accumulating the pooling results to obtain the multi-dimensional polarization attention feature A P :
[0085]
[0086] where represents the i-th channel of the global attention information.
[0087] Specifically, step S30 may include the following execution process:
[0088] S3031. Parallelly perform dilated convolution with a kernel size of 1×1, dilated convolution with a kernel size of 3×3, and dilated convolution with a kernel size of 5×5 on the global spatial attention feature to obtain three-way output features;
[0089] S3032. Concatenate the three-way output features to obtain the local spatial attention feature.
[0090] After obtaining the approximate position of the target from the global spatial attention, the local spatial attention module will further refine the target position using the attention mechanism. The local spatial attention module adopts a dilated convolution pyramid and uses 1*1 convolution, 3*3 dilated convolution, and 5*5 dilated convolution to obtain the local spatial attention feature A from the global attention feature L :
[0091]
[0092] where holoconvi Indicates the dilated convolution with a convolution kernel size of i.
[0093] Figure 5 This is the drone detection result diagram of the embodiment of the present invention. As Figure 5 shown, this is the drone detection result diagram of the embodiment of the present invention. The input infrared polarization mosaic image, after passing through the polarization decoding student network, the detection backbone network, and the polarization attention feature network, finally outputs the detection result of the drone target. It can be seen from the results that the detection results of this patent accurately detect the drone target in challenging scenarios such as interference situations, multiple targets, small volume, and false targets.
[0094] Refer to Figure 6 , on the basis of the above embodiments, the present application further provides a drone detection device based on polarization information enhancement and polarization attention features to solve other technical problems. The drone detection device 100 may include a polarization information generation module 101, a preliminary feature generation module 102, and a detection module 103:
[0095] The polarization information generation module 101 is used to perform multi-channel splitting and demosaicing processing on the infrared polarization mosaic image through the polarization decoding student network to generate a multi-dimensional polarization information map;
[0096] The preliminary feature generation module 102 is used to extract features from the multi-dimensional polarization information map based on the backbone network of YOLOv8 and output preliminary features including semantic and texture levels;
[0097] The detection module 103 is used to sequentially input the preliminary features into the global spatial attention module, the multi-dimensional polarization attention module, and the local spatial attention module to generate target localization features, and input the target localization features into the detection head to obtain the drone detection result.
[0098] On the basis of the above embodiments, the present application further provides a computer-readable storage medium, which includes instructions. When it runs on a computer, it enables the computer to perform the drone detection method based on polarization information enhancement and polarization attention features provided by any one of the above method embodiments.
[0099] On the basis of the above embodiments, the present application further provides an electronic device, which includes:
[0100] At least one processor, a memory, and an input-output unit;
[0101] Among them, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the drone detection method based on polarization information enhancement and polarization attention features provided by any one of the above method embodiments.
[0102] The above are only the preferred embodiments of the present application, which do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A UAV detection method based on polarization information enhancement and polarization attention features, characterized in that Including: Performing multi-channel splitting and demosaicing on the infrared polarization mosaic image through a polarization decoding student network to generate a multi-dimensional polarization information map; Extracting features from the multi-dimensional polarization information map based on the backbone network of YOLOv8 and outputting preliminary features including semantic and texture levels; Sequentially inputting the preliminary features into a global spatial attention module, a multi-dimensional polarization attention module, and a local spatial attention module to generate target localization features, and inputting the target localization features into a detection head to obtain the UAV detection result.
2. The UAV detection method based on polarization information enhancement and polarization attention features according to claim 1, wherein The process of performing multi-channel splitting and demosaicing on the infrared polarization mosaic image through the polarization decoding student network to generate a multi-dimensional polarization information map includes: Reconstructing the infrared polarization mosaic image into sub-images of four polarization channels through the first layer of the polarization decoding student network; Splitting the sub-images of each polarization channel through the second layer of the polarization decoding student network to obtain split images with both length and width being half of the original image; Constructing denoising and enhanced information of four polarization channels from each split image through the third layer of the polarization decoding student network; Constructing four-channel feature information from all the denoising and enhanced information and all the sub-images through the second layer of the polarization decoding student network; Reconstructing a mosaic multi-dimensional polarization information map from the four-channel feature information and the infrared polarization mosaic image through a backbone network composed of multiple convolutional modules.
3. The UAV detection method based on polarization information enhancement and polarization attention features according to claim 1, characterized in that The training process of the polarization decoding student network includes: Obtaining a 4-channel demosaicing ground truth map captured by an infrared camera, calculating the weighted sum of the 4-channel demosaicing ground truth map to obtain an S0 information map; Calculating the first Stokes vector and the second Stokes vector based on the 4-channel demosaicing ground truth map, and obtaining a Dolp information map based on the ratio of the root of the sum of the squares of the first Stokes vector and the second Stokes vector to the S0 information map; Obtaining an Aolp information map based on the weighted value of the arccotangent value of the ratio of the second Stokes vector to the first Stokes vector; Obtaining a multi-dimensional ground truth information map based on the S0 information map, the Dolp information map, and the Aolp information map; Inputting the infrared polarization mosaic image into the polarization decoding student network to obtain a first multi-dimensional polarization information map; Inputting the infrared polarization mosaic image into a teacher network to obtain a second multi-dimensional polarization information map; Calculating the first L1 loss between the first multi-dimensional polarization information map and the multi-dimensional ground truth information map; Calculating the second L1 loss between the first multi-dimensional polarization information map and the second multi-dimensional polarization information map; Using the sum of the first L1 loss and the second L1 loss as the objective function to train the polarization decoding student network.
4. The UAV detection method based on polarization information enhancement and polarization attention features according to claim 1, wherein Sequentially inputting the preliminary features into a global spatial attention module, a multi-dimensional polarization attention module, and a local spatial attention module to generate target localization features, and inputting the target localization features into a detection head to obtain the UAV detection result, including: Extracting global spatial attention features in the preliminary features based on the global spatial attention module; Extracting multi-dimensional polarization attention features in the global spatial attention features based on the multi-dimensional polarization attention module; Extracting local spatial attention features in the global spatial attention features based on the local spatial attention module; Input the global spatial attention feature, multi-dimensional polarization attention feature, and local spatial attention feature into the detection head, and output the UAV detection result.
5. The UAV detection method based on polarization information enhancement and polarization attention features according to claim 4, wherein The global spatial attention feature in the preliminary feature is extracted based on the global spatial attention module, including: Perform three groups of independent convolutions on the preliminary feature to obtain a query feature Q containing the approximate position information of the target, a key feature K containing different hierarchical information, and a convolutional preliminary value V containing the preliminary attention information of the target; Perform matrix multiplication on the query feature Q and the key feature K and then normalize it by Softmax to obtain a normalized feature; Perform weighted summation on the normalized feature and the convolutional preliminary value V to obtain the global spatial attention feature.
6. The UAV detection method based on polarization information enhancement and polarization attention features according to claim 4, wherein The multi-dimensional polarization attention feature in the global spatial attention feature is extracted based on the multi-dimensional polarization attention module, including: Perform average pooling and max pooling on each channel of the global spatial attention feature respectively to obtain a pooling result; Accumulate the pooling results to obtain the multi-dimensional polarization attention feature.
7. The UAV detection method based on polarization information enhancement and polarization attention features according to claim 4, characterized in that The local spatial attention feature in the global spatial attention feature is extracted based on the local spatial attention module, including: Parallelly perform dilated convolutions with a convolution kernel size of 1×1, a convolution kernel size of 3×3, and a convolution kernel size of 5×5 on the global spatial attention feature to obtain three-way output features; Concatenate the three-way output features to obtain the local spatial attention feature.
8. An unmanned aerial vehicle detection device based on polarization information enhancement and polarization attention features, characterized in that, Including: A polarization information generation module for performing multi-channel splitting and demosaicing on the infrared polarization mosaic image through a polarization decoding student network to generate a multi-dimensional polarization information map; A preliminary feature generation module for extracting features from the multi-dimensional polarization information map based on the backbone network of YOLOv8 and outputting preliminary features containing semantic and texture levels; A detection module for sequentially inputting the preliminary features into the global spatial attention module, the multi-dimensional polarization attention module, and the local spatial attention module to generate target localization features, and inputting the target localization features into the detection head to obtain the UAV detection result.
9. A computer-readable storage medium, characterized in that, It includes instructions that, when running on a computer, cause the computer to execute the UAV detection method based on polarization information enhancement and polarization attention features according to any one of claims 1 to 7.
10. An electronic device, characterized in that, The electronic device includes: At least one processor, a memory, and an input / output unit; Wherein, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the UAV detection method based on polarization information enhancement and polarization attention features according to any one of claims 1 to 7.