Infrared flapping wing type unmanned aerial vehicle identification method based on multi-dimensional attention fusion

The multi-dimensional attention fusion method enhances infrared flapping-wing drone recognition by leveraging a pre-trained model with bneck, SE, and BAM modules, addressing the challenge of identifying micro-sized drones in complex environments with improved accuracy and efficiency for resource-constrained devices.

CN120318563APending Publication Date: 2025-07-15XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510359104.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify flapping-wing drones in complex environments, especially in low light, night or inclement weather conditions. Infrared image recognition methods lack sufficient texture details and are susceptible to background interference, resulting in limited recognition accuracy.

Method used

The infrared flapping-wing drone recognition method based on multi-dimensional attention fusion is adopted, and the feature extraction and classification is extracted and classified through infrared image preprocessing, bneck network unit and output unit, combined with the SE module and BAM module, to enhance the recognition ability of the flapping-wing drone.

Benefits of technology

Improves recognition accuracy and efficiency in complex environments, suitable for deployment in edge devices and embedded devices, maintaining high recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318563A_ABST
    Figure CN120318563A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared flapping-wing unmanned aerial vehicle identification method based on multi-dimensional attention fusion, and solves the problem that a flapping-wing unmanned aerial vehicle cannot be accurately identified in a complex scene in the prior art. The method comprises the following steps: acquiring a to-be-recognized infrared image; inputting the to-be-recognized infrared image into a pre-trained infrared flapping wing type unmanned aerial vehicle recognition model to obtain a recognition result; the infrared image preprocessing unit is used for extracting local edge features of an infrared image to be recognized to obtain an input feature map; the bcheck network unit is used for obtaining an updated refined feature map according to the input feature map; the output unit is used for classifying the updated refined feature map and outputting a classification result as an identification result; according to the method, the identification precision and efficiency of an infrared small target identification system on the flapping wing type unmanned aerial vehicle in a complex environment are improved, and particularly, high identification performance can still be kept under the night or severe weather condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to an infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion. Background Art

[0002] Flapping-wing drones are bionic aircraft that generate lift and propulsion by simulating the flapping of wings of birds or insects. They are highly concealed and flexible and have been widely used in military reconnaissance, covert surveillance, and search and rescue operations. Due to their bionic structure and miniaturized design, flapping-wing drones are very similar to birds in appearance and movement patterns during flight. Especially in complex environments, the small size and high dynamics of the target further increase the difficulty of identification. Existing technologies are susceptible to environmental interference when accurately identifying such targets with similar structures and small sizes, and it is difficult to effectively distinguish flapping-wing drones from birds, which increases the risk of false alarms and directly affects the reliability of reconnaissance and monitoring tasks.

[0003] The current small target recognition technology has been widely used in fields such as monitoring and reconnaissance, but most of the recognition methods based on visible light show a significant decrease in accuracy under extreme lighting conditions such as low light, night or strong light. Infrared imaging technology has obvious advantages in such environments. It can capture the thermal radiation of the surface of the object to generate images and provide clear target images under poor lighting conditions. Therefore, it has important applications in small target recognition tasks. However, since infrared images often lack sufficient texture details, existing recognition methods are difficult to extract effective features under complex backgrounds, thus affecting the accuracy and stability of recognition. In addition, the thermal radiation characteristics of flapping-wing drones and birds under infrared imaging are similar, and small targets are easily affected by background interference, which seriously limits the recognition accuracy of existing technologies.

[0004] In recent years, a variety of models combining deep learning and attention mechanisms have been proposed to optimize the small target feature extraction and recognition performance of infrared images. However, these methods are usually accompanied by high computational overhead and are difficult to meet the needs of resource-constrained scenarios such as embedded or edge computing. Therefore, there is an urgent need for an infrared flapping-wing UAV recognition method that can combine a lightweight deep learning model and a multi-dimensional attention mechanism to meet the needs of accurate recognition of flapping-wing UAVs in complex scenarios. Summary of the invention

[0005] The present invention solves the problem that flapping-wing UAVs cannot be accurately identified in complex scenes in the prior art by providing an infrared flapping-wing UAV identification method based on multi-dimensional attention fusion, thereby improving the recognition accuracy and efficiency of flapping-wing UAVs by the infrared small target recognition system in complex environments, and can still maintain a high recognition performance, especially at night or under severe weather conditions.

[0006] The present invention provides an infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion, and the method includes:

[0007] Obtain the infrared image to be recognized;

[0008] Input the infrared image to be recognized into a pre-trained infrared flapping-wing UAV recognition model to obtain a recognition result; wherein, the infrared flapping-wing UAV recognition model includes: an infrared image preprocessing unit, a bneck network unit, and an output unit;

[0009] The infrared image preprocessing unit is used to extract local edge features of the infrared image to be recognized to obtain an input feature map;

[0010] The bneck network unit is used to: perform feature extraction on the input feature map based on a first convolutional module to obtain first output feature maps corresponding to each channel; perform feature compression, weight excitation, and weight recalibration on the first output feature maps corresponding to each channel in sequence based on an SE module to obtain second output feature maps corresponding to each channel; perform spatial dimension processing and channel dimension processing on the second output feature maps respectively based on the spatial attention branch and the channel attention branch in a BAM module to obtain a refined feature map; perform dimension restoration on the refined feature map based on a second convolutional module to obtain an updated refined feature map;

[0011] The output unit is used to classify the updated refined feature map and output the classification result as the recognition result.

[0012] In a possible implementation manner, the first convolutional module includes: a 1×1 convolutional layer and a depth convolutional layer;

[0013] The bneck network unit performs feature extraction on the input feature map based on the first convolutional module to obtain first output feature maps corresponding to each channel, including:

[0014] Perform channel expansion on the input feature map through a 1×1 convolutional layer, and then map the expanded input feature map to a high-dimensional space to obtain a high-dimensional feature map;

[0015] Perform spatial feature extraction of each channel on the high-dimensional feature map through a depth convolutional layer to obtain first output feature maps corresponding to each channel.

[0016] In a possible implementation manner, the SE module includes: a global average pooling layer, a first fully connected layer, a ReLU activation function, a second fully connected layer, and a Sigmoid activation function;

[0017] The bneck network unit performs feature compression, weight excitation, and weight recalibration on the first output feature map corresponding to each channel based on the SE module to obtain the second output feature map corresponding to each channel, including:

[0018] Performing a feature compression operation on the first output feature map of each channel using a global average pooling layer to obtain a global feature vector;

[0019] Performing dimensionality reduction and non-linearity introduction operations on the global feature vector using a first fully connected layer and a ReLU activation function to obtain a dimensionality-reduced global feature vector;

[0020] Performing dimensionality recovery and feature mapping operations on the dimensionality-reduced global feature vector using a second fully connected layer and a Sigmoid activation function to obtain an excitation channel weight vector;

[0021] Performing weight recalibration on the first output feature map of each corresponding channel according to the excitation channel weight vector to obtain the second output feature map corresponding to each channel.

[0022] In a possible implementation manner, the BAM module includes: a spatial attention branch and a channel attention branch;

[0023] The bneck network unit performs spatial dimension processing and channel dimension processing on the second output feature map based on the spatial attention branch and the channel attention branch in the BAM module to obtain a refined feature map, including:

[0024] Based on the spatial attention branch in the BAM module, performing processing on the second output feature map corresponding to each channel in the spatial dimension to obtain a spatial attention map;

[0025] Based on the channel attention branch in the BAM module, performing processing on the second output feature map corresponding to each channel in the channel dimension to obtain a channel attention map;

[0026] Performing element-wise summation of the spatial attention map and the channel attention map to obtain a bottleneck attention map;

[0027] Performing element-wise multiplication of the bottleneck attention map and the second output feature map corresponding to each channel respectively to obtain an initial refined feature map;

[0028] Performing element-wise addition of the initial refined feature map and the second output feature map corresponding to each channel respectively to obtain a refined feature map.

[0029] In a possible implementation manner, the spatial attention branch includes: a first convolutional layer, two dilated convolutional layers, a second convolutional layer, and a first batch normalization layer;

[0030] The bneck network unit processes the second output feature map corresponding to each channel in the spatial dimension based on the spatial attention branch in the BAM module to obtain a spatial attention map, including:

[0031] Using the first convolutional layer to compress the second output feature map corresponding to each channel to obtain a compressed feature map; wherein, the number of channels of the compressed feature map is calculated according to the number of second output feature maps and the contraction ratio;

[0032] Using the two dilated convolutional layers to perform feature extraction on the compressed feature map to obtain a dilated convolutional feature map;

[0033] Using the second convolutional layer and the first batch normalization layer to perform channel number recovery and normalization operations on the dilated convolutional feature map to obtain a spatial attention map.

[0034] In a possible implementation manner, the channel attention branch includes: a global average pooling layer, a third fully connected layer, a fourth fully connected layer, and a second batch normalization layer;

[0035] The bneck network unit processes the second output feature map corresponding to each channel in the channel dimension based on the channel attention branch in the BAM module to obtain a channel attention map, including:

[0036] Using the global average pooling layer to perform global information extraction on the second output feature map corresponding to each channel to obtain a multi-channel global feature vector;

[0037] Using the third fully connected layer to perform dimensionality reduction processing on the multi-channel global feature vector to obtain a multi-channel global feature vector after dimensionality reduction; wherein, the dimension of the multi-channel global feature vector after dimensionality reduction is calculated according to the number of channels of the second output feature map and the contraction ratio;

[0038] Using the fourth fully connected layer and the second batch normalization layer to perform dimension recovery and normalization on the multi-channel global feature vector after dimensionality reduction to obtain a channel attention map.

[0039] In a possible implementation manner, the step of adding the spatial attention map and the channel attention map element-wise to obtain a bottleneck attention map includes:

[0040] Adjusting the three-dimensional dimensions of the spatial attention map and the channel attention map respectively to obtain a preprocessed spatial attention map and a channel attention map;

[0041] Adding the preprocessed spatial attention map and the channel attention map element-wise to obtain a fused feature map;

[0042] The fused feature map is mapped using the Sigmoid function to obtain a bottleneck attention map.

[0043] In a possible implementation, the output unit includes: an average pooling layer, a fifth fully connected layer, an hswish activation function, and a sixth fully connected layer;

[0044] The output unit is used to classify the updated refined feature map and output the classification result as the recognition result, including:

[0045] Using the average pooling layer, the refined feature map is reduced in dimension to a first feature vector;

[0046] Using the fifth fully connected layer to extract features from the first feature vector to obtain a second feature vector, and then using the hswish activation function to perform a non-linear mapping on the second feature vector to obtain a high-dimensional feature vector;

[0047] Using the sixth fully connected layer to perform a linear transformation on the high-dimensional feature vector to obtain the classification result corresponding to the refined feature map, and output the classification result as the recognition result.

[0048] In a possible implementation, when training the infrared flapping-wing UAV recognition model, the loss function used is expressed as:

[0049]

[0050] where a represents a scaling factor; m represents a preset value; N represents the number of training sample sets in the current batch; i represents the i-th sample in the training sample set of the current batch; k represents the number of target categories in the training sample set of the current batch; yi represents the true category of the i-th sample; W yi represents the weight vector of the true category of the i-th sample, and f(x i ) represents the feature vector of the i-th sample; represents the angle between the feature vector of the i-th sample and the weight vector of the true category of this sample; j represents the j-th target category; W j represents the weight vector corresponding to the j-th target category; represents the angle between the weight vector corresponding to the j-th target category and the feature vector of the i-th sample, and the j-th target category is not the true category of the i-th sample.

[0051] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0052] Through the first convolution module in the bneck network unit, the present invention maps the input feature map to a high-dimensional space, then performs depth convolution in the high-dimensional space to capture spatial features, and then reduces the number of channels back to the original dimension while retaining the input features; the SE module sequentially performs feature compression, weight excitation, and weight recalibration on the first output feature map corresponding to each channel, and then adjusts the weight of each channel to enhance the network's perception ability of important features; while allowing the network to adaptively adjust the weights of features in different channels, the bneck network unit introduces the BAM module, which processes the second output feature map in the spatial dimension and the channel dimension according to the spatial attention branch and the channel attention branch to obtain a refined feature map; enhances the expression ability of the infrared flapping-wing UAV recognition model for flapping target features and improves the accuracy of target recognition in complex environments; the SE module and the BAM module in the bneck network unit increase the number of model parameters less. Therefore, the model has the characteristics of light weight, fast inference speed, and is suitable for deployment in edge devices and embedded devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of the steps of the multi-dimensional attention fusion infrared flapping-wing UAV recognition method provided by the embodiment of the present invention;

[0054] Figure 2 It is a schematic diagram of the bneck network unit provided by the embodiment of the present invention;

[0055] Figure 3 It is a schematic diagram of the processing flow of the first convolution module and the Squeeze and Excitation Module (SE) module provided by the embodiment of the present invention;

[0056] Figure 4 It is a schematic diagram of the processing flow of the Bottleneck Attention Module (BAM) module provided by the embodiment of the present invention;

[0057] Figure 5 It is a schematic diagram of the infrared flapping-wing UAV recognition model provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] The present invention provides an infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion. Refer to Figure 1 , the method includes the following steps S101 to S102.

[0060] S101, obtaining an infrared image to be recognized.

[0061] Exemplarily, the infrared image to be recognized is an image with a size of 3×224×224, representing a 3-channel infrared image with a resolution of 224×224.

[0062] S102, inputting the infrared image to be recognized into a pre-trained infrared flapping-wing UAV recognition model to obtain a recognition result; wherein, the infrared flapping-wing UAV recognition model includes: an infrared image preprocessing unit, a bneck network unit, and an output unit.

[0063] The infrared image preprocessing unit is used to extract local edge features of the infrared image to be recognized to obtain an input feature map;

[0064] Exemplarily, the infrared image to be recognized is an image with a size of 3×224×224, representing a 3-channel infrared image with a resolution of 224×224. First, the infrared image to be recognized passes through the first layer: a convolutional layer to extract low-level features and downsample at the same time. In the convolutional layer, first pass through a 3×3 convolutional layer with a stride of 2 to adjust the output channels to 16, reduce the input size to 112×112, reduce the computational amount, and extract low-level features such as local edges and textures. Then apply batch normalization (BN) to improve the training stability, accelerate convergence, and use the hswish activation function to enhance the non-linear expression ability;

[0065] Refer to Figure 2 , the first convolutional module includes: a 1×1 convolutional layer and a depth convolutional layer; the bneck network unit is used to: extract features from the input feature map based on the first convolutional module to obtain a first output feature map corresponding to each channel; perform feature compression, weight excitation, and weight recalibration on the first output feature map corresponding to each channel based on the SE module to obtain a second output feature map corresponding to each channel; perform spatial dimension processing and channel dimension processing on the second output feature map based on the spatial attention branch and the channel attention branch in the BAM module respectively to obtain a refined feature map; perform dimension restoration on the refined feature map based on the second convolutional module to obtain an updated refined feature map;

[0066] Specifically, in the bneck network unit, refer to Figure 3 , extracting features from the input feature map based on the first convolutional module (F tr ), to obtain a first output feature map corresponding to each channel, including:

[0067] (1) Expand the channels of the input feature map through a 1×1 convolutional layer, and then map the expanded input feature map into a high-dimensional space to obtain a high-dimensional feature map;

[0068] (2) Extract the spatial features of each channel of the high-dimensional feature map through a depthwise convolutional layer to obtain the first output feature map corresponding to each channel.

[0069] Exemplarily, see Figure 3 , first expand the channels through a 1×1 convolution to map the low-dimensional input features to a high-dimensional space, then perform depthwise convolution in the high-dimensional space to capture spatial features, and finally reduce the number of channels back to the original dimension through another 1×1 convolution, while introducing skip connections to retain the input features.

[0070] Specifically, expand the channels of the input feature map through a 1×1 convolutional layer, and then map the expanded input feature map into a high-dimensional space to obtain a high-dimensional feature map; depthwise separable convolution first applies a convolutional kernel to each channel of the high-dimensional feature map to generate a feature map with the same depth as the high-dimensional feature map; subsequently, through a pointwise convolution operation with a size of 1×1, map the generated feature map to a new feature space to obtain the first output feature map corresponding to each channel. Assume the high-dimensional feature map is X', and the calculation formulas for depthwise convolution and pointwise convolution are respectively:

[0071] F depthwise (X') = X' * K depthwise (1);

[0072] F pointwise (F depthwise (X')) = F depthwise (X') * K pointwise (2);

[0073] where, * represents the convolution operation, K depthwise and K pointwise are the depthwise convolutional kernel and the pointwise convolutional kernel respectively.

[0074] That is, for an infrared image X to be recognized with a size of C×W×H, after passing through F tr convolution operation, the first output feature map U is obtained. The process here is expressed by the formula:

[0075] F tr : X → U (3);

[0076] where, X ∈ R W′×H′×C′ , U ∈ R W×H×C .

[0077] Specifically, by applying a convolutional kernel to each channel x sPerform a convolution operation using the c-th convolution kernel V c s Perform a convolution operation channel by channel and sum the results to finally obtain the c-th channel u of the first output feature map c , and the process here is expressed by the formula as follows:

[0078]

[0079] where, * represents the convolution operation, here is a 2D spatial convolution kernel used to perform a local convolution operation on the input feature map; X represents the input feature map; x s represents the s-th channel of the input feature map; V c represents the c-th convolution kernel used to process the c-th channel, and c' represents the number of channels of the input feature map of the SE module (which can be different from the number of channels of the output feature map).

[0080] Since each convolution unit can only focus on the spatial information within its local receptive field, the information outside this receptive field cannot be utilized. After the F tr operation, for each channel of the corresponding first output feature map, an SE module is introduced.

[0081] Specifically, refer to Figure 3 and Figure 4 , the SE module includes: a global average pooling layer, a first fully connected layer, a ReLU activation function, a second fully connected layer, and a Sigmoid activation function;

[0082] In the bneck network unit, based on the SE module, for each channel of the corresponding first output feature map, perform feature compression (Squeeze F sq ), weight excitation (Excitation F ex ), and weight recalibration (Scale F scale ) in sequence to obtain the corresponding second output feature map for each channel, including:

[0083] (1) Use the global average pooling layer to perform a feature compression operation on the first output feature map of each channel to obtain a global feature vector;

[0084] (2) Use the first fully connected layer and the ReLU activation function to perform a dimensionality reduction and non-linear introduction operation on the global feature vector to obtain a dimensionality-reduced global feature vector;

[0085] (3) Use the second fully connected layer and the Sigmoid activation function to perform a dimensionality recovery and feature mapping operation on the dimensionality-reduced global feature vector to obtain an excitation channel weight vector;

[0086] (4) Weight calibration is performed on the first output feature map of each channel according to the excitation channel weight vector to obtain the second output feature map corresponding to each channel.

[0087] Exemplarily, perform an F sq operation on the first output feature map. The specific formula for the F sq operation is:

[0088]

[0089] where z c represents the global feature vector obtained through the F sq operation; F sq represents the global average pooling process; u c (i, j) represents the pixel value of the first output feature map of the c-th channel at the position (i, j), H represents the height of the first output feature map, W represents the width of the first output feature map, and H×W represents the spatial size of the first output feature map, which is used to average the pixel values of each channel to obtain the global feature vector.

[0090] After F sq global average pooling, the H×W×C feature map containing global information is directly compressed into a 1×1×C global feature vector, thereby calculating the global feature vector of each channel. The network can obtain the global feature vector of each channel, which overcomes the limitation that traditional convolution operations can only capture local dependencies. This feature vector will be used to readjust the channel weights in the subsequent excitation operation.

[0091] Then, through the F ex weight excitation, the purpose of this layer is to calculate the weights of each channel through a non-linear mechanism and redistribute these weights to the corresponding channels. Here, the global feature vector Z generated by the F sq layer is reduced from dimension C to C / r through the first fully connected layer W1 (r represents the compression ratio), and then W1Z passes through the ReLU activation function δ to introduce non-linearity. After that, δ(W1Z) passes through the second fully connected layer W2 to restore the dimension from C / r to C to ensure that the output vector matches the original number of channels. Finally, through the Sigmoid activation function Ω, the weight s of each channel is mapped between 0 and 1 to generate the excitation channel weight vector.

[0092] Its formula is:

[0093] s = F ex (Z, W) = Ω(W2δ(W1Z)) (6);

[0094] where s represents the weight vector of the channel, and s is obtained through F exIt is calculated by operation; Z represents the global feature vector; W1 represents the weight matrix of the first fully connected layer; W2 represents the weight matrix of the second fully connected layer; δ refers to the ReLU activation function, and Ω refers to the Sigmoid activation function.

[0095] Further, apply the generated excitation channel weight vector to each channel u of the first output feature map c in, that is, perform F scale weight recalibration operation, and its formula is:

[0096]

[0097] Among them, F scale (u c , s c ) represents the channel-level weighting operation, u c represents the first output feature map corresponding to the c-th channel; s c represents the c-th element in the channel weight vector s generated by the excitation operation through F ex , that is, the weight value corresponding to channel c. The two are multiplied to obtain the weighted second output feature map. That is, multiply channel by channel to complete the recalibration of the original features.

[0098] Specifically, the BAM module includes: a spatial attention branch and a channel attention branch; the spatial attention branch includes: a first convolutional layer, two dilated convolutional layers, a second convolutional layer, and a first batch normalization layer; the channel attention branch includes: a global average pooling layer, a third fully connected layer, a fourth fully connected layer, and a second batch normalization layer;

[0099] In the bneck network unit, based on the spatial attention branch and the channel attention branch in the BAM module, perform spatial dimension processing and channel dimension processing on the second output feature map respectively to obtain a refined feature map, including:

[0100] (1) Based on the spatial attention branch in the BAM module, process the second output feature map corresponding to each channel in the spatial dimension to obtain a spatial attention map;

[0101] Here, based on the spatial attention branch in the BAM module, process the second output feature map corresponding to each channel in the spatial dimension to obtain a spatial attention map, including:

[0102] (1.1) Use the first convolutional layer to compress the second output feature map corresponding to each channel to obtain a compressed feature map; among them, the number of channels of the compressed feature map is calculated according to the number of the second output feature map and the shrinkage ratio;

[0103] (1.2) Use two dilated convolutional layers to extract features from the compressed feature map to obtain a dilated convolutional feature map;

[0104] (1.3) Use the second convolutional layer and the first batch normalization layer to perform channel number recovery and normalization operations on the dilated convolutional feature map to obtain the spatial attention map.

[0105] (2) Based on the channel attention branch in the BAM module, process the second output feature map corresponding to each channel in the channel dimension to obtain the channel attention map;

[0106] Here, based on the channel attention branch in the BAM module, processing the second output feature map corresponding to each channel in the channel dimension to obtain the channel attention map includes:

[0107] (2.1) Use the global average pooling layer to extract global information from the second output feature map corresponding to each channel to obtain a multi-channel global feature vector;

[0108] (2.2) Use the third fully connected layer to perform dimensionality reduction on the multi-channel global feature vector to obtain a multi-channel global feature vector after dimensionality reduction; among them, the dimension of the multi-channel global feature vector after dimensionality reduction is calculated according to the number of channels of the second output feature map and the contraction ratio;

[0109] (2.3) Use the fourth fully connected layer and the second batch normalization layer to perform dimension recovery and normalization on the multi-channel global feature vector after dimensionality reduction to obtain the channel attention map.

[0110] (3) Perform element-wise summation of the spatial attention map and the channel attention map to obtain the bottleneck attention map;

[0111] Here, performing element-wise summation of the spatial attention map and the channel attention map to obtain the bottleneck attention map includes:

[0112] (3.1) Adjust the three-dimensional dimensions of the spatial attention map and the channel attention map respectively to obtain the preprocessed spatial attention map and channel attention map;

[0113] (3.2) Perform element-wise summation of the preprocessed spatial attention map and the channel attention map to obtain the fused feature map;

[0114] (3.3) Use the Sigmoid function to map the fused feature map to obtain the bottleneck attention map.

[0115] (4) Multiply the bottleneck attention map element-wise with the second output feature map corresponding to each channel to obtain the initial refined feature map;

[0116] (5) Add the initial refined feature map element-wise to the second output feature map corresponding to each channel to obtain the refined feature map.

[0117] Exemplarily, the second output feature maps corresponding to each channel output by the recalibration operation layer passing through the SE module will be input into the Bottleneck Attention Module (BAM). The BAM module consists of a spatial attention branch and a channel attention branch. See Figure 4 .

[0118] This layer uses a dilation value d and a contraction ratio r to regulate the spatial attention branch. The dilation value determines the size of the receptive field, which helps to aggregate the context information of the spatial attention branch; the contraction ratio controls the capabilities of the spatial attention branch and the channel attention branch. In this embodiment, the dilation value d = 4 and the contraction ratio r = 16 are set.

[0119] The channel attention branch first extracts a multi-channel global feature vector through global average pooling, and then reduces the number of channels from C to C / r through the first layer of weights W0 in the multi-layer perceptron in the third fully-connected layer to make the features more compact. Then, it restores the number of channels from C / r to the original dimension through the second layer of weights W1 in the multi-layer perceptron in the fourth fully-connected layer to correspond to the input feature map, and then adjusts the output of the spatial attention branch to a normalized feature map Mc(F) through batch normalization. The formula is as follows:

[0120] Mc(F) = BN(MLP(AvgPool(F))) = BN(W1(W0AvgPool(F) + b0) + b1) (8);

[0121] where, F represents the input feature map, that is, the second output feature maps corresponding to each channel; W0 represents the first layer of weights in the multi-layer perceptron in the third fully-connected layer; W1 represents the second layer of weights in the multi-layer perceptron in the fourth fully-connected layer, b0 represents the bias in the multi-layer perceptron in the third fully-connected layer; b1 represents the bias in the multi-layer perceptron in the fourth fully-connected layer, BN represents batch normalization; AvgPool(·) represents the average pooling operation; BN(·) represents the batch normalization operation.

[0122] The spatial attention branch first compresses the channels of the second output feature maps corresponding to each channel using a 1×1 convolution, further extracts features using two 3×3 dilated convolutions, and then restores the number of channels to the original dimension through a 1×1 convolution operation. Finally, it normalizes the output spatial attention map Ms(F) through batch normalization. The calculation formula of the spatial attention branch is as follows:

[0123]

[0124] where, F represents the input feature map, that is, the second output feature maps corresponding to each channel; f represents the convolution operation; f 1×1 represents using a 1×1 convolution; f 3×3It means using a 3×3 dilated convolution. The subscript of each f represents the order of different convolution operations, and BN represents the batch normalization operation.

[0125] Adjust the three-dimensional dimensions of the spatial attention map and the channel attention map respectively to obtain the preprocessed spatial attention map and channel attention map, with the size of

[0126] Perform element-wise summation on the preprocessed spatial attention map and channel attention map to obtain the fused feature map. After summation, use the Sigmoid function to obtain the final bottleneck attention map M(F) in the range of 0 to 1.

[0127] Its formula is as follows:

[0128] M(F) = σ(Mc(F) + Ms(F)) (10);

[0129] Among them, F represents the input feature map, that is, the second output feature map corresponding to each channel; σ is the Sigmoid activation function.

[0130] For a given second output feature map F ∈ R corresponding to each channel processed by the SE module C×H×W , the BAM module infers a bottleneck attention map M(F) ∈ R C×H×W . The calculation formula for the refined feature map F' is:

[0131]

[0132] Among them, represents element-wise multiplication.

[0133] The output unit is used to classify the updated refined feature map and output the classification result as the recognition result.

[0134] Specifically, the output unit includes: an average pooling layer, a fifth fully connected layer, an hswish activation function, and a sixth fully connected layer; the output unit is used to classify the updated refined feature map and output the classification result as the recognition result, including:

[0135] (1) Use the average pooling layer to reduce the dimension of the refined feature map to the first feature vector;

[0136] (2) Use the fifth fully connected layer to extract features from the first feature vector to obtain the second feature vector, and then use the hswish activation function to perform a non-linear mapping on the second feature vector to obtain a high-dimensional feature vector;

[0137] (3) Use the sixth fully connected layer to perform a linear transformation on the high-dimensional feature vector to obtain the classification result corresponding to the refined feature map, and output the classification result as the recognition result.

[0138] Exemplarily, through a 7×7 global average pooling layer, the refined feature map is reduced to a single vector (the first feature vector) to reduce the computational amount and extract global information;

[0139] Through a fifth fully connected layer with 1024 dimensions, a second feature vector is obtained;

[0140] Using hswish as the activation function, the second feature vector is mapped to obtain a high-dimensional feature vector to enhance the feature expression ability;

[0141] Then, through a final sixth fully connected layer with 2 dimensions, the high-dimensional feature vector is mapped into a 2-dimensional output space to obtain the corresponding classification result, and the classification result is output as the recognition result.

[0142] In the present invention, the pre-trained infrared flapping-wing UAV recognition model is obtained through the following training method, and the specific process is as follows:

[0143] In this embodiment, the input data set is selected to be divided according to the ratio that the training set accounts for 70%, and the validation set and the test set each account for 15%. Specifically, randomly select 70% of the images of flapping-wing UAVs and birds in the infrared image data set as the training set to input into the infrared flapping-wing UAV recognition model for training, select half of the remaining 30% of the images of flapping-wing UAVs and birds as the validation set to adjust the parameters of the model, and finally the test recognition result is completed in the remaining 15% of the images.

[0144] In this step, the ArcFace loss function is adopted in the training of the infrared flapping-wing UAV recognition model. ArcFace first normalizes the extracted feature vector and the classification weight vector through L2 normalization, calculates the included angle between them, uses the included angle to measure the distance between the sample and the class center, and requires the included angle θ yi between the feature and the correct class must be increased by a margin on the original basis, which makes the same-class features closer and the angular distance of different-class features larger, thereby enhancing the discrimination between different classes. The traditional Softmax loss is based on the inner product between the feature and the weight, and does not constrain the inter-class difference in the angular space. By introducing the ArcFace Loss, the model can enhance the inter-class discrimination in terms of angle, and can improve the accuracy of the model without adding an additional deep network structure, effectively distinguishing the subtle differences between flapping-wing UAVs and birds.

[0145] When training the infrared flapping-wing UAV recognition model, the loss function used is expressed as:

[0146]

[0147] Among them, a represents a scaling factor; m represents a preset value; N represents the number of training sample sets in the current batch; i represents the i-th sample in the training sample set of the current batch; k represents the number of target categories in the training sample set of the current batch; yi represents the true category of the i-th sample; W yi represents the weight vector of the true category of the i-th sample, f(x i ) represents the feature vector of the i-th sample; represents the angle between the feature vector of the i-th sample and the weight vector of the true category of this sample; j represents the j-th target category; W j represents the weight vector corresponding to the j-th target category; represents the angle between the weight vector corresponding to the j-th target category and the feature vector of the i-th sample, and the j-th target category is not the true category of the i-th sample.

[0148]

[0149] See Figure 5 , in a specific embodiment provided by the present invention, the input of the infrared flapping-wing UAV recognition model is an image with a size of 3×224×224, representing a 3-channel infrared image with a resolution of 224×224. First, it passes through the first layer: a convolutional layer to extract low-level features and downsample at the same time. In the convolutional layer, it first passes through a 3×3 convolutional layer with a stride of 2 to adjust the output channels to 16, reducing the input size to 112×112, reducing the computational amount, and extracting low-level features such as local edges and textures. Then, batch normalization (BN) is applied to improve the training stability, accelerate convergence, and use the hswish activation function to enhance the non-linear expression ability; then it enters the backbone network, and feature extraction is performed through 11 improved bottleneck blocks, thereby greatly improving the computational efficiency and classification ability of the model; then it passes through a 7×7 global average pooling layer to reduce the feature map to a single vector to reduce the computational amount and extract global information; then it passes through a fully connected layer with 1024 dimensions, and uses hswish as the activation function to map the pooled low-dimensional features to obtain a feature vector and enhance the feature expression ability; then it passes through the final 2D fully connected layer to map the feature vector to a 2D output space to obtain the corresponding classification result, and the classification result is output as the recognition result.

[0150] The present invention effectively overcomes the disadvantages of traditional visible light recognition systems performing poorly in environments with insufficient light and complex backgrounds by obtaining the input infrared image and constructing a data set based on the image, ensuring high-precision recognition of flapping-wing UAVs and birds in complex environments.

[0151] Based on the convolutional neural network, the present invention uses depthwise separable convolutions, splitting the convolution operation into depthwise convolution and pointwise convolution, greatly reducing the computational amount and the number of parameters. The inverted residual structure is introduced. First, the low-dimensional features are expanded to a high-dimensional space for convolution and then reduced back to the low-dimensional space. At the same time, skip connections are introduced to retain the input features. The SEB multi-dimensional attention fusion module is embedded in the inverted residual block as part of the feature extraction space. While allowing the network to adaptively adjust the weights of features in different channels, channel and spatial synchronous inference attention maps are introduced to further enhance the network's ability to express the infrared features of flapping-wing drones and birds, improving the recognition accuracy in complex environments.

[0152] Based on the convolutional neural network model, the present invention uses depthwise separable convolutions and average pooling to reduce the additional computational overhead of the convolution operation and enhance the neural network's ability to extract key feature channels of flapping-wing drones and birds. At the same time, the increase in the number of model parameters in the SEB attention module proposed above is relatively small. Therefore, this model has the characteristics of being lightweight, with a fast inference speed, and is suitable for deployment in edge devices and embedded devices. To further reduce the computational amount, the model proposed in the present invention is trained using the ArcFace loss function. This angle-based optimization method avoids the per-channel multiplication calculation of features and weights compared with the traditional Softmax loss function, reducing the computational complexity. At the same time, the angle margin mechanism introduced by this loss function effectively enhances the discrimination between features of different categories (i.e., flapping-wing drones and birds), and the model accuracy can be improved without adding an additional deep network structure. It ensures that the model can operate more efficiently while maintaining high performance and is suitable for resource-constrained device environments.

[0153] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. The key points described in each embodiment are the differences from other embodiments. All or part of the present invention can be used in many general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on.

[0154] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the present invention; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present invention.

Claims

1. An infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion, characterized in that, Including: Obtain the infrared image to be recognized; Input the infrared image to be recognized into a pre-trained infrared flapping UAV recognition model to obtain a recognition result; wherein, the infrared flapping UAV recognition model includes: an infrared image preprocessing unit, a bneck network unit, and an output unit; The infrared image preprocessing unit is used to extract local edge features of the infrared image to be recognized to obtain an input feature map; The bneck network unit is used to: perform feature extraction on the input feature map based on the first convolution module to obtain first output feature maps corresponding to each channel; perform feature compression, weight excitation, and weight recalibration on the first output feature maps corresponding to each channel in sequence based on the SE module to obtain second output feature maps corresponding to each channel; perform spatial dimension processing and channel dimension processing on the second output feature maps respectively based on the spatial attention branch and the channel attention branch in the BAM module to obtain a refined feature map; perform dimension restoration on the refined feature map based on the second convolution module to obtain an updated refined feature map; The output unit is used to classify the updated refined feature map and output the classification result as the recognition result.

2. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 1, wherein The first convolution module includes: a 1×1 convolution layer and a depth convolution layer; The bneck network unit performs feature extraction on the input feature map based on the first convolution module to obtain first output feature maps corresponding to each channel, including: Expand the channels of the input feature map through a 1×1 convolution layer, and then map the expanded input feature map into a high-dimensional space to obtain a high-dimensional feature map; Perform spatial feature extraction on each channel of the high-dimensional feature map through a depth convolution layer to obtain first output feature maps corresponding to each channel.

3. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 1, wherein The SE module includes: a global average pooling layer, a first fully connected layer, a ReLU activation function, a second fully connected layer, and a Sigmoid activation function; The bneck network unit performs feature compression, weight excitation, and weight recalibration on the first output feature maps corresponding to each channel in sequence based on the SE module to obtain second output feature maps corresponding to each channel, including: Perform a feature compression operation on the first output feature maps corresponding to each channel using the global average pooling layer to obtain a global feature vector; Perform dimension reduction and non-linear introduction operations on the global feature vector using the first fully connected layer and the ReLU activation function to obtain a dimension-reduced global feature vector; Perform dimension restoration and feature mapping operations on the dimension-reduced global feature vector using the second fully connected layer and the Sigmoid activation function to obtain an excitation channel weight vector; Perform weight recalibration on the first output feature maps of each channel according to the excitation channel weight vector to obtain second output feature maps corresponding to each channel.

4. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 1, wherein, The BAM module includes: a spatial attention branch and a channel attention branch; The bneck network unit performs spatial dimension processing and channel dimension processing on the second output feature maps respectively based on the spatial attention branch and the channel attention branch in the BAM module to obtain a refined feature map, including: Based on the spatial attention branch in the BAM module, the second output feature maps corresponding to each channel are processed in the spatial dimension to obtain a spatial attention map; Based on the channel attention branch in the BAM module, the second output feature maps corresponding to each channel are processed in the channel dimension to obtain a channel attention map; The spatial attention map and the channel attention map are summed element-wise to obtain a bottleneck attention map; The bottleneck attention map is multiplied element-wise with the second output feature maps corresponding to each channel to obtain an initial refined feature map; The initial refined feature map is added element-wise to the second output feature maps corresponding to each channel to obtain a refined feature map.

5. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 4, wherein The spatial attention branch includes: a first convolutional layer, two dilated convolutional layers, a second convolutional layer, and a first batch normalization layer; The bneck network unit, based on the spatial attention branch in the BAM module, processes the second output feature maps corresponding to each channel in the spatial dimension to obtain a spatial attention map, including: Using the first convolutional layer to compress the second output feature maps corresponding to each channel to obtain a compressed feature map; wherein, the number of channels of the compressed feature map is calculated according to the number of second output feature maps and the shrinkage ratio; Using the two dilated convolutional layers to perform feature extraction on the compressed feature map to obtain a dilated convolutional feature map; Using the second convolutional layer and the first batch normalization layer to perform channel number recovery and normalization operations on the dilated convolutional feature map to obtain a spatial attention map.

6. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 4, characterized in that, The channel attention branch includes: a global average pooling layer, a third fully-connected layer, a fourth fully-connected layer, and a second batch normalization layer; The bneck network unit, based on the channel attention branch in the BAM module, processes the second output feature maps corresponding to each channel in the channel dimension to obtain a channel attention map, including: Using the global average pooling layer to extract global information from the second output feature maps corresponding to each channel to obtain a multi-channel global feature vector; Using the third fully-connected layer to perform dimensionality reduction on the multi-channel global feature vector to obtain a dimensionality-reduced multi-channel global feature vector; wherein, the dimension of the dimensionality-reduced multi-channel global feature vector is calculated according to the number of channels of the second output feature maps and the shrinkage ratio; Using the fourth fully-connected layer and the second batch normalization layer to perform dimensionality recovery and normalization on the dimensionality-reduced multi-channel global feature vector to obtain a channel attention map.

7. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 4, wherein, The step of summing the spatial attention map and the channel attention map element-wise to obtain a bottleneck attention map includes: Adjusting the three-dimensional dimensions of the spatial attention map and the channel attention map respectively to obtain a preprocessed spatial attention map and a preprocessed channel attention map; Summing the preprocessed spatial attention map and the preprocessed channel attention map element-wise to obtain a fused feature map; Using the Sigmoid function to map the fused feature map to obtain a bottleneck attention map.

8. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 1, wherein The output unit includes: an average pooling layer, a fifth fully-connected layer, an hswish activation function, and a sixth fully-connected layer; The output unit is used to classify the updated refined feature map and output the classification result as the recognition result, including: Using the average pooling layer, reducing the dimension of the refined feature map to a first feature vector; Using the fifth fully-connected layer to extract features from the first feature vector to obtain a second feature vector, and then using the hswish activation function to perform a non-linear mapping on the second feature vector to obtain a high-dimensional feature vector; Using the sixth fully-connected layer to perform a linear transformation on the high-dimensional feature vector to obtain the classification result corresponding to the refined feature map, and outputting the classification result as the recognition result.

9. The infrared flapping-wing UAV recognition method based on multi-dimensional attention fusion according to claim 1, wherein When training the infrared flapping-wing UAV recognition model, the loss function used is expressed as: Among them, a represents the scaling factor; m represents the preset value; N represents the number of training sample sets in the current batch; i represents the i-th sample in the training sample set of the current batch; k represents the number of target categories in the training sample set of the current batch; yi represents the true category of the i-th sample; W yi represents the weight vector of the true category of the i-th sample, f(x i ) represents the feature vector of the i-th sample; represents the angle between the feature vector of the i-th sample and the weight vector of the true category of this sample; j represents the j-th target category; W j represents the weight vector corresponding to the j-th target category; represents the angle between the weight vector corresponding to the j-th target category and the feature vector of the i-th sample, and the j-th target category is not the true category of the i-th sample.