Unmanned aerial vehicle infrared small target detection method based on attention mechanism

By introducing DSEAM and S_RepViTBlock modules in the YOLO11 model, the accuracy and timeliness of drone infrared small target detection in complex backgrounds are solved, and efficient and accurate infrared small target detection effect is achieved.

CN119942379AActive Publication Date: 2025-05-06LIAONING TECHNICAL UNIVERSITY

Patent Information

Application Number
CN202510041578.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-06
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

The existing drone infrared small target detection technology is difficult to accurately detect infrared small targets under complex backgrounds, and the limitation of computing resources leads to insufficient detection timeliness, which makes it difficult to meet real-time requirements.

Method used

Using an improved YOLO11 model based on attention mechanism, the DSEAM and S_RepViTBlock modules are added to the backbone network to enhance the infrared small target feature information and reduce unnecessary forward propagation information. The deep separation convolution and NWD loss function are used to accelerate model training and improve detection quality.

Benefits of technology

In complex environments, the accuracy and efficiency of infrared small object detection are significantly improved, the calculation amount and parameter amount of the model are reduced, the detail extraction ability of background and target features is enhanced, and the requirements of real-time detection are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942379A_ABST
    Figure CN119942379A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of infrared small target detection, in particular to an unmanned aerial vehicle infrared small target detection method based on an attention mechanism, and the method comprises the following steps: detection model determination, model revision and infrared small target detection. According to the method, improvement is carried out on the basis of a YOLO11 model, that is, standardized data distribution is achieved through deep separable convolution and a DSEAM module, training of the model is accelerated, and unnecessary forward spreading information can be remarkably reduced while infrared small target feature information can be enhanced in the operation process; and meanwhile, the SRepViTBlock module can be used for capturing the infrared small target and enhancing the detail extraction capability of background and target characteristics by using a larger receptive field so as to fully reduce errors and improve the detection quality of the infrared small target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrared small target detection, and in particular to a method for detecting infrared small targets of unmanned aerial vehicles based on an attention mechanism. Background Art

[0002] With the rapid development of deep learning and visual inspection technology, the application of visual inspection systems on drones is becoming more and more widespread in the civil and military fields. Small drones have the advantages of low cost, strong portability, high flight flexibility, good concealment and high efficiency, making them used in many fields such as traffic monitoring, accident search and rescue, medical imaging and military intelligence collection. At present, most drones use visible light vision systems and have achieved good results, but they are limited by computing resources, which has a serious impact on the detection tasks under external conditions such as long-distance detection, bad weather, dark light at night, and difficulty in distinguishing from the surroundings. Infrared imaging technology is not affected by external light sources and has strong anti-interference ability, long-distance detection ability and all-weather stable imaging detection ability.

[0003] However, when using UAV infrared imaging, the detection targets usually appear as dense small infrared targets due to the flight altitude, and are limited by low resolution, lack of color features, and texture features, which increases the difficulty of detection. Therefore, improving the infrared small target detection performance in complex environments based on UAV platforms with limited computing resources has received widespread attention in the research field.

[0004] With the continuous efforts of scientific researchers at home and abroad, infrared small target detection algorithms based on deep learning have gradually made breakthroughs from low-level to high-level, from local to universal. In the infrared small target detection algorithm based on multi-layer convolution fusion, a deep learning model combining multi-layer convolution and multi-receptive field fusion is proposed. The infrared small target is characterized by multi-layer feature extraction and fusion, and the multi-receptive field module is used to adapt to its complex form. The residual learning and CBAM are combined to adjust the importance of the network. Finally, the predicted image of the matching size is output through the full convolution module;

[0005] Infrared weak infrared small target detection method based on cross-connection and fusion attention mechanism A single-stage infrared weak infrared small target detection algorithm based on cross-connection and fusion attention mechanism is proposed, which improves the infrared small weak target detection ability under various complex backgrounds;

[0006] The target detection network with dynamic selection of infrared and visible light image features proposes to add dynamic fusion layer and dynamic selection layer to the network structure, perform feature fusion on multi-source image feature maps multiple times, and use three attention mechanisms to enhance the multi-scale feature maps to screen useful features.

[0007] Although the above algorithm has improved the detection accuracy of small infrared targets in specific environments, in complex backgrounds, the inspection target may overlap or be obscured by other objects such as buildings and trees, and lacks the ability to distinguish in densely populated areas, making it difficult to accurately detect small infrared targets. In addition, drones will obtain a large amount of dynamic data in real time during flight, which will cause large delays when computing resources are limited, reducing the timeliness of detection and making it difficult to meet the real-time requirements of certain usage scenarios, so improvements are needed. Summary of the invention

[0008] The purpose of the present invention is to solve the shortcomings of the prior art and propose a method for detecting small infrared targets of drones based on the attention mechanism.

[0009] In order to achieve the above object, the present invention adopts the following technical solutions:

[0010] A method for detecting small infrared targets of unmanned aerial vehicles based on an attention mechanism includes the following steps:

[0011] S1. Detection model determination: The YOLO11 model is used, which consists of three parts: backbone network (Backbone), neck network (Neck), and head network (Head);

[0012] S2. Model revision: Add the efficient dual-channel feature enhancement module DSEAM module to the third and sixth layers of the YOLO11 backbone network, which can significantly reduce unnecessary forward propagation information while enhancing the feature information of infrared small targets; replace the standard convolution of the 19th and 22nd layers of the original head structure Head with the S_RepViTBlock module; use the NWD loss function as a whole in the model;

[0013] S3. Detection of small infrared targets: The input image is divided into S×S grids, and each grid predicts a fixed number of bounding boxes and their confidence and category probabilities; the confidence represents the probability that the bounding box contains an object, which is composed of the probability of the target existing and the degree of overlap between the predicted box and the true box; the output vector contains the bounding box coordinates, confidence and category probability; the improved YOLO11 model uses a comprehensive loss function to measure the prediction error; finally, non-maximum suppression (NMS) removes overlapping bounding boxes and retains the boxes with the highest confidence; and detects small infrared targets.

[0014] Compared with the prior art, the present application makes improvements based on the YOLO11 model, namely, using deep separable convolution and DSEAM modules to achieve standardized data distribution and accelerate the training of the model, so that during the operation, the feature information of small infrared targets can be enhanced while significantly reducing unnecessary forward propagation information; at the same time, the S_RepViTBlock module can be used to enhance the ability to extract details of background and target features while capturing small infrared targets using a larger receptive field, so as to fully reduce errors and improve the detection quality of small infrared targets.

[0015] Preferably, the design idea of ​​the efficient dual-channel feature enhancement module DSEAM module in S2 includes the following steps:

[0016] First, add a depth-wise separable convolution before the input of ECA, perform depth-wise convolution first, and batch normalize the output. This operation can standardize data distribution and accelerate model training. The batch normalization calculation formula is as follows:

[0017]

[0018] Where: μ is the mean of the current batch, σ 2 is the variance of the current batch, γ and β are learnable scaling and translation parameters, and ∈ is the zero division parameter;

[0019] Second, after output, the result is ReLU activated to increase the nonlinear expression of the model. The ReLU activation function formula is as follows:

[0020] x=ReLU(x)=max(0,x)ReLU maps all negative values ​​to 0, and keeps positive values ​​unchanged, which can reduce the gradient vanishing phenomenon; then point-by-point convolution is performed to adjust the number of channels, realize the linear combination of channel information, and perform batch normalization and ReLU activation on the output result again, and save the current feature tensor x as x_orig for subsequent residual connection;

[0021] Third, the information processing results are input into the ECA module for global average pooling to obtain the global features of each channel; the formula is:

[0022] y=AvgPool(x)

[0023] The shape of the output tensor y is (B, c2, 1, 1). Then, y is transformed into a one-dimensional convolution input format. A one-dimensional convolution is performed on y to capture the local dependencies between channels. Then, y is restored to its original dimension so that it can be multiplied with the input feature x. A temperature-adjusted Sigmoid activation is performed on y to generate the channel attention weight. The formula of the Sigmoid activation function with temperature adjustment is as follows:

[0024]

[0025] Where T is the temperature parameter, which is set to 0.5 by default. The shape of the final output attention channel weight CA is (B, c2, 1, 1); CA is expanded to the shape (B, c2, H, W) by element-by-element multiplication with the feature x; the input feature x (B, C, H, W) is average pooled (AvgOut) to obtain (B, 1, H, W) and the maximum pooled (MaxOut) to obtain (B, 1, H, W); the formula is:

[0026] AvgOut=Mean(x,dim=1,keepdim=True)

[0027] MaxOut=Max(x,dim=1,keepdim=True)

[0028] The output average pooling result AvgOut has a shape of (B, 1, H, W), and the maximum pooling result MaxOut has a shape of (B, 1, H, W); the former can extract global information, and the latter can extract significant features within the channel;

[0029] Fourth, the two pooled feature maps are fused with learnable weight parameters α and β; the fusion formula is as follows:

[0030] Fusion=α×AvgOut+β×MaxOut

[0031] The initial values ​​of α and β are both 0.5. The fused feature Fusion is convolved to extract the spatial attention feature. The convolution kernel size is set to 5, which can take into account the extraction of features at all scales. The convolution result is activated by Sigmoid, and the weight is compressed between 0 and 1. The network can suppress unimportant areas and highlight important areas, generating weights with spatial attention. The formula is:

[0032]

[0033] The output SA shape is (B, 1, H, W), and then it is expanded to make its shape become (B, c2, H, W); the learnable parameters α and β are used to perform weighted fusion with the twice expanded tensors. In order to ensure the normalization of the weights, the Softmax function is used to normalize [α, β]; formula:

[0034] [ω α ,ω β ]=Softmax([α,β])

[0035] CombineOut = ω α × ca +ωβ × sa

[0036] Among them, xca and xsa are the output tensors of the previous feature x after channel weight expansion and spatial weight expansion, and the shape of the fused feature CombinedOut is (B, c2, H, W);

[0037] Fifth, add the fused features to the previously saved features x_orig to realize residual connection, which can enhance the information flow and prevent the gradient from disappearing. The final output feature OutPut shape is (B, c2, H, W). In practical applications, the DSEAM module can enhance the features of small and densely detected targets in complex environments and reduce the amount of calculation and parameters of the model.

[0038] Furthermore, the features of small and densely detected targets can be enhanced in complex environments, and the computational complexity and parameter amount of the model can be reduced.

[0039] Preferably, the design steps of the S_RepViTBlock module in S2 include:

[0040] First, RepViT achieves a good balance between accuracy and latency on mobile devices by integrating the efficient Vit architecture into CNN. However, based on the infrared detection images of drones, the images may have noise or pixel quality degradation due to factors such as long distance, lens contamination, and complex mission scenarios, making it difficult to separate the target from the background when mixed together. The RepViT module is improved and added to YOLO11.

[0041] Second, the expansion convolution is introduced in the input part to increase the receptive field. While capturing small infrared targets, the larger receptive field can be used to enhance the ability to extract details of background and target features. In order to enhance the generalization ability of the model and improve the model's ability to deal with noise and fuzzy features, DropBlock regularization is introduced. DropbBock randomly blocks continuous rectangular areas on the feature map, allowing the model to learn features in a wider context. The formula is:

[0042]

[0043] Where B is the batch size, C is the number of channels, H and W are the height and width respectively, p is the zero probability, and the occlusion size is k×k;

[0044] Third, the improved RepViT module is named S_RepViT.

[0045] Furthermore, the S_RepViT module can extract small infrared targets in the infrared detection images of drones, and enhance the ability to extract details of background and target features through a larger receptive field.

[0046] Preferably, the error in S3 includes position error (MSE calculation), confidence error and category error (Cross-Entropy metric).

[0047] Furthermore, the measurement error is fully reduced to improve the quality of subsequent inspections.

[0048] The beneficial effects of the present invention are:

[0049] 1. The YOLO series model is a multi-scale single-stage target detection algorithm based on regression. Based on a fully convolutional network, the input image is divided into S×S grids. Each grid predicts a fixed number of bounding boxes and their confidence and category probabilities, which can help reduce the impact of other values ​​on infrared small target detection and improve the quality of infrared small target detection.

[0050] 2. Through the depth-wise separable convolution, the output is batch normalized. This operation can standardize the data distribution and accelerate the training of the model. It can enhance the features of small and dense detection targets in complex environments and reduce the amount of calculation and parameters of the model.

[0051] 3. The S_RepViT module can be used to extract small infrared targets in the infrared detection image of the drone, and the ability to extract details of the background and target features can be enhanced through a larger receptive field. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a YOLO11 model structure diagram in a UAV infrared small target detection method based on an attention mechanism proposed in the present invention;

[0053] Figure 2 This is a diagram of the deep separable convolution structure in the UAV infrared small target detection method based on the attention mechanism proposed in the present invention;

[0054] Figure 3 This is a structural diagram of the DSEAM module in the UAV infrared small target detection method based on the attention mechanism proposed in the present invention;

[0055] Figure 4 This is the structure diagram of S_RepViTBlock in the UAV infrared small target detection method based on attention mechanism proposed in this invention. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0057] Reference Figure 1-4 ,A method for detecting small infrared targets of UAV based on attention mechanism, includes the following steps:

[0058] S1. Detection model determination: The YOLO11 model is used, which consists of three parts: backbone network (Backbone), neck network (Neck), and head network (Head);

[0059] S2. Model revision: Add the efficient dual-channel feature enhancement module DSEAM module to the third and sixth layers of the YOLO11 backbone network, which can significantly reduce unnecessary forward propagation information while enhancing the feature information of infrared small targets; replace the standard convolution of the 19th and 22nd layers of the original head structure Head with the S_RepViTBlock module; use the NWD loss function as a whole in the model;

[0060] S3. Detection of small infrared targets: The input image is divided into S×S grids, and each grid predicts a fixed number of bounding boxes and their confidence and category probabilities; the confidence indicates the probability that the bounding box contains an object, which is composed of the probability of the target existence and the degree of overlap between the predicted box and the true box; the output vector contains the bounding box coordinates, confidence and category probability; the improved YOLO11 model uses a comprehensive loss function to measure the prediction error, which includes position error (MSE calculation), confidence error and category error (Cross-Entropy metric); finally, non-maximum suppression (NMS) removes overlapping bounding boxes and retains the box with the highest confidence; and detects small infrared targets.

[0061] In the present invention, the design idea of ​​the efficient dual-channel feature enhancement module DSEAM module in S2 includes the following steps:

[0062] First, add a depth-wise separable convolution before the input of ECA, perform depth-wise convolution first, and batch normalize the output. This operation can standardize data distribution and accelerate model training. The batch normalization calculation formula is as follows:

[0063]

[0064] Where: μ is the mean of the current batch, σ 2 is the variance of the current batch, γ and β are learnable scaling and translation parameters, and ∈ is the zero division parameter;

[0065] Second, after output, the result is ReLU activated to increase the nonlinear expression of the model. The ReLU activation function formula is as follows:

[0066] x = ReLU(x) = max(0, x)

[0067] ReLU maps all negative values ​​to 0, and keeps positive values ​​unchanged, which can reduce the gradient vanishing phenomenon. Then, point-by-point convolution is performed to adjust the number of channels, realize the linear combination of channel information, and perform batch normalization and ReLU activation on the output results again, saving the current feature tensor x as x_orig for subsequent residual connection.

[0068] Third, the information processing results are input into the ECA module for global average pooling to obtain the global features of each channel; the formula is:

[0069] y=AvgPool(x)

[0070] The shape of the output tensor y is (B, c2, 1, 1). Then, y is transformed into a one-dimensional convolution input format. A one-dimensional convolution is performed on y to capture the local dependencies between channels. Then, y is restored to its original dimension so that it can be multiplied with the input feature x. A temperature-adjusted Sigmoid activation is performed on y to generate the channel attention weight. The formula of the Sigmoid activation function with temperature adjustment is as follows:

[0071]

[0072] Where T is the temperature parameter, which is set to 0.5 by default. The shape of the final output attention channel weight CA is (B, c2, 1, 1); CA is expanded to the shape (B, c2, H, W) by element-by-element multiplication with the feature x; the input feature x (B, C, H, W) is average pooled (AvgOut) to obtain (B, 1, H, W) and the maximum pooled (MaxOut) to obtain (B, 1, H, W); the formula is:

[0073] AvgOut=Mean(x,dim=1,keepdim=True)

[0074] MaxOut=Max(x,dim=1,keepdim=True)

[0075] The output average pooling result AvgOut has a shape of (B, 1, H, W), and the maximum pooling result MaxOut has a shape of (B, 1, H, W); the former can extract global information, and the latter can extract significant features within the channel;

[0076] Fourth, the two pooled feature maps are fused with learnable weight parameters α and β; the fusion formula is as follows:

[0077] Fusion=α×AvgOut+β×MaxOut

[0078] The initial values ​​of α and β are both 0.5. The fused feature Fusion is convolved to extract the spatial attention feature. The convolution kernel size is set to 5, which can take into account the extraction of features at all scales. The convolution result is activated by Sigmoid, and the weight is compressed between 0 and 1. The network can suppress unimportant areas and highlight important areas, generating weights with spatial attention. The formula is:

[0079]

[0080] The output SA shape is (B, 1, H, W), and then it is expanded to make its shape become (B, c2, H, W); the learnable parameters α and β are used to perform weighted fusion with the twice expanded tensors. In order to ensure the normalization of the weights, the Softmax function is used to normalize [α, β]; formula:

[0081] [ω α ,ω β ]=Softmax([α,β])

[0082] CombineOut = ω α × ca +ω β × sa

[0083] Among them, xca and xsa are the output tensors of the previous feature x after channel weight expansion and spatial weight expansion, and the shape of the fused feature CombinedOut is (B, c2, H, W);

[0084] Fifth, add the fused features to the previously saved features x_orig to realize residual connection, which can enhance the information flow and prevent the gradient from disappearing. The final output feature OutPut shape is (B, c2, H, W). In practical applications, the DSEAM module can enhance the features of small and densely detected targets in complex environments and reduce the amount of calculation and parameters of the model.

[0085] In the present invention, the design steps of the S_RepViTBlock module in S2 include:

[0086] First, RepViT achieves a good balance between accuracy and latency on mobile devices by integrating the efficient Vit architecture into CNN. However, based on the infrared detection images of drones, the images may have noise or pixel quality degradation due to factors such as long distance, lens contamination, and complex mission scenarios, making it difficult to separate the target from the background when mixed together. The RepViT module is improved and added to YOLO11.

[0087] Second, the expansion convolution is introduced in the input part to increase the receptive field. While capturing small infrared targets, the larger receptive field can be used to enhance the ability to extract details of background and target features. In order to enhance the generalization ability of the model and improve the model's ability to deal with noise and fuzzy features, DropBlock regularization is introduced. DropbBock randomly blocks continuous rectangular areas on the feature map, allowing the model to learn features in a wider context. The formula is:

[0088]

[0089] Where B is the batch size, C is the number of channels, H and W are the height and width respectively, p is the zero probability, and the occlusion size is k×k;

[0090] Third, the improved RepViT module is named S_RepViT.

[0091] In the present invention, the drone takes aerial photos of the ground conditions, and after the aerial photos are obtained, they are transmitted to the improved YOLO11 model to process the aerial photos. First, the YOLO11 model can divide the aerial photos to form an S×S grid. At the same time, the staff revise the YOLO11 model. For example, the DSEAM module can enhance the features of small and dense detection targets in complex environments, and reduce the calculation amount and parameter amount of the model to improve the calculation efficiency.

[0092] At the same time, the S_RepViT module uses a larger receptive field to enhance the ability to extract details of background and target features, so as to improve the ability to capture small infrared targets and more comprehensively obtain small infrared targets from drone aerial images.

[0093] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A method for detecting small infrared targets of unmanned aerial vehicles based on attention mechanism, characterized in that: The following steps are involved: S1. Detection model determination: The YOLO11 model is used, which consists of three parts: backbone network (Backbone), neck network (Neck), and head network (Head); S2. Model revision: Add the efficient dual-channel feature enhancement module DSEAM module to the third and sixth layers of the YOLO11 backbone network, which can significantly reduce unnecessary forward propagation information while enhancing the feature information of infrared small targets; replace the standard convolution of the 19th and 22nd layers of the original head structure Head with the S_RepViTBlock module; use the NWD loss function as a whole in the model; S3. Detection of small infrared targets: The input image is divided into S×S grids, and each grid predicts a fixed number of bounding boxes and their confidence and category probabilities; the confidence represents the probability that the bounding box contains an object, which is composed of the probability of the target existing and the degree of overlap between the predicted box and the true box; the output vector contains the bounding box coordinates, confidence and category probability; the improved YOLO11 model uses a comprehensive loss function to measure the prediction error; finally, non-maximum suppression (NMS) removes overlapping bounding boxes and retains the boxes with the highest confidence; and detects small infrared targets.

2. According to the attention mechanism-based UAV infrared small target detection method of claim 1, it is characterized in that: The design idea of ​​the efficient dual-channel feature enhancement module DSEAM module in S2 includes the following steps: First, add a depth-wise separable convolution before the input of ECA, perform depth-wise convolution first, and batch normalize the output. This operation can standardize data distribution and accelerate model training. The batch normalization calculation formula is as follows: Where: μ is the mean of the current batch, σ 2 is the variance of the current batch, γ and β are learnable scaling and translation parameters, and ∈ is the zero division parameter; Second, after output, the result is ReLU activated to increase the nonlinear expression of the model. The ReLU activation function formula is as follows: x=ReLU(x)=max(0,x) ReLU maps all negative values ​​to 0, and keeps positive values ​​unchanged, which can reduce the gradient vanishing phenomenon. Then, point-by-point convolution is performed to adjust the number of channels, realize the linear combination of channel information, and perform batch normalization and ReLU activation on the output results again, saving the current feature tensor x as x_orig for subsequent residual connection. Third, the information processing results are input into the ECA module for global average pooling to obtain the global features of each channel; the formula is: y=AvgPool(x) The shape of the output tensor y is (B, c2, 1, 1). Then, y is transformed into a one-dimensional convolution input format. A one-dimensional convolution is performed on y to capture the local dependencies between channels. Then, y is restored to its original dimension so that it can be multiplied with the input feature x. A temperature-adjusted Sigmoid activation is performed on y to generate the channel attention weight. The formula of the Sigmoid activation function with temperature adjustment is as follows: Where T is the temperature parameter, which is set to 0.5 by default. The shape of the final output attention channel weight CA is (B, c2, 1, 1); CA is expanded to the shape (B, c2, H, W) by element-by-element multiplication with the feature x; the input feature x (B, C, H, W) is average pooled (AvgOut) to obtain (B, 1, H, W) and the maximum pooled (MaxOut) to obtain (B, 1, H, W); the formula is: AvgOut=Mean(x, dim=1, keepdim=True) MaxOut=Max(x, dim=1, keepdim=True) The output average pooling result AvgOut has a shape of (B, 1, H, W), and the maximum pooling result MaxOut has a shape of (B, 1, H, W); the former can extract global information, and the latter can extract significant features within the channel; Fourth, the two pooled feature maps are fused with learnable weight parameters α and β; the fusion formula is as follows: Fusion=α×AvgOut+β×MaxOut The initial values ​​of α and β are both 0.

5. The fused feature Fusion is convolved to extract the spatial attention feature. The convolution kernel size is set to 5, which can take into account the extraction of features at all scales. The convolution result is activated by Sigmoid, and the weight is compressed between 0 and 1. The network can suppress unimportant areas and highlight important areas, generating weights with spatial attention. The formula is: The output SA shape is (B, 1, H, W), and then it is expanded to make its shape become (B, c2, H, W); the learnable parameters α and β are used to perform weighted fusion with the twice expanded tensors. In order to ensure the normalization of the weights, the Softmax function is used to normalize [α, β]; formula: [oh α Oh, oh β ]=softmax([α,β]) CombineOut=ω a ×x ca +oh β ×x sa Among them, xca and xsa are the output tensors of the previous feature x after channel weight expansion and spatial weight expansion, and the shape of the fused feature CombinedOut is (B, c2, H, W); Fifth, add the fused features to the previously saved features x_orig to realize residual connection, which can enhance information flow and prevent gradient disappearance; the final output feature OutPut shape is (B, c2, H, W); in practical applications, the DSEAM module can enhance the features of small and densely detected targets in complex environments and reduce the amount of calculation and parameters of the model.

3. The method for detecting small infrared targets of unmanned aerial vehicles based on an attention mechanism according to claim 1, characterized in that: The design steps of the S_RepViTBlock module in S2 include: First, RepViT achieves a good balance between accuracy and latency on mobile devices by integrating the efficient Vit architecture into CNN. However, based on the infrared detection images of drones, the images may have noise or pixel quality degradation due to factors such as long distance, lens contamination, and complex mission scenarios, making it difficult to separate the target from the background when mixed together. The RepViT module is improved and added to YOLO11. Second, the expansion convolution is introduced in the input part to increase the receptive field. While capturing small infrared targets, the larger receptive field can be used to enhance the ability to extract details of background and target features. In order to enhance the generalization ability of the model and improve the model's ability to deal with noise and fuzzy features, DropBlock regularization is introduced. DropbBock randomly blocks continuous rectangular areas on the feature map, allowing the model to learn features in a wider context. The formula is: Where B is the batch size, C is the number of channels, H and W are the height and width respectively, p is the zero probability, and the occlusion size is k×k; Third, the improved RepViT module is named S_RepViT.

4. The method for detecting small infrared targets of unmanned aerial vehicles based on an attention mechanism according to claim 1, characterized in that: The errors in S3 include position error (MSE calculation), confidence error and category error (Cross-Entropy metric).

Citation Information

Patent Citations

  • Lightweight target detection method

    CN114120019A

  • Remote sensing image target detection method based on multi-scale feature extraction

    CN118230180A

  • Method for detecting small target in aerial image of unmanned aerial vehicle

    CN118762168A

  • Intelligent high-altitude camera cleaning method based on unmanned aerial vehicle

    CN119216290A

  • Defective pixel detection model training method, defective pixel detection method, and defective pixel repair method

    WO2024087163A1

Cited By

  • Malaria pathogen detection system based on improved YOLO algorithm

    CN120376009A

  • Driver distraction detection and identification method based on wavelet transform

    CN120852403A

  • A driver distraction detection and recognition method based on wavelet transform

    CN120852403B