Infrared image small target detection method and system based on adaptive receptive field and cross-scale fusion network

By constructing an adaptive receptive field and cross-scale fusion network, the complex background interference and regression ratio perturbation problems in infrared small object detection are solved, and high-precision detection in different scenarios is achieved, reducing the frequency of network adjustment.

CN120388186APending Publication Date: 2025-07-29HARBIN INST OF TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510456311.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The precise detection problem of infrared small object detection under complex background interference and regression perturbation is especially in different detection scenarios. The network performance is unstable, and existing methods require frequent adjustment and retraining.

Method used

Build an adaptive receptive field module, a two-way cross attention module, a cross-scale feature encoding fusion module and a dynamic balance loss function, establish an infrared small object detection network model, and use training data to optimize network parameters to realize adaptive feature capture and cross-scale fusion.

Benefits of technology

It effectively improves the accuracy of infrared small object detection, can detect stably in different scenarios, reduce network adjustment frequency, and improve bounding box regression effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388186A_ABST
    Figure CN120388186A_ABST
Patent Text Reader

Abstract

The invention discloses an infrared image small target detection method based on an adaptive receptive field and a cross-scale fusion network, and relates to the technical field of infrared image data processing. The objective of the invention is to solve the problems of complex background interference and regression frame intersection-to-parallel ratio disturbance during infrared small target detection. The method comprises the following steps: (1) constructing a self-adaptive receptive field module; (2) constructing a bidirectional cross attention module; (3) constructing a cross-scale feature coding fusion module; (4) constructing a dynamic balance loss function; (5) establishing an infrared small target detection network model; (6) training the network by using the existing data to obtain a network model; and (7) detecting an infrared small target by using the trained model. According to the method, inherent unique priori knowledge in different scenes can be effectively utilized, the loss of small target details is made up, and the detection precision of the infrared small target is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of infrared image processing, and particularly to an infrared image small target detection method and system based on an adaptive receptive field and a cross-scale fusion network. Background Art

[0002] Infrared imaging technology has a wide range of applications in the fields of national defense, industry, medical treatment, etc. Especially in the field of national defense, high-performance infrared imaging systems have advantages such as long operating range and all-weather operation. In images, long-distance interesting targets usually appear as small targets. Therefore, infrared small target detection is crucial in many applications. Infrared small targets usually consist of only a few pixels, and due to the long imaging distance, they lack specific shape and texture features. In addition, these small targets are often hidden in complex backgrounds, such as clouds. In recent years, due to the need for real-time warning, single-frame detection tasks have received much attention. Therefore, it is of great significance to achieve accurate detection of small targets in single-frame infrared images under the interference of complex backgrounds.

[0003] Currently, some researchers have tried to model the detection task as a bounding box regression problem and capture the features of infrared small targets through feature fusion. However, due to the limited utilization of background semantic information, when the detection scenario changes (such as in a maritime, urban or rural environment), the performance of the network may be affected. In some cases, the network even needs to replace the dataset and retrain to achieve the best detection effect. In addition, since the true bounding boxes of infrared small targets usually have only a few pixels, pixel-level errors that occur during the bounding box generation process may cause significant fluctuations in the intersection over union (IoU). The bounding box regression of infrared small targets faces greater challenges compared to conventional object detection tasks. Therefore, it is necessary to design an infrared image small target detection method to overcome the current challenges. Summary of the Invention

[0004] The technical problem to be solved by the present invention is:

[0005] The purpose of the present invention is to provide an infrared image small target detection method and system based on an adaptive receptive field and a cross-scale fusion network to solve the problems of complex background interference and regression box IoU perturbation that occur during infrared small target detection.

[0006] The technical solution adopted by the present invention to solve the above problems is:

[0007] (1) Construct an adaptive receptive field module;

[0008] (2) Construct a bidirectional cross-attention module;

[0009] (3) Construct a cross-scale feature encoding and fusion module;

[0010] (4) Construct a dynamic balance loss function;

[0011] (5) Establish an infrared small target detection network model;

[0012] (6) Use the existing data to train the network to obtain the network model;

[0013] (7) Use the trained model to detect infrared small targets.

[0014] Furthermore, the construction of the adaptive receptive field module in the above step (2) can be carried out according to the following steps:

[0015] For a given input, the input is divided into two parts by channel. One part is used for residual operation, and the other part is used for continuous downward conduction. Subsequently, the continuously transmitted part is successively passed through convolutional kernels of 3×3, 5×5, and 7×7, and the results of each convolution are retained. The obtained results are merged, and then the results after max pooling and average pooling are combined. Subsequently, a convolutional layer is connected to convert the pooled features into N spatial attention maps, and the formula is as follows:

[0016] F space =f 2-N ([AvgPool(F); MaxPool(F)])(1)

[0017] where F space represents the spatial attention map. f 2-N represents the convolutional operation that converts 2 channels into N channels. AvgPool represents average pooling. MaxPool represents max pooling. F represents the splicing of results under different receptive fields. For each spatial attention map, the sigmoid activation function is used to respectively obtain the spatial selection masks under different receptive fields:

[0018] M i =σ(F space (i))(2)

[0019] where σ(g) represents the sigmoid function. F space (i) represents the i-th spatial attention map. Multiply the feature maps of different receptive fields by their corresponding spatial selection masks, and then fuse them through a convolutional operation to obtain the final attention features. Finally, multiply the attention features by the original input, as shown in the following formula:

[0020]

[0021] where F input represents the original input features.

[0022] The result after spatial kernel selection is residually connected to the input partitioned at the start of the step to preserve the spatial structure of the gradient. Finally, a 1×1 convolutional kernel is used to adjust the output channels of the entire module.

[0023] Furthermore, the construction of the bidirectional cross-attention module in step (2) above can be carried out according to the following steps:

[0024] For a given input, multi-scale context information is encoded through multiple one-dimensional convolutions in the x and y directions. Subsequently, the obtained results are fused. The specific calculation is as follows:

[0025]

[0026] Among them, Conv 1×1 represents a 1×1 convolution operation. Conv1D i x and Conv1D i y represent one-dimensional convolutions along the x-axis and y-axis dimensions respectively. Norm represents the normalization operation. F represents the input.

[0027] Cross-axis calculations are performed on the results obtained for the x and y axes, and finally the output feature map is obtained. The specific calculation is as follows:

[0028] F out = σ(Conv 1×1 (MHCA y (F y , F x , F x )) + Conv 1×1 (MHCA x (F x , F y , F y )))*F(6)

[0029] Among them, MHCA x and MHCA y represent multi-head cross-attention along the x-axis and y-axis respectively, and σ represents the sigmoid function.

[0030] Furthermore, for the construction of the cross-scale feature encoding and fusion module in step (3) above, multiple encodings are first performed. Based on the feature map of the scale to be processed, the feature maps of its previous scale and the next scale are selected for encoding. The large-size feature map is downsampled through a straight-through layer. Subsequently, a 1×1 convolutional kernel is used to adjust the number of channels after downsampling to make it consistent with that of the base-scale feature map. The specific calculation is as follows:

[0031] F p = Conv1×1 (PTL(F l ))(7)

[0032] Among them, PTL represents the direct layer. F l represents the large-scale feature map. Conv 1×1 represents the 1×1 convolution operation.

[0033] The small-scale feature map is upsampled by transposed convolution to transfer the local features of the low-resolution image. Subsequently, the number of channels after downsampling is adjusted by a 1×1 convolution kernel so that the number of channels is consistent with that of the base-scale feature map. The specific calculation is as follows:

[0034] F t =Conv 1×1 (Conv trans (F s ))(8)

[0035] Among them, Conv trans represents the transposed convolution operation. F s represents the small-scale feature map. Three feature maps of different scales are combined to obtain the multi-coded feature map. The specific calculation is as follows:

[0036] F multi =Concat(F p ,F m ,F t )(9)

[0037] Among them, F m represents the reference feature map.

[0038] Secondly, cross-scale feature extraction is performed. The large-scale feature map is enlarged by nearest neighbor upsampling so that its size is consistent with that of the small-scale feature map. The number of channels of the large-scale feature map is adjusted to be the same as that of the small-scale feature Figure 1 map through a 1×1 convolution kernel. The two feature maps are combined into a four-dimensional tensor, and a scale dimension is added. Then, feature extraction is performed through two three-dimensional convolution modules. The three-dimensional convolution module includes a three-dimensional convolution kernel, batch normalization, an activation function, and an inter-layer residual connection. Finally, the cross-scale feature map to be extracted is obtained.

[0039] Finally, a feature pyramid is built in a top-down manner. The input of the feature map for building the feature pyramid is the result of multi-feature coding. And at the bottom layer of the pyramid, the cross-scale feature and the highest-level output are fused through CBAM attention. CBAM is a hybrid attention that can serially generate attention feature map information in both the channel and spatial dimensions.

[0040] M c(F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))(10)

[0041] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))(11)

[0042] Among them, AvgPool represents average pooling. MaxPool represents max pooling. MLP is a multi-layer perceptron. σ represents the sigmoid function, and f 7×7 represents a convolutional operation with a filter size of 7×7. F represents the input. Then, bottom-up fusion is performed to construct a feature pyramid. The resulting feature pyramid will be fed into a decoupler to establish prediction boxes.

[0043] Furthermore, in step (4), a dynamic balance loss function is constructed. Focaler-IoU is a method for reconstructing the intersection over union loss through linear interval mapping, which enables the network to autonomously determine whether to pay more attention to difficult samples or easy samples. It is defined as follows:

[0044]

[0045] Among them, B represents the prediction box. B gt represents the ground truth box. d and u are balance parameters. The new bounding box loss function Focaler-CIoU is calculated as follows:

[0046]

[0047]

[0048] L Focaler-CIoU = 1 - CIoU + IoU - IoU Focaler (17)

[0049] Among them, b and b gt represent the centers of the prediction box and the ground truth box. ρ() represents the Euclidean distance. w and w gt represent the widths of the prediction box and the GT box. h and h gt represent the heights of the prediction box and the ground truth box.

[0050] The regression loss is calculated by combining Focaler-CIoU and the distribution focal loss. The formula for the distribution focal loss is as follows:

[0051] L DFL = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 ))(18)

[0052] Among them, S i and S i+1 are the predicted values output by the network, adjacent predicted values. y, y i , y i+1 are the actual values of the labels, label integral values, adjacent label integral values.

[0053] The classification loss adopts binary cross-entropy. The calculation formula is as follows:

[0054]

[0055] Among them, y i is a binary label. That is, 0 or 1. p(yi) is the probability that the output belongs to the label. N represents the number of groups of model prediction objects. Finally, the calculation formula of the final loss function is as follows:

[0056] Loss = w1L Focaler-CIoU +w2L BCE +w3L DFL (20)

[0057] Among them, w1, w2, and w3 are the weight coefficients of each loss.

[0058] Furthermore, in step (5), an infrared small target detection network model is established, and the adaptive receptive field module is added to the front of the network backbone. Subsequently, different levels of features are extracted through the convolution module C2f in YOLOv8. At the tail of the network backbone, a bidirectional cross-attention module is added. Subsequently, the multi-scale feature maps obtained by the network backbone are input into the cross-scale feature encoding and fusion module to construct a feature pyramid. The feature pyramid is decoupled using a decoupled head, that is, the regression branch and the class branch are decoupled. The decoupled result is used to update the network parameters in reverse using the dynamic balance loss function.

[0059] Furthermore, in the above step (6), the constructed network is trained using the infrared small target detection dataset to obtain the optimal network parameters for subsequent infrared image small target detection. The dataset is divided into a training set, a validation set, and a test set according to 5:2:3. The training learning rate is 0.001, and the learning rate is decayed using the cosine annealing algorithm. The training stops when the model is fully converged or reaches 300 epochs to obtain the optimal network parameters.

[0060] Furthermore, in the above step (7), the infrared image to be detected is input into the optimal network for detection to obtain the small target detection frame.

[0061] The present invention has the following beneficial technical effects:

[0062] The present invention effectively solves the problems of complex background interference and regression box intersection over union (IoU) perturbation in infrared small target detection. Technical key points of the present invention: (1) constructing an adaptive receptive field module; (2) constructing a bidirectional cross-attention module; (3) constructing a cross-scale feature encoding and fusion module; (4) constructing a dynamic balance loss function; (5) establishing an infrared small target detection network model; (6) training the network using existing data to obtain the network model; (7) using the trained model to detect infrared small targets. The present invention can effectively utilize the unique prior knowledge inherent in different scenarios, make up for the loss of small target details, and effectively improve the detection accuracy of infrared small targets. When facing problems such as complex background interference and regression box IoU perturbation in infrared small target detection, the present invention can adaptively capture different receptive fields and fuse the features of infrared small targets at different scales to improve the bounding box regression effect, thereby effectively enhancing the detection accuracy of infrared small targets.

[0063] The present invention is an infrared image small target detection method based on an adaptive receptive field and cross-scale fusion network. The infrared image small target detection method based on an adaptive receptive field and cross-scale fusion network provided by the present invention can automatically detect targets from infrared image data. The algorithm proposed by the present invention can handle infrared small target detection tasks in different scenarios and has strong applicability. Experiments show that compared with existing infrared small target detection methods, the present invention can effectively utilize the unique prior knowledge inherent in different scenarios, make up for the loss of small target details, and significantly improve the detection accuracy of infrared small targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a flowchart of the method of the present invention.

[0065] Figure 2 is a schematic diagram of the overall structure of the network proposed by the present invention (schematic diagram of the overall structure of the infrared image small target detection method based on an adaptive receptive field and cross-scale fusion network proposed by the present invention);

[0066] Figure 3 is a schematic diagram of the structure of the adaptive receptive field module;

[0067] Figure 4 is a schematic diagram of the structure of the bidirectional cross-attention module;

[0068] Figure 5 is a comparison diagram of the detection results of the method of the present invention and other currently advanced methods. DETAILED DESCRIPTION OF THE INVENTION

[0069] The following combines the attached Figures 1-5 and examples to describe the present invention in detail.

[0070] The flowchart of the method of the present invention is shown in Figure 1 , and the specific implementation steps are as follows:

[0071] (1) Construct an adaptive receptive field module;

[0072] (2) Construct a bidirectional cross-attention module;

[0073] (3) Construct a cross-scale feature encoding and fusion module;

[0074] (4) Construct a dynamic balance loss function;

[0075] (5) Establish an infrared small target detection network model;

[0076] (6) Use the existing data to train the network to obtain the network model;

[0077] (7) Use the trained model to detect infrared small targets.

[0078] The above step (1) is carried out as follows:

[0079] For the given input, the input is bisected by channel, one part is used for residual operation, and the other part is used for continuous downward conduction. Subsequently, the continuously transmitted part is successively passed through convolutional kernels of 3×3, 5×5, and 7×7, and the results of each convolution are retained. The obtained results are merged, and then the results after max pooling and average pooling are combined. Subsequently, a convolutional layer is connected to convert the pooled features into N spatial attention maps, and the formula is as follows:

[0080] F space = f 2-N ([AvgPool(F); MaxPool(F)])(1)

[0081] where F space represents the spatial attention map. f 2-N represents the convolutional operation that converts 2 channels into N channels. AvgPool represents average pooling. MaxPool represents max pooling. F represents the concatenation of results under different receptive fields. For each spatial attention map, the sigmoid activation function is used to obtain the spatial selection masks under different receptive fields respectively:

[0082] M i = σ(F space (i))(2)

[0083] where σ(g) represents the sigmoid function. F space (i) represents the i-th spatial attention map. Multiply the feature maps of different receptive fields by their corresponding spatial selection masks, and then fuse them through a convolutional operation to obtain the final attention features. Finally, multiply the attention features by the original input, as shown in the following formula:

[0084]

[0085] Among them, F input represents the initial input feature.

[0086] The result after spatial kernel selection is residually connected to the input divided at the start of the step to preserve the spatial structure of the gradient. Finally, a 1×1 convolutional kernel is used to adjust the output channels of the entire module.

[0087] The above step (2) is carried out as follows:

[0088] For a given input, multi-scale context information is encoded through multiple one-dimensional convolutions in the x and y directions. Subsequently, the obtained results are fused. The specific calculation is expressed as follows:

[0089]

[0090] Among them, Conv 1×1 represents a 1×1 convolution operation. Conv1D i x and Conv1D i y represent one-dimensional convolutions along the x-axis and y-axis dimensions respectively. Norm represents the normalization operation. F represents the input.

[0091] Cross-axis calculations are performed on the results obtained for the x and y axes, and finally the output feature map is obtained. The specific calculation is expressed as follows:

[0092] F out = σ ( Conv 1×1 (MHCA y (F y , F x , F x )) + Conv 1×1 (MHCA x (F x , F y , F y )))*F(6)

[0093] Among them, MHCA x and MHCA y represent multi-head cross-attention along the x-axis and y-axis respectively, and σ represents the sigmoid function.

[0094] The above step (3) is carried out as follows:

[0095] First, perform multi - coding. Based on the feature map of the scale to be processed, select the feature maps of its previous scale and next scale for coding. Downsample the large - size feature map through the identity layer. Subsequently, use a 1×1 convolutional kernel to adjust the number of channels after downsampling to be the same as that of the base - scale feature map. The specific calculation is as follows:

[0096] F p =Conv 1×1 (PTL(F l ))(7)

[0097] Among them, PTL represents the identity layer. F l represents the large - size feature map. Conv 1×1 represents the 1×1 convolutional operation.

[0098] Upsample the small - size feature map through transposed convolution to transfer the local features of the low - resolution image. Subsequently, use a 1×1 convolutional kernel to adjust the number of channels after downsampling to be the same as that of the base - scale feature map. The specific calculation is as follows:

[0099] F t =Conv 1×1 (Conv trans (F s ))(8)

[0100] Among them, Conv trans represents the transposed convolutional operation. F s represents the small - size feature map. Combine the feature maps of three different scales to obtain the multi - coded feature map. The specific calculation is as follows:

[0101] F multi =Concat(F p ,F m ,F t )(9)

[0102] Among them, F m represents the reference feature map.

[0103] Secondly, perform cross - scale feature extraction. Enlarge the large - scale feature map through nearest - neighbor upsampling to make its size the same as that of the small - scale feature map. Use a 1×1 convolutional kernel to adjust the number of channels of the large - scale feature map to be the same as that of the small - scale Figure 1 feature map. Combine the two feature maps into a four - dimensional tensor and add a scale dimension. Then perform feature extraction through two three - dimensional convolutional modules. The three - dimensional convolutional module includes a three - dimensional convolutional kernel, batch normalization, activation function, and inter - layer residual skip connection. Finally, obtain the cross - scale feature map to be extracted.

[0104] Finally, a feature pyramid is built in a top-down manner. The input feature maps for constructing the feature pyramid are the results after multi-feature encoding. And at the bottom layer of the pyramid, cross-scale features and the highest-level output are fused through CBAM attention. CBAM is a hybrid attention mechanism that can sequentially generate attention feature map information in both the channel and spatial dimensions.

[0105] M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F)))(10)

[0106] M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)]))(11)

[0107] Among them, AvgPool represents average pooling. MaxPool represents max pooling. MLP is a multi-layer perceptron. σ represents the sigmoid function, and f 7×7 represents a convolutional operation with a filter size of 7×7. F represents the input. Subsequently, bottom-up fusion is performed to construct the feature pyramid. The obtained feature pyramid will be fed into the decoupler for the establishment of prediction boxes.

[0108] The above step (4) is carried out in the following manner:

[0109] Focaler-IoU is a method for reconstructing the intersection over union loss through linear interval mapping, which enables the network to autonomously determine whether to pay more attention to difficult samples or easy samples. It is defined as follows:

[0110]

[0111] Among them, B represents the prediction box. B gt represents the ground truth box. d and u are balance parameters. The new bounding box loss function Focaler-CIoU is calculated as follows:

[0112]

[0113] L Focaler-CIoU = 1 - CIoU + IoU - IoU Focaler (17)

[0114] Among them, b and b gt represent the centers of the prediction box and the ground truth box. ρ() represents the Euclidean distance. w and w gt represent the widths of the prediction box and the GT box. h and h gt represent the heights of the prediction box and the ground truth box.

[0115] The regression loss is calculated by combining Focaler-CIoU and distribution focal loss. The formula for the distribution focal loss is as follows:

[0116] L DFL = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 ))(18)

[0117] where S i and S i+1 are the predicted values and adjacent predicted values output by the network. y, y i , y i+1 are the actual values of the labels, the integral values of the labels, and the adjacent integral values of the labels.

[0118] The classification loss uses binary cross-entropy. The formula is as follows:

[0119]

[0120] where y i is the binary label, that is, 0 or 1. p(yi) is the probability that the output belongs to the label. N represents the number of groups of model prediction objects. Finally, the formula for the final loss function is as follows:

[0121] Loss = w1L Focaler-CIoU + w2L BCE + w3L DFL (20)

[0122] where w1, w2, and w3 are the weight coefficients of each loss.

[0123] The above step (5) is carried out in the following manner:

[0124] The adaptive receptive field module is added to the front of the network backbone. Subsequently, different levels of features are extracted through the convolutional module C2f in YOLOv8. At the end of the network backbone, a bidirectional cross-attention module is added. Subsequently, the multi-scale feature maps obtained from the network backbone are input into the cross-scale feature encoding and fusion module to construct a feature pyramid. The feature pyramid is decoupled using a decoupled head, that is, the regression branch and the classification branch are decoupled. The decoupled results are used to update the network parameters in reverse using the dynamic balance loss function.

[0125] The above step (6) is carried out in the following manner:

[0126] The constructed network is trained using an infrared small target detection dataset to obtain the parameters of the optimal network for subsequent infrared image small target detection. The dataset is divided into a training set, a validation set, and a test set in the ratio of 5:2:3. The training learning rate is 0.001, and the learning rate is decayed using the cosine annealing algorithm. The training stops when the model is fully converged or reaches 300 epochs, obtaining the parameters of the optimal network.

[0127] Step (7) above is carried out in the following manner:

[0128] The infrared image to be detected is input into the optimal network for detection to obtain a small target detection box.

[0129] To quantitatively evaluate the performance of the method proposed in the present invention, three quantitative evaluation indicators, namely Precision, Recall, and Average Precision (AP), are used in the experiment to evaluate the infrared small target detection effect. Precision is defined as follows:

[0130]

[0131] Among them, TP is the number of positive classes predicted as positive classes, and FP is the number of negative classes predicted as positive classes. Recall is defined as follows:

[0132]

[0133] Among them, FN is the number of positive classes predicted as negative classes. Average Precision (AP) is the area enclosed by the graph with the correct rate as the y-axis and the recall rate as the x-axis, and is usually calculated by integrating the area under the precision-recall (P-R) curve.

[0134] Figure 5 Shows a comparison chart of the detection results of the method of the present invention and other currently advanced methods. From left to right, it shows the input image, and the detection results of YOLOv7, YOLOv8, RT-DETR, MMRFF-Net, OSCAR, and the method proposed in the present invention. The white circles in the figure represent the missed detection targets and the misdetected targets. In a complex background, not all content that conforms to the model appearance prior is a real target. By extracting the semantics of the picture, the method proposed in the present invention can more accurately identify the target. Under different background interferences, the method proposed in the present invention obtains the best detection results.

[0135] Table 1 Quantitative results of the detection results of different methods on the infrared small target test dataset

[0136]

[0137]

[0138] Table 1 lists the quantitative comparison results of the method of the present invention with other currently advanced detection methods, including Precision, Recall, and Average Precision (AP). From the quantitative comparison results, it can be seen that the method proposed in the present invention exhibits the best performance and can significantly improve the detection accuracy of infrared small targets.

Claims

1. An infrared image small target detection method based on an adaptive receptive field and cross-scale fusion network, characterized in that, By using an adaptive receptive field to utilize the unique prior knowledge inherent in different scenarios and alleviating the low tolerance of the network to bounding box perturbations through cross-scale fusion, thereby improving the accuracy of infrared small target detection. The method includes the following steps: (1) Construct an adaptive receptive field module; (2) Construct a bidirectional cross-attention module; (3) Construct a cross-scale feature encoding and fusion module; (4) Construct a dynamic balance loss function; (5) Establish an infrared small target detection network model; (6) Use existing data to train the network to obtain a network model; (7) Use the trained model to detect infrared small targets.

2. The method according to claim 1, wherein The step (1) is carried out according to the following steps: For a given input, the input is divided into two parts by channel. One part is used for residual operation, and the other part is used for continuous downward transmission. Subsequently, the continuously transmitted part is successively passed through convolutional kernels of 3×3, 5×5, and 7×7, and the results of each convolution are retained. The obtained results are merged, and then the results after maximum pooling and average pooling are combined. Subsequently, a convolutional layer is connected to convert the pooled features into N spatial attention maps. The formula is as follows: F space = f 2-N ([AvgPool(F); MaxPool(F)]) (1) Among them, F space represents the spatial attention map; f 2-N represents the convolution operation that converts 2 channels into N channels; AvgPool represents average pooling; MaxPool represents max pooling; F represents the result splicing under different receptive fields; for each spatial attention map, the sigmoid activation function is used to obtain the spatial selection masks under different receptive fields respectively: M i = σ(F space (i)) (2) where, σ(g) represents the sigmoid function; F space (i) represents the i-th spatial attention map; the feature maps with different receptive fields are multiplied by their corresponding spatial selection masks, and then fused through a convolution operation to obtain the final attention feature; finally, the attention feature is multiplied by the original input, as shown in the following formula: Among them, F input represents the initial input feature; The result after spatial kernel selection is connected with the input divided at the beginning of the step through a residual connection to retain the spatial structure of the gradient. Finally, a 1×1 convolutional kernel is used to adjust the output channels of the entire module.

3. The method according to claim 1, characterized in that, The step (2) is carried out according to the following steps: For a given input, multi-scale context information is encoded in the x and y directions through multiple one-dimensional convolutions, and then the obtained results are fused. The specific calculation is as follows: Among them, Conv 1×1 represents a 1×1 convolution operation, and represent one-dimensional convolutions along the x-axis and y-axis dimensions respectively, Norm represents the normalization operation, and F represents the input; Cross-axis calculation is performed on the results obtained on the x and y axes, and finally the output feature map is obtained. The specific calculation is as follows: F out = σ(Conv 1×1 (MHCA y (F y , F x , F x )) + Conv 1×1 (MHCA x (F x , F y , F y )))*F(6) Among them, MHCA x and MHCA y represent multi-head cross attention along the x-axis and y-axis respectively, and σ represents the sigmoid function.

4. The method according to claim 1, wherein The step (3) is carried out in the following manner: First, multiple encodings are performed. Based on the feature map of the scale to be processed, the feature maps of its previous scale and next scale are selected for encoding. The large-size feature map is downsampled through a straight-through layer, and then a 1×1 convolutional kernel is used to adjust the number of channels after downsampling to be the same as that of the base-scale feature map. The specific calculation is as follows: F p = Conv 1×1 (PTL(F l )) (7) Among them, PTL represents the through layer, and F l represents the large-size feature map, and Conv 1×1 represents the 1×1 convolution operation; The small-size feature map is upsampled through a transposed convolution to transfer the local features of the low-resolution image. Subsequently, a 1×1 convolutional kernel is used to adjust the number of channels after downsampling to be the same as that of the base-scale feature map. The specific calculation is as follows: F t = Conv 1×1 (Conv trans (F s )) (8) Among them, Conv trans represents the transposed convolution operation, and F s represents the small-sized feature map. The feature maps of three different scales are merged to obtain the feature map after multiple encodings. The specific calculation is shown as follows: F multi = Concat(F p , F m , F t ) (9) Among them, F m represents the reference feature map; Secondly, cross-scale feature extraction is performed. The large-scale feature map is enlarged through nearest neighbor upsampling to make its size the same as that of the small-scale feature map. The number of channels of the large-scale feature map is adjusted to be the same as that of the small-scale feature map through a 1×1 convolutional kernel. The two feature maps are combined into a four-dimensional tensor, and a scale dimension is added. Then, feature extraction is performed through two three-dimensional convolutional modules. The three-dimensional convolutional module includes a three-dimensional convolutional kernel, batch normalization, an activation function, and an inter-layer residual connection. Finally, the cross-scale feature map to be extracted is obtained; Finally, a feature pyramid is built in a top-down manner. The feature maps input for constructing the feature pyramid are the results after multi-feature encoding. At the bottom layer of the pyramid, cross-scale features and the highest-level output are fused through CBAM attention. CBAM is a hybrid attention mechanism that can serially generate attention feature map information in both the channel and spatial dimensions; M c (F) = σ(MLP(AvgPool(F)) + MLP(MaxPool(F))) (10) M s (F) = σ(f 7×7 ([AvgPool(F); MaxPool(F)])) (11) Among them, AvgPool represents average pooling, MaxPool represents max pooling, MLP is a multi-layer perceptron, σ represents the sigmoid function, and f 7×7 represents a convolution operation with a filter size of 7×7, and F represents the input; subsequently, bottom-up fusion is performed to construct a feature pyramid, and the obtained feature pyramid will be passed into a decoupler to establish prediction boxes.

5. The method according to claim 4, characterized in that The step (4) is carried out in the following manner: Focaler-IoU is a method for reconstructing the intersection over union loss through linear interval mapping, which enables the network to autonomously determine whether to pay more attention to difficult samples or easy samples, and is defined as follows: Among them, B represents the predicted bounding box, and B gt represents the ground truth bounding box. d and u are balance parameters, and the new bounding box loss function Focaler-CIoU is calculated as follows: L Focaler-CIoU = 1 - CIoU + IoU - IoU Focaler (17) where b and b gt represent the center points of the predicted box and the ground truth box, ρ() represents the Euclidean distance, w and w gt represent the widths of the predicted box and the GT box, h and h gt represent the heights of the predicted box and the ground truth box; The regression loss is calculated by combining Focaler-CIoU and the distribution focal loss. The formula for the distribution focal loss is as follows: L DFL = -((y i+1 - y) log(S i ) + (y - y i ) log(S i+1 )) (18) Among them, S i and S i+1 are the predicted values output by the network, adjacent predicted values, y, y i , y i+1 are the actual values of the labels, label integral values, adjacent label integral values; The classification loss uses binary cross-entropy, and the formula is as follows: where y i is a binary label, i.e., 0 or 1, p(yi) is the probability that the output belongs to the label, N represents the number of groups of model prediction objects, and finally the calculation formula of the final loss function is as follows: Loss=w1L Focaler-CIoU +w2L BCE +w3L DFL (20) Among them, w1, w2, and w3 are the weight coefficients of each loss.

6. The method according to claim 1, wherein The step (5) is carried out in the following manner: The adaptive receptive field module is added to the front of the network backbone. Subsequently, different levels of features are extracted through the convolutional module C2f in YOLOv8. At the end of the network backbone, a bidirectional cross-attention module is added. Then, the multi-scale feature maps obtained from the network backbone are input into the cross-scale feature encoding and fusion module to construct a feature pyramid. The feature pyramid uses a decoupled head, that is, the regression branch and the classification branch are decoupled, and the dynamic balance loss function is used to inversely update the network parameters for the decoupled results.

7. The method according to claim 1, characterized in that, The step (6) is carried out in the following manner: The constructed network can be trained using the infrared small target detection dataset to obtain the parameters of the optimal network for subsequent infrared image small target detection. The dataset is divided into a training set, a validation set, and a test set in a ratio of 5:2:

3. The training learning rate is 0.001, and the learning rate is decayed using the cosine annealing algorithm. The training stops when the model is fully converged or reaches 300 epochs to obtain the parameters of the optimal network.

8. The method according to claim 1, characterized in that The step (7) is carried out in the following manner: The infrared image to be detected is input into the optimal network for detection to obtain the small target detection box.

9. An infrared image small target detection system based on an adaptive receptive field and a cross-scale fusion network, characterized in that: This system has program modules corresponding to the steps of any one of the above claims 1-8, and when running, it executes the steps in the above infrared image small target detection method based on the adaptive receptive field and cross-scale fusion network.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of a method for infrared image small target detection based on an adaptive receptive field and cross-scale fusion network as described in any one of claims 1-8 when called by a processor.

Citation Information

Cited By

  • Infrared target detection method based on bidirectional receptive field attention feature network

    CN121305046A

  • An infrared target detection method based on a bidirectional receptive field attention feature network

    CN121305046B

  • Infrared unmanned aerial vehicle target detection method based on multi-scale self-enhancement cross-layer fusion

    CN121861523A

  • Infrared small target detection method supporting set driving and significance priori guidance

    CN122289827A