An infrared dim target detection method based on a space-time feature enhancement network

CN118674944BActive Publication Date: 2026-09-29SHANGHAI INSTITUTE OF TECHNICAL PHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410767533.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2026-09-29
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

[0003]天基场景下空间红外目标占据像素少且淹没在背景杂波中,仅凭单帧图像信息很难检测出红外暗弱目标,为了克服现有的算法对于低信噪比目标的检测能力不足的问题,本发明充分利用图像序列中的时空特性,利用目标的运动特性对红外暗弱目标进行增强

Benefits of technology

[0045]1.本发明针对低信噪比目标在静态图像中难以检测的情况,单帧图像时域信息不足的问题,提出时域目标增强模块TEM,该模块利用多帧数据提取时域上下文,扩大模型时域感受野,利用多帧数据中目标的运动信息沿着目标运动轨迹积累能量,利用目标的运动特性来增强目标和抑制背景。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674944B_ABST
    Figure CN118674944B_ABST
Patent Text Reader

Abstract

The application discloses an infrared dim and weak target detection method based on a space-time feature enhancement network, relates to the technical field of image processing and target detection, and has the technical points that first, an infrared dim and weak target image sequence dataset is constructed, an anchor frame suitable for the infrared dim and weak target is defined, and an image is preprocessed, the dimension of an image input model is unified, then, a space-time feature enhancement network model TSF-Net is constructed, the model specifically comprises an input end, a time-domain target enhancement module, a space-time feature extraction module, a feature fusion module and a detection head, then, the TSF-Net network is trained and tested, and finally, the dim and weak target detection capability of the TSF-Net network is evaluated. The application can effectively suppress the background and improve the detection rate, has strong robustness and feasibility, and can enhance the target by using the time-domain information in multiple frames of images for the case that a single frame of image is difficult to detect a low signal-to-noise ratio target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and target detection technology, specifically to a method for detecting faint infrared targets based on a spatiotemporal feature enhancement network. Background Technology

[0002] Infrared detection systems, with their advantages of strong concealment, wide field of view, and long-range detection, are widely used in aerospace, counter-reconnaissance, and target detection. However, due to the great distance from the target, background clutter, and atmospheric radiation interference, the target occupies only a few pixels on the imaging plane, has a low grayscale response, and lacks morphological and textural features, making it easily obscured by clutter. While single-frame algorithms are simple to implement, the limited information in a single frame makes them prone to generating numerous false alarms and missed detections in complex backgrounds. Multi-frame detection algorithms, while utilizing temporal information in the image to improve detection rates, are complex to implement, computationally intensive, and traditional algorithms still cannot adapt to changes in target background and shape, leading to many false alarms. Deep learning algorithms, while adaptable to background and target size variations and significantly improving detection capabilities in complex environments, are currently designed for targets with high signal-to-noise ratios (SNR), resulting in a significant decrease in detection capability for targets with low SNRs. Therefore, accurately and efficiently detecting infrared targets with extremely low SNRs remains a challenge. Summary of the Invention

[0003] In space-based scenarios, infrared targets occupy few pixels and are often obscured by background clutter, making it difficult to detect faint infrared targets using only a single frame of image information. To overcome the limitations of existing algorithms in detecting low signal-to-noise ratio (SNR) targets, this invention fully utilizes the spatiotemporal characteristics of image sequences and leverages the motion characteristics of targets to enhance the detection of faint infrared targets. This invention aims to provide a faint infrared target detection method based on a spatiotemporal feature enhancement network. This method effectively enhances faint infrared targets using temporal information from multiple frames of data, fully utilizing spatiotemporal features to improve the network's detection performance for low SNR targets.

[0004] To achieve the above objectives, the technical solution of the present invention is as follows: A method for detecting faint infrared targets based on a spatiotemporal feature enhancement network, comprising the following steps:

[0005] S1: Obtain the infrared image sequence dataset;

[0006] S2: Construct TSF-Net, an infrared weak target detection network with spatiotemporal feature enhancement for low signal-to-noise ratio targets; The TSF-Net network consists of an input terminal, a temporal target enhancement module (TEM), a spatiotemporal feature extraction module, a feature fusion module, and a detection head;

[0007] S3: Train the TSF-Net model using the training dataset constructed in step S1;

[0008] S4: Input the test dataset constructed in step S1 into the TSF-Net trained in step S3 to test the detection performance of TSF-Net.

[0009] S5: Evaluate the performance of the TSF-Net network for detecting low signal-to-noise ratio infrared targets.

[0010] Preferably, in step S1, acquiring the infrared image sequence dataset includes the following steps:

[0011] S101: Obtain training and test samples; construct M infrared weak target image sequences using public datasets and real-world datasets, each image with a resolution of R×R and 3 channels, and label the infrared weak targets on each frame image. Then, use H infrared weak target image sequences and their corresponding labels as training set X, and use MH infrared weak target image sequences and their corresponding labels as validation set V.

[0012] S102: The infrared image sequence is sampled using a sliding window sampling method, selecting the current N frames and the previous N frames to ensure the temporal continuity of the input TSF-Net image.

[0013] Preferably, in step S2, the constructed spatiotemporally enhanced infrared weak target detection network TSF-Net includes the following steps:

[0014] S201: The input end performs time-series reading of the infrared image sequence, reading in the current frame and its adjacent N frames before and after to ensure the temporal sequence of the input image; sets the detection anchor box for weak targets and preprocesses the input image;

[0015] S202: The temporal target enhancement module TEM consists of a three-dimensional convolutional layer, a three-dimensional max pooling layer, and a three-dimensional average pooling layer, which enhances the target using the input image sequence;

[0016] S203: The spatiotemporal feature extraction module consists of a global context module GCBlock and residual units, and is divided into five stages. It fully extracts spatiotemporal features from the temporally enhanced feature layer and outputs feature layers of different sizes.

[0017] S204: The feature fusion module consists of a multi-scale shallow feature enhancement module and a channel attention mechanism. It enhances the receptive field of shallow features using dilated convolutions of different scales, and adds or subtracts a simple channel attention mechanism to restrict spatial details for deep features. Then, the resulting deep feature maps and shallow feature maps are added and fused together.

[0018] S205: The detection head has four branches, which detect infrared weak targets of different scales respectively. Each branch is composed of a 1×1 convolution. The feature layers of different scales obtained in step S204 are used to predict the regression parameters of the target, the predicted target box and the confidence of the target through the 1×1 convolution layer.

[0019] S206: Based on the preset anchor box set in step S201, adjust the position of the predicted box to obtain a target candidate region that is closer to the true box. The calculation formula is as follows:

[0020] P x =2*σ(p x )-0.5+c x

[0021] P y =2*σ(p y )-0.5+c y

[0022] P w =A w (2σ(p w )) 2

[0023] P h =A h (2σ(p h )) 2

[0024] In the formula p x p y ,p w ,p h , where A is the predicted offset. w A h c is the preset length and width of the anchor frame. x c y This represents the coordinates of the top-left corner of the grid in the feature layer;

[0025] S207: Map the target candidate regions obtained in step S206 to feature layers of different scales of the detection head in step S205, and calculate the loss between the candidate target boxes and the prediction results; the loss consists of two parts: target box loss and confidence loss. The target box loss is calculated by CIOU loss, and the confidence loss is calculated by binary cross-entropy loss. The loss calculation is as follows:

[0026] L Loss =a*L ciou +b*L conf

[0027]

[0028] Where ρ represents the distance between the centroid of the predicted bounding box and the ground truth bounding box, c is the diagonal distance of the minimum bounding rectangle, and v is used as a correction factor to further adjust the loss function; in addition, w t ,h t ,w p ,h p Let A represent the width and length of the ground truth box and the predicted box, respectively. Let IOU be the intersection-union ratio of the predicted box and the ground truth box. Let A be the area of ​​the ground truth box and B be the area of ​​the predicted box. Weight factors a and b are used to weight the confidence loss and the localization loss to avoid imbalance between the two parts of the loss during training.

[0029] Preferably, in step S3, training the TSF-Net model includes the following steps:

[0030] S301: Train the TSF-Net model constructed in step S2. The training environment is NVIDIA GeForce RTX3090 GPU, the network training is Epoch=300, the initial learning rate Lr=0.0001, the learning rate optimizer is adamw, the batch size=16, and training starts from 0.

[0031] S302: Retain the optimal weights obtained from training in step S301 for model detection and evaluation.

[0032] Preferably, in step S4, testing the detection performance of TSF-Net includes the following steps:

[0033] S401: Input the test dataset constructed in step S1 into the optimal weights obtained after training in step S3, predict the data, and obtain the target bounding box and confidence score of the target prediction result;

[0034] S402: In the prediction results, results with a confidence level greater than the set threshold T1 are judged as correct results, and the corresponding detection box is the predicted target location;

[0035] S403: Perform NMS suppression on the prediction results based on the NMS threshold T2 to remove duplicate prediction boxes and obtain the final prediction results.

[0036] Preferably, in step S5, evaluating the TSF-Net network's performance in detecting low signal-to-noise ratio infrared targets includes the following steps:

[0037] S501: Accuracy is used to evaluate the precision of the network, calculated using the following formula:

[0038]

[0039] S502: Recall is used to evaluate the network's ability to detect all targets. The calculation formula is as follows:

[0040]

[0041] S503: The F1 score is used to evaluate the overall detection capability of the network. The calculation formula is as follows:

[0042]

[0043] Where TP represents the number of positive targets predicted as positive, FP represents the number of negative targets incorrectly classified as positive, and FN represents the number of positive targets incorrectly classified as negative.

[0044] Compared with existing technologies, the beneficial effects of this solution are:

[0045] 1. This invention addresses the problem of low signal-to-noise ratio targets being difficult to detect in static images and the lack of temporal information in a single frame image. It proposes a temporal target enhancement module (TEM). This module extracts temporal context from multi-frame data, expands the model's temporal receptive field, accumulates energy along the target's motion trajectory using the target's motion information from multi-frame data, and enhances the target and suppresses the background by utilizing the target's motion characteristics.

[0046] 2. This invention constructs a global context residual network as the backbone network to extract target features after temporal enhancement. The residual network not only enhances the ability to extract features of weak infrared targets, but also avoids the problem of target features being submerged by clutter due to excessive network depth, thus solving the problem of strong clutter interference. The global context module improves the model's ability to capture background features, suppresses clutter, and increases the detection rate.

[0047] 3. This invention uses a multi-scale feature fusion module to fuse features from multiple feature layers at different scales, adds a small target detection layer, and integrates contextual information of targets at different scales by enhancing the shallow feature layers at multiple scales, thus solving the problem of insufficient semantics in shallow features. For deep feature layers, we add a simple channel attention mechanism to constrain deep features using the spatial details of the shallow feature maps, thereby enriching the deep feature maps. Attached Figure Description

[0048] Figure 1 This is a flowchart of the detection method in an embodiment of the present invention;

[0049] Figure 2 This is a network structure diagram of the temporal target enhancement module in an embodiment of the present invention;

[0050] Figure 3 This is an overall network structure diagram in an embodiment of the present invention;

[0051] Figure 4 This is a network structure diagram of the multi-scale feature fusion module in an embodiment of the present invention;

[0052] Figure 5 This is a diagram showing the detection effect of low signal-to-noise ratio targets in an embodiment of the present invention. Detailed Implementation

[0053] To enable those skilled in the art to better understand the present invention, the technical solution of the present invention will be described in further detail below with reference to the embodiments and accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0054] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the embodiments.

[0055] Example:

[0056] like Figure 1-5 As shown, a method for detecting faint infrared targets based on spatiotemporal feature enhancement networks is presented. The overall algorithm principle and flow are as follows: Figure 1 As shown, the method includes the following steps:

[0057] 1) Obtain the infrared image sequence dataset.

[0058] 1.1 Obtaining Training and Testing Samples. M infrared weak target image sequences are constructed using public and real-world datasets. Each image has a resolution of R×R and 3 channels. The infrared weak targets in each frame are labeled. Then, H infrared weak target image sequences and their corresponding labels are used as the training set X, and MH infrared weak target image sequences and their corresponding labels are used as the validation set V. In this embodiment, M = 90, H = 72.

[0059] 1.2 The infrared image sequence is sampled using a sliding window sampling method. The current N frames and the previous N frames are selected to ensure that the input TSF-Net image has temporal continuity. In this embodiment, N=2.

[0060] 2) Construct TSF-Net, an infrared weak target detection network with spatiotemporal feature enhancement for low signal-to-noise ratio. The SF-Net network consists of an input terminal, a temporal target enhancement module (TEM), a spatiotemporal feature extraction module, a feature fusion module, and a detection head.

[0061] 2.1 The input end performs time-series reading of the infrared image sequence, reading in the current frame and its adjacent N frames before and after to ensure the temporal order of the input images. Detection anchor boxes for dark and small targets are set, and preprocessing of the input images is performed.

[0062] The temporal target enhancement module (TEM) in section 2.2 consists of a 3D convolutional layer, a 3D max pooling layer, and a 3D average pooling layer. It enhances the target using the input image sequence. A 3D convolutional layer with a kernel size of 3×3 extracts temporal information from multiple frames to expand the temporal receptive field. 3D max pooling extracts the temporal maximum value, and 3D average pooling extracts the temporal background. Finally, the temporal background obtained from 3D average pooling and the target enhancement feature map obtained from 3D max pooling are multiplied together to achieve target enhancement and background suppression along the target's motion trajectory.

[0063] 2.3 The spatiotemporal feature extraction module consists of a global context module GCBlock and residual units, divided into five stages to fully extract spatiotemporal features from the temporally enhanced feature layers. It outputs feature layers of different sizes. Each stage consists of seven 3×3 convolutional layers and two global context modules. Residual connections prevent the disappearance of infrared weak target features as the network deepens, while the global context modules extract global background information. By aggregating contextual features to enhance each location in the feature layer, long-range dependencies can be effectively modeled as a simplified nonlocal module. Combining the residual network and the context modules significantly improves the model's ability to distinguish between targets and backgrounds, and enhances the network's ability to extract features from weak targets.

[0064] The feature fusion module in section 2.4 consists of a multi-scale shallow feature enhancement module and a channel attention mechanism. For shallow features, it uses dilated convolutions of different scales to enhance the receptive field. For deep features, it adds or subtracts a simple channel attention mechanism to restrict spatial details. Then, the resulting deep and shallow feature maps are added and fused. The multi-scale feature fusion module first enhances the shallow features at multiple scales, using dilated convolutions with dilation rates of 1, 2, and 3 and a kernel size of 3×3, and ordinary convolutions with a kernel size of 3×3 to enhance the shallow feature layers, increasing the receptive field of the shallow features. The feature maps after convolution operations are merged into a unified tensor, and then dimensionality reduction is performed using a convolution with a kernel size of 1×1, ensuring that the feature dimension remains unchanged.

[0065] 2.5 The detection head has four branches, which detect infrared targets of different scales. Each branch consists of a 1×1 convolution. The feature layers of different scales obtained in step 2.4 are used to predict the regression parameters of the target, the predicted target box, and the target confidence through the 1×1 convolution layer.

[0066] 2.6 Based on the preset anchor boxes set in step 2.1, adjust the position of the predicted bounding box to obtain a target candidate region that is closer to the true bounding box. The calculation formula is as follows:

[0067] P x =2*σ(px )-0.5+c x

[0068] P y =2*σ(p y )-0.5+c y

[0069] P w =A w (2σ(p w )) 2

[0070] P h =A h (2σ(p h )) 2

[0071] In the formula p x p y ,p w ,p h , where A is the predicted offset. w A h c is the preset length and width of the anchor frame. x c y The coordinates are the top-left corner coordinates of the grid in the feature layer.

[0072] 2.7 Map the candidate target regions obtained in step 2.6 to feature layers of different scales in the detection head of step 2.5, and calculate the loss between the candidate target boxes and the prediction results. The loss consists of two parts: the target box loss and the confidence loss. The target box loss is calculated using the CIOU loss, and the confidence loss is calculated using the binary cross-entropy loss. The loss calculation is as follows:

[0073] L Loss =a*L ciou +b*L conf

[0074]

[0075] Where ρ represents the distance between the centroid of the predicted bounding box and the ground truth bounding box, c is the diagonal distance of the minimum bounding rectangle, and v is used as a correction factor to further adjust the loss function. Additionally, w t ,h t ,w p ,h p Let A represent the width and length of the ground truth bounding box and the predicted bounding box, respectively. IOU is the intersection-union ratio of the predicted and ground truth bounding boxes. A is the area of ​​the ground truth bounding box, and B is the area of ​​the predicted bounding box. Weighting factors a and b are used to weight the confidence loss and the localization loss to avoid imbalance between the two parts of the loss during training. In this implementation, a = 0.05 and b = 1.

[0076] 3) Train the TSF-Net model using the training dataset constructed in step 1.

[0077] 3.1 Train the TSF-Net model constructed in step 2. The training environment is NVIDIA GeForce RTX3090 GPU, the network training is Epoch=300, the initial learning rate Lr=0.0001, the learning rate optimizer is adamw, the batch size=16, and training starts from 0.

[0078] 3.2 Retain the optimal weights obtained from training in step 3.1 for model detection and evaluation.

[0079] 4) Input the test dataset constructed in step 1 into the TSF-Net trained in step 3 to test the detection performance of TSF-Net.

[0080] 4.1 Input the test dataset constructed in step 1 into the optimal weights obtained after training in step 3, and make predictions on the data to obtain the target bounding box and confidence score of the target prediction results.

[0081] 4.2 In the prediction results, results with a confidence level greater than the set threshold T1 are judged as correct results, and the corresponding detection box is the predicted target location.

[0082] 4.3 Perform NMS suppression on the prediction results based on the NMS threshold T2 to remove duplicate prediction boxes and obtain the final prediction results.

[0083] 5) Evaluate the performance of the TSF-Net network in detecting low signal-to-noise ratio infrared targets.

[0084] 5.1 Accuracy is used to evaluate the precision of the network, and the calculation formula is as follows:

[0085]

[0086] 5.2 Recall is used to evaluate the network's ability to detect all targets. The calculation formula is as follows:

[0087]

[0088] 5.3 The F1 score is used to evaluate the overall detection capability of the network. The calculation formula is as follows:

[0089]

[0090] Where TP represents the number of positive targets predicted as positive, FP represents the number of negative targets incorrectly classified as positive, and FN represents the number of positive targets incorrectly classified as negative.

[0091] The experimental results of this embodiment on a real infrared image sequence are as follows: Figure 5 As shown;

[0092] To demonstrate the detection effectiveness of the embodiments of the present invention, the detection results of the embodiments of the present invention for different signal-to-noise ratios are compared with those of existing detection networks. The results of various experimental indicators are shown in the table below:

[0093] Table 1F1

[0094]

[0095] Table 2Precision

[0096]

[0097] Table 3 Recall

[0098]

[0099] Referring to Tables 1, 2, and 3, the performance of various detection algorithms was compared. This invention achieved the best detection results across all metrics. Compared to infrared weak targets under extremely low signal-to-noise ratio (SNR) conditions, deep learning algorithms such as ACM, LPNet, and DNA-Net showed a significant decrease in detection performance after the SNR fell below 1.5. However, the method of this invention maintained a detection rate of over 90% even under an extremely low SNR of 0.83. This demonstrates the detection capability of this invention for targets with extremely low SNR.

[0100] The above specific embodiments are merely explanations of the present invention and are not intended to limit the present invention. After reading this specification, those skilled in the art can make modifications to these embodiments without contributing any inventive step, but as long as they are within the scope of the claims of the present invention, they are protected by patent law.

Claims

1. A method for detecting faint infrared targets based on a spatiotemporal feature enhancement network, characterized in that, Includes the following steps: S1: Obtain the infrared image sequence dataset; S2: Construct TSF-Net, an infrared target detection network with spatiotemporal feature enhancement for low signal-to-noise ratio targets; The SF-Net network consists of an input terminal, a temporal target enhancement module (TEM), a spatiotemporal feature extraction module, a feature fusion module, and a detection head; S3: Train the TSF-Net model using the training dataset constructed in step S1; S4: Input the test dataset constructed in step S1 into the TSF-Net trained in step S3 to test the detection performance of TSF-Net. S5: Evaluate the performance of the TSF-Net network for detecting low signal-to-noise ratio, faint infrared targets; In step S2, the constructed spatiotemporally enhanced infrared weak target detection network TSF-Net includes the following steps: S201: The input end performs time-series reading of the infrared image sequence, reading in the current frame and its adjacent N frames before and after to ensure the temporal sequence of the input image; sets the detection anchor box for weak targets and preprocesses the input image; S202: The temporal target enhancement module TEM consists of a three-dimensional convolutional layer, a three-dimensional max pooling layer, and a three-dimensional average pooling layer, which enhances the target using the input image sequence; S203: The spatiotemporal feature extraction module consists of a global context module GCBlock and residual units, and is divided into five stages. It fully extracts spatiotemporal features from the temporally enhanced feature layer and outputs feature layers of different sizes. S204: The feature fusion module consists of a multi-scale shallow feature enhancement module and a channel attention mechanism. It enhances the receptive field of shallow features using dilated convolutions of different scales, and adds or subtracts a simple channel attention mechanism to restrict spatial details for deep features. Then, the resulting deep feature maps and shallow feature maps are added and fused together. S205: The detection head has four branches, which detect infrared weak targets of different scales respectively. Each branch is composed of a 1×1 convolution. The feature layers of different scales obtained in step S204 are used to predict the regression parameters of the target, the predicted target box and the confidence of the target through the 1×1 convolution layer. S206: Based on the preset anchor box set in step S201, adjust the position of the predicted box to obtain a target candidate region that is closer to the true box. The calculation formula is as follows: In the formula , , , , represents the predicted offset. , The length and width of the preset anchor frame, , This represents the coordinates of the top-left corner of the grid in the feature layer; S207: Map the target candidate regions obtained in step S206 to feature layers of different scales of the detection head in step S205, and calculate the loss between the candidate target boxes and the prediction results; the loss consists of two parts: target box loss and confidence loss. The target box loss is calculated by CIOU loss, and the confidence loss is calculated by binary cross-entropy loss. The loss calculation is as follows: in This represents the distance between the centroid of the predicted bounding box and the ground truth bounding box. It is the diagonal distance of the minimum bounding rectangle. Used as a correction factor to further adjust the loss function; in addition, , , , Let A represent the width and length of the ground truth box and the predicted box, respectively. Let IOU be the intersection-union ratio of the predicted box and the ground truth box. Let A be the area of ​​the ground truth box and B be the area of ​​the predicted box. Weight factors a and b are used to weight the confidence loss and the localization loss to avoid imbalance between the two parts of the loss during training.

2. The infrared weak target detection method based on spatiotemporal feature enhancement network according to claim 1, characterized in that, In step S1, obtaining the infrared faint target image sequence dataset includes the following steps: S101: Obtain training and test samples; construct M infrared weak target image sequences using public datasets and real-world datasets, each image with a resolution of R×R and 3 channels, and label the infrared weak targets on each frame image. Then, use H infrared weak target image sequences and their corresponding labels as training set X, and use MH infrared weak target image sequences and their corresponding labels as validation set V. S102: The infrared image sequence is sampled using a sliding window sampling method, selecting the current N frames and the previous N frames to ensure the temporal continuity of the input TSF-Net image.

3. The method for detecting faint targets based on a spatiotemporal feature enhancement network according to claim 1, characterized in that, In step S3, training the TSF-Net model includes the following steps: S301: Train the TSF-Net model constructed in step S2. The training environment is NVIDIA GeForce RTX 3090 GPU, the network training is Epoch=300, the initial learning rate Lr=0.0001, the learning rate optimizer is adamw, the batch size=16, and training starts from 0. S302: Retain the optimal weights obtained from training in step S301 for model detection and evaluation.

4. The infrared weak target detection method based on spatiotemporal feature enhancement network according to claim 1, characterized in that, In step S4, testing the detection performance of TSF-Net includes the following steps: S401: Input the test dataset constructed in step S1 into the optimal weights obtained after training in step S3, predict the data, and obtain the target bounding box and confidence score of the target prediction result; S402: In the prediction results, results with a confidence level greater than the set threshold T1 are judged as correct results, and the corresponding detection box is the predicted target location; S403: Perform NMS suppression on the prediction results based on the NMS threshold T2 to remove duplicate prediction boxes and obtain the final prediction results.

5. The infrared weak target detection method based on spatiotemporal feature enhancement network according to claim 1, characterized in that, In step S5, evaluating the TSF-Net network's performance in detecting low signal-to-noise ratio, dimly lit targets includes the following steps: S501: Accuracy is used to evaluate the precision of the network, calculated using the following formula: S502: Recall is used to evaluate the network's ability to detect all targets. The calculation formula is as follows: S503: The F1 score is used to evaluate the overall detection capability of the network. The calculation formula is as follows: Where TP represents the number of positive targets predicted as positive, FP represents the number of negative targets incorrectly classified as positive, and FN represents the number of positive targets incorrectly classified as negative.

Citation Information

Patent Citations

  • Space-based infrared dim small moving target detection method

    CN114373130A

  • Infrared target detection method applied to complex environment

    CN116229217A