Infrared small target tracking method under complex conditions
Through the TCS segmentation network model, the problem of fast and accurate tracking of airborne infrared small targets in complex backgrounds is solved, and the efficient tracking effect under complex conditions is achieved, which improves the tracking success rate and stability of infrared small targets.
Patent Information
- Application Number
- CN202510459947.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional infrared small-target tracking methods are difficult to achieve fast and accurate target recognition and tracking in complex backgrounds, especially on airborne platforms, affected by complex backgrounds, fast target movement and lens changes, resulting in reduced tracking accuracy or lost targets.
The TCS segmentation network model is adopted, including a multi-directional attention module, a target feature condition module and a multi-scale receptive field inference module. Combined with a full convolutional network, the feature discrimination and robustness are enhanced through the cross attention mechanism and noise perturbation strategy, and the trained TCS segmentation network model is used for small-object tracking.
On the basis of ensuring real-time performance, the tracking success rate of infrared small targets under complex backgrounds has been greatly improved, and the accuracy and stability of tracking have been improved.
Smart Images

Figure CN120298455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to infrared small target tracking, and particularly to a method for tracking infrared small targets under complex conditions. Background Art
[0002] Infrared small target tracking technology plays a crucial role in many fields. In the military field, it is widely used in military reconnaissance, missile guidance, etc., and can help military personnel accurately identify and track potential targets at long distances, providing important support for military decision-making; in the civilian field, it is also indispensable in scenarios such as security monitoring and industrial inspection. For example, in security monitoring, it can achieve intelligent monitoring of specific areas and timely detect abnormal targets.
[0003] In the data of the airborne infrared small target tracking scenario, a series of complex characteristics are presented. First of all, "complex background" is one of the significant features. The environment where the airborne platform is located is complex and diverse. From vast mountains, oceans to various backgrounds such as urban buildings, all may serve as the background of the target. These backgrounds not only have rich textures and diverse structures, but may also contain various interference sources, such as clouds, atmospheric turbulence, etc., which greatly increases the difficulty of accurately identifying and tracking small targets from the background.
[0004] Secondly, "fast moving target" brings huge challenges to tracking. Due to the movement of the aircraft itself and the rapid movement of the target itself, the position of the infrared small target in the image changes rapidly. Traditional tracking algorithms often have difficulty quickly and accurately capturing the dynamic changes of the target, resulting in a decline in tracking accuracy or even loss of the target.
[0005] Furthermore, "changing lens" further exacerbates the complexity of tracking. During the flight of the aircraft, it is affected by factors such as air flow and attitude adjustment, making the lens inevitably jitter, rotate and zoom. These changes in the lens will cause scale changes, position offsets and perspective changes of the target in the image, which poses extremely high requirements for the stability and adaptability of the tracking algorithm.
[0006] Facing these complex and challenging problems in airborne infrared small target tracking, traditional tracking methods are difficult to meet the actual needs. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a method for tracking infrared small targets under complex conditions, which can greatly improve the success rate of tracking infrared small targets under complex backgrounds on the basis of ensuring real-time tracking.
[0008] The purpose of the present invention is achieved by the following technical solutions: A method for tracking infrared small targets under complex conditions, characterized by comprising the following steps:
[0009] Construct a TCS segmentation network model, including a multi-directional attention module, a target feature conditioning module, a multi-scale receptive field inference module, and a fully convolutional network;
[0010] The multi-directional attention module consists of a set of 8 convolutional kernels with different directions and fixed parameters. Its function is to model the direction information in the image and is used to process two input images, generating feature maps of the two input images respectively;
[0011] The two images include the image to be segmented and the reference image, also known as the frame to be segmented and the reference frame.
[0012] The reference image is obtained by calculating the foreground mask from the first frame of the sequence to be tracked and has the same resolution as the image to be segmented.
[0013] The target feature conditioning module is used to fuse the feature maps of the two input images to obtain a fused feature map;
[0014] The target feature conditioning module adopts a cross-attention mechanism to achieve the fusion of feature layers:
[0015] The cross-attention mechanism dynamically adjusts the feature weights by calculating the interaction relationship between the query Qref of the reference frame feature map and the key Kcur and value Vcur of the current frame to be segmented feature map, thereby enhancing the discriminability of the features:
[0016] Among them, Qref represents the query Q of the reference frame feature map, and Kcur and Vcur represent the key K and value V of the frame to be segmented feature map respectively;
[0017] First, calculate the similarity score between the query Qref and the key Kcur. This score is achieved through a dot product operation; normalize the score through the Softmax function to obtain the attention weight; multiply the attention weight by Vcur and sum to obtain the fused feature representation, denoted as:
[0018]
[0019] The multi-scale receptive field inference module is used to realize the inference of different scale information, and the feature map output by the multi-scale receptive field inference module generates the final segmentation result through a fully convolutional network. The multi-scale receptive field inference module adopts the MRFFI module in the MobileMamba network; MobileMamba proposes a lightweight multi-receptive field vision Mamba network. Through a three-stage network design and the MRFFI (Multi-Receptive Field Feature Interaction) module, while improving the model inference speed, it achieves higher accuracy, surpassing the CNN, ViT, and Mamba structures.
[0020] Construct a sample set for training the TCS segmentation network model, and use the sample set to train the TCST segmentation network model to obtain a trained TCS segmentation network model;
[0021] The samples in the sample set for training the TCS segmentation network model are:
[0022] Each frame image in the continuous image sequence; the label of each sample is the preset expected segmentation result of the sample.
[0023] When using the sample set to train the TCS segmentation network model, the loss function adopted is SoftIoULoss.
[0024] Use the trained TCS segmentation network model to perform small target tracking on the sequence to be tracked, including
[0025] A1. Input the first frame image of the sequence to be tracked and the initial ground truth point coordinates of the target;
[0026] A2. Use the truth point coordinates as the initialized center point center, expand k preset pixels outward based on this coordinate to form a foreground mask mask of the target; subsequently, perform mask processing on the entire image, and the generated mask image is used as a reference image;
[0027] A3. For each frame in the sequence, input the current image frame and the reference image frame into the TCS model at the same time; the model outputs multiple segmentation results, and each segmentation result is a potential target region mask;
[0028] Suppose the number of potential target region masks is n. For the potential n target region masks, calculate the centroid center_pred respectively, and select the point closest to the initial center as the new center point;
[0029] A4. Based on the new center point, perform convolution operations through the Siamese network SiamFC to optimize the coordinates of the center point; SiamFC convolution determines the optimal position of the target in the current frame by calculating the response map; if the global response value is low in SiamFC, it is determined as a false alarm and the center point is discarded.
[0030] The beneficial effect of the present invention is that the present invention can greatly improve the success rate of infrared small target tracking in complex backgrounds on the basis of ensuring real-time tracking. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flowchart of the method of the present invention;
[0032] Figure 2 is a diagram of the TCS segmentation network architecture;
[0033] Figure 3 It is a flowchart for tracking pipeline inference. Specific implementation manners
[0034] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the following.
[0035] First, the application will introduce the overall structure of the TCST framework in detail, including its core components and their interaction mechanisms. The TCST framework obtains a segmentation model TCS (Target-Conditioned Segmentation) that is more suitable for the tracking task by fusing the Target Feature Condition Module (TFCM). The TFCM module is a core component of the TCST framework. It extracts the feature information of the target in the first frame and embeds it into the segmentation model to achieve dynamic tracking of the target. The design of this module not only considers the spatio-temporal consistency of the target features, but also enhances the robustness of the model to the reference frame through a noise perturbation strategy. This design shows significant superiority in complex background and interference scenarios, effectively improving the accuracy and stability of tracking.
[0036] Secondly, the application will introduce the tracking pipeline process of the TCST framework in detail. From the initial frame target localization to feature extraction, and then to the dynamic update during the tracking process. By introducing and referring to the Hungarian matching algorithm, precise matching of multiple targets output by the TCS is achieved. The TCST framework realizes efficient tracking of infrared small targets through a series of carefully designed steps.
[0037] As Figure 1 shown, a method for tracking infrared small targets under complex conditions includes the following steps:
[0038] Construct a TCS segmentation network model, including a multi-directional attention module, a target feature condition module, a multi-scale receptive field inference module, and a fully convolutional network;
[0039] In the field of target tracking, traditional tracking methods based on detection / segmentation results usually adopt the method of single-frame independent detection. This method fails to fully utilize the feature information of the target in consecutive frames, resulting in limited tracking accuracy. In addition, although some methods extract the features of the previous target during the tracking inference process to guide subsequent tracking, this feature extraction method often lacks adaptability to the dynamic changes of the target and is difficult to effectively meet the target representation requirements in complex scenarios.
[0040] To address the above problems, the TCST (Target-Conditioned Segmentation Tracking) framework proposed in this application introduces a Target Feature Conditioning Module (TFCM) during the training and inference stages of the segmentation model. By dynamically embedding the target features of the first frame into the network structure, this module enables the model to learn the correlation between the current frame and the reference frame, thereby enhancing the model's target representation ability under complex conditions. Specifically, the TFCM module can effectively capture the spatio-temporal consistency features of the target in consecutive frames and incorporate them into the decision-making process of the segmentation model, providing more accurate feature guidance for the tracking task.
[0041] In addition, to further improve the robustness of the tracking model in complex scenarios, the TCS framework specially designs a noise perturbation strategy during the training process. By simulating the noise interference that the reference frame may be subjected to in practical applications, this strategy enhances the model's adaptability to noise, thereby improving the robustness of the model to the reference frame features during the training stage. In this way, the TCST framework can not only effectively utilize the target feature information but also significantly improve the tracking effect in complex scenarios, providing a more efficient and robust solution for the target tracking task in computer vision.
[0042] As Figure 2 shown, the TCS segmentation network receives two inputs: the target image to be segmented and the reference image (ref_img). The reference image is obtained by calculating the foreground mask from the first frame of the sequence to be tracked and has the same resolution as the target image. Drawing on the design concept of the RDIAN network framework, the two input images are first fed into a Multi-directional Attention Module (MDA) with shared parameters. This module consists of a set of 8 convolutional kernels with different fixed directions, and its main function is to model the direction information in the image. After being processed by the MDA module, the two input images respectively generate feature maps with 16 channels and a size of 256×256.
[0043] Subsequently, these two feature maps are input into the Target Feature Conditioning Module (TFCM) for feature fusion. The fused feature map is further fed into the Multi-scale Receptive Field Estimation Module (MRFE) of the RDIAN network to achieve reasoning about information at different scales. Finally, the fused feature map passes through a Fully Convolutional Network (FCN) module to generate the final segmentation result.
[0044] In the TCST framework, the Target Feature Conditioning Module (TFCM) adopts a Cross-Attention Mechanism to achieve the fusion of feature layers. Specifically, the Cross-Attention Mechanism dynamically adjusts the feature weights by calculating the interaction relationship between the reference frame feature (denoted as Qref) and the current frame feature (denoted as Kcur, Vcur), thereby enhancing the discriminability of the features.
[0045] Given the reference frame feature Qref and the current frame features Kcur, Vcur, first calculate the similarity score between the Query (Q) and the Key (K), usually achieved through a dot product operation. Subsequently, normalize the score through the Softmax function to obtain the attention weights. Finally, multiply the attention weights by the Value (V) and sum them to obtain the fused feature representation. This process can be formalized as:
[0046]
[0047] where d is the dimension of the Query Q, Key K, and Value V, which is used to limit the scale of the dot product within a reasonable range to ensure the stability and reasonable distribution of the Softmax input.
[0048] Construct a sample set for training the TCS segmentation network model, and use the sample set to train the TCST segmentation network model to obtain a trained TCS segmentation network model;
[0049] The samples in the sample set for training the TCS segmentation network model are:
[0050] Each frame image in the continuous image sequence; the label of each sample is the expected segmentation result of the sample set in advance.
[0051] During the training process of the TCS segmentation network, the SoftIoU Loss is adopted as the loss function. The SoftIoU Loss is a continuously differentiable loss function based on the Intersection over Union (IoU), which can directly optimize the IoU metric in the segmentation task. Here, P is the predicted mask, G is the GroundTruth mask, and ε = 1 is the smoothing term. This loss alleviates the small-object pixel-level imbalance problem by directly optimizing the coverage of the predicted region and the ground-truth region.
[0052]
[0053] The trained TCS segmentation network model is used to track small objects in the sequence to be tracked.
[0054] As Figure 3 shown in the inference pipeline of the TCST model, the entire tracking process is designed as a systematic multi-step process aimed at efficiently achieving continuous tracking of the target. The specific steps are as follows:
[0055] Input: The input of the inference pipeline includes the first-frame image of the sequence to be tracked and the initial ground truth (GT) point coordinates of the target.
[0056] Initialization (Step 1): Using the GT point coordinates as the initialized center point, k preset pixels are extended outward based on this coordinate to form the foreground mask of the target. Subsequently, the entire image is masked, and the generated masked image is used as the reference image (ref_img) for tracking subsequent frames.
[0057] Object Detection and Feature Fusion (Step 2): For each frame in the sequence, the current image frame and the reference image frame are simultaneously input into the TCS model. The model outputs n potential target region masks through its internal feature extraction and fusion module (such as the TFCM module). For each connected region, its centroid (center_pred) is calculated, and the point closest to the initial center is selected as the new center point.
[0058] Center Point Optimization (Step 3): Based on the updated center value, convolution operations are performed through the Siamese network (SiamFC) to further optimize the coordinates of the center point. The SiamFC convolution determines the optimal position of the target in the current frame by calculating the response map. If the global response value is low, it is determined as a false alarm and the center point is discarded.
[0059] In the embodiments of the present application, for the first frame, when tracking a target that is occluded, the model will lose the target, and such occluded targets often appear in scenarios such as urban buildings and forests. Facing this situation where the features completely disappear, this paper designs a tracking method integrating a detection module to handle this situation: using the current lost frame as the search original frame, detecting several frames through the detection module until the target reappears, and then proceeding with tracking.
[0060] In the embodiments of the present application, in a multi-object tracking scenario, the Hungarian Matching Algorithm can also be used to match different targets. This algorithm realizes the efficient update of target trajectories by minimizing the cost matrix between targets. The Hungarian Matching Algorithm is widely used in multi-object tracking because it can effectively handle problems such as changes in the number of targets and occlusions.
[0061] In the embodiments of the present application, the following experiments are carried out to analyze and verify the present application:
[0062] (I) Dataset
[0063] The datasets used include: 1. The dataset for detecting and tracking small aircraft targets in the ground-air background, hereinafter referred to as dataset-small. This dataset is for the application of detecting and tracking small aircraft targets flying at low altitudes. Through field shooting and data preparation and processing, it provides a set of algorithm test datasets with one or more fixed-wing UAV targets as the detection objects. The dataset acquisition scenarios cover backgrounds such as the sky and the ground, as well as various scenarios, including a total of 22 segments of data, 30 flight tracks, 16,177 frames of images, and 16,944 targets. Each target corresponds to a marked position, and each segment of data corresponds to a marked file. This dataset can provide basic data for research on detecting small targets, precision guidance, and infrared target characteristics, etc. Among the dataset, there are 17 for training and 5 for testing. The marked data only includes the number of targets and the target center coordinates of each frame, including a small amount of multi-target data, with the characteristics of real data and low-quality data with obvious camera movement. 2. The semi-simulation dataset for detecting small infrared moving targets in complex backgrounds, hereinafter referred to as dataset-big. Various conditions are configured during the image capture process of this dataset, including relative height (looking down, looking straight, and looking up), scenarios (vegetation, water area, and buildings), platform movement, weather, time, etc. At the same time, the movement of the imaging platform and the movement of the synthetic targets are considered to ensure that the semi-synthetic image sequence is as close as possible to the real application scenario. In addition, we vary the target characteristics during the synthesis process, including shape, intensity, and movement. This dataset includes 350 image sequences, 150,185 images, and a marked file that details the target positions in the images. This dataset can be used for research on detecting and tracking small infrared moving targets. Among the dataset, there are 175 for training and 175 for testing. The marked data includes the targets, center coordinates, and signal-to-noise ratio of each frame, including a small amount of multi-target data, with the characteristics of complex backgrounds, camera shake problems, and high task difficulty.
[0064] (2) Experimental settings
[0065] 1. Evaluation metrics
[0066] 1.1 Evaluation metrics for detection tasks
[0067] Since the two target datasets used only provide the center point coordinates, in the evaluation of the detection task, a square area with a 10-pixel expansion from the center point is used as the ground truth box. The specific evaluation metrics are as follows:
[0068] PixAcc (Pixel Accuracy): Calculate the proportion of correctly classified pixels to the total number of pixels, which is used to evaluate the classification accuracy of the model at the pixel level.
[0069] mIOU (Mean Intersection over Union): Calculates the average intersection over union of small target bounding boxes, used to measure the overlap degree between the predicted bounding box and the ground truth bounding box.
[0070] PD (Probability of Detection): Equivalent to recall rate, representing the detection rate of valid small targets, used to evaluate the target detection ability of the model.
[0071] FA (False Alarm): Represents the probability that the background is misjudged as a target, used to measure the anti-false detection ability of the model.
[0072] 1.2 Evaluation Metrics for Tracking Tasks
[0073] The main evaluation metric for the tracking task is the Success Rate (average tracking success rate). In this study, if the distance between the center point of the tracking result and the annotated center point is less than 5 pixels, then this tracking is considered a successful tracking.
[0074] 2. Experimental Configuration
[0075] 2.1 Hardware Configuration
[0076] The experiment uses the following hardware configuration:
[0077] GPU: 1×NVIDIA GeForce RTX 4090
[0078] CPU: 96×Intel / AMD CPU
[0079] 2.2 Software Environment
[0080] The software environment configuration of the experiment is as follows:
[0081] Operating System: Ubuntu 22.04 LTS
[0082] Deep Learning Framework: PyTorch 2.5.0
[0083] CUDA Version: CUDA 12.2, used for GPU accelerated computing
[0084] 3. Data Processing
[0085] 3.1 Data Standardization
[0086] To ensure data consistency between different datasets and different imaging devices and to promote the convergence of the model, the original data was standardized. The specific method is to use Z-score standardization, and calculate the mean and standard deviation of pixel values according to the distribution of the full dataset.
[0087] 3.2 Data Augmentation
[0088] During the training process, data augmentation is performed by means of random cropping, and image patches of size 256×256 are randomly cut out for training.
[0089] 3.3 Hyperparameter Settings
[0090] Batch Size: 128
[0091] Patch Size: 256×256
[0092] Optimizer: Adam
[0093] Learning Rate: 5×10 -4
[0094] 4. Training Strategy
[0095] 4.1 Construction of Training Data Pairs
[0096] The form of the training data pair is: {input image, reference image; predicted mask}.
[0097] 4.2 Method for Constructing Reference Images
[0098] v1: For a sequence containing N frames, starting from the first frame, for the current frame T, the following data are input simultaneously:
[0099] The full image (img) of frame T
[0100] The GT mask (Ground Truth Mask) of frame T
[0101] The target reference image (ref_img) processed based on the GT mask of frame T - 1
[0102] v1-aug: Considering that there may be errors in the detection results of the previous frame during the tracking process, perturbations are added to the ref_img of v1 during the training phase. The specific method is to input a completely black reference image with a probability of 50% to simulate the noise interference in actual tracking.
[0103] v2: The reference frame is fixed as the target reference image processed based on the GT mask of the first frame, and a completely black reference image is input with a probability of 50%.
[0104] 4.3 Data Input in the Inference Phase
[0105] In the inference phase, the input data includes:
[0106] The full image of the current frame T
[0107] The target reference image processed based on the GT mask
[0108] Output the predicted mask result map of the T frame
[0109] (3) Experimental results and analysis
[0110] 1. Overall effect
[0111] 1) Segmentation task
[0112] dataset_big
[0113]
[0114]
[0115] 2) Tracking task
[0116] dataset_big
[0117]
[0118] dataset_small
[0119]
[0120] 3. Ablation experiment
[0121] 1. Selection of backbone
[0122] When reproducing and selecting existing open-source infrared target segmentation algorithms, based on the characteristics of the tracking task, this study focused on two key performance indicators: PD (Probability of Detection) and inference speed. The PD indicator directly determines the success rate of the tracking task, while the inference speed affects the real-time performance of the tracking system. Based on this, we comprehensively evaluated various algorithms to screen out the most suitable algorithm for real-time tracking tasks.
[0123] However, it is worth noting that some relatively new algorithms based on the Transformer architecture (such as SCTransNet) have high computational complexity, resulting in a large model size and slow inference speed, making it difficult to meet the speed requirements of real-time tracking tasks. Therefore, these algorithms were not included in the scope of this comparative analysis.
[0124] Comparison of accuracy effects
[0125]
[0126]
[0127] Comparison of speed and number of parameters effects
[0128]
[0129] In the process of reproducing and selecting various existing open-source infrared target segmentation algorithms, we comprehensively considered two key factors: detection accuracy (PD metric) and model computational efficiency (inference speed). Through systematic comparative analysis, the RDIAN network was finally selected as the backbone of the TCST method due to its excellent performance in high-precision target segmentation tasks and lightweight model design. The RDIAN network not only meets the requirements of high accuracy for real-time tracking tasks in terms of detection accuracy, but its lightweight architecture also ensures efficient inference ability in practical applications, thus providing a solid foundation for the efficient implementation of the TCST method in infrared small target tracking tasks. This selection provides reliable performance guarantee for subsequent tracking tasks and lays a good starting point for further optimization in complex scenarios in the future.
[0130] In addition, conduct the effectiveness verification of the TFCM module
[0131] To further verify the effectiveness of the Target Feature Condition Module (TFCM) in the TCST framework, we designed a comparative experiment. By removing the TFCM module, the method was degraded to a tracking framework based on the segmentation results of the current frame (IGST: IRSTDGuidance Siamese-based Tracking). The inference pipeline of this framework is as follows:
[0132] Input: The first-frame image of the sequence to be tracked and the initial ground truth (GT) point coordinates of the target.
[0133] Initialization (Step 1): Use the GT point coordinates as the initialized center point (center), and expand k pixels outward based on this coordinate to form the foreground region of the target. Subsequently, use the feature extraction network to extract features from the expanded region, and the obtained feature representation is used as the query feature for target matching in subsequent frames.
[0134] Target detection and segmentation (Step 2): For each frame in the sequence, input the current frame into the IRSTD model for global search to obtain n potential target region masks. For each connected region, calculate its centroid (center_pred) and use it as the candidate target center point.
[0135] Center Point Optimization and False Alarm Suppression (Step 3): Based on the updated center value, expand outward within a preset range to form a search area. Use the query feature and the Siamese network (SiamFC) for convolution operations to calculate the response map, and select the position with the highest response as the coordinates of the optimized center point. If the global response value is low, indicating that no reliable target is detected in the current frame, it is determined as a false alarm and the center point is discarded.
[0136] Method Success_rate Inference speed (fps) IGST 89.14% 32 TCST v2 99.38% 32
[0137] Verify the training strategy
[0138] Method Success_rate Inference speed (fps) TCST v1 94.51% 32 TCST v1-aug 95.82% 32 TCST v2 99.38% 32
[0139] The augmented training strategy with noise perturbation can significantly improve the stability of target tracking. Specifically, this strategy increases the tracking success rate by 1.31 percentage points. Further, the scheme (v2) of replacing the reference frame with the first frame shows a more significant improvement effect, with the tracking success rate increasing by more than 3 percentage points. It is analyzed that this improvement is mainly attributed to the fact that in the tracking update process, the tracking result of the previous frame may itself be biased, leading to incorrect feature input and thus affecting the tracking results of subsequent frames. In contrast, fixedly using the first frame as the reference frame can provide more stable and accurate feature information, thereby enhancing the robustness of tracking. In addition, from the theoretical basis of target tracking, tracking algorithms usually rely on appearance modeling and motion information modeling. In terms of appearance modeling, accurate target feature extraction is crucial for tracking stability. By introducing the noise perturbation strategy, the model can better adapt to changes in target features and reduce tracking failures caused by feature drift. In terms of motion modeling, fixedly using the first frame as the reference frame can provide a more stable input for the motion model and avoid cumulative errors caused by incorrect features in the previous frame.
[0140] Verify the Hungarian matching effect
[0141]
[0142] In multi-object tracking tasks, the Hungarian Matching Algorithm is adopted to solve the problem of object association, significantly improving the success rate of tracking. Specifically, compared with the baseline method without using this algorithm, the tracking success rate has increased by 23 percentage points. This significant improvement is attributed to the fact that the Hungarian Matching Algorithm can effectively handle the complex interaction relationships between objects, especially in complex scenarios such as changes in the number of objects, occlusion, and object overlap. By minimizing the global cost function, the continuity and consistency of object trajectories are ensured. However, the introduction of the Hungarian Matching Algorithm also brings certain computational overhead. Since this algorithm needs to calculate the cost matrix between objects and solve the optimal matching through the Hungarian algorithm, this process has a relatively high computational complexity. Therefore, in practical applications, the inference performance of the model decreases slightly. Specifically, the inference speed is affected to a certain extent, especially when dealing with large-scale objects or high-resolution images, and the consumption of computing resources is more significant.
[0143] Verify the loss function
[0144] This part further compares the performance differences of various segmentation loss functions in small object tracking tasks. The experimental results show that after comprehensively considering the accuracy effect and training stability, SoftIoU Loss is the most suitable loss function for small object tracking tasks in the current complex background.
[0145] Accuracy improvement: SoftIoU Loss performs excellently in processing small object segmentation. Especially in the case of class imbalance, it can effectively improve the segmentation accuracy. This feature is particularly important for small object tracking in complex backgrounds because the feature information of small objects is relatively sparse and is easily interfered by background noise.
[0146] Training stability: Compared with the traditional cross-entropy loss, SoftIoU Loss shows more stable convergence characteristics during the training process. By optimizing the intersection over union (IoU), an index directly related to the segmentation performance, it avoids the common problems of gradient explosion or disappearance during training, thus improving the training stability of the model.
[0147] Comparison with other loss functions: In semantic segmentation tasks, SoftIoU Loss can handle the boundary problems of small objects more effectively compared to Dice Loss and Focal Loss. In addition, compared with Lovász-Softmax Loss, SoftIoU Loss has a lower computational complexity in small object tracking tasks while maintaining a high segmentation accuracy.
[0148]
[0149]
[0150] The above are the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in related fields. Any changes and modifications made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. An infrared small target tracking method under complex conditions, characterized in that: Including the following steps: Construct a TCS segmentation network model, including a multi-directional attention module, a target feature conditioning module, a multi-scale receptive field inference module, and a fully convolutional network; Construct a sample set for training the TCS segmentation network model, and use the sample set to train the TCS segmentation network model to obtain a trained TCS segmentation network model; Use the trained TCS segmentation network model to perform small target tracking on the sequence to be tracked.
2. The infrared small target tracking method under complex conditions according to claim 1, wherein: The multi-directional attention module consists of a set of 8 convolutional kernels with different directions and fixed parameters. Its function is to model the direction information in the image and is used to process two input images to generate feature maps of the two input images respectively; The two images include the image to be segmented and the reference image, also known as the frame to be segmented and the reference frame.
3. The infrared small target tracking method under complex conditions according to claim 2, wherein: The reference image is obtained by calculating the foreground mask from the first frame of the sequence to be tracked and has the same resolution as the image to be segmented.
4. A complex-condition infrared small target tracking method according to claim 1, characterized in that: The target feature conditioning module is used to fuse the feature maps of the two input images to obtain a fused feature map; The target feature conditioning module adopts a cross-attention mechanism to achieve the fusion of feature layers: The cross-attention mechanism dynamically adjusts the feature weights by calculating the interaction relationship between Qref of the reference frame feature map and Kcur, Vcur of the current frame to be segmented feature map, thereby enhancing the discriminability of the features: Among them, Qref represents the query Q of the reference frame feature map, and Kcur, Vcur represent the key K and value V of the frame to be segmented feature map respectively; First, calculate the similarity score between the query Qref and the key Kcur. This score is achieved through a dot product operation; normalize the score through the Softmax function to obtain the attention weight; multiply the attention weight by Vcur and sum to obtain the fused feature representation, denoted as: Among them, d is the dimension of the query Q, key K, and value V.
5. A complex-condition infrared small target tracking method according to claim 1, characterized in that: The multi-scale receptive field inference module is used to realize the inference of different scale information, and the feature map output by the multi-scale receptive field inference module generates the final segmentation result through a fully convolutional network.
6. The infrared small target tracking method under complex conditions according to claim 1, wherein: The samples in the sample set for training the TCS segmentation network model are: Each frame image in the continuous image sequence; the label of each sample is the expected segmentation result of the sample set in advance.
7. A method for tracking infrared small targets under complex conditions according to claim 1, characterized in that: When using the sample set to train the TCS segmentation network model, the loss function used is SoftIoULoss.
8. A method for tracking infrared small targets under complex conditions according to claim 1, characterized in that: The small target tracking of the sequence to be tracked using the trained TCS segmentation network model: A1. Input the first frame image of the sequence to be tracked and the initial ground truth point coordinates of the target; A2. Use the truth point coordinates as the initialized center point center, and expand k preset pixels outward based on this coordinate to form the foreground mask mask of the target; Subsequently, perform mask processing on the entire image, and the generated mask image is used as the reference image; A3. For each frame in the sequence, input the current image frame and the reference image frame into the TCS model at the same time; The model outputs multiple segmentation results, and each segmentation result is a potential target area mask; Let the number of potential target region masks be n. For the n potential target region masks, calculate the centroid center_pred respectively, and select the point closest to the initial center as the new center point; A4. Based on the new center point, perform convolution operations through the Siamese network SiamFC to optimize the coordinates of the center point.