An infrared dim and small target tracking method, system, device and medium

By using the Transformer feature fusion module and the resnet50 network in infrared weak target tracking for feature enhancement and fusion, the problem of poor infrared weak target detection and tracking performance in the prior art is solved, and higher detection accuracy and tracking stability are achieved.

CN118887257BActive Publication Date: 2025-05-30JIANGNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411075562.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-07
Publication Date
2025-05-30
Estimated Expiration
2044-08-07

AI Technical Summary

Technical Problem

The prior art has poor performance in infrared weak target detection and tracking, is susceptible to background interference, and lacks real-time performance.

Method used

An infrared weak target tracking method is adopted. By selecting the first frame of the image sequence as a template frame, the target area features are extracted, and the features are enhanced and fused with the features of subsequent frames. The Transformer feature fusion module and the resnet50 network are used for detection, and the target position and size are output.

Benefits of technology

It improves the detection accuracy and tracking stability of weak infrared targets, enhances the semantic understanding of small targets, improves the accuracy of recognition and positioning, and solves the problem of insufficient information capture at different distances and perspectives of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887257B_ABST
    Figure CN118887257B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, system, device and medium for tracking infrared dim and small targets. Among them, the method includes: selecting the first frame image in the infrared image sequence to be tracked as the template frame image, and performing feature extraction on the template frame image to obtain the feature map of the template frame image; sequentially loading the T-th (T>1) frame image of the infrared image sequence to be tracked as the detection frame image, and performing feature extraction to obtain the feature map of the detection frame image; performing feature enhancement and fusion on the feature map of the T-th detection frame image and the feature map of the template frame image, and then performing infrared dim and small target detection to output the position and size of the infrared dim and small target in the T-th detection frame image; judging whether the T-th detection frame image is the last frame image; if so, completing the tracking of the infrared dim and small target; if not, detecting the last frame image and completing the target tracking. The present invention can effectively track infrared dim and small targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and in particular to an infrared dim and small target tracking method, system, device and medium. Background Art

[0002] Infrared imaging technology has become increasingly mature, and application scenarios have been updated. With the development of artificial intelligence, the demand for infrared target detection and tracking has gradually increased. Most of the infrared dim and small target detection and tracking applications are in security monitoring, military guidance, and meteorological analysis, etc. Due to its weak energy, it is easily submerged in the clutter background, and the target edge is blurred, etc., making the difficulties faced by the detection and tracking of infrared dim and small targets urgent to be solved.

[0003] In the prior art, for example, Luca Bertinetto et al. from Harvard University disclosed a video target tracking method based on a fully convolutional Siamese network in the published literature "Fully-Convolutional Siamese Networks for Object Tracking". This method is based on the ImageNet2015 database and uses a pre-online learning method to train a neural network to solve the problem of generative similarity learning. This similarity matching function is simply evaluated during the tracking process. Then, a pre-trained deep convolutional network is used as a feature encoder to improve the tracking performance. However, in the process of detecting the target, due to the weak energy of the infrared dim and small target and the poor imaging situation, it is easily interfered by the similar color background, resulting in large errors in feature extraction of the target, and at the same time, it is easy to cause missed detection, and the real-time performance of detecting infrared dim and small targets is poor. Summary of the Invention

[0004] Therefore, the technical problem to be solved by the present invention is to overcome the problem of poor detection and tracking performance of infrared dim and small targets in the prior art.

[0005] To solve the above technical problem, the present invention provides an infrared dim and small target tracking method, including:

[0006] Step S1: Select the first frame image in the infrared image sequence to be tracked as the template frame image, extract the target image area to be tracked in the template frame image, and perform feature extraction on the target image area to be tracked to obtain the feature map of the template frame image;

[0007] Step S2: Sequentially load the T-th frame image except the first frame image in the infrared image sequence to be tracked as the detection frame image; perform feature extraction on the detection frame image to obtain the feature map of the detection frame image;

[0008] Step S3: Perform feature enhancement and fusion on the feature map of the T-th frame detection frame image and the feature map of the template frame image, and then perform infrared small target detection on the fused features to output the position and size of the infrared small target in the T-th frame detection frame image, where T > 1;

[0009] Step S4: Determine whether the T-th frame detection frame image is the last frame image of the infrared image sequence to be tracked; if so, complete the tracking of the infrared small target; if not, return to Step S2 until the last frame image is detected to complete the target tracking.

[0010] In an embodiment of the present invention, the method for obtaining the feature map of the template frame image in Step S1 is as follows:

[0011] Select the first frame image of the infrared image sequence to be tracked as the template frame image;

[0012] According to the label of the template frame image, calculate the center coordinates of the infrared small target, and use the prediction box formed by the center coordinates as the target image area to be tracked, and use the target image area to be tracked as the initial position of the tracking target; use the center of the prediction box as the cropping center, and perform template cropping on a square area twice the area of the prediction box to obtain the processed template frame image;

[0013] Input the processed template frame image into the infrared small target tracking model after side window filtering, and use the resnet50 network to extract features from the processed template frame image to obtain the feature map of the template frame image.

[0014] In an embodiment of the present invention, the method for obtaining the feature map of the detection frame image in Step S2 is as follows:

[0015] Extract the T-th frame image except the first frame image of the infrared image sequence to be tracked as the detection frame image, determine the detection area of the target in the T-th frame detection frame image according to the position information of the target in the previous frame image, use the center of the prediction box corresponding to the detection area as the cropping center, perform template cropping and image filling on a square area four times the area of the prediction box, and perform side window filtering to form a detection frame image of a fixed size;

[0016] Input the detection frame image of the fixed size into the infrared small target tracking model initialized by the template frame image, and use the resnet50 network to extract features from the detection frame image of the fixed size to obtain the feature map of the detection frame image.

[0017] In an embodiment of the present invention, inputting the feature map of the T-th frame detection frame image and the feature map of the template frame image into the Transformer feature fusion module for feature enhancement and fusion in Step S3 is specifically as follows:

[0018] Preprocessing: After the feature extraction of the detected frame image and the template frame image, the corresponding feature maps f z and f x are obtained. Use 1*1 convolution to reduce the channel dimensions of f z and f x to obtain f z1 and f x1 ;

[0019] Feature fusion: Take f z1 and f x1 as two branches and input them into the dual attention module in the Transformer feature fusion module respectively. The dual attention modules of the two branches adaptively focus on and extract the respective feature maps of f z1 and f x1 ; Then, the two CFA modules in the Transformer feature fusion module simultaneously receive the feature maps output by the dual attention modules of their respective branches and the dual attention modules of the other branch. The CFA module fuses the two feature maps through the built-in multi-head cross-attention mechanism;

[0020] The two dual attention modules and the two CFA modules form a fusion layer. The fusion layer is repeated N times, and then a CFA module is passed through to obtain the fusion result.

[0021] In an embodiment of the present invention, the dual attention module includes a multi-scale attention module and a self-attention module connected in sequence, where,

[0022] The multi-scale attention module satisfies:

[0023] X = L(C(g i ))

[0024] where X represents the output result of the multi-scale attention, C() represents the concat function, g i are concatenated together, and then feature aggregation is performed through the linear layer L(), and g i = F(Q i , K i , V i , r i ), F() is the multi-scale attention function, the dilation rate r i is the i-th head, Q, K, V are the query, key, and value of the multi-scale attention function respectively, Q i , K i , V i represent the fragments of the feature maps f z1 and f x1 ;

[0025] When using multi-scale attention operation, the feature map is processed using multiple heads, each with a different dilation rate. Let the dilation rate r take values of 1, 2, 3. Let x ij be f z1 and f x1 The pixel sets at position (i, j) in different dilation rates in form a feature block. Then x ij in X satisfies:

[0026]

[0027] where q ij represents the query at position (i, j) in f z1 and f x1 , K r and V r represent the key and value selected from K and V respectively. represents the key-value dimension, represents the selected key;

[0028] The self-attention module satisfies:

[0029] For the input X of the multi-scale attention module, feature extraction is performed through the self-attention module, and the output X E is expressed as:

[0030] X E = X + M(X + P x , X + P x , X)

[0031] where P x represents the spatial position encoding, M() represents the multi-head attention mechanism of the self-attention module and satisfies:

[0032] M(Q * , K * , V * ) = C(H i )W o

[0033]

[0034] where H i represents the attention of the i-th head, W o represents the weight matrix, represents the key-value dimension, represents the selected key, Q * , K * , V * represent the query, key, and value of the improved self-attention module respectively. C() represents the concat operation, and S() represents the softmax operation.

[0035] In one embodiment of the present invention, it further includes training the infrared small and weak target tracking model, and the model parameters are optimized using the moving exponential average (EMA) during the training process. Specifically:

[0036] Define the set Θ of model parameters, which includes all the training parameters in the Transformer feature fusion module and the ResNet50 network;

[0037] Initialize the EMA parameter set Θ EMA as a copy of Θ to ensure that at the beginning of model training, the EMA parameters in Θ EMA are the same as the initial parameters of the model;

[0038] In each iteration process, first update the parameters to obtain a new parameter set Θ t ;

[0039] Update the parameters in Θ EMA For each parameter θ belonging to Θ, its corresponding EMA parameter Θ EMA is updated according to a preset rule, and the formula is:

[0040]

[0041] where t represents the current iteration number, ln represents the natural logarithm calculation with base e, represents the parameter belonging to Θ in the t-th iteration EMA , represents the parameter belonging to Θ in the (t - 1)-th iteration EMA , θ t represents the parameter in the new parameter set Θ t , and α is a decay coefficient between 0 and 1, which is used to control the fusion degree of historical parameters and current parameters;

[0042] In different stages of training, by adjusting the value of α, the sensitivity of EMA is adjusted to adapt to different training dynamics;

[0043] During model validation or testing, Θ EMA is used as the model parameter for target tracking.

[0044] In one embodiment of the present invention, a data augmentation mechanism is introduced during the training process of the infrared small and weak target tracking model. The data augmentation mechanism perturbs the bounding boxes collected during the target tracking process, including:

[0045] Random perturbation of the bounding box: Randomly perturb the bounding box of a given target, and this random perturbation is used to simulate the possible position changes and scale changes of the target in consecutive frames;

[0046] Scale jitter: For each bounding box, multiply both its width and height by a random variable drawn from a log-normal distribution with a mean of 0, where the standard deviation of the log-normal distribution is controlled by a first parameter. This scale jitter is used to reduce the scale variation of the target due to distance changes in the field of view;

[0047] Center point jitter: Determining the maximum offset of the center jitter is achieved by calculating half of the jitter size and then multiplying it by a second parameter. Then, multiply the calculated maximum center point offset by a random vector to obtain the final center point position, where the random vector is uniformly distributed in the interval [0, 1] and subtract 0.5 from it to generate a value in the interval [-0.5, 0.5];

[0048] Combined jitter bounding box: Combine the results of randomly perturbing the jittered bounding box, scale jitter, and center point jitter to form a new jittered bounding box, which is used for target localization during the training process.

[0049] To solve the above technical problems, the present invention provides an infrared dim and small target tracking system, including:

[0050] First feature map acquisition module: Used to select the first frame image in the infrared image sequence to be tracked as the template frame image, extract the target image area to be tracked in the template frame image, perform feature extraction on the target image area to be tracked, and obtain the feature map of the template frame image;

[0051] Second feature map acquisition module: Used to sequentially load the T-th frame image in the infrared image sequence to be tracked except the first frame image as the detection frame image; perform feature extraction on the detection frame image to obtain the feature map of the detection frame image;

[0052] Fusion and detection module: Used to perform feature enhancement and fusion on the feature map of the T-th frame detection frame image and the feature map of the template frame image, and then perform infrared dim and small target detection on the fused features, and output the position and size of the infrared dim and small target in the T-th frame detection frame image, where T > 1;

[0053] Judgment module: Used to judge whether the T-th frame detection frame image is the last frame image of the infrared image sequence to be tracked; if so, complete the infrared dim and small target tracking; if not, return to the second feature acquisition module until the last frame image is detected to complete the target tracking.

[0054] To solve the above technical problems, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above infrared dim and small target tracking method are implemented.

[0055] To solve the above technical problems, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above infrared dim and small target tracking method are implemented.

[0056] The above technical solution of the present invention has the following advantages compared with the prior art:

[0057] The infrared dim and small target tracking model constructed by the present invention has a good target detection effect on infrared dim and small targets, and the generalization performance of the infrared dim and small target tracking model is good, so that it can maintain a stable tracking performance for target changes in practical applications;

[0058] The dual attention module constructed by the present invention incorporates a multi-scale attention module and a self-attention module, which can simultaneously integrate the detailed features and global information of the target, enhance the semantic understanding of small targets, improve the accuracy of recognition and positioning, and solve the problem of insufficient information capture by traditional single-scale feature extraction at different distances and perspectives;

[0059] To address the problem of parameter fluctuations in model training, in this embodiment, the moving exponential average (EMA) method is used to optimize the network parameters, and the update rule of the EMA parameters is optimized, so as to prevent the model from being too close to the latest training data and improve the prediction performance of the model for unseen data. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to specific embodiments of the present invention in conjunction with the accompanying drawings.

[0061] Figure 1 is the flowchart of the method of the present invention;

[0062] Figure 2 is the structure diagram of the infrared dim and small target tracking model in the embodiment of the present invention;

[0063] Figure 3 is the dilation rate schematic diagram of the multi-scale attention module in the embodiment of the present invention;

[0064] Figure 4 is the schematic diagram of the label and prediction box of the infrared dim and small target image sequence after tracking in the embodiment of the present invention;

[0065] Figure 5 is the comparison diagram of the infrared small target tracking results between the method of the present invention and other methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.

[0067] Embodiment 1

[0068] Referring to Figure 1 As shown, the present invention relates to a method for tracking infrared small and weak targets, including:

[0069] Step 1: Select the first frame image in the infrared image sequence to be tracked as the template frame image, obtain the position, width, and height of the current infrared small and weak target according to the label of the template frame image; preprocess the template frame image; perform edge-preserving denoising on the processed image using side window filtering (SWF), and input the filtered template frame image into the resnet50 network of the infrared small and weak target tracking model for feature extraction to obtain the feature map of the template frame image.

[0070] Step 2: Sequentially load the T-th frame image except the first frame in the infrared image sequence to be tracked as the detection frame image, where T is an integer greater than 1; preprocess the detection frame image, and perform edge-preserving denoising on the detection frame image through side window filtering, and input the detection frame image after side window filtering (SWF) into the infrared small and weak target tracking model, and extract the feature map of the detection frame image through the resnet50 network.

[0071] Step 3: Input the feature map of the T-th frame detection frame image and the feature map of the template frame image into the Transformer feature fusion module in the infrared small and weak target tracking model, and then use the box regression network to detect the infrared small and weak target; obtain the prediction box of the infrared small and weak target detected in the detection frame image, and determine the position and size of the infrared small and weak target in the current detection frame image to be output.

[0072] Step 4: Determine whether the current detection frame image is the last frame of the infrared small and weak target sequence; if the current detection frame image is the last frame, the tracking of the infrared small and weak target is completed; if not, return to execute Step 2 to Step 3 until the last frame image is detected and the tracking is completed.

[0073] Specifically, in step 1, the first frame image of the infrared image sequence to be tracked is selected as the template frame image. According to the label of the template frame image, the central coordinates of the small and weak infrared target are calculated, and the prediction box formed by the central coordinates is used as the image area of the target to be tracked. The image area of the target to be tracked is used as the initial position of the tracking target. Taking the center of the prediction box as the cropping center, a square area twice the area of the prediction box is cropped for the template to obtain the processed template frame image. The processed template frame image is input into the small and weak infrared target tracking model after side window filtering. The resnet50 network is used to extract features from the processed template frame image to obtain the feature map of the template frame image.

[0074] Specifically, in step 2, after the template frame image is processed, the T-th frame image of the infrared image sequence to be tracked except the first frame image is extracted as the detection frame image. According to the position information of the target in the previous frame image, the detection area of the target in the T-th frame detection frame image is determined. Taking the center of the prediction box corresponding to the detection area as the cropping center, a square area four times the area of the prediction box is cropped for the template and image filling (the image filling in this embodiment means that if the size of the cropped image is not square or there are missing edges, the image needs to be filled to maintain the square shape), and side window filtering is performed to form a detection frame image of a fixed size. The detection frame image of the fixed size is input into the small and weak infrared target tracking model initialized by the template frame image. The resnet50 network is used to extract features from the detection frame image of the fixed size to obtain the feature map of the detection frame image.

[0075] In the above process, the process of using side window filtering to process the small and weak infrared target image sequence specifically includes: selecting eight directions of the filtering window, namely up, down, left, right, northeast, southeast, northwest, and southwest for processing; among them, the first four directions place the pixel to be processed on a certain side of the filtering window, and the last four directions place the pixel to be processed at a certain corner of the filtering window; given pixel i and the filter kernel F( ), the result of side window filtering is:

[0076]

[0077] where q i represents the input image of the filtering window where pixel i is located, r represents the radius of the filtering window, ρ represents the azimuth of pixel i in the window, ρ has two values (ρ = 0, r), when ρ is 0, it means pixel i is on a certain side of the window, when ρ takes r, it means pixel i is at a certain corner of the window, θ represents the direction of the filtering window, θ has four values of θ = 0, π / 2, π, 3π / 2, and different combinations of ρ and θ constitute eight possible window directions;

[0078] During the filtering process, to preserve the edge information and minimize the gap between the input and the output, select the one that is the same as the input qi Output with the minimum Euclidean distance As the final filtering result:

[0079]

[0080] where argmin represents the function of minimization, denotes for any given value, |||| 2 denotes the Euclidean distance.

[0081] After the detection of side window filtering for the current frame image is completed, the filtering output corresponding to the direction of the final filtering result is selected as the input of side window filtering for the next frame image.

[0082] To deepen the effect of feature fusion, the feature map of the T-th frame detection frame image and the feature map of the template frame image are input into the Transformer feature fusion module for feature enhancement and fusion. Please refer to Figure 2 and the method of feature enhancement and fusion is as follows:

[0083] Preprocessing: After the detection frame image and the template frame image are subjected to feature extraction, their corresponding feature maps f z and f x are obtained. Use 1*1 convolution to reduce the channel dimensions of f z and f x to obtain f z1 and f x1 ;

[0084] Feature fusion: Take f z1 and f x1 as two branches and input them into the dual attention module in the Transformer feature fusion module respectively. The dual attention modules of the two branches adaptively focus on and extract the feature maps of f z1 and f x1 respectively; then, through the two CFA (cross-feature augment) modules in the Transformer feature fusion module, simultaneously receive the feature maps output by the dual attention modules of their respective branches and the dual attention modules of the other branch. The CFA module fuses the two feature maps through its built-in multi-head cross-attention mechanism;

[0085] The two dual attention modules and the two CFA modules form a fusion layer. The fusion layer is repeated N times, and then a fusion result is obtained through a CFA module.

[0086] It should be noted that the CFA module promotes the complementarity and fusion between feature maps through the multi-head cross-attention mechanism, thereby improving the overall expression ability of features.

[0087] The dual attention module in this implementation enables the tracking system to effectively capture the multi-scale representations of the target, integrate the detailed features and global information of the target simultaneously, enhance the semantic understanding of small targets, improve the accuracy of recognition and localization, and solve the problem of insufficient information capture in traditional single-scale feature extraction at different distances and perspectives.

[0088] In this embodiment, a deep feature fusion network is constructed by stacking two dual attention modules and two CFA modules multiple times. This network refines the feature representation in multiple iterations and maintains the coherence and consistency of the features during the fusion process. Finally, through the fused feature map output by the additional CFA module, the information of the detection frame image and the template frame image is integrated, providing high-precision feature information for infrared small target tracking.

[0089] The dual attention module in this embodiment includes a multi-scale attention module and a self-attention module connected in sequence, where,

[0090] (1) The multi-scale attention module satisfies:

[0091] X = L(C(g i ))

[0092] where X represents the output result of the multi-scale attention, C() represents the concat function, and g i are concatenated together and then feature aggregation is performed through the linear layer L(), and g i = F(Q i , K i , V i , r i ), F() is the multi-scale attention function, the dilation rate r i is the i-th head, Q, K, V are the query, key, and value of the multi-scale attention function respectively, and Q i , K i , V i represent the fragments of the feature maps f z1 and f x1 ;

[0093] Please refer to Figure 3 , when using the multi-scale attention operation, the feature map is processed with multiple heads, each head having a different dilation rate. Let the dilation rate r take values of 1, 2, 3. Among them, r = 1 corresponds to the central pixel and the 3*3 pixel edge (including the central pixel and the 8 pixels around the central pixel), r = 2 corresponds to 8 pixels among the central pixel and the 5*5 edge pixels (equivalent to the dilated convolution with a dilation rate of 2), r = 3 corresponds to 8 pixels among the central pixel and the 7*7 edge pixels (equivalent to the dilated convolution with a dilation rate of 3). Let x ij be f z1 and fx1 The pixel sets at the middle position (i, j) under different dilation rates form a feature block, then x in X ij satisfies (x ij represents the feature block at the position (i, j), and the feature block is a total of 9 pixels corresponding to different dilation rates):

[0094]

[0095] where q ij represents the query at the position (i, j) in f z1 and f x1 , K r and V r respectively represent the key and value selected from K and V, represents the key-value dimension, represents the selected key;

[0096] (II) The self-attention module satisfies:

[0097] For the input X of the multi-scale attention module, feature extraction is performed through the self-attention module, and the output X E is expressed as:

[0098] X E = X + M(X + P x , X + P x , X)

[0099] where P x represents the spatial position encoding, and M() represents the multi-head attention mechanism of the self-attention module and satisfies:

[0100] M(Q * , K * , V * ) = C(H i )W o

[0101]

[0102] where H i represents the attention of the i-th head, and W o represents the weight matrix, represents the key-value dimension, represents the selected key, Q * , K * , V * respectively represent the query, key, and value of the improved self-attention module, C() represents the concat operation, and S() represents the softmax operation.

[0103] Further, perform object prediction on the fused feature map: input the fused feature map into an MLP model (Multi-Layer Perceptron) for bounding box processing. This MLP model contains multiple hidden layers, and each hidden layer is used to extract abstract features at different levels to enhance the model's representation ability. In this method, the input features first undergo a series of weighted sum and activation operations through the hidden layers of the MLP model, and then are transformed into the bounding box coordinates of the object through the output layer. This output can directly represent the position of the target object, including but not limited to the center point coordinates, height, and width of the bounding box. This method utilizes the powerful function approximation ability of the MLP model to overcome the limitations of traditional regression methods and improve the accuracy and robustness of object localization. In addition to using the MLP model to determine the size and position of the prediction box in this embodiment, it also determines whether there is an object in the prediction box through, for example, the softmax function.

[0104] To address the problem of parameter fluctuations in model training, in this embodiment, the Exponential Moving Average (ema) method is used to optimize the network parameters, as follows:

[0105] (a) Define the set Θ of model parameters, which includes all training parameters in the resnet50 network and the Transformer feature fusion module;

[0106] (b) Initialize the EMA parameter set Θ EMA as a copy of Θ, ensuring that at the beginning of model training, the EMA parameters in Θ EMA are the same as the initial parameters of the model;

[0107] (c) In each iteration process, first perform parameter update to obtain a new parameter set Θ t ;

[0108] (d) Update the parameters in Θ EMA , for each parameter θ belonging to Θ, its corresponding EMA parameter Θ EMA is updated according to a preset rule, and the formula is:

[0109]

[0110] where t represents the current iteration number, ln represents the natural logarithm calculation with base e, represents the parameter belonging to Θ in the t-th iteration EMA , represents the parameter belonging to Θ in the (t - 1)-th iteration EMA , θ t represents the new parameter set Θ tAmong the parameters, α is a decay coefficient between 0 and 1, which is used to control the degree of fusion between historical parameters and current parameters. For example, α takes 0.9 or 0.99 to give greater weight to historical parameters and achieve smooth updates. The advantage of the EMA method lies in the smoothing effect on parameter updates, reducing spikes and noise interference during training. In addition, θ EMA tends to accumulate more information over multiple training steps, thus avoiding the model being too close to the most recent training data and improving the model's prediction performance for unseen data;

[0111] (e) During different stages of training, by adjusting the value of α, the sensitivity of EMA is adjusted to adapt to different training dynamics;

[0112] (f) During model validation or testing, use Θ EMA as the model's parameters for target tracking.

[0113] To address the problem of large fluctuations in infrared small target tracking, a bounding box processing method based on random perturbation is used in the method, as follows:

[0114] During the training process of the infrared small target tracking model, a data augmentation mechanism is introduced. The data augmentation mechanism perturbs the bounding boxes collected during the target tracking process, including:

[0115] Bounding box random perturbation: Randomly perturb the bounding box of a given target, which is used to simulate the possible position and scale changes of the target in consecutive frames;

[0116] Scale jitter: For each bounding box, multiply its width and height by a random variable drawn from a log-normal distribution with a mean of 0. The standard deviation of the log-normal distribution is controlled by the scale_jitter_factor parameter (the first parameter), and this scale jitter is used to reduce the scale variation of the target due to distance changes in the field of view;

[0117] Center point jitter: Determine the maximum offset of the center jitter by calculating half of the jitter size and then multiplying it by a center_jitter_factor parameter (the second parameter). Then multiply the calculated maximum center point offset by a random vector to obtain the final center point position. Among them, the random vector is uniformly distributed in the interval [0,1] and subtract 0.5 to produce a value in the interval [-0.5,0.5];

[0118] Combine the jittered bounding box, random perturbation of the bounding box, scale jitter, and center point jitter results to form a new jittered bounding box, and the new bounding box is used for target localization during the training process.

[0119] Finally, output the position coordinates and size of the target in the current frame image. Determine whether the detected frame image of the current detection is the last frame image of the infrared dim and small target image sequence. If it is the last frame image, complete the tracking of the infrared dim and small target in this image sequence; if not, load the (T + 1)-th frame image of the infrared dim and small target image sequence, return to execute the above steps 2 to 3 until the processing of the last frame image of this image sequence is completed. And after the current detected frame image is detected and the position coordinates and size of the infrared dim and small target are output, take this frame image as the template frame image, and input it together with the next frame image into the infrared dim and small target tracking model for detection until the detection of the input infrared dim and small target image sequence is completed.

[0120] Refer to Figure 4 As shown, the infrared dim and small target image sequence is processed by the method of the present invention to obtain the prediction result of the infrared dim and small target. Figure 4 In [the figure], number 1 is the target position predicted by the infrared dim and small target tracking model, and number 2 is the label obtained from the template frame image. Compare the prediction result of the present invention with the label preset in the template frame image. The prediction box output by the present invention contains the infrared dim and small target to be tracked, and has a large coincidence range with the label, and the prediction effect is good.

[0121] This embodiment uses precision, frames per second, and success rate to evaluate the performance. Precision can be described as the proportion of frames in which the center location error (CLE) between the prediction box and the ground truth is lower than a certain threshold. Frames per second, that is, the number of frames per second, describes how many images the video can display in one second. The success rate, also known as the overlap rate, represents the ratio of the number of frames in which the intersection / union between the predicted bounding box and the ground truth is above a certain threshold to the total number of frames. In addition, the area under the curve (AUC) of the success curve represents the area under the performance curve with respect to the horizontal axis.

[0122] This embodiment is compared with other related algorithms on the NIR dataset where each video contains 300 frames, and it is found that the algorithm of this embodiment is superior to other tracking algorithms in terms of precision and overlap rate, and is also remarkable in terms of success rate and accuracy.

[0123] In addition, Figure 5Shows the performance comparison of all NIR videos, including accuracy, the AUC of the success rate curve, and frames per second (FPS). It can be seen that STRCD has a higher tracking accuracy for small targets, and the AUC is the second highest (0.612). In addition, the FPS of STRCF is as high as 145, but this model performs poorly when tracking small targets. The main reason is that the model fails to handle well the parts with strong edges. In addition, due to SWF, the baseline algorithm has a higher FPS than the algorithm in this article. However, the method in this embodiment performs excellently in terms of overall accuracy (0.881) and AUC (0.660), and the stability of the tracking box is better.

[0124] It is not difficult to find that the tracking performance of the present invention is good.

[0125] Embodiment 2

[0126] This embodiment provides an infrared small and weak target tracking system, including:

[0127] The first feature map acquisition module: used to select the first frame image in the infrared image sequence to be tracked as the template frame image, extract the target image area to be tracked in the template frame image, perform feature extraction on the target image area to be tracked, and obtain the feature map of the template frame image;

[0128] The second feature map acquisition module: used to sequentially load the T-th frame image except the first frame image in the infrared image sequence to be tracked as the detection frame image; perform feature extraction on the detection frame image to obtain the feature map of the detection frame image;

[0129] The fusion and detection module: used to perform feature enhancement and fusion on the feature map of the T-th frame detection frame image and the feature map of the template frame image, and then perform infrared small and weak target detection on the fused features, and output the position and size of the infrared small and weak target in the T-th frame detection frame image, where T>1;

[0130] The judgment module: used to judge whether the T-th frame detection frame image is the last frame image of the infrared image sequence to be tracked; if so, complete the infrared small and weak target tracking; if not, return to the second feature acquisition module until the last frame image is detected to complete the target tracking.

[0131] Embodiment 3

[0132] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the infrared small and weak target tracking method described in Embodiment 1 are implemented.

[0133] Embodiment 4

[0134] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the infrared small and weak target tracking method described in Embodiment 1 are implemented.

[0135] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0136] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0137] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0138] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0139] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0140] Obviously, the above embodiments are merely examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. And the obvious changes or variations derived therefrom still fall within the protection scope of the present invention.

Claims

1. A method for tracking small infrared targets, characterized in that: include: Step S1: selecting the first frame image in the infrared image sequence to be tracked as a template frame image, extracting the target image region to be tracked of the template frame image, performing feature extraction on the target image region to be tracked, and obtaining a feature map of the template frame image; Step S2: sequentially loading the T-th frame images except the first frame image in the infrared image sequence to be tracked as detection frame images; performing feature extraction on the detection frame images to obtain a feature map of the detection frame images; Step S3: Enhance and fuse the feature map of the T-th detection frame image with the feature map of the template frame image, perform infrared weak small target detection on the fused features, and output the position and size of the infrared weak small target in the T-th detection frame image, where T>1; In step S3, the feature map of the T-th detection frame image and the feature map of the template frame image are input into the Transformer feature fusion module for feature enhancement and fusion, specifically: Preprocessing: After feature extraction of the detection frame image and the template frame image, their corresponding feature maps f are obtained. z and f x , use 1*1 convolution to reduce f z and f x The channel dimension of f z1 and f x1 ; Feature fusion: f z1 and f x1 As two branches, they are input into the dual attention module in the Transformer feature fusion module respectively. The dual attention module of the two branches adaptively focuses on and extracts f z1 and f x1 The two CFA modules in the Transformer feature fusion module simultaneously receive the feature maps output by the dual attention modules of their respective branches and the dual attention modules of another branch. The CFA module fuses the two feature maps through its own multi-head cross attention mechanism. Two dual attention modules and two CFA modules form a fusion layer, which is repeated N times and then passes through a CFA module to obtain the fusion result; The dual attention module includes a multi-scale attention module and a self-attention module connected in sequence, wherein: The multi-scale attention module satisfies: X=L(C(g i )) Among them, X represents the multi-scale attention output result, C() represents the concat function, and g i are connected together, and then feature aggregation is performed through the linear layer L() and g i =F(Q i ,K i ,V i ,r i ), F() is the multi-scale attention function, and the expansion rate r i is the i-th head, Q, K, and V are the query, key, and value of the multi-scale attention function respectively, Q i ,K i ,V i Represents the feature map f corresponding to the i-th head z1 and f x1 fragments; When using multi-scale attention operation, the feature map is processed using multiple heads, each with a different expansion rate. Let the expansion rate r be 1, 2, 3, and let x ij f z1 and f x1 The pixel set at position (i, j) in X under different expansion rates constitutes a feature block, then x in X ij satisfy: Among them, q ij represents f z1 and f x1 The query at position (i,j) in K r and V r Represents the key and value selected from K and V respectively, Represents the key-value dimension, A key indicating selection; The self-attention module satisfies: For the input X of the multi-scale attention module, feature extraction is performed through the self-attention module, and the output X E It is expressed as: X E =X+M(X+P x ,X+P x ,X) Among them, P x represents the spatial position encoding, M() represents the multi-head attention mechanism of the self-attention module and satisfies: M(Q * ,K * ,V * )=C(H i )W o Among them, H i represents the attention of the i-th head, W o represents the weight matrix, Represents the key-value dimension, The key indicating selection, Q * ,K * ,V * They represent the query, key, and value of the improved self-attention module respectively, C() represents the concat operation, and S() represents the softmax operation; Step S4: determine whether the Tth detected frame image is the last frame image of the infrared image sequence to be tracked; if so, complete the infrared weak target tracking; if not, return to step S2 until the last frame image is detected and the target tracking is completed.

2. The infrared small target tracking method according to claim 1, characterized in that: The method for obtaining the feature map of the template frame image in step S1 is: Selecting the first frame image of the infrared image sequence to be tracked as a template frame image; According to the label of the template frame image, the center coordinates of the infrared weak target are calculated, and the prediction frame formed by the center coordinates is used as the image area of ​​the target to be tracked, and the image area of ​​the target to be tracked is used as the initial position of the tracking target; Taking the center of the prediction box as the cropping center, a square area twice the area of ​​the prediction box is cropped to obtain a processed template frame image; The processed template frame image is filtered by the side window and then input into the infrared weak small target tracking model. The resnet50 network is used to extract the features of the processed template frame image to obtain the feature map of the template frame image.

3. The infrared small target tracking method according to claim 2, characterized in that: The method for obtaining the feature map of the detection frame image in step S2 is: Extract the Tth frame image except the first frame image of the infrared image sequence to be tracked as the detection frame image, determine the detection area of ​​the target in the Tth detection frame image according to the position information of the target in the previous frame image, take the center of the prediction box corresponding to the detection area as the cropping center, perform template cropping and image filling on a square area four times the area of ​​the prediction box, and perform side window filtering to form a detection frame image of a fixed size; The fixed-size detection frame image is input into the infrared dim small target tracking model initialized by the template frame image, and the resnet50 network is used to extract the features of the fixed-size detection frame image to obtain the feature map of the detection frame image.

4. The infrared small target tracking method according to claim 1, characterized in that: The invention also includes training the infrared small target tracking model, wherein the training process uses the moving exponential average EMA to optimize the model parameters, specifically: Define a set of model parameters Θ, which includes all training parameters of the Transformer feature fusion module and the resnet50 network; Initialize EMA parameter set Θ EMA is a copy of Θ, ensuring that at the beginning of model training, Θ EMA The EMA parameters in are the same as the initial parameters of the model; In each iteration, the parameters are updated first to obtain a new parameter set Θ t ; Update Θ EMA For each parameter θ belonging to Θ, the corresponding EMA parameter Θ EMA Update according to the preset rules, the formula is: Among them, t represents the current number of iterations, ln represents the logarithmic calculation with e as the base, Indicates that the tth iteration belongs to Θ EMA Parameters, Indicates that the t-1th iteration belongs to Θ EMA The parameter θ t Represents the new parameter set Θ t The parameter ɑ in is a decay coefficient between 0 and 1, which is used to control the degree of fusion between historical parameters and current parameters; At different stages of training, the sensitivity of EMA can be adjusted by adjusting the value of α to adapt to different training dynamics; When validating or testing a model, use Θ EMA As a parameter of the model to track the target.

5. The infrared small target tracking method according to claim 4, characterized in that: A data enhancement mechanism is introduced during the training process of the infrared small target tracking model. The data enhancement mechanism perturbs the bounding box collected during the target tracking process, including: Bounding box random perturbation: The bounding box of a given target is randomly perturbed to simulate the position and scale changes that may occur in consecutive frames. Scale jitter: For each bounding box, multiply its width and height by a random variable drawn from a log-normal distribution with a mean of 0. The standard deviation of the log-normal distribution is controlled by the first parameter. This scale jitter is used to reduce the scale change of the target caused by the distance change in the field of view; Center point jitter: The maximum offset of the center jitter is determined by calculating half of the jitter size and multiplying it by the second parameter. The calculated maximum center point offset is then multiplied by a random vector to obtain the final center point position, where the random vector is uniformly distributed in the interval [0,1] and 0.5 is subtracted from it to produce a value in the interval [-0.5,0.5]; Combined jittered bounding box: The results of random perturbation, scale jittering and center point jittering of the jittered bounding box are combined to form a new jittered bounding box, which is used for target positioning during training.

6. An infrared small target tracking system, used to implement the infrared small target tracking method according to any one of claims 1 to 5, characterized in that: include: The first feature map acquisition module is used to select the first frame image in the infrared image sequence to be tracked as a template frame image, extract the target image area to be tracked of the template frame image, perform feature extraction on the target image area to be tracked, and obtain a feature map of the template frame image; The second feature map acquisition module is used to sequentially load the T-th frame image except the first frame image in the infrared image sequence to be tracked as the detection frame image; perform feature extraction on the detection frame image to obtain a feature map of the detection frame image; Fusion and detection module: used to enhance and fuse the feature map of the T-th detection frame image with the feature map of the template frame image, and then perform infrared weak target detection on the fused features, and output the position and size of the infrared weak target in the T-th detection frame image, where T>1; Judgment module: used to judge whether the T-th detection frame image is the last frame image of the infrared image sequence to be tracked; if so, complete the infrared weak target tracking; if not, return to the second feature acquisition module until the last frame image is detected and the target tracking is completed.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the infrared weak small target tracking method as described in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the infrared weak small target tracking method as claimed in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Infrared weak and small target tracking method based on style re-calibration and improved twin network

    CN116630373A