A method and system for infrared tracking of unmanned aerial vehicle targets in complex scenes

By acquiring time-continuous infrared images in complex scenarios, extracting features and dynamic convolution, combining appearance similarity, position measurement and motion information matching, the fast, accurate and robustness of drone target tracking in complex scenarios is solved, and efficient infrared tracking effect is achieved.

CN119810141BActive Publication Date: 2025-05-23NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510271034.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-05-23
Estimated Expiration
2045-03-07

AI Technical Summary

Technical Problem

In complex scenarios, it is difficult for the prior art to achieve fast, accurate and robust infrared tracking of drone targets, especially when background heat sources are interfering with a lot.

Method used

By obtaining the time-continuous infrared image containing the target, extracting the features of the template image and searching the image set, dynamic convolution is used to perform dynamic convolution, generating detection boxes, and optimizing target boxes selection and trajectory prediction through appearance similarity, position measurement and motion information matching.

Benefits of technology

Fast, accurate and robust infrared tracking of drone targets in complex scenarios is achieved, reducing the problem of poor tracking drift and generalization, and improving the stability and accuracy of tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810141B_ABST
    Figure CN119810141B_ABST
Patent Text Reader

Abstract

The invention discloses an infrared tracking method for unmanned aerial vehicle targets facing complex scenes, comprising: obtaining a set of search images containing targets and being continuous in time; cropping the earliest image to a preset size containing the target to obtain a template image; extracting image features to obtain template features and a search feature set; obtaining dynamic convolution parameters according to the template features; performing dynamic convolution according to the search feature set and the dynamic convolution parameters to obtain all detection frames of each image; selecting the detection frame with the greatest possibility of the target category as a candidate frame; performing appearance similarity, position measurement and motion information matching on the candidate frame of each image to obtain a prediction frame of each image; calculating the Euclidean distance of the center point between the prediction frame of each image and all the detection frames, selecting the detection frame with the smallest Euclidean distance as the target frame of each image; connecting the center points of all the target frames to obtain the target trajectory. The invention also discloses an infrared tracking system for unmanned aerial vehicle targets facing complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle target tracking, and in particular to an infrared tracking method and system for unmanned aerial vehicle targets facing complex scenes. Background Art

[0002] How to effectively monitor and track drones and determine their trajectory and location is crucial to security. Infrared thermal imagers are not affected by lighting conditions and can image in low-light conditions. Detecting and tracking drone targets with infrared thermal imagers has become a hot research topic.

[0003] When tracking UAV targets through infrared thermal imagers, the background of the UAV targets is usually very complex, including complex scenes such as buildings, woods and urban areas. There are many heat sources in these complex scenes that interfere with the targets, which makes it very difficult to track UAV targets robustly and accurately. Existing infrared UAV target tracking methods that consider complex scenes mainly include methods based on correlation filter tracking and methods based on deep network model tracking. In the method based on correlation filter tracking, a filter template is first designed, and then the template is used to perform correlation operations on the target candidate area. However, this method relies on manually designed features and is sensitive to background noise. Among the methods based on deep network model tracking, there are usually matching-based deep twin networks, classification-based online tracking methods and Transformer architecture-based methods. The matching-based deep twin network extracts the features of the template and the search area respectively through a weight-sharing two-stream backbone network, and then fuses and interacts the two features. The classification-based online tracking method uses an online learning classifier to enhance the perception of appearance changes to improve the robustness of target tracking. The method based on the Transformer architecture establishes long-distance dependencies between features to improve the feature representation ability of the model. However, due to the extremely small size of UAV targets, limited extractable features, and the presence of a large amount of background clutter and thermal cross noise and other challenging scenarios, the classification-based online tracking method and the Transformer architecture-based method are difficult to achieve satisfactory tracking results when tracking UAV targets; in addition, the existing matching-based deep twin network uses a two-stage detector. Although it has achieved good UAV target tracking results, it requires a lot of computing resources and time, and cannot meet the real-time requirements of UAV target tracking. The tracking accuracy and robustness also need to be improved, and there are problems of tracking drift and poor generalization in complex scenarios.

[0004] Therefore, a new technical solution is urgently needed to solve the technical problem of how to achieve fast, accurate and robust infrared tracking of UAV targets in complex scenarios. Summary of the invention

[0005] The present invention provides a method and system for infrared tracking of unmanned aerial vehicle targets in complex scenes, so as to solve the technical problem of how to achieve fast, accurate and robust infrared tracking of unmanned aerial vehicle targets in complex scenes.

[0006] To achieve the above object, the present invention provides a method for infrared tracking of unmanned aerial vehicle targets in complex scenes, comprising:

[0007] A preset number of infrared images containing the target and continuous in time are acquired to obtain a search image set.

[0008] The earliest image in the search image set is cropped to a preset size that includes the target to obtain a template image.

[0009] The features of the template image and the search image set are extracted to obtain the template features and the search feature set respectively.

[0010] The dynamic convolution parameters are obtained according to the template features; the dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain the entire detection frame of each image.

[0011] The detection box with the highest probability of the target category among all the detection boxes is selected as the candidate box; the candidate boxes of each image are matched with appearance similarity, position metric matching and motion information matching to obtain the predicted box of each image.

[0012] Calculate the Euclidean distance between the center point of the prediction box of each image and all the detection boxes, select the detection box with the smallest Euclidean distance as the target box of each image; connect the center points of the target boxes of each image to obtain the target trajectory.

[0013] Preferably, extracting features of the template image and the search image set to obtain the template features and the search feature set respectively comprises:

[0014] The template image is extracted through Swin Transformer to obtain multi-scale template features; the P 2 , P 3 and P 4 The multi-scale template features are aggregated in the layer to obtain the template features.

[0015] The Swin Transformer is used to extract features from the images in the search image set to obtain a multi-scale search feature set. 2 , P 3 and P 4 The layer performs multi-scale feature aggregation on the features in the multi-scale search feature set to obtain the search feature set.

[0016] Preferably, obtaining the dynamic convolution parameters according to the template features includes:

[0017] The template features are cropped again to retain the target and obtain the target area information.

[0018] pass The convolution adjusts the number of channels of the target area information to a preset number to obtain the first information; an m-dimensional vector is generated according to the first information through global average pooling, where m is a positive integer, to obtain the dynamic convolution parameters.

[0019] Preferably, dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain the entire detection frame of each image including:

[0020] Dynamic convolution includes classification processing part and regression processing part.

[0021] In the classification processing part, the features in the search feature set are multiplied by the dynamic convolution parameters and then go through four basic Convolution, get the first feature; after encoding the first feature, multiply it with the dynamic convolution parameter, and go through a set of basic After convolution, the target category information is output through 2 channels and activation function.

[0022] In the regression processing part, the features in the search feature set are multiplied by the dynamic convolution parameters and then go through four basic Convolution, get the first feature; after encoding the first feature, multiply it with the dynamic convolution parameter, and go through a set of basic After convolution, the target rectangular frame coordinate information is output through 4 channels, and the target center point offset information is output through 2 channels.

[0023] All the rough selected detection frames of each image are obtained according to the target rectangular frame coordinate information and the target center point offset information; all the rough selected detection frames are sorted based on the target category possibility according to the target category information, and the first preset number of rough selected detection frames with larger target category possibility are selected as all the detection frames.

[0024] Preferably, appearance similarity matching, position metric matching and motion information matching are performed on the candidate boxes of each image respectively, and the predicted boxes of each image are obtained, including:

[0025] Perform appearance similarity matching, position metric matching and motion information matching on the candidate box of a single image to obtain the first output, second output and third output respectively; perform Hadamard multiplication on the first output, second output and third output to obtain the predicted box of a single image C , which can be expressed as:

[0026] ;

[0027] in, represents the first output; represents the second output; represents the third output; represents Hadamard multiplication;

[0028] Solve each image to get the predicted box of each image.

[0029] Preferably, the appearance similarity matching includes:

[0030] The appearance similarity matching adopts a cross-fusion structure. j The selection boxes of each image are passed through three Convolution, corresponding to the generation of vectors Q, K and V; the first output It can be expressed as:

[0031] ;

[0032] in, Indicates The selection box of the image; Indicates The selection box of the image; Represents the self-attention obtained by multiplying vectors Q, K and V.

[0033] Preferably, the location metric matching includes:

[0034] The second output obtained by position metric matching It can be expressed as:

[0035] ;

[0036] in, and Represent the predicted box and the selected box respectively, and Respectively represent the union area and intersection area of ​​the predicted box and the selected box, Represents the Euclidean distance between the center point of the selected box and the predicted box. is the weight hyperparameter, Indicates aspect ratio similarity.

[0037] Preferably, the motion information matching includes:

[0038] The third output obtained by motion information matching It can be expressed as:

[0039] ;

[0040] in, and Respectively represent the center positions of the predicted box and the selected box; Represents the covariance matrix S The inverse of , T means transpose.

[0041] The present invention also provides an infrared tracking system for unmanned aerial vehicle targets in complex scenes, which is used in the method of the present invention. The system includes an image processing module, a feature extraction module, a feature association module, a trajectory prediction module and a trajectory association module.

[0042] The image processing module is used to obtain a preset number of infrared images that contain the target and are continuous in time to obtain a search image set; the earliest image in the search image set is cropped to a preset size that contains the target to obtain a template image.

[0043] The feature extraction module is used to extract the features of the template image and the search image set to obtain the template features and the search feature set respectively.

[0044] The feature association module is used to obtain dynamic convolution parameters based on template features; dynamic convolution is performed based on the search feature set and the dynamic convolution parameters to obtain the entire detection frame of each image.

[0045] The trajectory prediction module is used to select the detection box with the highest probability of the target category among all the detection boxes as the candidate box; the appearance similarity matching, position metric matching and motion information matching are performed on the candidate boxes of each image to obtain the predicted box of each image.

[0046] The trajectory association module is used to calculate the Euclidean distance between the center points of the prediction box of each image and all the detection boxes, and select the detection box with the smallest Euclidean distance as the target box of each image; the center points of the target boxes of each image are connected to obtain the target trajectory.

[0047] Preferably, the system also includes a system testing module.

[0048] The system testing module is used to obtain a preset number of infrared images of pre-labeled targets and continuous in time to obtain a training image set; based on the training image set, the image processing module, feature extraction module, feature association module, trajectory prediction module and trajectory association module in the system are trained; the tracking result indicators of the training results are calculated and whether the system meets the usage requirements is determined based on the tracking result indicators.

[0049] The present invention has the following beneficial effects:

[0050] The infrared tracking method for unmanned aerial vehicle targets facing complex scenes of the present invention provides a target infrared image data basis with temporal continuity for subsequent target tracking by acquiring infrared images containing the target and being continuous in time, so that the method can track the target more accurately. The earliest image is cropped to a preset size containing the target, providing a template for subsequent target detection of the method of the present invention, so that the method can detect targets in all images based on the template. By extracting the features of the template image and the search image set, the method obtains the features of the template image and all images, providing a data basis for subsequent target detection. Dynamic convolution parameters are obtained according to the template features, providing parameters for subsequent dynamic convolution. Dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain all detection frames of each image, so that the method obtains all possible target detection frames in all images, providing a selection range for subsequent target frame selection, and establishing a feature association between the template image and all images, making the use of the template more flexible and reducing the complexity of the method. The candidate boxes of each image are matched for appearance similarity, position metric matching and motion information matching to obtain a prediction box. By utilizing the appearance feature information, position information and motion information of the target, the tracking drift problem existing in the global instance search scheme is reduced, so that a smoother and more accurate target trajectory can be obtained in the future, and a more stable target tracking can be achieved. By calculating the Euclidean distance of the center point between the prediction box and all the detection boxes, the detection box with the smallest Euclidean distance is selected as the target box of each image, so that the method of the present invention obtains the target box and realizes the target tracking of a single image. By connecting the center points of the target boxes of each image, the target trajectory is obtained, and the trajectory tracking of the target is realized. The method of the present invention has good robustness and high precision, and can realize fast, accurate and robust infrared tracking of drone targets in complex scenes.

[0051] The complex scene-oriented unmanned aerial vehicle target infrared tracking system of the present invention is used in the method of the present invention and has the same technical effect as the method of the present invention.

[0052] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:

[0054] Figure 1 It is a schematic diagram of a method flow of a preferred embodiment of the present invention.

[0055] Figure 2It is a system schematic diagram of a preferred embodiment of the present invention.

[0056] Figure 3 This is a diagram of infrared image experimental results of a scene according to a preferred embodiment of the present invention.

[0057] Figure 4 This is a diagram of infrared image experimental results under scene 2 of a preferred embodiment of the present invention.

[0058] Figure 5 This is a diagram of infrared image experimental results under scene three of the preferred embodiment of the present invention.

[0059] Figure 6 This is a diagram of infrared image experimental results in a field scene according to a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0060] The embodiments of the present invention are described in detail below with reference to the accompanying drawings, but the present invention can be implemented in many different ways as defined and covered by the claims.

[0061] See also Figure 1 In a preferred embodiment of the present invention, a method for infrared tracking of unmanned aerial vehicle targets in complex scenes is provided, comprising:

[0062] S1. Acquire a preset number of infrared images that contain the target and are continuous in time to obtain a search image set.

[0063] S2. Crop the earliest image in the search image set to a preset size that includes the target, and obtain a template image.

[0064] S3, extracting features of the template image and the search image set to obtain template features and search feature sets, respectively. Specifically including:

[0065] The template image is extracted through Swin Transformer to obtain multi-scale template features; the P 2 , P 3 and P 4 The layer aggregates the multi-scale features of the multi-scale template to obtain the template features;

[0066] The Swin Transformer is used to extract features from the images in the search image set to obtain a multi-scale search feature set. 2 , P 3 and P 4 The layer performs multi-scale feature aggregation on the features in the multi-scale search feature set to obtain the search feature set.

[0067] S4. Dynamic convolution parameters are obtained according to the template features; dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain the entire detection frame of each image.

[0068] In a preferred embodiment of the present invention, obtaining dynamic convolution parameters according to template features includes:

[0069] Perform secondary cropping of the template features to retain the target and obtain the target area information;

[0070] pass The convolution adjusts the number of channels of the target area information to a preset number to obtain the first information; an m-dimensional vector is generated according to the first information through global average pooling, where m is a positive integer, to obtain the dynamic convolution parameters.

[0071] In a preferred embodiment of the present invention, dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain the entire detection frame of each image including:

[0072] Dynamic convolution includes classification processing part and regression processing part;

[0073] In the classification processing part, the features in the search feature set are multiplied by the dynamic convolution parameters and then go through four basic Convolution, get the first feature; after encoding the first feature, multiply it with the dynamic convolution parameter, and go through a set of basic After convolution, the target category information is output through 2 channels and activation function;

[0074] In the regression processing part, the features in the search feature set are multiplied by the dynamic convolution parameters and then go through four basic Convolution, get the first feature; after encoding the first feature, multiply it with the dynamic convolution parameter, and go through a set of basic After convolution, the target rectangle coordinate information is output through 4 channels, and the target center point offset information is output through 2 channels;

[0075] All the rough selected detection frames of each image are obtained according to the target rectangular frame coordinate information and the target center point offset information; all the rough selected detection frames are sorted based on the target category possibility according to the target category information, and the first preset number of rough selected detection frames with larger target category possibility are selected as all the detection frames.

[0076] S5. Select the detection box with the highest probability of the target category among all the detection boxes as the candidate box; perform appearance similarity matching, position metric matching and motion information matching on the candidate boxes of each image respectively to obtain the predicted box of each image.

[0077] In a preferred embodiment of the present invention, appearance similarity matching, position metric matching and motion information matching are performed on the candidate boxes of each image respectively, and the predicted boxes of each image are obtained, including:

[0078] Perform appearance similarity matching, position metric matching and motion information matching on the candidate box of a single image to obtain the first output, second output and third output respectively; perform Hadamard multiplication on the first output, second output and third output to obtain the predicted box of a single image C , which can be expressed as:

[0079] ;

[0080] in, represents the first output; represents the second output; represents the third output; represents Hadamard multiplication;

[0081] Solve each image to get the predicted box of each image.

[0082] In a preferred embodiment of the present invention, the appearance similarity matching includes:

[0083] The appearance similarity matching adopts a cross-fusion structure. j The selection boxes of each image are passed through three Convolution, corresponding to the generation of vectors Q, K and V; the first output It can be expressed as:

[0084] ;

[0085] in, Indicates The selection box of the image; Indicates The selection box of the image; Represents the self-attention obtained by multiplying vectors Q, K and V.

[0086] In a preferred embodiment of the present invention, the location metric matching includes:

[0087] The second output obtained by position metric matching It can be expressed as:

[0088] ;

[0089] in, and Represent the predicted box and the selected box respectively, and Respectively represent the union area and intersection area of ​​the predicted box and the selected box, Represents the Euclidean distance between the center point of the selected box and the predicted box. is the weight hyperparameter, Indicates aspect ratio similarity.

[0090] In a preferred embodiment of the present invention, motion information matching includes:

[0091] The third output obtained by motion information matching It can be expressed as:

[0092] ;

[0093] in, and Respectively represent the center positions of the predicted box and the selected box; Represents the covariance matrix S The inverse of , T means transpose.

[0094] S6. Calculate the Euclidean distance between the center point of the prediction box of each image and all the detection boxes, select the detection box with the smallest Euclidean distance as the target box of each image; connect the center points of the target boxes of each image to obtain the target trajectory.

[0095] The infrared tracking method for unmanned aerial vehicle targets facing complex scenes of the present invention provides a target infrared image data basis with temporal continuity for subsequent target tracking by acquiring infrared images containing the target and being continuous in time, so that the method can track the target more accurately. The earliest image is cropped to a preset size containing the target, providing a template for subsequent target detection of the method of the present invention, so that the method can detect targets in all images based on the template. By extracting the features of the template image and the search image set, the method obtains the features of the template image and all images, providing a data basis for subsequent target detection. Dynamic convolution parameters are obtained according to the template features, providing parameters for subsequent dynamic convolution. Dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain all detection frames of each image, so that the method obtains all possible target detection frames in all images, providing a selection range for subsequent target frame selection, and establishing a feature association between the template image and all images, making the use of the template more flexible and reducing the complexity of the method. The candidate box of each image is matched for appearance similarity, position metric matching and motion information matching to obtain a prediction box. By utilizing the appearance feature information, position information and motion information of the target, the tracking drift problem existing in the global instance search scheme is reduced, so that a smoother and more accurate target trajectory can be obtained in the future, and a more stable target tracking can be achieved. By calculating the Euclidean distance of the center point between the prediction box and all the detection boxes, the detection box with the smallest Euclidean distance is selected as the target box of each image, so that the method of the present invention obtains the target box and realizes the target tracking of a single image. By connecting the center points of the target boxes of each image, the target trajectory is obtained, and the trajectory tracking of the target is realized. The method of the present invention has good robustness and high precision, can be applied to complex scenes, and is practical.

[0096] See also Figure 2 ,The present invention also provides an infrared tracking system for unmanned aerial vehicle targets facing complex scenes, which is used in the method of the present invention. The system includes an image processing module, a feature extraction module, a feature association module, a trajectory prediction module and a trajectory association module;

[0097] The image processing module is used to obtain a preset number of infrared images that contain the target and are continuous in time to obtain a search image set; the earliest image in the search image set is cropped to a preset size that contains the target to obtain a template image;

[0098] The feature extraction module is used to extract the features of the template image and the search image set, and obtain the template features and the search feature set respectively;

[0099] The feature association module is used to obtain dynamic convolution parameters based on template features; dynamic convolution is performed based on the search feature set and dynamic convolution parameters to obtain all detection frames of each image;

[0100] The trajectory prediction module is used to select the detection box with the highest probability of the target category among all the detection boxes as the candidate box; the candidate boxes of each image are matched by appearance similarity, position metric matching and motion information matching to obtain the predicted box of each image;

[0101] The trajectory association module is used to calculate the Euclidean distance between the center points of the prediction box of each image and all the detection boxes, and select the detection box with the smallest Euclidean distance as the target box of each image; the center points of the target boxes of each image are connected to obtain the target trajectory.

[0102] In a preferred embodiment of the present invention, the system further comprises a system testing module;

[0103] The system testing module is used to obtain a preset number of infrared images of pre-labeled targets and continuous in time to obtain a training image set; based on the training image set, the image processing module, feature extraction module, feature association module, trajectory prediction module and trajectory association module in the system are trained; the tracking result indicators of the training results are calculated and whether the system meets the usage requirements is determined based on the tracking result indicators.

[0104] The complex scene-oriented unmanned aerial vehicle target infrared tracking system of the present invention is used in the method of the present invention and has the same technical effect as the method of the present invention.

[0105] Verification part:

[0106] In a preferred embodiment of the present invention, the test area is The wavelength of the infrared camera used in the test is between 8 and 14 μm. The search image set is obtained by shooting an infrared video with a frame rate of 30 frames and saving it as an infrared image. The resolution is Pixels. A total of 200 original infrared videos were collected, with a total length of about 3,000 seconds and a total of 90,000 frames of infrared images.

[0107] In a preferred embodiment of the present invention, target tracking experiments are carried out based on a total of 8 methods including ATOM, DiMP18, PrDiMP50, ToMP101, TaMOs, SiamCAR, DMtrack and the method of the present invention, and the tracking effects and index performances are compared.

[0108] The experimental results can be found in Figures 3 to 5 , Figures 3 to 5 These are the experimental results of a single infrared image under scene one, scene two, and scene three. Figures 3 to 5 The true value in is the manually marked target location box. Figures 3 to 5It can be seen that, except for the method of the present invention, other methods cannot achieve stable tracking when dealing with complex scenes such as thermal crosstalk, scale changes, targets entering and exiting the field of view, and dynamic background clutter. The method of the present invention can always continuously and stably track the target in the image, indicating that the method of the present invention can well cope with these challenging complex scenes.

[0109] For a comparison of performance indicators, see Table 1. The performance indicators include Precision, Success and SA. The ↑ next to the indicator indicates that the larger the value of the indicator, the better the tracking performance. As can be seen from Table 1, the method of the present invention shows better performance than other methods in all evaluation indicators Precision, Success and SA. Although UAV targets often shuttle through buildings, woods, mountains and other backgrounds with serious clutter interference, making it difficult to track infrared UAV targets, the method of the present invention can still stably track UAV targets, handle these situations well, and effectively deal with these challenging scenarios.

[0110] Table 1 Performance index comparison table

[0111] Methods Precision↑ Success↑ SA↑ Methods Precision↑ Success↑ SA↑ ATOM 0.617 0.412 0.4190 TaMOs 0.789 0.469 0.4760 DiMP18 0.715 0.482 0.4901 SiamCAR 0.535 0.338 0.3426 PrDiMP50 0.764 0.510 0.5189 DMtrack 0.780 0.504 0.5117 ToMP101 0.752 0.506 0.5150 Method of the present invention 0.838 0.561 0.5647

[0112] The results of target tracking in a certain field according to the method of the present invention are shown in Figure 6 ,Depend on Figure 6 It can be seen that the method of the present invention can meet the actual application needs, can realize the UAV target tracking function in complex scenarios, and has broad anti-UAV application prospects.

[0113] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for infrared tracking of unmanned aerial vehicle targets in complex scenes, characterized in that: include: Acquire a preset number of infrared images that contain the target and are continuous in time to obtain a search image set; The earliest image in the search image set is cropped to a preset size that includes the target, to obtain a template image; Extracting features of the template image and the search image set to obtain template features and a search feature set, respectively; Obtaining dynamic convolution parameters according to the template features; performing dynamic convolution according to the search feature set and the dynamic convolution parameters to obtain all detection frames of each image; Select the detection box with the highest probability of the target category among all the detection boxes as the box to be selected; Perform appearance similarity matching, position metric matching, and motion information matching on the candidate boxes of each image to obtain the predicted box of each image; Calculate the Euclidean distance between the center point of the prediction box of each image and all the detection boxes, and select the detection box with the smallest Euclidean distance as the target box of each image; Connect the center points of the target boxes of each image to obtain the target trajectory.

2. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 1 is characterized in that: Extracting features of the template image and the search image set to obtain template features and search feature sets respectively includes: Extracting features of the template image through Swin Transformer to obtain multi-scale template features; performing multi-scale feature aggregation on the multi-scale template features through P2, P3 and P4 layers of a feature pyramid network to obtain the template features; The images in the search image set are subjected to feature extraction through Swin Transformer to obtain a multi-scale search feature set; the features in the multi-scale search feature set are subjected to multi-scale feature aggregation through the P2, P3 and P4 layers of the feature pyramid network to obtain the search feature set.

3. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 2 is characterized in that: The dynamic convolution parameters obtained according to the template features include: Performing secondary cropping of the template features to retain the target, to obtain target area information; pass Convolution adjusts the number of channels of the target area information to a preset number to obtain first information; an m-dimensional vector is generated through global average pooling according to the first information, where m is a positive integer, to obtain the dynamic convolution parameter.

4. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 3 is characterized in that: Dynamic convolution is performed according to the search feature set and the dynamic convolution parameters to obtain all detection frames of each image including: The dynamic convolution includes a classification processing part and a regression processing part; In the classification processing part, the features in the search feature set are multiplied by the dynamic convolution parameters and then go through four basic Convolution, obtain the first feature; encode the first feature and multiply it with the dynamic convolution parameter, and then pass a set of basic After convolution, the target category information is output through 2 channels and activation function; In the regression processing part, the features in the search feature set are multiplied by the dynamic convolution parameters and then go through four basic Convolution, obtain the first feature; encode the first feature and multiply it with the dynamic convolution parameter, and then pass a set of basic After convolution, the target rectangle coordinate information is output through 4 channels, and the target center point offset information is output through 2 channels; All rough-selected detection frames of each image are obtained according to the target rectangular frame coordinate information and the target center point offset information; all rough-selected detection frames are sorted based on the target category possibility according to the target category information, and a preset number of rough-selected detection frames with greater target category possibility are selected as all detection frames.

5. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 4 is characterized in that: The appearance similarity matching, position metric matching and motion information matching are performed on the candidate boxes of each image respectively, and the predicted boxes of each image are obtained including: Perform appearance similarity matching, position metric matching and motion information matching on the candidate box of a single image to obtain the first output, the second output and the third output respectively; perform Hadamard multiplication on the first output, the second output and the third output to obtain the predicted box of the single image C , which can be expressed as: ; in, represents the first output; represents the second output; represents the third output; represents Hadamard multiplication; Solving each image, that is, obtaining the prediction box of each image.

6. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 5 is characterized in that: The appearance similarity matching includes: The appearance similarity matching adopts a cross fusion structure. j The selection boxes of each image are passed through three Convolution, corresponding to the generation of vectors Q, K and V; the first output It can be expressed as: ; in, Indicates The selection box of the image; Indicates The selection box of the image; Represents the self-attention obtained by multiplying vectors Q, K and V.

7. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 6 is characterized in that: The position metric matching includes: The second output obtained by position metric matching It can be expressed as: ; in, and Represent the predicted box and the selected box respectively, and Respectively represent the union area and intersection area of ​​the predicted box and the selected box, Represents the Euclidean distance between the center point of the selected box and the predicted box. is the weight hyperparameter, Indicates aspect ratio similarity.

8. The infrared tracking method for unmanned aerial vehicle targets in complex scenes according to claim 7 is characterized in that: The motion information matching includes: The third output obtained by motion information matching It can be expressed as: ; in, and Respectively represent the center positions of the predicted box and the selected box; Represents the covariance matrix S The inverse of , T means transpose.

9. An infrared tracking system for unmanned aerial vehicle targets in complex scenes, used in the method described in any one of claims 1 to 8, characterized in that: The system includes an image processing module, a feature extraction module, a feature association module, a trajectory prediction module and a trajectory association module; The image processing module is used to obtain a preset number of infrared images that contain the target and are continuous in time to obtain a search image set; the earliest image in the search image set is cropped to a preset size that contains the target to obtain a template image; The feature extraction module is used to extract features of the template image and the search image set to obtain template features and a search feature set respectively; The feature association module is used to obtain dynamic convolution parameters according to the template features; perform dynamic convolution according to the search feature set and the dynamic convolution parameters to obtain all detection frames of each image; The trajectory prediction module is used to select the detection box with the highest probability of the target category among all the detection boxes as the to-be-selected box; Perform appearance similarity matching, position metric matching, and motion information matching on the candidate boxes of each image to obtain the predicted box of each image; The trajectory association module is used to calculate the Euclidean distance between the center points of the prediction frame of each image and all the detection frames, select the detection frame with the smallest Euclidean distance as the target frame of each image; connect the center points of the target frames of each image to obtain the target trajectory.

10. The complex scene-oriented UAV target infrared tracking system according to claim 9, characterized in that: The system also includes a system testing module; The system testing module is used to obtain a preset number of infrared images of pre-labeled targets and continuous in time to obtain a training image set; based on the training image set, the image processing module, feature extraction module, feature association module, trajectory prediction module and trajectory association module in the system are trained; the tracking result indicators of the training results are calculated and whether the system meets the usage requirements is determined based on the tracking result indicators.

Citation Information

Patent Citations

  • Lightweight infrared unmanned aerial vehicle target tracking method based on Siamese network

    CN115909110A

  • Unmanned aerial vehicle visual identification method and device

    CN119152335A