Target identification tracking algorithm and device

By combining the target recognition and tracking algorithm of static image recognition and dynamic video tracking, the problems of limited detection accuracy and missed detection in low-resolution infrared images are solved, and efficient target recognition and tracking are achieved, meeting the real-time and accuracy requirements of the embedded platform.

CN120673027APending Publication Date: 2025-09-19SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510650229.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing target detection algorithms have limited detection accuracy in low-resolution infrared images and are prone to missed detections. In addition, traditional non-maximum suppression processing is prone to mistakenly deleting correct detection frames in scenes with dense or overlapping multiple targets, resulting in missed detections.

Method used

Build a target recognition and tracking algorithm that combines static image recognition and dynamic video tracking. Utilizing the YOLOv11 and TCTrack algorithms, through a lightweight backbone network and feature fusion network, combined with a deformable decoder and an end-to-end prediction head, it replaces the traditional NMS processing to achieve efficient candidate target screening and continuous frame tracking.

Benefits of technology

It improves the real-time recognition and tracking accuracy of infrared targets, reduces the model calculation complexity, meets the power consumption and real-time requirements of the embedded platform, and ensures the target detection accuracy and consistency in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673027A_ABST
    Figure CN120673027A_ABST
Patent Text Reader

Abstract

The invention discloses a target identification tracking algorithm and device, and the algorithm comprises the steps: constructing an initial target identification model, obtaining a target identification weight parameter, and initializing the target identification model based on the target identification weight parameter; the target identification weight parameter is obtained based on static image data training; constructing a target tracking model and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter; the target tracking weight parameter is obtained based on dynamic video data training; acquiring to-be-detected image data, identifying the to-be-detected image data based on the target identification model, and acquiring candidate targets; and determining an interested target in the candidate targets, and tracking the interested target based on the target tracking model to complete target identification and tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition and detection, and in particular to a target recognition and tracking algorithm and device. Background Art

[0002] With the rapid development of drone technology, using drone-mounted infrared imaging equipment to acquire information about ground targets has become an important method. Compared to visible light images, infrared images are more adaptable to complex lighting environments and are widely used in military reconnaissance, security monitoring, disaster search and rescue, and other fields. Due to the maneuverability of drones and the complex backgrounds they capture, traditional target detection and tracking algorithms face limitations in terms of accuracy and real-time performance. Among target detection algorithms, the YOLO series has garnered widespread attention due to its end-to-end detection approach and high real-time performance. With the continuous evolution of deep learning and network architectures, the YOLO series has seen multiple versions and improvements. However, when dealing with low-resolution infrared images, it still suffers from limited detection accuracy and a high risk of missed detections. Furthermore, to meet the power and computing resource constraints of embedded platforms, further lightweighting and improving real-time performance are current research hotspots.

[0003] Existing object detection networks typically require non-maximum suppression (NMS) post-processing to remove redundant detection frames. However, in target detection scenarios where multiple objects are densely packed or overlapping, or their features are not distinct enough, NMS may mistakenly delete correct detection frames, resulting in missed detections. Summary of the Invention

[0004] The present invention provides a target recognition and tracking algorithm and device to improve the accuracy of target recognition and tracking in complex motion scenes.

[0005] In order to solve the above technical problems, the present invention provides a target recognition and tracking algorithm, comprising:

[0006] Constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters; the target recognition weight parameters are obtained based on training of static image data;

[0007] Constructing a target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters; the target tracking weight parameters are obtained based on training of dynamic video data;

[0008] Acquire image data to be detected, identify the image data to be detected based on the target recognition model, and obtain a candidate target data set;

[0009] An interested target is determined in the candidate target data set, and the interested target is tracked based on the target tracking model to complete target recognition and tracking.

[0010] By combining static image recognition with dynamic video tracking, the present invention achieves an end-to-end recognition-tracking closed loop. The target detection model is initialized using target recognition weight parameters obtained through static image training, effectively screening candidate targets from the image to be detected. The target tracking model is then initialized using tracking weight parameters obtained through dynamic video training, and continuous frame tracking of the selected targets of interest is performed. This not only improves the accuracy of target detection, but also effectively ensures the consistency and robustness of the tracking process, avoiding the problem of missed detection or loss of targets that can occur with a single model in complex scenarios. This significantly enhances the accuracy of real-time recognition and tracking of ground-based infrared targets.

[0011] Furthermore, the constructing of the initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters, includes:

[0012] Build a target recognition model based on the YOLOv11 algorithm, obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters;

[0013] The target recognition model includes a first backbone network, a feature fusion network and a detection head network;

[0014] The first backbone network includes a CBS module, an MNv4 module, an SPPF module and a C2PSA_2 module;

[0015] The feature fusion network includes 5 Fusion modules;

[0016] The detection head network includes a deformable decoder and a prediction head.

[0017] This paper modularizes the existing backbone network, feature fusion network, and detection head network. It uses CBS, MNv4, SPPF, and C2PSA_2 modules to form a lightweight backbone network with multi-scale perception capabilities. Five Fusion modules are used to achieve efficient feature fusion. A deformable decoder and end-to-end prediction head replace the traditional NMS post-processing process. This ensures small target detection accuracy in low-resolution infrared images while effectively reducing model computational complexity, meeting the dual requirements of power consumption and real-time performance for embedded deployment.

[0018] Furthermore, the target recognition weight parameters are obtained based on static image data training, including:

[0019] Generating static image data based on the collection of multi-scene static images by the UAV, and annotating the static image data to obtain an infrared target data set;

[0020] The target recognition model is trained based on the infrared target data set until a preset condition is met, and a target recognition weight parameter is output.

[0021] This method uses drones to collect static infrared images of multiple scenes and combines them with semi-automated annotation to construct a professional dataset covering complex backgrounds and multiple viewpoints. Based on this dataset, a recognition model integrating YOLOv11 and MobileNetV4 is iteratively trained, enabling the model to fully learn the thermal signature distribution of infrared targets in different scenarios. This process not only improves the diversity of training samples and the quality of annotations, but also ensures the generalization and detection accuracy of the recognition model in real-world scenarios by selecting optimal weights based on pre-defined conditions.

[0022] Furthermore, the constructing of the target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters, include:

[0023] Building a target tracking model based on the TCTrack algorithm; and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter;

[0024] The target tracking model includes a second backbone network, an ATT-TAda module, an AT-Trans module and a prediction head;

[0025] The second backbone network is a MobileNetV4 module, which includes a convolution module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling module;

[0026] The ATT-TAda module includes a maximum pooling layer, a cross attention module, an average pooling layer and a convolutional layer;

[0027] The AT-Trans module includes five multi-head attention mechanism modules and a temporal information filter module;

[0028] The prediction head includes a classification module and a regression module arranged in parallel.

[0029] The target tracking model of this invention includes a complete architecture consisting of a MobileNetV4 backbone network, an ATT-TAda spatiotemporal context fusion module, an AT-Trans multi-head attention temporal refinement module, and a parallel prediction head. MobileNetV4 provides lightweight deep feature extraction. The ATT-TAda module incorporates multi-frame contextual information into convolutional weights through cross-attention, enabling dynamic adaptation to target motion and background changes. The AT-Trans module utilizes an encoder-decoder structure to refine the similarity graph, improving positioning accuracy in scenes with occlusion or rapid motion.

[0030] Furthermore, the target tracking weight parameters are obtained based on training of a dynamic video dataset and include:

[0031] Collecting continuous video data of multiple scenes based on a drone, performing sequence tracking on the continuous video data of multiple scenes to generate an infrared tracking dataset;

[0032] The target tracking model is trained based on the infrared tracking data set until a preset condition is met, and a target tracking weight parameter is output.

[0033] This paper constructs a high-quality infrared tracking dataset by continuously capturing and annotating multi-scene videos from drones. This dataset covers a wide range of motion patterns, occlusion conditions, and background variations, making the training process based on the aforementioned tracking model more realistic for real-world applications. Through iterative training until preset accuracy requirements are met, the resulting tracking weights enable stable and continuous tracking of infrared targets in complex dynamic scenes, significantly reducing offline parameter adjustment costs and providing a reliable real-time tracking solution for drone platforms.

[0034] In a second aspect, the present invention provides a target recognition and tracking device, comprising: a recognition model initialization module, a tracking model initialization module, a target recognition module, and a target tracking module;

[0035] The recognition model initialization module is used to construct an initialized target recognition model and obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters; the target recognition weight parameters are obtained based on static image data training;

[0036] The tracking model initialization module is used to build a target tracking model and obtain target tracking weight parameters, and initialize the target tracking model based on the target tracking weight parameters; the target tracking weight parameters are obtained based on dynamic video data training;

[0037] The target recognition module is used to obtain image data to be detected, identify the image data to be detected based on the target recognition model, and obtain a candidate target data set;

[0038] The target tracking module is used to determine the target of interest in the candidate target data set, track the target of interest based on the target tracking model, and complete target recognition and tracking.

[0039] Furthermore, the constructing of the initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters, further includes:

[0040] Build a target recognition model based on the YOLOv11 algorithm, obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters;

[0041] The target recognition model includes a first backbone network, a feature fusion network and a detection head network;

[0042] The first backbone network includes a CBS module, an MNv4 module, an SPPF module and a C2PSA_2 module;

[0043] The feature fusion network includes 5 Fusion modules;

[0044] The detection head network includes a deformable decoder and a prediction head.

[0045] Furthermore, the target recognition weight parameters are obtained based on static image data training, including:

[0046] Generating static image data based on the collection of multi-scene static images by the UAV, and annotating the static image data to obtain an infrared target data set;

[0047] The target recognition model is trained based on the infrared target data set until a preset condition is met, and a target recognition weight parameter is output.

[0048] Furthermore, the constructing of the target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters, include:

[0049] Building a target tracking model based on the TCTrack algorithm; and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter;

[0050] The target tracking model includes a second backbone network, an ATT-TAda module, an AT-Trans module and a prediction head;

[0051] The second backbone network is a MobileNetV4 module, which includes a convolution module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling module;

[0052] The ATT-TAda module includes a maximum pooling layer, a cross attention module, an average pooling layer and a convolutional layer;

[0053] The AT-Trans module consists of five multi-head attention mechanism modules and a temporal information filter module.

[0054] Furthermore, the target tracking weight parameters are obtained based on training of a dynamic video dataset and include:

[0055] Collecting continuous video data of multiple scenes based on a drone, performing sequence tracking on the continuous video data of multiple scenes to generate an infrared tracking dataset;

[0056] The target tracking model is trained based on the infrared tracking data set until a preset condition is met, and a target tracking weight parameter is output. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram of a flow chart of a target recognition and tracking algorithm provided in an embodiment of the present invention;

[0058] Figure 2 A schematic diagram of the structure of a target recognition model provided by an embodiment of the present invention;

[0059] Figure 3 A schematic diagram of the structure of a CBS module provided in an embodiment of the present invention;

[0060] Figure 4 A schematic structural diagram of a FIB module provided in an embodiment of the present invention;

[0061] Figure 5 A schematic structural diagram of a C2PSA_2 module provided in an embodiment of the present invention;

[0062] Figure 6 A schematic diagram of the structure of the Fusion module provided in an embodiment of the present invention;

[0063] Figure 7 A schematic diagram of the structure of the Simplify Rep module provided in an embodiment of the present invention;

[0064] Figure 8 A schematic diagram of the structure of a target tracking model provided by an embodiment of the present invention;

[0065] Figure 9 A schematic structural diagram of a second backbone network provided in an embodiment of the present invention;

[0066] Figure 10 A schematic structural diagram of the ATT-TAda module provided in an embodiment of the present invention;

[0067] Figure 11 A schematic structural diagram of an AT-Trans module provided in an embodiment of the present invention;

[0068] Figure 12 Another flowchart of a target recognition and tracking algorithm provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0069] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0070] The terms "first," "second," and the like in the specification, claims, and drawings of this application are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0071] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0072] First, some terms in this application are explained to facilitate understanding by those skilled in the art.

[0073] Example 1

[0074] See also Figure 1 , Figure 1 The present invention provides a target recognition and tracking algorithm according to an embodiment of the present invention. The present invention provides a target recognition and tracking algorithm, including steps 101 to 104, as follows:

[0075] Step 101: constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters; the target recognition weight parameters are obtained based on training of static image data;

[0076] In this embodiment, the step of constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters, includes:

[0077] Build a target recognition model based on the YOLOv11 algorithm, obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters;

[0078] The target recognition model includes a first backbone network, a feature fusion network and a detection head network;

[0079] The first backbone network includes a CBS module, an MNv4 module, an SPPF module and a C2PSA_2 module;

[0080] The feature fusion network includes 5 Fusion modules;

[0081] The detection head network includes a deformable decoder and a prediction head.

[0082] Please refer to Figure 2 , Figure 2 A schematic diagram of the structure of the target recognition model provided by an embodiment of the present invention.

[0083] In this embodiment, a target recognition model is constructed based on YOLOv11. The network after the first CBS module and before the SPPF module in the YOLOv11 backbone network is completely replaced with MobileNetV4, and an EMA attention mechanism is added after the FIB module and after the ConvNext module. The Neck module of YOLOv11 is replaced as a whole with the Efficient RepGFPN network to improve the feature representation capability of infrared targets. The DQ-DETR Decoder is used to replace the detection head of YOLOv11 to complete the construction of the target recognition model.

[0084] In this embodiment, Efficient RepGFPN efficiently integrates multi-scale features, using different numbers of channels for features at different scales. This allows for flexible control of the expressive power of both high- and low-level features while maintaining a low computational footprint. Flexible configuration optimizes accuracy and speed while approximating FLOPS and latency. By eliminating upsampling operations, accuracy is not significantly compromised while reducing computational overhead.

[0085] In this embodiment, in the MNv4 module, two layers of FIB (Fused Inverted Bottleneck) modules, EMA attention mechanism modules, Extra DW (Extra DepthWise) modules, four layers of IB (MobileNet Inverted Bottleneck) modules, one layer of ConvNext module, EMA attention mechanism modules, two layers of Extra DW modules, and four layers of IB modules are connected. The feature information extracted by the EMA (Efficient Multi-Scale Attention) module and the C2PSA_2 module is input into the Neck model for feature information fusion, and the feature information of different sizes obtained is input into the DQ-DETRDecoder detection head for target recognition and detection.

[0086] In this embodiment, the first backbone network in the target identification module includes CBS→MNv4 (including FIB×2→EMA→Extra DW→IB×4→ConvNext→EMA→Extra DW×2→IB×4)→SPPF→C2PSA_2.

[0087] Please refer to Figure 3 , Figure 3 A structural diagram of a CBS module provided in an embodiment of the present invention.

[0088] In this embodiment, the CBS module includes a normal convolution layer, a Batch Normalization layer, a ReLU activation layer, and a maximum pooling layer.

[0089] Please refer to Figure 4 , Figure 4 A structural diagram of a FIB module provided in an embodiment of the present invention.

[0090] In this embodiment, the FIB module is composed of a stack of a normal convolution layer, a Batch Normalization layer, an activation function layer, and a separable point convolution layer.

[0091] In this embodiment, the Extra DW module, ConvNext module and IB module are all instances of the UIB module under different configuration parameters. UIB is composed of an optional separable depth convolution layer (Optional DepthWise Convolution), a separable point convolution layer (PointWise Convolution), an optional separable depth convolution layer, and a separable point convolution layer. For the sake of distinction, the former point convolution layer is recorded as the point expansion layer, and the latter point convolution layer is recorded as the point projection layer. ConvNext only enables the first optional depth convolution layer, and by performing spatial mixing before the point expansion layer, it allows the use of a larger kernel size to achieve more cost-effective spatial mixing. IB only enables the second optional depth convolution layer, and performs spatial mixing on the point expansion layer features to improve the model capacity. The Extra DW module enables two optional depth convolution layers, which increases the network depth and receptive field at a lower computational cost, taking into account the advantages of both ConvNext and IB modules.

[0092] Please refer to Figure 5 , Figure 5 A structural diagram of the C2PSA_2 module provided in an embodiment of the present invention.

[0093] In this embodiment, the C2PSA_2 module is composed of a convolutional layer, several PSABlocks stacked and connected with residuals.

[0094] Please refer to Figure 6 , Figure 6 A structural diagram of a Fusion module provided in an embodiment of the present invention.

[0095] In this embodiment, the feature fusion network is composed of five stacked feature fusion modules. These modules fuse feature maps of different scales across different channel dimensions, avoiding additional upsampling operations and reducing computational complexity. The fusion module consists of multiple 1×1 convolutional layers, N Simplify Rep modules, and a Concat module.

[0096] Please refer to Figure 7 , Figure 7 A schematic diagram of the structure of the Simplify Rep module provided in an embodiment of the present invention.

[0097] In this embodiment, the Simplify Rep module is divided into two modes: Train and Infer. In Train mode, it consists of a 3×3 convolution, a 1×1 convolution, a Batch Normalization layer, and an activation function, while in Infer mode, it consists only of a 3×3 convolution and an activation function.

[0098] In this embodiment, the detection head network is a DQ-DETR Decoder detection head, which includes a deformable decoder and a prediction head.

[0099] In this embodiment, the original backbone network, feature fusion network, and detection head network are modularly replaced. CBS, MNv4, SPPF, and C2PSA_2 modules are used to form a lightweight backbone network with multi-scale perception capabilities. Five Fusion modules are used to achieve efficient feature fusion. A deformable decoder and end-to-end prediction head replace the traditional NMS post-processing process. This not only ensures the accuracy of small target detection in low-resolution infrared images, but also effectively reduces the model's computational complexity, meeting the dual requirements of power consumption and real-time performance for embedded deployment.

[0100] In this embodiment, the target recognition weight parameters are obtained based on static image data training, including:

[0101] Generating static image data based on the collection of multi-scene static images by the UAV, and annotating the static image data to obtain an infrared target data set;

[0102] The target recognition model is trained based on the infrared target data set until a preset condition is met, and a target recognition weight parameter is output.

[0103] In this embodiment, an infrared camera is mounted on a drone to collect static image data of the target of interest from multiple angles, multiple fields of view, and multiple scenes. The open source image data annotation tool X-AnyLabeling is used to annotate the collected data and construct an infrared target dataset.

[0104] In this embodiment, the collected moral static images can also be input into a pre-trained target recognition model to perform semi-automatic annotation on the collected data to construct an infrared target data set.

[0105] In this embodiment, several open-source authoritative infrared target datasets can be collected on the Internet, and appropriate ones can be selected and merged with the infrared target dataset collected based on drones to construct a richer and more generalized dataset.

[0106] In this embodiment, the fused network is trained using the constructed infrared target recognition training data set, and after obtaining the current model weights, the network performance is tested using the test data set, and the model weights with better test results are continuously saved until the optimal model weight parameter file is obtained.

[0107] In this embodiment, the infrared target dataset is preprocessed to reduce the image size to 640×640×3, and the infrared target dataset is divided into a training set and a test set in a ratio of 8:2.

[0108] In this embodiment, the target recognition model is tested using a training set and a test set to obtain the optimal target recognition weight parameters.

[0109] In this example, a specialized dataset covering complex backgrounds and multiple viewpoints was constructed by collecting multi-scene static infrared images using drones and combining them with semi-automated annotation. This dataset was used to iteratively train a recognition model integrating YOLOv11 and MobileNetV4, enabling the model to fully learn the thermal signature distribution of infrared targets in different scenarios. This process not only improves the diversity of training samples and the quality of annotations, but also ensures the generalization and detection accuracy of the recognition model in real-world scenarios by selecting optimal weights based on pre-defined conditions.

[0110] Step 102: constructing a target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters; the target tracking weight parameters are obtained based on training of dynamic video data;

[0111] In this embodiment, the step of constructing a target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters, includes:

[0112] Building a target tracking model based on the TCTrack algorithm; and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter;

[0113] The target tracking model includes a second backbone network, an ATT-TAda module, an AT-Trans module and a prediction head;

[0114] The second backbone network is a MobileNetV4 module, which includes a convolution module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling module;

[0115] The ATT-TAda module includes a maximum pooling layer, a cross attention module, an average pooling layer and a convolutional layer;

[0116] The AT-Trans module includes five multi-head attention mechanism modules and a temporal information filter module;

[0117] The prediction head includes a classification module and a regression module arranged in parallel.

[0118] Please refer to Figure 8 , Figure 8 A schematic diagram of the structure of a target tracking model provided by an embodiment of the present invention.

[0119] In this embodiment, the target tracking model includes a second backbone network based on MobileNetV4 for feature extraction, an ATT-TAda module for contextual information fusion of the extracted features, and an AT-Trans module for adaptive temporal refinement of the obtained similarity graph.

[0120] Please refer to Figure 9 , Figure 9 A schematic structural diagram of a second backbone network provided in an embodiment of the present invention.

[0121] In this embodiment, the second backbone network is composed of a convolutional layer Conv2D module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling AvgPool module.

[0122] In this embodiment, in the original MobileNetV4 network, the last two traditional convolutional layers are replaced by two ATT-TAda modules.

[0123] Please refer to Figure 10 , Figure 10 A structural diagram of the ATT-TAda module provided in an embodiment of the present invention.

[0124] In this embodiment, the ATT-TAda module consists of a maximum pooling layer (Max Pool), a cross-attention layer (Cross-attention), an average pooling layer (Avg Pool), and a convolutional layer. Compared with the traditional convolutional layer, the cross-attention layer incorporates the contextual information of the sequence, and uses this information to correct the convolutional weights, making the features more representative of the current target and enhancing robustness and generalization capabilities.

[0125] Please refer to Figure 11 , Figure 11 A schematic structural diagram of the AT-Trans module provided in an embodiment of the present invention.

[0126] In this embodiment, the AT-Trans module consists of five multi-head attention mechanism modules, a temporal information filter module, etc.

[0127] In this embodiment, the final prediction part includes a classification branch and a regression branch connected in parallel to predict the position of the target.

[0128] In this embodiment, To represent the features of the t-th frame proposed by the second backbone network, the output after the ATT-TAda module is It can be expressed as:

[0129] X t =W t *X t +b t (1)

[0130] Where * represents the convolution operation, X t and b t Represents the weight and bias parameters of the ATT-TAda module.

[0131] In this embodiment, unlike the traditional standard convolution, the learnable parameters X t and b t Will be adjusted based on the time context information:

[0132]

[0133] in, and Represents the parameter correction factor of the ATT-TAda module. With a small correction factor, the ATT-TAda module can efficiently integrate temporal context information.

[0134] In this embodiment, in order to avoid storing the feature information of historical frames, improve the utilization efficiency of memory, and reduce the uncontrollable influence of hyperparameters, the ATT-TAda module uses a fixed-size time information To accumulate historical features, and use adaptive max pooling to reduce feature parameters and convolution to reduce the number of feature channels, the cross attention operation used can be expressed as:

[0135] X′ t =MaxPool(X t ) (3)

[0136]

[0137] in, Represents a convolution operation.

[0138] In this embodiment, in order to improve efficiency, adaptive global average pooling (GAP) is used to process time information, and in order to ensure the parameter correction factor To improve the effectiveness of the proposed method, we introduce the residual connection structure. and It can be expressed as:

[0139] X t =GAP(Xt * ) (5)

[0140]

[0141] When the first frame is initialized, the weights can be learned and are initialized to 0, then W t =W b ,b t =bb. Because the first frame does not have feature data of historical frames, convolution is used to initialize the temporal context information

[0142] In this embodiment, Represents the output of the second backbone network (backbone network) after considering the time context in the feature extraction process, and the similarity graph F at the tth frame t It can be expressed as:

[0143]

[0144] Among them, Z represents the established target template and ★ represents depth-related operations.

[0145] In this embodiment, the AT-Trans module is used to combine the temporal context information to analyze the similarity graph F. t To refine the image, the AT-Trans module uses an encoder-decoder structure. In the encoder stage, the temporal context information is integrated, and in the decoder stage, the similarity graph is refined and positioned based on the temporal context information. In the encoder-decoder structure, a multi-head attention mechanism is used:

[0146]

[0147] In this embodiment, in the encoder part of the AT-Trans module, two multi-head attention mechanism modules are used before the input time information filter. represents the previous time prior information, F t Represents the similarity graph of the current frame, the output of the two multi-head attention modules of the t-th frame It can be expressed as:

[0148]

[0149] In this embodiment, After being processed by the time information filter, the final time prior knowledge is obtained:

[0150]

[0151] In this embodiment, in the decoder part, the final output It can be expressed as:

[0152]

[0153] In this embodiment, the loss function combines two classification loss functions and two position loss functions:

[0154]

[0155] In this embodiment,

[0156] is the main classification loss, using the cross-entropy loss function to measure the classification output and the true label The difference between them is used to strengthen the discrimination ability of the main classification head, so that the model can accurately identify the target category.

[0157] In this embodiment, It is an auxiliary classification loss that uses the binary cross-entropy loss function to train the model's binary judgment ability on "whether it is a target" to improve the model's ability to distinguish foreground / background or main target / interference target and enhance robustness.

[0158] In this embodiment, It is a positioning loss based on the normalized Euclidean distance, taking into account the deviation between the predicted center point and the true target center point in terms of width and height ratio. Where D represents the distance between the predicted position and the true position, M is a mask used to select the valid position, and The model is encouraged to accurately predict the specific location of the target in the image, especially the precise location of the center point.

[0159] In this embodiment, This loss term is based on IoU (Intersection over Union) and directly measures the degree of overlap between the predicted bounding box and the ground-truth box. It is used to optimize the fit of the target bounding box and improve the model's accurate perception of the target size and boundaries.

[0160] Step 103: Acquire image data to be detected, identify the image data to be detected based on the target recognition model, and obtain a candidate target data set;

[0161] In this embodiment,

[0162] Please refer to Figure 12 , Figure 12 Another flowchart of a target recognition and tracking algorithm provided by an embodiment of the present invention.

[0163] In this embodiment, the recognition module in the drone is deployed based on the improved Y0L0v11 and D0-DETR fusion network, that is, the target recognition model; the tracking module in the drone is deployed based on the improved TCTrack tracking algorithm model, that is, the target tracking model.

[0164] Step 104: Determine a target of interest in the candidate target data set, track the target of interest based on the target tracking model, and complete target recognition and tracking.

[0165] In this embodiment, the tracking module in the drone is deployed based on an improved TCTrack tracking algorithm model, namely a target tracking model. Thus, the recognition module in the drone captures the target, and then the tracking module stably tracks the target.

[0166] In this embodiment, the target recognition model processes image data, while the target tracking model processes continuous video data. The target recognition model detects several targets, while the target tracking model tracks the selected target of interest.

[0167] In one embodiment, after a natural disaster occurs, such as an earthquake, flood, or landslide, roads are often blocked, communications are interrupted, and manual search and rescue becomes difficult. UAVs carry infrared imaging equipment for post-disaster search and rescue, and can detect trapped people or vehicles on the ground in real time, especially in low light or bad weather conditions, as infrared imaging can penetrate obstacles such as smoke, dust, and haze. The target recognition method based on YOLOv11 and DQ-DETR and the target tracking method based on TCTrack can achieve efficient target detection and stable tracking of trapped people or vehicles on a fast-moving UAV platform. Especially in complex environments, the ability to accurately identify trapped people or moving targets and track their locations helps to monitor and dispatch rescue forces in real time, thereby improving the efficiency of post-disaster search and rescue.

[0168] In one embodiment, in border areas, especially in remote mountainous areas, forests, and desert areas, illegal border crossings and smuggling activities are often difficult to be detected in a timely manner by traditional monitoring systems. By using drones equipped with infrared cameras for border patrols, infrared heat sources can be monitored in real time. In particular, at night or in environments with poor visibility, infrared images can effectively capture abnormal activities. The present invention can efficiently identify and track different types of ground infrared targets, including illegal personnel, vehicles, animals, and other targets. In complex environments, through multi-perspective image data acquisition and spatiotemporal context fusion, the system can accurately track and mark targets, provide accurate real-time data support, and effectively improve the accuracy and efficiency of border patrols.

[0169] In one embodiment, in the intelligent security system of modern cities, the use of drones to patrol key areas can help monitor possible safety hazards in real time. Especially in crowded areas or at night, traditional video surveillance may cause difficulty in target recognition due to factors such as lighting and angle. The present invention uses drones equipped with infrared imaging equipment, combined with YOLOv11 and DQ-DETR networks, to achieve rapid recognition of abnormal behaviors (such as intrusions, illegal gatherings, etc.). In complex dynamic environments, the system can detect and track multiple targets in real time, and the TCTrack algorithm ensures that the target is not lost in a fast-moving environment, greatly enhancing the stability and response speed of the security monitoring system.

[0170] In one embodiment, in the perception system of an unmanned vehicle, infrared sensors are generally used to detect pedestrians, obstacles, and other potentially dangerous targets while driving at night. In particular, when visibility is limited, infrared images can provide more reliable information than visible light images. The present invention integrates the target recognition algorithms of YOLOv11 and DQ-DETR, combined with TCTrack target tracking technology, to achieve real-time detection and stable tracking of dynamic targets (such as pedestrians, other vehicles, etc.). Especially in complex traffic environments, it can ensure accurate identification and stable tracking of targets, thereby improving the safety and reliability of unmanned vehicles.

[0171] In one embodiment, in modern military reconnaissance, drones are often used for aerial reconnaissance to collect dynamic information about the enemy. Equipped with infrared imaging equipment, drones can conduct precise surveillance at night or in complex weather conditions. The present invention enables rapid and efficient target recognition and tracking, accurately identifying and tracking moving targets (such as enemy soldiers and equipment) during reconnaissance near enemy positions. Particularly in complex and dynamic environments, the system can avoid missed targets and mistracking, thereby improving the accuracy and effectiveness of reconnaissance.

[0172] The embodiment of the present invention further provides a target recognition and tracking device, comprising: a recognition model initialization module, a tracking model initialization module, a target recognition module and a target tracking module;

[0173] The recognition model initialization module is used to construct an initialized target recognition model and obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters; the target recognition weight parameters are obtained based on static image data training;

[0174] The tracking model initialization module is used to build a target tracking model and obtain target tracking weight parameters, and initialize the target tracking model based on the target tracking weight parameters; the target tracking weight parameters are obtained based on dynamic video data training;

[0175] The target recognition module is used to obtain image data to be detected, identify the image data to be detected based on the target recognition model, and obtain a candidate target data set;

[0176] The target tracking module is used to determine the target of interest in the candidate target data set, track the target of interest based on the target tracking model, and complete target recognition and tracking.

[0177] In this embodiment, the step of constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters, includes:

[0178] Build a target recognition model based on the YOLOv11 algorithm, obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters;

[0179] The target recognition model includes a first backbone network, a feature fusion network and a detection head network;

[0180] The first backbone network includes a CBS module, an MNv4 module, an SPPF module and a C2PSA_2 module;

[0181] The feature fusion network includes 5 Fusion modules;

[0182] The detection head network includes a deformable decoder and a prediction head.

[0183] In this embodiment, the target recognition weight parameters are obtained based on static image data training, including:

[0184] Generating static image data based on the collection of multi-scene static images by the UAV, and annotating the static image data to obtain an infrared target data set;

[0185] The target recognition model is trained based on the infrared target data set until a preset condition is met, and a target recognition weight parameter is output.

[0186] In this embodiment, the step of constructing a target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters, includes:

[0187] Building a target tracking model based on the TCTrack algorithm; and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter;

[0188] The target tracking model includes a second backbone network, an ATT-TAda module, an AT-Trans module and a prediction head;

[0189] The second backbone network is a MobileNetV4 module, which includes a convolution module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling module;

[0190] The ATT-TAda module includes a maximum pooling layer, a cross attention module, an average pooling layer and a convolutional layer;

[0191] The AT-Trans module consists of five multi-head attention mechanism modules and a temporal information filter module.

[0192] In this embodiment, the target tracking weight parameters are obtained based on training of a dynamic video dataset and include:

[0193] Collecting continuous video data of multiple scenes based on a drone, performing sequence tracking on the continuous video data of multiple scenes to generate an infrared tracking dataset;

[0194] The target tracking model is trained based on the infrared tracking data set until a preset condition is met, and a target tracking weight parameter is output.

[0195] In an embodiment of the present invention, a multi-device access platform processing device is also provided, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the above-mentioned multi-device access platform processing method is implemented.

[0196] In an embodiment of the present invention, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored computer program, wherein when the computer program runs, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned multi-device access platform processing method.

[0197] For example, a computer program may be divided into one or more modules, one or more of which are stored in a memory and executed by a processor to implement the present invention. One or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the multi-device access platform processing device.

[0198] The multi-device access platform processing device can be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The multi-device access platform processing device may include, but is not limited to, a processor, a memory, and a display. Those skilled in the art will appreciate that the aforementioned components are merely examples of the multi-device access platform processing device and do not constitute a limitation of the multi-device access platform processing device. The multi-device access platform processing device may include more or fewer components, or a combination of certain components, or different components. For example, the multi-device access platform processing device may also include input and output devices, network access devices, buses, and the like.

[0199] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the multi-device access platform processing device, and uses various interfaces and lines to connect various parts of the multi-device access platform processing device.

[0200] The memory can be used to store computer programs and / or modules. The processor implements various functions of the multi-device access platform processing device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function (such as a sound playback function, a text conversion function, etc.); the data storage area can store data created based on the use of the mobile phone (such as audio data, text message data, etc.). In addition, the memory can include a high-speed random access memory and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0201] Among them, if the module based on the processing of the multi-device access platform is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. Those skilled in the art can understand and implement it without paying any creative work.

[0202] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.

Claims

1. A target recognition and tracking algorithm, characterized in that: include: Constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters; The target recognition weight parameters are obtained based on static image data training; Constructing a target tracking model and obtaining target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters; the target tracking weight parameters are obtained based on training of dynamic video data; Acquire image data to be detected, identify the image data to be detected based on the target recognition model, and obtain a candidate target data set; An interested target is determined in the candidate target data set, and the interested target is tracked based on the target tracking model to complete target recognition and tracking.

2. A target recognition and tracking algorithm as claimed in claim 1, characterized in that: The constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters, includes: Build a target recognition model based on the YOLOv11 algorithm, obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters; The target recognition model includes a first backbone network, a feature fusion network and a detection head network; The first backbone network includes a CBS module, an MNv4 module, an SPPF module and a C2PSA_2 module; The feature fusion network includes 5 Fusion modules; The detection head network includes a deformable decoder and a prediction head.

3. A target recognition and tracking algorithm as claimed in claim 2, characterized in that: The target recognition weight parameters are obtained based on static image data training, and include: Generating static image data based on the collection of multi-scene static images by the UAV, and annotating the static image data to obtain an infrared target data set; The target recognition model is trained based on the infrared target data set until a preset condition is met, and a target recognition weight parameter is output.

4. A target recognition and tracking algorithm as claimed in claim 1, characterized in that: The constructing of the target tracking model and obtaining the target tracking weight parameters, and initializing the target tracking model based on the target tracking weight parameters comprises: Building a target tracking model based on the TCTrack algorithm; and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter; The target tracking model includes a second backbone network, an ATT-TAda module, an AT-Trans module and a prediction head; The second backbone network is a MobileNetV4 module, which includes a convolution module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling module; The ATT-TAda module includes a maximum pooling layer, a cross attention module, an average pooling layer and a convolutional layer; The AT-Trans module includes five multi-head attention mechanism modules and a temporal information filter module; The prediction head includes a classification module and a regression module arranged in parallel.

5. A target recognition and tracking algorithm as claimed in claim 4, characterized in that: The target tracking weight parameters are obtained based on training of a dynamic video dataset and include: Collecting continuous video data of multiple scenes based on a drone, performing sequence tracking on the continuous video data of multiple scenes to generate an infrared tracking dataset; The target tracking model is trained based on the infrared tracking data set until a preset condition is met, and a target tracking weight parameter is output.

6. A target recognition and tracking device, characterized in that: include: Recognition model initialization module, tracking model initialization module, target recognition module and target tracking module; The recognition model initialization module is used to construct an initialized target recognition model and obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters; the target recognition weight parameters are obtained based on static image data training; The tracking model initialization module is used to build a target tracking model and obtain target tracking weight parameters, and initialize the target tracking model based on the target tracking weight parameters; the target tracking weight parameters are obtained based on dynamic video data training; The target recognition module is used to obtain image data to be detected, identify the image data to be detected based on the target recognition model, and obtain a candidate target data set; The target tracking module is used to determine the target of interest in the candidate target data set, track the target of interest based on the target tracking model, and complete target recognition and tracking.

7. A target recognition and tracking device according to claim 6, characterized in that: The constructing an initialized target recognition model and obtaining target recognition weight parameters, and initializing the target recognition model based on the target recognition weight parameters, includes: Build a target recognition model based on the YOLOv11 algorithm, obtain target recognition weight parameters, and initialize the target recognition model based on the target recognition weight parameters; The target recognition model includes a first backbone network, a feature fusion network and a detection head network; The first backbone network includes a CBS module, an MNv4 module, an SPPF module and a C2PSA_2 module; The feature fusion network includes 5 Fusion modules; The detection head network includes a deformable decoder and a prediction head.

8. A target recognition and tracking device according to claim 7, characterized in that: The target recognition weight parameters are obtained based on static image data training, and include: Generating static image data based on the collection of multi-scene static images by the UAV, and annotating the static image data to obtain an infrared target data set; The target recognition model is trained based on the infrared target data set until a preset condition is met, and a target recognition weight parameter is output.

9. A target recognition and tracking device according to claim 1, characterized in that: The target tracking model is constructed and the target tracking weight parameter is obtained, and the target tracking model is initialized based on the target tracking weight parameter, including: Building a target tracking model based on the TCTrack algorithm; and obtaining a target tracking weight parameter, and initializing the target tracking model based on the target tracking weight parameter; The target tracking model includes a second backbone network, an ATT-TAda module, an AT-Trans module and a prediction head; The second backbone network is a MobileNetV4 module, which includes a convolution module, a FusedIB module, an Extra DW module, an IB module, a ConvNext module, and an average pooling module; The ATT-TAda module includes a maximum pooling layer, a cross attention module, an average pooling layer and a convolutional layer; The AT-Trans module includes five multi-head attention mechanism modules and a temporal information filter module; The prediction head includes a classification module and a regression module arranged in parallel.

10. The target recognition and tracking device according to claim 9, wherein: The target tracking weight parameters are obtained based on training of a dynamic video dataset and include: Collecting continuous video data of multiple scenes based on a drone, performing sequence tracking on the continuous video data of multiple scenes to generate an infrared tracking dataset; The target tracking model is trained based on the infrared tracking data set until a preset condition is met, and a target tracking weight parameter is output.