A Low-Altitude Intelligent Target Tracking Method Based on Multimodal Fusion

By employing a multimodal fusion-based low-altitude intelligent target tracking method, which utilizes visual images, infrared images, and radar echo data, combined with an improved TimeSformer model and a fusion controller, the problems of occlusion and identity interruption in low-altitude target tracking are solved. This achieves the continuity of target trajectory and the stability of identity, thereby improving the intelligence level of the tracking system.

CN121305220BActive Publication Date: 2026-04-03ANHUI FALCON WAVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing low-altitude target tracking methods are prone to identity interruption and trajectory jumps when subjected to dynamic occlusion, environmental interference, and changes in target shape. Furthermore, they lack effective fusion and adjustment mechanisms, resulting in decreased tracking continuity and identity stability, low feature matching accuracy, and severe modal interference.

Method used

A multimodal fusion low-altitude intelligent target tracking method is adopted, which utilizes visual images, infrared images and radar echo data, combined with an improved TimeSformer model and fusion controller. Through modal decoupling, rhythm perception and memory injection, structural continuity modeling is carried out to construct a consistency scoring matrix and confidence adjustment structure, so as to realize identity recovery and trajectory continuation after occlusion.

Benefits of technology

It improves the stability and accuracy of target tracking in complex low-altitude environments. Through modal enhancement and structure preservation, it achieves identity recovery and trajectory continuity after occlusion, ensuring the intelligence and robustness of the tracking system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305220B_ABST
    Figure CN121305220B_ABST
Patent Text Reader

Abstract

This invention discloses a low-altitude intelligent target tracking method based on multimodal fusion, comprising the following steps: S1, constructing a multimodal input sequence to generate a fused input dataset; S2, inputting a semantic structure feature map into an improved TimeSformer model to output a predicted structure map; S3, constructing a consistency scoring matrix; S4, extracting candidate target samples whose trajectory distances meet a set similarity threshold and generating anchor point feature vectors; S5, extracting target feature vectors and historical identity feature vectors to generate an identity similarity score; S6, performing identity regression judgment based on the identity similarity score; S7, recording the fusion output result and identity label status of each frame to generate a target tracking trajectory sequence and an identity record set. This invention achieves accurate tracking under complex occlusion and multimodal interference conditions in low-altitude scenes, possessing advantages such as high tracking robustness, strong identity recovery accuracy, and excellent anomaly response time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target tracking technology, and in particular to a low-altitude intelligent target tracking method based on multimodal fusion. Background Technology

[0002] With the increasing application of intelligent sensing tasks in low-altitude scenarios, target tracking technology in complex environments is gradually becoming a core component of unmanned systems, autonomous inspection, and security monitoring. Existing low-altitude target tracking methods mostly rely on temporal visual information or infrared thermal imaging data to construct target trajectories, combining appearance or positional information between consecutive frames for target status maintenance and identification. However, in practical applications with frequent dynamic occlusion, variable environmental interference, and drastic changes in target morphology, existing technologies still have the following problems:

[0003] During occlusion, targets are prone to identity interruption, trajectory jumps, or target loss, leading to a significant decrease in tracking continuity and identity stability. Especially when a target briefly disappears and reappears, existing methods often struggle to accurately trace back its historical identity. For abnormal changes in target state and occlusion events, existing methods lack effective judgment mechanisms and structured recording methods, making it impossible to perform trajectory repair and tracking reconstruction in a timely manner. At the same time, feature matching used for identity confirmation during tracking is mostly based on single-frame images or single modalities, ignoring the temporal feature continuity and structural change information of the target, resulting in low matching accuracy and a high false recognition rate.

[0004] When the target state changes drastically or the confidence level fluctuates wildly, traditional fusion strategies lag in adjusting modal contributions, easily leading to information redundancy or modal interference, which affects the accuracy of subsequent predicted structures and the reliability of fusion trajectories. Existing methods have not yet established an effective fusion adjustment mechanism to cope with the changing trends of dynamic modal contributions.

[0005] Therefore, how to provide a low-altitude intelligent target tracking method based on multimodal fusion is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose a low-altitude intelligent target tracking method based on multimodal fusion. This invention fully utilizes multimodal perception data from visual images, infrared images, and radar echoes, and constructs a fusion input dataset using a unified time index. It details the entire process of fusion feature extraction, structure prediction, consistency score calculation, confidence adjustment, anchor point construction, identity regression judgment, and tracking state generation. This invention employs an improved TimeSformer model to model the semantic structure feature map, constructing a multimodal decoupling unit, a rhythm-aware attention module, and a memory state injection unit to achieve efficient fusion of modality enhancement, structure preservation, and temporal awareness. Furthermore, it proposes a confidence adjustment structure based on a consistency score matrix to determine the prediction weighting period and activate the anchor point path mechanism. A fusion controller is constructed to execute an identity confirmation frame mechanism based on multi-frame features, achieving identity recovery and trajectory continuation after occlusion. During the tracking period, the fusion state and identity state are recorded, outputting a target tracking trajectory sequence containing the identity record set. An anomaly monitoring and reset mechanism is also constructed to ensure tracking stability.

[0007] A low-altitude intelligent target tracking method based on multimodal fusion according to an embodiment of the present invention includes the following steps:

[0008] S1. Collect multimodal perception data of the target, construct a multimodal input sequence, and generate a fused input dataset;

[0009] S2. Based on the fused input dataset, perform initial target detection and semantic structure extraction to obtain the semantic structure feature map and modality confidence distribution of the target; input the semantic structure feature map into the improved TimeSformer model to output the predicted structure map;

[0010] S3. Perform a matching operation between the predicted structure graph and the multimodal features extracted at the same time to construct a consistency score matrix; calculate the global consistency score based on the consistency score matrix; if the global consistency score is less than the first preset threshold, perform a modal confidence suppression operation on the fusion scaling factor corresponding to the predicted structure graph, and mark the corresponding time as the prediction weight reduction period.

[0011] S4. During the prediction weight reduction period, extract candidate target samples whose trajectory distance meets the set similarity threshold from the target's neighborhood, perform feature clustering operation, and generate anchor point feature vectors.

[0012] S5. Determine whether the occlusion state of the target has ended. If the occlusion state has ended and the target is detected to reappear in the perception area, then perform the identity verification operation, extract the target feature vector at the same time and match it with the historical identity feature vector to generate an identity similarity score.

[0013] S6. Perform identity regression judgment based on identity similarity score and second and third preset thresholds; if identity similarity score is greater than second preset threshold and less than third preset threshold, trigger identity confirmation frame mechanism, continuously extract target features from multiple frames for dense matching operation; if identity similarity score is greater than third preset threshold, or identity confirmation frame mechanism outputs matching confirmation result, restore the original identity trajectory of the target and terminate the target anchor point path.

[0014] S7. Record the fusion output result and identity tag status of each frame during the tracking period to generate the target tracking trajectory sequence and identity record set.

[0015] Preferably, S1 specifically comprises:

[0016] Visual images, infrared images, and radar echo data are acquired to construct a multimodal input sequence containing multimodal sensing data; frame numbering is uniformly processed on the multimodal input sequence, and a global sampling time interval and time index order are set;

[0017] The data from each modal sensing mode are interpolated, truncated, or filled according to the global time index until the number of frames in each modal sequence is consistent.

[0018] Based on a unified time index, data from visual images, infrared images, and radar echo data at corresponding time frames are combined into a fusion frame group. Spatial coordinate projection transformation is performed on each modal data in the fusion frame group to construct a spatial alignment result under a unified spatial reference. Based on the spatial alignment result, the image information of each modality is combined to generate a fusion input dataset.

[0019] Preferably, the step of performing initial target detection and semantic structure extraction based on the fused input dataset to obtain the semantic structure feature map and modality confidence distribution of the target specifically involves:

[0020] Perform feature extraction on the multimodal image data of each frame in the fused input dataset to extract visual image feature vectors, infrared image feature vectors, and radar echo feature vectors;

[0021] The visual image feature vector, infrared image feature vector, and radar echo feature vector are input into a unified feature fusion network to generate a fused feature map.

[0022] Based on the fused feature map, candidate region generation is performed to identify potential target regions and perform region cropping; target bounding box fitting and semantic mask segmentation are performed on the cropped regions to generate initial target detection results and semantic structure map;

[0023] Multimodal fusion features are extracted for each detected target in the semantic structure graph, the contribution weight of each modality is calculated, and the modality confidence distribution is generated.

[0024] Preferably, the improved TimeSformer model includes a multimodal decoupling unit, a rhythm-aware attention module, a temporal modeling structure, a memory state injection unit, and a fusion controller, specifically:

[0025] A multimodal decoupling unit is constructed, and convolutional coding operations are performed on the semantic structure feature maps of visual images, infrared images, and radar echo images to extract edge change features, hot spot distribution features, and scattering morphology features. Channel alignment and dimension normalization operations are performed on the feature maps of the edge change features, hot spot distribution features, and scattering morphology features to generate modality-enhanced feature maps.

[0026] A rhythm-aware attention module is constructed. The confidence change trend, motion displacement amplitude and occlusion state index of the target in continuous historical frames are normalized to generate a rhythm state vector. The size of the inter-frame attention window is adjusted according to the rhythm state vector, and dynamic window encoding is performed to generate a rhythm-aware position vector.

[0027] A temporal modeling structure is constructed, and the rhythm-aware position vector and modality-enhanced feature map are input into the structural continuous attention path and the semantic global attention path. The structural continuous attention path generates short-term structural change vectors based on boundary region masks, and the semantic global attention path performs cross-frame semantic attention operations based on full-map features to generate semantic evolution feature vectors.

[0028] A memory state injection unit is constructed. Based on the predicted structure map of historical continuous frames, the contour boundary point set, region mask map and center drift vector are extracted. A structural memory representation tensor is constructed and residually fused with the output of the structural continuous attention path to generate memory-enhanced structural features.

[0029] A fusion controller is constructed to concatenate memory-enhanced structural features and semantic evolution feature vectors, and performs weight allocation and fusion operations based on modality confidence distribution, outputting a predicted structure map and a structure confidence score map.

[0030] Preferably, S3 specifically includes:

[0031] Extract visual image feature vectors, infrared image feature vectors, and radar echo feature vectors under the same time index to construct a time-aligned multimodal feature set; project the predicted structure map onto a unified spatial coordinate system, perform boundary segmentation operation on each structural region in the predicted structure map, and generate a set of structural region boundaries;

[0032] Feature matching operations are performed on the corresponding regions in the structural region boundary set and visual image feature vector, infrared image feature vector and radar echo feature vector, respectively, to generate visual modal similarity matrix, infrared modal similarity matrix and radar modal similarity matrix; weighted calculation is performed on the visual modal similarity matrix, infrared modal similarity matrix and radar modal similarity matrix according to preset modal fusion weights to generate consistency score matrix;

[0033] The average score of all structural regions in the consistency score matrix is ​​calculated to generate a global consistency score. If the global consistency score is less than a first preset threshold, the fusion scaling factor associated with the predicted structural graph is input to the confidence adjustment structure. The confidence adjustment structure includes a confidence calculation unit, a scaling factor adjustment unit, and a time-series recording unit. The confidence calculation unit receives the structural confidence value output by the consistency score matrix, the scaling factor adjustment unit performs a suppression operation on the fusion scaling factor according to the confidence value, and the time-series recording unit records the time index when the fusion scaling factor suppression occurs.

[0034] After the confidence adjustment structure completes the suppression operation, it marks the corresponding time index as the predicted deweighting period and records the deweighting state label in the tracking state sequence.

[0035] Preferably, S4 specifically includes:

[0036] Under the time index corresponding to the predicted weight reduction period, a spatial neighborhood region with a fixed radius is constructed based on the target position coordinates of the previous moment in the fused input dataset;

[0037] Visual images, infrared images, and radar echo data are extracted from the spatially adjacent region. Convolutional feature extraction operations are performed on each of these data to generate visual image feature vectors, infrared image feature vectors, and radar echo feature vectors, which together form a candidate target sample set.

[0038] For each target sample in the candidate target sample set, calculate the Euclidean distance and the direction angle between the historical trajectory endpoint position of the target sample and the current candidate sample position, and jointly generate a trajectory distance score.

[0039] Target samples with trajectory distance scores less than a set similarity threshold are selected, and their corresponding visual image feature vectors, infrared image feature vectors, and radar echo feature vectors are extracted to construct an aligned multimodal joint feature set.

[0040] Based on the multimodal joint feature set, cluster center calculation and feature point classification operations are performed to generate multiple feature clusters; the feature set with the highest cluster density in the feature cluster is used to calculate the center point feature vector, which is output as the anchor point feature vector; the time index, spatial coordinates and modal feature number corresponding to the anchor point feature vector are recorded, the anchor point path information sequence is constructed and stored in the tracking auxiliary cache structure.

[0041] Preferably, S5 specifically includes:

[0042] Perform target detection on visual images, infrared images and radar echo data in a multi-frame fused input dataset, and determine whether there is a target candidate region associated with the target prediction structure map in the target detection region; if no corresponding target candidate region is detected in several consecutive frames, mark the target as entering the occlusion state and record the occlusion start time index.

[0043] Continue performing the target detection operation. If a new target candidate region is detected, and the center coordinate position, boundary contour features and boundary information of the predicted structure map of the new target candidate region meet the preset recovery conditions, then the occlusion state is determined to end, and the occlusion end time index is recorded. Extract the visual image feature vector, infrared image feature vector and radar echo feature vector from the frame corresponding to the occlusion end time index, and construct a multimodal joint feature vector for the re-identified target.

[0044] Extract the target's identity feature cache set from historical records. The identity feature cache set includes a labeled multimodal fusion feature vector sequence, location index sequence, and identity tag number. Perform a similarity calculation operation on the multimodal joint feature vector of the re-identified target and the fusion feature vector in the identity feature cache set. Generate an identity similarity score based on the modality weighted cosine similarity function. Record the time index and target identification information corresponding to the identity similarity score, output the identity confirmation candidate tag, and enter the identity regression judgment process.

[0045] Preferably, S6 specifically includes:

[0046] After receiving the identity similarity score, a threshold range judgment operation is performed, comparing the identity similarity score with the second preset threshold and the third preset threshold.

[0047] If the identity similarity score is greater than the second preset threshold and less than the third preset threshold, then visual image feature vector, infrared image feature vector and radar echo feature vector are extracted from the continuous time index to construct a multi-frame joint feature set.

[0048] For each frame feature vector in the multi-frame joint feature set, a matching calculation is performed with the fused feature vector in the historical identity feature cache set. A multi-frame matching score matrix is ​​generated based on the modality weighted cosine similarity function. The average matching score of all frames in the multi-frame matching score matrix is ​​calculated to generate an average matching score. If the average matching score is greater than a third preset threshold, the target identity is determined to be consistent with the historical identity, and the identity matching confirmation result is output. If the identity similarity score is greater than the third preset threshold, or the identity matching confirmation result is determined to be consistent, the identity identifier of the current frame is associated with the historical identity label to restore the original identity trajectory sequence.

[0049] After the identity trajectory is restored, the anchor point path tracking operation associated with the target is terminated, the fusion scaling factor is updated to the initial state parameters, and the identity restoration label is recorded.

[0050] Preferably, S7 specifically includes:

[0051] In each time indexed frame, visual image feature vectors, infrared image feature vectors, and radar echo feature vectors are extracted to construct a fusion input feature set. The fusion input feature set is input to the fusion controller, and feature fusion operation is performed by combining the prediction structure map of the previous frame and the fusion scaling factor to output a fusion state vector. The center position coordinates, bounding box parameters, and modal confidence distribution in the fusion state vector are recorded to generate the fusion output result of the current frame.

[0052] Based on the identity similarity score, the identity confirmation frame mechanism result and the trajectory number mapping relationship, the identity number of the target in the current frame is confirmed; after confirming the identity label, the fusion output result is combined with the identity label number to construct the target tracking status entry; the target tracking status entry is stored in the tracking trajectory cache set in time index order to generate the target tracking trajectory sequence;

[0053] An identity record index structure is established for each target identifier, and identity tag status changes, anchor path records, identity recovery operations and trajectory jump markers are summarized to generate an identity record set; the target tracking trajectory sequence and the identity record set are used as the output results of the tracking cycle.

[0054] Preferred options also include:

[0055] Perform state analysis on the target tracking trajectory sequence and identity record set generated within the tracking period; statistically analyze the duration of continuous occlusion, identity switching frequency, modality confidence fluctuation amplitude and fusion ratio factor variation range to generate an anomaly monitoring index set;

[0056] Each indicator in the abnormal monitoring indicator set is compared with its corresponding abnormal judgment threshold. If the duration of continuous occlusion is greater than the preset occlusion threshold, or the identity switching frequency is greater than the preset frequency threshold, or the modal confidence fluctuation amplitude is greater than the preset fluctuation threshold, or the fusion ratio factor change range is greater than the preset range threshold, then the tracking abnormality mark is triggered.

[0057] After triggering the tracking anomaly marker, terminate the current target's fusion state vector update operation and write the anomaly state to the anomaly label record table; send a reset request to the anchor path cache structure, identity feature cache set and fusion controller to perform a state reset operation; record the time index and target number corresponding to the reset operation and build an anomaly recovery process log.

[0058] The beneficial effects of this invention are:

[0059] This invention addresses the problems of frequent target occlusion, identity misjudgment, and severe modal interference in low-altitude scenes through joint modeling with an improved TimeSformer model and a fusion controller. It employs a structural continuity modeling mechanism combining modal decoupling, rhythm awareness, and memory injection to improve the accuracy of the predicted structure graph's response to semantic evolution trends. Furthermore, this invention proposes a consistency scoring matrix and a confidence adjustment structure. It performs score fitting between the prediction result and the current multimodal features, driving dynamic suppression of the fusion scaling factor through score values ​​to achieve timely weight reduction of untrusted modal information and correction of tracking deviations. Regarding occlusion handling, this invention… This invention constructs an anchor point path assistance mechanism and an identity confirmation frame mechanism, combining trajectory distance scoring and modal weighted cosine similarity to complete identity recovery and trajectory reconnection operations, improving the accuracy of occlusion recovery. During the identity preservation phase, target tracking state entries are constructed by combining the fusion state vector and identity tag number, and full-cycle identity state management is achieved through an identity record set. Finally, this invention also constructs an anomaly monitoring index set, combining key indicators such as continuous occlusion duration, identity switching frequency, modal confidence fluctuation amplitude, and fusion ratio factor change range for anomaly determination, triggering fusion state reset and resetting mechanisms to ensure the continuous stability of the tracking state. Ultimately, this achieves intelligent closed-loop processing for full-cycle fusion tracking and anomaly recovery of targets in complex low-altitude environments. Attached Figure Description

[0060] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0061] Figure 1 This is a flowchart of a low-altitude intelligent target tracking method based on multimodal fusion proposed in this invention;

[0062] Figure 2This is a data flow diagram of a low-altitude intelligent target tracking method based on multimodal fusion proposed in this invention;

[0063] Figure 3 This is a structural diagram of the improved TimeSformer model proposed in this invention. Detailed Implementation

[0064] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0065] refer to Figure 1-3 A low-altitude intelligent target tracking method based on multimodal fusion includes the following steps:

[0066] S1. Collect multimodal perception data of the target, construct a multimodal input sequence, and generate a fused input dataset;

[0067] S2. Based on the fused input dataset, perform initial target detection and semantic structure extraction to obtain the semantic structure feature map and modality confidence distribution of the target; input the semantic structure feature map into the improved TimeSformer model to output the predicted structure map;

[0068] S3. Perform a matching operation between the predicted structure graph and the multimodal features extracted at the same time to construct a consistency score matrix; calculate the global consistency score based on the consistency score matrix; if the global consistency score is less than the first preset threshold, perform a modal confidence suppression operation on the fusion scaling factor corresponding to the predicted structure graph, and mark the corresponding time as the prediction weight reduction period.

[0069] S4. During the prediction weight reduction period, extract candidate target samples whose trajectory distance meets the set similarity threshold from the target's neighborhood, perform feature clustering operation, and generate anchor point feature vectors.

[0070] S5. Determine whether the occlusion state of the target has ended. If the occlusion state has ended and the target is detected to reappear in the perception area, then perform the identity verification operation, extract the target feature vector at the same time and match it with the historical identity feature vector to generate an identity similarity score.

[0071] S6. Perform identity regression judgment based on identity similarity score and second and third preset thresholds; if identity similarity score is greater than second preset threshold and less than third preset threshold, trigger identity confirmation frame mechanism, continuously extract target features from multiple frames for dense matching operation; if identity similarity score is greater than third preset threshold, or identity confirmation frame mechanism outputs matching confirmation result, restore the original identity trajectory of the target and terminate the target anchor point path.

[0072] S7. Record the fusion output result and identity tag status of each frame during the tracking period to generate the target tracking trajectory sequence and identity record set.

[0073] This implementation constructs a multimodal input sequence by acquiring visual images, infrared images, and radar echo data, and performs time synchronization and spatial alignment processing to generate a fused input dataset, effectively improving the spatiotemporal consistency and input quality of multimodal data fusion. Based on the fused input dataset, initial target detection and semantic structure extraction are performed to obtain semantic structure feature maps and modal confidence distributions, improving the accuracy of target recognition and the completeness of structural representation. An improved TimeSformer model is introduced to predict the semantic structure feature maps, outputting predicted structure maps, enabling forward-looking modeling of target structural change trends and enhancing the ability to express structural continuity under occlusion conditions. The predicted structure maps are matched with multimodal features extracted at the same time to construct a consistency scoring matrix and calculate a global consistency score. If the score is lower than a first preset threshold, modal confidence suppression is performed, marking the prediction weight reduction period, effectively improving the robustness of abnormal structure prediction. In one step, during the prediction weighting period, candidate target samples whose trajectory distance meets the set similarity threshold are extracted, and feature clustering is performed to generate anchor point feature vectors, enhancing the ability to retain temporary features of occluded targets. If the occlusion state is determined to have ended and the target is detected to have reappeared, an identity verification operation is performed, generating an identity similarity score based on the matching of the current frame and historical identity features, improving the accuracy of target identity recognition after occlusion. Further, identity regression judgment is performed based on the identity similarity score and the preset threshold. If the identity verification condition is met, the identity verification frame mechanism is triggered to perform multi-frame dense matching operations, or the original identity trajectory is directly restored and the anchor point path is terminated, improving the accuracy and efficiency of identity trajectory restoration. At the same time, the fusion output results and identity tag status are recorded during the tracking period to generate a target tracking trajectory sequence and identity record set, realizing continuous identity tracking and trajectory traceability management under multimodal fusion, and improving the stability and intelligence level of the overall tracking system.

[0074] In this embodiment, S1 specifically refers to:

[0075] Visual images, infrared images, and radar echo data are acquired to construct a multimodal input sequence containing multimodal sensing data; frame numbering is uniformly processed on the multimodal input sequence, and a global sampling time interval and time index order are set;

[0076] The data from each modal sensing mode are interpolated, truncated, or filled according to the global time index until the number of frames in each modal sequence is consistent.

[0077] Based on a unified time index, data from visual images, infrared images, and radar echo data at corresponding time frames are combined into a fusion frame group. Spatial coordinate projection transformation is performed on each modal data in the fusion frame group to construct a spatial alignment result under a unified spatial reference. Based on the spatial alignment result, the image information of each modality is combined to generate a fusion input dataset.

[0078] In this embodiment, the step of performing initial target detection and semantic structure extraction based on the fused input dataset to obtain the semantic structure feature map and modality confidence distribution of the target specifically involves:

[0079] Convolutional coding operations are performed on the visual image, infrared image and radar echo data of each frame in the fused input dataset to extract local edge features, hot spot response features and echo contour features, which are used to construct visual image feature vector, infrared image feature vector and radar echo feature vector respectively.

[0080] Visual image feature vectors, infrared image feature vectors, and radar echo feature vectors are input into a unified feature fusion network. The unified feature fusion network includes a multimodal feature alignment module, a channel enhancement module, and a spatial fusion module. The multimodal feature alignment module performs dimension matching and position mapping operations on different modal feature vectors. The channel enhancement module assigns higher weights to high-response channels based on a channel attention mechanism. The spatial fusion module outputs a fused feature map based on convolutional kernel fusion operations.

[0081] Candidate region generation is performed based on the fused feature map. An anchor box set is constructed and the classification score and regression offset are calculated for each anchor box. The boundary offset prediction vector is obtained by fitting the bounding box regression loss function. Target confidence ranking and non-maximum suppression are performed on the candidate regions to screen out potential target regions and perform region cropping.

[0082] Boundary box fitting and semantic mask segmentation are performed on the cropped image regions respectively. Boundary box fitting is based on the position offset vector to fine-tune the position of the original anchor box. Semantic mask segmentation is based on the semantic category label of each pixel output by the branch convolutional network to generate a semantic segmentation map and localization bounding box, which constitute the initial target detection result and semantic structure map.

[0083] For each multimodal region of the detected target in the semantic structure graph, a fusion feature extraction operation is performed. The mean response intensity of the visual image channel, the hot spot focusing degree of the infrared image, and the variance of the radar echo scattering amplitude are calculated respectively. The above features are used as modal perception indicators and input into the modal confidence calculation function. The modal confidence calculation function outputs the confidence value of each modality through the Softmax function to generate the modal confidence distribution.

[0084] In this embodiment, the improved TimeSformer model includes a multimodal decoupling unit, a rhythm-aware attention module, a temporal modeling structure, a memory state injection unit, and a fusion controller, specifically:

[0085] A multimodal decoupling unit is constructed, and convolutional coding operations are performed on the semantic structure feature maps of visual images, infrared images, and radar echo images to extract edge variation features, hotspot distribution features, and scattering morphology features. Channel alignment and dimension normalization operations are then performed on the edge variation features, hotspot distribution features, and scattering morphology features to construct a modal feature mapping matrix. The modal alignment parameter vector is obtained by fitting the modal difference loss function by minimizing the modal alignment parameter vector. The modal alignment parameter vector is then applied to the modal feature mapping matrix to generate a modal enhancement feature map.

[0086] A rhythm-aware attention module is constructed. Time normalization and amplitude smoothing operations are performed on the confidence change trend, motion displacement amplitude and occlusion state index of the target in continuous historical frames to generate a rhythm state vector. The time change rate and inter-frame variance are calculated based on the rhythm state vector, and the rhythm response curve is obtained by fitting it with a Gaussian kernel function. The size of the inter-frame attention window is dynamically adjusted according to the rhythm response curve, the rhythm-aware encoding operation is performed, and the rhythm-aware position vector is output.

[0087] A temporal modeling structure is constructed, and rhythm-aware position vectors and modality-enhanced feature maps are input into structural continuous attention paths and semantic global attention paths. The structural continuous attention path calculates boundary similarity weights based on boundary region mask maps and generates short-term structural change vectors through weighted summation. The semantic global attention path performs cross-frame attention calculations on the semantic features of the entire frame, calculates the inter-frame semantic similarity matrix, and generates semantic evolution feature vectors through matrix multiplication. The short-term structural change vectors and semantic evolution feature vectors are concatenated and fused to form a temporal joint feature representation.

[0088] A temporal modeling structure is constructed, and rhythm-aware position vectors and modality-enhanced feature maps are input into structural continuous attention paths and semantic global attention paths. The structural continuous attention path calculates boundary similarity weights based on boundary region mask maps and generates short-term structural change vectors through weighted summation. The semantic global attention path performs cross-frame attention calculations on the semantic features of the entire frame, calculates the inter-frame semantic similarity matrix, and generates semantic evolution feature vectors through matrix multiplication. The short-term structural change vectors and semantic evolution feature vectors are concatenated and fused to form a temporal joint feature representation.

[0089] A fusion controller is constructed, which performs feature concatenation operation on memory-enhanced structural features and semantic evolution feature vectors, and inputs them into a fusion weighting layer. The fusion weighting layer assigns fusion weights according to the modality confidence distribution and generates the structural fusion result through attention-weighted summation. The fusion controller outputs a predicted structure map and a structure confidence score map. The structure confidence score map is fitted with a regression function to obtain a confidence distribution curve, which is used to characterize the stability of the target structure prediction.

[0090] This implementation method achieves dynamic perception, semantic continuous modeling, and memory enhancement fusion of target structure in time series through the above method, improving the stability, accuracy, and temporal consistency of target prediction structure under multimodal input, and providing a robust model foundation and reliable structure prediction capability for intelligent target tracking in low-altitude complex scenarios.

[0091] In this embodiment, S3 specifically refers to:

[0092] Extract visual image feature vectors, infrared image feature vectors, and radar echo feature vectors under the same time index to construct a time-aligned multimodal feature set; perform spatial coordinate transformation on the structural segmentation results in the predicted structural map and project them to a unified spatial coordinate system; perform boundary segmentation on each structural region in the predicted structural map, extract the structural boundary contour and region mask, and generate a set of structural region boundaries;

[0093] Feature matching is performed between the set of structural region boundaries and the corresponding regions in the visual image feature vector to construct a visual modal similarity matrix; feature matching is performed between the set of structural region boundaries and the corresponding regions in the infrared image feature vector to construct an infrared modal similarity matrix; feature matching is performed between the set of structural region boundaries and the corresponding regions in the radar echo feature vector to construct a radar modal similarity matrix.

[0094] The visual modal similarity matrix, infrared modal similarity matrix, and radar modal similarity matrix are input into the fusion scoring module. A weighted operation is performed according to the preset modal fusion weights to generate a consistency scoring matrix. The scoring values ​​of all structural regions in the consistency scoring matrix are averaged to generate a global consistency scoring value. The global consistency scoring value is obtained by fitting the average scoring value through the scoring mean calculation module.

[0095] If the global consistency score is less than the first preset threshold, the fusion scaling factor corresponding to the predicted structure graph is input to the confidence adjustment structure. The confidence adjustment structure includes a confidence calculation unit, a scaling factor adjustment unit, and a time series recording unit. The confidence calculation unit performs weighted regression analysis on the score data in the consistency score matrix and outputs the structure confidence value. The scaling factor adjustment unit receives the structure confidence value and the fusion scaling factor, calls the fusion control function to perform scaling factor suppression operation, and outputs the adjusted fusion scaling factor. The time series recording unit records the time index and modality number corresponding to the suppression operation in the adjustment index table, marking it as the prediction weight reduction period.

[0096] After the predicted weight reduction period is labeled, the current time index is added to the tracking state sequence, and the corresponding target's state label information is updated in the tracking state sequence to record the weight reduction state label.

[0097] In this embodiment, S4 specifically refers to:

[0098] Under the time index corresponding to the predicted weight reduction period, a spatial neighborhood region with a fixed radius is constructed based on the target position coordinates of the previous time step in the fused input dataset. Image cropping is performed on the spatial neighborhood region to extract image region fragments from visual images, infrared images, and radar echo data, respectively, as a candidate sensing image set. The candidate sensing image set is input into a multimodal shared convolutional path, and feature extraction operations are performed to generate visual image feature vectors, infrared image feature vectors, and radar echo feature vectors, thus constructing a candidate target sample set.

[0099] For each target sample in the candidate target sample set, the trajectory mapping function is called to calculate the Euclidean distance and the direction angle between the historical trajectory endpoint position at the previous moment and the current candidate sample position; the Euclidean distance and the direction angle are input into the distance weighting function, and weighted calculation is performed to generate a trajectory distance score; the trajectory distance score is obtained by fitting through the trajectory distance fitting module;

[0100] Filter all candidate target samples whose trajectory distance score is less than a set similarity threshold to construct a spatial candidate subset; extract visual image feature vectors, infrared image feature vectors and radar echo feature vectors from the spatial candidate subset, perform modal channel unification and time alignment operations, and construct a multimodal joint feature set;

[0101] The multimodal joint feature set is input into the density clustering structure, and similarity calculation and feature classification operations are performed to generate multiple feature clusters. Cluster density analysis is performed on all feature clusters, and the feature cluster with the highest cluster density is selected as the anchor feature set. A weighted average operation is performed on all feature vectors in the anchor feature set to generate centroid feature vectors, which are output as anchor feature vectors. The anchor feature vectors are obtained by fitting through the modal co-clustering weighted module.

[0102] The anchor point feature vector's time index, center position coordinates, and modality number label are combined to construct the anchor point path information entry; the anchor point path information entry is then written into the tracking auxiliary cache structure, and the anchor point source state is marked in the tracking state sequence.

[0103] In this embodiment, S5 specifically refers to:

[0104] A frame-by-frame traversal operation is performed on the fused input dataset under continuous time index. For each frame, the target detection structure is called for the visual image, infrared image, and radar echo data, and an initial detection bounding box sequence is output. A boundary matching judgment operation is performed on the initial detection bounding box sequence to determine whether the center position, boundary aspect ratio, and boundary information of the predicted structure map at the previous time step meet the condition for the continued existence of the target. Statistical accumulation is performed on the detection frames that do not meet the condition for continued existence. If the number of frames that do not meet the condition for the existence of the target within the continuous time window exceeds a set threshold, the target is marked as entering the occlusion state, and the occlusion start time index is recorded.

[0105] In subsequent time indices, detection and judgment operations continue to be performed. Boundary center extraction, region contour reconstruction and area comparison processing are performed on all target candidate regions in the detection frame, and the structure alignment score is output. The structure alignment score is compared with the historical boundary template of the target predicted structure map. If the score exceeds the recovery judgment threshold, the occlusion state is determined to end, and the occlusion end time index is recorded.

[0106] The visual image, infrared image, and radar echo image in the frame corresponding to the occlusion end time index are extracted, and channel normalization, scale resampling, and feature mapping construction operations are performed respectively to obtain three types of modal feature vectors with structure alignment. The feature splicing unit is called to fuse the three types of modal feature vectors according to the channel dimension to generate a multimodal joint feature vector of the re-identified target.

[0107] The system retrieves the corresponding target identifier's identity feature cache set from the historical state cache structure of the fusion controller. The identity feature cache set includes a fusion feature vector sequence, a spatial location index sequence, and an identity tag number set stored after time-series sampling. The system then calls the identity similarity calculation module to perform modal weighted cosine similarity calculation on the multimodal joint feature vector and each fusion feature vector in the identity feature cache set. The similarity value is obtained by fitting the modal weight function with the embedded spatial vector calculation module.

[0108] Perform a maximum value retrieval operation on all similarity values ​​to obtain the highest matching score and the corresponding identity tag number, and record the time index and fusion input data frame identifier corresponding to the matching score; output the matching score as the identity similarity score, and at the same time, write the identity tag associated with the matching score as the identity confirmation candidate tag and write it into the identity confirmation pending judgment queue to enter the next step of the identity regression judgment process.

[0109] This implementation method constructs a continuous occlusion detection mechanism and recovery judgment rules, and combines them with a multimodal joint feature re-identification process to achieve preliminary confirmation of the identity of the target after occlusion, effectively improving the accuracy of target re-identification and the continuity of label recovery in occluded scenarios.

[0110] In this embodiment, S6 specifically refers to:

[0111] After receiving the identity similarity score, perform a numerical range judgment operation; call the threshold judgment structure to compare the identity similarity score with the second preset threshold to determine whether it meets the condition of being greater than the second preset threshold; for identity similarity scores that meet the condition of being greater than the second preset threshold, continue to compare them with the third preset threshold to determine whether they simultaneously meet the condition of being less than the third preset threshold.

[0112] If the identity similarity score meets the condition of being greater than the second preset threshold and less than the third preset threshold, then the visual image feature vector, infrared image feature vector and radar echo feature vector of the corresponding frame of the continuous time index are retrieved; the extracted continuous multi-frame feature vectors are constructed into a multi-frame joint feature set according to the time index order, and channel stacking and scale normalization processing are performed to form a three-dimensional joint feature tensor.

[0113] For each frame in the multi-frame joint feature tensor, the three modal feature vectors are compared with the fused feature vectors in the historical identity feature cache set. The modal weighted cosine similarity function is called to calculate the matching score. The modal weighted cosine similarity function is obtained by fitting the multi-modal confidence adaptive weight function with a weighted inner product structure. The matching score of each frame is recorded in the matching score matrix, and a multi-frame matching score matrix is ​​constructed based on the time order.

[0114] The matching scores of all frames in the multi-frame matching score matrix are summed and averaged to generate an average matching score. The identity verification judgment logic structure is called to determine whether the average matching score is greater than the third preset threshold. If it is greater than the third preset threshold, the identity matching verification result is output and the matching tag number is recorded.

[0115] If the identity similarity score is greater than the third preset threshold, or the identity matching confirmation result is determined to be consistent, then extract the target number and historical identity tag number of the current frame, call the identity trajectory merging structure to update the target identity mapping relationship table; bind the target identifier of the current frame with the historical identity tag, and call the identity trajectory cache structure to restore the original trajectory state sequence of the corresponding target;

[0116] After the identity trajectory is restored, the anchor point tracking termination process is initiated. The anchor point path management structure is called to release the anchor point path tracking task of the corresponding target. The fusion scaling factor corresponding to the target is reset to the initial state parameters of the prediction structure graph, and the identity restoration label is recorded to the identity record set.

[0117] This implementation method introduces interval judgment, multi-frame joint matching and trajectory merging mechanisms to achieve accurate regression of identity tags and synchronous termination of anchor path in occlusion recovery scenarios, thereby improving the ability to maintain identity consistency and the accuracy and efficiency of tracking path management.

[0118] In this embodiment, S7 specifically refers to:

[0119] In each frame corresponding to a time index, visual image feature vectors, infrared image feature vectors, and radar echo feature vectors are extracted. A modal stitching structure is then used to construct a fusion input feature set. The fusion input feature set is input to the fusion controller, which receives the predicted structure map from the previous time index and the current fusion scaling factor. Cross-frame feature compensation and modal dynamic fusion operations are performed to generate a fusion state vector. The target's center position coordinates, bounding box parameters, contour change index, and modal confidence distribution are extracted from the fusion state vector. The state record structure is then used to write the above parameters into the fusion state cache table in time index order, generating the fusion output result for the current frame.

[0120] Based on the identification number of the target corresponding to the fusion output, and combining the identity similarity score, the identity confirmation frame mechanism output, and the trajectory number mapping rule, the identity number confirmation operation is performed; the identity labeling structure is called to match the target identifier of the current frame with the historical identity label, and the identity label number is output.

[0121] After the identity tag number is confirmed, the fusion output of the current frame is bound to the identity tag number to construct the target tracking status entry; the trajectory cache write structure is called to write the target tracking status entry into the tracking trajectory cache set in time index order, and the target tracking trajectory sequence is generated by dividing the time window;

[0122] An identity record index structure is established for each target identifier. The identity record index structure includes fields for label change trajectory, anchor path number, occlusion recovery label and trajectory jump marker. Multidimensional aggregation operation is performed on the fusion state cache table and the identity record index structure to generate an identity record set. The target tracking trajectory sequence and the identity record set are used as the complete output of the tracking cycle.

[0123] This implementation method improves the temporal consistency of state records and the full-cycle management capability of identity information during target tracking by integrating state vector generation, identity number binding, and index structure organization, thereby enhancing the integrity and traceability of tracking results.

[0124] In this embodiment, it also includes:

[0125] Perform state analysis on the target tracking trajectory sequence and identity record set generated within the tracking period; statistically analyze the duration of continuous occlusion, identity switching frequency, modality confidence fluctuation amplitude and fusion ratio factor variation range to generate an anomaly monitoring index set;

[0126] Each indicator in the abnormal monitoring indicator set is compared with its corresponding abnormal judgment threshold. If the duration of continuous occlusion is greater than the preset occlusion threshold, or the identity switching frequency is greater than the preset frequency threshold, or the modal confidence fluctuation amplitude is greater than the preset fluctuation threshold, or the fusion ratio factor change range is greater than the preset range threshold, then the tracking abnormality mark is triggered.

[0127] After triggering the tracking anomaly marker, terminate the current target's fusion state vector update operation and write the anomaly state to the anomaly label record table; send a reset request to the anchor path cache structure, identity feature cache set and fusion controller to perform a state reset operation; record the time index and target number corresponding to the reset operation and build an anomaly recovery process log.

[0128] Example 1:

[0129] To verify the feasibility of this invention in practice, it was applied to the expansion of a city's rail transit system to suburban areas. Several key sections were located near urban mountains and complex environments, resulting in blind spots in the ground-to-air connection points under unattended conditions. To improve the continuous tracking capability of operational targets (such as operational drones, inspection robots, and emergency vehicles) in low-altitude areas and ensure the timeliness and accuracy of remote monitoring, the management deployed an intelligent tracking system equipped with the capabilities of this invention at the B-area test base and conducted comparative experiments with mainstream methods.

[0130] In real-world scenarios, Area B is an overlapping area between high-altitude tracks and semi-open equipment areas. The low-altitude environment is limited by uneven lighting, high humidity, vegetation obstruction, and wind interference. Conventional video surveillance methods suffer from problems such as target loss due to occlusion, inaccurate re-identification, and inconsistent identification numbers. This invention employs a multimodal input sequence constructed by fusing visual images, infrared images, and radar echo data, effectively addressing performance fluctuations caused by environmental factors in different modalities. The system integrates an improved TimeSformer model, a modality consistency scoring mechanism, and an identity regression judgment strategy, enabling re-identification of occluded targets and anchor point path correction.

[0131] During the three-month trial period, the system completed 1387 task tracking cycles, covering three typical targets: fixed-track inspection drones, mobile emergency patrol vehicles, and track maintenance robots. To compare performance, the testing team set up four control groups: the method of this invention, the traditional Kalman filtering method, the YOLOv5+LSTM method, and the DeepSORT method. Test metrics focused on occlusion handling capability and stability under extreme environments, as detailed in the table below:

[0132] Table 1 Comparison of Occlusion Recovery Recognition Accuracy

[0133]

[0134] As shown in Table 1, the present invention performs optimally in target occlusion recovery scenarios. Specifically, the occlusion recovery accuracy of the method of the present invention reaches 94.6%, far exceeding the 77.2% of the traditional Kalman method, the 85.8% of the YOLOv5+LSTM method, and the 88.1% of the DeepSORT method. In terms of identity recovery latency, the present invention is only 218ms, which is about 20%-50% lower than the other three methods, achieving faster response. The false recognition rate is also controlled within 1.3%, significantly better than other methods. In terms of occlusion processing time, the present invention has an average processing time of 0.84s, which is the only solution less than 1s. This indicates that the present invention effectively solves the problems of target loss and false recognition during occlusion by fusing the dynamic suppression mechanism of the scaling factor and the auxiliary mechanism of the anchor path, significantly improving the tracking robustness and accuracy under occlusion conditions.

[0135] Furthermore, to verify the stability of the multimodal fusion scheme under different weather conditions, the testing team selected four typical weather conditions—sunny, cloudy, light rain, and heavy fog—and compared the differences in the percentage of lost frames among the various methods. Table 2 shows the frame loss situation of different methods under various weather conditions.

[0136] Table 2 Comparison of Multimodal Fusion Tracking Stability under Different Weather Conditions

[0137]

[0138] As shown in Table 2, all three methods performed well under clear weather conditions, but the present invention still leads with a frame loss rate of 0.6%, demonstrating its basic stability. In environments with increased interference, such as cloudy days, light rain, and heavy fog, the advantages are even more pronounced. Under light rain conditions, the frame loss rates of YOLOv5+LSTM and DeepSORT are 7.2% and 5.8%, respectively, while the present invention's is only 2.8%. In heavy fog scenarios, although the present invention is slightly affected, it is still controlled at 4.3%, far superior to YOLOv5+LSTM's 10.8%. This indicates that the modal confidence distribution mechanism and dynamic fusion controller constructed in this invention have stronger adaptability to low visibility conditions, especially when infrared and radar modal characteristics are emphasized, ensuring stable output of the overall system.

[0139] This embodiment achieves continuous, accurate, and robust target tracking in complex low-altitude environments through multimodal fusion modeling, temporal structure learning, occlusion state recognition, and identity regression mechanisms. It significantly outperforms existing mainstream algorithms and is suitable for low-altitude target management needs in urban rail transit, airport access roads, and mountainous facilities. The system not only solves the problems of identity confusion and trajectory jumps during target occlusion but also improves continuity and stability in extreme environments, providing a feasible solution and practical reference for multimodal target perception fusion applications.

[0140] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A low-altitude intelligent target tracking method based on multimodal fusion, characterized in that, Includes the following steps: S1. Collect multimodal perception data of the target, construct a multimodal input sequence, and generate a fused input dataset; S2. Based on the fused input dataset, perform initial target detection and semantic structure extraction to obtain the semantic structure feature map and modality confidence distribution of the target; input the semantic structure feature map into the improved TimeSformer model to output the predicted structure map; S3. Perform a matching operation between the predicted structure graph and the multimodal features extracted at the same time to construct a consistency score matrix; calculate the global consistency score based on the consistency score matrix; if the global consistency score is less than the first preset threshold, perform a modal confidence suppression operation on the fusion scaling factor corresponding to the predicted structure graph, and mark the corresponding time as the prediction weight reduction period. S4. During the prediction weight reduction period, extract candidate target samples whose trajectory distance meets the set similarity threshold from the target's neighborhood, perform feature clustering operation, and generate anchor point feature vectors. S5. Determine whether the occlusion state of the target has ended. If the occlusion state has ended and the target is detected to reappear in the perception area, then perform the identity verification operation, extract the target feature vector at the same time and match it with the historical identity feature vector to generate an identity similarity score. S6. Perform identity regression judgment based on identity similarity score and second and third preset thresholds; if identity similarity score is greater than second preset threshold and less than third preset threshold, trigger identity confirmation frame mechanism, continuously extract target features from multiple frames for dense matching operation; if identity similarity score is greater than third preset threshold, or identity confirmation frame mechanism outputs matching confirmation result, restore the original identity trajectory of the target and terminate the target anchor point path. S7. Record the fusion output result and identity tag status of each frame during the tracking period to generate the target tracking trajectory sequence and identity record set.

2. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 1, characterized in that, Specifically, S1 is: Visual images, infrared images, and radar echo data are acquired to construct a multimodal input sequence containing multimodal sensing data; frame numbering is uniformly processed on the multimodal input sequence, and a global sampling time interval and time index order are set; The data from each modal sensing mode are interpolated, truncated, or padded according to the global time index until the number of frames in each modal sequence is consistent. Based on a unified time index, data from visual images, infrared images, and radar echo data corresponding to the time frame are combined into a fused frame group. Perform spatial coordinate projection transformation on each modal data in the fused frame group to construct a spatial alignment result under a unified spatial reference. Based on the spatial alignment results, the information from each modality image is combined to generate a fused input dataset.

3. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 2, characterized in that, The initial target detection and semantic structure extraction are performed based on the fused input dataset to obtain the semantic structure feature map and modality confidence distribution of the target, specifically as follows: Perform feature extraction on each frame of multimodal image data in the fused input dataset to extract visual image feature vectors, infrared image feature vectors, and radar echo feature vectors; The visual image feature vector, infrared image feature vector, and radar echo feature vector are input into a unified feature fusion network to generate a fused feature map. Candidate region generation is performed based on the fused feature map to identify potential target regions and perform region cropping. Perform target bounding box fitting and semantic mask segmentation on the cropped region to generate initial target detection results and semantic structure map; Multimodal fusion features are extracted for each detected target in the semantic structure graph, the contribution weight of each modality is calculated, and the modality confidence distribution is generated.

4. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 3, characterized in that, The improved TimeSformer model includes a multimodal decoupling unit, a rhythm-aware attention module, a temporal modeling structure, a memory state injection unit, and a fusion controller, specifically: A multimodal decoupling unit is constructed, and convolutional coding operations are performed on the semantic structure feature maps of visual images, infrared images, and radar echo images respectively to extract edge change features, hot spot distribution features, and scattering morphology features; The feature maps of the edge variation features, hot spot distribution features and scattering morphology features are subjected to channel alignment and dimension normalization operations to generate modality-enhanced feature maps. A rhythm-aware attention module is constructed. The confidence change trend, motion displacement amplitude and occlusion state index of the target in continuous historical frames are normalized to generate a rhythm state vector. The size of the inter-frame attention window is adjusted according to the rhythm state vector, and dynamic window encoding is performed to generate a rhythm-aware position vector. A temporal modeling structure is constructed, and the rhythm-aware position vector and modality-enhanced feature map are input into the structural continuous attention path and the semantic global attention path. The structural continuous attention path generates short-term structural change vectors based on boundary region masks, and the semantic global attention path performs cross-frame semantic attention operations based on full-map features to generate semantic evolution feature vectors. A memory state injection unit is constructed. Based on the predicted structure map of historical continuous frames, the contour boundary point set, region mask map and center drift vector are extracted. A structural memory representation tensor is constructed and residually fused with the output of the structural continuous attention path to generate memory-enhanced structural features. A fusion controller is constructed to concatenate memory-enhanced structural features and semantic evolution feature vectors, and performs weight allocation and fusion operations based on modality confidence distribution, outputting a predicted structure map and a structure confidence score map.

5. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 4, characterized in that, Specifically, S3 is: Extract visual image feature vectors, infrared image feature vectors, and radar echo feature vectors under the same time index to construct a time-aligned multimodal feature set; project the predicted structure map onto a unified spatial coordinate system, perform boundary segmentation operation on each structural region in the predicted structure map, and generate a set of structural region boundaries; Feature matching operations are performed on the corresponding regions in the structural region boundary set and visual image feature vector, infrared image feature vector and radar echo feature vector, respectively, to generate visual modal similarity matrix, infrared modal similarity matrix and radar modal similarity matrix; weighted calculation is performed on the visual modal similarity matrix, infrared modal similarity matrix and radar modal similarity matrix according to preset modal fusion weights to generate consistency score matrix; Calculate the average score of all structural regions in the consistency score matrix to generate a global consistency score. If the global consistency score is less than the first preset threshold, the fusion scaling factor associated with the predicted structure graph is input to the confidence adjustment structure. The confidence adjustment structure includes a confidence calculation unit, a scaling factor adjustment unit, and a time-series recording unit. The confidence calculation unit receives the structure confidence value output by the consistency score matrix, the scaling factor adjustment unit performs a suppression operation on the fusion scaling factor according to the confidence value, and the time-series recording unit records the time index when the fusion scaling factor suppression occurs. After the confidence adjustment structure completes the suppression operation, it marks the corresponding time index as the predicted deweighting period and records the deweighting state label in the tracking state sequence.

6. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 5, characterized in that, Specifically, S4 is: Under the time index corresponding to the predicted weight reduction period, a spatial neighborhood region with a fixed radius is constructed based on the target position coordinates of the previous moment in the fused input dataset; Visual images, infrared images, and radar echo data are extracted from the spatially adjacent region. Convolutional feature extraction operations are performed on each of these data to generate visual image feature vectors, infrared image feature vectors, and radar echo feature vectors, which together form a candidate target sample set. For each target sample in the candidate target sample set, calculate the Euclidean distance and the direction angle between the historical trajectory endpoint position of the target sample and the current candidate sample position, and jointly generate a trajectory distance score. Target samples with trajectory distance scores less than a set similarity threshold are selected, and their corresponding visual image feature vectors, infrared image feature vectors, and radar echo feature vectors are extracted to construct an aligned multimodal joint feature set. Based on the multimodal joint feature set, cluster center calculation and feature point classification operations are performed to generate multiple feature clusters; the feature set with the highest cluster density in the feature cluster is used to calculate the center point feature vector, which is output as the anchor point feature vector; the time index, spatial coordinates and modal feature number corresponding to the anchor point feature vector are recorded, the anchor point path information sequence is constructed and stored in the tracking auxiliary cache structure.

7. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 6, characterized in that, Specifically, S5 is: Perform target detection on visual images, infrared images and radar echo data in a continuous multi-frame fused input dataset, and determine whether there are target candidate regions in the target detection region that are associated with the target prediction structure map; If no corresponding target candidate region is detected in several consecutive frames, the target is marked as entering an occlusion state, and the occlusion start time index is recorded. Continue performing the target detection operation. If a new target candidate region is detected, and the center coordinate position, boundary contour features and boundary information of the predicted structure map of the new target candidate region meet the preset recovery conditions, then the occlusion state is determined to end, and the occlusion end time index is recorded. Extract the visual image feature vector, infrared image feature vector and radar echo feature vector from the frame corresponding to the occlusion end time index, and construct a multimodal joint feature vector for the re-identified target. Extract the target's identity feature cache set from historical records. The identity feature cache set includes a labeled multimodal fusion feature vector sequence, location index sequence, and identity tag number. Perform a similarity calculation operation on the multimodal joint feature vector of the re-identified target and the fusion feature vector in the identity feature cache set. Generate an identity similarity score based on the modality weighted cosine similarity function. Record the time index and target identification information corresponding to the identity similarity score, output the identity confirmation candidate tag, and enter the identity regression judgment process.

8. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 7, characterized in that, Specifically, S6 is: After receiving the identity similarity score, a threshold range judgment operation is performed, comparing the identity similarity score with the second preset threshold and the third preset threshold. If the identity similarity score is greater than the second preset threshold and less than the third preset threshold, then visual image feature vector, infrared image feature vector and radar echo feature vector are extracted from the continuous time index to construct a multi-frame joint feature set. Perform matching calculations on the feature vector of each frame in the multi-frame joint feature set and the fused feature vector in the historical identity feature cache set, and generate a multi-frame matching score matrix based on the modality weighted cosine similarity function; perform an averaging operation on the matching scores of all frames in the multi-frame matching score matrix to generate an average matching score. If the average matching score is greater than the third preset threshold, the target identity is determined to be consistent with the historical identity, and the identity matching confirmation result is output. If the identity similarity score is greater than the third preset threshold, or the identity matching confirmation result is determined to be consistent, the identity identifier of the current frame will be associated with the historical identity label to restore the original identity trajectory sequence. After the identity trajectory is restored, the anchor point path tracking operation associated with the target is terminated, the fusion scaling factor is updated to the initial state parameters, and the identity restoration label is recorded.

9. The low-altitude intelligent target tracking method based on multimodal fusion according to claim 8, characterized in that, Specifically, S7 is: In each time indexed frame, visual image feature vectors, infrared image feature vectors, and radar echo feature vectors are extracted to construct a fusion input feature set. The fusion input feature set is input to the fusion controller, and feature fusion operation is performed by combining the prediction structure map of the previous frame and the fusion scaling factor to output a fusion state vector. The center position coordinates, bounding box parameters, and modal confidence distribution in the fusion state vector are recorded to generate the fusion output result of the current frame. Based on the identity similarity score, the identity confirmation frame mechanism result and the trajectory number mapping relationship, the identity number of the target in the current frame is confirmed; after confirming the identity label, the fusion output result is combined with the identity label number to construct the target tracking status entry; the target tracking status entry is stored in the tracking trajectory cache set in time index order to generate the target tracking trajectory sequence; For each target identifier, an identity record index structure is established, and identity tag status changes, anchor path records, identity recovery operations and trajectory jump markers are summarized to generate an identity record set. The target tracking trajectory sequence and the identity record set are used as the output of the tracking cycle.

Citation Information

Patent Citations

  • Visual target tracking method and terminal based on attention mechanism

    CN117372721A

  • Multi-target tracking method based on weak clue and trajectory prediction

    CN119904485A