Hierarchical multi-modal tracking method and system based on large model cognitive driving

CN121811326BActive Publication Date: 2026-08-21SOUTHWEST UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610011466.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-08-21
Estimated Expiration
2046-01-06

AI Technical Summary

Technical Problem

[0006]针对现有技术中的上述不足,本发明提供的基于大模型认知驱动的层级化多模态跟踪方法及系统解决了现有技术缺乏对场景语义和目标意图的高层认知推理能力、无法完全实现符合物理规律的轨迹修复与长时关联的问题

Benefits of technology

本发明通过整合采集到的可见光与红外热成像双模态视频流,利用采用基于视觉傅里叶提示的特征融合网络与包含不确定性感知能力的概率型检测网络,实现了“实时感知+按需认知”的高效跟踪模式。通过在主干网络中嵌入频域提示向量,有效结合了可见光与红外的全局纹理与轮廓特征,克服了单模态方案在低照度或热交叉掩盖下的感知缺陷。同时,引入多因子加权融合,实时监测定位协方差与运动突变,仅在必要时唤醒认知智能体,从而在保证高准确度的同时降低系统的平均算力消耗。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811326B_ABST
    Figure CN121811326B_ABST
Patent Text Reader

Abstract

The application discloses a hierarchical multi-modal tracking method and system based on a large model cognitive drive, relates to the technical field of computer vision and multi-modal intelligent perception, and comprises the following steps: generating enhanced fusion features based on visible light and infrared thermal imaging images in a monitoring scene; adopting a probabilistic single-stage target detection network to perform target tracking based on the enhanced fusion features, so as to obtain target positioning results; monitoring state monitoring signals of the tracking target in real time according to the target positioning results; performing adaptive decision-making according to the state monitoring signals; wherein the adaptive decision-making comprises a continuous tracking mechanism and a passive wake-up mechanism; and performing target tracking according to the adaptive decision-making. The application can perform spatio-temporal causal logic reasoning and semantic association on the 'disappearance and reappearance' process of a target, significantly improves the identity maintenance capability under long-time occlusion, effectively solves the problems of stiff and unreasonable trajectories caused by traditional linear interpolation, and reduces the average computing power consumption of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and multimodal intelligent perception technology, specifically to a hierarchical multimodal tracking method and system based on large model cognition. Background Technology

[0002] Multi-target tracking, as a core perception module in intelligent monitoring, autonomous driving, and human-machine interaction systems, aims to continuously locate multiple targets of interest in a video sequence and maintain consistent identification. With the development of sensor technology, multimodal tracking technology combining visible light and infrared thermal imaging has attracted widespread attention due to its all-weather perception advantages.

[0003] Currently, most mainstream trackers rely primarily on underlying motion coherence and appearance similarity. This "data-driven" approach suffers from identity switching or trajectory breakage when faced with long-tail scenarios such as severe occlusion, abrupt changes in target appearance, or complex interactions, due to a lack of understanding of scene semantics.

[0004] Although multimodal large models have demonstrated powerful semantic understanding and reasoning capabilities, their massive number of parameters results in extremely high inference latency, which cannot meet the stringent real-time requirements of video surveillance or autonomous driving. In addition, existing trajectory completion methods are mostly based on simple linear interpolation, and the generated trajectories lack physical realism and reasonableness in environmental interaction.

[0005] For example, the Chinese invention patent CN114663470 A provides an "Adaptive Cross-Modal Visual Tracking Method Based on Soft Selection," which solves the problem of information loss caused by modality switching by designing a soft selection module to adaptively predict the importance weights of visible light and near-infrared modes. However, this method only involves weighted fusion at the feature level and lacks high-level cognitive reasoning capabilities regarding scene semantics and target intent. When faced with severe occlusion or complex interactions leading to prolonged target loss, it cannot fully achieve trajectory repair and long-term association that conforms to physical laws. Summary of the Invention

[0006] To address the aforementioned shortcomings in existing technologies, the hierarchical multimodal tracking method and system based on large-model cognitive drive provided by this invention solves the problems of existing technologies lacking high-level cognitive reasoning capabilities for scene semantics and target intent, and being unable to fully achieve trajectory repair and long-term association that conform to physical laws.

[0007] To achieve the above-mentioned objectives, this invention provides a hierarchical multimodal tracking method based on large-model cognitive drive, comprising: Enhanced fusion features are generated based on visible light and infrared thermal imaging images in surveillance scenarios; Based on enhanced fusion features, a probabilistic single-stage target detection network is used for target tracking to obtain target localization results; Based on the target location results, monitor and track the target's status signals in real time. Adaptive decision-making is based on status monitoring signals; the adaptive decision-making includes a continuous tracking mechanism and a passive wake-up mechanism. Target tracking is performed based on adaptive decision-making.

[0008] This invention also provides a hierarchical multimodal tracking system based on large model cognition, comprising: The front-end fast execution module is used to generate enhanced fusion features based on visible light and infrared thermal imaging images in the monitoring scene; based on the enhanced fusion features, a probabilistic single-stage target detection network is used to track the target and obtain the target localization result; and the status monitoring signal of the tracked target is monitored in real time according to the target localization result. The backend inference and execution module is used to make adaptive decisions based on status monitoring signals; and to track targets based on the adaptive decisions.

[0009] The beneficial effects of this invention are as follows: This invention integrates acquired dual-modal video streams from visible light and infrared thermal imaging. Utilizing a feature fusion network based on visual Fourier cueing and a probabilistic detection network incorporating uncertainty perception capabilities, it achieves a highly efficient tracking mode of "real-time perception + on-demand cognition." By embedding frequency-domain cue vectors into the backbone network, it effectively combines global texture and contour features from both visible light and infrared sources, overcoming the perception limitations of single-modal schemes under low-light or thermal cross-masking conditions. Simultaneously, it introduces multi-factor weighted fusion to monitor localization covariance and motion mutations in real time, activating the cognitive agent only when necessary, thereby reducing the system's average computational power consumption while maintaining high accuracy.

[0010] To address the shortcomings of existing tracking algorithms, such as a lack of logical reasoning capabilities and susceptibility to identity loss or errors under severe occlusion or complex interactions, this invention introduces a cognitive agent based on a multimodal large model. Through a thought chain mechanism, the system can perform spatiotemporal causal logical reasoning and semantic association on the "disappearance and reappearance" process of targets, significantly improving identity maintenance capabilities under long-term occlusion. Furthermore, this invention utilizes a generative repair layer based on a diffusion model, guided by singular space projection and semantic conditions, to generate trajectory completion results that conform to physical motion inertia and scene constraints, effectively solving the problems of stiff and unreasonable trajectories caused by traditional linear interpolation. This method is also applicable to other scenarios requiring long-term continuous tracking and trajectory prediction of multiple targets. Attached Figure Description

[0011] Figure 1The flowchart of a hierarchical multimodal tracking method based on large model cognition is provided for the embodiment. Detailed Implementation

[0012] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0013] like Figure 1 As shown, in one embodiment of the present invention, a hierarchical multimodal tracking method based on large model cognition-driven methods includes the following steps: S1. Enhanced fusion features are generated based on visible light and infrared thermal imaging images in the monitoring scene.

[0014] The specific method is as follows: S1-1. Acquire and time-series align visible light and infrared thermal imaging images in the monitored scene; S1-2. Divide the time-aligned visible light and infrared thermal imaging images into image blocks of fixed size, and linearly project the image blocks into a sequence of feature tokens; S1-3. Use a feature extraction network based on visual Fourier prompts and modal fusion prompt generators to extract features from the feature token sequence to obtain enhanced fusion features.

[0015] Steps S1-3 specifically include: Learnable cue vectors are embedded in the Transformer backbone network, and two-dimensional fast Fourier transforms are performed on the learnable cue vectors in the channel and spatial dimensions. The vectors after fast Fourier transforms are then mapped to the frequency domain space to obtain frequency domain cue. The feature token sequence is input into a pre-trained Transformer backbone network containing frequency domain cues for multi-level feature extraction to obtain visible light features and infrared features. The channel dimensions of visible light and infrared features are compressed using a shared linear projection layer; the two compressed features are fused using an addition operation; and the fused features are then subjected to feature interaction and dimension restoration using a convolution operation to obtain a unified modality fusion cue, the expression of which is:

[0016] In the formula, This indicates a unified modal fusion suggestion. This represents the convolution operation. This represents a linear projection operation. Indicates the characteristics of visible light. Indicates infrared characteristics; The unified modal fusion cue is fused with visible light features and infrared features in residual form to obtain enhanced fusion features, including visible light enhanced fusion features and infrared enhanced fusion features, the expressions of which are:

[0017] In the formula, This indicates visible light enhanced fusion characteristics. Indicates infrared enhanced fusion characteristics, This represents the residual form of a unified modal fusion cues.

[0018] S2. Based on enhanced fusion features, a probabilistic single-stage target detection network is used for target tracking to obtain target localization results.

[0019] Specifically: S2-1. Pre-train the single-stage object detection network using negative log-likelihood loss. The expression for the negative log-likelihood loss function is:

[0020] In the formula, Indicates the loss value. This represents the mean coordinates of the predicted target bounding box. The bounding box representing the true target. This represents the target's localization covariance matrix; S2-2. Input the enhanced fusion features corresponding to the current frame into the pre-trained single-stage object detection network, and output the bounding box coordinates of all potential targets in the current enhanced fusion features, their corresponding confidence scores, and the localization covariance matrix; obtain the mean coordinates of the target bounding boxes; model the target localization results as a multivariate Gaussian distribution, the expression of which is:

[0021] In the formula, This represents a multivariate Gaussian distribution, i.e., the target localization result.

[0022] S3. Monitor and track the target's status signals in real time based on the target positioning results.

[0023] Specifically: S3-1. Obtain the target prediction trajectory based on the real-time target positioning results; S3-2. Based on the preset confidence threshold and the confidence level of each bounding box, divide the bounding boxes into a high-score set and a low-score set; S3-3. Match the high-scoring set with the target predicted trajectory once, and match the low-scoring set with the unmatched target predicted trajectory a second time. S3-4. Extract the matching quality features of the current frame, including motion mutation features containing cross-union ratio change rate and normalized velocity vector mutation value, occlusion overlap ratio features and cross-frame appearance consistency features; S3-5. A multi-factor weighted fusion model is used to map the extracted features into a unified anomaly risk score, the expression of which is:

[0024] In the formula, Indicates the abnormal risk score. This represents the weight coefficients corresponding to the matching quality features. Indicates matching quality features. This represents the weighting coefficients corresponding to the motion abrupt change features, which include the crossover ratio change rate and the normalized abrupt change value of the velocity vector. This represents the normalized abrupt change value of the velocity vector. Indicates the rate of change of the crossover ratio. Weighting coefficients representing the occlusion overlap ratio feature. This indicates the occlusion overlap ratio feature. The weighting coefficients represent the appearance consistency features across frames. Indicates cross-frame appearance consistency characteristics; S3-6. Real-time comparison of the abnormal risk score with the preset safety risk threshold, and output status signal; when the abnormal risk score is less than or equal to the preset safety risk threshold, the output status monitoring signal is 0, otherwise the output status monitoring signal is 1.

[0025] S4. Make adaptive decisions based on status monitoring signals; the adaptive decision-making includes a continuous tracking mechanism and a passive wake-up mechanism.

[0026] When the status monitoring signal is 0, the continuous tracking mechanism is triggered; when the status monitoring signal is 1, the passive wake-up mechanism is triggered.

[0027] ① The continuous tracking mechanism specifically involves updating the trajectory state of the current frame using an adaptive Kalman filter, including: The matching relationship between the bounding box of the current frame and the predicted trajectory of the target is locked. Using the bounding boxes in the high-resolution set as observations, the state vector of the Kalman filter is updated posteriorly. The expression is as follows:

[0028] In the formula, This represents the updated state vector posteriorly. For the prior prediction state, For Kalman gain, The observation matrix; By utilizing the statistical properties of the autocovariance of the sequences within a sliding window, a least-squares optimization problem concerning the process noise covariance and the observation noise covariance is constructed, and its expression is:

[0029] In the formula, This represents the optimized process noise covariance. and observation noise covariance The parameter set, Represents the coefficient matrix. This represents the parameter set before optimization, which includes the process noise covariance and the observation noise covariance. This represents the sample covariance vector calculated based on the observed residual sequence within the sliding window. Denotes the Euclidean norm; Solve the least squares optimization problem with respect to process noise covariance and observation noise covariance in real time, and dynamically update the process noise covariance and observation noise covariance; The error covariance matrix is ​​updated in real time, and its expression is:

[0030] in, This represents the updated error covariance matrix; Describes the identity matrix, whose dimensions and error covariance matrix are... The dimensions are consistent; This represents the error covariance matrix before the update. The spatial uncertainty of the current target's predicted trajectory is quantified using the updated error covariance matrix and passed to the next frame to support the state tracking of the target in the next frame; Reset the unmatched counter for the current target predicted trajectory to 0.

[0031] ②The passive wake-up mechanism is as follows: Freeze the state updates of the tracked target to prevent the accumulation of Kalman filter errors; The local image and neighborhood background of the tracked target in the current frame, the historical motion trajectory coordinates and appearance feature sequence of the tracked target in the previous N frames, and the global scene information of the current frame are packaged and constructed into a multimodal contextual cue. Specifically, obtaining the local image of the tracked target and its neighborhood background in the current frame involves: based on the predicted bounding box of the target in the previous frame. Construct a contextual region of interest that includes environmental semantics, and introduce an expansion coefficient. Calculate the context capture region The coordinates of are expressed as:

[0032] in, The x-coordinate of the bounding box center. This represents the y-coordinate of the bounding box center. Indicates the width of the bounding box. Indicates the height of the bounding box.

[0033] Synchronous backtracking Constructing a spatiotemporal trajectory sequence from frame historical data :

[0034] In the formula, Indicates the first Frame coordinates, Indicates the first Frame speed, Indicates the first The feature vector of the frame, Represents the first N frames, This indicates the previous frame.

[0035] Input multimodal contextual prompts into the cognitive agent; Semantic fingerprints of tracked targets are generated using a multimodal large language model through a cognitive agent; The CoT (Cooperative Thought Chain) mechanism is used to perform logical reasoning and semantic association on tracked targets that have disappeared or reappeared, specifically including: An autoregressive decoding strategy is used to generate K candidate thought chain reasoning paths in parallel. Analyze the target state hypothesis corresponding to each candidate thought chain reasoning path, calculate the logical consistency confidence score of each target state hypothesis, and select the target state hypothesis with the highest logical consistency confidence score as the final reasoning conclusion. Its expression is:

[0036] in, For the final conclusion, To track the set of assumptions for the target state, Indicates the first One candidate thought chain reasoning path; For from the first A mapping function for extracting conclusions from candidate thought chain reasoning paths. This indicates the assumption about the state of the target being tracked. For indicator functions, when When true, its value is 1; otherwise, it is 0. Indicates the first The generation probability of candidate thought chain reasoning paths, Indicates multimodal context hints; A weighted voting strategy is used to eliminate illusory reasoning generated by the large language model in the final reasoning conclusion; Based on the reasoning conclusions after eliminating hallucinations, infer the logical location and probability of the tracked target's identity.

[0037] A generative inpainting layer based on a diffusion model is introduced, which combines physical motion features and semantic conditions to generate missing trajectories that conform to scene constraints. Specifically, this includes: The historical motion trajectory of the tracked target is projected onto a low-rank singular space (i.e., a principal component subspace representing the main motion patterns formed by the first k largest singular values ​​of the trajectory matrix and their corresponding left singular vectors) using singular decomposition to extract motion pattern features. The extracted motion pattern features are then concatenated with the logical position and identity probability of the tracked target to obtain a multi-condition guidance vector. An activation-conditional diffusion generative model, using multi-conditional guiding vectors as constraints, generates the residual distribution of missing trajectory segments in the latent space through cascaded denoising, and reconstructs a complete trajectory sequence that conforms to physical laws. Specifically, this includes: Initial anchor point paths are generated in singular space based on extracted motion pattern features; By injecting multi-conditional guiding vectors into the conditional diffusion generation model, and progressively predicting and correcting the residuals relative to the initial anchor point path through inverse denoising, a complete trajectory sequence conforming to physical laws is generated, the expression of which is:

[0038] In the formula, This represents a complete trajectory sequence that conforms to physical laws. Indicates the initial anchor path. This is the residual scaling factor. This indicates the number of steps in the reverse denoising process. This indicates the learning denoising network. Indicates the first The potential state of the step, Represents a multi-condition guiding vector; The generated complete trajectory sequence is inversely mapped back to the original spatiotemporal coordinate system to fill the broken time windows, and global trajectory smoothing is performed using the B-spline interpolation algorithm.

[0039] S5. Track targets based on adaptive decision-making.

[0040] In another embodiment of the present invention, the following specific tracking method is implemented based on the specific steps of the above embodiments.

[0041] At a busy urban intersection, a binocular monitoring system integrating a visible light camera and an infrared thermal imager is deployed. Using timestamp synchronization technology, the system collects real-time traffic and pedestrian flow data at the intersection around the clock, and transmits the synchronized visible light and infrared thermal imaging video streams in real time to a "hierarchical multimodal tracking system" based on this invention, deployed in the traffic control center.

[0042] During a rush hour surveillance mission, a pedestrian wearing a gray coat attempted to cross a crosswalk but was completely obscured by a large bus for five seconds before reappearing from the other side of the bus. The system's goal is to continuously and accurately identify target A in this complex environment with low light and severe obstruction, and to complete the pedestrian's trajectory during the period of obstruction. The system's workflow is as follows: Step S1: The front-end fast execution layer performs real-time detection and status monitoring.

[0043] Step S2: When an occlusion is encountered, an exception is triggered, and the system activates the cognitive agent in the backend inference execution module.

[0044] This includes background images of buses and streetlights, and traces back the spatiotemporal trajectory sequence of past frames, packaging it into a multimodal contextual cue and sending it to the backend.

[0045] Step S3: The cognitive agent performs logical reasoning and identity association.

[0046] After receiving the prompt words, the cognitive reasoning module quickly generates the semantic fingerprint of target A: "wearing a gray hoodie, carrying a black backpack, and moving north." Then, the thought chain reasoning engine is activated, and the large model outputs the reasoning path: "Target A is obscured by a bus traveling from east to west at time t. Based on its previous pace and direction, it is expected to appear in the rear area of ​​the bus after t+5 seconds."

[0047] The candidate target B, currently appearing at the rear of the bus, is blurred due to motion, but its features of 'gray clothing' and 'black backpack' are highly consistent with those of target A, and its location matches logical predictions. Based on this high-confidence inference, the system successfully associated the newly appearing target B with target A, maintaining consistency in identity.

[0048] Step S4: Generative repair layer completes the occlusion trajectory and outputs the complete path.

[0049] Although the identity was established, the trajectory for the middle 5 seconds was missing. The system projected the historical movement pattern of target A and combined it with the semantic condition "obscured by the bus and walking in a straight line" output by the cognitive agent. Using a conditional diffusion generative model, it generated the trajectory residual of the missing segment in the latent space. After inverse mapping and B-spline smoothing, the system drew a smooth dashed trajectory on the monitoring screen that conformed to physical inertia and avoided the volume constraints of the bus, perfectly reproducing the movement route of target A in the blind spot.

[0050] This embodiment fully demonstrates that the present invention can achieve high real-time continuous tracking in complex traffic scenarios with insufficient lighting and prolonged severe occlusion, solving the problems of identity loss and trajectory breakage. Furthermore, the present invention is also applicable to other scenarios requiring long-term continuous tracking and trajectory prediction of multiple targets.

[0051] To address the shortcomings of existing tracking algorithms, such as a lack of logical reasoning capabilities and susceptibility to identity loss or errors under severe occlusion or complex interactions, this invention introduces a cognitive agent based on a multimodal large model. Through a thought chain mechanism, the system can perform spatiotemporal causal logical reasoning and semantic association on the "disappearance and reappearance" process of the target, significantly improving identity maintenance capabilities under long-term occlusion. Furthermore, this invention utilizes a generative repair layer based on a diffusion model, guided by singular space projection and semantic conditions, to generate trajectory completion results that conform to physical motion inertia and scene constraints, effectively solving the problems of stiff and unreasonable trajectories caused by traditional linear interpolation.

[0052] The technical solution provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of the claims of this invention.

Claims

1. A hierarchical multimodal tracking method based on large-model cognitive drive, characterized in that, include: Enhanced fusion features are generated based on visible light and infrared thermal imaging images in surveillance scenarios; Based on enhanced fusion features, a probabilistic single-stage target detection network is used for target tracking to obtain target localization results; Based on the target location results, monitor and track the target's status signals in real time. Make adaptive decisions based on status monitoring signals; Adaptive decision-making includes a continuous tracking mechanism and a passive wake-up mechanism; Target tracking is performed based on adaptive decision-making. Specifically, the continuous tracking mechanism involves using an adaptive Kalman filter to update the trajectory state of the current frame. The passive wake-up mechanism is as follows: Freeze the state updates of the tracked target to prevent the accumulation of Kalman filter errors; The local image and neighborhood background of the tracked target in the current frame, the historical motion trajectory coordinates and appearance feature sequence of the tracked target in the previous N frames, and the global scene information of the current frame are packaged and constructed into a multimodal contextual cue. Input multimodal contextual prompts into the cognitive agent; Semantic fingerprints of tracked targets are generated using a multimodal large language model through a cognitive agent; The CoT (Cooperative Thinking Trace) mechanism is used to perform logical reasoning and semantic association on the tracked targets that have disappeared or reappeared. A generative repair layer based on a diffusion model is introduced, which combines physical motion features and semantic conditions to generate missing trajectories that conform to scene constraints.

2. The method according to claim 1, characterized in that, Specific methods for generating enhanced fusion features based on visible light and infrared thermal imaging images in surveillance scenarios include: Acquire and time-series align visible light and infrared thermal imaging images in the monitored scene; The time-aligned visible light and infrared thermal imaging images are segmented into fixed-size image blocks, and the image blocks are linearly projected into a sequence of feature tokens; We use a feature extraction network based on visual Fourier prompts and modal fusion prompt generators to extract features from the feature token sequence, thereby obtaining enhanced fusion features.

3. The method according to claim 2, characterized in that, The feature extraction network based on the visual Fourier cue and modal fusion cue generator extracts features from the feature token sequence to obtain enhanced fusion features. The specific method is as follows: Learnable cue vectors are embedded in the Transformer backbone network, and two-dimensional fast Fourier transforms are performed on the learnable cue vectors in the channel and spatial dimensions. The vectors after fast Fourier transforms are then mapped to the frequency domain space to obtain frequency domain cue. The feature token sequence is input into a pre-trained Transformer backbone network containing frequency domain cues for multi-level feature extraction to obtain visible light features and infrared features. The channel dimensions of visible light and infrared features are compressed using a shared linear projection layer; the two compressed features are fused using an addition operation; and the fused features are then subjected to feature interaction and dimension restoration using a convolution operation to obtain a unified modality fusion cue, the expression of which is: In the formula, This indicates a unified modal fusion suggestion. This represents the convolution operation. This represents a linear projection operation. Indicates the characteristics of visible light. Indicates infrared characteristics; The unified modal fusion cue is fused with visible light features and infrared features in residual form to obtain enhanced fusion features, including visible light enhanced fusion features and infrared enhanced fusion features, the expressions of which are: In the formula, This indicates visible light enhanced fusion characteristics. Indicates infrared enhanced fusion characteristics, This represents the residual form of a unified modal fusion cues.

4. The method according to claim 3, characterized in that, Based on enhanced fusion features, a probabilistic single-stage target detection network is used for target tracking to obtain the target localization result, as follows: The single-stage object detection network is pre-trained using negative log-likelihood loss. The expression for the negative log-likelihood loss function is as follows: In the formula, Indicates the loss value. This represents the mean coordinates of the predicted target bounding box. The bounding box representing the true target. This represents the target's localization covariance matrix. Represents the natural logarithm function; The enhanced fusion features corresponding to the current frame are input into a pre-trained single-stage object detection network, which outputs the bounding box coordinates of all potential targets in the current enhanced fusion features, their corresponding confidence scores, and the localization covariance matrix; the mean coordinates of the target bounding boxes are obtained; and the target localization results are modeled as a multivariate Gaussian distribution, the expression of which is: In the formula, This represents a multivariate Gaussian distribution, i.e., the target localization result; Indicated by For mean vector, It is a multivariate Gaussian distribution of the covariance matrix.

5. The method according to claim 4, characterized in that, Based on the target location results, the target's status monitoring signals are monitored and tracked in real time, specifically as follows: The target's predicted trajectory is obtained based on the real-time target location results; Based on the preset confidence threshold and the confidence level of each bounding box, the bounding boxes are divided into a high-score set and a low-score set; The high-scoring set is matched with the target predicted trajectory once, and the low-scoring set is matched with the unmatched target predicted trajectory a second time. Extract the matching quality features of the current frame, motion mutation features including cross-union ratio change rate and normalized velocity vector mutation value, occlusion overlap ratio features and cross-frame appearance consistency features; A multi-factor weighted fusion model is used to map the extracted features into a unified anomaly risk score, the expression of which is: In the formula, Indicates the abnormal risk score. This represents the weight coefficients corresponding to the matching quality features. Indicates matching quality features. This represents the weighting coefficients corresponding to the motion abrupt change features, which include the crossover ratio change rate and the normalized abrupt change value of the velocity vector. This represents the normalized abrupt change value of the velocity vector. Indicates the rate of change of the crossover ratio. Weighting coefficients representing the occlusion overlap ratio feature. This indicates the occlusion overlap ratio feature. The weighting coefficients represent the appearance consistency features across frames. Indicates cross-frame appearance consistency characteristics; The system compares the abnormal risk score with the preset safety risk threshold in real time and outputs a status signal. When the abnormal risk score is less than or equal to the preset safety risk threshold, the output status monitoring signal is 0; otherwise, the output status monitoring signal is 1.

6. The method according to claim 5, characterized in that, When the status monitoring signal is 0, the continuous tracking mechanism is triggered; when the status monitoring signal is 1, the passive wake-up mechanism is triggered. The continuous tracking mechanism specifically involves updating the trajectory state of the current frame using an adaptive Kalman filter, including: The matching relationship between the bounding box of the current frame and the predicted trajectory of the target is locked. Using the bounding boxes in the high-resolution set as observations, the state vector of the Kalman filter is updated posteriorly. The expression is as follows: In the formula, This represents the updated state vector posteriorly. For the prior prediction state, For Kalman gain, The observation matrix; By utilizing the statistical properties of the autocovariance of the sequences within a sliding window, a least-squares optimization problem concerning the process noise covariance and the observation noise covariance is constructed, and its expression is: In the formula, This represents the optimized process noise covariance. and observation noise covariance The parameter set, Represents the coefficient matrix. This represents the parameter set before optimization, which includes the process noise covariance and the observation noise covariance. This represents the sample covariance vector calculated based on the observed residual sequence within the sliding window. Denotes the Euclidean norm; Solve the least squares optimization problem with respect to process noise covariance and observation noise covariance in real time, and dynamically update the process noise covariance and observation noise covariance; The error covariance matrix is ​​updated in real time, and its expression is: in, This represents the updated error covariance matrix; Describes the identity matrix, whose dimensions and error covariance matrix are... The dimensions are consistent; This represents the error covariance matrix before the update. The spatial uncertainty of the current target's predicted trajectory is quantified using the updated error covariance matrix and passed to the next frame to support the state tracking of the target in the next frame; Reset the unmatched counter for the current target predicted trajectory to 0.

7. The method according to claim 6, characterized in that, The CoT (Cooperative Thought Chain) mechanism is used to perform logical reasoning and semantic association on tracked targets that have disappeared or reappeared. Specifically: An autoregressive decoding strategy is used to generate K candidate thought chain reasoning paths in parallel. Analyze the target state hypothesis corresponding to each candidate thought chain reasoning path, calculate the logical consistency confidence score of each target state hypothesis, and select the target state hypothesis with the highest logical consistency confidence score as the final reasoning conclusion. Its expression is: in, For the final conclusion, To track the set of assumptions for the target state, Indicates the first One candidate thought chain reasoning path; For from the first A mapping function for extracting conclusions from candidate thought chain reasoning paths. This indicates the assumption about the state of the target being tracked. For indicator functions, when When true, its value is 1; otherwise, it is 0. Indicates the first The generation probability of candidate thought chain reasoning paths, Indicates multimodal context hints; A weighted voting strategy is used to eliminate illusory reasoning generated by the large language model in the final reasoning conclusion; Based on the reasoning conclusions after eliminating hallucinations, infer the logical location and probability of the tracked target's identity.

8. The method according to claim 7, characterized in that, A generative inpainting layer based on a diffusion model is introduced, which combines physical motion features and semantic conditions to generate missing trajectories that conform to scene constraints. Specifically: Singular decomposition is used to project the historical motion trajectory of the tracked target into a low-rank singular space to extract motion pattern features; the extracted motion pattern features are then concatenated with the logical position and identity probability of the tracked target to obtain a multi-condition guidance vector. The activation conditional diffusion generative model, constrained by multiple conditional guiding vectors, generates the residual distribution of missing trajectory segments in the latent space through cascaded denoising, and reconstructs a complete trajectory sequence that conforms to physical laws. Specifically, this includes: Initial anchor point paths are generated in singular space based on extracted motion pattern features; By injecting multi-conditional guiding vectors into the conditional diffusion generation model, and progressively predicting and correcting the residuals relative to the initial anchor point path through inverse denoising, a complete trajectory sequence conforming to physical laws is generated, the expression of which is: In the formula, This represents a complete trajectory sequence that conforms to physical laws. Indicates the initial anchor path. This is the residual scaling factor. This indicates the number of steps in the reverse denoising process. This indicates the learning denoising network. Indicates the first The potential state of the step, Represents a multi-condition guiding vector; The generated complete trajectory sequence is inversely mapped back to the original spatiotemporal coordinate system to fill the broken time windows, and global trajectory smoothing is performed using the B-spline interpolation algorithm.

9. A system based on the hierarchical multimodal tracking method driven by large model cognition as described in any one of claims 1 to 8, characterized in that, include: A front-end fast execution module is used to generate enhanced fusion features based on visible light and infrared thermal imaging images in the monitoring scene; Based on enhanced fusion features, a probabilistic single-stage target detection network is used for target tracking to obtain target localization results; Based on the target location results, monitor and track the target's status signals in real time. The backend inference execution module is used to make adaptive decisions based on status monitoring signals; Target tracking is performed based on adaptive decision-making.

Citation Information

Patent Citations

  • Adaptive cross-modal visual tracking method based on soft selection

    CN114663470A

  • Large and small model collaborative unmanned aerial vehicle target tracking method

    CN119598140A

  • Deep learning method for multiple object tracking from video

    US20240144489A1