Multi-modal remote sensing target tracking positioning and intention discrimination method and device
By combining frame alignment of visible and infrared image frames with a multimodal detection and tracking model in UAV remote sensing videos, the problems of unstable target tracking and inaccurate positioning in remote sensing videos are solved, achieving high-precision target behavior and intent recognition, and improving intelligent monitoring and task early warning capabilities.
Patent Information
- Application Number
- CN202511657162.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-20
AI Technical Summary
Existing technologies suffer from unstable target tracking and inaccurate positioning in UAV remote sensing videos, and lack intelligent modeling and judgment of target behavior and intentions, making it difficult to meet the needs of intelligent monitoring and task early warning in complex scenarios.
By acquiring visible light and infrared light image frames and aligning them, a multimodal detection and tracking model and a back-projection mapping function are used, combined with a behavior recognition model, to achieve cross-frame target identity association and high-precision geographic coordinate positioning, and output the target's behavioral intent.
It improves the accuracy and positioning precision of target tracking in remote sensing videos, can accurately identify target behavior and intentions, and enhances the intelligent monitoring and task early warning capabilities in complex remote sensing scenarios.
Smart Images

Figure CN121708046A_ABST
Abstract
Description
Technical Field
[0001] This application relates to remote sensing technology, and more particularly to a method and apparatus for multimodal remote sensing target tracking, localization, and intent discrimination. Background Technology
[0002] Currently, with the rapid development of new-generation low-altitude remote sensing technology, UAV platforms, due to their advantages such as high mobility, low cost, and rapid response, are gradually becoming an important sensing means in fields such as security defense, urban management, coastal patrol, environmental monitoring, disaster emergency response, and wilderness search and rescue.
[0003] Compared to static remote sensing images, remote sensing videos acquired by UAVs are characterized by their continuity, multi-angle perspective, high resolution, and dynamic changes. Therefore, how to accurately track targets and identify behavioral intentions in remote sensing videos is a pressing issue that needs to be addressed. Summary of the Invention
[0004] This application provides a multimodal remote sensing target tracking, localization, and intent discrimination method and apparatus, which can improve the accuracy of target tracking and behavioral intent recognition in remote sensing videos.
[0005] The technical solution of this application embodiment is implemented as follows: This application provides a multimodal remote sensing target tracking, localization, and intent determination method, the method comprising: Multiple visible light image frames and multiple infrared image frames are acquired, and frame alignment is performed on each visible light image frame and each infrared image frame to obtain multiple sets of valid image frame pairs. For each set of valid image frame pairs, the tracking identification information of each detected object in the valid image frame pair is determined based on the valid image frame pair and the pre-trained multimodal detection and tracking model. For each detected object, the latitude and longitude trajectory of the detected object is determined based on the tracking identification information and the back projection mapping function. The back projection mapping function is determined based on the camera imaging model and multi-source pose data. The behavioral intention of each detected object is determined based on the behavior recognition model and the latitude and longitude trajectories of each detected object.
[0006] This application provides a multimodal remote sensing target tracking, localization, and intent determination device, including: The acquisition module is used to acquire multiple visible light image frames and multiple infrared light image frames, and to perform frame alignment operations on each visible light image frame and each infrared light image frame to obtain multiple sets of valid image frame pairs; The first determining module is used to determine the tracking identification information of each detected object in each valid image frame pair based on the valid image frame pair and the pre-trained multimodal detection and tracking model. The second determination module is used to determine the latitude and longitude trajectory of each detected object based on the tracking identification information and the back projection mapping function; wherein the back projection mapping function is determined based on the camera imaging model and multi-source pose data. The third determination module is used to determine the behavioral intent of each detected object based on the behavior recognition model and the latitude and longitude trajectories of each detected object.
[0007] This application provides an electronic device, which includes: Memory is used to store executable instructions or computer programs. When a processor executes computer-executable instructions or computer programs stored in memory, it implements the multimodal remote sensing target tracking, localization, and intent discrimination method provided in the embodiments of this application.
[0008] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the multimodal remote sensing target tracking, localization, and intent determination method provided in this application.
[0009] This application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, they implement the multimodal remote sensing target tracking, positioning, and intent discrimination method provided in this application.
[0010] The embodiments of this application have the following beneficial effects: In the multimodal remote sensing target tracking, localization, and intent discrimination method provided in this application embodiment, multiple sets of high-quality, effective image frame pairs are constructed through frame alignment operations of visible light and infrared images, providing stable input for subsequent multimodal collaborative modeling. Secondly, a pre-trained multimodal detection and tracking model is used, combined with target detection and multi-target tracking mechanisms, to achieve cross-frame target identity association, improving occlusion recovery and tracking stability, thereby enhancing the accuracy of target tracking. Furthermore, pixel coordinates are converted to geographic coordinates using a back-projection mapping function to obtain latitude and longitude trajectories, achieving high-precision location reconstruction. Finally, based on the latitude and longitude trajectories, the target's behavioral intent is output through temporal modeling. This overcomes the problems of unstable target tracking and inaccurate localization in related remote sensing videos, and accurately identifies the target's behavioral intent, thereby improving intelligent monitoring and task early warning capabilities in complex remote sensing scenarios. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0012] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0013] Figure 1 A flowchart illustrating a multimodal remote sensing target tracking, localization, and intent determination method provided in this application embodiment. Figure 1 ; Figure 2 A flowchart illustrating a multimodal remote sensing target tracking, localization, and intent determination method provided in this application embodiment. Figure 2 ; Figure 3 A schematic diagram of the overall process of a multimodal remote sensing target tracking, localization, and intent discrimination method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the composition structure of a multimodal remote sensing target tracking, localization, and intent discrimination device provided in an embodiment of this application; Figure 5 This is a schematic diagram of the hardware entity of an electronic device provided in an embodiment of this application.
[0014] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation
[0015] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0016] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0018] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0019] In this embodiment, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0020] Currently, with the rapid development of next-generation low-altitude remote sensing technology, unmanned aerial vehicle (UAV) platforms, due to their advantages such as high mobility, low cost, and rapid response, are gradually becoming an important sensing tool in fields such as national security defense, urban management, coastal patrol, environmental monitoring, disaster emergency response, and wilderness search and rescue. Compared with traditional static remote sensing images, remote sensing videos acquired by UAVs have characteristics such as continuity, multi-angle, high resolution, and dynamic changes, providing possibilities for target behavior analysis and real-time intervention. However, in the process of moving from "image perception" to "dynamic understanding," existing technologies still face many challenges in target tracking, geolocation, and behavioral intent recognition.
[0021] In related technologies, research on remote sensing targets mainly focuses on detection and recognition tasks on static images, relying on predefined target categories and a large amount of manually labeled data. Typical methods achieve good detection accuracy in high-resolution remote sensing images, but often perform poorly in video stream data. This is because video data has characteristics such as continuous inter-frame interference, motion blur, drastic scale changes, and frequent occlusion, making it difficult for related technologies to model the temporal evolution characteristics of targets, leading to problems such as unstable detection and target loss. Especially in complex scenes, such as forest cover, water reflections, and densely built-up urban areas, targets are easily interfered with by the background or occluded due to changes in viewing angle, further exacerbating the detection difficulties.
[0022] Meanwhile, most target tracking methods in related technologies originate from the field of visual tracking in natural scenes. While these methods have achieved good results in video data, they are difficult to directly apply to remote sensing video scenarios. On the one hand, remote sensing videos contain a large number of objects with similar target types but vastly different scales, such as different models of vehicles or ships, making target discrimination difficult. On the other hand, the flight trajectories of UAV platforms are constantly changing, with frequent changes in perspective, resulting in continuous background changes. Traditional appearance trackers have significant shortcomings in modeling inter-frame consistency. Furthermore, common tracking algorithms do not incorporate spatial geographic information, only outputting pixel-level coordinates, making them difficult to directly use for position tracking and scheduling decisions in the real world.
[0023] More importantly, most remote sensing monitoring systems lack the ability to analyze the behavioral intentions of targets. In real-world scenarios, simply identifying the target's category or current coordinates is insufficient for practical applications. For example, border monitoring systems not only need to identify the presence of illegally crossing vehicles, but also determine whether they are in states such as "approaching the boundary," "attempting to conceal themselves," or "escaping at high speed," in order to achieve more intelligent decision support. This requires the system to have the ability to abstract behavioral patterns from visual information and perform intention recognition. Currently, intention recognition is mostly based on preset rules, which is difficult to adapt to the complexity and dynamism of target behavior.
[0024] On the other hand, video data collected by drones typically includes a large amount of metadata, such as timestamps, GPS location, attitude angles, heading, and camera intrinsic and extrinsic parameters. Current technologies often ignore this data when processing video, or treat it only as supplementary information for coarse processing, failing to construct rigorous geometric mapping models. This makes it impossible to achieve high-precision latitude and longitude location reconstruction of targets in the images. Especially in tasks such as surveillance and port identification, the precise location of targets is often the foundation for command execution and action deployment. Therefore, providing only pixel coordinates or coarse map positioning information cannot meet practical needs.
[0025] With breakthroughs in multimodal large-scale models in computer vision and natural language processing, new paradigms have emerged that integrate visual, linguistic, and temporal information for perception and reasoning. These models, through joint modeling of images and text, achieve cross-modal understanding from descriptive language to visual objects, providing new insights for querying, retrieving, and assisted localization of remotely sensed targets. However, these models still primarily rely on static images and single descriptions, lacking the ability to dynamically model target behavior in videos, and are almost entirely absent from fusing visual detection results with geospatial coordinates.
[0026] Despite the progress made in current remote sensing intelligent recognition technology, the following prominent technical challenges remain in applications for dynamic target monitoring using drone video: (1) Video-level target detection and tracking are unstable, and occlusion recovery and target consistency modeling are insufficient; (2) The target lacks a precise mapping mechanism from pixel coordinates to geographic coordinates; (3) The understanding of target behavior remains at the static recognition stage, lacking intelligent modeling and judgment of its potential intentions.
[0027] Based on this, embodiments of this application provide a multimodal remote sensing target tracking, localization, and intent determination method. The method includes: acquiring multiple visible light image frames and multiple infrared light image frames, and performing frame alignment operations on each visible light image frame and each infrared light image frame to obtain multiple sets of valid image frame pairs; for each set of valid image frame pairs, determining the tracking identification information of each detected object in the valid image frame pair based on the valid image frame pair and a pre-trained multimodal detection and tracking model; for each detected object, determining the latitude and longitude trajectory of the detected object based on the tracking identification information and a back projection mapping function; wherein, the back projection mapping function is determined based on a camera imaging model and multi-source pose data; and determining the behavioral intent of each detected object based on a behavior recognition model and the latitude and longitude trajectories of each detected object.
[0028] In this way, by aligning visible light and infrared images, multiple sets of high-quality, effective image frame pairs are constructed, providing stable input for subsequent multimodal collaborative modeling. Secondly, by utilizing a pre-trained multimodal detection and tracking model, combined with target detection and multi-target tracking mechanisms, cross-frame target identity association is achieved, improving occlusion recovery and tracking stability, thereby enhancing target tracking accuracy. Furthermore, a back-projection mapping function is used to convert pixel coordinates into geographic coordinates to obtain latitude and longitude trajectories, achieving high-precision location reconstruction. Finally, based on the latitude and longitude trajectories, temporal modeling outputs the target's behavioral intent. This approach overcomes the problems of unstable target tracking and inaccurate positioning in remote sensing videos using related technologies, while also accurately identifying the target's behavioral intent, thus improving intelligent monitoring and task early warning capabilities in complex remote sensing scenarios.
[0029] The technical solutions in the embodiments of this application will now be clearly and completely described with reference to the accompanying drawings.
[0030] It should be noted that the multimodal remote sensing target tracking, localization, and intent determination method provided in the embodiments of this application can be executed by an electronic device. This electronic device can be a drone platform, a ground monitoring terminal, or an edge computing device. In other words, the multimodal remote sensing target tracking, localization, and intent determination method in the embodiments of this application can be executed by a drone platform, a ground monitoring terminal, or through collaborative interaction between an edge computing device and the cloud.
[0031] Figure 1A flowchart illustrating a multimodal remote sensing target tracking, localization, and intent determination method provided in this application embodiment. Figure 1 The following will combine Figure 1 The steps shown will be explained. It should be noted that... Figure 1 The method described uses a drone platform as the execution subject as an example for illustration, such as... Figure 1 As shown, the method may include S101 to S104, wherein: S101: Acquire multiple visible light image frames and multiple infrared light image frames, and perform frame alignment operation on each visible light image frame and each infrared light image frame to obtain multiple sets of valid image frame pairs.
[0032] Here, a visible light image frame refers to multiple images extracted from a visible light video; visible light image frames can be acquired by an optical camera mounted on a drone.
[0033] Infrared image frames refer to multiple images extracted from infrared video; infrared image frames can be acquired by thermal imaging sensors.
[0034] It is understandable that the two sensors can collect data synchronously and ensure the temporal consistency of the two images by aligning the timestamps.
[0035] It can also be understood that frame alignment operation refers to making infrared image frames and visible light image frames at the same time completely correspond in spatial coordinates through geometric transformation; for example, frame alignment operation can include spatial registration and geometric correction, so that images of two different bands are kept in the same spatial position, thereby obtaining effective image frame pairs, which can also be called multimodal image frame pairs.
[0036] It can also be understood that a valid image frame pair refers to a pair of visible light image frames and infrared light image frames that are strictly corresponding in time and space after frame alignment; that is, a valid image frame pair can include aligned visible light image frames and aligned infrared light image frames.
[0037] For example, effective image frame pairs can provide more robust target detection capabilities under different environmental conditions (such as changes in lighting and weather effects). For instance, in nighttime or foggy conditions, visible light images may be blurry due to insufficient light, but infrared images can still clearly capture the target outline; while under strong sunlight, infrared images may be distorted due to overheating and reflection, in which case visible light images can provide more accurate boundary information. Therefore, the introduction of multimodal image frame pairs can improve the perception capabilities and robustness of UAV platforms in complex environments.
[0038] In one possible implementation, the above-mentioned S101 "acquiring multiple visible light image frames and multiple infrared light image frames" may further include the following steps: S1011 acquires visible light video data and infrared video data synchronously collected by the drone.
[0039] It is understandable that drones can be equipped with multimodal sensors that can simultaneously acquire video data from both visible and infrared light channels.
[0040] It should be noted that the basic requirements for video data are a frame rate of ≥25fps, a resolution of ≥1920×1080, and a duration of ≥10 seconds. Of course, custom settings can be made according to the scenario, and this application embodiment does not limit this.
[0041] Visible light video data can reflect the appearance and color information of a target, while infrared light video data can capture thermal radiation signals, which is helpful for identifying camouflaged or concealed targets.
[0042] In this way, by simultaneously acquiring visible light and infrared video data, complementary information from visible light and infrared images can be fused in subsequent processing, thereby improving the accuracy and stability of target recognition and maintaining good perception capabilities under complex lighting conditions (such as nighttime, foggy days, and strong reflections).
[0043] S1012, preprocess the visible light video data and infrared light video data to determine multiple visible light image frames corresponding to the visible light video data and multiple infrared light image frames corresponding to the infrared light video data.
[0044] Here, the preprocessing process may include video decoding, keyframe extraction, geometric registration, and low-quality frame removal.
[0045] In some embodiments, the original video stream can be decoded into single-frame images; then, keyframes are extracted at a set frequency (e.g., 5 frames per second) to reduce computational load and retain key dynamic information; next, all visible and infrared images are geometrically registered to ensure spatial consistency; finally, low-quality frames are removed based on factors such as image sharpness and occlusion level, resulting in a high-quality image frame sequence. This preprocessing of visible and infrared video data ensures data quality and spatiotemporal consistency for subsequent processing.
[0046] In this embodiment, by simultaneously acquiring visible light and infrared video data and preprocessing the data, the robustness and versatility of target detection can be significantly improved. Simultaneous acquisition and preprocessing of visible light and infrared video data effectively addresses illumination changes and occlusion issues in complex environments, thereby enhancing the system's environmental adaptability and providing more reliable foundational data for subsequent tracking, localization, and intent determination.
[0047] Of course, multiple visible light image frames and multiple infrared light image frames can also be acquired simultaneously according to a preset time interval, but this application embodiment does not limit this.
[0048] S102, for each pair of valid image frames, based on the valid image frame pair and the pre-trained multimodal detection and tracking model, determine the tracking identification information of each detected object in the valid image frame pair.
[0049] Here, the multimodal detection and tracking model is a deep learning model that combines visible light and infrared image features. The multimodal detection and tracking model can be pre-trained based on historical valid image frames and the tracking identifier information of each detected object within those historical valid image frames.
[0050] It is understandable that tracking identification information is used to uniquely identify the identity features of a target object (i.e., each detected object) in consecutive video frames.
[0051] For example, in one frame of an image, a vehicle is occluded, but when it reappears in the next frame, it can be matched with the vehicle's historical trajectory and current detection features to assign the same tracking identification information to the vehicle, thereby ensuring the continuity of the vehicle throughout the video sequence.
[0052] In one possible implementation, aligned valid image frame pairs can be input into a multimodal detection and tracking model, which then directly outputs the tracking identification information of each detected object.
[0053] In another possible implementation, the multimodal detection and tracking model may include a target detection model and a multi-target tracking model; aligned valid image frame pairs may be input to the target detection model, which outputs the target detection result; further, the target detection result may be input to the multi-target tracking model, which outputs the tracking identification information of each detected object.
[0054] S103, for each detected object, determine the latitude and longitude trajectory of the detected object based on the tracking identification information and the back projection mapping function.
[0055] Here, the back projection mapping function is a mapping relationship from image pixel coordinates to real geographic coordinates (also known as latitude and longitude coordinates) constructed based on camera imaging models (such as pinhole models) and multi-source pose data.
[0056] It is understandable that multi-source attitude data may include, but is not limited to, the UAV's GPS coordinates, flight altitude, attitude angles, etc.
[0057] Alternatively, latitude and longitude trajectories can be understood as indicating a sequence of locations of a monitored object (such as a vehicle or pedestrian) within a preset time period (the time it is continuously tracked) in the Earth's real geographic coordinate system. Each location is represented by latitude and longitude coordinates (usually also including elevation) and a timestamp.
[0058] In one possible implementation, multiple image pixel coordinates of the detected object can be determined based on tracking identification information, and the latitude and longitude trajectory of the detected object can be determined according to the back projection mapping function and the multiple image pixel coordinates.
[0059] In another possible implementation, multiple image pixel coordinates of the detected object can be determined based on the tracking identification information, and these multiple image pixel coordinates can be input into a pre-trained neural network model to obtain the latitude and longitude trajectory of the detected object.
[0060] S104, based on the behavior recognition model and the latitude and longitude trajectories of each detected object, determines the behavioral intent of each detected object.
[0061] Here, a behavior recognition model refers to a deep learning-based method that uses the trajectory features and temporal information of a target to identify the possible behavioral intentions of the detected object. For example, the structure of a behavior recognition model may include a bidirectional long short-term memory network and an attention mechanism to extract temporal features and output a probability distribution of behavior categories.
[0062] It is understandable that the behavioral intent of the detected object is used to indicate the potential behavioral purpose reflected by the movement trajectory of the detected target over a period of time, which may include, but is not limited to, "normal driving", "abnormal stop", "rapid escape", "approaching the boundary" and "attempting to conceal".
[0063] In one possible implementation, feature extraction can be performed on the latitude and longitude trajectory of the detected object. The kinematic features of the latitude and longitude trajectory can be extracted, including but not limited to the velocity, direction angle, acceleration, and uniform motion time period of the detected object. A time series feature vector sequence is constructed and input into the behavior recognition model to obtain the behavior intention of the detected object.
[0064] In another possible implementation, the latitude and longitude trajectory of the detected object can be directly input into the behavior recognition model. The behavior recognition model can generate multiple possible paths for the object in the future, with each path representing a potential intent hypothesis. The generated possible future paths are matched and probabilistically evaluated against a library of known typical behavior patterns. An intent confidence score is calculated for each hypothetical path, and the intent with the highest confidence score is selected as the most likely behavioral intent. As new trajectory data arrives, the prediction and intent judgment are continuously updated.
[0065] It is understandable that behavior recognition models can be pre-trained based on historical latitude and longitude trajectories and the corresponding behavioral intentions.
[0066] In another possible implementation, time series feature vectors and context label vectors can be determined based on latitude and longitude trajectories and the corresponding scene semantic labels. The time series feature vectors and context label vectors are then input into the behavior recognition model to determine the behavioral intent of each detected object.
[0067] In this embodiment, multiple high-quality, effective image frame pairs are constructed through frame alignment operations on visible light and infrared images, providing stable input for subsequent multimodal collaborative modeling. Secondly, a pre-trained multimodal detection and tracking model, combined with target detection and multi-target tracking mechanisms, is used to achieve cross-frame target identity association, improving occlusion recovery and tracking stability, thereby enhancing target tracking accuracy. Furthermore, pixel coordinates are converted to latitude and longitude coordinates through a back-projection mapping function to obtain latitude and longitude trajectories, achieving high-precision position reconstruction. Finally, based on the latitude and longitude trajectories, the target's behavioral intent is output through temporal modeling. This overcomes the problems of unstable target tracking and inaccurate positioning in remote sensing videos using related technologies, and accurately identifies the target's behavioral intent, thereby improving intelligent monitoring and task early warning capabilities in complex remote sensing scenarios.
[0068] In some embodiments, after step S101, "acquiring multiple visible light image frames and multiple infrared light image frames, and performing frame alignment operations on each visible light image frame and each infrared light image frame to obtain multiple sets of valid image frame pairs", the following steps may be included: Acquire metadata collected by the aircraft's sensors; For each pair of valid image frames, the metadata of the valid image frame pair under its corresponding timestamp is bound to obtain the sample sequence associated with the valid image frame pair.
[0069] Here, aircraft sensors may include, but are not limited to, GPS modules, inertial measurement units, barometers, attitude sensors, etc., for real-time acquisition of various status information of the aircraft during flight.
[0070] Metadata is used to indicate aircraft sensor data acquired during image acquisition, including timestamps, GPS coordinates, altitude, heading angle, pitch angle, roll angle, and camera intrinsic and extrinsic parameters.
[0071] By collecting this metadata, it is possible to model the spatiotemporal consistency between image content and the real world, thereby improving target localization accuracy and behavior understanding capabilities. For example, in low-altitude remote sensing scenarios, combining the aircraft's attitude and camera parameters can more accurately reconstruct the actual geographical location of targets in the image.
[0072] Here, the sample sequence associated with a valid image frame pair is a structured data set consisting of image frames and their corresponding metadata. Each image frame contains a timestamp of the time it was acquired, and the corresponding aircraft state information at that timestamp is bound to it. This binding method ensures that the spatial and temporal context of the image frame and its environment are consistent, thereby supporting subsequent target detection, tracking, geographic inversion, and intent determination tasks.
[0073] In this way, the establishment of associated sample sequences for effective image frame pairs allows for the uniform processing of each frame in the video stream, thereby enhancing data traceability and consistency. For example, during target tracking, the associated sample sequences for effective image frame pairs can provide information on the time intervals and spatial variations between consecutive frames during the tracking process, which helps the model better model the target's trajectory.
[0074] Correspondingly, the above-mentioned S102 "For each pair of valid image frames, based on the valid image frame pair and the pre-trained multimodal detection and tracking model, determine the tracking identification information of each detected object in the valid image frame pair" may include the following steps: Based on the sample sequence associated with effective image frame pairs and the pre-trained multimodal detection and tracking model, the tracking identification information of each detected object in the sample sequence associated with effective image frame pairs is determined.
[0075] In one possible implementation, the sample sequence associated with the effective image frame pairs can be input into the multimodal detection and tracking model, and then the multimodal detection and tracking model can directly output the tracking identification information of each detected object.
[0076] In another possible implementation, the multimodal detection and tracking model can include a target detection model and a multi-target tracking model; the sample sequence associated with effective image frame pairs can be input into the target detection model, and the target detection model outputs the target detection result; further, the target detection result is input into the multi-target tracking model, and the multi-target tracking model outputs the tracking identification information of each detected object.
[0077] In this embodiment, metadata collected by aircraft sensors is bound to image frames to form a sample sequence associated with valid image frame pairs, and multimodal detection and tracking are performed based on this sample sequence. This method is beneficial for recovering the identity of occluded targets under complex occlusion or rapidly changing viewpoint conditions, reducing ID switching issues, and improving the continuity and accuracy of target recognition, thereby enhancing the tracking effect of the detected object.
[0078] In some embodiments, the above-mentioned S102 "determining the tracking identification information of each detected object in the effective image frame pair based on the effective image frame pair and the pre-trained multimodal detection and tracking model" may further include the following steps: S1021, Target detection is performed on the effective image frame pairs based on the target detection model to obtain the target detection results of each detected object in the effective image frame pairs.
[0079] Here, the object detection model refers to a pre-trained deep learning model used to identify objects in an image and output the object category, bounding box (i.e., image pixel coordinates), and the confidence score corresponding to each object category. For example, the object detection model can be implemented based on convolutional neural networks, such as YOLO, Faster R-CNN, and other architectures.
[0080] It is understandable that the target detection model can process visible light and infrared dual-channel images collected by UAVs, and has strong robustness, especially in complex backgrounds, lighting changes or occlusion conditions, it can still maintain high detection accuracy.
[0081] Here, the object detection results may include the object category, image pixel coordinates, and the confidence score corresponding to the object category.
[0082] The target category can include common remote sensing target types such as vehicles, ships, pedestrians, and buildings. Image pixel coordinates refer to the position of the detected target on the image plane, which can be represented by a rectangular bounding box [x1, y1, x2, y2]. Confidence reflects the probability that the target detection model believes the detected target belongs to a certain category; a higher value indicates that the target detection model is more confident in the correctness of a particular category.
[0083] In some embodiments, visible light image frames and infrared light image frames can be input into the target detection model respectively. Then, the target detection model outputs the target detection result. The output format can be {category C, bounding box [x1,y1,x2,y2], confidence s}.
[0084] It's important to note the close data dependency between the object detection model and the subsequent multi-object tracking model. The object detection model provides all detected targets and their attributes in the current frame, while the multi-object tracking model utilizes this information, combined with the state feature vectors of each detected object from historical frames, to perform cross-frame target matching and identity association. For example, in consecutive video frames, if a car is detected in two frames at similar locations, the multi-object tracking model will treat it as the same target and assign it the same tracking ID. This ensures the continuity and consistency of the target throughout the video sequence, preventing target loss due to changes in viewpoint or brief occlusion.
[0085] It should also be noted that in practical applications, the target detection model can be deployed on edge computing devices or cloud servers to support real-time processing of video streams transmitted from drones. Through the initial screening by the target detection model, invalid areas can be quickly filtered out, allowing resources to be concentrated on processing potential targets, thereby improving the overall system's response speed and resource utilization.
[0086] S1022, based on the multi-target tracking model, associates the target detection results of each detected object with the state feature vector of each detected object in the historical N frames to determine the tracking identification information of each detected object.
[0087] Here, multi-target tracking models refer to models based on temporal modeling. These models are used to maintain the consistency of target identity and trajectory integrity across consecutive video frames. Multi-target tracking models typically combine bidirectional attention mechanisms with Hungarian matching algorithms to achieve dynamic association of targets across frames.
[0088] Here, the state feature vector is used to characterize the set of behavioral features of the detected object in the time dimension. The state feature vector may include the appearance feature vector and motion feature vector of the detected object.
[0089] The appearance feature vector can include visual information such as color histogram, texture features, local binary pattern (LBP), depth features (such as features extracted by ResNet), and size, used to characterize the appearance stability of the detected object. The motion feature vector includes the target's velocity, orientation angle, acceleration, curvature of the movement path, etc., used to describe the dynamic behavior of the detected object.
[0090] It's understandable that combining appearance feature vectors and motion feature vectors enables multi-target tracking models to more accurately identify whether targets are the same entity in complex dynamic scenes. For example, when two cars of the same model appear in the same frame, it might be difficult to distinguish them using only appearance feature vectors. However, by adding motion feature vectors, the movement trajectories of the two cars can be used to further confirm their independence. The combination of appearance and motion feature vectors not only improves the tracking accuracy of multi-target tracking models but also enhances their ability to recover from target occlusion.
[0091] In some embodiments, the target detection results of each detected object and the state feature vectors of each detected object in the historical N frames can be input into the multi-target tracking model to determine the similarity matrix between each detected object and each detected object in the historical N frames; and the tracking identification information of each detected object can be determined based on the similarity matrix.
[0092] It should be noted that in practical applications, multi-target tracking models can also use Kalman filters to predict and smooth the state of the detected objects, reducing the impact of inter-frame jitter. Furthermore, trajectory length limits and confidence thresholds can be introduced to eliminate false detections or low-quality tracking results, thereby ensuring that the final output tracking identification information has high reliability and stability.
[0093] In this embodiment, the basic attributes of the detected objects are extracted by the target detection model, and the state feature vectors of each detected object in historical frames are dynamically correlated by the multi-target tracking model. This efficiently determines the tracking identification information of each detected object. In this way, the consistency of the target's identity in the video sequence can be stably maintained, thereby accurately constructing the target's spatiotemporal trajectory, and providing reliable basic data for subsequent behavior intent determination and geolocation.
[0094] In some embodiments, the above-mentioned S1022 "based on the multi-target tracking model, associating the target detection results of each detected object with the state feature vectors of each detected object's historical N frames to determine the tracking identification information of each detected object" may further include the following steps: The target detection results of each detection object and the state feature vector of each detection object in the historical N frames are input into the multi-target tracking model to determine the similarity matrix between each detection object and each detection object in the historical N frames. Here, the similarity matrix is used to indicate a two-dimensional matrix generated by calculating the similarity between each detected target in the current frame and all tracked targets in the past N frames. Each element in the similarity matrix represents the probability of a match between two targets, with larger values indicating that they are more likely to be the same target.
[0095] It is understandable that cosine similarity, Euclidean distance, or other distance metrics can be used to quantify the degree of matching between targets during the construction of the similarity matrix. The similarity matrix provides the basic data support for the subsequent Hungarian matching algorithm, thereby making the allocation of targets and tracking IDs more accurate.
[0096] In some embodiments, the target detection result can be compared with the state feature vector of each detected object in the target detection result for N historical frames, thereby effectively distinguishing whether the target is occluded, lost, or newly appeared, and adjusting the tracking process according to the judgment result to optimize the tracking result.
[0097] Furthermore, after determining the similarity matrix, the tracking identification information of each detected object can be determined based on the similarity matrix and the Hungarian matching algorithm.
[0098] Here, the Hungarian matching algorithm is a graph theory algorithm that can be applied to the maximum weight matching problem in bipartite graphs.
[0099] For example, the Hungarian matching algorithm can be used to optimally match the target detected in the current frame with the tracked target that already exists in the historical frames. In other words, the Hungarian matching algorithm can find an optimal set of matching combinations by optimizing the similarity matrix, thereby minimizing the total matching cost.
[0100] Using the Hungarian matching algorithm, a unique tracking identifier (Track ID) can be assigned to each detected target in each frame. The tracking identifier is used to uniquely identify the detected target throughout the entire video sequence. Even if the detected target is temporarily occluded or reappears after leaving the field of view, the identity information corresponding to the tracking identifier (Track ID) can be recovered.
[0101] In this embodiment, the multi-target tracking model and Hungarian matching algorithm can be used to associate the target with historical frames. By assigning a unique tracking identifier to each detected target, the ID switching problem common in traditional tracking methods is avoided. This approach can adapt to the frequent changes in viewpoint and target occlusion in UAV videos, ensuring continuous tracking of the target in complex environments and improving the consistency and stability of target tracking.
[0102] In some embodiments, the above-mentioned S1021 "performing target detection on effective image frame pairs based on the target detection model to obtain the target detection results of each detected object in the effective image frame pairs" may further include the following steps: The aligned visible light image frame is input into the target detection model to obtain the first detection result of each detection object in the aligned visible light image frame; The aligned infrared image frame is input into the target detection model to obtain the second detection result of each detected object in the aligned infrared image frame; The first and second detection results are fused together to obtain the target detection result.
[0103] It is understood that a valid image frame pair may include an aligned visible light image frame and an aligned infrared light image frame.
[0104] Here, the first detection result refers to the identification result of all detected objects in the aligned visible light image frame based on the object detection model. The second detection result refers to the identification result of all detected objects in the aligned infrared light image frame based on the object detection model.
[0105] It is understandable that both the first and second detection results contain the target category C, the bounding box [x1,y1,x2,y2], and the confidence level s.
[0106] Here, fusion processing refers to the comprehensive analysis and integration of the first detection result from the visible light image and the second detection result from the infrared light image to eliminate the uncertainty present in the two independent detection results from the visible light and infrared light images, and to extract more stable and accurate target information. For example, fusion processing methods may include strategies such as weighted averaging, confidence threshold screening, and non-maximum suppression.
[0107] As can be understood, the object detection result is the final detection result after fusion processing. This result includes the category, location, and confidence information of all credible objects in the image. Compared to single-modal detection results, object detection results have higher accuracy and robustness.
[0108] In this embodiment, visible light and infrared light images are input into the target detection model respectively, and the detection results of the visible light and infrared light images are fused to achieve more accurate target identification. This method overcomes the limitations of single-modal images in specific scenarios, thereby improving adaptability and detection performance in complex environments, and providing high-quality data support for subsequent trajectory modeling and behavioral intent judgment.
[0109] In some embodiments, Figure 2 A flowchart illustrating a multimodal remote sensing target tracking, localization, and intent determination method provided in this application embodiment. Figure 2 ,like Figure 2 As shown, the above-mentioned S104 "determining the behavioral intent of each detected object based on the behavior recognition model and the latitude and longitude trajectories of each detected object" may include the following steps: S1041, extract features from the latitude and longitude trajectories of each detected object to obtain the time series feature vector of each detected object.
[0110] Here, latitude and longitude trajectories are records of the changing positions of each detected object in geographic space over time, usually consisting of a series of points represented in the WGS-84 coordinate system.
[0111] In some embodiments, kinematic feature extraction can be performed on latitude and longitude trajectories to extract key features such as velocity, direction angle, acceleration, and uniform motion time intervals, thereby forming a time series feature vector F=[f1,f2,...,fn].
[0112] It is understandable that time series feature vectors can reflect the dynamic behavioral characteristics of each detected object and serve as the basic input for subsequent intent determination.
[0113] S1042 performs spatial matching between latitude and longitude trajectories and scene semantic labels to determine the context label vector for each detected object. The scene semantic labels include roads, bushes, buildings, and boundary lines.
[0114] Here, scene semantic labels can be used to describe the geographical environment in which each detected object is located. For example, a road indicates that each detected object is in a passable area, and a boundary line indicates that the boundary line is close to a restricted area.
[0115] In some embodiments, a spatial matching algorithm can be used to compare the latitude and longitude trajectories of each detected object with the scene semantic labels on the map, thereby determining the current context label vector S=[s1,s2,...,sn] of each detected object.
[0116] The context label vector is a structured representation of the spatial matching results of each detected object, providing additional environmental information for behavior recognition. For example, in urban security scenarios, if a detected object lingers near bushes for an extended period, its context label vector may be flagged as an attempt to conceal itself.
[0117] S1043, based on multi-channel embedding, the time series feature vectors of each detected object and the context label vectors of each detected object are projected into the same embedding space to obtain the target vectors of each detected object.
[0118] Here, multichannel embedding is a deep learning network that can map features from different sources (such as time series features and context labels) into a shared high-dimensional vector space, enabling features from different sources to be compared and fused within a unified framework.
[0119] The target vector is the final feature representation, which includes the motion features of each detected object and the context label vector of the detected object.
[0120] S1044, Input the target vector of each detected object into the behavior recognition model to determine the behavioral intent of each detected object.
[0121] In some embodiments, the target vector of each detected object can be input into the behavior recognition model to output the behavioral intent of each detected object.
[0122] Behavior recognition models can identify various behavioral intention states. For example, "normal driving" means that each detected object moves along a regular route, while "attempt to conceal" means that each detected object tries to hide its position.
[0123] In some embodiments, to improve the discriminative robustness of the behavior recognition model, a behavior prior graph can be introduced to model the joint probability of historical behaviors between the target and the scene, and the KL divergence (also known as relative entropy) can be used to guide the convergence of the behavior recognition model output. Furthermore, the behavior intent prediction results can be post-processed, including but not limited to strategies such as time smoothing, confidence filtering, and category jump restrictions, thereby improving prediction stability.
[0124] As can be understood, a behavior prior graph is a quantified, historical data-based empirical knowledge base used to characterize the probability of different types of target (i.e., the object being detected) behavioral intentions in a specific scenario. By constructing a behavior prior graph, behavior recognition models can acquire prior knowledge about target behavioral patterns, thereby enabling more accurate and robust predictions on new observation data.
[0125] It can also be understood that joint probabilistic modeling is used to indicate the probability of a target exhibiting a certain behavior in a specific scenario. For example, what is the probability of a vehicle "rapidly fleeing" in a specific traffic intersection scenario? This can be achieved by collecting a large amount of historical data, statistically analyzing the frequency of various behaviors in different scenarios, and then calculating the corresponding probabilities. This probabilistic information is stored in a behavior prior graph, forming a probability distribution about the relationship between target and scenario behaviors.
[0126] Here, KL divergence, also known as relative entropy, is a non-negative measure used to measure the difference between two probability distributions.
[0127] In some embodiments, in the behavior recognition model, the behavior probability distribution P output by the model is... model It is predicted by the model based on the current input data, while the line prior map provides a prior probability distribution P about the target and scene behavior. prior By calculating P mode and P prior The KL divergence between the two parameters is calculated and incorporated into the model's training process as part of the loss function or as a regularization term. During training, the model continuously adjusts its parameters to make P... model Gradually approaching P prior This means guiding the model's output to converge to the behavioral pattern reflected by the prior probability distribution. This allows us to use prior knowledge to correct potential erroneous predictions by the model in the presence of insufficient data or noise interference, thus improving the model's robustness.
[0128] For example, if the prior graph shows that the probability of a vehicle “driving normally” is high in a certain scenario, while the model predicts that the probability of other behaviors is high, the model will adjust its parameters through the guidance of KL divergence, so that the probability of “driving normally” increases, thereby making the prediction results more consistent with prior knowledge.
[0129] In some embodiments, after obtaining the behavioral intent of each detected object, the target category, image pixel coordinates, latitude and longitude coordinates (i.e., geographic coordinates), tracking identification information and behavioral intent of each detected object can be formatted and output, supporting spatial formats such as GeoJSON and Shapefile; at the same time, a structured log file containing target trajectory, speed and behavioral status is generated.
[0130] Furthermore, it can also connect to map systems or drone platforms to display target trajectories and changes in intent status in real time on the map, and support issuing alerts for high-risk targets.
[0131] In this embodiment, the fusion modeling method of time series feature vectors and context label vectors can improve the accuracy and robustness of behavior recognition, which is conducive to enhancing the reliability of judging the behavioral intent of each detected object, thereby providing strong technical support for application scenarios such as security monitoring, border management, and traffic scheduling.
[0132] In some embodiments, the above-mentioned step S103, "determining the latitude and longitude trajectory of the detected object based on tracking identification information and back projection mapping function," may include the following steps: S1031, determine the image pixel trajectory of the detected object based on the tracking identification information.
[0133] Here, image pixel trajectory refers to the set of positional information of the detected object in each frame of the video sequence.
[0134] The image pixel trajectory can include multiple image pixel coordinates of the detected object and a timestamp corresponding to each image pixel coordinate. In other words, the image pixel trajectory can be composed of a series of image pixel coordinates, each image pixel coordinate carrying a corresponding timestamp, which indicates at what moment in the video stream the image pixel coordinate was captured.
[0135] Image pixel coordinates can be understood as being represented in two-dimensional space as (x, y), and are used to describe the specific location of the detected object on the image plane. Timestamps are used to record the time points when the detected object appears or moves, facilitating subsequent temporal modeling and behavior analysis.
[0136] In some embodiments, multiple image pixel coordinates of the detected object and the timestamp corresponding to each image pixel coordinate can be determined based on the tracking identification information of the detected object in multiple valid image frames. By organizing the multiple image pixel coordinates in chronological order, the motion path of the detected object in the video, i.e., the image pixel trajectory, can be clearly reflected. The trajectory data generated by the image pixel trajectory can not only be used for localization but also serve as one of the input features for subsequent intent determination.
[0137] S1032, based on the back projection mapping function, converts multiple image pixel coordinates into multiple latitude and longitude coordinates.
[0138] Here, the back projection mapping function is a mathematical model that converts image plane coordinates (i.e., image pixel coordinates) into geographic coordinates (such as latitude and longitude under the WGS-84 standard).
[0139] Latitude and longitude coordinates are the standard geographic representation of a location on the Earth's surface, typically expressed using the WGS-84 coordinate system. By converting image pixel coordinates to latitude and longitude coordinates, it is possible to more intuitively understand the changes in the detected object's location in actual geographic space, which is helpful for subsequent spatial analysis, map fusion, and decision support.
[0140] In some embodiments, a back-projection mapping function from image coordinates to geographic coordinates can be constructed based on the relationship between the camera imaging model and remote sensing projection.
[0141] For example, the back projection mapping function can be expressed as: (1) in, Represents the homogeneous coordinates of pixels on an image frame. This represents the column coordinate of a pixel in the image coordinate system; This represents the row coordinate of a pixel in the image coordinate system; The intrinsic parameters matrix of the camera includes the focal length and principal point position; h represents the elevation value, which represents the height of the ground point corresponding to the pixel relative to a reference surface (such as an ellipsoid); R represents the camera's attitude rotation matrix, which is a 3x3 matrix that describes the rotation transformation from the geographic coordinate system (or world coordinate system) to the camera coordinate system; T represents the camera's translation vector, which is a 3x1 vector that describes the displacement from the origin of the geographic coordinate system to the camera's optical center.
[0142] It should be noted that the elevation value in the above formula (1) can be iteratively optimized through a query-inversion-correction process to update the latitude and longitude coordinates.
[0143] In some embodiments, an initial elevation value can be obtained, and latitude and longitude coordinates can be determined based on the initial elevation value and the back-projection mapping function. Further, the elevation value can be determined from a mathematical elevation model based on the latitude and longitude coordinates, the average ground elevation of the detected target's location can be estimated using linear interpolation, and the average ground elevation can be used to update the latitude and longitude coordinates in the back-projection mapping function. That is, the above geometric inversion is performed on the center pixel of the detected target in each frame to obtain its geographical location (latitude and longitude coordinates) in the WGS-84 coordinate system. The inversion error does not exceed 2 meters in flat terrain areas.
[0144] S1033 merges multiple latitude and longitude coordinates to construct the latitude and longitude trajectory of the detected object.
[0145] Here, the latitude and longitude trajectory is a sequence of multiple latitude and longitude coordinates arranged in chronological order, used to describe the movement path of the detected object in real geographic space.
[0146] In some embodiments, multiple latitude and longitude coordinates can be interpolated and connected into a continuous trajectory line to form a spatiotemporal behavior map of the detected object. The latitude and longitude trajectory can not only be used to visualize the movement of the detected object, but also serve as an important input to the behavior intent recognition module, helping the remote sensing video intelligent processing system predict the next behavior of the detected object.
[0147] In this embodiment, a high-precision latitude and longitude trajectory is constructed by combining an image pixel trajectory set with a back-projection mapping function. The method of this embodiment can accurately map the visual information of the detected object into geospatial space, thereby supporting multi-object tracking, trajectory analysis, and intent determination.
[0148] In some embodiments, Figure 3 This is a schematic diagram of the overall process of a multimodal remote sensing target tracking, localization, and intent determination method provided in an embodiment of this application, as shown below. Figure 3 As shown, the method includes S301 to S310, wherein: S301, acquires visible light video data and infrared video data synchronously collected by the drone; S302, preprocess the visible light video data and infrared light video data to determine multiple visible light image frames corresponding to the visible light video data and multiple infrared light image frames corresponding to the infrared light video data. S303, perform frame alignment operation on each visible light image frame and each infrared light image frame to obtain multiple sets of valid image frame pairs; S304, For each pair of valid image frames, target detection is performed on the valid image frame pairs based on the target detection model to obtain the target detection results of each detected object in the valid image frame pairs; S305, based on a multi-target tracking model, associates the target detection results of each detected object with the state feature vector of each detected object in the historical N frames to determine the tracking identification information of each detected object; S306, For each detected object, the latitude and longitude trajectory of the detected object is determined based on the tracking identification information and the back projection mapping function; S307, extract features from the latitude and longitude trajectories of each detected object to obtain the time series feature vector of each detected object; S308 spatially matches latitude and longitude trajectories with scene semantic labels to determine the context label vector of each detected object; S309, Based on multi-channel embedding, the time series feature vectors of each detected object and the context label vectors of each detected object are projected into the same embedding space to obtain the target vectors of each detected object; S310, input the target vector of each detected object into the behavior recognition model to determine the behavioral intent of each detected object.
[0149] The following describes the application of the multimodal remote sensing target tracking, localization, and intent discrimination method provided in the embodiments of this application in real-world scenarios.
[0150] This application belongs to the interdisciplinary field of artificial intelligence and remote sensing technology, specifically involving a multimodal remote sensing target tracking, localization, and intent discrimination method based on UAV video data.
[0151] With the rapid development of next-generation low-altitude remote sensing technology, unmanned aerial vehicle (UAV) platforms, due to their advantages of high mobility, low cost, and rapid response, are gradually becoming an important sensing tool in fields such as national security defense, urban management, coastal patrol, environmental monitoring, disaster emergency response, and wilderness search and rescue. Compared with traditional static remote sensing images, remote sensing videos acquired by UAVs are characterized by continuity, multi-angle, high resolution, and dynamic changes, providing possibilities for target behavior analysis and real-time intervention. However, in the process of moving from "image perception" to "dynamic understanding," existing technologies still face many challenges in target tracking, geolocation, and behavioral intent recognition.
[0152] To address the aforementioned issues, this application proposes a multimodal remote sensing target tracking, localization, and intent discrimination method for UAV video. This method integrates three core capabilities: visual perception, temporal modeling, and spatial reasoning, constructing a complete technical system from video-level continuous target tracking and accurate geolocation inversion to intelligent behavioral intent discrimination. Unlike traditional methods that only focus on static image detection or coarse-grained recognition, this application introduces a multimodal collaborative modeling mechanism into the original UAV video, fully analyzing the target's appearance features, motion trajectory, geographical constraints, and environmental semantics, achieving a higher level of target understanding and task perception. It not only supports continuous multi-target tracking and occlusion recovery but also, by combining metadata information, can invert the target's true latitude and longitude trajectory in real time and discriminate and predict its behavioral state, providing crucial support for intelligent monitoring and task early warning in complex remote sensing scenarios.
[0153] This application possesses excellent versatility, scalability, and deployment capabilities, making it suitable for various drone platforms and mission scenarios. It shows broad prospects in practical applications such as security patrols, border surveillance, smuggling investigations, and urban security. Through deep decoding and multimodal semantic fusion of video information, this application effectively overcomes the technical shortcomings of existing technologies, such as unstable tracking, inaccurate positioning, and difficulty in behavior judgment in dynamic environments. It lays a crucial foundation for intelligent remote sensing video processing and provides important support for the future intelligent development of remote sensing "from perception to cognition."
[0154] The technical solution proposed in this application is as follows: A multimodal remote sensing target tracking, localization, and intent determination method for UAV video includes the following steps: Step 1: First, acquire video data from the multimodal sensors on the UAV, covering both visible light and infrared channels to ensure sensing capabilities under different lighting and weather conditions. The basic requirements for video data are a frame rate ≥ 25fps, resolution ≥ 1920×1080, and duration ≥ 10 seconds. Simultaneously, acquire metadata such as GPS coordinates, flight altitude, attitude angles (pitch, roll, yaw), and camera intrinsic and extrinsic parameters to ensure spatiotemporal consistency. Then, perform video decoding and keyframe extraction, removing blurry and occluded frames to form a quality-controlled sample sequence, laying the foundation for subsequent processing.
[0155] Step 2: In each frame, target detection is performed on both the visible light and infrared images using a pre-trained detection model (i.e., the target detection model in the above embodiment), and a fusion strategy is used to improve detection robustness and diversity. Cross-frame tracking employs a time-based multi-target tracking model, combining target state vectors from several historical frames (i.e., state feature vectors in the above embodiment), and utilizing a bidirectional attention mechanism to model temporal correlation. Target identity is associated using the Hungarian matching algorithm (i.e., the Hungarian matching algorithm in the above embodiment), and a Kalman filter is used to smooth the trajectory, effectively mitigating interference caused by occlusion, target loss, and viewpoint changes.
[0156] Step 3: Based on the pinhole model and multi-source attitude data, a back-projection mapping function from image coordinates to geographic coordinates (WGS-84) is constructed. This process considers sensor position, heading, and camera intrinsic parameters, and uses DEM data (i.e., the digital elevation model in the above embodiment) to correct terrain height. To address the geometric inconsistencies of multimodal images, joint calibration and registration ensure consistent spatial accuracy. Finally, the image center point of each detected target can be converted into actual latitude and longitude coordinates. The inversion accuracy is better than 2.5 meters in flat areas, providing precise location information for subsequent spatiotemporal modeling and surveillance tasks.
[0157] Step 4: Construct a trajectory set based on the latitude and longitude sequence of each target, and further extract kinematic features (velocity, direction angle, acceleration, dwell time, etc.) to form a time-series feature vector. Combined with the regional semantic map, spatially match the target trajectory with scene semantic labels (such as roads, woodlands, buildings, boundaries, etc.) to obtain contextual information (i.e., the context label vector in the above embodiment). The above motion features and semantic labels are projected onto a unified feature space through multi-channel embedding to form a complete input vector (i.e., the target vector in the above embodiment), for use by the intent discrimination module.
[0158] Step 5: Based on a hybrid network structure of BiLSTM and attention mechanism, the aforementioned temporal feature vector is input, and the current target's behavior category probability distribution is output. Candidate categories include typical behavior states such as "normal driving," "rapid escape," "abnormal stop," "approaching the boundary," and "attempting to conceal." To improve generalization ability, a prior behavior graph is introduced to model the joint probability relationship between historical behavior and the scene, and the network output is constrained by KL divergence to improve recognition accuracy and stability. In addition, smoothing and constraint strategies are used to suppress prediction jumps, achieving stable output of behavior states.
[0159] Step 6: Integrate and output the results of detection, tracking, localization, and intent determination in a unified GeoJSON or Shapefile format for easy integration with map systems or drone platforms. The system supports structured logging of multi-target states and displays their trajectories, behavioral states, and risk levels in real time on the map. High-risk targets will trigger an alarm mechanism to assist command personnel in making response decisions. The entire system is real-time and deployable, and can be applied to various complex remote sensing scenarios such as coastal defense patrols and urban security.
[0160] The beneficial effects of this application may include the following: 1. Support for continuous moving target tracking and localization: This application employs a multi-target tracking module that combines temporal models and appearance feature matching, effectively solving the ID loss problem caused by viewpoint changes and target occlusion in UAV videos. Compared with traditional frame-level detection methods, this application significantly improves the continuity of target recognition and trajectory integrity in video streams, making it particularly suitable for multi-target recognition and association in complex dynamic scenes.
[0161] 2. Achieving a high-precision geographic coordinate inversion mechanism: By fusing camera intrinsic and extrinsic parameters, UAV attitude data, and a digital elevation model, a back-projection function from image pixels to geographic coordinates is constructed to accurately invert the true latitude and longitude position of targets in the image. This module is widely applicable to complex imaging conditions such as oblique photography and low-altitude squint, significantly improving the spatial locationability and decision-making usability of remote sensing data in practical deployments.
[0162] 3. Introduction of a target intent discrimination module: This module not only identifies the target's location and category but also extracts its trajectory changes and spatial semantic environment, training a sequence model to automatically identify target behavior and classify intent. This achieves a leap from the perception layer to the cognitive layer, enabling remote sensing systems to identify and respond to risky targets in advance, meeting the practical needs of intelligent early warning and auxiliary command.
[0163] 4. Multimodal Collaborative Perception Design: This design integrates heterogeneous information such as image features, trajectory dynamics, geographic location, and spatial semantics into a fusion model, introducing an attention mechanism for feature enhancement and interaction modeling. Even under conditions of multi-source interference, target camouflage, or complex backgrounds, it maintains high recognition accuracy and behavioral judgment stability, demonstrating good generalization ability and environmental adaptability.
[0164] Based on the above embodiments, this application also provides a multimodal remote sensing target tracking, localization, and intent determination device. Figure 4 This is a schematic diagram of the composition structure of a multimodal remote sensing target tracking, localization, and intent discrimination device provided in an embodiment of this application, as shown below. Figure 4As shown, the multimodal remote sensing target tracking, localization, and intent determination device 400 includes an acquisition module 401, a first determination module 402, a second determination module 403, and a third determination module 404, wherein: The acquisition module 401 is used to acquire multiple visible light image frames and multiple infrared light image frames, and to perform frame alignment operation on each visible light image frame and each infrared light image frame to obtain multiple sets of valid image frame pairs. The first determining module 402 is used to determine the tracking identification information of each detected object in each effective image frame pair based on the effective image frame pair and the pre-trained multimodal detection and tracking model; The second determining module 403 is used to determine the latitude and longitude trajectory of each detected object based on the tracking identification information and the back projection mapping function; wherein the back projection mapping function is determined based on the camera imaging model and multi-source attitude data. The third determination module 404 is used to determine the behavioral intent of each detected object based on the behavior recognition model and the latitude and longitude trajectories of each detected object.
[0165] In some embodiments of this application, the multimodal detection and tracking model includes a target detection model and a multi-target tracking model; the first determination module 402 includes a target detection unit and a determination unit; the target detection unit is used to perform target detection on valid image frame pairs based on the target detection model to obtain the target detection results of each detected object in the valid image frame pairs; the target detection results include the target category, image pixel coordinates, and the confidence level corresponding to the target category; the determination unit is used to associate the target detection results of each detected object with the state feature vectors of each detected object in the historical N frames based on the multi-target tracking model to determine the tracking identification information of each detected object; the state feature vectors include the appearance feature vectors and motion feature vectors of the detected object.
[0166] In some embodiments of this application, the determining unit is further configured to input the target detection results of each detected object and the state feature vector of each detected object in the historical N frames into the multi-target tracking model to determine the similarity matrix between each detected object and each detected object in the historical N frames; and to determine the tracking identification information of each detected object based on the similarity matrix and the Hungarian matching algorithm.
[0167] In some embodiments of this application, the effective image frame pair includes an aligned visible light image frame and an aligned infrared light image frame; the target detection unit is further configured to input the aligned visible light image frame into the target detection model to obtain a first detection result for each detected object in the aligned visible light image frame; input the aligned infrared light image frame into the target detection model to obtain a second detection result for each detected object in the aligned infrared light image frame; and fuse the first detection result and the second detection result to obtain a target detection result.
[0168] In some embodiments of this application, the third determining module 404 is further configured to extract features from the latitude and longitude trajectories of each detected object to obtain the time-series feature vector of each detected object; perform spatial matching between the latitude and longitude trajectories and scene semantic labels to determine the context label vector of each detected object; the scene semantic labels include roads, bushes, buildings and boundary lines; project the time-series feature vector of each detected object and the context label vector of each detected object into the same embedding space based on multi-channel embedding to obtain the target vector of each detected object; input the target vector of each detected object into the behavior recognition model to determine the behavioral intent of each detected object.
[0169] In some embodiments of this application, the second determining module 403 is further configured to determine the image pixel trajectory of the detected object based on the tracking identification information; the image pixel trajectory includes multiple image pixel coordinates of the detected object and a timestamp corresponding to each image pixel coordinate; based on the back projection mapping function, the multiple image pixel coordinates are converted into multiple latitude and longitude coordinates; the multiple latitude and longitude coordinates are merged to construct the latitude and longitude trajectory of the detected object.
[0170] In some embodiments of this application, the acquisition module 401 is further configured to acquire visible light video data and infrared light video data synchronously collected by the UAV; preprocess the visible light video data and infrared light video data to determine multiple visible light image frames corresponding to the visible light video data and multiple infrared light image frames corresponding to the infrared light video data.
[0171] In some embodiments of this application, after the acquisition module 404, the multimodal remote sensing target tracking, localization, and intent determination device 400 further includes: a first acquisition module and a binding module, wherein: The first acquisition module is used to acquire metadata collected by the aircraft's sensors; The binding module is used to bind the metadata of each valid image frame pair to its corresponding timestamp for each valid image frame pair, so as to obtain the sample sequence associated with the valid image frame pair. Correspondingly, the first determining module 402 is also used to determine the tracking identification information of each detected object in the sample sequence based on the sample sequence and the pre-trained multimodal detection and tracking model.
[0172] It should be noted that, in the embodiments of this application, if the above methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of software products. These software products are stored in a storage medium and include several instructions to cause an electronic device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0173] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the above-described method.
[0174] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. The computer-readable storage medium can be transient or non-transient.
[0175] This application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof.
[0176] In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0177] It should be noted that, Figure 5 This is a hardware entity diagram of an electronic device provided in an embodiment of this application, such as... Figure 5 As shown, the hardware entity of the electronic device 500 includes: a processor 501, a communication interface 502, and a memory 503, wherein: Processor 501 typically controls the overall operation of electronic device 500.
[0178] Communication interface 502 enables electronic device 500 to communicate with other terminals or servers via a network.
[0179] The memory 503 is configured to store instructions and applications executable by the processor 501, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 501 and various modules in the electronic device 500. It can be implemented using flash memory or RAM. Data transfer between the processor 501, the communication interface 502, and the memory 503 can be performed via bus 504.
[0180] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0181] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0182] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0183] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components may be combined, or integrated into another system, or some features may be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0184] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0185] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0186] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory, magnetic disks, or optical disks.
[0187] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks. The above are merely embodiments of this application and are not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A multimodal remote sensing target tracking, localization, and intent determination method, characterized in that, The method includes: Multiple visible light image frames and multiple infrared light image frames are acquired, and frame alignment is performed on each of the visible light image frames and each of the infrared light image frames to obtain multiple sets of valid image frame pairs; For each pair of valid image frames, based on the valid image frame pair and the pre-trained multimodal detection and tracking model, the tracking identification information of each detected object in the valid image frame pair is determined; For each detected object, the latitude and longitude trajectory of the detected object is determined based on the tracking identification information and the back projection mapping function; wherein, the back projection mapping function is determined based on the camera imaging model and multi-source pose data; Based on the behavior recognition model and the latitude and longitude trajectories of each detected object, the behavioral intent of each detected object is determined.
2. The method according to claim 1, characterized in that, The multimodal detection and tracking model includes a target detection model and a multi-target tracking model; the determination of tracking identification information for each detected object in the effective image frame pair based on the effective image frame pair and the pre-trained multimodal detection and tracking model includes: Based on the target detection model, target detection is performed on the effective image frame pairs to obtain the target detection results for each detected object in the effective image frame pairs; the target detection results include the target category, image pixel coordinates, and the confidence score corresponding to the target category; Based on the multi-target tracking model, the target detection results of each detected object are associated with the state feature vectors of each detected object in the historical N frames to determine the tracking identification information of each detected object; the state feature vectors include the appearance feature vectors and motion feature vectors of the detected object.
3. The method according to claim 2, characterized in that, The step, based on the multi-target tracking model, associates the target detection results of each detected object with the state feature vectors of each detected object across N historical frames to determine the tracking identification information of each detected object, including: The target detection results of each detected object and the state feature vector of each detected object in the historical N frames are input into the multi-target tracking model to determine the similarity matrix between each detected object and each detected object in the historical N frames; Based on the similarity matrix and the Hungarian matching algorithm, the tracking identification information of each detected object is determined.
4. The method according to claim 2, characterized in that, The effective image frame pair includes aligned visible light image frames and aligned infrared light image frames; The step of performing target detection on the effective image frame pairs based on the target detection model to obtain the target detection results of each detected object in the effective image frame pairs includes: The aligned visible light image frame is input into the target detection model to obtain the first detection result of each detection object in the aligned visible light image frame; The aligned infrared image frame is input into the target detection model to obtain the second detection result of each detected object in the aligned infrared image frame; The first detection result and the second detection result are fused together to obtain the target detection result.
5. The method according to any one of claims 1 to 4, characterized in that, The determination of the behavioral intent of each detected object based on the behavior recognition model and the latitude and longitude trajectories of each detected object includes: Feature extraction is performed on the latitude and longitude trajectories of each detected object to obtain the time series feature vector of each detected object; The latitude and longitude trajectories are spatially matched with scene semantic labels to determine the context label vector of each detected object; the scene semantic labels include roads, bushes, buildings, and boundary lines; Based on multi-channel embedding, the time series feature vectors of each detected object and the context label vectors of each detected object are projected into the same embedding space to obtain the target vectors of each detected object; The target vectors of each detected object are input into the behavior recognition model to determine the behavioral intent of each detected object.
6. The method according to any one of claims 1 to 4, characterized in that, Determining the latitude and longitude trajectory of the detected object based on the tracking identification information and the back projection mapping function includes: The image pixel trajectory of the detected object is determined based on the tracking identification information; the image pixel trajectory includes multiple image pixel coordinates of the detected object and a timestamp corresponding to each image pixel coordinate. Based on the back projection mapping function, the multiple image pixel coordinates are converted into multiple latitude and longitude coordinates; The multiple latitude and longitude coordinates are merged to construct the latitude and longitude trajectory of the detected object.
7. The method according to any one of claims 1 to 4, characterized in that, The acquisition of multiple visible light image frames and multiple infrared light image frames includes: Acquire visible light video data and infrared video data synchronously collected by the drone; The visible light video data and infrared light video data are preprocessed to determine multiple visible light image frames corresponding to the visible light video data and multiple infrared light image frames corresponding to the infrared light video data.
8. The method according to any one of claims 1 to 4, characterized in that, After performing frame alignment operations on each of the visible light image frames and each of the infrared light image frames to obtain multiple sets of valid image frame pairs, the method further includes: Acquire metadata collected by the aircraft's sensors; For each pair of valid image frames, the metadata of the valid image frame pair under its corresponding timestamp is bound to obtain the sample sequence associated with the valid image frame pair; Correspondingly, determining the tracking identification information of each detected object in the effective image frame pair based on the effective image frame pair and the pre-trained multimodal detection and tracking model includes: Based on the sample sequence and the pre-trained multimodal detection and tracking model, the tracking identification information of each detected object in the sample sequence is determined.
9. A multimodal remote sensing target tracking, localization, and intent determination device, characterized in that, The device includes: The acquisition module is used to acquire multiple visible light image frames and multiple infrared light image frames, and to perform frame alignment operations on each of the visible light image frames and each of the infrared light image frames to obtain multiple sets of valid image frame pairs; The first determining module is used to determine the tracking identification information of each detected object in each effective image frame pair based on the effective image frame pair and the pre-trained multimodal detection and tracking model. The second determining module is used to determine the latitude and longitude trajectory of each detected object based on the tracking identification information and the back projection mapping function; wherein the back projection mapping function is determined based on the camera imaging model and multi-source pose data. The third determining module is used to determine the behavioral intent of each detected object based on the behavior recognition model and the latitude and longitude trajectories of each detected object.
10. The apparatus according to claim 9, characterized in that, The multimodal detection and tracking model includes a target detection model and a multi-target tracking model; the first determination module includes a target detection unit and a determination unit; The target detection unit is used to perform target detection on the effective image frame pair based on the target detection model, and obtain the target detection result of each detected object in the effective image frame pair; The target detection result includes the target category, image pixel coordinates, and the confidence level corresponding to the target category; The determining unit is used to associate the target detection results of each detected object with the state feature vectors of each detected object in N historical frames based on the multi-target tracking model, and determine the tracking identification information of each detected object; the state feature vectors include the appearance feature vectors and motion feature vectors of the detected objects.
Citation Information
Cited By
An unmanned aerial vehicle intelligent perception and intention reasoning system and device
CN122175020A
An unmanned aerial vehicle intelligent perception and intention reasoning system and device
CN122175020B
A method for locating obstructed piers and identifying and monitoring project progress in railway construction based on UAV monocular vision.
CN122416322B