Satellite video dense vehicle tracking method and system
By constructing the VDD-VEH dataset and using DSFNet, improving the motion position graph MPG of the RAFT network, and optimizing ILP, the problems of missed detection, false detection, and trajectory breakage of vehicle targets in satellite video were solved, achieving high-precision long-term time-series tracking, which is suitable for intelligent traffic monitoring and remote sensing video analysis.
Patent Information
- Application Number
- CN202510698494.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Satellite video suffers from problems such as missed detection, false detection, and trajectory breakage due to the small scale, dense distribution, low contrast, and environmental interference of vehicle targets. Existing methods lack semantic features for small targets, leading to frequent missed detections and false detections. Dense targets are easily confused, short-term motion prediction is difficult to maintain long-term temporal consistency, and the lack of continuous motion information annotation limits spatiotemporal modeling.
We constructed a satellite video dense vehicle tracking dataset VDD-VEH, and adopted the spatiotemporal target detector DSFNet and the improved RAFT network. Through motion location graph MPG and global optimization strategy, combined with multi-feature edge weight MFEW and integer linear programming ILP, we removed abnormal trajectories, repaired trajectory breaks, and achieved long-term time-series tracking.
It significantly improves the MOTA and IDF1 metrics for vehicle tracking, reduces the number of identity switching times by 55%, and has high accuracy and cross-scenario generalization capabilities, making it suitable for intelligent traffic monitoring and remote sensing video analysis.
Smart Images

Figure CN120598996B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and remote sensing technology, specifically relating to a method and system for tracking dense vehicles in satellite video based on motion position map (MPG) and global optimization, which is suitable for continuous tracking and trajectory optimization of small, dense, low-contrast vehicle targets in satellite video. Background Technology
[0002] With the widespread adoption of video satellites, satellite video has demonstrated enormous potential in fields such as traffic monitoring and urban planning. However, vehicle targets in satellite video are characterized by their high imaging altitude and low resolution, resulting in small size (10-30 pixels), dense distribution, and low contrast. They are also susceptible to cloud cover, changes in lighting, and dynamic background interference. Traditional multi-target tracking methods (such as SORT and DeepSORT) rely on target appearance features or local motion modeling. In satellite scenarios, these methods face challenges such as insufficient semantic features of small targets, leading to frequent missed and false detections, easy confusion of dense targets, difficulty in maintaining long-term temporal consistency in short-term motion prediction, trajectory fragmentation due to occlusion or missed detection, and frequent ID switching.
[0003] Existing satellite video datasets (such as AIR-MOT and SAT-MTB) often lack continuous motion information annotations, which limits the training and evaluation of spatiotemporal modeling methods. Therefore, a complete solution combining data construction, spatiotemporal modeling, and global optimization is needed. Summary of the Invention
[0004] To address the issues of missed detections, false detections, and trajectory breaks caused by the small scale, dense distribution, low contrast, and environmental interference of vehicles in satellite video, this invention proposes a dense vehicle tracking method based on Motion Position Map (MPG) and global optimization. By constructing a spatiotemporal graph model and employing a global optimization strategy, it solves the long-term tracking challenge of small, dense vehicles. Specifically, the VDD-VEH dataset is constructed to include detailed annotations of vehicle targets, such as horizontal bounding boxes, inter-frame motion vectors (forward and backward velocities), occlusion states, and background dynamic attributes. Simultaneously, the original satellite video frames are cropped into 512×512 pixel blocks to adapt to the deep learning model. This dataset is large in scale, containing 70,665 frames and 15,461 instances, with a target density exceeding 100 per frame and covering multiple scenes. The spatiotemporal target detector DSFNet adopts a dual-stream architecture: the static stream uses the DLA-34 network to extract single-frame spatial features to enhance the details of small targets, while the dynamic stream uses a lightweight 3D... Convolutional processing captures temporal motion patterns, and deformable convolution aligns multi-scale features, combining them with feature pyramids to output detection boxes and confidence scores. Motion Location Map (MPG) modeling maps DSFNet detection results to spatiotemporal nodes. Based on an improved RAFT network, inter-frame optical flow is estimated to generate cross-frame candidate association edges. The Multi-Feature Edge Weighting (MFEW) strategy is used to calculate edge weights by integrating multiple factors. The Global Optimization and Trajectory Optimization (TRM) module maximizes the sum of global matching edge weights using integer linear programming and constrains the uniqueness of node entry and exit edges. Abnormal trajectories are eliminated based on trajectory point distribution density and total displacement. Kalman filtering is used to predict virtual nodes and splice broken trajectories according to motion consistency scores.
[0005] The technical solution adopted in this invention is: a satellite video dense vehicle tracking method, comprising the following steps:
[0006] (1) Construct a satellite video dense vehicle tracking dataset VDD-VEH, which contains multiple frames of satellite video images and their corresponding vehicle target spatial location annotations and inter-frame motion vector annotations;
[0007] (2) Based on the spatiotemporal target detector DSFNet, the detection information of vehicle targets in continuous video frames is extracted to generate a set of nodes containing target location and confidence level;
[0008] (3) The improved RAFT network is used as a motion flow estimator to estimate the motion flow of vehicle targets between adjacent frames, and cross-frame candidate association edges are constructed based on the motion flow;
[0009] (4) Based on the estimation results obtained from the spatiotemporal target detector DSFNet and the motion flow estimator, a motion location map is constructed, that is, the position of each detected target is mapped to a node in the motion location map, and then the estimated motion flow is used to establish connection edges between nodes in adjacent frames.
[0010] (5) Globally optimize the motion position map based on integer linear programming to generate the initial target trajectory;
[0011] (6) Filter and stitch the initial trajectory, remove abnormal trajectories and repair trajectory breaks, and output the final vehicle tracking result.
[0012] Furthermore, the construction of the VDD-VEH dataset in step (1) includes: cropping image blocks of a certain size from satellite video, labeling the horizontal bounding box of the vehicle target and the corresponding motion vector, wherein the motion vector includes forward velocity and backward velocity, and labeling the occlusion state of the target and the dynamic interference attributes of the background.
[0013] Furthermore, the spatiotemporal target detector DSFNet described in step (2) adopts a dual-stream feature extraction architecture, in which a static feature stream extracts spatial information of a single frame through 2D convolution; a dynamic feature stream extracts time series information through lightweight 3D convolution; and a hierarchical feature fusion strategy is adopted to make features of different scales complement each other, and finally outputs the target's bounding box, category information and confidence score through the detection head.
[0014] Furthermore, the improved RAFT network described in step (3) first determines the location of the foreground region based on the dataset annotation information during the training process and marks its corresponding region in the optical flow map. During loss calculation, the foreground region has a higher weight, while the background region has a lower weight. The loss weights for sample n are defined as follows:
[0015]
[0016] in, It is a manually set hyperparameter used to enhance the loss weights of foreground targets. It is a dynamic weight allocation strategy. The loss weight of sample n at position p is defined, and the robustness of optical flow estimation is optimized by distinguishing between foreground and background and reasonable abnormal motion. This is a binary mask used to determine whether a given location belongs to a foreground object. This indicates that it belongs to the prospective goals, while This indicates the range of motion at that point; for target points with excessively large ranges of motion, i.e., exceeding the set maximum range of motion... The motion flow estimator considers it an outlier and therefore excludes it from the loss calculation; based on this, the final motion flow loss function is constructed as follows:
[0017]
[0018] Where N represents the number of samples in the training batch, and M represents the number of iterations in the RAFT network. P represents the set of all pixel locations in sample n; in the context of satellite video, each sample n corresponds to a frame pair, and P covers all pixel locations in that frame pair for which optical flow loss needs to be calculated. It is the loss weight of sample n at position p. This represents the optical flow value predicted by RAFT. ) represents the motion vector of the target bounding box labeled in the VDD-VEH dataset; These are the weight parameters.
[0019] Furthermore, the motion flow estimator optimizes the gradient allocation for the background region. Specifically, it limits the update magnitude of the background region when calculating the gradient to reduce its impact on foreground target prediction. In this specific case, an adaptive gradient clipping mechanism is employed. First, the derivative of the loss function is obtained to acquire the original gradients of all parameters. L Then based on the mask The parameter gradients belonging to the background region are selected and denoted as follows. The structure is as follows:
[0020] =▽ L ⊙(1- )
[0021] Where ⊙ represents element-wise multiplication. It's a binary mask, with foreground values of 1 and background values of 0; finally, the gradient values are updated. Apply constraints: ,in It is a dynamically adjusted upper limit value for the gradient.
[0022] Furthermore, the target position information output by the spatiotemporal target detector in step (4) will be mapped to nodes of the motion position map. Specifically, for each object i detected in time frame t, a node will be added to the motion position map to represent the object, and its node set is as follows:
[0023]
[0024] The matching relationships between objects are then established through edges in the graph. Edge generation is based on the node motion flow predicted by the motion flow estimator. During edge construction, a method based on motion flow prediction of position offset is used for determination: the optical flow information extracted by RAFT is used to predict the nodes of the current frame. Motion position in the next frame :
[0025]
[0026] in Let i be the position of object i in frame t. The motion flow vector obtained through RAFT is the optical flow estimate obtained using the improved RAFT. The inter-frame time interval is defined when the distance between the predicted and actual positions meets the following condition:
[0027]
[0028] This indicates that objects i and j have a potential spatiotemporal relationship, and a path will be established from node j in the graph. Pointing to node The edges are added to the candidate edge set E;
[0029] A multi-feature edge weighting (MFEW) strategy is further designed to assign association weights to each candidate edge. This strategy integrates spatial consistency, temporal consistency, appearance consistency, and detection confidence to comprehensively evaluate the association strength between nodes. The calculation formula is as follows:
[0030]
[0031] in Represents a node and The correlation weight between them The node confidence is directly provided by the output of the spatiotemporal target detector. , and Δp is the coefficient of three components: spatial consistency, temporal consistency, and appearance consistency. It is used to measure the contribution of different components to improving the accuracy of node association. Δp represents the spatial location difference. Δv represents the spatial Gaussian kernel standard deviation, controlling for distance sensitivity, and Δv represents the temporal motion difference. The Gaussian kernel standard deviation represents the variation in motion, adjusted for velocity consistency tolerance. The cosine similarity between nodes i and j is used to measure the appearance consistency between nodes.
[0032] Furthermore, the processing procedure for the integer linear programming (ILP) described in step (5) is as follows:
[0033] In the motion location graph, any node p or q represents the detected target, and each edge... To represent possible inter-frame correlations, ILP defines a binary variable to find the optimal trajectory match globally. Used to characterize edges Is it selected?
[0034]
[0035] when When the value is 1, it indicates that the edge is selected as part of the final trajectory; otherwise, the edge will not appear in the final result. The set of all candidate edges is denoted as E. The objective function of ILP is defined as maximizing the global score of the entire trajectory matching, which is the weighted sum of all matching edges:
[0036]
[0037] in, The matching weights are set for the edges. To ensure the rationality of the trajectory, ILP introduces constraints, including uniqueness constraints and binary variable constraints. The uniqueness constraint requires that each target can have at most one incoming edge and one outgoing edge, preventing the same target from being matched repeatedly in multiple frames.
[0038]
[0039]
[0040] Furthermore, binary variable constraints ensure It can only take the value 0 or 1:
[0041]
[0042] These constraints ensure that the trajectory generated after ILP optimization meets the singleness requirement, that is, each target can correspond to at most one trajectory node at any given time.
[0043] Furthermore, in step (6), the initial trajectory is filtered and spliced by the trajectory optimization module. The trajectory optimization module consists of two parts: trajectory filtering (TF) and trajectory splicing (FS). TF eliminates false trajectories caused by false detection by analyzing the distribution characteristics of trajectory points and trajectory displacement. Specifically, it calculates the spatial concentration and total displacement of the trajectory point sequence. The process is as follows: Let the trajectory point sequence be... in Let k represent the coordinates of the midpoint of the trajectory, L be the length of the trajectory, and the point distribution characteristics of the trajectory be defined as follows:
[0044]
[0045] in It is located in Centered on, with radius The number of points within the circle; the total displacement of the trajectory is calculated as follows:
[0046]
[0047] Setting a spatial discrete threshold in TF With minimum displacement threshold If the spatial dispersion of the trajectory exceeds Or the total displacement is less than If the trajectory is low quality, it will be identified as a low-quality trajectory and removed.
[0048] Furthermore, FS continuously generates virtual nodes for trajectory fragments. These virtual nodes are used to stitch together trajectory fragments belonging to the same target, improving trajectory integrity. Specifically, by setting the total number of frames in the current video sequence to N, FS generates virtual nodes for all initial trajectory sets. Each trajectory segment The start and end frames are respectively denoted as The start and end position coordinates are denoted as follows: , The set of trajectory fragments to be pieced together. Defined as:
[0049]
[0050] in This indicates the image edge region, meaning the trajectory endpoint is not in the last frame of the video sequence. Only trajectories whose endpoints are not located in the image edge region are considered potentially broken trajectory fragments and proceed to the subsequent stitching process. The selected candidate trajectory fragments are then stitched together, and the spatial distance and velocity difference between the predicted location and the actual trajectory starting point are calculated. Specifically, assuming from the set... Extract a trajectory segment to be spliced from the data. The trajectory ends at time [time]. The final position and velocity are denoted as follows: and In order to determine whether the trajectory can be compared with another trajectory splicing, including the trajectory The starting time and starting position are denoted as follows: FS uses a Kalman filter to analyze the trajectory Generate continuous virtual nodes and predict their trajectory. Position of the starting time With speed :
[0051]
[0052] Subsequently, FS calculates the trajectory. Predicted location and trajectory Spatial distance between actual starting positions and the difference between predicted speed and actual speed :
[0053]
[0054]
[0055] Based on the above difference values, a trajectory splicing matching scoring function is defined:
[0056]
[0057] in, and The normalized thresholds and weighting coefficients represent the differences in spatial location and velocity, respectively. and Used to balance the importance of position and velocity, if rated Exceeding the threshold Then, virtual nodes are inserted between the breakpoints to form a continuous trajectory.
[0058] The present invention also provides a satellite video dense vehicle tracking system, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the program to implement the satellite video dense vehicle tracking method as described above.
[0059] Compared with existing technologies, the advantages and beneficial effects of this invention are as follows: It constructs a dense vehicle dataset (VDD-VEH) from satellite video containing motion vector annotations, providing supervised information for spatiotemporal modeling; it designs a motion position graph (MPG) to map the target spatial location and motion flow into a three-dimensional spatiotemporal graph structure, and utilizes a multi-feature edge weighting (MFEW) strategy to fuse spatiotemporal consistency, appearance features, and detection confidence to quantify node association strength; it employs integer linear programming (ILP) to achieve globally optimal trajectory association, and combines this with a trajectory optimization module (TRM) to remove abnormal trajectories and repair trajectory breaks, enhancing long-term tracking stability. This method significantly outperforms existing technologies in metrics such as MOTA and IDF1, reduces identity switching times by 55%, and is suitable for intelligent traffic monitoring and remote sensing video analysis, possessing high accuracy and cross-scenario generalization capabilities. Attached Figure Description
[0060] Figure 1 The overall flowchart of the MPG method includes data input, detection, optical flow estimation, graph construction, global optimization, and trajectory output.
[0061] Figure 2 : DSFNet network structure diagram, showing the two-stream feature extraction and fusion process.
[0062] Figure 3 : Schematic diagram of the training process and optical flow prediction of the improved RAFT network.
[0063] Figure 4 Example of constructing nodes and edges in a Motion Position Graph (MPG).
[0064] Figure 5 Comparison of the effects of trajectory filtering (TF) and trajectory stitching (FS). Detailed Implementation
[0065] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0066] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for dense vehicle tracking based on motion position maps and global optimization using satellite video, comprising the following steps:
[0067] (1) Construct a satellite video dense vehicle tracking dataset VDD-VEH, which contains multiple frames of satellite video images and their corresponding vehicle target spatial location labels and inter-frame motion vector labels.
[0068] (2) Based on the spatiotemporal target detector DSFNet, which employs dynamic feature fusion technology, the detection accuracy of small targets is effectively improved, and a confidence score is provided, providing reliable basic information for subsequent target association. DSFNet is used to extract vehicle target detection information from continuous video frames, generating a target TF. By analyzing the trajectory point distribution characteristics and trajectory displacement, false trajectories caused by false detections are eliminated; specifically, the spatial concentration of the trajectory point sequence (such as the dispersion of trajectory points around the starting point) and the total displacement (Euclidean distance from the starting point to the ending point) are calculated. The process is as follows: Let the trajectory point sequence be... in Let k represent the coordinates of the midpoint of the trajectory, L be the length of the trajectory, and the point distribution characteristics of the trajectory be defined as follows:
[0069]
[0070] in It is located in Centered on, with radius The number of points inside the circle. The total displacement of the trajectory is calculated as follows:
[0071]
[0072] In the experiment, TF defines and sets a threshold. (Spatial discrete threshold) and (Minimum displacement threshold). If the spatial dispersion of the trajectory exceeds... Or the total displacement is less than If a trajectory is identified as low-quality (e.g., a brief false detection or stationary noise) and removed, the stability and consistency of the tracking results are improved. FS (Flash Frame) continuously generates virtual nodes for trajectory fragments. These virtual nodes are used to stitch together trajectory fragments belonging to the same target, improving trajectory integrity. Specifically, by setting the total number of frames in the current video sequence to N, all initial trajectory sets... Each trajectory segment The start and end frames are respectively denoted as The start and end position coordinates are denoted as follows: , The set of trajectory fragments to be pieced together. Defined as:
[0073]
[0074] in This indicates the edge region of the image. That is, the trajectory endpoint is not in the last frame of the video sequence. Only trajectories whose endpoints are not located in the image edge region are considered potentially broken trajectory fragments and proceed to the subsequent stitching process. The selected candidate trajectory fragments are then stitched together, and the spatial distance and velocity difference between the predicted location and the actual trajectory starting point are calculated. Specifically, assuming from the set... Extract a trajectory segment to be spliced from the data. The trajectory ends at time [time]. The final position and velocity are denoted as follows: and To determine whether this trajectory can be compared with another trajectory. splicing (including the trajectory) The starting time and starting position are denoted as follows: ),FS uses a Kalman filter to analyze the trajectory Generate continuous virtual nodes and predict their trajectory. Position of the starting time With speed :
[0075]
[0076] Subsequently, FS calculates the trajectory. Predicted location and trajectory Spatial distance between actual starting positions and the difference between predicted speed and actual speed :
[0077]
[0078]
[0079] Based on the above difference values, a trajectory splicing matching scoring function is defined:
[0080]
[0081] in, and The normalized thresholds and weighting coefficients represent the differences in spatial location and velocity, respectively. and Used to balance the importance of position and velocity. (As a rating) Exceeding the threshold Virtual nodes (connected by dashed lines) are inserted between breakpoints to form a continuous trajectory, thereby significantly improving the integrity and identity consistency of long-term tracking and reducing trajectory fragmentation problems.
[0082] A set of nodes with location and confidence level.
[0083] (3) The improved RAFT network is used as a motion flow estimator to estimate the motion flow of vehicle targets between adjacent frames, and cross-frame candidate association edges are constructed based on the motion flow.
[0084] (4) A motion location graph is constructed based on the estimation results obtained from the DSFNet target detector and the motion flow estimator. That is, the position of each detected target is mapped to a node in the graph by spatiotemporal target detection, and then the motion flow estimated by RAFT is used to establish the connection edges between nodes in adjacent frames. The detection result of each target in a certain frame constitutes a graph node, and the edge weights of targets in adjacent frames are calculated by the MFEW strategy. The MFEW strategy integrates the spatial position relationship between targets, motion pattern, appearance similarity and detection confidence, and assigns edge weights between nodes to improve the accuracy and robustness of matching.
[0085] (5) Global optimization of the motion location map (the motion location map obtained by the target detector and motion flow estimator) based on integer linear programming (ILP) is performed to generate the initial target trajectory.
[0086] (6) The target trajectory is generated based on the global optimization method of ILP solution. However, the generated trajectory still has some problems: the detection results of the spatiotemporal target detector inevitably have some false detections, which leads to some abnormal trajectories containing false detection nodes in the trajectory solved by ILP, reducing the overall trajectory quality; MPG only defines the possible association between adjacent frame nodes, so it still relies on frame-by-frame matching relationship for construction, but it cannot establish the association between nodes across multiple frames. This causes the trajectory to break due to the detector missing detection or the target being occluded, forming many trajectory fragments. Therefore, based on the global optimization of ILP, a trajectory optimization module (TRM) was further designed.
[0087] The construction of the VDD-VEH dataset in step (1) includes: cropping a 512×512 pixel image patch from satellite video; labeling the horizontal bounding box (HBB) of the vehicle target and its corresponding motion vector, the motion vector including forward velocity and backward velocity, as well as labeling the target's occlusion state and background dynamic interference attributes.
[0088] The part described in step (2) aims to improve the target detection accuracy in satellite video and ensure stability in complex environments. Therefore, DSFNet is introduced as a detector to detect targets. Figure 2 The DSFNet architecture employs a dual-stream feature extraction system, extracting target features from both static and dynamic dimensions, and combining multi-scale feature fusion to improve target detection accuracy. The static feature stream extracts spatial information from a single frame through 2D convolution, effectively extracting the target's appearance features and background context. The dynamic feature stream extracts temporal series information through lightweight 3D convolution, capturing the target's motion information across consecutive frames. Furthermore, DSFNet uses a hierarchical feature fusion strategy, allowing features at different scales to complement each other. Finally, the detection head outputs the target's bounding box, category information, and confidence score. The process is as follows: For a satellite video sequence with N frames, before extracting features from the nth frame (1 ≤ n ≤ N), DSFNet samples a sequence of n consecutive frames backward from the current frame n, constructing the input to the feature extraction network.
[0089]
[0090] Where m is the temporal window length defined by DSFNet. To ensure the validity of boundary frames, when nm + 1 < 1, the starting frame index will be pruned. This ensures that the input always falls within a valid frame number range. The static streaming network uses DLA-34 (Deep Layer Aggregation) as the backbone network for feature extraction. DLA-34 establishes direct connections between feature maps at different levels through a hierarchical feature aggregation mechanism, allowing high-level features to better utilize the spatial details of the lower levels. For the input sequence... DLA-34 frame capture ,in This represents a 3D tensor, where H represents the vertical number of pixels in the image (e.g., the height of a satellite video frame), W represents the horizontal number of pixels in the image (e.g., the width of a satellite video frame), and 3 channels represent the RGB three color channels (red, green, and blue) of the corresponding color image. Each channel stores the pixel intensity value of the corresponding color. DLA-34 takes the captured frames as input to a static stream, and the calculated feature map is represented as follows: Among them 2 d This indicates that this is a two-dimensional feature map, and i-1 indicates that the feature map belongs to the (i-1)th layer of the network. This represents the number of channels in the i-th layer. To maintain detection accuracy and enhance focus on small targets, DSFNet employs deformable convolution for feature alignment to reduce offset errors between feature layers. The specific calculation method is as follows:
[0091]
[0092] in, Represents deformable convolution. This represents transposed convolution. Deformable convolution can adaptively adjust the receptive field, allowing the detector to focus on small targets of varying scales, thereby improving detection performance. The input to the dynamic feature stream is a continuous sequence of m frames within a temporal window, i.e. To reduce computational overhead while maintaining the expressive power of temporal features, DSFNet employs decomposition convolution, which involves first performing 1D temporal convolution to extract temporal information, and then performing 2D spatial convolution to extract spatial features.
[0093]
[0094] Meanwhile, to further reduce redundant information interference, DSFNet employs 3D max pooling to further reduce the feature complexity in the temporal dimension, and uses deformable convolution to align spatiotemporal features, thereby enhancing the continuous detection capability of targets and enabling stable detection of moving targets even in long time sequences. To effectively integrate static and dynamic features, DSFNet adopts a multi-scale feature fusion strategy, that is, weighted fusion of static and dynamic features on feature maps at different levels. The fusion process is as follows:
[0095]
[0096] in, Representing dynamic features, extracted by a 3D convolutional network. Representing static features, extracted by a 2D convolutional network (such as DLA-34). This represents element-wise addition, ultimately yielding the comprehensive feature. This fusion approach fully utilizes the spatial detail information provided by the static stream and combines it with the temporal consistency of the dynamic stream, thereby improving the stability and accuracy of target detection. Furthermore, the fused features are further processed through a feature pyramid to enhance the detection capability for targets at different scales, enabling DSFNet to meet the detection needs of both large-scale and small-scale targets in satellite video scenarios. Finally, the final output of the target detection is generated by the detection head. The detection head consists of three prediction branches: heatmap, offset, and bounding box, used to predict the target's center point coordinates, subtle offset of the target's center point, and the target's bounding box size, respectively. The training process of the detection head is optimized using binary cross-entropy (BCE), L1, and GIoU loss, with the objective function defined as follows:
[0097]
[0098]
[0099] in, and These are GIoU loss and L1 loss, respectively. These are derived from the output of the model regression branch, representing the center coordinates and width / height of the predicted bounding box of the target, and from the labeled data, representing the center coordinates and width / height of the true bounding box of the target. and Let represent the confidence level of the predicted target and the confidence level of the actual target, respectively. This is actually the output of the model's classification branch, representing the probability that the model believes a candidate region contains the target. The score is converted into a probability value using the Sigmoid function.
[0100]
[0101] in In the detection head, each candidate region is binary classified (foreground / background) through a classification branch, and the raw score (logits) is output. These are binary labels generated from labeled data, indicating whether a candidate region actually contains the target. By jointly optimizing these loss functions, DSFNet can effectively reduce false positives and false negatives, improving detection stability.
[0102] The improved RAFT network in step (3) addresses the challenge of accurately annotating motion flow in satellite video data, a common difficulty in training optical flow networks. This study improves upon the RAFT network as a motion flow estimator. Specifically, this is reflected in two aspects: First, during training, the target motion vectors annotated in the VDD-VEH dataset are directly used as motion flow supervision for the foreground region, replacing the pixel-level optical flow annotations commonly found in simulation data. Second, during loss calculation, the improved RAFT employs different gradient optimization strategies for the foreground and background regions, enabling the network to learn foreground target motion features while minimizing the interference of background noise on overall optical flow prediction. Figure 3 As shown, during training, the improved RAFT first determines the location of the foreground region based on the dataset annotation information and marks its corresponding region in the optical flow map. In loss calculation, the foreground region has a higher weight to make the network focus more on the motion estimation of the foreground object, while the background region has a lower weight to avoid background noise affecting the stability of optical flow prediction. (Sample) n The loss weights are defined as follows:
[0103]
[0104] in, It is a manually set hyperparameter used to enhance the loss weights of foreground targets. It is a dynamic weight allocation strategy, sample n The loss weights at position p are defined to optimize the robustness of optical flow estimation by distinguishing between foreground / background and reasonable anomalous motion. This is a binary mask used to determine whether a given location belongs to a foreground object. This indicates that it belongs to the prospective goals, while This indicates the range of motion at that point. For target points with excessively large ranges of motion (i.e., exceeding the set maximum range of motion), ... The motion flow estimator considers this an outlier and therefore excludes it from the loss calculation. Based on this, the final motion flow loss function is constructed as follows:
[0105]
[0106] Where N represents the number of samples in the training batch, and M represents the number of iterations in the RAFT network. P represents the set of all pixel locations in sample n. In the context of satellite video, each sample n corresponds to a frame pair (such as the current frame and the next frame), and P covers all pixel locations in that frame pair where optical flow loss needs to be calculated. It is the loss weight of sample n at position p. This represents the optical flow value predicted by RAFT. ) represents the motion vector of the target bounding box labeled in the VDD-VEH dataset. The weighting parameters are set to give greater weight to later iterations, ensuring the stability of the final optical flow estimation. Furthermore, to further improve the reliability of the optical flow estimation, the motion flow estimator optimizes the gradient allocation for the background region. Specifically, the update magnitude of the background region is limited when calculating the gradient to reduce its impact on foreground target prediction. This study employs an adaptive gradient clipping mechanism, first obtaining the original gradients ∇∇ of all parameters by differentiating the loss function. L Then based on the mask ( ), and filter out the parameter gradients belonging to the background region, denoted as The structure is as follows:
[0107] =▽ L ⊙(1- )
[0108] Where ⊙ represents element-wise multiplication. It is a binary mask (foreground is 1, background is 0).
[0109] Finally, update the gradient values. Apply constraints: ,in This is a dynamically adjusted upper bound on the gradient. This prevents abnormal amplification of gradients in background regions, which could affect the overall stability of optical flow prediction. In summary, the improved RAFT, through dynamic weight masking and gradient optimization strategies, organically combines spatiotemporal consistency modeling of foreground targets in satellite video with background noise suppression. The loss function guides the network to focus on key regions, gradient optimization ensures training stability, and RAFT's iterative mechanism gradually refines the prediction results. These three elements synergistically improve the robustness and accuracy of optical flow estimation.
[0110] The construction part of MPG in step (4) is as follows Figure 4 As shown, this primarily relies on the collaborative work of a spatiotemporal target detector and a motion flow estimator. The target position information output by the spatiotemporal target detector is mapped to nodes in the motion position map. Specifically, for each object i detected in time frame t, a node is added to the motion position map to represent that object. The set of nodes is shown in the formula:
[0111]
[0112] The matching relationships between objects are then established through edges in the graph. Edge generation is based on the node motion flow predicted by the motion flow estimator. During edge construction, a method based on motion flow prediction of position offset is used for determination: the optical flow information extracted by RAFT is used to predict the nodes of the current frame. Motion position in the next frame :
[0113]
[0114] in Let i be the position of object i in frame t. The motion flow vector obtained through RAFT is the optical flow estimate obtained using the improved RAFT. The inter-frame time interval is typically considered to be 1 unit of time. The distance between the predicted and actual positions is considered equal to the distance between the predicted and actual positions when the following conditions are met:
[0115]
[0116] This indicates that objects i and j have a potential spatiotemporal relationship, and a path will be established from node j in the graph. Pointing to node Edges are added to the candidate edge set E. To address the challenges posed by small, dense targets and external interference in satellite video and to improve the accuracy of target association, a MFEW strategy is further designed to assign association weights to each candidate edge. This strategy integrates spatial consistency, temporal consistency, appearance consistency, and detection confidence to comprehensively evaluate the association strength between nodes. Its calculation formula is as follows:
[0117]
[0118] in Represents a node and The correlation weight between them. The node confidence is directly provided by the output of the spatiotemporal target detector. , and These are the coefficients of three components: spatial consistency, temporal consistency, and appearance consistency. They are used to measure the contribution of different components to improving the accuracy of node association. Δp represents the spatial location difference, usually the Euclidean distance between nodes.
[0119]
[0120] The spatial Gaussian kernel standard deviation controls distance sensitivity, and Δv represents the temporal motion difference.
[0121]
[0122] The Gaussian kernel standard deviation represents the variation in motion, adjusted for velocity consistency tolerance. The cosine similarity between nodes i and j is used to measure the appearance consistency between nodes and is usually calculated using feature vectors.
[0123]
[0124] in Let represent the appearance feature vector of the i-th target in frame t. Finally, the spatial consistency embodied by the MFEW strategy is quantified by an exponential function to determine the reasonableness of the target position offset between adjacent frames. Smaller values have higher weights), and time consistency changes with speed ( To measure the continuity of movement trends, appearance consistency is based on cosine similarity (…). ) assesses the degree of matching of target appearance features, while node confidence ( As a global weighting factor, it amplifies the association priority of high-confidence detection targets, and ultimately determines the correlation priority through coefficients. Balancing the contributions of different features enables robust correlation of small, densely packed vehicles in complex satellite video scenarios.
[0125] The optimization objective of Integer Linear Programming (ILP) in step (5) is to maximize the weighted sum of the globally matched edges. After the MPG is constructed, global optimization is further performed based on this graph model to complete the full multi-target tracking task. First, Integer Linear Programming (ILP) is used to solve for the optimal matching path of the target. Integer Linear Programming (ILP) is a mathematical modeling method for global optimization problems. In multi-target tracking tasks, ILP is mainly used to solve for the optimal matching of target trajectories. Traditional multi-target tracking tasks usually rely on heuristic association strategies, such as the Hungarian algorithm or bipartite graph matching. These methods can achieve high tracking accuracy when dealing with simple scenes, but when targets are dense, occlusion is severe, or detectors make false detections, strategies that rely solely on local matching often cannot guarantee the global optimal solution. In contrast, ILP, as a mathematical optimization method, can establish the association relationship of targets at the global scale, thereby improving the integrity of the trajectory and the robustness of the matching. In multi-target tracking tasks, the core idea of ILP modeling is to transform the trajectory optimization problem into a graph optimization problem, where each detected target corresponds to a node in the graph, and the target matching relationship between adjacent frames corresponds to the edge in the graph. ILP constructs global constraints and optimizes objectives to ensure that the final generated trajectory conforms to the laws of physical motion while maximizing the overall matching confidence. Specifically, any node p or q in the motion location graph represents the detected target, and each edge... This represents possible inter-frame correlations. To find the optimal trajectory match globally, ILP defines a binary variable. Used to characterize edges Is it selected?
[0126]
[0127] when When the score is 1, it indicates that the edge is selected as part of the final trajectory; otherwise, the edge will not appear in the final result. The set of all candidate edges is denoted as E. The objective function of ILP is defined as maximizing the global score of the entire trajectory matching, which is the weighted sum of all matched edges:
[0128]
[0129] in, The matching weights for edges are typically determined by the target's position, motion pattern, appearance features, and detection confidence. To ensure the reasonableness of the trajectory, ILP needs to introduce constraints, including uniqueness constraints and binary variable constraints. Uniqueness constraints require that each target can have at most one incoming edge and one outgoing edge, preventing the same target from being matched repeatedly in multiple frames.
[0130]
[0131]
[0132] Furthermore, binary variable constraints ensure It can only take the value 0 or 1:
[0133]
[0134] These constraints ensure that the trajectory generated after ILP optimization meets the singleness requirement, that is, each target can correspond to at most one trajectory node at any given time.
[0135] The trajectory optimization module (TRM) mentioned in step (6) consists of two parts: trajectory filtering (TF) and trajectory stitching (FS).
[0136] This invention also provides a satellite video dense vehicle tracking system, including a memory, a processor, and a computer program stored in the memory. When the processor executes the program, it implements the satellite video dense vehicle tracking method as described above.
[0137] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the satellite video dense vehicle tracking method as described above.
[0138] This invention can be applied to smart city traffic monitoring, emergency vehicle dispatching, border vehicle tracking, and other fields, providing high-precision data support for macro-level traffic management. It primarily achieves three technical benefits: on the VDD-VEH dataset, MOTA reaches 71.4%, IDF1 reaches 70.5%, and IDs are reduced to 60; on datasets such as SAT-MTB and SDM-Car, it demonstrates stable generalization performance (MOTA>73%); and it supports long-term trajectory repair in complex occlusion scenarios, improving traffic flow analysis accuracy by more than 30%. The following specific embodiments further illustrate this invention.
[0139] The first step in constructing the VDD-VEH dataset was to carefully extract continuous frame sequences from video data collected by the Jilin-1 satellite. Jilin-1 satellite video was chosen because of its high resolution and wide coverage, providing rich and accurate ground information. The extracted frame sequences, to meet the requirements of subsequent data processing and model input, needed to be meticulously cropped to a strict 512×512 pixel specification. Professional image editing software was used during the cropping process to ensure the accurate size of each frame, laying the foundation for subsequent analysis. Next, a professional annotation tool was used to annotate vehicle targets in the cropped images, accurately delineating the vehicle's HBB bounding boxes. This annotation tool has an intuitive interface and high-precision annotation functions, ensuring accuracy and efficiency. Simultaneously, complex and precise algorithms were used to calculate the vehicle displacement between frames, generating motion vectors to represent the vehicle's motion state. This process required precise matching and calculation of vehicle feature points across multiple frames. Furthermore, the occlusion status of vehicles needs to be labeled, clearly distinguishing between three states: 0 (no occlusion), 1 (partial occlusion), and 2 (complete occlusion). This provides crucial information for subsequent vehicle detection in complex scenarios. In addition, dynamic background elements, such as shifting clouds and undulating water, are also labeled with appropriate dynamic tags to comprehensively record various information within the video, thus constructing a complete and high-quality VDD-VEH dataset.
[0140] In the MPG tracking process, the first stage is detection and optical flow estimation. This stage inputs five consecutive frames of images into the DSFNet network. Leveraging its unique architecture and training mechanism, this network accurately outputs the target object's location information and provides a corresponding confidence score, indicating the reliability of the detected target's location accuracy. Parallel to this, an improved RAFT network uses two closely consecutive frames as input data. By optimizing its algorithm structure, the improved RAFT network efficiently predicts the target object's motion flow, clearly showing the target's trajectory and direction between frames. It also possesses the ability to suppress background gradient interference, effectively eliminating the influence of complex and variable factors in the background on the extraction of target motion features, thus making the target's motion features more prominently presented.
[0141] The next step is graph construction and optimization. During this process, a series of key parameters need to be set, such as the node spacing threshold. The pixel threshold was determined based on in-depth research and extensive experiments on the motion patterns of the target object and the distribution density of objects in real-world scenes. The aim is to reasonably control the distance between nodes in the graph structure, ensuring that the graph structure accurately reflects the spatial relationships of the target object without affecting subsequent analysis due to excessively dense or sparse nodes. Simultaneously, hyperparameters... These hyperparameters correspond to different feature weights, used to comprehensively consider the importance of various characteristics of the target object in graph construction. After setting the parameters, the professional Gurobi solver is used to handle the integer linear programming (ILP) problem. With its powerful computing capabilities and efficient solution algorithm, Gurobi can generate the initial target object trajectory based on the previously set parameters and the information obtained in the detection and optical flow estimation stages, providing a basic framework for subsequent fine-grained optimization.
[0142] Finally, we arrive at the trajectory optimization stage. This stage also relies on carefully set parameters, specifically the TF parameters. These parameters are used to filter out unreasonable parts of the trajectory. By analyzing the distribution characteristics of trajectory points and motion trends, abnormal trajectory points caused by noise interference or false detection are eliminated, making the trajectory smoother and more consistent with actual motion logic. Regarding FS parameters... These parameters play a crucial role in the trajectory stitching process, measuring the degree of matching in position and velocity between different trajectory segments. Simultaneously, a stitching scoring threshold is set. Only when the splicing scores of different trajectory segments exceed this threshold will they be spliced together, thus ensuring that the spliced trajectory achieves a high standard in terms of motion continuity and accuracy, and realizing precise optimization of the target object's motion trajectory, such as... Figure 5 As shown.
[0143] In system deployment, the first step is to meticulously encapsulate the previously developed algorithms. This process utilizes advanced software encapsulation technology to package complex algorithmic logic into independent and fully functional software modules. Each module performs its specific function while working collaboratively to ensure the algorithm's efficient and stable operation in subsequent applications. After encapsulation, these software modules are deployed to the satellite ground processing system. This system is equipped with high-performance hardware and an optimized operating system, providing powerful computing capabilities and a stable operating environment for the algorithms. Through ingenious system integration technology, the algorithm modules are perfectly integrated with the satellite ground processing system, achieving seamless integration. The deployed system possesses excellent real-time processing capabilities, capable of instantly analyzing and processing data transmitted from the satellite at a frame rate of up to 32 FPS, quickly responding to various data requests. Simultaneously, to facilitate operators' intuitive understanding of the target object's movement, the system also integrates advanced trajectory visualization functions. By utilizing professional graphics rendering algorithms and user interface design technology, the trajectory of the target object calculated by the algorithm is displayed on the operation interface in an intuitive and clear graphical way. Whether it is the direction of the trajectory, the change of the target's position, or the motion state at different time periods, it can be seen at a glance, which greatly improves the readability and operability of the data and provides strong support for decision-making and analysis in related fields.
Claims
1. A method for tracking dense vehicles using satellite video, characterized in that, Includes the following steps: (1) Construct a satellite video dense vehicle tracking dataset VDD-VEH, which contains multiple frames of satellite video images and their corresponding vehicle target spatial location annotations and inter-frame motion vector annotations; (2) Based on the spatiotemporal target detector DSFNet, the detection information of vehicle targets in continuous video frames is extracted to generate a set of nodes containing target location and confidence level; (3) The improved RAFT network is used as a motion flow estimator to estimate the motion flow of vehicle targets between adjacent frames, and cross-frame candidate association edges are constructed based on the motion flow; (4) Based on the estimation results obtained from the spatiotemporal target detector DSFNet and the motion flow estimator, a motion location map is constructed, that is, the position of each detected target is mapped to a node in the motion location map, and then the estimated motion flow is used to establish connection edges between nodes in adjacent frames. In step (4), the target position information output by the spatiotemporal target detector is mapped to nodes of the motion position map. Specifically, for each object i detected in time frame t, a node is added to the motion position map to represent the object, and its node set is as follows: The matching relationships between objects are then established through edges in the graph. Edge generation is based on the node motion flow predicted by the motion flow estimator. During edge construction, a method based on motion flow prediction of position offset is used for determination: the optical flow information extracted by RAFT is used to predict the nodes of the current frame. Motion position in the next frame : in Let i be the position of object i in frame t. The motion flow vector obtained through RAFT is the optical flow estimate obtained using the improved RAFT. The inter-frame time interval is defined when the distance between the predicted and actual positions meets the following condition: This indicates that objects i and j have a potential spatiotemporal relationship, and a path will be established from node j in the graph. Pointing to node The edges are added to the candidate edge set E; A multi-feature edge weighting (MFEW) strategy is further designed to assign association weights to each candidate edge. This strategy integrates spatial consistency, temporal consistency, appearance consistency, and detection confidence to comprehensively evaluate the association strength between nodes. The calculation formula is as follows: in Represents a node and The correlation weight between them The node confidence is directly provided by the output of the spatiotemporal target detector. , and Δp is the coefficient of three components: spatial consistency, temporal consistency, and appearance consistency. It is used to measure the contribution of different components to improving the accuracy of node association. Δp represents the spatial location difference. Δv represents the spatial Gaussian kernel standard deviation, controlling for distance sensitivity, and Δv represents the temporal motion difference. The Gaussian kernel standard deviation represents the variation in motion, adjusted for velocity consistency tolerance. The cosine similarity between nodes i and j is used to measure the appearance consistency between nodes. (5) Globally optimize the motion position map based on integer linear programming to generate the initial target trajectory; (6) Filter and stitch the initial trajectory, remove abnormal trajectories and repair trajectory breaks, and output the final vehicle tracking result.
2. The satellite video dense vehicle tracking method according to claim 1, characterized in that: The construction of the VDD-VEH dataset in step (1) includes: cropping image blocks of a certain size from satellite video, labeling the horizontal bounding box of the vehicle target and the corresponding motion vector, the motion vector including forward velocity and backward velocity, and labeling the occlusion state of the target and background dynamic interference attributes.
3. The satellite video dense vehicle tracking method according to claim 1, characterized in that: The spatiotemporal target detector DSFNet described in step (2) adopts a dual-stream feature extraction architecture, in which a static feature stream extracts spatial information of a single frame through 2D convolution; a dynamic feature stream extracts time series information through lightweight 3D convolution; and a hierarchical feature fusion strategy is adopted to make features of different scales complement each other, and finally outputs the bounding box, category information and confidence score of the target through the detection head.
4. The satellite video dense vehicle tracking method according to claim 1, characterized in that: The improved RAFT network described in step (3) first determines the location of the foreground region based on the dataset annotation information during training and marks its corresponding region in the optical flow map. During loss calculation, the foreground region has a higher weight, while the background region has a lower weight. The loss weights for sample n are defined as follows: in, It is a manually set hyperparameter used to enhance the loss weights of foreground targets. It is a dynamic weight allocation strategy. The loss weight of sample n at position p is defined, and the robustness of optical flow estimation is optimized by distinguishing between foreground and background and reasonable abnormal motion. This is a binary mask used to determine whether a given location belongs to a foreground object. This indicates that it belongs to the prospective goals, while This indicates the range of motion at that point; for target points with excessively large ranges of motion, i.e., exceeding the set maximum range of motion... The motion flow estimator considers it an outlier and therefore excludes it from the loss calculation; based on this, the final motion flow loss function is constructed as follows: Where N represents the number of samples in the training batch, and M represents the number of iterations in the RAFT network. P represents the set of all pixel locations in sample n; in the context of satellite video, each sample n corresponds to a frame pair, and P covers all pixel locations in that frame pair for which optical flow loss needs to be calculated. It is the loss weight of sample n at position p. This represents the optical flow value predicted by RAFT. ) represents the motion vector of the target bounding box labeled in the VDD-VEH dataset; These are the weight parameters.
5. The satellite video dense vehicle tracking method according to claim 1, characterized in that: The motion flow estimator also optimizes the gradient allocation for the background region. Specifically, it limits the update magnitude of the background region when calculating the gradient to reduce its impact on foreground target prediction. Specifically, it employs an adaptive gradient clipping mechanism, first obtaining the original gradients of all parameters by differentiating the loss function. L Then based on the mask The parameter gradients belonging to the background region are selected and denoted as follows. The structure is as follows: = L ⊙(1 ) Where ⊙ represents element-wise multiplication. It's a binary mask, with foreground values of 1 and background values of 0; finally, the gradient values are updated. Apply constraints: ,in It is a dynamically adjusted upper limit value for the gradient.
6. The satellite video dense vehicle tracking method according to claim 1, characterized in that, The process of integer linear programming (ILP) described in step (5) is as follows: In the motion location graph, any node p or q represents the detected target, and each edge... To represent possible inter-frame correlations, ILP defines a binary variable to find the optimal trajectory match globally. Used to characterize edges Is it selected? when When the value is 1, it indicates that the edge is selected as part of the final trajectory; otherwise, the edge will not appear in the final result. The set of all candidate edges is denoted as E. The objective function of ILP is defined as maximizing the global score of the entire trajectory matching, which is the weighted sum of all matching edges: in, The matching weights are set for the edges. To ensure the rationality of the trajectory, ILP introduces constraints, including uniqueness constraints and binary variable constraints. The uniqueness constraint requires that each target can have at most one incoming edge and one outgoing edge, preventing the same target from being matched repeatedly in multiple frames. Furthermore, binary variable constraints ensure It can only take the value 0 or 1: These constraints ensure that the trajectory generated after ILP optimization meets the singleness requirement, that is, each target can correspond to at most one trajectory node at any given time.
7. The satellite video dense vehicle tracking method according to claim 1, characterized in that: In step (6), the initial trajectory is filtered and stitched together by the trajectory optimization module. The trajectory optimization module consists of two parts: trajectory filtering (TF) and trajectory stitching (FS). TF eliminates false trajectories caused by false detection by analyzing the distribution characteristics of trajectory points and trajectory displacement. Specifically, it calculates the spatial concentration and total displacement of the trajectory point sequence. The process is as follows: Let the trajectory point sequence be... in Let k represent the coordinates of the midpoint of the trajectory, L be the length of the trajectory, and the point distribution characteristics of the trajectory be defined as follows: in It is located in Centered on, with radius The number of points within the circle; the total displacement of the trajectory is calculated as follows: Setting a spatial discrete threshold in TF With minimum displacement threshold If the spatial dispersion of the trajectory exceeds Or the total displacement is less than If the trajectory is low quality, it will be identified as a low-quality trajectory and removed.
8. The satellite video dense vehicle tracking method according to claim 7, characterized in that: FS continuously generates virtual nodes for trajectory fragments. These virtual nodes are used to stitch together trajectory fragments belonging to the same target, improving trajectory integrity. Specifically, by setting the total number of frames in the current video sequence to N, FS generates virtual nodes for all initial trajectory sets. Each trajectory segment The start and end frames are respectively denoted as The start and end position coordinates are denoted as follows: , The set of trajectory fragments to be pieced together. Defined as: in This indicates the image edge region, meaning the trajectory endpoint is not in the last frame of the video sequence. Only trajectories whose endpoints are not located in the image edge region are considered potentially broken trajectory fragments and proceed to the subsequent stitching process. The selected candidate trajectory fragments are then stitched together, and the spatial distance and velocity difference between the predicted location and the actual trajectory starting point are calculated. Specifically, assuming from the set... Extract a trajectory segment to be spliced from the data. The trajectory ends at time [time]. The final position and velocity are denoted as follows: and In order to determine whether the trajectory can be compared with another trajectory splicing, including the trajectory The starting time and starting position are denoted as follows: FS uses a Kalman filter to analyze the trajectory Generate continuous virtual nodes and predict their trajectory. Position of the starting time With speed : Subsequently, FS calculates the trajectory. Predicted location and trajectory Spatial distance between actual starting positions and the difference between predicted speed and actual speed : Based on the above difference values, a trajectory splicing matching scoring function is defined: in, and The normalized thresholds and weighting coefficients represent the differences in spatial location and velocity, respectively. and Used to balance the importance of position and velocity, if rated Exceeding the threshold Then, virtual nodes are inserted between the breakpoints to form a continuous trajectory.
9. A satellite video-dense vehicle tracking system, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: When the processor executes the program, it implements the satellite video dense vehicle tracking method as described in any one of claims 1-8.
Citation Information
Patent Citations
Multi-target tracking method and system based on graph neural network
CN111161315A
Multi-vehicle target tracking method based on deep learning
CN114862910A