An Accelerated Method for Generating Multi-Temporal Individual Trajectories Based on Satellite Video Data

By processing the position and motion state prediction of satellite video data in parallel at the main node, and combining feature extraction on the domestic graphical computing chip, multi-node parallel computing and Hungarian algorithm are used to match, the problem of low trajectory generation efficiency of multi-time sensitive individuals with limited resources in traditional methods is solved, and fast and accurate trajectory generation is achieved.

CN119027458BActive Publication Date: 2025-07-08XIAN SPACE STAR TECH IND GRP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411506281.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-07-08
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

The traditional multi-time sensitive individual trajectory generation method is inefficient when processing high-resolution satellite video data, and has high hardware resource requirements, making it difficult to meet practical application needs, especially when using domestic chips.

Method used

The prediction of the position and motion state of the main node in parallel time-sensitive individual and motion feature extraction are used, and feature extraction is performed on the domestic graph computing chip with adaptive acceleration Kalman filtering and SpeedNet detector. The Mahalanobis distance and Cosine distance calculation are used to calculate the Mahalanobis distance and Cosine distance, and match it through the Hungarian algorithm.

Benefits of technology

It realizes rapid and accurate multi-time sensitive individual trajectory generation under limited resource conditions, improves processing speed and reduces the amount of parameters, and solves the problem of processing bottlenecks when resource constraints are encountered.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119027458B_ABST
    Figure CN119027458B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for accelerating the generation of multi-temporal sensitive individual trajectories based on satellite video data. In the present invention, the prediction of the positions and motion states of multi-temporal sensitive individuals and the extraction of motion features are completed in parallel on the master node, and the time-consuming Mahalanobis distance calculation and Cosine distance calculation are completed in parallel on multiple nodes (one master node + multiple slave nodes), making full use of the characteristics of the multi-node hardware architecture of the terminal device, correctly tracking more multi-temporal sensitive individuals, and improving the overall processing timeliness of the multi-temporal sensitive individual trajectory generation algorithm while ensuring accuracy; the SpeedNet detector is used to complete the extraction of multi-temporal sensitive individual motion features, the core of which is the adoption of a self-supervised lightweight and compact detection head, and a grouped convolution parallel network structure is adopted to fuse features with different receptive fields, enabling a single detection head to also use features of different scales, simplifying the entire processing flow, significantly reducing the number of parameters, and significantly improving the processing speed, ultimately achieving the acceleration of the generation of multi-temporal sensitive individual trajectories in satellite video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of near-real-time target tracking based on satellite video data, and particularly relates to an acceleration method for generating multi-temporal sensitive individual trajectories based on satellite video data. Background Art

[0002] In recent years, various methods for generating multi-temporal sensitive individual trajectories from satellite video data have emerged continuously, promoting great development in this field. However, with the higher resolution, larger size, and deeper bit depth of satellite video data, traditional methods for generating multi-temporal sensitive individual trajectories have bottleneck problems such as complex business processes, a large number of memory reads and writes, and processing of a huge number of parameters, resulting in low processing timeliness. To meet the demand for rapid generation of multi-temporal sensitive individual trajectories based on satellite video data, extremely high requirements are imposed on hardware resources, and many devices in the production environment are difficult to meet these hardware resource requirements, leading to the difficulty in generating multi-temporal sensitive individual trajectories based on satellite video data to meet actual application scenarios. Therefore, conducting research on acceleration methods for generating multi-temporal sensitive individual trajectories based on satellite video data has important practical significance for promoting the development of the field.

[0003] With the rapid development of domestic chips, the timeliness of using traditional methods for rapidly generating multi-temporal sensitive individual trajectories from satellite video data is greatly limited. Therefore, how to utilize the limited computing power of domestic chips to rapidly and efficiently generate multi-temporal sensitive individual trajectories from satellite video data within an acceptable accuracy loss range has become a frontier and practical issue in current research. Summary of the Invention

[0004] The purpose of the present invention is to provide an acceleration method for generating multi-temporal sensitive individual trajectories based on satellite video data. By means of acceleration such as parallel prediction of the positions and motion states of multi-temporal sensitive individuals, extraction of motion features, and initial matching of multi-temporal sensitive individuals at multiple nodes on the main node, the acceleration of generating multi-temporal sensitive individual trajectories in satellite video data is achieved.

[0005] The technical solution adopted by the present invention is an acceleration method for generating multi-temporal sensitive individual trajectories based on satellite video data, including the following steps:

[0006] S1, the main node receives the position and confidence extraction results of multi-temporal sensitive individuals in the current frame of satellite video data;

[0007] S2, create 2 threads and start working simultaneously. The working mode of one thread is to use the historical trajectory information of multi-temporal sensitive individuals to apply adaptive acceleration Kalman filtering to complete the prediction of positions and motion states on the CPU, and the working mode of the other thread is to use the SpeedNet detector to complete the extraction of motion features of multi-temporal sensitive individuals on the domestic graphics computing chip;

[0008] The adaptive acceleration Kalman filter consists of 12-dimensional motion parameters and an adaptive noise covariance matrix R, where (C x , C y , Z, L) are the horizontal and vertical positions of the center point of the target position D in the current frame of satellite video data, the aspect ratio of the target, and the width of the target, represents the velocity vector S of the target in the current frame of satellite video data, represents the acceleration vector P;

[0009] The core of the SpeedNet detector is to adopt a self-supervised lightweight compact detection head to extract the target motion features in the current frame of satellite video data nearly in real time with a small number of parameters under the low-power condition of domestic chips;

[0010] S3. According to the prediction results of the time-sensitive individual position and motion state, it is judged whether it is in the tracking state. If it is in the tracking state, combined with the extraction results of the time-sensitive individual motion features, multi-node initial matching of time-sensitive individuals is performed; if it is not in the tracking state or the initial matching of the tracking state fails, EfficientIou time-sensitive individual re-matching is performed;

[0011] The multi-node initial matching of time-sensitive individuals adopts a multi-node distributed parallel computing method to implement the calculation of the minimum association cost distance matrix MinCostMatching. The value of each element in the matrix MinCostMatching is obtained by weighting the Mahalanobis distance and the Cosine distance. The master node divides the calculation of the Mahalanobis distance and the Cosine distance into multiple tasks through a multi-task distribution mode, leaves one task to be executed on the master node, and distributes the remaining tasks to the slave nodes for execution. After the tasks of the slave nodes are completed, the results of each slave node are merged to the master node to complete the reconstruction of the matrix MinCostMatching;

[0012] The EfficientIou time-sensitive individual re-matching implements the calculation of the EfficientIou time-sensitive individual re-matching cost matrix EffiouCost in a multi-threaded manner on the master node and completes the optimal matching through the Hungarian algorithm, where each element in the cost matrix EffiouCost is the EfficientIou distance between the extraction results in the current frame of satellite video data and the targets that failed in the multi-node initial matching of time-sensitive individuals or the trajectories of non-tracking states;

[0013] S4. If the multi-node initial matching of time-sensitive individuals is successful or the EfficientIou time-sensitive individual re-matching is successful, the trajectory of the matched extraction result is updated; if the EfficientIou time-sensitive individual re-matching fails, a new trajectory ID is assigned to the unmatched extraction result, and at the same time, the trajectory generation result of the time-sensitive individual is output.

[0014] Furthermore, the specific steps of S2 are as follows:

[0015] B1. Based on the state information of time-sensitive individuals in the (t - 1)th frame, use adaptive acceleration Kalman filtering to predict the state of time-sensitive individuals in the current frame of satellite video data. The specific calculation formula is as follows:

[0016]

[0017] where t represents the current frame of satellite video data; t - 1 represents the previous frame of the current frame of satellite video data; Q t is the motion state of the tth frame, which is composed of the target position D, velocity vector S, and acceleration vector P; J t-1 represents the prediction process noise; Δt is the imaging time difference between the tth frame and the (t - 1)th frame of satellite video data;

[0018] B2. Use the position information of time-sensitive individuals to extract the pixel matrix corresponding to the current frame of satellite video data, perform Lanczos interpolation size scaling on the pixel matrix to complete the unification of the input size, and use the SpeedNet detector to extract the motion features of time-sensitive individuals on the domestic graphics computing chip;

[0019] The order of steps B1 and B2 can be adjusted and there is no priority.

[0020] Furthermore, the core of the SpeedNet detector is to adopt a self-supervised lightweight compact detection head. The specific implementation steps of the self-supervised lightweight compact detection head are as follows:

[0021] b1. Input the features of time-sensitive individuals at different scales;

[0022] b2. Use depthwise separable convolution to unify the number of input channels of each branch;

[0023] b3. Use 2 groups of different convolution methods to serially form the SpeedHost module. The first group of convolution consists of 1 depthwise separable convolution + HSwish activation function, and the second group of convolution adopts 2 depthwise separable convolutions + HSwish activation function. Finally, use projection convolution to achieve the fusion of the two groups of convolutions at the channel scale;

[0024] b4. Perform projection convolution + HSwish activation function again to complete the further compression and dimensionality reduction of the features;

[0025] b5. Use 3 groups of different convolutions to map the feature map to a higher dimension to obtain stronger fitting ability. Each group of convolutions consists of a different number of depthwise separable convolutions + HSwish activation function in series. Use feature concatenation to achieve the fusion of features with different receptive fields, and can more accurately align and extract the key features in the image, especially when the shape and size of the unified target vary in different frames of satellite video data.

[0026] Further, in step S3, the specific steps for initial matching of multi-node time-sensitive individuals are as follows:

[0027] c11. The master node divides the time-sensitive individual extraction results, the trajectories of time-sensitive individuals in the tracking state, and the appearance features extracted by the SpeedNet detector into multiple tasks. One task is left at the master node, and the remaining tasks are assigned to each slave node for execution;

[0028] c12. The master node and the slave nodes respectively calculate the Mahalanobis distance cost matrix and the Cosine distance cost matrix on their respective nodes;

[0029] The calculation formula of the Mahalanobis distance cost matrix is:

[0030] M(m,n)=(P m -Q n ) T H -1 (P m -Q n )

[0031] Where P m is the extraction result matrix of the m-th time-sensitive individual; Q n is the trajectory matrix of the n-th time-sensitive individual in the tracking state; H is the covariance matrix of the time-sensitive individual predicted by the adaptive acceleration Kalman filter;

[0032] The calculation formula of the Cosine distance cost matrix is:

[0033]

[0034] Where r n represents the appearance feature matrix of the n-th detection box; represents the appearance feature of the m-th trajectory extracted by the SpeedNet detector;

[0035] c13. The master node receives the results of the Mahalanobis distance cost matrix and the Cosine distance cost matrix from the slave nodes and completes the aggregation of the two cost matrices to reconstruct the minimum association cost distance matrix MinCostMatching. The specific calculation formula is:

[0036] MinCostMatching=glg g DisMah+(1 - g)lg 1-g DisCos

[0037] Where DisMah is the overall Mahalanobis distance cost matrix; DisCos is the overall Cosine distance cost matrix; g represents the weight coefficient;

[0038] c14. Screen the minimum correlation cost distance matrix MinCostMatching according to the maximum threshold MaxDistance, assign infinity to the elements in the matrix that exceed the threshold, and use the Hungarian matching algorithm to complete the initial matching of multi-node time-sensitive individuals for MinCostMatching.

[0039] Furthermore, in the step S3, the specific steps for performing EfficientIou time-sensitive individual re-matching are as follows:

[0040] c21. Calculate the EfficientIou time-sensitive individual re-matching cost matrix EffiouCost;

[0041] c22. Use the Hungarian matching algorithm to perform an optimal matching on the cost matrix EffiouCost.

[0042] Furthermore, in the step c21, the formula for calculating the EfficientIou time-sensitive individual re-matching cost matrix EffiouCost is:

[0043]

[0044] where EffiouCost m,n represents the intersection over union of the extraction result of the m-th time-sensitive target in the current frame of satellite video data and the target or non-tracked state trajectory that fails the initial matching of the n-th time-sensitive individual of the multi-node; det m represents the extraction result of the m-th time-sensitive target in the current frame of satellite video data; represents the target or non-tracked state trajectory that fails the initial matching of the n-th time-sensitive individual of the multi-node; k c and g c respectively represent the width and height of the minimum bounding rectangle of the two frames of the extraction result of the satellite video data in the current frame and the trajectory of the time-sensitive individual of the multi-node that fails the initial matching or is in the non-tracked state; k m and respectively represent the width of the m-th extraction result of the satellite video data in the current frame and the width of the target or non-tracked state that fails the initial matching of the n-th time-sensitive individual of the multi-node; g m and respectively represent the height of the m-th extraction result of the satellite video data in the current frame and the height of the target or non-tracked state that fails the initial matching of the n-th time-sensitive individual of the multi-node; Φ 2 represents the Euclidean distance.

[0045] The beneficial effects of the present invention are as follows:

[0046] (1) The present invention realizes a method for accelerating the generation of multi-temporal sensitive individual trajectories based on satellite video data. It splits the steps of the multi-temporal sensitive individual trajectory generation algorithm and reconstructs the logic adaptively. By predicting the positions and motion states of the time-sensitive individuals and extracting motion features in parallel on the master node, and calculating the Mahalanobis distance and Cosine distance, which are time-consuming, in parallel on multiple nodes, it makes full use of the characteristics of the multi-node hardware architecture of the terminal device, and solves the bottleneck problems of low tracking accuracy and poor processing timeliness of existing traditional processing algorithms under resource constraints. At the same time, it can quickly and correctly track more time-sensitive individuals, and also provides strong support for the engineering deployment and application of the multi-temporal sensitive individual trajectory generation algorithm.

[0047] (2) Use the SpeedNet detector to extract the motion features of the time-sensitive individuals. Its core is to adopt a self-supervised lightweight and compact detection head, which uses a grouped convolution parallel network structure to fuse the features of different receptive fields, enabling a single detection head to also use features of different scales, simplifying the entire processing flow, significantly reducing the number of parameters, and significantly improving the processing speed, ultimately achieving the acceleration of the generation of multi-temporal sensitive individual trajectories in satellite video data. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic diagram of the overall process of the method of the present invention.

[0049] Figure 2 It is a partial display of the trajectory generation result of the method of the present invention Figure 1 。

[0050] Figure 3 It is a partial display of the trajectory generation result of the method of the present invention Figure 2 。 DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings.

[0052] A method for accelerating the generation of multi-temporal sensitive individual trajectories based on satellite video data aims to, under the condition of limited terminal computing resources, meet the high timeliness and high-precision requirements for the generation of time-sensitive individual trajectories, and realize the acceleration of the processing of generating multi-temporal sensitive individual trajectories based on satellite video data, as Figure 1 shown, including the following steps:

[0053] S1. The master node receives the position and confidence extraction results of the time-sensitive individuals in the current frame of the satellite video data.

[0054] S2. Create 2 threads and start working simultaneously. The working mode of one thread is to use the historical trajectory information of time-sensitive individuals to apply adaptive acceleration Kalman filtering to complete the prediction of position and motion state on the CPU (FT-D2000). The working mode of the other thread is to use the SpeedNet detector to complete the extraction of time-sensitive individual motion features on the domestic graphics computing chip (Atlas200).

[0055] The adaptive acceleration Kalman filter consists of 12-dimensional motion parameters and an adaptive noise covariance matrix R, where (C x , C y , Z, L) are the horizontal and vertical positions of the center point of the target position D in the current frame of satellite video data, the aspect ratio of the target, and the width of the target, represents the velocity vector S of the target in the current frame of satellite video data, represents the acceleration vector P.

[0056] The core of the SpeedNet detector is to adopt a self-supervised lightweight compact detection head to complete the near-real-time extraction of the motion features of the target in the current frame of satellite video data with a small number of parameters under the low-power condition of domestic chips.

[0057] The specific steps to use the historical trajectory information of time-sensitive individuals to apply adaptive acceleration Kalman filtering to complete the prediction of position and motion state on the CPU and use the SpeedNet detector to complete the extraction of time-sensitive individual motion features on the domestic graphics chip are as follows:

[0058] B1. Based on the state information of the time-sensitive individual in the (t - 1)th frame, use the adaptive acceleration Kalman filter to predict the state of the time-sensitive individual in the current frame of satellite video data. The specific calculation formula is:

[0059]

[0060] where t represents the current frame of satellite video data, t - 1 represents the previous frame of the current frame of satellite video data, Q t is the motion state of the tth frame, which consists of the target position D, the velocity vector S, and the acceleration vector P; J t-1 represents the process noise of the prediction; Δt is the imaging time difference between two consecutive frames of images;

[0061] B2. Use the position information of the time-sensitive individual to extract the corresponding pixel matrix of the current frame of satellite video data, perform Lanczos interpolation size scaling on the pixel matrix, uniformly scale it to the pixel size of 1024×1024, and use the SpeedNet detector to complete the extraction of time-sensitive individual motion features on the domestic graphics computing chip.

[0062] The core of the SpeedNet detector is to adopt a self-supervised lightweight compact detection head. The specific implementation steps of the self-supervised lightweight compact detection head are as follows:

[0063] b1, Input the features of time-sensitive individuals at different scales;

[0064] b2, Use dilated convolutions to unify the number of input channels for each branch, significantly improving the feature extraction ability and reducing the network parameters;

[0065] b3, Use two different convolution methods in series to form the SpeedHost module. The first group of convolutions consists of 1 depthwise separable convolution + HSwish activation function, and the second group of convolutions consists of 2 depthwise separable convolutions + HSwish activation function. Finally, use projection convolution to achieve the fusion of the two groups of convolutions at the channel scale. Through split feature extraction and channel fusion, the number of parameters is significantly reduced;

[0066] b4, Perform projection convolution + HSwish activation function again to complete further compression and dimensionality reduction of the features, effectively reducing the number of model parameters and computational complexity;

[0067] b5, Use three different convolutions to map the feature map to a higher dimension to obtain stronger fitting ability. Each group of convolutions consists of a different number of depthwise separable convolutions + HSwish activation function in series. Use feature concatenation to achieve the fusion of features with different receptive fields, and can more accurately align and extract key features in the image, especially when the shape and size of the unified target vary in different-frame satellite video data.

[0068] It should be noted that the order of steps B1 and B2 can be adjusted and there is no priority.

[0069] S3, Judge whether it is in the tracking state according to the prediction results of the time-sensitive individual's position and motion state. If it is in the tracking state, combine the time-sensitive individual's motion feature extraction results to perform multi-node time-sensitive individual initial matching; if it is not in the tracking state or the initial matching of the tracking state fails, perform EfficientIou time-sensitive individual re-matching.

[0070] When there are multiple nodes, the initial matching of time-sensitive individuals adopts a multi-node distributed parallel computing method to implement the calculation of the minimum association cost distance matrix MinCostMatching. The value of each element in the matrix MinCostMatching is obtained by weighting the Mahalanobis distance and the Cosine distance. The master node divides the calculation of the Mahalanobis distance and the Cosine distance into multiple tasks through a multi-task distribution mode, leaving one task to be executed on the master node and distributing the remaining tasks to the slave nodes for execution. After the tasks of the slave nodes are completed, the results of each slave node are merged to the master node to complete the reconstruction of the matrix MinCostMatching. The specific steps are as follows:

[0071] c11. The master node divides the time-sensitive individual extraction results, the trajectories of time-sensitive individuals in the tracking state, and the appearance features extracted by the SpeedNet detector into multiple tasks, leaving one task on the master node and distributing the remaining tasks to each slave node for execution;

[0072] c12. The master node and the slave nodes respectively calculate the Mahalanobis distance cost matrix and the Cosine distance cost matrix on their respective nodes;

[0073] The calculation formula of the Mahalanobis distance cost matrix is:

[0074] M(m,n)=(P m -Q n ) T H -1 (P m -Q n )

[0075] Among them, P m is the extraction result matrix of the m-th time-sensitive individual; Q n is the trajectory matrix of the n-th time-sensitive individual in the tracking state; H is the covariance matrix of the time-sensitive individual predicted by the adaptive acceleration Kalman filter;

[0076] The calculation formula of the Cosine distance cost matrix is:

[0077]

[0078] Among them, r n represents the appearance feature matrix of the n-th detection frame; represents the appearance feature of the m-th trajectory extracted by the SpeedNet detector;

[0079] c13. The master node receives the results of the Mahalanobis distance cost matrix and the Cosine distance cost matrix from the slave nodes and completes the summary of the two cost matrices to reconstruct the minimum association cost distance matrix MinCostMatching. The specific calculation formula is:

[0080] MinCostMatching = glg g DisMah+(1 - g)lg 1-g DisCos

[0081] Among them, DisMah is the overall Mahalanobis distance cost matrix; DisCos is the overall Cosine distance cost matrix; g represents the weight coefficient;

[0082] c14. Screen the minimum associated cost distance matrix MinCostMatching according to the maximum threshold MaxDistance, assign infinity to the elements in the matrix that exceed the threshold, and use the Hungarian matching algorithm to complete the initial matching of multi-node time-sensitive individuals for MinCostMatching.

[0083] EfficientIou time-sensitive individual re-matching implements the calculation of the EfficientIou time-sensitive individual re-matching cost matrix EffiouCos in a multi-threaded manner at the master node, and completes the optimal matching through the Hungarian algorithm. Each element in the cost matrix EffiouCos is the EfficientIou distance between the extraction result of the current frame of satellite video data and the target that fails the initial matching of multi-node time-sensitive individuals or the trajectory in the non-tracking state. The specific steps are as follows:

[0084] c21. Calculate the EfficientIou time-sensitive individual re-matching cost matrix EffiouCos, and the calculation formula is:

[0085]

[0086] Among them, EffiouCost m,n represents the intersection over union of the extraction result of the m-th time-sensitive target in the current frame of satellite video data and the target that fails the initial matching of the n-th multi-node time-sensitive individual or the trajectory in the non-tracking state; det m represents the extraction result of the m-th time-sensitive target in the current frame of satellite video data; represents the target that fails the initial matching of the n-th multi-node time-sensitive individual or the trajectory in the non-tracking state; k c 、g c respectively represent the width and height of the minimum bounding rectangle of the two bounding boxes of the extraction result of the current frame of satellite video data and the target that fails the initial matching of multi-node time-sensitive individuals or the trajectory in the non-tracking state; k m 、 respectively represent the width of the m-th extraction result of the current frame of satellite video data and the width of the target that fails the initial matching of the n-th multi-node time-sensitive individual or the trajectory in the non-tracking state; g m 、 respectively represent the height of the extraction result of the current frame of the m-th satellite video data, the target of the initial matching failure of the n-th time-sensitive individual of multiple nodes, or the height of the non-tracking state; Φ 2 represents the Euclidean distance.

[0087] c22, use the Hungarian matching algorithm to perform optimal matching on the cost matrix EffiouCos.

[0088] S4, when the initial matching of the time-sensitive individual of multiple nodes is successful or the re-matching of the time-sensitive individual of EfficientIou is successful, update the trajectory of the matched extraction result; when the re-matching of the time-sensitive individual of EfficientIou fails, assign a new trajectory ID to the unmatched extraction result, and at the same time output the trajectory generation result of the time-sensitive individual. The trajectory generation processing time of the time-sensitive individual is 2.1 seconds (processing 100 frames of satellite video data), which is about 2.5 times more efficient than the traditional method. The output trajectory generation result of the time-sensitive individual is as Figure 2 、 Figure 3 shown.

[0089] Contents not described in detail in this specification are all well-known prior arts to those skilled in the art.

Claims

1. A method for accelerating the generation of multi-temporal sensitive individual trajectories based on satellite video data, characterized in that, It includes the following steps: S1. The master node receives the position and confidence extraction results of the time-sensitive individuals in the current frame of satellite video data; S2. Create 2 threads and start working simultaneously. The working mode of one thread is to use the historical trajectory information of the time-sensitive individuals to apply adaptive acceleration Kalman filtering to complete the prediction of the position and motion state on the CPU, and the working mode of the other thread is to use the SpeedNet detector to complete the extraction of the motion features of the time-sensitive individuals on the domestic graphics computing chip; Adaptive acceleration Kalman filter consists of 12-dimensional motion parameters and an adaptive noise covariance matrix R, where (C x , C y , Z, L) are the horizontal and vertical positions of the center point of the target position D in the current frame of satellite video data, the aspect ratio of the target, and the width of the target, represents the velocity vector S of the target in the current frame of satellite video data, represents the acceleration vector P; The core of the SpeedNet detector is to adopt a self-supervised lightweight compact detection head to complete the near-real-time extraction of the target motion features in the current frame of satellite video data with a small number of parameters under the low-power condition of the domestic chip; S3. Judge whether it is in the tracking state according to the prediction results of the position and motion state of the time-sensitive individuals. If it is in the tracking state, combine the extraction results of the motion features of the time-sensitive individuals to perform the initial matching of the time-sensitive individuals among multiple nodes; if it is not in the tracking state or the initial matching of the tracking state fails, perform the EfficientIou re-matching of the time-sensitive individuals; The initial matching of the time-sensitive individuals among multiple nodes is implemented by using the multi-node distributed parallel computing method to calculate the minimum association cost distance matrix MinCostMatching. The value of each element in the matrix MinCostMatching is obtained by weighting the Mahalanobis distance and the Cosine distance. The master node divides the calculation of the Mahalanobis distance and the Cosine distance into multiple tasks through the multi-task distribution mode, leaves one task to be executed on the master node, and distributes the remaining tasks to the slave nodes for execution. After the tasks of the slave nodes are completed, the results of each slave node are merged to the master node to complete the reconstruction of the matrix MinCostMatching; The EfficientIou re-matching of the time-sensitive individuals is to calculate the EfficientIou re-matching cost matrix EffiouCost of the time-sensitive individuals in a multi-threaded manner on the master node and complete the optimal matching through the Hungarian algorithm, where each element in the cost matrix EffiouCost is the EfficientIou distance between the extraction results in the current frame of satellite video data and the targets that failed in the initial matching of the time-sensitive individuals among multiple nodes or the trajectories of non-tracking states; S4. If the initial matching of the time-sensitive individuals among multiple nodes is successful or the EfficientIou re-matching of the time-sensitive individuals is successful, update the trajectories of the matched extraction results; if the EfficientIou re-matching of the time-sensitive individuals fails, assign new trajectory IDs to the unmatched extraction results, and at the same time output the trajectory generation results of the time-sensitive individuals.

2. The multi-temporal sensitive individual trajectory generation acceleration method based on satellite video data according to claim 1, wherein The specific steps of S2 are as follows: B1. Based on the state information of the time-sensitive individuals in the (t - 1)th frame, use adaptive acceleration Kalman filtering to predict the state of the time-sensitive individuals in the current frame of satellite video data. The specific calculation formula is: Among them, t represents the current frame of satellite video data; t-1 represents the previous frame of the current frame of satellite video data; Q t is the motion state of the t-th frame, which is composed of the target position D, the velocity vector S, and the acceleration vector P; J t-1 represents the predicted process noise; Δt is the imaging time difference between the t-th frame and the (t-1)-th frame of satellite video data; B2. Extract the pixel matrix corresponding to the current frame of satellite video data using the position information of the time-sensitive individual, perform Lanczos interpolation size scaling on the pixel matrix to complete the unification of the input size, and use the SpeedNet detector to extract the motion features of the time-sensitive individual on the domestic graphics computing chip; The order of steps B1 and B2 can be adjusted and there is no priority.

3. The multi-temporal sensitive individual trajectory generation acceleration method based on satellite video data according to claim 2, wherein, The core of the SpeedNet detector is to adopt a self-supervised lightweight compact detection head. The specific implementation steps of the self-supervised lightweight compact detection head are as follows: b1. Input the features of time-sensitive individuals at different scales; b2. Use dilated convolution to unify the number of input channels of each branch; b3. Use 2 groups of different convolution methods in series to form the SpeedHost module. The first group of convolution consists of 1 depthwise separable convolution + HSwish activation function, and the second group of convolution consists of 2 depthwise separable convolution + HSwish activation function. Finally, use projection convolution to achieve the fusion of the two groups of convolution at the channel scale; b4. Perform projection convolution + HSwish activation function again to complete the dimensionality reduction and compression of the features; b5. Use 3 groups of different convolutions to map the feature map to a higher dimension to obtain stronger fitting ability. Each group of convolution consists of a different number of depthwise separable convolution + HSwish activation function in series. Use feature concatenation to achieve the fusion of features with different receptive fields, and be able to more accurately align and extract the key features in the image.

4. The method for accelerating the generation of multi-temporal sensitive individual trajectories based on satellite video data according to claim 2, characterized in that, In S3, the specific steps for performing initial matching of time-sensitive individuals among multiple nodes are as follows: c11. The master node divides the time-sensitive individual extraction results, the trajectories of time-sensitive individuals in the tracking state, and the appearance features extracted by the SpeedNet detector into multiple tasks. Leave one task at the master node and allocate the remaining tasks to each slave node for execution; c12. The master node and the slave nodes respectively calculate the Mahalanobis distance cost matrix and the Cosine distance cost matrix on their respective nodes; The calculation formula for the Mahalanobis distance cost matrix is: M(m,n) = (P m -Q n ) T H -1 (P m -Q n ) Among them, P m is the extraction result matrix of the m-th time-sensitive individual; Q n is the trajectory matrix of the time-sensitive individual in the n-th tracking state; H is the covariance matrix of the time-sensitive individual predicted by the adaptive acceleration Kalman filter; The calculation formula for the Cosine distance cost matrix is: Among them, r n represents the appearance feature matrix of the nth detection box; represents the appearance feature of the mth trajectory extracted by the SpeedNet detector; c13. The master node receives the results of the Mahalanobis distance cost matrix and the Cosine distance cost matrix from the slave nodes and completes the aggregation of the two cost matrices, and reconstructs the minimum associated cost distance matrix MinCostMatching. The specific calculation formula is: MinCostMatching = glg g DisMah+(1 - g)lg 1-g DisCos Where DisMah is the overall Mahalanobis distance cost matrix; DisCos is the overall Cosine distance cost matrix; g represents the weight coefficient; c14. Screen the minimum associated cost distance matrix MinCostMatching according to the maximum threshold MaxDistance, assign infinity to the elements in the matrix that exceed the threshold, and use the Hungarian matching algorithm to complete the initial matching of time-sensitive individuals among multiple nodes for MinCostMatching.

5. The method for accelerating the generation of multi-temporal sensitive individual trajectories based on satellite video data according to claim 1, characterized in that In S3, the specific steps for performing EfficientIou time-sensitive individual re-matching are as follows: c21. Calculate the EfficientIou time-sensitive individual re-matching cost matrix EffiouCost; In c22, the Hungarian matching algorithm is used to perform optimal matching on the cost matrix EffiouCost.

6. The multi-temporal sensitive individual trajectory generation acceleration method based on satellite video data according to claim 5, wherein, In the step c21, the formula for calculating the re-matching cost matrix EffiouCost of the EfficientIou time-sensitive individuals is: Among them, EffiouCost m,n represents the intersection over union of the extraction result of the m-th time-sensitive target in the current frame of satellite video data and the target that fails the initial matching of the n-th time-sensitive individual of multiple nodes or the trajectory in the non-tracking state; det m represents the extraction result of the m-th time-sensitive target in the current frame of satellite video data; represents the target that fails the initial matching of the n-th time-sensitive individual of multiple nodes or the trajectory in the non-tracking state; k c and g c respectively represent the width and height of the minimum bounding rectangle of the two bounding boxes of the extraction result of the current frame of satellite video data and the trajectory of the time-sensitive individual of multiple nodes that fails the initial matching or is in the non-tracking state; k m and respectively represent the width of the extraction result of the m-th current frame of satellite video data and the width of the target that fails the initial matching of the n-th time-sensitive individual of multiple nodes or the non-tracking state; g m and respectively represent the height of the extraction result of the m-th current frame of satellite video data and the height of the target that fails the initial matching of the n-th time-sensitive individual of multiple nodes or the non-tracking state; Φ 2 represents the Euclidean distance.

Citation Information

Patent Citations

  • Multi-target tracking positioning and motion state estimation method based on unmanned aerial vehicle

    CN113269098A

  • Multi-target tracking method for intelligent driving

    CN116402850A