Elliptical multi-extension target tracking method and system based on dynamic heterogeneous graph neural network
Through the dynamic heterogeneous graph neural network modeling multi-scaling target tracking process, the problems of large amount of computing and insufficient dynamic change capabilities in traditional methods are solved, and flexible modeling and efficient tracking of multi-scaling targets in radar systems are achieved, which is suitable for a variety of sensor platforms.
Patent Information
- Application Number
- CN202510346946.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing multi-scaling target tracking methods have problems in radar systems such as model dependence, large computing volume, and insufficient ability to handle dynamic changes. Especially in multi-scaling target tracking, it is difficult to effectively utilize the characteristics of radar data and task structure.
Using a method based on dynamic heterogeneous graph neural network, the multi-scaling target tracking process is modeled as discrete time dynamic heterogeneous graphs, end-to-end calculation is performed through graph neural networks, and the multi-head attention mechanism and self-attention timing aggregation module are used to jointly optimize the object detection, data association and state filtering tasks, and design multi-task loss function to guide the global optimization of the model.
It realizes flexible modeling and efficient tracking of multiple expansion targets, improves tracking performance, and is suitable for a variety of sensor platforms such as ground surveillance radar, air early warning and sea surface search, and has good versatility and real-timeness.
Smart Images

Figure CN120298463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of radar target tracking, and in particular, to an elliptical multi-extended target tracking method and system based on a dynamic heterogeneous graph neural network. Background Art
[0002] Target tracking is one of the key tasks of modern radar systems, aiming to estimate the states (such as positions, velocities, etc.) of an unknown number of targets in real time based on the measurement data with noise and false alarms obtained by a radar receiver, and track their continuous motion trajectories. Traditional multi-target tracking methods are mainly based on Bayesian filtering theory, such as probability hypothesis density (PHD) filtering, multiple hypothesis tracking (MHT), etc. These methods model the target state as a random finite set (RFS), and propagate the posterior probability density through recursive Bayesian estimation to characterize the random appearance and disappearance of targets.
[0003] With the rapid development of modern sensing technologies, another more challenging task is multiple extended target tracking (METT). Different from idealized point targets, high-resolution radars can obtain the fine structures and scattering characteristics of targets. For targets with large volumes and complex structures, which usually cover multiple resolution cells, one target will generate multiple radar measurement information. In addition to estimating their motion parameters, it is also necessary to estimate the shape appearance (size, orientation, etc.). METT is an important research direction in the field of radar target tracking, aiming to solve the problems of detection, association, and state estimation when multiple extended targets with geometric shapes exist simultaneously in a real environment. For the problem of multiple extended target tracking, researchers have proposed various solutions based on different ideas. Most traditional methods are built on the RFS framework and introduce shape parameters to characterize the characteristics of extended targets. Et al. proposed the Gaussian inverse Wishart PHD (GIW-PHD) filter, which uses the Gaussian inverse Wishart distribution to model the kinematics and shape expansion of extended targets. Beard et al. proposed the labeled GIW-PHD filter, which can output the trajectories of extended targets. And the GLMB filter was extended to elliptical extended target tracking, modeling the joint distribution of target existence probability, motion state, and shape parameters. Xia et al. proposed a novel variant of the PMBM filter, which models the scattering points of the target as a random finite set and can flexibly adapt to extended targets of different shapes. Although these extended target tracking methods based on the RFS theory have solved the representation and estimation problems of multiple extended targets to a certain extent, they still inevitably face limitations such as high state dimensions, large computational complexity, and model dependence. Another approach is to regard the extended target as a set of interrelated scattering points. For example, the method proposed by Daniyan et al. based on the hierarchical Bayesian model first clusters the scattering points and then uses elliptical curve fitting to obtain the shape of the extended target. These methods simplify the problem to a certain extent but do not make full use of the spatial correlation of the scattering points.
[0004] In recent years, deep learning methods have achieved great success in the field of computer vision target tracking, which inspires researchers to explore their introduction into the end-to-end solution of multi-target tracking. Convolutional neural networks, recurrent neural networks, etc. are used to detect and track targets end-to-end from image sequences. Further, some scholars have proposed combining deep learning with random finite set filtering to learn the density function representing the multi-target state. Although these works demonstrate the potential of the data-driven paradigm, they mainly focus on visual tracking and are difficult to be directly applied to the extended target tracking problem. In addition, most of the existing deep tracking architectures adopt a "black box" design, lacking interpretability and consideration of the characteristics of radar data and task structure, resulting in insufficient generalization ability.
[0005] Generally speaking, multi-extended target tracking has long mainly relied on statistical modeling and Bayesian inference, suffering from inherent limitations such as model dependence, high computational cost, and insufficient ability to handle dynamic changes. Deep learning provides a new solution idea for this problem with its powerful feature extraction and function fitting capabilities, but there is still a lack of targeted and interpretable end-to-end methods. Summary of the Invention
[0006] To solve the above technical problems existing in the prior art, the present invention proposes an elliptical multi-extended target tracking method and system based on a dynamic heterogeneous graph neural network, which is a new end-to-end multi-extended target tracking paradigm based on a dynamic heterogeneous graph and integrating task structure and data representation.
[0007] On the one hand, to achieve the above object, the present invention provides an elliptical multi-extended target tracking method based on a dynamic heterogeneous graph neural network, including:
[0008] Obtain the information of the target to be tracked;
[0009] Construct a multi-extended target tracking model;
[0010] Input the information of the target to be tracked into the multi-extended target tracking model for processing to obtain the tracking result;
[0011] Among them, the multi-extended target tracking model is used to model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and use a graph neural network to perform end-to-end calculations on the discrete-time dynamic heterogeneous graph to guide the global optimization of the model.
[0012] Preferably, constructing the multi-extended target tracking model includes:
[0013] Model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and construct a heterogeneous graph slice containing measurement nodes and target nodes at each time step, where the heterogeneous graph slice includes measurement-measurement edges, measurement-target edges, and target-target edges;
[0014] Encode the heterogeneous graph slice at each moment, aggregate the embedding representations of different types of nodes through a multi-head attention mechanism, and generate updated node embeddings;
[0015] Perform temporal modeling on the embedding sequence of the target nodes, and generate target node representations containing long-term spatio-temporal dependencies in combination with the historical state memory matrix;
[0016] Jointly decode the measurement-target association probability matrix, target number prediction, and target state estimation through an association prediction decoding module, where the association prediction decoding module includes an association calculation layer, a target number prediction layer, and a target state prediction layer;
[0017] Design a multi-task loss function, jointly optimize the data association, target number estimation, and state filtering tasks, guide the global optimization of the model, and obtain the multi-extended target tracking model.
[0018] Preferably, the weight of the measurement-measurement edge is calculated by a Gaussian kernel function, which is used to reflect the spatial proximity of the measurement nodes; the weight of the measurement-target edge is calculated by a normalized Mahalanobis distance, which is used to measure the association probability between the measurement and the target nodes; the weight of the target-target edge is calculated by a normalized Gaussian Wasserstein distance, which is used to measure the similarity of the target nodes in terms of position and shape.
[0019] Preferably, encoding the heterogeneous graph slice at each moment includes:
[0020] Perform type-specific linear transformations on the measurement nodes and target nodes respectively through the heterogeneous graph attention network, and aggregate the embedded representations of heterogeneous neighbor nodes through the multi-head attention mechanism;
[0021] Among them, the multi-head attention mechanism uses edge type-specific attention vectors to calculate the attention coefficients between nodes, and generates the updated node embeddings through weighted aggregation.
[0022] Preferably, the processing process of the multi-head attention mechanism is as follows:
[0023]
[0024] In the formula, is the embedding of the i-th node in the l-th layer, ‖ represents concatenation, is the set of edge types, is the set of neighbors of node i under the r-th type of edge, is the attention weight, is the node type φ j is the transformation matrix, K is the number of attention heads, is the embedding of the j-th neighbor node in the (l - 1)-th layer.
[0025] Preferably, the temporal modeling of the embedded sequence of the target nodes includes:
[0026] Use the self-attention temporal aggregation module to perform temporal modeling on the embedded sequence of the target nodes. Among them, the self-attention temporal aggregation module is used to add sine position encoding to the target node embedding to generate a position encoding vector, interact with the historical state memory matrix through the dot product attention mechanism for the position encoding vector, dynamically fuse the historical temporal information, and generate the target node representation containing long-term spatio-temporal dependencies; the historical state memory matrix is used to store the historical embeddings of the target nodes and is updated through the sliding window mechanism to limit the computational complexity.
[0027] Preferably, the association calculation layer is used to map the embeddings of the measurement nodes and target nodes to a common space through linear transformation, and generate the measurement-target association probability matrix through matrix multiplication and the sigmoid function;
[0028] The target quantity prediction layer is used to predict the total number of targets at the next moment according to the current measurement information and historical target information;
[0029] The target state prediction layer is used to map the predicted state vector through a multi-layer perceptron by combining the association probability weighted measurement information, target node embedding and temporal aggregation embedding.
[0030] Preferably, the multi-task loss function includes binary cross-entropy loss, categorical cross-entropy loss, mean absolute error loss, and Gaussian Wasserstein distance loss;
[0031] Among them, the calculation steps of the Gaussian Wasserstein distance loss are as follows:
[0032] Convert the discrete measurement point set into a probability distribution through the Dirac measure;
[0033] For each discrete point, calculate the Wasserstein distance between the Gaussian distribution with the discrete point as the mean and the identity matrix as the variance and the target elliptical Gaussian distribution;
[0034] Take the average of the Wasserstein distances of all points as the metric between the point set and the elliptical distribution:
[0035] Based on the Gaussian Wasserstein distance formula, calculate the distribution difference between each measurement point and the target ellipse, and take the average as the final loss term.
[0036] On the other hand, to achieve the above object, the present invention also provides an elliptical multi-extended target tracking system based on a dynamic heterogeneous graph neural network, including:
[0037] Information acquisition unit: used to acquire the information of the target to be tracked;
[0038] Model construction unit: used to construct a multi-extended target tracking model;
[0039] Tracking processing unit: used to input the information of the target to be tracked into the multi-extended target tracking model for processing to obtain a tracking result; among them, the multi-extended target tracking model is used to model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and guide the global optimization of the model by designing an end-to-end multi-task loss function.
[0040] Preferably, the model construction unit includes:
[0041] Dynamic heterogeneous graph modeling module: used to construct a heterogeneous graph slice containing measurement nodes and target nodes in real time, where the heterogeneous graph slice includes measurement-measurement edges, measurement-target edges, and target-target edges;
[0042] Heterogeneous graph encoding module: used to encode the heterogeneous graph slice at each moment, aggregate the embedding representations of different types of nodes through the multi-head attention mechanism, and generate updated node embeddings;
[0043] Temporal aggregation module: used to perform temporal modeling on the embedding sequence of the target nodes, and generate a target node representation containing long-term spatio-temporal dependencies in combination with the historical state memory matrix;
[0044] Association prediction decoding module: used to jointly output measurement-target association probability, target quantity, and state estimation;
[0045] Multi-task optimization module: used to design a multi-task loss function, jointly optimize data association, target quantity estimation, and state filtering tasks, guide the global optimization of the model, and obtain the multi-extended target tracking model.
[0046] Compared with the prior art, the present invention has the following advantages and technical effects:
[0047] (1) The present invention first proposes a dynamic heterogeneous graph representation for modeling the multi-extended target tracking process, and adaptively depicts the complex evolution of heterogeneous information such as measurements, extended target states, and associations in time and space through the graph structure, overcoming the limitations of traditional vectorized representations;
[0048] (2) The present invention first applies the graph attention neural network to the field of target tracking, realizes end-to-end multi-target tracking, jointly optimizes links such as target detection, data association, and state filtering, fully explores the coupling dependencies between different tasks, and improves the tracking performance. In particular, the heterogeneous graph encoding uses type-specific attention aggregation to flexibly model the interaction patterns of measurements and targets under different semantics. The self-attention temporal aggregation mechanism enhanced by external memory enables the current target state estimation to focus on historical state information, making up for the deficiencies of existing depth tracking methods based on time sequential processing;
[0049] (3) The present invention provides a brand-new technical idea for the modeling, solution, and evaluation of multi-extended target tracking problems, and is expected to promote the theoretical innovation and engineering practice in the field of radar target tracking. The developed end-to-end tracking algorithm is universal and is not only applicable to ground surveillance radars, but can also be extended to broader sensor platforms and application scenarios such as airborne early warning and sea surface search. Description of the Drawings
[0050] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0051] Figure 1 Schematic diagram of the dynamic heterogeneous graph modeling for multi-extended target tracking in the embodiment of the present invention;
[0052] Figure 2 Schematic diagram of the graph encoding module of the multi-extended target tracking model in the embodiment of the present invention;
[0053] Figure 3 Schematic diagram of the temporal aggregation module of the multi-extended target tracking model in the embodiment of the present invention;
[0054] Figure 4Schematic diagram of the association prediction decoding module of the multi-extended target tracking model according to the embodiment of the present invention;
[0055] Figure 5 Experimental comparison chart of the GWD index for the matching metric between the measurement point set and the elliptical target in an elliptical multi-extended target tracking algorithm according to the embodiment of the present invention. Among them, (a) is a schematic diagram of the point set distributed outside the elliptical target with 100 measurement points, (b) is a schematic diagram of the point set distributed both inside and outside the elliptical target with 100 measurement points, (c) is a schematic diagram when the point set is only inside the elliptical target with 100 measurement points, (d) is a schematic diagram of the point set distributed inside the elliptical target with a quantity of 1000, (e) is the GWD index of the measurement point set with a quantity of 100 under three distribution methods, and (f) is a schematic diagram of the numerical change of the index when the point set is distributed inside the elliptical target and the number of measurement points is continuously increased;
[0056] Figure 6 Comparison chart of the actual tracking results and evaluation indexes of the multi-extended target tracking method according to the embodiment of the present invention. Among them, (a) is a schematic diagram of the trajectory, (b) is a schematic diagram of the GWD index, and (c) is a schematic diagram of the IoU index;
[0057] Figure 7 Flowchart of an elliptical multi-extended target tracking method based on a dynamic heterogeneous graph neural network according to the embodiment of the present invention. Detailed implementation manners
[0058] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0059] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0060] This embodiment proposes an elliptical multi-extended target tracking method based on a dynamic heterogeneous graph neural network, as Figure 7 , including:
[0061] Obtain the information of the target to be tracked;
[0062] Construct a multi-extended target tracking model;
[0063] Input the information of the target to be tracked into the multi-extended target tracking model for processing to obtain the tracking result;
[0064] Among them, the multi-extended target tracking model is used to model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and use a graph neural network to perform end-to-end calculations on the discrete-time dynamic heterogeneous graph to guide the global optimization of the model.
[0065] Specifically, in this embodiment, a tracking scenario including M extended targets is considered. The kinematic state and shape parameters of the j-th target can be represented by a state vector as follows:
[0066]
[0067] where and represent the position coordinates of target j at time t, and are the corresponding velocity components, is the direction angle of the target, and and are the major axis and minor axis of the elliptical shape respectively.
[0068] At each time t, the sensor returns a set of measurement values where represents the position coordinates of the i-th measurement, and n (t) is the total number of measurements at time t. The measurement set usually contains true measurements generated by M real targets (each target may generate multiple measurements) and false measurements (i.e., false alarms) generated by environmental clutter, but they cannot be directly distinguished. Multi-extended target tracking aims to estimate the state of each target at each time according to the measurement sets at a series of time steps and establish the correspondence between measurements and targets. Among them, is the total number of estimated targets at time t,
[0069] and
[0070]
[0071]
[0072] Model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and construct a heterogeneous graph slice containing measurement nodes and target nodes in each time step, where the heterogeneous graph slice includes measurement-measurement edges, measurement-target edges, and target-target edges; Encode the heterogeneous graph slice at each time, and aggregate the embedding representations of different types of nodes through a multi-head attention mechanism to generate updated node embeddings;Perform temporal modeling on the embedding sequence of the target nodes, and generate the target node representation containing long-term spatio-temporal dependencies by combining with the historical state memory matrix;
[0073] Jointly decode the measurement-target association probability matrix, target number prediction, and target state estimation through the association prediction decoding module, where the association prediction decoding module includes an association calculation layer, a target number prediction layer, and a target state prediction layer;
[0074] Design a multi-task loss function to jointly optimize the data association, target number estimation, and state filtering tasks, guide the global optimization of the model, and obtain a multi-extended target tracking model.
[0075] Specifically, as Figure 1 , the entire tracking process is modeled as a discrete-time dynamic heterogeneous graph where T is the total number of tracking time steps. For each time step t, the dynamic graph contains a heterogeneous subgraph called a graph slice, representing the tracking state at the current moment. Each graph slice consists of two types of nodes: measurement nodes and state nodes corresponding to the sensor measurements and estimated target states at the current moment, respectively; at the same time, it contains three types of edges: measurement-measurement edges, measurement-state edges, and state-state edges, which respectively depict the interaction relationships between different nodes.
[0076] For the graph slice its node set is containing two types of heterogeneous nodes, defined as follows:
[0077] Measurement node set where, represents the i-th measurement node, and n (t) is the total number of measurements at time step t.
[0078] Each measurement node carries a feature vector initialized to the corresponding measurement position coordinates:
[0079]
[0080] Target node set where represents the state estimate of the j-th target node at time step t, and m (t) is the total number of estimated targets at time step t. Each target node carries a feature vector initialized to the corresponding target state estimate value:
[0081]
[0082] In the formula, and respectively represent the estimated position coordinates of target j at time t. and are respectively the estimated velocity component values in the x-axis direction and the y-axis direction. are respectively the direction angle, major semi-axis, and minor semi-axis of the elliptical target. is the state estimate of the j-th target at time t.
[0083] Figure slice contains three types of semantic edges, which connect different subsets of nodes respectively:
[0084] Measurement-measurement edge set Among them, represents an undirected edge between measurement node and measurement node .
[0085] Weight of the measurement-measurement edge is defined as a Gaussian kernel function:
[0086]
[0087] Among them, ‖·‖2 represents the Euclidean distance. The physical meaning of this weight design is that the closer two measurements are in space, the greater the probability that they come from the same target.
[0088] Measurement-target edge set Among them; represents an undirected edge between measurement node and target node .
[0089] Weight of the measurement-target edge is defined as the normalized Mahalanobis distance:
[0090]
[0091] Among them, the Mahalanobis distance is defined as:
[0092]
[0093] Among them, is the estimated position covariance matrix of the j-th target at time t, which can be calculated according to its shape parameter estimate as follows:
[0094]
[0095] Among them, Denote a two-dimensional rotation matrix:
[0096]
[0097] Mahalanobis distance is a weighted Euclidean distance that takes into account the uncertainty of the target. When the position of the measurement is closer to the position component of the target state estimate in the Mahalanobis sense, the edge weight between them is larger, indicating a higher degree of association between the two.
[0098] Target-target edge set where denotes the target node and the undirected edge between them.
[0099] Weight of the target-target edge is defined as the normalized Gaussian Wasserstein Distance (GWD):
[0100]
[0101] where the definition of GWD(·,·) is:
[0102]
[0103] where s1 and s2 are two target state vectors, s1[[x,y]] and s2[[x,y]] represent their position components, and Σ1 and Σ2 are the corresponding position covariance matrices.
[0104] The Gaussian Wasserstein Distance comprehensively considers the differences in position and shape between two targets and is a robust similarity measure. Intuitively, the closer the spatial positions and the more similar the shapes of two targets, the stronger the interaction between them.
[0105] In summary, the edge set of the slice can be expressed as:
[0106]
[0107] Correspondingly, three adjacency matrices and are defined to represent the weights of different types of edges:
[0108]
[0109] Here the graph slice is a fully connected graph. Therefore, the adjacency matrices and are n (t) ×n(t) , n (t) ×m (t) and m (t) ×m (t) dense matrix of
[0110] Furthermore, encoding each time slice of the heterogeneous graph includes:
[0111] Performing type-specific linear transformations on measurement nodes and target nodes respectively through a heterogeneous graph attention network, and aggregating the embedded representations of heterogeneous neighbor nodes through a multi-head attention mechanism;
[0112] Among them, the multi-head attention mechanism uses an edge type-specific attention vector to calculate the attention coefficients between nodes, and generates updated node embeddings through weighted aggregation.
[0113] Specifically, in this embodiment, a heterogeneous graph attention network (HGAT) is designed as the main tool for graph encoding. Compared with traditional graph neural networks, the advantage of HGAT is that it can distinguish different types of nodes and edges, model their feature evolution and interaction patterns respectively, so as to better adapt to the complex data structure in the multi-target tracking scenario.
[0114] Construction of the heterogeneous graph:
[0115] Define the graph The node set of is Among them, the measurement node set Each node corresponds to a measurement vector The target node set Each node corresponds to a target state estimate m t is the total number of targets estimated at time t-1.
[0116] Define the graph The edge set of is Among them, is the measurement-measurement edge, connecting all pairs of measurement nodes, and the edge weight is the Gaussian kernel function, which measures the geometric similarity between nodes; is the measurement-target edge, connecting all measurement-target node pairs, and the edge weight is the normalized Mahalanobis distance, which measures the association strength between nodes; is the target-target edge, connecting all pairs of target nodes, and the edge weight is the normalized Gaussian Wasserstein distance, which measures the interaction pattern between nodes.
[0117] Heterogeneous graph attention layer:
[0118] The graph encoding module adopts stacked L-layer Heterogeneous Graph Multi-head Attention Layers (HGMALs) to learn node embeddings. The input of the l-th layer HGMAL is the node embedding matrix of the previous layer and the output is the updated node embedding matrix The initial embedding matrix H (t,0) is the original feature of the node
[0119] The calculation process of HGMAL is as follows
[0120] Type-specific transformation: For node i of different types φ ∈ {z, s}, use the type-specific transformation matrix to map its embedding to a K-dimensional feature space with dimension d l :
[0121]
[0122] Heterogeneous graph attention coefficients: For each type of edge r ∈ {zz, zs, ss}, use the type-specific attention vector to calculate the attention coefficient between nodes i and j, and then normalize it through softmax
[0123]
[0124] where is the set of neighbor nodes of node i under the r-th type of edge, is the l-th layer attention score of nodes i and j with edge type r in the feature space k, is the feature node j participating in the attention score calculation in the l-th layer in the feature space k at time t, is all feature nodes connected by the r-type edge at time t, is the attention score between node i and its neighbor node n in the l-th layer in the feature space k with edge type r;
[0125] Heterogeneous neighbor aggregation: Node i aggregates the neighbor information of all K heads and all r-type edges through the attention weight to obtain the updated node representation
[0126]
[0127] where ‖ represents vector concatenation
[0128] In summary, the core of the graph encoding module is the Heterogeneous Graph Multi-head Attention Layer (HGMAL), which fuses local structure and semantic information by distinguishing the interaction patterns of different types of nodes and edges. As Figure 2 。
[0129] The calculation of HGMAL can be summarized as follows:
[0130]
[0131] Among them, is the embedding of the i-th node in the l-th layer, ‖ represents concatenation, is the set of edge types, is the set of neighbors of node i under the r-th type of edge, is the attention weight, is the transformation matrix of node type φ j K is the number of attention heads, is the embedding of the j-th neighbor node in the (l - 1)-th layer.
[0132] After stacking L layers of HGMAL, the final node embedding matrix can be obtained:
[0133]
[0134] Among them, is the measurement node embedding sub-matrix, is the target node embedding sub-matrix, H (t ,L) is the node embedding matrix of the L-th layer at time t. These embedding vectors fuse the structural information of the graph and the heterogeneous interaction patterns, and the output of the graph encoding will be fed into the subsequent temporal aggregation and association prediction decoding modules for subsequent temporal modeling and tracking tasks.
[0135] Furthermore, the temporal modeling of the embedding sequence of the target node includes:
[0136] Using the self-attention temporal aggregation module to perform temporal modeling on the embedding sequence of the target node, where the self-attention temporal aggregation module is used to add sine position encoding to the target node embedding to generate a position encoding vector, and interact with the historical state memory matrix through the dot product attention mechanism to dynamically fuse historical temporal information to generate the target node representation containing long-term spatio-temporal dependencies; the historical state memory matrix is used to store the historical embeddings of the target node and is updated through a sliding window mechanism to limit the computational complexity.
[0137] Specifically, as Figure 3, in the multi-extended target tracking task, there are often complex dependency relationships and evolution patterns in the target state in the time dimension. To better model and utilize this temporal information, based on the graph encoding module, the self-attention temporal aggregation module dynamically aggregates the target state node embeddings at different times through the self-attention mechanism to generate node representations containing historical context information.
[0138] The self-attention temporal aggregation module mainly consists of three parts: positional encoding, memory matrix update, and self-attention aggregation.
[0139] Positional encoding:
[0140] To introduce the time step information into the node embedding, first use the Sinusoidal Positional Encoding (SPE) function SPE(t, m, d) to generate a unique d-dimensional positional encoding vector for the m-th target node at time t:
[0141] SPE(t, m, 2i) = sin(t m / 10000 2i / d );
[0142] SPE(t, m, 2i + 1) = cos(t m / 10000 2i / d );
[0143] where t m = t + m / M t is the timestamp of node m, and i is the dimension index. Add the positional encoding vector pointwise to the node embedding to obtain the node representation matrix that fuses the temporal information
[0144]
[0145] Memory matrix update:
[0146] Before performing attention aggregation, it is necessary to append the node embedding at the current moment to the memory matrix to construct the complete temporal context:
[0147]
[0148] where [·; ·] represents the concatenation operation in the first dimension (time dimension). Here, the historical node embeddings in X (t-1) are not recalculated, but the previous results are directly reused to improve the calculation efficiency.
[0149] Self-attention aggregation:
[0150] The goal of self-attention aggregation is to enable each node at the current moment to adaptively focus on the node information related to it in history, forming a context-aware node representation.
[0151] Specifically, for the m-th node at the t-th moment, the query vector the key vector and the value vector are calculated as follows:
[0152]
[0153] where are all learnable weight matrices.
[0154] Then, node m aggregates information from historical nodes through the attention mechanism:
[0155]
[0156]
[0157] where and are the key sequence and value sequence of the memory matrix X (t) respectively, and is the attention weight between node m and historical node n. Finally, the aggregated embedding of the m-th node
[0158] Concatenating the aggregated embeddings of all M t nodes, we obtain the target state node embedding matrix that fuses historical information at the t-th moment:
[0159]
[0160] In the formula, is the aggregated embedding of the M-th t node at the t-th moment.
[0161] After the above calculations, the output contains the spatio-temporal dependence relationship between the current moment and historical moments, which can better guide subsequent association and prediction tasks. At the same time, this calculation process is differentiable, allowing end-to-end gradient propagation and optimization.
[0162] Assume that at the t-th moment, the average number of targets per moment is Then the time complexity of the memory matrix X (t) is and the space complexity is The time complexity of self-attention calculation is and the space complexity is The output matrix The time and space complexity are both O(M t d v ).
[0163] Therefore, the overall time complexity is The space complexity is Although the computational cost increases linearly with the time step t, through parallelization and optimized coding implementation, the real-time requirements can still be met. In addition, d q , d k , d v is generally much smaller than d, and the number of heads can also be controlled within a small constant range. Therefore, the overall complexity is acceptable.
[0164] Furthermore, the association calculation layer is used to map the measurement nodes and target nodes into a common space through linear transformation, and generate the measurement-target association probability matrix through matrix multiplication and the sigmoid function;
[0165] The target number prediction layer is used to predict the total number of targets at the next moment according to the current measurement information and historical target information;
[0166] The target state prediction layer is used to map the predicted state vector through a multi-layer perceptron by combining the association probability weighted measurement information, target node embedding and temporal aggregation embedding.
[0167] Specifically, as Figure 4 , the association prediction and decoding module (Association Prediction and Decoding Module, APDM) includes three sub-layers: the association calculation layer (Association Computation Layer, ACL), the target number prediction layer (Target Number Prediction Layer, TNPL) and the target state prediction layer (Target State Prediction Layer, TSPL).
[0168] Association Computation Layer (ACL):
[0169] The purpose of the association calculation layer is to calculate the association probability between each pair of measurement-target nodes at the current moment, and obtain a two-dimensional association probability matrix Intuitively, the (i, j)-th item of the matrix represents the possibility that the i-th measurement is generated by the j-th target. This information is crucial for subsequent target state updates, because only by correctly associating the measurement with the target can the measurement information be used to correct and improve the estimation of the target state.
[0170] ACL embeds the matrix of measurement nodes output by the graph encoding module and the target node embedding matrix As inputs, they are first mapped to a common d a -dimensional space through a linear transformation, and then the association probability is calculated through matrix multiplication and the sigmoid function:
[0171]
[0172] where and are the learnable mapping parameters of the measurement and target embedding respectively. The sigmoid function compresses the original association score to the interval (0, 1) to obtain a normalized probability value.
[0173] This calculation process can be regarded as using the measurement node to perform attention weighting on the target node, and the association probability matrix is the attention weight. Different from the prior association method such as the global nearest neighbor (GNN), the association matrix learned by ACL is data-driven, can model complex association patterns, and has stronger adaptability.
[0174] Target Number Prediction Layer (TNPL):
[0175] The task of the Target Number Prediction Layer is to predict the total number of targets at the next moment based on the current measurement information and historical target information Accurately estimating the number of targets is crucial for multi-target tracking, because it determines how many target states the tracker needs to maintain and how to handle newly emerging and disappearing targets.
[0176] TNPL adopts a classification idea to predict the number of targets. Specifically, in this embodiment, a prior maximum number of targets M max is first set, and the possible number of targets is discretized into M max +1 categories (including 0). Then, the measurement node embedding at time t the target node embedding and the temporally aggregated target node embedding are averaged in the node dimension, concatenated into a 3d-dimensional feature vector, and the probability distribution of each number category is calculated through a multi-layer perceptron (MLP) and the softmax function:
[0177]
[0178] where ColMean(·) represents column mean pooling of the input matrix, || represents vector concatenation, and MLP M (·) represents a multi-layer perceptron, and the learnable parameters are and are the parameters of the number classification head, is the probability distribution of the target quantity category at time t, is the probability of the target quantity at time t+1.
[0179] Finally, select the category with the highest probability as the predicted target quantity:
[0180]
[0181] The input of TNPL includes not only the local features extracted by the graph encoding module and also incorporates the global context information learned by the temporal aggregation module This enables the model to not only consider the observational evidence at the current moment but also refer to the information of the entire tracking history when predicting the target quantity, having a certain degree of foresight and stability. At the same time, transforming the continuous quantity estimation problem into a classification problem also reduces the learning difficulty of the model.
[0182] Target State Prediction Layer (TSPL):
[0183] The purpose of the Target State Prediction Layer is to estimate the state vector of each target at the next moment where Compared with classical tracking methods, the characteristic of TSPL is that it does not predict each target in isolation but fully utilizes the target-measurement association and target-target interaction learned by the graph encoding module, as well as the cross-time dependencies captured by the temporal aggregation module, comprehensively considering the all-to-all spatio-temporal context, thereby obtaining more accurate and consistent prediction results.
[0184] Specifically, for the j-th target, TSPL first performs a weighted sum on the measurement node embedding matrix according to the j-th column of the association probability matrix to obtain the aggregated measurement information of this target
[0185]
[0186] Then, is concatenated with the graph embedding of target node j and the temporal aggregation embedding and mapped to the state space through another multi-layer perceptron to obtain the predicted state vector
[0187]
[0188] where MLP s (·) represents another multi-layer perceptron with learnable parameters W s ,b s .
[0189] It can be seen that the prediction of the target state comprehensively utilizes three aspects of information: through weighted measurement information reflects the current observation evidence related to target j; the graph embedding of target j itself encodes its interaction pattern with other targets and measurements at time t; the temporal aggregation embedding of target j contains its state evolution and global dependencies at past times.
[0190] The convergence of these three information flows enables TSPL to achieve a balance between current observations and past trajectories, making more robust and coherent predictions. In addition, the multi-layer perceptron endows the model with the ability of non-linear modeling, enabling it to fit complex state transition functions.
[0191] Finally, the output of APDM includes: the association probability matrix P (t) 、the predicted number of targets and the state estimation of each target These results will be fed back to the graph encoding module at the next time step to construct a new heterogeneous graph, forming a cyclic tracking process.
[0192] Furthermore, the multi-task loss function includes binary cross-entropy loss, categorical cross-entropy loss, mean absolute error loss, and Gaussian Wasserstein distance loss;
[0193] Among them, the steps of the Gaussian Wasserstein distance loss are calculated as follows:
[0194] Convert the discrete point set into a probability distribution through the Dirac measure;
[0195] For each discrete point, calculate the Wasserstein distance between the Gaussian distribution with the discrete point as the mean and the identity matrix as the variance and the target elliptical Gaussian distribution;
[0196] Take the average of the Wasserstein distances of all points as the metric between the point set and the elliptical distribution:
[0197] Based on the Gaussian Wasserstein distance formula, calculate the distribution difference between each measurement point and the target ellipse, and take the average as the final loss term.
[0198] Specifically, the association loss:
[0199] For the association probability matrix output by the association calculation layer (ACL) This embodiment uses binary cross-entropy loss to measure its difference from the true association matrix :
[0200]
[0201] wherein, indicates that the i-th measurement is truly associated with the j-th target, indicating no association. This loss function encourages the model to learn a soft probability matrix consistent with the true associations.
[0202] Target quantity loss:
[0203] For the target quantity probability distribution output by the Target Number Prediction Layer (TNPL), the Cross-Entropy Loss is used to measure the difference between it and the one-hot encoding of the true target quantity
[0204]
[0205] wherein, is a one-hot vector, with only the element corresponding to the true target quantity being 1 and the rest being 0. This loss function prompts the model to learn to accurately predict the total number of targets at the next moment.
[0206] Target state loss:
[0207] For each target state estimation output by the Target State Prediction Layer (TSPL), the Mean Absolute Error (MAE) and the Gaussian Wasserstein Distance (GWD) are combined to measure the difference between it and the true target state
[0208]
[0209] wherein, is the MAE loss:
[0210]
[0211] is the GWD loss:
[0212]
[0213] wherein, respectively represent the position components of the predicted and true target states, respectively represent the position covariance matrices of the predicted and true target states, tr(·) represents the trace of the matrix, λ mae and λgwd The weight coefficients for balancing the two losses.
[0214] Multi-task loss:
[0215] By weighted summing the above three loss functions, the multi-task loss of the association prediction decoding module at time t is obtained:
[0216]
[0217] Among them, λ ac , λ num , λ state are the weight coefficients of each loss term, which can be adjusted according to the importance of the task and the numerical scale. During training, the losses over the entire time series are accumulated, and an L2 regularization term for all learnable parameters (denoted as Θ) is added to obtain the final optimization objective:
[0218]
[0219] Among them, λ Θ is the regularization coefficient, and T is the total number of time steps.
[0220] This embodiment also provides an elliptical multi-extended target tracking system based on a dynamic heterogeneous graph neural network, including:
[0221] Information acquisition unit: used to acquire the information of the target to be tracked;
[0222] Model construction unit: used to construct a multi-extended target tracking model;
[0223] Tracking processing unit: used to input the information of the target to be tracked into the multi-extended target tracking model for processing to obtain a tracking result; among them, the multi-extended target tracking model is used to model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and guide the global optimization of the model by designing an end-to-end multi-task loss function.
[0224] Specifically, the model construction unit includes:
[0225] Dynamic heterogeneous graph modeling module: used to construct a heterogeneous graph slice containing measurement nodes and target nodes in real time, where the heterogeneous graph slice includes measurement-measurement edges, measurement-target edges, and target-target edges;
[0226] Heterogeneous graph encoding module: used to encode the heterogeneous graph slice at each moment, and aggregate the embedding representations of different types of nodes through a multi-head attention mechanism to generate updated node embeddings;
[0227] Temporal Aggregation Module: It is used to perform temporal modeling on the embedding sequence of the target node, and generate the target node representation containing long-term spatio-temporal dependencies by combining the historical state memory matrix;
[0228] The temporal aggregation module also includes an external storage sub-module, which is used to dynamically maintain the embedding matrix of historical target nodes, and realize information interaction across time steps through the self-attention mechanism
[0229] Association Prediction Decoding Module: It is used to jointly output the measurement-target association probability, the number of targets and the state estimation;
[0230] Multi-Task Optimization Module: It is used to design a multi-task loss function, jointly optimize the data association, target number estimation and state filtering tasks, guide the global optimization of the model, and obtain a multi-extended target tracking model.
[0231] To more clearly express the technical solution of the present invention, specific embodiments are provided below for introducing the solution:
[0232] Input: Measurement sequence Initial state estimation S (0) .
[0233] Step 1: Initialize the memory matrix Initial hidden state
[0234] Step 2: for t = 1 to T;
[0235] Step 2.1: The graph encoding module processes the measurement sequence at time t and the state estimation S (t-1) , and outputs the target measurement node embedding at this moment and the state node embedding
[0236] Step 2.2: The temporal aggregation module processes the target state node embedding and the historical moment node embedding memory matrix X (t-1) , and outputs the t-time target state node embedding matrix that fuses historical information and updates the memory matrix X (t) .
[0237] Step 2.3: The association prediction decoding module processes and outputs the association probability matrix P (t) , the target number prediction and the state estimation S of each target (t+1) ., S (t+1) .
[0238] Step 3: Output the tracking result
[0239] Calculation process of the graph encoding module:
[0240] Input: Measurement set State estimation S (t-1) .
[0241] Step 1: Construct a heterogeneous graph
[0242] Step 1.1: Add measurement nodes and target nodes
[0243] Step 1.2: Add measurement-measurement edges Calculate the weight
[0244] Step 1.3: Add measurement-target edges Calculate the weight
[0245] Step 1.4: Add target-target edges Calculate the weight
[0246] Step 2: Initialize node features
[0247] Step 3: for l = 1 to L;
[0248] Step 3.1: Type-specific transformation Wherein, is the transformation matrix, and are the node embeddings before and after the transformation respectively, is the node set type, and K is the number of attention channels.
[0249] Step 3.2: Edge type-specific attention:
[0250] Wherein, is the attention score, is the attention weight, is the edge type, is the node type, and K is the number of attention channels.
[0251] Step 3.3: Heterogeneous neighbor aggregation Aggregate the node embeddings through the attention score to obtain is the aggregated embedding, is the node type.
[0252] Output: Measurement node embedding Target node embedding
[0253] Calculation process of the time series aggregation module:
[0254] Input: Target node embedding Memory matrix X (t-1) .
[0255] Step 1: Position encoding In the formula, is the embedding of m state nodes at time t, and SPE(t, m, d) represents the d-dimensional time position encoding of m state nodes at time t.
[0256] Step 2: Memory storage
[0257] Step 3: Self-attention aggregation
[0258] Step 3.1: W is a learnable weight matrix, K (t) and V (t) are the key sequence and value sequence of the memory matrix X (t) respectively.
[0259] Step 3.2: for m = 1 to m t ;
[0260] Step 3.2.1: In the formula, q is the query vector and W is the weight matrix.
[0261] Step 3.2.2:
[0262] Step 3.2.3: In the formula, represents the aggregated embedding of the m-th node.
[0263] Step 3.3: In the formula, is the target state node embedding matrix that fuses historical information at time t.
[0264] Output: Aggregated embedding Updated memory matrix X (t) .
[0265] Calculation process of the association prediction decoding module:
[0266] Input: Measurement node embedding Target node embedding Aggregated embedding
[0267] Step 1: Association calculation layer:
[0268] Step 1.1: W az and b az are learnable mapping parameters for measurement embedding.
[0269] Step 1.2: W as and b as are learnable mapping parameters for target embedding.
[0270] Step 1.3: P (t) is the association probability value
[0271] Step 2: Target quantity prediction layer:
[0272] Step 2.1: is the probability distribution of the target quantity category at time t.
[0273] Step 2.2: is the probability of the target quantity at time t+1, W pM and b pM are learnable parameters.
[0274] Step 2.3: is the predicted target quantity.
[0275] Step 3: Target state prediction layer:
[0276] Step 3.1: for j = 1 to m t ;
[0277] Step 3.1.1: represents the aggregated measurement information of the j-th target, is the j-th column of the association probability matrix.
[0278] Step 3.1.2:
[0279] Step 3.1.3:
[0280] where MLP(·) represents another multi-layer perceptron, || represents the concatenation operation, W s and b s are learnable parameters, is the intermediate result, is the predicted state vector.
[0281] Step 3.2:
[0282] Output: Associated probability matrix P (t) , target quantity estimation Target state estimation S (t+1) .
[0283] Model hyperparameter configuration:
[0284] Graph encoding module: Original feature dimension d of measurement nodes z = 2, original feature dimension d of target nodes s = 7, embedding dimension d = 256, number of attention heads K = 8, number of heterogeneous graph attention layers L = 3;
[0285] Temporal aggregation module: Position encoding dimension d en = 256, query / key / value matrix dimension d q , d k , d v = 64, 64, 256, maximum number of time steps T of the memory matrix mem = 100;
[0286] Associated prediction decoding module: Mapping matrix dimension d az = 256, d as = 256, maximum number of targets M max = 20, the MLP structure of target quantity prediction is 256→128→64→M max + 1, the MLP structure of target state prediction is 256→128→64→d s , the MLP activation function is ReLU;
[0287] Model training hyperparameters: The optimizer is Adam, the initial learning rate is 0.001, Batch size = 32, the number of training epochs is 1000, and the loss function weights λ ac = 1.0, λ num = 0.1, λ state = 1.0, regularization coefficient λ Θ = 0.0001, λ mae = 0.5, λ gwd = 0.5.
[0288] The embodiment of the present invention also provides a metric for the elliptical multi-extended target tracking algorithm. Please refer to Figure 5 , which is used for the matching metric between the measurement point set and the elliptical target.
[0289] Specifically:
[0290] Let the point set be The ellipse is E: The goal is to calculate the Wasserstein distance between a discrete point set \(P\) and a continuous Gaussian (elliptical) distribution \(N(\mu,\Sigma)\). The point set represents \(n\) points in a two-dimensional space, while \(N(\mu,\Sigma)\) represents a two-dimensional Gaussian distribution with mean \(\mu\) and covariance matrix \(\Sigma\), which geometrically corresponds to an ellipse.
[0291] First, represent the discrete point set \(P\) as a probability distribution using the Dirac measure:
[0292]
[0293] where represents the Dirac measure located at point \(x\) i . This representation evenly distributes the probability mass of each point over all points, with the probability mass of each point being \(1 / n\).
[0294] The Gaussian distribution of the ellipse is represented as:
[0295]
[0296] where \(|\Sigma|\) represents the determinant of the covariance matrix \(\Sigma\).
[0297] The definition of the Wasserstein distance between two Gaussian distributions is:
[0298]
[0299] where \(N_1\sim N(\mu_1,\Sigma_1)\) and \(N_2\sim N(\mu_2,\Sigma_2)\) are two Gaussian distributions, and \(tr(\cdot)\) represents the trace of a matrix.
[0300] Since the point set \(P\) is not a truly continuous distribution, the above formula cannot be directly used to calculate the Wasserstein distance between it and the Gaussian distribution \(N(\mu,\Sigma)\). So each point \(x\) i is regarded as a Gaussian distribution \(N(x\) i ,\(\epsilon I)\) with mean \(x\) i and covariance matrix \(\epsilon I\), where \(\epsilon\) is a very small positive number and \(I\) is the identity matrix. Then, the Wasserstein distance between each such Gaussian distribution and the target Gaussian distribution \(N(\mu,\Sigma)\) can be calculated, and their average value is taken as an approximation of the Wasserstein distance between the point set \(P\) and \(N(\mu,\Sigma)\):
[0301]
[0302] For each point \(x\) i , the above formula can be used to calculate
[0303]
[0304] Further simplification gives:
[0305]
[0306] Adding up the contributions of each point gives:
[0307]
[0308] When ∈ approaches zero, the final approximate expression is obtained:
[0309]
[0310] The simulation conditions of this embodiment are:
[0311] System: 64-bit Windows 11 operating system;
[0312] CPU: 12th Gen Intel(R) Core(TM) i5-12400 2.50GHz;
[0313] GPU: NVIDIA GeForce RTX 3090Ti;
[0314] Python: 3.7.4;
[0315] Pytorch: 1.6.0;
[0316] Torchvision: 0.7.0.
[0317] Figure 5 It is a Gaussian Wasserstein distance metric between a point set and an elliptical target provided by an embodiment of the present invention, used to measure the matching degree between the point set and the elliptical target. Figure 5 There are a total of six subgraphs, among which Figure 5 (a), (b), and (c) are respectively schematic diagrams of the point set outside the elliptical target, the point set distributed both inside and outside the elliptical target, and the point set only inside the elliptical target with 100 measurement points. Figure 5 (d) of Figure 5 is a schematic diagram of the point set distributed inside the elliptical target with a quantity of 1000. Figure 5 (e) of Figure 5 is the index value corresponding to (a), (b), and (c) of
[0318] Figure 6It is the tracking result of the model with a target number of 3 provided by the embodiments of the present invention. Figure 6 (a) of Figure 6 is a schematic diagram of the trajectory. Figure 6 (b) of Figure 6 is the GWD index. Figure 6 (c) of Figure 6 is the IoU index. The total tracking duration is 60 time steps. It can be seen that the tracking error decreases with time. Correspondingly, GWD converges, IoU increases and finally the change tends to be flat. It is not difficult to see from the results that the elliptical multi-extended target tracking algorithm based on the discrete-time dynamic heterogeneous graph neural network proposed in this embodiment has good tracking accuracy and can generate target trajectories in real time.
[0319] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An elliptical multi-extended target tracking method based on a dynamic heterogeneous graph neural network, characterized in that Including: Obtain the information of the target to be tracked; Construct a multi-extended target tracking model; Input the information of the target to be tracked into the multi-extended target tracking model for processing to obtain a tracking result; Among them, the multi-extended target tracking model is used to model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and use a graph neural network to perform end-to-end calculations on the discrete-time dynamic heterogeneous graph to guide the global optimization of the model.
2. The method for tracking elliptical multi-extended targets based on a dynamic heterogeneous graph neural network according to claim 1, wherein Constructing the multi-extended target tracking model includes: Model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and construct a heterogeneous graph slice containing measurement nodes and target nodes at each time step, where the heterogeneous graph slice includes measurement-measurement edges, measurement-target edges, and target-target edges; Encode the heterogeneous graph slice at each moment, aggregate the embedding representations of different types of nodes through a multi-head attention mechanism, and generate updated node embeddings; Perform temporal modeling on the embedding sequence of the target node, and combine the historical state memory matrix to generate a target node representation containing long-term spatio-temporal dependencies; Jointly decode the measurement-target association probability matrix, target number prediction, and target state estimation through an association prediction decoding module, where the association prediction decoding module includes an association calculation layer, a target number prediction layer, and a target state prediction layer; Design a multi-task loss function, jointly optimize the data association, target number estimation, and state filtering tasks, guide the global optimization of the model, and obtain the multi-extended target tracking model.
3. The method for tracking elliptical multi-extended targets based on a dynamic heterogeneous graph neural network according to claim 2, wherein The weight of the measurement-measurement edge is calculated by a Gaussian kernel function and is used to reflect the spatial proximity of the measurement nodes; the weight of the measurement-target edge is calculated by a normalized Mahalanobis distance and is used to measure the association probability between the measurement and the target node; the weight of the target-target edge is calculated by a normalized Gaussian Wasserstein distance and is used to measure the similarity of the target nodes in terms of position and shape.
4. The method for tracking elliptical multi-extended targets based on a dynamic heterogeneous graph neural network according to claim 2, wherein Encoding the heterogeneous graph slice at each moment includes: Perform type-specific linear transformations on the measurement nodes and target nodes respectively through a heterogeneous graph attention network, and aggregate the embedding representations of heterogeneous neighbor nodes through a multi-head attention mechanism; Among them, the multi-head attention mechanism uses an edge type-specific attention vector to calculate the attention coefficients between nodes, and generates updated node embeddings through weighted aggregation.
5. The elliptical multi-extended target tracking method based on a dynamic heterogeneous graph neural network according to claim 4, characterized in that The processing process of the multi-head attention mechanism is: wherein, is the embedding of the i-th node in the l-th layer, ‖ represents concatenation, is the set of edge types, is the set of neighbors of node i under the r-th type of edge, is the attention weight, is the node type φ j is the transformation matrix of, K is the number of attention heads, is the embedding of the j-th neighbor node in the (l-1)-th layer.
6. The method for tracking elliptical multi-extended targets based on a dynamic heterogeneous graph neural network according to claim 2, wherein Performing temporal modeling on the embedding sequence of the target node includes: Use a self-attention temporal aggregation module to perform temporal modeling on the embedding sequence of the target node, where the self-attention temporal aggregation module is used to add sinusoidal position encoding to the target node embedding to generate a position encoding vector, interact with the historical state memory matrix through a dot product attention mechanism for the position encoding vector, dynamically fuse historical temporal information, and generate the target node representation containing long-term spatio-temporal dependencies; the historical state memory matrix is used to store the historical embeddings of the target nodes and is updated through a sliding window mechanism to limit the computational complexity.
7. The method for tracking elliptical multi-extended targets based on a dynamic heterogeneous graph neural network according to claim 2, wherein The associated calculation layer is used to embed and map measurement nodes and target nodes into a common space through linear transformation, and generate the measurement-target association probability matrix through matrix multiplication and sigmoid function; The target quantity prediction layer is used to predict the total number of targets at the next moment according to the current measurement information and historical target information; The target state prediction layer is used to map and obtain a predicted state vector through a multi-layer perceptron by combining the measurement information weighted by the association probability, the target node embedding and the temporal aggregation embedding.
8. The method for tracking elliptical multi-extended targets based on a dynamic heterogeneous graph neural network according to claim 2, wherein The multi-task loss function includes binary cross-entropy loss, categorical cross-entropy loss, mean absolute error loss and Gaussian Wasserstein distance loss; Among them, the steps of the Gaussian Wasserstein distance loss are calculated as follows: Convert the discrete measurement point set into a probability distribution through the Dirac measure; For each discrete point, calculate the Wasserstein distance between the Gaussian distribution with the discrete point as the mean and the unit matrix as the variance and the target elliptical Gaussian distribution; Take the average of the Wasserstein distances of all points as the metric between the point set and the elliptical distribution: Based on the Gaussian Wasserstein distance formula, calculate the distribution difference between each measurement point and the target ellipse, and take the average as the final loss term.
9. An elliptical multi-extended target tracking system based on a dynamic heterogeneous graph neural network, characterized in that, Including: Information acquisition unit: used to acquire the information of the target to be tracked; Model construction unit: used to construct a multi-extended target tracking model; Tracking processing unit: used to input the information of the target to be tracked into the multi-extended target tracking model for processing to obtain a tracking result; among them, the multi-extended target tracking model is used to model the multi-extended target tracking process as a discrete-time dynamic heterogeneous graph, and guide the global optimization of the model by designing an end-to-end multi-task loss function.
10. The elliptical multi-extended target tracking system based on the dynamic heterogeneous graph neural network according to claim 9, wherein, The model construction unit includes: Dynamic heterogeneous graph modeling module: used to construct a heterogeneous graph slice containing measurement nodes and target nodes in real time, where the heterogeneous graph slice includes measurement-measurement edges, measurement-target edges and target-target edges; Heterogeneous graph encoding module: used to encode the heterogeneous graph slice at each moment, and aggregate the embedding representations of different types of nodes through the multi-head attention mechanism to generate updated node embeddings; Temporal aggregation module: used to perform temporal modeling on the embedding sequence of target nodes, and generate a target node representation containing long-term spatio-temporal dependencies by combining the historical state memory matrix; Association prediction decoding module: used to jointly output the measurement-target association probability, target quantity and state estimation; Multi-task optimization module: used to design a multi-task loss function, jointly optimize the data association, target quantity estimation and state filtering tasks, guide the global optimization of the model, and obtain the multi-extended target tracking model.
Citation Information
Patent Citations
Single program splitting method and system for multi-channel attention map neural network clustering
CN114647465A
Multi-extended target tracking method based on deep neural network
CN117788511A
Multi-robot cluster control method and system based on deep reinforcement learning
CN118795893A
Training and prediction of hybrid graph neural network model
US20240152732A1