Maneuvering target intelligent tracking method based on self-attention mechanism and physical constraint
By introducing a self-attention mechanism and physical constraints in maneuvering target tracking, combined with TCN-Transformer and insensitive Kalman filter, the problem of reduced tracking accuracy when targets are maneuvered is solved, achieving higher tracking accuracy and efficiency.
Patent Information
- Application Number
- CN202510542346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When the target is accelerated, decelerated or turned, the preset model mismatches with the real movement, resulting in a decrease or divergence in tracking accuracy, especially in the tracking of strong maneuverable targets.
Using an intelligent tracking method based on self-attention mechanism and physical constraints, the TCN-Transformer motion state prediction model combined with insensitive Kalman filters, dynamically model the target motion mode, correct state estimation in real time, and configure physical constraints in the loss function to improve the accuracy and efficiency of prediction.
It significantly improves the tracking accuracy and prediction efficiency of three-dimensional maneuvering targets, and can operate stably under complex environments and noise interference, avoid lag, and meet real-time monitoring and decision-making needs.
Smart Images

Figure CN120065200A_ABST
Abstract
Description
Background Art
[0002] Target tracking is a technology that estimates the continuous motion trajectory of a target using noisy discrete sensor data. Its core is to establish a motion model through dynamic priors and perform iterative optimization by combining spatio-temporal observation laws to output the target state estimation. However, in practical applications, the target often undergoes maneuver mutations such as acceleration, deceleration, or turning, resulting in a mismatch between the preset model and the real motion, leading to a decrease or even divergence in tracking accuracy. In particular, the precise tracking of highly maneuverable targets has always been a key technical problem in radar data processing.
[0003] Currently, the commonly used maneuvering target tracking methods mainly include two categories: one is the method of state prediction based on a prior model, and the other is the method of state prediction based on a neural network. The first category of methods is based on the Kalman filter, and its tracking effect depends on whether the set target motion model is accurate. For example, the multiple model method and the interacting multiple model (IMM) method work synchronously with multiple models and perform weighted summation on their respective state estimations to obtain a better state estimation. However, due to the inaccuracy of the model and the lag in the calculation and update of the weights, at the moment when the target maneuver occurs, the probabilities of each model are not accurate, resulting in a lag in the tracking result and a decrease in tracking accuracy. Summary of the Invention
[0004] The present invention provides an intelligent tracking method for maneuvering targets based on a self-attention mechanism and physical constraints. The method improves the accuracy of the tracking performance of three-dimensional maneuvering targets compared with the IMM interacting multiple model method, and improves the prediction accuracy and efficiency compared with the existing intelligent tracking methods for maneuvering targets.
[0005] The method includes: Step S101: Perform coordinate transformation on the original radar measurement data to obtain the traces in the Cartesian coordinate system, and use the interacting multiple model algorithm for preliminary tracking, and take the data of the first δ time step lengths as the initial window; Step S102: Calculate the Sigma point set of the target state vector at each time step of the initial window; Step S103: Calculate the number of Sigma points for n sampling points; Step S104: Normalize the Sigma point sequence and save the maximum and minimum values; Step S105: Input the normalized Sigma point sequence into the TCN-Transformer motion state prediction model to generate the main prediction, speed prediction, and acceleration prediction results; Step S106: Configure physical constraints in the loss function, and the physical constraints include speed constraint loss and acceleration constraint loss; Step S107: Combine the prediction result with the unscented Kalman filter and utilize the real-time correction ability of the unscented Kalman filter; Step S108: Output the estimated values of the final position, velocity, and acceleration for display in the radar tracking system.
[0006] It should be further noted that Step S101 also includes: converting the original radar measurement data from the distance, azimuth angle, and elevation angle in the polar coordinate system into a nine-dimensional target state vector based on three-dimensional position, velocity, and acceleration in the Cartesian coordinate system as a trace.
[0007] It should be further noted that in Step S102: when calculating the Sigma point set of the target state vector at each time step in the initial window, it is determined according to the generation method of Sigma points in the unscented transformation.
[0008] It should be further noted that the generation method of Sigma points in the unscented transformation includes the following methods: For an n-dimensional random variable with a mean and a covariance the Sigma points are generated by the following formula: The first point is the mean itself
[0009] The remaining 2n points are obtained by offsetting the mean and the square root of the covariance matrix:
[0010]
[0011] where S is the square root after the Cholesky decomposition of the covariance matrix P, is the scaling parameter; , and are the parameters that control the distribution of Sigma points, controls the distance between the control point and the mean value, while is the adjustment parameter; The weights of the Sigma points are calculated as follows:
[0012]
[0013]
[0014] where, represents the average weight of the i-th Sigma point, represents the covariance weight of the i-th Sigma point, is a parameter for adjusting covariance estimation.
[0015] It should be further noted that step S107 also includes: in the prediction stage, requiring the Sigma points to be transformed into a set of points in the observation space through the observation model; Each Sigma point is propagated through a non - linear function and incorporated with the prediction information of the model to obtain:
[0016] The estimated mean and covariance matrix are calculated through weighted averaging:
[0017]
[0018] By continuously performing the prediction and update steps, and utilizing the real - time correction ability of the unscented Kalman filter, the state estimation is gradually optimized.
[0019] It should be further noted that in step S103, the method for calculating the number of Sigma points is: , where c represents the dimension of the target state vector; Each sampling point corresponds to a set of Sigma points, and n sampling points are concatenated to form an input sequence of m×c.
[0020] It should be further noted that in step S105, the TCN - Transformer motion state prediction model has two TCN layers and one Transformer layer; The dilation rates of the two TCN layers are 1 and 2 respectively, and the padding is set to 8 and 16 in the first and second layers respectively; a Relu activation function is configured after each convolutional layer.
[0021] It should be further noted that the physical constraint L in step S106 main is the mean square error between the predicted state and the true label: ; The velocity constraint loss is expressed as: ; The acceleration constraint loss is expressed as: ; where N is the total number of time steps within the time window, i is the loop variable representing the current calculated sample or time step sequence number, ranging from 1 to N; v t is the three - dimensional velocity of the target at time step t, p t is the three - dimensional position of the target at time step t, v pred is the predicted target velocity component, y trueis the true value of the target motion state, y pred is the predicted value of the target motion state, a pred is the predicted target acceleration component.
[0022] It should be further noted that the method also involves balancing the loss terms by hyperparameters λ1 and λ2 in the following way: T Z =L main +λ 1 L v +λ 2 L a .
[0023] It should be further noted that step S108 also uses the normalization parameters saved in step S104 to restore the prediction result from the [0, 1] interval to the actual physical dimension; Output the corrected nine-dimensional state vector as the track data for real-time display of the radar tracking system.
[0024] It can be seen from the above technical solutions that the present invention has the following advantages: The intelligent tracking method for maneuvering targets based on self-attention mechanism and physical constraints provided by this application provides a unified data basis through coordinate transformation, combines the interactive multi-model algorithm to adaptively match the target motion mode, and improves the initial tracking accuracy.
[0025] The TCN-Transformer motion state prediction model of this application uses the parallel processing structure of TCN to shorten the response time, combines the self-attention mechanism of Transformer to dynamically model the historical trajectory correlation, effectively captures the trajectory evolution under complex motion modes such as acceleration and turning, and adapts to diverse, high-dimensional and fast-changing tracking scenarios. The unscented Kalman filter is combined with the deep learning model to overcome the problem of prediction error accumulation in a pure data-driven model, and incorporates prior dynamic knowledge to enhance the interpretability and reliability of the prediction; physical constraints reduce the sensitivity to noise data and improve the generalization ability of the model, enabling the tracking system to still operate stably under complex environments and noise interference. The parallel data processing structure of TCN significantly shortens the response time to the dynamic changes of the target, ensures the real-time nature of the tracking process, avoids lag phenomena, and meets the requirements for real-time monitoring and decision-making of the target. The finally output position, velocity and acceleration estimation values intuitively display the real-time state of the target, providing a clear decision-making basis for the operator. Description of the Drawings In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 Flowchart of an intelligent tracking method for maneuvering targets based on self-attention mechanism and physical constraints; Figure 2 Schematic diagram of the Transformer structure; Figure 3 Schematic diagram of the network framework of the TCN-Transformer motion state prediction model; Figure 4 Flowchart of the TTU; Figure 5 Flowchart for constructing the Sigma point set. Detailed implementation manners
[0027] The intelligent tracking method for maneuvering targets based on self-attention mechanism and physical constraints provided by this application uses a temporal convolutional neural network (TCN) that captures long-term dependence relationships to extract target motion features, uses a Transformer network based on self-attention mechanism to achieve accurate regression of the maneuvering model, and uses an unscented Kalman filter (UKF) to introduce measurements and achieve target motion state estimation. On this basis, to improve the accuracy of trajectory prediction, a physical constraint mechanism is introduced. Simulation analysis and the processing results of radar measured data show that the three-dimensional maneuvering target tracking performance of the proposed method has been greatly improved compared with the traditional interacting multiple model (IMM) method, and has also been significantly improved compared with the existing intelligent tracking methods for maneuvering targets.
[0028] The TCN involved in this application is a special convolutional neural network for processing time series data, and it performs excellently in extracting long-term dependence relationship features of sequences and parallel computing. Its core is divided into three parts in total: causal convolution, dilated convolution, and residual connection.
[0029] The Transformer network involved in this application is a deep learning architecture based on self-attention mechanism, which realizes sequence data modeling through global attention mechanism. As Figure 2 shown, its working principle can be divided into the following four links: 1) Self-attention mechanism: Dynamically capture long-distance dependence relationships by calculating the correlation weights between sequence elements. This process enables the model to focus on context information at different positions.
[0030] 2) Multi-head attention: Execute multiple independent self-attention calculations in parallel, splice the outputs of each head, and fuse them through linear transformation. This design expands the ability of the model to capture diverse feature patterns in different subspaces.
[0031] 3) Position Encoding: To make up for the position-agnostic defect of the self-attention mechanism, position embedding vectors are generated through sine functions and added to the input embeddings to inject position information. This enables the model to distinguish the sequential relationship of sequence elements and breaks through the serial computing limitation of traditional RNNs.
[0032] 4) Feed-Forward Network: Two fully connected layers are applied after attention calculation, and a ReLU activation function is used to achieve non-linear transformation. This module independently processes the features of each position, forms a functional complement with the attention mechanism, and enhances the expressive power of the model.
[0033] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to provide a thorough understanding of the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details.
[0034] It should be understood that when used in the specification of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0035] It should be understood that "one or more" mentioned in the present application refers to one, two or more than two, and "multiple" mentioned in the present application refers to two or more than two. In the description of the present application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. The "and / or" herein is merely a description of the association relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone.
[0036] The statement "in one embodiment" or "in some embodiments" etc. described in the present application means that the specific features, structures or characteristics described in the embodiment are included in one or more embodiments of the present application. Thus, the statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different parts of the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0037] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0038] Please refer to Figure 1 The figure shows a flowchart of an intelligent tracking method for maneuvering targets based on self-attention mechanism and physical constraints in a specific embodiment. The method includes: Step S101: Perform coordinate transformation on the original radar measurement data to obtain the tracks in the Cartesian coordinate system, and use the interactive multiple model algorithm for preliminary tracking. Take the data of the first δ time step lengths as the initial window.
[0039] In this embodiment, the original radar measurement data may exist in the form of polar coordinates, etc. For the convenience of subsequent processing, the distance, azimuth angle, and elevation angle in the polar coordinate system can be converted into the three-dimensional position in the Cartesian coordinate system, which can be the tracks.
[0040] In three-dimensional space, after completing the coordinate transformation, use the interactive multiple model algorithm (IMM) for preliminary tracking. The IMM algorithm will run multiple different motion models simultaneously, and each model corresponds to a possible target motion mode, such as uniform linear motion, uniformly accelerated motion, turning motion, etc. According to the radar measurement data and the state estimates of each model, calculate the probability of each model. Then, based on these probabilities, weight and fuse the estimation results of each model to obtain the preliminary tracking result. Finally, take the data of the first δ time step lengths as the initial window, and these data will be used as the input for the subsequent steps.
[0041] The interactive multiple model algorithm in this embodiment can adaptively select the most suitable motion model according to the actual motion situation of the target, improving the accuracy of preliminary tracking. Taking the data of the first δ time steps as the initial window can obtain a segment of historical motion information of the target for analyzing the motion trend and characteristics of the target.
[0042] Step S102: Calculate the Sigma point set of the target state vector at each time step of the initial window.
[0043] In this embodiment, for the target state vector at each time step in the initial window, calculate the Sigma point set. The target state vector usually contains information such as position, velocity, and acceleration, and is a nine-dimensional vector in this embodiment. The mean and covariance matrix of the target state vector can be determined first. The mean can be obtained by statistical averaging of the target state vector over a period of time, and the covariance matrix reflects the correlation and uncertainty between the dimensions of the target state vector.
[0044] For this embodiment, it can be processed based on the Unscented Kalman Filter (UKF), where the Unscented Kalman Filter is an advanced filtering technique based on the Unscented Transform. The implementation of UKF first involves the selection of sigma points. For an n-dimensional random variable with a mean and covariance , the Sigma points are generated by the following formula: The first point is the mean itself
[0045] The remaining 2n points are obtained by offsetting the mean and the square root of the covariance matrix:
[0046]
[0047] where S is the square root of the covariance matrix P after Cholesky decomposition, is the scaling parameter. , and are parameters that control the distribution of Sigma points. controls the distance between the points and the mean value (usually less than 1), while is the adjustment parameter.
[0048] The weights of the Sigma points are calculated as follows:
[0049]
[0050]
[0051] where, represents the average weight of the i-th Sigma point, while represents the covariance weight of the i-th Sigma point. is a parameter used to adjust the covariance estimate and is usually set to 2 in the Gaussian distribution.
[0052] In this embodiment, in the intelligent tracking method of maneuvering targets based on the self-attention mechanism and physical constraints, the relevant content of the unscented transform and the unscented Kalman filter is incorporated, making the entire tracking process more perfect and accurate. In step S102, by generating the Sigma point set through the unscented transform and calculating its weights, a data basis that can better reflect the uncertainty and distribution characteristics of the target state is obtained. In step S107, by using the prediction and update mechanisms of the unscented Kalman filter, the neural network prediction result is combined with the unscented Kalman filter, enabling the tracking system to better fuse the model prediction information and the latest measurement information in complex maneuvering target tracking scenarios and correct the state estimate in real time. This combination method makes up for the defects of the tracking algorithm based solely on deep learning, such as the accumulation of prediction errors and the lack of utilization of classical dynamics prior knowledge, and improves the tracking accuracy and robustness of the system when dealing with highly maneuvering targets.
[0053] Step S103: For n sampling points, calculate the number of Sigma points.
[0054] In this embodiment, according to the relevant principle of the unscented transform, the number of Sigma points is calculated for n sampling points. Here, the n sampling points refer to the number of time steps in the initial window. The purpose of calculating the number of Sigma points is to determine how many Sigma points need to be generated for each time step to fully reflect the changes in the target state.
[0055] The method for calculating the number of Sigma points in this embodiment is as follows: , where c represents the dimension of the target state vector; each sampling point corresponds to a set of Sigma points, and the n sampling points are concatenated to form an input sequence of m×c.
[0056] The specific calculation method of this embodiment is associated with the generation of the Sigma point set in step S102. Sigma points are generated according to the same rule for each time step, and finally the total number of Sigma points is determined.
[0057] Exemplarily speaking, for the nine-dimensional target state vector of each time step, the corresponding number of Sigma points is generated, and the total number of Sigma points for n time steps is the number of Sigma points to be calculated. In this way, by reasonably calculating the number of Sigma points, while ensuring the calculation efficiency, the change information of the target state can be fully captured, ensuring that the tracking algorithm can effectively handle the complex motion situation of the target.
[0058] Step S104: Normalize the Sigma point sequence and save the maximum and minimum values.
[0059] In this embodiment, the obtained Sigma point sequence is normalized. The normalization method usually scales the data between 0 and 1.
[0060] When performing normalization, record the maximum and minimum values in the data. Through normalization, data in different dimensions have the same scale, avoiding the situation where data in certain dimensions dominate in model training due to their large numerical ranges and affecting the performance of the model.
[0061] Step S105: Input the normalized Sigma point sequence into the TCN-Transformer motion state prediction model to generate the main prediction, speed prediction, and acceleration prediction results.
[0062] In this embodiment, before using the neural network for maneuvering target trajectory prediction, a convolutional layer is used to extract motion features. TCN can capture long-range temporal dependencies in a shallow network, with high computational efficiency. And due to its property of maintaining causality, the model prediction only uses current and previous information, avoiding information leakage, which is particularly important for the maneuvering target state prediction task.
[0063] In actual situations, due to the dynamic complexity of the environment, targets often maneuver, which requires the radar tracking system to respond to the dynamic changes of the target in real time. The parallel data processing structure of TCN shortens the response time, ensuring the real-time nature of the tracking process and avoiding the occurrence of lag phenomena. Secondly, the Transformer module dynamically models the temporal correlation of historical trajectories through the multi-head self-attention mechanism, and uses the attention weight distribution to quantify the contribution of states at different time steps to the current prediction. This explicit correlation modeling is more adept at capturing the trajectory evolution laws of maneuvering targets in complex dynamic modes such as acceleration and turning compared to the implicit memory mechanism of traditional LSTM.
[0064] As Figure 3 shows the schematic diagram of the network framework of the TCN-Transformer motion state prediction model involved in this embodiment. The TCN-Transformer motion state prediction model can handle the diversified and complicated problems that occur in the maneuvering target tracking task, especially performing well in large-scale, high-dimensional, and fast-changing tracking scenarios. Figure 3 This is the specific network framework, which consists of two layers of TCN and one layer of Transformer. The input is the preprocessed historical track data in the Cartesian coordinate system, and the structure is b*n*c. Among them, b represents the batch size (i.e., the number of samples that the model processes simultaneously in one forward and backward propagation). n represents the sequence length (in the maneuvering target tracking task, it represents the track time steps within a window), and c represents the number of features.
[0065] In some embodiments, the normalized Sigma point sequence is input into the TCN-Transformer motion state prediction model. The TCN-Transformer motion state prediction model consists of two TCN layers and one Transformer layer. Before entering the TCN layers, the data structure can be transformed because convolutional operations require the number of features in the middle. The dilation rates of the two TCN layers are 1 and 2 respectively. To maintain the time dimension, the padding is set to 8 and 16 in the first and second layers respectively.
[0066] Behind each convolutional layer, there is a Relu activation function, and 0.2 Dropout is applied to prevent overfitting.
[0067] The TCN layer can capture long-range temporal dependencies and extract target motion features. Then the sequence length quantity is transformed to the first dimension and then enters the Transformer layer. The Transformer part processes the 128-dimensional features output by the TCN. The Transformer encoder network consists of 6 layers in total. The number of heads in the multi-head attention mechanism for each layer is set to 8. The number of hidden units in the feed-forward network is 512. A unidirectional structure is adopted to extract features from the forward sequence. The output dimension of the Transformer encoder is 256, and then it is mapped to a 9-dimensional output through a fully connected layer. The main prediction, speed prediction, and acceleration prediction results are output by the multi-task output head. The main prediction result reflects the main states such as the position of the target, and the speed prediction and acceleration prediction provide data for subsequent physical constraints.
[0068] It can be seen that the TCN-Transformer motion state prediction model combines the advantages of TCN and Transformer, and can effectively handle the diversified and complicated problems that occur in the maneuvering target tracking task. The parallel data processing structure of TCN shortens the response time, ensures the real-time nature of the tracking process, and can capture long-range temporal dependencies at the same time. Transformer dynamically models the temporal correlation of historical trajectories through the multi-head self-attention mechanism, and is better at capturing the trajectory evolution law of maneuvering targets in complex dynamic patterns.
[0069] Step S106: Configure physical constraints in the loss function. The physical constraints include velocity constraint loss and acceleration constraint loss.
[0070] In this embodiment, physical constraints are configured in the loss function. The physical constraints include velocity constraint loss and acceleration constraint loss. The purpose of the velocity constraint loss is to constrain the velocity predicted by the model to be consistent with the physical law, that is, the acceleration should be equal to the difference in velocity at adjacent times divided by the time interval. The purpose of the acceleration constraint loss is to constrain the acceleration predicted by the model to be consistent with the physical law, and it is also calculated according to the relationship between the difference in velocity at adjacent times and the time interval.
[0071] In some specific embodiments, in step S106, the physical constraint L main is the mean square error between the predicted state and the true label: ; The physical constraint is to ensure that the position, velocity, and acceleration in the predicted state are overall close to the true values.
[0072] The velocity constraint loss is expressed as: ; The acceleration constraint loss is expressed as: ; where N is the total number of time steps within the time window, i is the loop variable representing the sequence number of the currently calculated sample or time step, ranging from 1 to N; v t is the three-dimensional velocity of the target at time step t, p t is the three-dimensional position of the target at time step t, and in the velocity constraint loss, it is the true physical value used to calculate the velocity. Δt is the radar sampling time interval, which determines the discretization accuracy of the derivative.
[0073] v pred is the predicted target velocity component, y true is the true value of the target motion state, and y true serves as the supervision signal for calculating the main loss to measure the deviation between the predicted value and the true value.
[0074] y pred is the predicted value of the target motion state. In the TCN-Transformer model, after inputting the historical track data, the main prediction result output by the network is y pred。 a pred is the predicted target acceleration component.
[0075] The way to balance the loss terms through hyperparameters λ1 and λ2 is: T Z =L main +λ 1 L v +λ 2 L a 。
[0076] By jointly optimizing the three losses, while fitting the data, the model adheres to Newton's kinematic laws, improving the physical rationality and generalization ability of the prediction.
[0077] Step S107: Combine the prediction result with the unscented Kalman filter to utilize the real-time correction ability of the unscented Kalman filter.
[0078] In some embodiments, the prediction results of the TCN-Transformer model are combined with the unscented Kalman filter. In the unscented Kalman filter, the initial state of the Sigma points is determined according to the prediction results. The latest measurement information is combined with the Sigma points and propagated through a non-linear function to obtain a new set of Sigma points.
[0079] In this embodiment, the estimated mean and covariance matrix are calculated according to the weights of the Sigma points. By continuously performing prediction and update steps, the real-time correction ability of the unscented Kalman filter is used to correct the prediction results of the TCN-Transformer model.
[0080] As an example of this application, when combining the prediction results with the unscented Kalman filter, the prediction and update mechanisms of the unscented Kalman filter are utilized. In the prediction stage, it is required that the Sigma points are transformed into a set of points in the observation space through the observation model; Each Sigma point is propagated through a non-linear function and incorporated with the prediction information of the model to obtain:
[0081] The estimated mean and covariance matrix are calculated through weighted averaging:
[0082]
[0083] By continuously performing prediction and update steps, the real-time correction ability of the unscented Kalman filter is used to gradually optimize the state estimation.
[0084] It can be seen that the unscented Kalman filter can effectively handle the state estimation problem of non-linear systems. By combining with the prediction results of the TCN-Transformer model, it makes up for the defects of the prediction error accumulation and the lack of utilization of classical dynamics prior knowledge in the tracking algorithm based solely on deep learning. The real-time correction ability enables the tracking system to timely adjust the state estimation when the target undergoes complex situations such as high maneuverability, improving the tracking accuracy and robustness.
[0085] Step S108: Output the estimated values of the final position, velocity, and acceleration for display in the radar tracking system.
[0086] In some embodiments, the final estimated values of the target's position, velocity, and acceleration are obtained through multiple steps such as coordinate transformation, preliminary tracking, Sigma point calculation, normalization, model prediction, physical constraint, and unscented Kalman filter correction. The estimated values are output, transmitted to the radar tracking system, and displayed to visually present the real-time state of the target. The display method can be to graphically show the target's position on the radar screen while displaying the numerical information of the velocity and acceleration. This provides intuitive target state information for the operator, facilitating their understanding of the target's movement and enabling them to make corresponding decisions.
[0087] In one embodiment of the present invention, for further illustration of the method, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0088] This application constructs a TCN-transformer-UKF intelligent tracking framework for maneuvering target intelligent tracking method based on self-attention mechanism and physical constraints. The TCN-transformer-UKF intelligent tracking framework combines modern neural networks and traditional Kalman filters, integrating the powerful prediction and model construction capabilities of neural networks with the real-time correction capabilities of Kalman filters and applying physical constraints during training. This further improves the tracking accuracy and robustness of the system when dealing with highly maneuvering targets.
[0089] In the unscented Kalman filter, the change of Sigma points can be regarded as a time series reflecting the dynamic behavior of the system. Therefore, the TCN-Transformer network model proposed in this application can be trained to understand the change pattern of Sigma points over time to predict future states. Based on the TTU maneuvering target intelligent tracking algorithm based on unscented filtering constructed in this application, the specific process is as Figure 4 shown.
[0090] At the beginning stage, appropriate data preprocessing is first performed. The original radar measurement data is subjected to coordinate transformation to obtain the point traces in the Cartesian coordinate system, and the interactive multiple model algorithm is used for preliminary tracking. The data of the first n time step lengths is taken as the initial window, and the Sigma point set of the target state vector at each time step of the initial window is calculated, with the shape of s*n*c, where s represents the number of selected Sigma points, and c represents the dimension of the target state vector. According to the general practice of unscented transformation, there is:
[0091] As Figure 5As shown, in the TTU, since the network-based one-step prediction requires all the information of the target state at the previous n time sampling points, each Sigma point set constructed should also reflect the propagation characteristics during this period and the motion characteristics of the target. Therefore, the number of Sigma points for n sampling points is as follows:
[0092] Before inputting the Sigma point sequence into the TCN-Transformer network for prediction, a normalization module must be introduced to ensure that the algorithm can converge faster and more effectively. Through normalization, the importance attached to each feature by the model is balanced, thus avoiding the problem of poor model training effect caused by the large variation range of position features.
[0093]
[0094] This normalization method scales the data between 0 and 1, where max(x) and min(x) are the minimum and maximum values in the data, and the data of both are saved for subsequent denormalization.
[0095] Compared with existing tracking algorithms, this method provides a new technical approach for the accuracy and reliability of intelligent tracking systems through a collaborative optimization strategy of data-driven + model-driven + physical constraints.
[0096] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0097] The intelligent tracking method for maneuvering targets based on the self-attention mechanism and physical constraints involved in this application combines the units and algorithm steps of each example described in the embodiments disclosed herein, and can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0098] Those skilled in the art can understand that various aspects of the intelligent tracking method for maneuvering targets based on the self-attention mechanism and physical constraints can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0099] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent tracking method for maneuvering targets based on self-attention mechanism and physical constraints, characterized in that: Methods include: Step S101: coordinate transformation is performed on the original radar measurement data to obtain point traces in a Cartesian system, and an interactive multi-model algorithm is used for preliminary tracking, taking the data of the first δ time steps as the initial window; Step S102: Calculate the Sigma point set of the target state vector at each time step in the initial window; Step S103: Calculate the Sigma points for n sampling points; Step S104: normalize the Sigma point sequence and save the maximum and minimum values; Step S105: input the normalized Sigma point sequence into the TCN-Transformer motion state prediction model to generate the main prediction, velocity prediction and acceleration prediction results; Step S106: configuring physical constraints in the loss function, where the physical constraints include velocity constraint loss and acceleration constraint loss; Step S107: combining the prediction result with the insensitive Kalman filter, and utilizing the real-time correction capability of the insensitive Kalman filter; Step S108: Output the final position, velocity, and acceleration estimates for display in the radar tracking system.
2. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 1, characterized in that: Step S101 also includes: converting the original radar measurement data from the distance, azimuth, and pitch angle in the polar coordinate system into a target state vector based on three-dimensional position, velocity, and acceleration in the Cartesian coordinate system, with a total of nine dimensions, as a point trace.
3. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 1, characterized in that: In step S102: when calculating the Sigma point set of the target state vector at each time step in the initial window, it is determined according to the generation method of the Sigma point in the insensitive transformation.
4. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 3, characterized in that: The generation methods of Sigma points in insensitive transformation include the following methods: For a and covariance If n-dimensional random variables are generated, the Sigma points are generated by the following formula: The first point is the mean itself The remaining 2n points are obtained by shifting the mean and the square root of the covariance matrix: Among them, S is the square root of the covariance matrix P after Cholesky decomposition, is the scaling parameter; , and is the parameter that controls the distribution of Sigma points, The distance between the control point and the mean, and is the adjustment parameter; The weight of Sigma points is calculated as follows: in, represents the average weight of the i-th Sigma point, represents the covariance weight of the i-th Sigma point, is a parameter used to adjust the covariance estimate.
5. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 1, characterized in that: Step S107 also includes: in the prediction stage, requiring the Sigma point to be converted into a point set in the observation space through the observation model; Each Sigma point is propagated through a nonlinear function and incorporated into the model’s prediction information to obtain: The estimated mean and covariance matrix are calculated by weighted averaging: By continuously performing prediction and updating steps, the state estimation is gradually optimized by utilizing the real-time correction capability of the insensitive Kalman filter.
6. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 1, characterized in that: In step S103, the Sigma points are calculated as follows: , c represents the dimension of the target state vector; Each sampling point corresponds to a set of Sigma points, and n sampling points are connected in series to form an m×c input sequence.
7. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 1, characterized in that: In step S105, the TCN-Transformer motion state prediction model has two layers of TCN and one layer of Transformer; The dilation rates of the two TCN layers are 1 and 2 respectively, and the padding is set to 8 and 16 in the first and second layers respectively; a ReLU activation function is configured after each convolutional layer.
8. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 1, characterized in that: In step S106, the physical constraint L main is the mean square error between the predicted state and the true label: ; The speed constraint loss is expressed as: ; The acceleration constraint loss is expressed as: ; Where N is the total number of time steps in the time window, i is a loop variable, which represents the sample or time step number currently being calculated, ranging from 1 to N; v t is the three-dimensional velocity of the target at time step t, p t is the 3D position of the target at time step t, v pred is the predicted target velocity component, y true is the true value of the target motion state, y pred is the predicted value of the target motion state, a pred is the predicted target acceleration component.
9. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 8, characterized in that: The method also involves balancing the loss terms through hyperparameters λ1 and λ2 as follows: T Z =L main +λ1L v +λ2L a 。 10. The method for intelligent tracking of maneuvering targets based on self-attention mechanism and physical constraints according to claim 3, characterized in that: Step S108 also uses the normalization parameters saved in step S104 to restore the prediction results from the [0,1] interval to the actual physical dimension; The corrected nine-dimensional state vector is output as the track data displayed in real time by the radar tracking system.
Citation Information
Patent Citations
Long-time target tracking method based on space-time constraint
CN110942471A
Vehicle track deep learning prediction method considering physical constraint
CN116495007A
Geothermal energy production intelligent prediction method and system based on PINNs and physical boundary condition constraint
CN117744527A
Peripheral vehicle track prediction method and device, medium and product
CN118247956A
Intelligent maneuvering target tracking method and device based on long and short term memory and UT conversion
CN118799358A
Cited By
Initial track construction method and device based on double-attention time sequence convolutional network
CN121091238A
A method and apparatus for constructing the initial trajectory based on a dual-attention temporal convolutional network
CN121091238B
Unmanned aerial vehicle tracking and trajectory prediction method based on radar plot information
CN121959458A
A method for tracking and trajectory prediction of a UAV based on radar plot information
CN121959458B
Method and device for starting track of high-maneuvering target in low signal-to-noise ratio
CN122652537A