Multi-model maneuvering target tracking method based on Transform architecture

Through the multi-model maneuver target tracking method based on the Transformer architecture, combined with the interactive multi-model algorithm and the Transformer auxiliary identification module, the problems of improper model matching and poor algorithm stability in maneuver target tracking are solved, and high-precision and stable target tracking are achieved.

CN120405694APending Publication Date: 2025-08-01HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510528454.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

When facing the random and unpredictable motion patterns of maneuverable targets, existing maneuverable target tracking algorithms have problems such as improper model matching, poor algorithm stability, lack of interpretability and insufficient noise data processing capabilities.

Method used

The multi-model maneuverable target tracking method based on the Transformer architecture is adopted, and through interactive multi-model algorithm and Transformer assisted recognition module, combined with traceless Kalman filtering, the target motion state is accurately modeled, and effective features are extracted and state estimated using Transformer to prevent model interference and improve algorithm stability and interpretability.

Benefits of technology

It significantly improves the target tracking performance, improves practicality and generalization capabilities in various scenarios, ensures accurate modeling and tracking accuracy of target motion state, and reduces the complexity of network learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120405694A_ABST
    Figure CN120405694A_ABST
Patent Text Reader

Abstract

The invention relates to a maneuvering target tracking method, in particular to a multi-model maneuvering target tracking method based on a Transform framework. The method comprises the following steps: step 1, measuring a distance and an azimuth angle of a target, acquiring k target positions for observation, and obtaining a continuous target observation track; pre-processing: carrying out non-overlapping segmented time domain partitioning processing on the target observation trajectory to generate local dynamic features, and obtaining a state change fragment set by adopting an interactive multi-model algorithm; step 2, setting a target maneuvering model library, and using a Transform auxiliary identification module to obtain a target motion category and specific maneuvering parameters; and step 3, establishing a target accurate motion model according to a target motion category and specific maneuvering parameters, estimating target position and speed information by using unscented Kalman filtering according to the target observation trajectory data preprocessed in the step 1, and obtaining a final target trajectory after trajectory reconstruction, so as to effectively capture the maneuvering change of the target and improve the maneuvering accuracy of the target. And the target tracking performance is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a maneuvering target tracking method, and more particularly to a multi-model maneuvering target tracking method based on the Transformer architecture. Background Art

[0002] Maneuvering target tracking is crucial in both civilian and military fields, including air traffic control and air defense reconnaissance. For a long time, radar has been widely used for target detection and tracking. Radar has the ability to detect all-weather, at long distances and provide target range and angle information. However, it is sensitive to clutter interference and has low detection accuracy.

[0003] With the development and application of vision technology, optoelectronic tracking systems (OETS) have gradually emerged in the field of target tracking. OETS combines optical sensors and ranging sensors to provide more accurate target position information. Different from fixed-view optical detection (high-speed or infrared cameras), OETS can use the target miss distance provided by the optical sensor to adjust the two-axis turntable so that the target remains within the camera's field of view. In addition, by superimposing the target miss distance on the angle output by the high-precision encoder of the turntable, high-precision target angle information can be obtained. However, even though the observation accuracy has been improved, the interference of observation noise still exists. Estimating the target motion state using noisy data is an important task.

[0004] So far, a large number of target tracking algorithms have been proposed, such as the Kalman filter (KF), extended Kalman filter (EKF), unscented Kalman filter (UKF), and particle filter (PF), etc. These techniques use a pre-determined single motion model to simulate the actual trajectory of the target, and then use the observed values to correct the model prediction results. However, maneuvering targets exhibit random and unpredictable motion patterns, which limits the effective tracking of a single model. Subsequently, many multi-model-based methods have been developed to improve the accuracy of moving target tracking, such as the interacting multiple model (IMM) and its variants. In such methods, multiple motion models are combined with various filters to track the target simultaneously. Then, the tracking results of each model are mixed with specific weights to produce the best tracking performance. Since these methods apply multiple target motion models, their effectiveness depends on the quality of motion modeling and the reasonable allocation of weights between models. In addition, such techniques require collecting a certain amount of observation data to determine changes in the target situation, which may cause estimation delays. End-to-end tracking algorithms based on deep learning can directly map the original measurement sequence to the state estimate, thus avoiding errors related to motion modeling and nonlinear processing. However, the algorithm has poor stability when facing scene changes, and such methods lack interpretability.

[0005] Although some multi-model-based algorithms and deep learning-based algorithms have achieved reliable object tracking results, there are still the following problems: 1. For multi-model-based algorithms, the rationality of probability allocation among models is the basis of the algorithm's estimation performance. Due to the interaction mechanism, the correct motion model at each moment will be contaminated by other models. This problem is more obvious when the probabilities of multiple models are similar. In addition, fixed-parameter motion models cannot adapt to the random maneuvers of the target. After the target's motion state changes, there will be a problem of mismatch between the motion model and the target's motion. 2. The LSTM-based method can provide good tracking effects. However, as the sequence interval becomes longer and the forget gate in the model is used, it still ignores the key state information of the target components. 3. End-to-end algorithms based on deep learning have recently received extensive attention and achieved gratifying results. However, this method also has some disadvantages. First, deep learning is not suitable for directly extracting effective information from noisy measurements. In addition, the training data also has a great impact on the model performance. If the dataset is unbalanced or the working scenario changes, the stability, accuracy, and generalization ability of the algorithm will be greatly reduced, or even completely lose value. Finally, it lacks interpretability because it regards the tracking process as a completely black-box process. Summary of the Invention

[0006] The present invention provides a multi-model maneuvering target tracking method based on the Transformer architecture, aiming to effectively capture the maneuvering changes of the target and significantly improve the target tracking performance.

[0007] The above object is achieved by the following technical solutions:

[0008] A multi-model maneuvering target tracking method based on the Transformer architecture includes the following steps:

[0009] Step 1: Measure the distance and azimuth angle of the target. After collecting k target position observations, obtain a continuous target observation trajectory; Preprocessing: Perform non-overlapping time-domain block processing on the target observation trajectory to generate local dynamic features, and use the interactive multi-model algorithm to obtain a set of state change segments;

[0010] Step 2: Set up a target maneuver model library, and use the Transformer-assisted recognition module to obtain the target motion category and specific maneuver parameters;

[0011] Step 3: Establish an accurate target motion model based on the target motion category and specific maneuver parameters. For the target observation trajectory data after preprocessing in Step 1, use the unscented Kalman filter to estimate the target position and velocity information, and obtain the final target trajectory after trajectory reconstruction.

[0012] The beneficial effects of a multi-model maneuvering target tracking method based on the Transformer architecture in the present invention are as follows:

[0013] Through the preprocessing method proposed in the present invention, the practicability of the algorithm in various scenarios with different positions, speeds, and turning rates is improved. Through the Transformer-based motion state recognition module proposed in the present invention, that is, the motion model classification module and the turning rate estimation module, accurate modeling of the motion state of the maneuvering target at each stage can be achieved. Only a unique motion model is selected to describe the target motion at each stage, thus ensuring that there is no interaction between multiple models and preventing model interference. By combining deep learning and the multi-model tracking idea, the stability of the algorithm is improved while enhancing the interpretability of the algorithm. Using the state change between adjacent sampling points as the feature input sequence can prevent the influence of target distance differences and different trajectory initial values on the algorithm performance, reduce the complexity of network learning, and improve its generalization ability. These solutions can effectively capture the maneuvering changes of the target and significantly improve the target tracking performance. Description of the Drawings

[0014] Figure 1 is the overall flowchart of a multi-model maneuvering target tracking method based on the Transformer architecture in the present invention;

[0015] Figure 2 is the specific structure diagram of the TARM model;

[0016] Figure 3 is the flowchart of using TARM for motion model modeling;

[0017] Figure 4 is the parameter diagram of the customized trajectory dataset;

[0018] Figure 5 is the flowchart of trajectory segmentation and reconstruction;

[0019] Figure 6 is the confusion matrix diagram of TARM_1 on the test set;

[0020] Figure 7 is according to Figure 4 the parameter range of, 9 trajectories containing various maneuvering forms are generated to evaluate the tracking accuracy and stability of multiple algorithms, and the specific parameter diagrams of the nine trajectories;

[0021] Figures 8 to 16 are the test result diagrams of the nine trajectories respectively;

[0022] Figures 17 to 34 Among them, the odd-numbered diagrams respectively correspond to the position RMSE diagrams of the nine trajectories, and the even-numbered diagrams respectively correspond to the speed RMSE diagrams of the nine trajectories;

[0023] Figure 35 Enumeration diagram of position / speed ARMSE mean root mean square error;

[0024] Figure 36 Enumeration diagram of maximum ARMSE. Detailed implementation manners

[0025] A multi-model maneuvering target tracking method based on the Transformer architecture, combined with Figure 1 , includes the following steps:

[0026] 1. In a two-dimensional coordinate system, use an optoelectronic tracking turntable to measure the distance and azimuth angle of the target, and set Z t = [r, θ] t , where r and θ respectively represent the distance and azimuth angle between the target and the optoelectronic tracking system at time t;

[0027] After continuously collecting k target position observations at equal intervals at a frequency of 10 Hz, a continuous target observation trajectory is obtained

[0028] 1-1. Preprocessing 1 (preprocessing of the left branch): Use a sliding window with a fixed length of m (5 seconds) to divide the observation trajectory containing k sampling points into a set of numbered observation trajectory segments without overlap, where n represents the trajectory serial number and m represents the length of the trajectory segment. The above process is expressed as: where Φ

[0029]

[0030] where Φ seg represents the trajectory segmentation process;

[0031] 1-2. Preprocessing 2 (preprocessing of the right branch): S1. Use the Interacting Multiple Model (IMM) algorithm to process Z 1:k , and the motion model set of IMM includes three motion forms: CV, CTL, and CTR. The turning rate of the left-turning model is set to 5° / s, and the turning rate of the right-turning model is set to -5° / s. The initial model probability and probability transition matrix are set as:

[0032]

[0033] Denote the trajectory processed by the IMM algorithm as

[0034]

[0035] where, Φ IMMRepresenting the IMM algorithm;

[0036] Segment it in the same way as 1-1 to obtain and calculate the state change amount between adjacent time points and use it to recombine the trajectory sequence to obtain the state change segment set

[0037]

[0038] where, Φ seg represents the trajectory segmentation process;

[0039] S2. Set the three most common motion patterns of maneuvering targets as the target maneuver model library: constant velocity (CV), constant turn left (CTL), and constant turn right (CTR); Transformer-aided recognition module: Input the target trajectory data preprocessed by the right branch of S1 into the pre-trained motion model classification form (TARM_1) and turn rate estimation form (TARM_2) to obtain the target motion category and specific maneuver parameters respectively. Among them, Transformer-aided recognition module is briefly referred to as TARM;

[0040] Define the state transition equation and observation equation based on the state space model (SSM) as follows:

[0041] State transition equation:

[0042] X t = FX t-1 + n;

[0043] where, X t = [x, y, V x , V y t , [x, y] t represents the two-dimensional position of the target at time t, and [V x , V y t represents the two-dimensional velocity of the target at time t. F is the transition matrix and n is the transition noise;

[0044] The transition matrices corresponding to the three motion models are:

[0045]

[0046] ​​where ω is the turning rate of the maneuvering target, and Δt is the sampling interval of the trajectory, which is set to 0.1 s. The transition noise n is simulated by Gaussian noise:

[0047]

[0048] where σ d = 0.5σ a ·Δt 2 and σ v = σ a ·Δt are the standard deviations of the distance and velocity transition noises. σ a is the standard deviation of the acceleration noise, nd is the target distance process noise, and nv is the target velocity process noise.

[0049] Observation equation:

[0050] Z t = h(X t ) + m;

[0051] where h is the non-linear observation and m is the observation noise;

[0052] For target tracking in a two-dimensional plane, Z t can be defined as:

[0053]

[0054] where and are the distance noise and azimuth noise. σ r and σ θ represent the standard deviations of the corresponding distance noise and azimuth noise;

[0055] 2-1. As Figure 2 shown, build TARM, which includes four modules from bottom to top, namely the input processing module, the encoder module, the decoder module, and the output module;

[0056] 2-1-1. The input processing module solves the problem that the azimuth input data features are "submerged" by the distance data features, and normalizes each trajectory segment so that represents the trajectory after normalization,

[0057]

[0058] where represents the largest absolute value among all state values of the trajectory segment;

[0059] Use a linear transformation with an output dimension of d m to extract the shallow features of the input sequence to obtain:

[0060]

[0061] where Φ L represents linear transformation processing;

[0062] Add absolute position encodings of the same size, and use sine and cosine functions with different frequencies to construct the position encodings. At time t, the position encoding can be expressed as:

[0063]

[0064] where The final input data added with the position encoding can be expressed as

[0065] 2-1-2. The encoder module uses N stacked encoders to extract self-attention features in each layer;

[0066] Each encoder layer contains a multi-head self-attention sub-layer and a feed-forward network sub-layer, and there is a normalization layer and a residual connection after each sub-layer;

[0067] For the multi-head self-attention sub-layer, in each head, the input sequence is mapped to three different matrices, namely the query matrix the key matrix and the value matrix By calculating the dot product of the query and the key, the similarity matrix can be obtained:

[0068]

[0069] where Softmax is the normalized exponential function, and d k is the scaling factor;

[0070] By calculating A KQ ·V, the attention information in a single head can be obtained. Connecting the attention information of multiple single heads can obtain the multi-head attention information. The data processing process of the i-th encoder can be expressed as:

[0071]

[0072] where and are the input and output of the i-th encoder layer respectively, LN represents layer normalization, and FFN represents the feed-forward network;

[0073] The j-th single-head self-attention module and the multi-head self-attention module of the i-th encoder layer can be expressed as:

[0074]

[0075] Among them represents the similarity matrix calculated by the j-th single-head attention module in the i-th encoder layer. and are learnable projection matrices. d k = d m / h, where h is the number of single-head attentions;

[0076] The feed-forward network sub-layer contains two fully connected layers to perform linear transformation on the attention information at each position. The output of the encoder layer is defined as Then it can be expressed as:

[0077]

[0078] Among them, Φ fully represents the processing of the fully connected layer. is the final output after being processed by N encoder layers.

[0079] 2-1-3. The decoder module uses two 1D convolutional layers as the decoder. Let be the output of the decoder:

[0080]

[0081] Among them, Φ 1D represents the 1D convolution operation;

[0082] 2-1-4. The output module uses a linear transformation layer as the output module. Let be the output of the output module:

[0083]

[0084] Among them, Φ L represents the linear transformation processing;

[0085] For the motion model classification form (TARM_1), the output dimension of the linear layer is set to 3, corresponding to the probabilities of the three motion models of the n-th trajectory segment respectively and Among them, corresponds to the probability of the constant velocity straight-line motion (CV) model, corresponds to the probability of the constant velocity left-turn motion (CTL) model, corresponds to the probability of the constant velocity right-turn motion (CTR) model;

[0086] As Figure 3 shown, select the motion model corresponding to the maximum value for each trajectory segment as its corresponding motion model category That is, the target motion state only corresponds to one of the three motion models in each stage. Corresponding to the transition matrix F CV , Corresponding to the transition matrix F CTL , Corresponding to the transition matrix F CTR ;

[0087] By processing the trajectory segment set through TARM_1 , a set of transition matrices {F n} corresponding to each segment can be obtained;

[0088] For the classification task, the loss function is defined as follows:

[0089]

[0090] where, represents the output of the model, C represents the true label, K represents the number of categories, and i represents the i-th category.

[0091] For the turning rate estimation form (TARM_2), the output dimension of the linear layer is set to 1, corresponding to the target turning rate ω n of the n-th trajectory segment, as Figure 3 shown;

[0092] When the target motion recognized by TARM_1 is uniform linear motion, the turning rate is directly set to 0 and does not need to be estimated by TARM_2.

[0093] For the regression task, the loss function is defined as follows:

[0094]

[0095] where, represents the estimated turning rate of the model, ω n represents the true label of the turning rate, and N represents the number of samples.

[0096] Combining the transition matrix F n of each trajectory segment with the corresponding turning rate ω n can obtain a set of motion models {M n} corresponding to the trajectory segments, as Figure 3 shown;

[0097] 2-2. According to the different tasks, TARM is divided into the motion model classification form (TARM_1) and the turning rate estimation form (TARM_2). The basic structures of the two are the same, but the number of parameters is different. Different loss functions are used to optimize the two model forms respectively.

[0098] In the classification task (TARM_1), the output dimension of the first linear transformation layer is set to 64. The depth of the encoder is set to 2, the output dimension of the multi-head attention layer is set to 64, 32 heads are used, and the dimension of each single-head attention layer is 2. The output dimensions of the one-dimensional convolutional layer in the decoder are set to 32 and 16. The output dimension of the second linear transformation layer is set to 3 to obtain the probabilities corresponding to the three types of motion models; the above parameters are preferred parameters, and other parameters are also acceptable, but a small number of parameters will affect the accuracy, and a large number of parameters will not significantly improve the effect and will be time-consuming.

[0099] In the regression task (TARM_2), the output dimension of the first linear transformation layer is set to 64. The depth of the encoder is set to 3, the output dimension of the multi-head attention layer is set to 64, 32 heads are used, and the dimension of each single-head attention layer is 2. The output dimensions of the one-dimensional convolutional layer in the decoder are set to 32 and 16. The output dimension of the second linear transformation layer is set to 1 to obtain the turning rate corresponding to the trajectory segment;

[0100] 2-3. Specifically, TARM_1 and TARM_2 are obtained through the following datasets and training methods:

[0101] Set the initial state X0, time step k, noises n and m, transition matrix F, and the minimum interval between turning rates is set to 0.1° / s. The parameters of the customized trajectory dataset are as Figure 4 shown.

[0102] According to Figure 4 , obtain the trajectory dataset for target tracking from the state transition equation and the observation equation. Among them, 150,000 trajectories are generated for training and testing the classification ability of TARM, and 300,000 trajectories are generated for training and testing the estimation ability of TARM;

[0103] Each group of data contains the target true trajectory, measurement trajectory, processed state change trajectory segment, and the corresponding classification label (for TARM_1) or turning rate label (for TARM_2);

[0104] Each target true trajectory segment and observation trajectory segment contain 50 sampling points, and the sampling interval is 0.1;

[0105] Divide this dataset into a training set and a test set according to a ratio of 4:1;

[0106] Train TARM_1 and TARM_2 100 times each with a batch size of 100 on a single NVIDIA GeForce GTX 3090 GPU;

[0107] Use the ADAM optimizer, set the initial learning rate to 0.001, and the decay factor for every ten epochs to 0.1;

[0108] S3. Establish a target accurate motion model based on the target motion category and specific maneuver parameters obtained in S2. Based on the target observation trajectory data preprocessed by the left branch in S1, use the unscented Kalman filter (UKF) to estimate the target position and velocity information, and obtain the target final trajectory after trajectory reconstruction;

[0109] 3-1. As Figure 3 shown, use the UKF algorithm combined with the set of motion models {M n} obtained in the previous section to track each trajectory segment in the set of observation trajectories preprocessed by the left branch in step (1-1). Let the set of trajectories processed by the UKF algorithm be

[0110] 3-2. As Figure 5 shown, splice the set of trajectories processed by the UKF algorithm according to the trajectory segment numbers obtained by the preprocessing operation in S1. Let the final complete target trajectory after splicing be where Φ

[0111]

[0112] represents the UKF algorithm, and F represents the transition matrix; UKF

[0113] Test results:

[0114] To illustrate the effectiveness of the method of the present invention and the improvement in the tracking accuracy and stability of maneuvering targets, the motion model classification ability of TARM_1 was verified on the test set generated in 2-3, and the generated confusion matrix is as Figure 6 shown;

[0115] where 0 is the label of the CV model, 1 is the label of the CTL model, and 2 is the label of the CTR model. The recognition accuracies of the three models are as high as 98.77%, 99.00%, and 99.13% respectively, and the average recognition accuracy is 98.97%. The extremely low false detection rate proves the feasibility of using TARM_1 for target motion model recognition;

[0116] In addition, according to Figure 4 the parameter range of, nine trajectories containing various maneuver forms were generated to evaluate the tracking accuracy and stability of multiple algorithms. The specific parameters of the nine trajectories are as Figure 7 shown:

[0117] All experiments have undergone 100 Monte Carlo simulations; ​

[0118] Figures 8 to 16 The position tracking trajectory comparison between the present invention and other methods is shown. It can be clearly seen that the present invention is more consistent with the true maneuvering trajectory of the target compared with other methods. During the entire tracking process, the trajectory does not show obvious vibrations or deviations, and it can still accurately track after the target makes a maneuvering movement;

[0119] The evaluation indexes involved in the test of this method include: the position / velocity RMSE (root mean square error) graph, which evaluates the error distance of all points between the tracking trajectory and the true trajectory. The smaller the value, the higher the accuracy of the method, and the smaller the change range of this value, the better the stability of the method. As Figures 17 to 34 shown; the position / velocity ARMSE (average root mean square error) graph, which evaluates the average error distance of all points between the tracking trajectory and the true trajectory. The smaller the value, the higher the comprehensive tracking accuracy of the method. As Figure 35 shown; the maximum ARMSE graph, which evaluates the maximum error distance among all points between the tracking trajectory and the true trajectory. The smaller the value, the higher the tracking stability of the method. As Figure 36 shown.

Claims

1. A multi-model maneuvering target tracking method based on the Transformer architecture, comprising the following steps: Step 1: Measure the distance and azimuth angle of the target. After collecting k target position observations, a continuous target observation trajectory is obtained. Preprocessing: Perform non-overlapping time-domain block processing on the target observation trajectory to generate local dynamic features, and use the interactive multi-model algorithm to obtain a set of state change segments. Step 2: Set up a target maneuver model library, and use the Transformer-assisted recognition module to obtain the target motion category and specific maneuver parameters. Step 3: Establish an accurate target motion model based on the target motion category and specific maneuver parameters. For the target observation trajectory data after preprocessing in Step 1, use the unscented Kalman filter to estimate the target position and velocity information, and obtain the final target trajectory after trajectory reconstruction.

2. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 1, wherein the preprocessing: 1-1. Use a sliding window with a fixed length of m and a step size of p = m to non-overlappingly segment the observation trajectory containing k sampling points into a set of numbered observation trajectory segments where n represents the trajectory number, and Φ represents the trajectory segmentation process; seg ​ The trajectory after being processed by the IMM algorithm is Among them, Φ IMM represents the IMM algorithm; Segment in the same way as 1-1 to obtain and calculate the state change amount between adjacent time points and use it to recombine the trajectory sequence to obtain a set of state change segments 3. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 1, wherein the target maneuver model library includes uniform motion, uniform left turn, and uniform right turn.

4. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 3, input the target trajectory data after preprocessing the right branch of Step 1 into the pre-trained TARM_1 and TARM_2, and obtain the target motion category and specific maneuver parameters respectively. The state transition equation is as follows: X t = FX t-1 + n; Among them, X t = [x, y, V x , V y t , [x, y] t represents the two-dimensional position of the target at time t, [V x , V y t represents the two-dimensional velocity of the target at time t, F is the transition matrix, and n is the transition noise;​​ The transition matrices corresponding to the three motion models are: ω is the turning rate of the maneuvering target, and Δt is the sampling interval of the trajectory, which is set to 0.1 s. The transition noise n is simulated by Gaussian noise: σ d = 0.5σ a ·Δt 2 and σ v = σ a ·Δt is the standard deviation of the distance and speed transition noise, and σ a is the standard deviation of the acceleration noise; Observation equation: Z t = h(X t ) + m; Where h is the non-linear observation and m is the observation noise; For object tracking in a two-dimensional plane, Z t can be defined as: where and are the range noise and the azimuth noise, and σ r and σ θ represent the standard deviations of the corresponding range noise and azimuth noise, respectively.

5. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 2, in Step 2, construct TARM, which includes an input processing module, an encoder module, a decoder module, and an output module from bottom to top.

6. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 5, for each trajectory segment is normalized, and let represent the trajectory after normalization Among them represents the number with the largest absolute value among all the state values of the trajectory segment; Extract the shallow features of the input sequence using a linear transformation with an output dimension of d m to obtain: where Φ L represents linear transformation processing; Add absolute position encodings of the same size, and use sine and cosine functions with different frequencies to construct the position encodings. At time t, the position encoding can be expressed as: Among them The final input data with positional encoding added is represented as The encoder module uses N stacked encoders to extract self-attention features in each layer. Each encoder layer contains a multi-head self-attention sub-layer and a feed-forward network sub-layer. After each sub-layer, there is a normalization layer and a residual connection.

7. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 6. For the multi-head self-attention sub-layer, in each head, the input sequence is mapped to three different matrices, namely the query matrix the key matrix and the value matrix By calculating the dot product of the query and the key, a similarity matrix can be obtained: Among them, Softmax is a normalized exponential function, and d k is a scaling factor; By calculating A KQ ·V, the attention information in a single head can be obtained. Connecting the attention information of multiple single heads can obtain the multi-head attention information. The data processing process of the i-th encoder can be expressed as: where i = 1, 2, ..., N and are the input and output of the i-th encoder layer, respectively, LN represents layer normalization, and FFN represents a feed-forward network; The j-th single-head self-attention module and multi-head self-attention module of the i-th encoder layer can be expressed as: represents the similarity matrix calculated by the j-th single-head attention module in the i-th encoder layer; and are parameter-learnable projection matrices; d k = d m / h, where h is the number of single-head attentions; The feed-forward network sub-layer contains two fully connected layers to perform linear transformations on the attention information at each position, and the output of the encoder layer is defined as which can be expressed as: Φ fully Indicates the processing of the fully connected layer, which is the final output after being processed by N encoder layers; The decoder module uses two 1D convolutional layers as the decoder, and let be the output of the decoder: Φ 1D represents a 1D convolution operation; The output module uses a linear transformation layer as the output module. Let be the output of the output module: Φ L represents linear transformation processing; For the motion model classification form (TARM_1), the output dimension of the linear layer is set to 3, corresponding to the probabilities of the three motion models of the nth trajectory segment respectively and where The probability corresponding to the uniform linear motion model, The probability corresponding to the uniform left-turning motion model, The probability corresponding to the uniform right-turning motion model; Select the motion model corresponding to the maximum value for each trajectory segment as its corresponding motion model category That is, the target motion state only corresponds to one of the three motion models in each stage. Corresponding to the transition matrix F CV , Corresponding to the transition matrix F CTL , Corresponding to the transition matrix F CTR ; By processing the set of trajectory segments through TARM_1 a set of transition matrices {F n} corresponding to each segment can be obtained; For the classification task, the loss function is defined as follows: Among them, represents the output of the model, C represents the true label, K represents the number of categories, and i represents the i-th category. For TARM_2, the output dimension of the linear layer is set to 1, corresponding to the target turning rate ω of the n-th trajectory segment n ; When the target motion recognized by TARM_1 is uniform straight-line motion, the turning rate is directly set to 0; For the regression task, the loss function is defined as follows: Among them, represents the estimated turning rate of the model, ω n represents the true label of the turning rate, N represents the number of samples, and the transition matrix F of each trajectory segment n is combined with the corresponding turning rate ω n to obtain a set of motion models {M n} corresponding to the trajectory segment.

8. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 7, the basic structures of TARM_1 and TARM_2 are the same, only the number of parameters is different.

9. The multi-model maneuvering target tracking method based on the Transformer architecture according to claim 7, step 3, using the UKF algorithm to combine the obtained set of motion models {M n} to track each trajectory segment in the preprocessed observation trajectory set in step 1. Let the trajectory set processed by the UKF algorithm be 10. For the multi-model maneuvering target tracking method based on the Transformer architecture according to claim 9, splice the trajectory set processed by the UKF algorithm according to the trajectory segment numbers obtained in the preprocessing operation in step 1. Suppose the final complete trajectory of the target after splicing is Among which Φ UKF represents the UKF algorithm, and F represents the transition matrix.