Distributed target tracking method and system based on transformer interaction multi-model and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies lack stability and consistency in tracking maneuvering targets under sensor network degradation conditions. In particular, when there is packet loss, sudden noise increase, outliers, or local sensor failure, traditional methods are prone to switching lag, probability oscillation, and filter divergence, which reduce the consistency of distributed fusion and tracking accuracy.
A distributed target tracking method based on Transformer interactive multi-model is adopted. Node-level estimation fusion is performed through the TAMPN model, and node reliability is evaluated on the network side. Reliability-weighted information consistency fusion is used to suppress the spread of unreliable node information and improve the timeliness and stability of maneuver pattern discrimination.
Achieving stable and robust distributed target state estimation under sensor degradation environments improves cooperative tracking performance and enhances the stability and accuracy of maneuvering target tracking.
Smart Images

Figure CN122265880A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed maneuvering target tracking. In particular, it relates to a distributed target tracking method, system, and medium based on a Transformer interactive multi-model approach. Technical Field
[0002] This invention relates to the field of distributed maneuvering target tracking. In particular, it relates to a distributed target tracking method, system, and medium based on a Transformer interactive multi-model approach. Background Technology
[0003] With the development of sensor network and unmanned system collaborative technologies, maneuvering target tracking has significant application value in scenarios such as collaborative perception, collaborative monitoring, and integrated air-ground protection. Existing target tracking and distributed state estimation technologies can generally be divided into two categories based on network structure: centralized fusion and distributed fusion. Centralized schemes rely on a central node to aggregate data from the entire network for unified estimation, but this carries a heavy communication burden and the risk of single-point failure. Distributed schemes achieve estimation fusion through neighborhood interactions between nodes, exhibiting better scalability and robustness. Regarding filtering frameworks, existing methods are mostly based on Bayesian recursive estimation, employing linear / nonlinear filtering or sampling-based random filtering to achieve online estimation. Due to the high maneuverability and strong uncertainty of aerial targets, multi-model estimation frameworks are often used in engineering to achieve robust tracking under maneuver switching through parallel sub-filtering and model probability weighting. However, the key lies in model probability allocation and switching discrimination. When the network experiences degradation such as packet loss, sudden noise increases, outliers, or local sensor failures, traditional updates based on residuals and likelihoods are prone to misinterpreting sensor anomalies as maneuver changes, leading to switching lag, probability oscillations, or even filter divergence, thereby reducing distributed fusion consistency and tracking accuracy. In recent years, tracking approaches using neural networks have been introduced for temporal feature modeling or end-to-end regression, but their stability and consistency under degradation conditions remain insufficient. How to improve the cooperative tracking performance under sensor degradation environments is an urgent problem to be solved. This paper proposes a distributed tracking method that can stably discriminate maneuvers and fuse estimates at the node level, and evaluate node reliability and perform weighted fusion at the network level. Summary of the Invention
[0004] This invention provides a distributed target tracking method, system, and medium based on the Transformer interactive multi-model, to address the insufficient stability and consistency of existing technologies under degradation conditions such as network packet loss, noise spikes, outliers, or local sensor failures.
[0005] To achieve the above objectives, in a first aspect, the present invention relates to a distributed target tracking method based on a Transformer interactive multi-model approach, comprising: Step 1: Observational Data Acquisition and Transformer Auxiliary Model Probabilistic Network Window Feature Construction: Periodically obtain nodes Obtain the two-dimensional observation vector of the target.
[0006] in, and Representing nodes respectively exist Always keep the target in mind , The measured value of direction, The sampling period; Constructing differential features based on observations from adjacent frames and defining nodes. At any moment Input feature vector:
[0007] in, , , and Representing nodes respectively exist Always keep the target in mind , Measurement of direction; Let the short-time series window length be... For each node Recently The single-frame features of each frame are stacked in temporal order to form the input matrix of the Transformer auxiliary model probabilistic network. : Step 2: Establish a multi-model IMM framework and execute parallel EKF sub-filters: Assume the motion pattern of the maneuvering target is as follows: It consists of 1 candidate model, with the model index as . The motion mode switching of the maneuvering target satisfies a first-order Markov chain, and its transition matrix is defined as follows: , is a 3×3 matrix, where Indicates the candidate model Transfer to candidate model The probability of; Obtain candidate models nodes At any moment Predicted probability Covariance ; Step 3: Fusion of TAMPN model probability calculation and node-level estimation: Step 31: Place the node At any moment Input feature vector Perform linear embedding and then superimpose positional encoding to obtain the encoder input sequence; for each node Recently The single-frame features of the frames are stacked in chronological order to obtain the encoder input matrix of the window; Step 32: TAMPN's encoder is... Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries, keys, and values respectively; Based on the generated query, key, and value, for the first... Each attention head is used to compute the scaled dot product attention, where... ; The outputs of the scaled dot product attention described in each head are concatenated and linearly mapped to obtain the output of the attention sublayer: Both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. Stack the residual connections with the layer normalized results. The layer obtains the final output of the encoder. .
[0008] Step 33: Perform average pooling on the time dimension to obtain a fixed-length representation:
[0009] in, For at any time From the matrix Selected from All columns of the row are linearly mapped to obtain the candidate models. :
[0010] in , The parameters are learnable, and the final output model posterior probability vector is obtained through softmax.
[0011] The probabilities of each candidate model are
[0012] During the training phase, the TAMPN model constructs a training set. ,in To represent the supervision labels for the real maneuver model categories, the cross-entropy loss function is used:
[0013] in, For the model to the first Each node at time... The class probability; Step 34: Obtain the node from Step 2 At any moment posterior estimation With covariance The model probabilities output in step 3.3 are used to weight and fuse the results of each candidate model to obtain a node-level single-state estimate. With covariance :
[0014] in Thus, the node is obtained. Local output ; Step 4: Includes: Step 41 will node Local posterior estimation With covariance Convert to information matrix and information vector, and set initial values for consistent iteration; Step 42: Let the fundamental weight matrix corresponding to the communication topology be... , No. During round iteration, nodes The instantaneous state estimate is as follows:
[0015] Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights, and nodes are defined. In the The neighborhood reference residual of the wheel is used, and a Huber-type suppression function is introduced to weaken outlier biases, using a length of... Sliding window statistics definition node At any moment observation loss rate The model confidence level is taken as the maximum posterior probability output in step 3.3, and finally the confidence level is obtained at the 1st step. Round-by-round iteration to construct node reliability; Step 43: Based on the fact that the reliability of a node is proportional to its contribution, for edges... By constructing symmetric edge scores, a double-random consistency weight matrix is obtained using symmetric normalization. : Step 44 node According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector. Based on this step, the information consistency iteration is completed for a preset number of rounds to obtain the node. The fusion posterior estimate.
[0016] Preferably, when ≥ hour, ,in, for The probability network input at time t, for The probability network input at time t, for The probability network input at time t, when When the model probability is initialized using a given initial probability vector, the model probability is initialized.
[0017] Preferably, candidate models are obtained in step 2. Predicted probability Covariance Specifically, it includes: Step 21: Interaction phase of IMM: Calculating candidate models Prior probability
[0018] Based on the prior probability, calculate from arrive Mixed weights The initial conditions of each sub-filter are mixed using a mixing weight, where... Indicates the first Each sensor node at time For the model The prior probability, Indicates the first Each sensor node at time For the model The posterior probability; Set nodes At any moment The The posterior estimates of the candidate models are: covariance is Then the candidate model The initial values for the mixture are:
[0019] in Indicates at time , No. Each sensor node selected the model Post-state estimation; The mixed covariance is:
[0020] in , Indicates time Next, node After selecting the model The covariance matrix after that, In the model The covariance matrix under the following conditions Representation Model With model Differences in state estimation between them For nodes At any moment For the model Posterior state estimation, For nodes At any moment For the model Initial mixing values at the input of the sub-filter; Step 22: For each model Establish state equations With observation equation For candidate models The predicted probabilities and covariances are:
[0021] in, Represents the state transition function. Represents the state transition matrix. Represents the process noise covariance matrix; When node At any moment Obtain observation At that time, for each model Calculate residuals : Based on residual covariance, observation Jacobian matrix, and model Predicting residual covariance Based on residual covariance and the model Predictive calculation of Kalman gain Based on the Kalman gain and each candidate model Candidate models based on residuals Update the predicted probabilities and covariance; Node At any moment Input feature vector Linear embedding followed by positional encoding yields the encoder input sequence:
[0022] in , For learnable parameters, For position encoding vectors, For model dimensions; based on length of Short time series window, for each node Recently The single-frame features of each frame are stacked in chronological order to obtain the encoder input matrix for the window: ; TAMPN's encoder is Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries respectively ,key Sum :
[0023] in, For the first in the Transformer network Layers are used to generate queries. The weight matrix, For the first in the Transformer network Layers are used to generate keys The weight matrix, For the first in the Transformer network Layers are used to generate values The weight matrix; For the A person's attention Calculate the scaled dot product attention:
[0024] in For key dimensions and query dimensions, Indicates the first The node at the th Layer to the first A query matrix with attention heads The key matrix, It is a value matrix; The outputs of each head are concatenated and linearly mapped to obtain the output of the attention sublayer:
[0025] in These are learnable parameters; both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization.
[0026] in Indicates the first Layered feedforward network, The layer normalization function is represented. Represents a node At any moment Enter the number The feature sequence of the layer, Indicates the first The output of a multi-head self-attention sublayer. For intermediate features after the attention sublayer, stacking The layer obtains the final output of the encoder. .
[0027] Preferably, step 41 specifically comprises: Node Local posterior estimation With covariance Convert to an information matrix and information vector, and set the initial value for the consistency iteration to: ,in, Represents a node At any moment The initialization information matrix, Represents a node At any moment The initialization information vector, Represents a node The original information pairs are directly calculated from the local filtering results. Step 42 specifically involves: Let the fundamental weight matrix corresponding to the communication topology be... , ,
[0028] Satisfies topological consistency and non-negativity, and ; No. During round iteration, nodes The instantaneous state estimate is as follows:
[0029] Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights:
[0030] in, These are elements of the weight matrix, representing nodes. When constructing neighborhood references, nodes The weighting coefficients of the information, For nodes In the Information matrix during round iteration For nodes In the Information vector during round iteration, For nodes At any moment , No. The neighborhood reference information matrix constructed by the wheel. For nodes At any moment , No. The neighborhood reference information vector constructed by the wheel, For nodes At any moment , No. Neighborhood reference state estimation for wheel construction; Define nodes In the The neighborhood reference residual of the wheel is:
[0031] in, For nodes At any moment , No. The state estimate obtained from rounds of iteration, For nodes At any moment , No. Neighborhood reference state estimation of wheel structure, For nodes In the The information matrix during round iterations is used, and a Huber-type suppression function is introduced to reduce outlier bias:
[0032] in For threshold; Define nodes At any moment observation loss rate Using a length of Sliding window statistics:
[0033] in To indicate whether an observation was received, let 1 represent received and 0 represent lost. The model confidence level is the maximum posterior probability output in step 3.3.
[0034] in, Represents a node At any moment The target is determined to belong to the first The probability of the candidate model is finally obtained at the th . Reliability of nodes constructed through round-iteration construction:
[0035] in To adjust the index, As a lower bound for reliability, It is a residual suppression factor.
[0036] Preferably, step 43 specifically includes: Define a symmetric fusion function:
[0037] in, For nodes At any moment , No. Reliability of round iteration, For nodes At any moment , No. The reliability of round iterations is assessed, and edge scoring is defined based on this:
[0038] in, and define nodes Ratings and:
[0039] Constructing a double random consistency weight matrix using symmetric normalization It is a 9x9 matrix. Represents a node In the During round fusion, nodes Information weighting coefficients: .
[0040] Preferably, step 44 specifically includes: In the Round iteration, node According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector:
[0041] in . Represents a node In the During round fusion, nodes The weighting coefficients of the information, For nodes The self-weight, the number of iterations is finite. ,in This satisfies the trade-off between communication overhead and accuracy; in completing After rounds of information consistency iteration, the nodes The fusion posterior estimate is: ,in For nodes At any moment go through The information matrix after rounds of consistency iterations For nodes At any moment go through The information vector after each round of consistency iteration For nodes At any moment Fusion of posterior state estimates, For nodes At any moment Fusion posterior covariance matrix.
[0042] To achieve the above objectives, in a second aspect, the present invention relates to a distributed target tracking system based on a Transformer interactive multi-model architecture, comprising: The observation data acquisition and feature construction module is used for: Periodically obtain nodes Obtain the two-dimensional observation vector of the target.
[0043] in, and Representing nodes respectively exist Always keep the target in mind , The measured value of direction, The sampling period; Constructing differential features based on observations from adjacent frames and defining nodes. At any moment Input feature vector:
[0044] in, , , and Representing nodes respectively exist -1 time for the target , The measured value of the direction; let the short-time series window length be... For each node Recently The single-frame features of each frame are stacked in temporal order to form the input matrix of the Transformer auxiliary model probabilistic network. : The framework establishment and filtering module is used to establish a multi-model IMM framework and perform parallel EKF sub-filters. Assume the motion pattern of the maneuvering target is as follows: It consists of 1 candidate model, with the model index as . The motion mode switching of the maneuvering target satisfies a first-order Markov chain, and its transition matrix is defined as follows: , is a 3×3 matrix, where Indicates the candidate model Transfer to candidate model The probability of obtaining candidate models nodes At any moment Predicted probability Covariance ; The TAMPN model probability calculation and node-level estimation fusion module includes an encoder input submodule and a stacking submodule. The layer encoder output submodule, the TAMPN model training phase submodule, and the weighted fusion submodule; The encoder input submodule is used to input nodes At any moment Input feature vector Perform linear embedding and then superimpose positional encoding to obtain the encoder input sequence; for each node Recently The single-frame features of the frames are stacked in chronological order to obtain the encoder input matrix of the window; The stack Layer encoder output submodule: The encoder for TAMPN consists of... Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries, keys, and values respectively; based on the generated queries, keys, and values, perform operations on the first... Each attention head is used to compute the scaled dot product attention, where... The outputs of the scaled dot product attention described in each head are concatenated and linearly mapped to obtain the output of the attention sublayer. Both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. The results of the residual connections and layer normalization are stacked. The layer obtains the final output of the encoder. ; The TAMPN model, in its training phase submodule, is used for: Average pooling is applied to the time dimension to obtain a fixed-length representation:
[0045] in, For at any time From the matrix Selected from All columns of the row are linearly mapped to obtain the candidate models. :
[0046] in , The parameters are learnable, and the final output model posterior probability vector is obtained through softmax.
[0047] The probabilities of each candidate model are
[0048] During the training phase, the TAMPN model constructs a training set. ,in To represent the supervision labels for the real maneuver model categories, the cross-entropy loss function is used:
[0049] in, The model represents the first Each node at time... The class probability; The weighted fusion submodule is used to obtain the node from step 2. At any moment posterior estimation With covariance The model probabilities output in step 3.3 are used to weight and fuse the results of each candidate model to obtain a node-level single-state estimate. With covariance :
[0050] in , obtain node Local output ; The reliability-weighted information consistency fusion module includes a transformation and iteration initial value submodule, an iterative construction node reliability submodule, a consistency weight matrix construction submodule, and a fusion posterior estimation submodule. The transformation and iteration initial value submodule is used to convert nodes Local posterior estimation With covariance Convert to information matrix and information vector, and set initial values for consistent iteration; The iterative construction of node reliability submodule is used to set the basic weight matrix corresponding to the communication topology as follows: , No. During round iteration, nodes The instantaneous state estimate is as follows:
[0051] Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights, and nodes are defined. In the The neighborhood reference residual of the wheel is used, and a Huber-type suppression function is introduced to weaken outlier biases, using a length of... Sliding window statistics definition node At any moment observation loss rate The model confidence level is taken as the maximum posterior probability output in step 3.3, and finally the confidence level is obtained at the 1st step. Round-by-round iteration to construct node reliability; The submodule for constructing the consistency weight matrix is used to assign weights to edges based on the principle that the reliability of a node is proportional to its contribution. By constructing symmetric edge scores, a double-random consistency weight matrix is obtained using symmetric normalization. : The fusion posterior estimation submodule is used for nodes According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector. Based on this step, the information consistency iteration is completed for a preset number of rounds to obtain the node. The fusion posterior estimate.
[0052] To achieve the above objectives, in a third aspect, the present invention also relates to a computer-readable storage medium storing instructions that, when executed, perform the aforementioned distributed target tracking method based on a Transformer interactive multi-model.
[0053] The present invention relates to a distributed target tracking method, system, and medium based on Transformer interactive multi-model, which has the following advantages compared with the prior art: This invention provides a distributed maneuvering target tracking method (TIMM-RWIC) based on Transformer-enhanced Interactive Multi-Model (TIMM) and reliability-weighted information consensus (RWIC). This distributed tracking method can stably identify and fuse maneuvers at the node side and evaluate node reliability and perform weighted fusion at the network side. By improving the timeliness and stability of maneuver pattern identification at the node side and suppressing the spread of unreliable node information at the network side, it achieves stable and robust distributed target state estimation under degradation scenarios such as measurement packet loss, noise surges, and local sensor failures, thereby improving cooperative tracking performance under sensor degradation environments. Attached Figure Description
[0054] Figure 1 This is a TAMPN model structure diagram of a distributed target tracking method based on Transformer interactive multi-model in Example 1; Figure 2 The flowchart of the TIMM-RWIC algorithm, a distributed target tracking method based on Transformer interactive multi-model, is shown in Example 1. Figure 3 This is a diagram showing the expected trajectory of the target and the distribution of sensor node positions in Example 1 of a distributed target tracking method based on a Transformer interactive multi-model, as described in Embodiment 1. Figure 4 This is a distributed sensor network communication topology diagram of Example 1 of a distributed target tracking method based on Transformer interactive multi-model in Embodiment 1. Figure 5 This is a confusion matrix diagram of Example 1 of a distributed target tracking method based on Transformer interactive multi-model in Embodiment 1; Figure 6 This is a model probability curve diagram of Example 1 of a distributed target tracking method based on Transformer interactive multi-model in Embodiment 1; Figure 7 This is the position / velocity RMSE diagram of Example 1 of a distributed target tracking method based on Transformer interactive multi-model in Embodiment 1; Figure 8 This is a TIMM-RWIC fusion network-level RMSE and single-node local RMSE graph for Example 1 of a distributed target tracking method based on Transformer interactive multi-model in Embodiment 1. Figure 9This is an equivalent experimental scenario diagram of the multi-robot collaborative platform flight, which is an example of a distributed target tracking method based on a Transformer interactive multi-model in Example 1 of Embodiment 1. Figure 10 This is a flight equivalent experiment result diagram of Example 1 of a distributed target tracking method based on a Transformer interactive multi-model in Embodiment 1; Figure 11 This is a schematic diagram of the structure of a distributed target tracking system based on a Transformer interactive multi-model in Embodiment 2. Detailed Implementation
[0055] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention and not the entire structure.
[0056] Example 1 A distributed target tracking method based on Transformer interactive multi-model, please refer to [link / reference]. Figure 1-2 As shown, the present invention provides a distributed target tracking method based on Transformer interactive multi-model, comprising the following steps 1 to 4.
[0057] Step 1: Observational data acquisition and construction of window features for the Transformer auxiliary model probabilistic network.
[0058] In a distributed sensor network, let there be a total of There are 1 observation node, with node number as... At discrete time (Sampling period is) When ), node The two-dimensional observation vector of the target is denoted as .
[0059] (1) in, and Representing nodes respectively At any moment For the target , The measured value of the direction.
[0060] To characterize the local motion changes of the target within a short time window and reduce the impact of measurement bias on pattern discrimination, this invention constructs differential features based on observations from adjacent frames and defines nodes. At any moment Input feature vector: (2) in, , .
[0061] Let the short-time series window length be... For each node Recently The single-frame features of each frame are stacked in temporal order to form the input matrix of the Transformer-aided model probability network (TAMPN). : (3) when When the model probability is initialized using a given initial probability vector, the model probability is initialized.
[0062] Step 2: Establish a multi-model IMM framework and execute parallel EKF sub-filters Assume the motion pattern of the maneuvering target is as follows: It consists of 1 candidate model, with the model index as . The mode switching follows a first-order Markov chain, and its transition matrix is defined as follows: , is a 3×3 matrix, where Indicated by model Transfer to model The probability of.
[0063] Step 21: The interaction phase of IMM first calculates the model. Prior probability (predicted probability): (4) Then calculate from Mixed weights (conditional mixture probabilities): (5) The initial conditions of each sub-filter are mixed using a mixing weight, where... Indicates the first Each sensor node at time For the model The prior probability, Indicates the first Each sensor node at time For the model The posterior probability. Let node... At any moment The The posterior estimate of the model is: covariance is Then the model The initial values for the mixture are: (6) in Indicates at time , No. Each sensor node selected the model Post-state estimation; The mixed covariance is: (7) in , Indicates time Next, node After selecting the model The covariance matrix after that, In the model The covariance matrix under the following conditions Representation Model With model Differences in state estimation between them For nodes At any moment For the model Posterior state estimation, For nodes At any moment For the model The initial mixed values of the sub-filter input.
[0064] Step 22: For each model Establish state equations and observation equations (8) in , For the model The prediction is: (9) in, Represents the state transition function. Represents the state transition matrix. This represents the process noise covariance matrix.
[0065] When node At any moment Obtain observation At that time, for each model Calculate the residuals: (10) in To observe the Jacobian matrix, the residual covariance is: (11) The Kalman gain is: (12) State update and covariance update are as follows: (13) node At any moment If observations are missing, this embodiment can adopt a "predict only, do not update" strategy: (14) Step 3: Fusion of TAMPN model probability calculation and node-level estimation: Step 31 will node At any moment Input feature vector Linear embedding and then superimposed positional encoding are performed to obtain the encoder input sequence. (15) in , For learnable parameters, For position encoding vectors, For model dimensions. Stacked to obtain (16) Step 32: TAMPN's encoder is... It consists of layers of Transformers. For the first... layer Given the output of the previous layer Generate queries, keys, and values respectively: (17) in, For the first in the Transformer network Layers are used to generate queries. The weight matrix, For the first in the Transformer network Layers are used to generate keys The weight matrix, For the first in the Transformer network Layers are used to generate values The weight matrix; For the A person's attention Calculate the scaled dot product attention: (18) in For the key (query) dimension, Indicates the first The node at the th Layer to the first A query matrix with attention heads The key matrix, The outputs of each head are concatenated and linearly mapped to obtain the output of the attention sublayer: (19) in These are learnable parameters. Both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. (20) in Indicates the first Layered feedforward network, The layer normalization function is represented. Represents a node At any moment Enter the number The feature sequence of the layer, Indicates the first The output of a multi-head self-attention sublayer. For intermediate features after the attention sublayer, stacking The layer obtains the final output of the encoder. .
[0066] Step 33: Perform average pooling on the time dimension to obtain a fixed-length representation: (twenty one) in, For at any time From the matrix Selected from All columns of the row are linearly mapped to obtain the models for each. : (twenty two) in , The parameters are learnable, and the final output is the model's posterior probability vector through softmax: (twenty three) The probabilities of each model are (twenty four) The TAMPN model structure is as follows: Figure 1 As shown. During the training phase, a training set is constructed. ,in The supervision labels (real maneuver model categories) are used. The cross-entropy loss function is employed. (25) in, For the model to the first Each node at time... The class probability.
[0067] Step 34: Obtain the node from step 22 At any moment posterior estimation With covariance The model probabilities output in step 33 are used to weight and fuse the results of each model to obtain the node-level single-state estimates and covariance: (26) in Thus, the node is obtained. Local output This information is used for consistency fusion of reliability-weighted information in subsequent step 4.
[0068] Step 4: Reliability-weighted information consistency fusion Step 41 will node Local posterior estimation With covariance Convert it into an information matrix and information vector: (27) Set the initial value for the consistency iteration to: (28) in, Represents a node At any moment The initialization information matrix, Represents a node At any moment The initialization information vector, Represents a node The original information pairs are directly calculated from the local filtering results.
[0069] Step 42: Let the fundamental weight matrix corresponding to the communication topology be... Satisfying topological consistency and nonnegativity (if) but ),and . No. During round iteration, nodes The instantaneous state estimate is as follows: (29) Construct neighborhood reference information pairs (weighting neighborhood information with basic weights): (30) in, These are elements of the weight matrix, representing nodes. When constructing neighborhood references, nodes The weighting coefficients of the information, For nodes In the Information matrix during round iteration For nodes In the Information vector during round iteration, For nodes At any moment , No. The neighborhood reference information matrix constructed by the wheel. For nodes At any moment , No. The neighborhood reference information vector constructed by the wheel, For nodes At any moment , No. Neighborhood reference state estimation for wheel construction.
[0070] Define nodes In the The neighborhood reference residual (NRR) of the wheel is: (31) in, For nodes At any moment , No. The state estimate obtained from rounds of iteration, For nodes At any moment , No. Neighborhood reference state estimation of wheel structure, For nodes In the The information matrix during round iterations is used, and a Huber-type suppression function is introduced to reduce outlier bias: (32) in Define the threshold. Define the node. At any moment observation loss rate Using a length of Sliding window statistics: (33) in This represents the indication of whether an observation was received (1 for received, 0 for lost). The model confidence level is the maximum posterior probability output from step 3.3: (34) in, Represents a node At any moment The target is determined to belong to the first The probability of the candidate model is finally obtained at the th . Reliability of nodes constructed through round-iteration construction: (35) in To adjust the index, This serves as a lower bound for reliability, used to ensure the weak connectivity and numerical stability of the weight matrix. It is a residual suppression factor.
[0071] Step 43, to achieve the fusion principle of "reliable nodes contributing more and degraded nodes contributing less," involves processing the edges... Construct a symmetric edge score. Define a symmetric fusion function: (36) in, For nodes At any moment , No. Reliability of round iteration, For nodes At any moment , No. The reliability of round iterations is assessed, and edge scoring is defined based on this: (37) in, and define nodes Ratings and: (38) A symmetric normalization method is used to construct a double random (DS) consensus weight matrix. It is a 9x9 matrix. Represents a node In the During round fusion, nodes Information weighting coefficients: (39) This construction makes It is topologically consistent, non-negative, and symmetric, and can achieve good numerical stability and consistency fusion effect in engineering implementation.
[0072] Step 44 in Round iteration, node according to The neighborhood information is weighted and fused to update the information matrix and information vector: (40) in , Represents a node In the During round fusion, nodes The weighting coefficients of the information, For nodes The self-weight, the number of iterations is finite. This satisfies the trade-off between communication overhead and accuracy. (After completing...) After rounds of information consistency iteration, the nodes The fusion posterior estimate is: (41) in For nodes At any moment go through The information matrix after rounds of consistency iterations For nodes At any moment go through The information vector after each round of consistency iteration For nodes At any moment Fusion of posterior state estimates, For nodes At any moment Fusion posterior covariance matrix.
[0073] This concludes the TIMM-RWIC maneuvering target tracking method, and its process is as follows: Figure 2 As shown.
[0074] To better illustrate the solution of the present invention, an example is given below, such as... Figure 3 As shown: The method proposed in this invention is demonstrated through Python simulation experiments. Discrete-time system model. (42) in, Let be the system state vector. The measurement vector is defined as follows. The target state vector is defined as follows: ,in These are the planar position coordinates, To correspond to the velocity components, three motion models were adopted: constant velocity motion (CV), coordinated left turn (CTL), and coordinated right turn (CTR). The CV state transition matrix and process noise matrix are as follows:
[0075] The state transition matrix and process noise matrix of the CT model are as follows:
[0076] The total simulation duration is N = 300 seconds, the sampling time is T = 1 second, and the target motion follows a five-stage model sequence: seconds 1-50 are constant velocity (CV) motion, seconds 51-100 are constant left turn (CTL) motion, and the angular rate... rad / s, 101-200 seconds is CV motion, 201-250 seconds is constant right turn (CTR) motion, angular rate The target's trajectory is measured in rad / s, with a CV motion lasting 251-300 seconds. The positions of the nine sensor nodes are (0,0)m, (4000,0)m, (8000,0)m, (0,-6000)m, (4000,-6000)m, (8000,-6000)m, (0,-12000)m, (4000,-12000)m, and (8000,-12000)m. The expected trajectory of the target and the distribution of the sensor node positions are as follows: Figure 3 As shown, the communication topology of the distributed sensor network is as follows: Figure 4 As shown. Model probability and Markov transition matrix The initial value is defined as
[0077] The initial values of the filtering parameters are shown in Table 1.
[0078] Table 1 Filter Parameter Configuration
[0079] The Transformer network in this invention employs a lightweight encoder. Training data is derived from the LAST dataset, with corresponding model labels automatically annotated based on predefined motion patterns. We generated 120,000 trajectories for training and 30,000 trajectories for testing (10,000 trajectories for each of the three motion patterns). The Transformer network and training parameters are shown in Table 2.
[0080] Table 2 Transformer Network and Training Parameter Configuration
[0081] This invention uses root mean square error (RMSE) and average RMSE (ARMSE) to evaluate the performance of the algorithm, and their mathematical definitions are as follows:
[0082] in, Indicates the number of Monte Carlo simulations. This represents the total duration of each simulation. For the first In the second simulation, time... The true state This represents the corresponding estimated state. In this invention, the number of Monte Carlo simulations is set to... .
[0083] To quantify the motion model classification ability of TAMPN Figure 5 The confusion matrix for the test set containing 30,000 trajectories is shown, where label 0 corresponds to the CV model, label 1 to the CTL model, and label 2 to the CTR model. The recognition accuracies for each class are 98.72%, 98.23%, and 98.15%, respectively, with an average accuracy of 98.37%. The misclassification rate for all classes is less than 2%, indicating that TAMPN provides reliable motion model recognition for subsequent maneuvering target tracking.
[0084] Next, we investigate the impact of TAMPN on tracking by integrating TAMPN into the proposed TIMM filter and comparing it with three single-node baselines (IMM, ATPM-PIMM, and LSTM-IMM). The model probability curves are shown below. Figure 6 As shown, the position / velocity RMSE is as follows: Figure 7 As shown. Figure 6 As shown, both TIMM and LSTM-IMM can quickly drive the posterior probability of the true mode to a maneuver transition close to 1, while keeping the non-true mode at a low probability; among them, TIMM exhibits a smoother oscillating probability trajectory. In contrast, IMM shows significant lag and more frequent increases in the probability of the non-true mode during the transition interval, while ATPM-PIMM exhibits short-term oscillations and transient spikes over several cycles, reducing the stability of mode selection. These results indicate that learning-based model selection utilizing temporal context improves the robustness of mode recognition, while traditional instantaneous likelihood normalization is more prone to producing unstable probability estimates under complex maneuvers and noise perturbations. Consistently, Figure 7 This shows that other IMM-based algorithms produce increasingly large estimation errors because they fail to adequately track model probabilities.
[0085] Table 3 summarizes the single-node errors. The proposed TIMM achieves state-of-the-art performance on two metrics: position ARMSE of 16.0170 m and velocity ARMSE of 2.7558 m / s, representing reductions of approximately 19.4% and 66.6% respectively compared to IMM. The performance improvement primarily stems from TIMM's temporal-context modeling, which allows for faster convergence to the correct mode during transformations and maintains a low probability of non-true modes, thereby reducing estimation errors.
[0086] Table 3 Comparison of ARMSE for Position and Velocity
[0087] Figure 8 The fused network-level RMSE and the local RMSE of individual nodes for TIMM-RWIC are presented. The disturbance time for sensor 2 is 20–90 seconds, for sensor 3 it is 70–150 seconds, for sensor 4 it is 100–180 seconds, for sensor 5 it is 120–200 seconds, for sensor 7 it is 140–220 seconds, and for sensor 8 it is 200–280 seconds. These disturbance ranges reduce the corresponding local estimates, leading to a significant increase in the local RMSE curve. However, RWIC suppresses corrupted information through reliability-weighted consensus and utilizes information from reliable neighbors, thus limiting the spread of errors in the network. As a result, despite continuous perturbations, the fused network-level RMSE remains small across the entire horizon, demonstrating the robustness of TIMM-RWIC in sensor-degraded networks.
[0088] To verify the practical application value of the proposed framework, such as Figure 9 As shown, we conducted a flight equivalence experiment using a multi-robot collaborative platform. We built a collaborative target tracking experimental platform consisting of four robots and one mobile UAV. Each robot was equipped with an ultra-wideband (UWB) system to collect target-related data, while the actual target position was provided by an optical motion capture system. The integration and data synchronization of the entire system were achieved based on the Robot Operating System (ROS). The main experimental results are as follows: Figure 10 As shown: the left column corresponds to maneuver scenario 1, and the right column corresponds to maneuver scenario 2. In each column, the first row displays the estimated values and actual target trajectories of the TIMM-RWIC and DEIF-IMM algorithms; the middle row presents the relative distance measurements between the target and the four robots; and the last row compares the RMSE trends over time. It can be seen that the RMSE value of the TIMM-RWIC algorithm is the smallest among the two distributed algorithms, indicating that the TIMM-RWIC algorithm's estimation results are superior.
[0089] Example 2 A distributed target tracking system based on the Transformer interactive multi-model architecture is implemented in hardware using an electronic device with a central processing unit. It can be implemented on a personal computer, smart terminal, local area network, server, etc. For implementation details in this example, please refer to [link to relevant documentation]. Figure 11 It includes an observation data acquisition and feature construction module 61, a framework establishment and filtering module 62, a TAMPN model probability calculation and node-level estimation fusion module 63, and a reliability weighted information consistency fusion module 64.
[0090] The observation data acquisition and feature construction module 61 is used for: Periodically obtain nodes Obtain the two-dimensional observation vector of the target.
[0091] in, and Representing nodes respectively exist Always keep the target in mind , The measured value of direction, The sampling period; Constructing differential features based on observations from adjacent frames and defining nodes. At any moment Input feature vector:
[0092] in, , , and Representing nodes respectively exist Always keep the target in mind , The measured value of the direction; let the short-time series window length be... For each node Recently The single-frame features of each frame are stacked in temporal order to form the input matrix of the Transformer auxiliary model probabilistic network. : The framework establishment and filtering module 62 is used to establish a multi-model IMM framework and perform parallel EKF sub-filters. Assume the motion pattern of the maneuvering target is as follows: It consists of 1 candidate model, with the model index as . The motion mode switching of the maneuvering target satisfies a first-order Markov chain, and its transition matrix is defined as follows: , is a 3×3 matrix, where Indicates the candidate model Transfer to candidate model The probability of obtaining candidate models nodes At any moment Predicted probability Covariance ; The TAMPN model probability calculation and node-level estimation fusion module 63 includes an encoder input submodule 631 and a stack. The layer encoder output submodule 632, the TAMPN model training phase submodule 633, and the weighted fusion submodule 634; The encoder input submodule 631 is used to input nodes At any moment Input feature vector Perform linear embedding and then superimpose positional encoding to obtain the encoder input sequence; for each node Recently The single-frame features of the frames are stacked in chronological order to obtain the encoder input matrix of the window; The stack Layer encoder output submodule 632: The encoder for TAMPN is composed of... Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries, keys, and values respectively; based on the generated queries, keys, and values, perform operations on the first... Each attention head is used to compute the scaled dot product attention, where... The outputs of the scaled dot product attention described in each head are concatenated and linearly mapped to obtain the output of the attention sublayer. Both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. The results of the residual connections and layer normalization are stacked. The layer obtains the final output of the encoder. ; The TAMPN model, in its training phase submodule 633, is used for: Average pooling is applied to the time dimension to obtain a fixed-length representation:
[0093] in, For at any time From the matrix Selected from All columns of the row are linearly mapped to obtain the candidate models. :
[0094] in , The parameters are learnable, and the final output is the model's posterior probability vector through softmax:
[0095] The probabilities of each candidate model are
[0096] During the training phase, the TAMPN model constructs a training set. ,in To represent the supervision labels for the real maneuver model categories, the cross-entropy loss function is used:
[0097] in, The model represents the first Each node at time... The class probability; The weighted fusion submodule 634 is used to obtain nodes from the framework establishment and filtering module 62. At any moment posterior estimation With covariance The model probabilities output by submodule 633 of the TAMPN model during the training phase are used to weight and fuse the results of each candidate model to obtain a node-level single-state estimate. With covariance :
[0098] in , obtain node Local output ; The reliability-weighted information consistency fusion module 64 includes a transformation and iteration initial value submodule 641, an iterative node reliability construction submodule 642, a consistency weight matrix construction submodule 643, and a fusion posterior estimation submodule 644. The transformation and iteration initial value submodule 641 is used to convert the nodes Local posterior estimation With covariance Convert to information matrix and information vector, and set initial values for consistent iteration; The iteratively constructed node reliability submodule 642 is used to set the basic weight matrix corresponding to the communication topology as follows: , No. During round iteration, nodes The instantaneous state estimate is as follows:
[0099] Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights, and nodes are defined. In the The neighborhood reference residual of the wheel is used, and a Huber-type suppression function is introduced to weaken outlier biases, using a length of... Sliding window statistics definition node At any moment observation loss rate The model confidence level is taken as the maximum posterior probability output in step 3.3, and finally the confidence level is obtained at the 1st step. Round-by-round iteration to construct node reliability; The constructed consistency weight matrix submodule 643 is used to construct edges based on the principle that the reliability of a node is proportional to its contribution. By constructing symmetric edge scores, a double-random consistency weight matrix is obtained using symmetric normalization. : The fusion posterior estimation submodule 644 is used for nodes According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector. Based on this step, the information consistency iteration is completed for a preset number of rounds to obtain the node. The fusion posterior estimate.
[0100] The distributed target tracking system based on the Transformer interactive multi-model described in this embodiment has the same implementation process, method and effect as the distributed target tracking method based on the Transformer interactive multi-model described in Embodiment 1, and will not be repeated here.
[0101] Example 3 This invention relates to a computer-readable storage medium storing instructions. When the instructions are executed, they perform a distributed target tracking method based on a Transformer interactive multi-model, as described in Embodiment 1. The runtime implementation process and effects are the same as those of the distributed target tracking method based on a Transformer interactive multi-model described in Embodiment 1, and will not be repeated here.
[0102] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0103] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A distributed target tracking method based on Transformer interactive multi-model, characterized in that, include: Step 1: Observational Data Acquisition and Transformer Auxiliary Model Probabilistic Network Window Feature Construction: Periodic acquisition Point to obtain the two-dimensional observation vector of the target ; in, and Representing nodes respectively exist Always keep the target in mind , The measured value of direction, The sampling period; Constructing differential features based on observations from adjacent frames and defining nodes. At any moment Input feature vector: , in, , , and Representing nodes respectively exist Always keep the target in mind , Measurement of direction; Let the short-time series window length be... For each node Recently The single-frame features of each frame are stacked in temporal order to form the input matrix of the Transformer auxiliary model probabilistic network. : Step 2: Establish a multi-model IMM framework and execute parallel EKF sub-filters: Assume the motion pattern of the maneuvering target is as follows: It consists of 1 candidate model, with the model index as . The motion mode switching of the maneuvering target satisfies a first-order Markov chain, and its transition matrix is defined as follows: , is a 3×3 matrix, where Indicates the candidate model Transfer to candidate model The probability of; Obtain candidate models nodes At any moment Predicted probability Covariance ; Step 3: Fusion of TAMPN model probability calculation and node-level estimation: Step 31: Place the node At any moment Input feature vector Perform linear embedding and then superimpose positional encoding to obtain the encoder input sequence; for each node Recently The single-frame features of the frames are stacked in chronological order to obtain the encoder input matrix of the window; Step 32: TAMPN's encoder is... Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries, keys, and values respectively; Based on the generated query, key, and value, for the first... Each attention head is used to compute the scaled dot product attention, where... ; The outputs of the scaled dot product attention described in each head are concatenated and linearly mapped to obtain the output of the attention sublayer; Both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. Stack the residual connections with the layer normalized results. The layer obtains the final output of the encoder. ; Step 33: Perform average pooling on the time dimension to obtain a fixed-length representation: , in, For at any time From the matrix Selected from All columns of the row are linearly mapped to obtain the candidate models. : , in , The parameters are learnable, and the final output model posterior probability vector is obtained through softmax. ; The probabilities of each candidate model are ; During the training phase, the TAMPN model constructs a training set. ,in To represent the supervision labels for the real maneuver model categories, the cross-entropy loss function is used: , in, For the model to the first Each node at time... The class probability; Step 34: Obtain the node from Step 2 At any moment posterior estimation With covariance The model probabilities output in step 3.3 are used to weight and fuse the results of each candidate model to obtain a node-level single-state estimate. With covariance : , in Thus, the node is obtained. Local output ; Step 4: Includes: Step 41 will node Local posterior estimation With covariance Convert to information matrix and information vector, and set initial values for consistent iteration; Step 42: Let the fundamental weight matrix corresponding to the communication topology be... , No. During round iteration, nodes The instantaneous state estimate is as follows: ; Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights, and nodes are defined. In the The neighborhood reference residual of the wheel is used, and a Huber-type suppression function is introduced to weaken outlier biases, using a length of... Sliding window statistics definition node At any moment observation loss rate The model confidence level is taken as the maximum posterior probability output in step 3.3, and finally the confidence level is obtained at the 1st step. Round-by-round iteration to construct node reliability; Step 43: Based on the fact that the reliability of a node is proportional to its contribution, for edges... By constructing symmetric edge scores, a double-random consistency weight matrix is obtained using symmetric normalization. : Step 44 node According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector. Based on this step, the information consistency iteration is completed for a preset number of rounds to obtain the node. The fusion posterior estimate.
2. The distributed target tracking method based on Transformer interactive multi-model as described in claim 1, characterized in that, when ≥ hour, ,in, for The probability network input at time t, for The probability network input at time t, for The probability network input at time t, when When the model probability is initialized using a given initial probability vector, the model probability is initialized.
3. The distributed target tracking method based on Transformer interactive multi-model as described in claim 1, characterized in that, Obtain candidate models in step 2 Predicted probability Covariance Specifically, it includes: Step 21: Interaction phase of IMM: Calculating candidate models Prior probability , Based on the prior probability, calculate from arrive Mixed weights The initial conditions of each sub-filter are mixed using a mixing weight, where... Indicates the first Each sensor node at time For the model The prior probability, Indicates the first Each sensor node at time For the model The posterior probability; Set nodes At any moment The The posterior estimates of the candidate models are: covariance is Then the candidate model The initial mixed values are: , in Indicates at time , No. Each sensor node selected the model Post-state estimation; The mixed covariance is: , in , Indicates time Next, node After selecting the model The covariance matrix after that, In the model The covariance matrix under the following conditions Representation Model With model Differences in state estimation between them For nodes At any moment For the model Posterior state estimation, For nodes At any moment For the model Initial mixing values at the input of the sub-filter; Step 22: For each model Establish state equations x(k) With observation equation z j (k) For candidate models The predicted probabilities and covariances are: , in, Represents the state transition function. Represents the state transition matrix. Represents the process noise covariance matrix; When node At any moment Obtain observation At that time, for each model Calculate residuals : The residual covariance is calculated based on the residual covariance, the observation Jacobian matrix, and the prediction of model m. The Kalman gain is calculated based on the residual covariance and the prediction of model m. Based on the Kalman gain and each candidate model Candidate models based on residuals Update the predicted probabilities and covariance; Node At any moment Input feature vector Linear embedding followed by positional encoding yields the encoder input sequence: in, , For learnable parameters, For position encoding vectors, For model dimensions; based on length of Short time series window, for each node Recently The single-frame features of each frame are stacked in chronological order to obtain the encoder input matrix for the window: ; TAMPN's encoder is Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries respectively ,key Sum : ; in, For the first in the Transformer network Layers are used to generate queries. The weight matrix, For the first in the Transformer network Layers are used to generate keys The weight matrix, For the first in the Transformer network Layers are used to generate values The weight matrix; For the A person's attention Calculate the scaled dot product attention: , in For key dimensions and query dimensions, Indicates the first The node at the th Layer to the first A query matrix with attention heads The key matrix, It is a value matrix; The outputs of each head are concatenated and linearly mapped to obtain the output of the attention sublayer: , in These are learnable parameters; both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. , in Indicates the first Layered feedforward network, The layer normalization function is represented. Represents a node At any moment Enter the number The feature sequence of the layer, Indicates the first The output of a multi-head self-attention sublayer. For intermediate features after the attention sublayer, stacking The layer obtains the final output of the encoder. .
4. The distributed target tracking method based on Transformer interactive multi-model according to claim 1, characterized in that, Step 41 is as follows: Node Local posterior estimation With covariance Convert to an information matrix and information vector, and set the initial values for the consistency iteration to: ,in, Represents a node At any moment The initialization information matrix, Represents a node At any moment The initialization information vector, Represents a node The original information pairs obtained directly from the local filtering results; Step 42 specifically involves: Let the fundamental weight matrix corresponding to the communication topology be... , , , Satisfies topological consistency and non-negativity, and ; No. During round iteration, nodes The instantaneous state estimate is as follows: , Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights: ; in, These are elements of the weight matrix, representing nodes. When constructing neighborhood references, nodes The weighting coefficients of the information, For nodes In the Information matrix during round iteration For nodes In the Information vector during round iteration, For nodes At any moment , No. The neighborhood reference information matrix constructed by the wheel. For nodes At any moment , No. The neighborhood reference information vector constructed by the wheel, For nodes At any moment , No. Neighborhood reference state estimation for wheel construction; Define nodes In the The neighborhood reference residual of the wheel is: , in, For nodes At any moment , No. The state estimate obtained from rounds of iteration, For nodes At any moment , No. Neighborhood reference state estimation of wheel structure, For nodes In the The information matrix during round iterations is used, and a Huber-type suppression function is introduced to reduce outlier bias: ; in For the threshold; Define nodes At any moment observation loss rate Using a length of Sliding window statistics: , in To indicate whether an observation was received, let 1 represent received and 0 represent lost. The model confidence level is the maximum posterior probability output in step 3.
3. , in, Represents a node At any moment The target is determined to belong to the first The probability of the candidate model is finally obtained at the th . Reliability of nodes constructed through round-iteration construction: , in To adjust the index, As the lower bound of reliability, It is a residual suppression factor.
5. A distributed target tracking method based on Transformer interactive multi-model according to claim 4, characterized in that, Step 43 specifically involves: Define a symmetric fusion function: , in, For nodes At any moment , No. Reliability of round iteration, For nodes At any moment , No. The reliability of round iterations is assessed, and edge scoring is defined based on this: , in, and define nodes Ratings and: , Constructing a double random consistency weight matrix using symmetric normalization It is a 9x9 matrix. Represents a node In the During round fusion, nodes Information weighting coefficients: 。 6. The distributed target tracking method based on Transformer interactive multi-model according to claim 5, characterized in that, Step 44 is as follows: In the Round iteration, node According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector: ; in , Represents a node In the During round fusion, nodes The weighting coefficients of the information, For nodes The self-weight, the number of iterations is finite. ,in This satisfies the trade-off between communication overhead and accuracy; in completing After rounds of information consistency iteration, the nodes The fusion posterior estimate is: ,in For nodes At any moment go through The information matrix after rounds of consistency iterations For nodes At any moment go through The information vector after each round of consistency iteration For nodes At any moment Fusion of posterior state estimates, For nodes At any moment Fusion posterior covariance matrix.
7. A distributed target tracking system based on Transformer interactive multi-model, characterized in that, include: The observation data acquisition and feature construction module is used for: Periodically obtain nodes Obtain the two-dimensional observation vector of the target. , in, and Representing nodes respectively exist Always keep the target in mind , The measured value of direction, The sampling period; Constructing differential features based on observations from adjacent frames and defining nodes. At any moment Input feature vector: , in, , , and Representing nodes respectively exist Always keep the target in mind , The measured value of the direction; let the short-time series window length be... For each node Recently The single-frame features of each frame are stacked in temporal order to form the input matrix of the Transformer auxiliary model probabilistic network. : The framework establishment and filtering module is used to establish a multi-model IMM framework and perform parallel EKF sub-filters. Assume the motion pattern of the maneuvering target is as follows: It consists of 1 candidate model, with the model index as . The motion mode switching of the maneuvering target satisfies a first-order Markov chain, and its transition matrix is defined as follows: , is a 3×3 matrix, where Indicates the candidate model Transfer to candidate model The probability of obtaining candidate models nodes At any moment Predicted probability Covariance ; The TAMPN model probability calculation and node-level estimation fusion module includes an encoder input submodule and a stacking submodule. The layer encoder output submodule, the TAMPN model training phase submodule, and the weighted fusion submodule; The encoder input submodule is used to input nodes At any moment Input feature vector Perform linear embedding and then superimpose positional encoding to obtain the encoder input sequence; for each node Recently The single-frame features of the frames are stacked in chronological order to obtain the encoder input matrix of the window; The stack Layer encoder output submodule: The encoder for TAMPN consists of... Composed of layers of Transformer, for the first Layer, in which Given the output of the previous layer Generate queries, keys, and values respectively; based on the generated queries, keys, and values, perform operations on the first... Each attention head is used to compute the scaled dot product attention, where... The outputs of the scaled dot product attention described in each head are concatenated and linearly mapped to obtain the output of the attention sublayer. Both the attention sublayer and the feedforward sublayer employ residual connections and layer normalization. The results of the residual connections and layer normalization are stacked. The layer obtains the final output of the encoder. ; The TAMPN model, in its training phase submodule, is used for: Average pooling is applied to the time dimension to obtain a fixed-length representation: ; in, For at any time From the matrix Selected from All columns of the row are linearly mapped to obtain the candidate models. : , in , The parameters are learnable, and the final output is the model's posterior probability vector through softmax: , The probabilities of each candidate model are , During the training phase, the TAMPN model constructs a training set. ,in To represent the supervision labels for the real maneuver model categories, the cross-entropy loss function is used: , in, The model represents the first Each node at time... The class probability; The weighted fusion submodule is used to obtain the node from step 2. At any moment posterior estimation With covariance The model probabilities output in step 3.3 are used to weight and fuse the results of each candidate model to obtain a node-level single-state estimate. With covariance : ; in , obtain node Local output ; The reliability-weighted information consistency fusion module includes a transformation and iteration initial value submodule, an iterative construction node reliability submodule, a consistency weight matrix construction submodule, and a fusion posterior estimation submodule. The transformation and iteration initial value submodule is used to convert nodes Local posterior estimation With covariance Convert to information matrix and information vector, and set initial values for consistent iteration; The iterative construction of node reliability submodule is used to set the basic weight matrix corresponding to the communication topology as follows: , No. During round iteration, nodes The instantaneous state estimate is as follows: , Neighborhood reference information pairs are constructed by weighting neighborhood information with basic weights, and nodes are defined. In the The neighborhood reference residual of the wheel is used, and a Huber-type suppression function is introduced to weaken outlier biases, using a length of... Sliding window statistics definition node At any moment observation loss rate The model confidence level is taken as the maximum posterior probability output in step 3.3, and finally the confidence level is obtained at the 1st step. Round-by-round iteration to construct node reliability; The submodule for constructing the consistency weight matrix is used to assign weights to edges based on the principle that the reliability of a node is proportional to its contribution. By constructing symmetric edge scores, a double-random consistency weight matrix is obtained using symmetric normalization. : The fusion posterior estimation submodule is used for nodes According to the double random consistency weight matrix The neighborhood information is weighted and fused to update the information matrix and information vector. Based on this step, the information consistency iteration is completed for a preset number of rounds to obtain the node. The fusion posterior estimate.
8. A computer-readable storage medium, characterized in that: The storage medium stores instructions that, when executed, perform a distributed target tracking method based on a Transformer interactive multi-model as described in any one of claims 1-6.