Air combat target maneuvering multi-mode trajectory prediction method and system based on space-time joint attention mechanism

Through a codec neural network based on the joint attention mechanism of space-time, space-time, joint air combat attention mechanism, the problems of large errors and physical distortion of existing air combat trajectory prediction technology under high nonlinear maneuver are solved, multimodal trajectory generation and aerodynamic constraints are realized, and the accuracy and reliability of air combat decisions are improved.

CN120470264AActive Publication Date: 2025-08-12POLIXIR TECH LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510613680.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-12
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing air combat trajectory prediction technology has large errors in the face of high nonlinear maneuver and prolongs in calculations, insufficient feature expression dimensions, insufficient multimodal output capabilities, and the generation of trajectories may violate the basic laws of aerodynamics, resulting in tactical decision-making risks.

Method used

A codec neural network based on the joint attention mechanism of space-time is adopted to generate multiple possible trajectories through translation invariance, rotation invariance, and stationary modeling, combining spatiotemporal correlation modeling and multimodal decoding, and aerodynamic penalty terms are introduced to verify the physical feasibility of the trajectory.

Benefits of technology

It improves the accuracy and reliability of trajectory prediction, can generate multimodal trajectories that meet the limits of the fighter aircraft body, reduces the risk of physical distortion, and enhances the support ability of air combat decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470264A_ABST
    Figure CN120470264A_ABST
Patent Text Reader

Abstract

The invention provides an air combat target maneuvering multi-modal trajectory prediction method and system based on a space-time joint attention mechanism. The method comprises the following steps: modeling an air combat scene; modeling translation invariance and rotation invariance; carrying out stability modeling; modeling space-time correlation; performing multi-mode decoding; and verifying feasibility. According to the method, synchronous capture of time mutation and spatial game is realized through a double-branch attention mechanism; multi-modal prediction is used to generate a plurality of possible trajectories for a target in a heterogeneous complex high-dynamic environment, and uncertainty and variability in trajectory prediction are considered; an aerodynamic penalty term is introduced into a loss function, it is ensured that the trajectory conforms to the warplane body limit, physical distortion is eliminated, and the accuracy and reliability of trajectory prediction are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data processing and relates to air combat simulation technology, and specifically to a method and system for predicting multimodal trajectories of air combat targets based on a spatiotemporal joint attention mechanism. Background Art

[0002] In air combat, airborne sensors, command platforms, and data link systems form a real-time information network that continuously transmits multi-dimensional state information such as the target's three-dimensional coordinates, velocity vector, and attitude angle. Accurately predicting the enemy's maneuvering trajectory is a core prerequisite for achieving a tactical advantage, directly influencing intercept path planning and weapon positioning decisions. Current technology systems primarily employ four modeling approaches for trajectory prediction: kinematic deduction based on physical models, pattern matching based on machine learning, strategy simulation based on reinforcement learning, and end-to-end mapping based on deep learning. These approaches differ significantly in their architectural design and actual combat effectiveness, and each also possesses corresponding drawbacks.

[0003] Prediction systems based on physical models typically consist of a Newtonian kinematic equation derivation module, a radar measurement preprocessing module, and a trajectory extrapolation display. Their core relies on pre-set target dynamic parameters and implements probability weighting under multiple motion assumptions through a Kalman filter array within an interactive multi-model algorithm. However, when faced with highly nonlinear maneuvers, these methods increase the error between the predicted and true trajectories because the model set cannot cover the actual maneuvering patterns. Furthermore, computational latency increases exponentially with the number of models.

[0004] The machine learning-based prediction system utilizes a cascaded architecture of feature engineering and classification regression. It first constructs a situation description vector based on expert experience, including handcrafted features such as relative distance and azimuth angle change rate. It then uses a combination of support vector machines and decision trees to identify predefined maneuver pattern categories. Finally, a Gaussian process regression unit is used to generate a single predicted trajectory. This approach suffers from a sharp drop in recognition rate when faced with the complex maneuvers of new fighter aircraft due to insufficient feature representation dimensionality. Furthermore, the trajectory generation module lacks multimodal output capabilities and is unable to represent the uncertainty of the target's tactical intent.

[0005] The reinforcement learning-based prediction system constructs an environmental simulator that captures the confrontational situation between the Red and Blue teams. It uses a deep value network or proximal policy optimization algorithm to drive the agent to explore the optimal maneuver strategy. Its reward function library integrates evaluation metrics such as energy maneuver advantage and weapon launch conditions. Although trajectory sequences that conform to the logic of air combat games can be generated in the simulation environment, such systems require repeated interactive training to converge.

[0006] Deep learning-based prediction methods currently show the greatest potential for application. Their typical architecture employs a bidirectional recurrent neural network as a temporal encoder, mapping historical trajectory sequences into latent space feature vectors, which are then decoded through fully connected layers to generate deterministic trajectory coordinates. These methods automatically extract spatiotemporal features through end-to-end training and can integrate environmental constraints (such as terrain elevation data) with adversarial situational information (such as the relative azimuth of the enemy and friendly forces). However, existing systems suffer from three fundamental flaws: First, the network structure processes temporal and spatial features in stages (e.g., temporal encoding followed by spatial encoding), disrupting the spatiotemporal coupling of maneuvers; second, the mean squared error loss function forces the network to output an averaged trajectory, failing to represent the multimodal maneuver strategies that a target might adopt in tactical games; and third, the generated trajectories may violate fundamental aerodynamic laws, such as when instantaneous overload exceeds the structural limits of the aircraft or when the velocity vector undergoes discontinuous changes. Such physical distortions pose significant decision-making risks during deployment. The above bottlenecks are further exacerbated in complex electromagnetic interference environments. When the signal-to-noise ratio is lower than the threshold, the prediction confidence of the existing long short-term memory network (LSTM) method drops sharply, exposing major hidden dangers in the battlefield robustness of the existing technical system. Summary of the Invention

[0007] In response to the problems existing in the prior art, the present invention provides a method and system for predicting the multimodal trajectory of air combat target maneuvers based on a spatiotemporal joint attention mechanism. The solution builds an encoding and decoding neural network with an embedded spatiotemporal attention module, and uses the network's powerful self-learning ability to complete the prediction of future air combat target trajectories.

[0008] To achieve the above object, the technical solution of the present invention is as follows:

[0009] The present invention provides a method for predicting multimodal trajectories of air combat target maneuvers based on a spatiotemporal joint attention mechanism, comprising the following steps:

[0010] Step 1: Air combat scenario modeling: Based on the multi-agent trajectory data, extract the historical trajectory information of each agent;

[0011] Step 2: Modeling translational and rotational invariance: Based on the trajectory data obtained in step 1, a local spatiotemporal coordinate system is established for each scene element. This allows us to obtain the relative spatiotemporal information of the agent and reconstruct the absolute position of one element from another.

[0012] Step 3, stationarity modeling: perform instance normalization based on the local coordinate system trajectory data obtained in step 2, calculate the position and velocity normalization parameters based on the motion state of all targets in the current scene, and normalize the position and velocity characteristics to obtain stationary trajectory data;

[0013] Step 4: Spatiotemporal correlation modeling: Based on the stationary trajectory data obtained in step 3, the spatiotemporal context of the scene is fused based on the spatiotemporal joint encoder, and the spatiotemporal joint scene encoding is adaptively output;

[0014] Step 5: Multimodal decoding: Based on the scene encoding output from step 4, a mixed density decoder is used to output multiple future trajectories for each target agent and predict the optimal trajectory.

[0015] Furthermore, the following steps are included:

[0016] Step 6: Feasibility verification: Perform physical feasibility verification based on the optimal trajectory predicted in step 5 to achieve trajectory prediction.

[0017] Furthermore, the step 1 specifically includes:

[0018] Given the historical observations o of all agents in the previous H time steps, we extract the historical trajectory sequence of the past H time steps with a sliding window and construct the input tensor X, which is defined as follows:

[0019]

[0020] Among them: o includes the target input source information;

[0021] The set of K possible future T time-step trajectories of all agents is expressed as:

[0022]

[0023] Where: Y k represents the set of states predicted by the agent at time steps 1 to T.

[0024] Furthermore, the step 1 further includes the following steps:

[0025] Taking the center of the tactical scene as the origin, the airborne sensor data are uniformly converted to the North-East coordinate system.

[0026] Furthermore, step 2 includes the following sub-steps:

[0027] Step 2.1: For a i ,v i ,t i ) and another element with an absolute spatiotemporal position (p j ,v j ,t j ) elements, using a three-dimensional descriptor to summarize the relative relationship of the intelligent body. The three-dimensional descriptor components include the relative posture angle α i→j , relative azimuth β i→j and relative position p i→j =pi -p j , where p is the position, v is the velocity, and t is the timestamp;

[0028] Relative attitude angle α i→j Expressed as:

[0029]

[0030] Relative azimuth β i→j Expressed as:

[0031]

[0032] Therefore, the relative spatiotemporal information r i→j Expressed as:

[0033] r i→j =[sin(α i→j ),cos(α i→j ),sin(β i→j ),cos(β i→j ),||p i→j ||];

[0034] Step 2.2: Combine the geometric attributes and semantic attributes of the agent and generate relative position embeddings through the multi-layer perceptron MLP i→j .

[0035] Furthermore, the step 3 specifically includes the following sub-steps:

[0036] Step 3.1: Based on the motion states of all targets in the current scene, perform the following calculations:

[0037]

[0038]

[0039] Where: μ and σ are the mean and variance of a particular variable;

[0040] Step 3.2: Normalize the position and velocity features according to the scene mean and variance. The calculation formula is as follows:

[0041]

[0042] in: Represents the normalized feature variable.

[0043] Furthermore, the step 4 specifically includes the following sub-steps:

[0044] Step 4.1: Time-space fusion branch: First, use the LSTM network as the encoder of the historical trajectory data, and each agent outputs its own temporal features. Then, treat each agent as a token, encode it and add the spatial position encoding. Then, use a two-layer self-attention encoder as the agent-agent interaction encoder, where the query, key, and value are the historical trajectory embeddings of the encoded entity, and finally output the spatial interaction information.

[0045] Step 4.2: Spatial-temporal fusion branch: First, treat each agent as a token, encode it, and add spatial position encoding. Then, use a two-layer self-attention encoder as the agent-agent interaction encoder, where the query, key, and value are the current position embeddings of the encoded entity, and finally output the spatial interaction information. Then, use the LSTM network as the encoder for the historical trajectory data, and each agent outputs its own spatiotemporal features.

[0046] Step 4.3: The features output by the time-space fusion branch and the space-time fusion branch are concatenated and then gating weights are generated by the multi-layer perceptron MLP. The weighted fusion outputs the spatiotemporal joint scene coding. The calculation formula is as follows:

[0047]

[0048]

[0049] Among them: Sigmoid(·) represents the normalized nonlinear transformation unit, is the spatial interaction information output by the time-space fusion branch, It is the spatiotemporal feature output by the space-time fusion branch.

[0050] Furthermore, the step 5 specifically includes the following sub-steps:

[0051] Step 5.1: Expand the spatiotemporal encoding features to a multidimensional latent space through a multi-layer perceptron to generate multimodal latent variables covering the diversity of tactical intent

[0052] Step 5.2: The hybrid density network decodes and outputs multimodal multi-entity multi-step prediction position parameters, multimodal multi-entity multi-step prediction variance parameters, and multimodal multi-entity modal probabilities based on the multimodal representation, where the variance parameters represent trajectory uncertainty.

[0053] For each modality k, the decoder outputs Gaussian distribution parameters and modality probability π k , the calculation formula is as follows:

[0054] μ k ,σ k =MLP(ek )

[0055] π k =Softmax(MLP(e k ))

[0056] Among them: Softmax(·) represents the normalized nonlinear transformation unit, μ k is the mean, σ k is the variance;

[0057] Step 5.3: Calculate the matching error between the candidate trajectory and the true trajectory, and select the pattern with the smallest error as the optimal prediction.

[0058] Furthermore, the step 6 specifically includes the following steps:

[0059] Step 6.1: Introduce the penalty term L for the heading angle change rate |Δψ| and overload η in the training loss phy , the calculation formula is as follows:

[0060] L phy =λ1∑max(Δψ|-ψ max ,0)+λ2∑max(η-η max ,0)

[0061] Among them: λ1, λ2 represent the penalty weights, ψ max Indicates the maximum tolerance limit of the heading angle change rate, η max Indicates the maximum tolerance limit of overload;

[0062] Step 6.2: Perform threshold determination on the overload, speed, and heading angle change rate of the predicted trajectory;

[0063] Step 6.3: Perform local interpolation correction on the predicted trajectory that violates the physical rules to ensure that it complies with the body dynamics constraints.

[0064] The present invention also provides a computer system comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method for predicting multimodal trajectory of air combat target maneuvers based on a spatiotemporal joint attention mechanism provided by the present invention.

[0065] The beneficial effects of the present invention are:

[0066] The multimodal trajectory prediction method for air combat target maneuvers combined with a spatiotemporal joint attention mechanism provided by the present invention realizes the synchronous capture of temporal mutations and spatial games through a dual-branch attention mechanism; uses multimodal prediction to generate multiple possible trajectories for targets in heterogeneous, complex, and highly dynamic environments, taking into account the uncertainty and variability in trajectory prediction; introduces an aerodynamic penalty term into the loss function to ensure that the trajectory meets the limits of the fighter body and eliminates physical distortion, thereby further improving the accuracy and reliability of trajectory prediction and providing more effective support for decision-making in air combat. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 Schematic diagram of the implementation architecture of the multimodal trajectory prediction method for air combat target maneuvers based on the spatiotemporal joint attention mechanism provided by the present invention.

[0068] Figure 2 A schematic flow chart of the steps of the multimodal trajectory prediction method for air combat target maneuvers based on the spatiotemporal joint attention mechanism provided by the present invention. DETAILED DESCRIPTION

[0069] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0070] like Figure 1 、 Figure 2 As shown, the method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism provided by the present invention includes the following steps:

[0071] S1: Air combat scenario modeling: Based on the given multi-agent trajectory data, extract the historical trajectory information of each agent; specifically, the following steps are included:

[0072] S1.1: This invention assumes that the air combat scenario can be described as a continuous space-discrete space system, involving the self (denoted as A0) and other N intelligent agents (denoted as A1 to A2). N Given the historical observations o of all agents in the previous H time steps, we extract the historical trajectory sequence of the past H time steps with a sliding window and construct the input tensor X, which is defined as follows:

[0073]

[0074] Among them: o includes airborne sensors, command platforms, data links and other measurement equipment that transmit a large amount of target input source information, including observation data such as position p, speed v, and timestamp t.

[0075] The present invention expresses the set of K possible future T time-step trajectories of all agents as:

[0076]

[0077] Where: Y k represents the set of states predicted by the agent at time steps 1 to T.

[0078] S1.2: With the center of the tactical scene as the origin, uniformly transform the airborne sensor data into the North-East-East (NED) coordinate system to eliminate global motion interference.

[0079] For each agent, the present invention constructs its local scene context by aggregating potential interaction scene contexts within a specified range to help the model understand the changes in the agent's position over time.

[0080] S2: Modeling translational and rotational invariance: To achieve generalization and robustness in air combat trajectory prediction, a local spatiotemporal coordinate system is established for each scene element based on the trajectory data obtained in S1. This eliminates global motion interference and ensures that the model remains invariant to the overall translation and rotation of the scene. This includes the following steps:

[0081] S2.1: For a i ,v i ,t i ) and another element with an absolute spatiotemporal position (p j ,v j ,t j ), a three-dimensional descriptor is used to summarize the relative relationship of the agents, whose components include the relative posture angle α i→j , relative azimuth β i→j and relative position p i→j =p i -p j To improve numerical stability and avoid gradient explosion caused by direction jumps, angles are represented using sine and cosine values.

[0082] The relative attitude angle α of the present invention i→j Expressed as:

[0083]

[0084] The relative azimuth angle β of the present invention i→j Expressed as:

[0085]

[0086] Therefore, the relative spatiotemporal information r i→j Expressed as:

[0087] r i→j =[sin(α i→j ),cos(α i→j),sin(β i→j ),cos(β i→j ),||p i→j ||]

[0088] Since the absolute position of one element can be easily reconstructed from another element with the help of descriptors, all spatiotemporal position information of pairs of scene elements is preserved;

[0089] S2.2: Combine the geometric attributes of the agent with its semantic attributes (such as category, friend-or-foe identification) and generate relative position embeddings through a multi-layer perceptron (MLP) i→j , the calculation formula is as follows:

[0090]

[0091] Where: h i ,h j Represents the attribute feature vector of agent i, j, Represents vector concatenation.

[0092] S3: Stationarity Modeling: To eliminate the differences in motion scales of different targets, instance normalization is performed based on the local coordinate system trajectory data obtained in S2 to improve the model's convergence speed and generalization ability. The specific steps include the following:

[0093] S3.1: Stationarity refers to the fact that the statistical properties of a time series remain unchanged over time. However, due to the complex dynamics of entity trajectories in air combat scenarios and the random errors of sensors, non-stationarity is widely present in trajectory spatiotemporal series data, posing a severe challenge to accurate prediction. Considering that time series data are usually collected over a long period of time, these non-stationary series will inevitably cause the distribution of the prediction model to change over time. This shift can lead to performance degradation during testing due to covariate shift or conditional shift. To address this issue, the present invention uses an instance normalization method on the time series input X;

[0094] S3.2: Calculate the normalized position and velocity parameters based on the motion states of all agents in the current scene. The calculation formula is as follows:

[0095]

[0096] Where: μ and σ are the mean and variance of a particular variable.

[0097] S3.3: Normalize the position and velocity features according to the scene mean and variance. The calculation formula is as follows:

[0098]

[0099] in: Represents the normalized feature variable.

[0100] S4: Spatiotemporal Correlation Modeling: To capture the spatiotemporal coupling characteristics of the target's maneuvering behavior and the enemy-friendly game logic, the scene spatiotemporal context is fused based on the spatiotemporal joint encoder based on the stationary trajectory data obtained in S3. The specific steps include:

[0101] S4.1: This invention uses dual-branch parallel modeling and dynamic gating fusion to overcome the problem of insufficient modeling of complex game scenarios by traditional methods by separating temporal and spatial priorities.

[0102] S4.2: Time-space fusion branch: First, use the LSTM network as the encoder of historical trajectory data, and each agent outputs its own time series features. The calculation formula is as follows:

[0103]

[0104] Among them: LSTM(·) represents the time series fusion calculation of data.

[0105] Each agent is then treated as a token, encoded and spatially encoded, and a two-layer self-attention encoder is used as the agent-agent interaction encoder, where the query, key, and value are embedded as the historical trajectory of the encoded entity, and finally the spatial interaction information is output. The calculation formula is as follows:

[0106]

[0107] Among them: SelfAttention(·) is the self-attention calculation of the token, and LayerNorm(·) is the layer normalization calculation.

[0108] S4.3: Spatial-temporal fusion branch: Each agent is first treated as a token, encoded and spatially encoded, and then a two-layer self-attention encoder is used as the agent-agent interaction encoder, where the query, key, and value are the current position embeddings of the encoded entity, and finally the spatial interaction information is output. The calculation formula is as follows:

[0109]

[0110]

[0111] Among them: EnityEmbed(·) maps each entity plus its position encoding to a high-dimensional vector through a trainable linear projector,

[0112] Then, the LSTM network is used as the encoder of the historical trajectory data, and each agent outputs its own spatiotemporal features. The calculation formula is as follows:

[0113]

[0114] S4.4: After the dual-branch features are concatenated, a multi-layer perceptron is used to generate gating weights, and weighted fusion is used to output a spatiotemporal joint scene code. The calculation formula is as follows:

[0115]

[0116] Wherein: Sigmoid(·) represents the normalized nonlinear transformation unit.

[0117] This paper establishes an interaction model centered on each agent for the future time system. The interaction graph of each agent is independently constructed based on its characteristics and surrounding environment, which can more accurately capture its unique behavior patterns and interaction needs.

[0118] S5: Multimodal decoding: To generate a set of multiple trajectories that are physically plausible and cover tactical possibilities, the mixed density decoder outputs multiple future trajectories for each target agent based on the scenario encoding output by S4. This includes the following steps:

[0119] S5.1: Due to the complexity and diversity of possible trajectory intentions, the actual trajectory of a target in a combat scenario is often uncertain. To address the uncertainty and variability in trajectory prediction, the multimodal mapping module uses a multi-layer perceptron to expand the spatiotemporal encoding features into a multidimensional latent space, generating multimodal latent variables that cover the diversity of tactical intentions.

[0120] S5.2: The hybrid density network decodes and outputs multimodal multi-entity multi-step prediction position parameters, multimodal multi-entity multi-step prediction variance parameters, and multimodal multi-entity modal probabilities based on the multimodal representation, where the variance parameters represent trajectory uncertainty.

[0121] For each modality k, the decoder outputs Gaussian distribution parameters (mean μ k , variance σ k ) and modal probability π k , the calculation formula is as follows:

[0122] μ k ,σ k =MLP(e k )

[0123] π k =Softmax(MLP(e k ))

[0124] Where: Softmax(·) represents the normalized nonlinear transformation unit.

[0125] S5.3: By calculating the matching error between the candidate trajectory and the true trajectory, the pattern with the smallest error is selected as the optimal prediction.

[0126] S6: Feasibility Verification: To ensure that the generated trajectory meets aerodynamic and structural constraints, a physical feasibility verification is performed based on the optimal trajectory predicted in S5, thereby achieving trajectory prediction. The specific steps include the following:

[0127] S6.1: Introduce the penalty term L for the heading angle change rate |Δψ| and overload η in the training loss phy , the calculation formula is as follows:

[0128] L phy =λ1∑max(|Δψ|-ψ max ,0)+λ2∑max(η-η max ,0)

[0129] Among them: λ1, λ2 represent the penalty weights, ψ max Indicates the maximum tolerance limit of the heading angle change rate, η max Indicates the maximum tolerance limit for overload.

[0130] S6.2: Threshold determination is performed on the overload, speed, and heading angle change rate of the predicted trajectory. Trajectories exceeding the preset thresholds are treated as trajectory segments that violate physical rules and then processed in S6.3.

[0131] S6.3: Perform local interpolation corrections on trajectory segments that violate physical rules to ensure compliance with the body dynamics constraints.

[0132] The present invention also provides a computer system comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method for predicting multimodal trajectory of air combat target maneuvers based on a spatiotemporal joint attention mechanism provided by the present invention.

[0133] It should be noted that the above content merely illustrates the technical idea of the present invention and cannot be used to limit the scope of protection of the present invention. For ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications all fall within the scope of protection of the claims of the present invention.

Claims

1. A multimodal trajectory prediction method for air combat target maneuvers based on a spatiotemporal joint attention mechanism, characterized by: The steps include: Step 1: Air combat scenario modeling: Based on the multi-agent trajectory data, extract the historical trajectory information of each agent; Step 2: Modeling translational and rotational invariance: Based on the trajectory data obtained in step 1, a local spatiotemporal coordinate system is established for each scene element. This allows us to obtain the relative spatiotemporal information of the agent and reconstruct the absolute position of one element from another. Step 3, stationarity modeling: perform instance normalization based on the local coordinate system trajectory data obtained in step 2, calculate the position and velocity normalization parameters based on the motion state of all targets in the current scene, and normalize the position and velocity characteristics to obtain stationary trajectory data; Step 4: Spatiotemporal correlation modeling: Based on the stationary trajectory data obtained in step 3, the spatiotemporal context of the scene is fused based on the spatiotemporal joint encoder, and the spatiotemporal joint scene encoding is adaptively output; Step 5: Multimodal decoding: Based on the scene encoding output from step 4, a mixed density decoder is used to output multiple future trajectories for each target agent and predict the optimal trajectory.

2. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1 is characterized in that The following steps are also included: Step 6: Feasibility verification: Perform physical feasibility verification based on the optimal trajectory predicted in step 5 to achieve trajectory prediction.

3. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1, characterized in that: The step 1 specifically includes: Given the historical observations o of all agents in the previous H time steps, we extract the historical trajectory sequence of the past H time steps with a sliding window and construct the input tensor X, which is defined as follows: Among them: o includes the target input source information; The set of K possible future T time-step trajectories of all agents is expressed as: Where: Y k represents the set of states predicted by the agent at time steps 1 to T.

4. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1, characterized in that: The step 1 further comprises the following steps: Taking the center of the tactical scene as the origin, the airborne sensor data are uniformly converted to the North-East coordinate system.

5. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1, characterized in that: The step 2 includes the following sub-steps: Step 2.1: For a i ,v i ,t i ) and another element with an absolute spatiotemporal position (p j ,v j ,t j ) elements, using a three-dimensional descriptor to summarize the relative relationship of the intelligent body. The three-dimensional descriptor components include the relative posture angle α i→j , relative azimuth β i→j and relative position p i→j =p i -p j , where p is the position, v is the velocity, and t is the timestamp; Relative attitude angle α i→j Expressed as: Relative azimuth β i→j Expressed as: Therefore, the relative spatiotemporal information r i→j Expressed as: r i→j =[sin(α i→j ),cos(α i→j ),sin(β i→j ),cos(β i→j ),||p i→j ||]; Step 2.2: Combine the geometric attributes and semantic attributes of the agent and generate relative position embeddings through the multi-layer perceptron MLP i→j .

6. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1, characterized in that: The step 3 specifically includes the following sub-steps: Step 3.1: Based on the motion states of all targets in the current scene, perform the following calculations: Where: μ and σ are the mean and variance of a particular variable; Step 3.2: Normalize the position and velocity features according to the scene mean and variance. The calculation formula is as follows: in: Represents the normalized feature variable.

7. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1, characterized in that: The step 4 specifically includes the following sub-steps: Step 4.1: Time-space fusion branch: First, use the LSTM network as the encoder of the historical trajectory data, and each agent outputs its own time series features; Each agent is then treated as a token, encoded and spatially encoded, and a two-layer self-attention encoder is used as the agent-agent interaction encoder, where the query, key, and value are embedded as the historical trajectory of the encoded entity, and finally the spatial interaction information is output; Step 4.2: Spatial-temporal fusion branch: First, treat each agent as a token, encode it, and add spatial position encoding. Then, use a two-layer self-attention encoder as the agent-agent interaction encoder, where the query, key, and value are the current position embeddings of the encoded entity, and finally output the spatial interaction information. Then, use the LSTM network as the encoder for the historical trajectory data, and each agent outputs its own spatiotemporal features. Step 4.3: The features output by the time-space fusion branch and the space-time fusion branch are concatenated and then gating weights are generated by the multi-layer perceptron MLP. The weighted fusion outputs the spatiotemporal joint scene coding. The calculation formula is as follows: Among them: Sigmoid(·) represents the normalized nonlinear transformation unit, is the spatial interaction information output by the time-space fusion branch, It is the spatiotemporal feature output by the space-time fusion branch.

8. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 1, characterized in that: The step 5 specifically includes the following sub-steps: Step 5.1: Expand the spatiotemporal encoding features to a multidimensional latent space through a multi-layer perceptron to generate multimodal latent variables covering the diversity of tactical intent Step 5.2: The hybrid density network decodes and outputs multimodal multi-entity multi-step prediction position parameters, multimodal multi-entity multi-step prediction variance parameters, and multimodal multi-entity modal probabilities based on the multimodal representation, where the variance parameters represent trajectory uncertainty. For each modality k, the decoder outputs Gaussian distribution parameters and modality probability π k , the calculation formula is as follows: m k ,s k =MLP(e k ) π k =Softmax(MLP(e k )) Among them: Softmax(·) represents the normalized nonlinear transformation unit, μ k is the mean, σ k is the variance; Step 5.3: Calculate the matching error between the candidate trajectory and the true trajectory, and select the pattern with the smallest error as the optimal prediction.

9. The method for predicting multimodal trajectory of air combat target maneuvers based on spatiotemporal joint attention mechanism according to claim 2, characterized in that: The step 6 specifically includes the following steps: Step 6.1: Introduce the penalty term L for the heading angle change rate |Δψ| and overload η in the training loss phy , the calculation formula is as follows: L phy =λ1∑max(Δψ|-ψ max ,0)+λ2∑max(η-η max ,0) Among them: λ1, λ2 represent the penalty weights, ψ max Indicates the maximum tolerance limit of the heading angle change rate, η max Indicates the maximum tolerance limit of overload; Step 6.2: Perform threshold determination on the overload, speed, and heading angle change rate of the predicted trajectory; Step 6.3: Perform local interpolation correction on the predicted trajectory that violates the physical rules to ensure that it complies with the body dynamics constraints.

10. A computer system comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the multimodal trajectory prediction method for air combat target maneuvers based on a spatiotemporal joint attention mechanism provided in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Flight simulation environment-oriented hand motion trail prediction method

    CN116611027A

  • Target trajectory prediction method under battlefield task planning background

    CN116663384A

  • Multi-modal trajectory prediction method based on vehicle interaction diagram space-time decoupling coding

    CN119329558A

  • Vehicle optimal lane changing opportunity prompting method and system based on pre-training LLMs trajectory prediction

    CN119682776A

  • Aircraft control device, aircraft, and method for computing aircraft trajectory

    US20180074524A1