Complex game confrontation task-oriented agent training method and system

Through sliding window data processing, trajectory prediction of Transformer model, as well as the intention prediction of LSTM model, and the model fusion builds an opponent modeling system, the problem of low accuracy of opponent trajectory and intention prediction in the existing technology is solved, and higher prediction accuracy and scientificity of air combat decision-making are achieved.

CN120146091APending Publication Date: 2025-06-13NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510230035.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the opponent's trajectory and intention in air game confrontation, especially in complex actions and long-term span scenarios, resulting in low prediction accuracy and unscientific decision-making.

Method used

The sliding window data processing method and the powerful Transformer model are used for trajectory prediction, combined with the LSTM model to train tactics and maneuver intention prediction models, and a comprehensive adversary modeling system is built through model fusion.

Benefits of technology

It significantly improves the accuracy of prediction of opponent trajectory and intentions, enhances the adaptability to complex tactical changes, improves the scientificity and effectiveness of air combat decisions, and improves the winning rate of our own aircraft in air combat.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146091A_ABST
    Figure CN120146091A_ABST
Patent Text Reader

Abstract

The invention discloses a complex game confrontation task-oriented agent training method and system, and belongs to the technical field of agent training in game confrontation. The method comprises the following steps: firstly, designing and training a reinforcement learning agent on a game confrontation simulation platform; secondly, the agents are used for mutual game confrontation and data collection, a sliding window method is used for preprocessing track data, and opponent intention data are labeled and cleaned; then, respectively training a trajectory prediction model and an intention prediction model by utilizing Transform and LSTM (Long Short Term Memory); and finally, carrying out feature extraction and fusion on the multi-modal opponent information in the two models to form a new decision network with an opponent modeling function for retraining so as to obtain a more advanced agent. According to the method, in complex scenes such as air game confrontation and the like, through deep research on multi-modal opponent modeling, the capability of constructing a comprehensive and effective strategy based on opponent modeling is enhanced, and the performance of the intelligent agent in complex tasks and the capability of dynamically adjusting the strategy in real time for the opponent are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent agent training in game confrontation, and particularly relates to an intelligent agent training method and system for complex game confrontation tasks. Background Art

[0002] In the complex scenario of air game confrontation, whether it is traditional air combat or modern UAV air combat, accurate cognition of the opponent is a key factor in achieving victory. However, there are many obvious defects in the current existing opponent modeling technologies.

[0003] In terms of trajectory prediction, most traditional methods rely on simple dynamic models or statistical analysis based on limited historical data. For example, some trajectory prediction methods based on Kalman filtering, although they can have a certain effect in relatively stable flight states, when facing sudden maneuvers of the opponent's aircraft, such as sharp turns, dives, climbs, etc., since these methods are difficult to capture the drastic changes in flight states and complex dynamic information, the prediction accuracy will drop significantly. And prediction algorithms based on time series, such as LSTM, although they can handle time series data to a certain extent, for the prediction of long sequence trajectories, due to their own structural limitations, it is difficult to effectively capture long-distance dependencies, and they cannot accurately predict the future trajectory of the opponent when facing air combat scenarios with a long time span.

[0004] In the field of intention prediction, existing technologies mainly judge the opponent's intention through simple rule matching or classification methods based on a small number of features. For example, only judging whether the opponent is in an attack intention based on a few features such as the speed change and heading change of the opponent's aircraft, this way ignores many other key factors in air combat, such as the attitude of the aircraft, the weapon loading state, the relative position and distance from one's own aircraft, etc. In addition, for some complex tactical intentions, such as feigning an attack, threat maintenance, etc., the existing intention prediction methods are even more difficult to accurately identify, resulting in a lack of sufficient accurate information support in air combat decision-making and being easily in a passive position.

[0005] In addition, currently there are few methods that can organically integrate trajectory prediction and intention prediction to construct a comprehensive and systematic opponent modeling system. This makes it impossible to comprehensively analyze and judge the behavior and intention of the opponent in actual air combat confrontation, affecting the scientificity and effectiveness of combat decision-making. Therefore, it is urgent to develop an intelligent agent training method and system for complex game confrontation tasks that can effectively solve the above problems. Summary of the Invention

[0006] The core objective of the present invention is to provide an innovative intelligent agent training method and system for complex game confrontation tasks. Through unique data processing methods, advanced model training techniques, and effective model fusion strategies, the ability to predict the opponent's trajectory and intention is significantly improved, thereby providing strong decision-making support for complex game confrontation.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is: an intelligent agent training method for complex game confrontation tasks, the method comprising:

[0008] S1. In a high-fidelity air game confrontation simulation platform, a reinforcement learning intelligent agent with initial intelligence is trained based on the proximal policy optimization algorithm;

[0009] S2. In the data collection stage, the reinforcement learning intelligent agents with initial intelligence are used to conduct mutual game confrontations, collect confrontation data, and perform all-round data collection; all-round collection of key information such as the flight trajectory, speed, acceleration, weapon firing state, and relative position relationship with the opponent of the intelligent agent, providing a high-quality data set for subsequent opponent modeling;

[0010] S3. Training of the trajectory prediction model, using the sliding window method to preprocess the collected trajectory data and convert it into time series samples suitable for model training; at the same time, the collected game data is finely annotated, and the opponent's maneuver intention is divided into altitude action abstractions (climb, maintain, descend), speed action abstractions (accelerate, maintain, decelerate), and direction action abstractions (turn left, maintain, turn right), and the tactical intention is divided into threat maintenance, feigned attack, full attack, and breakaway and evasion;

[0011] S4. Training of the intention prediction model, using Transformer to train the trajectory prediction model, and using LSTM to train the tactical and maneuver intention prediction models respectively;

[0012] S5. In the decision-making process, the trajectory prediction model and intention prediction model trained in S4 are used for deployment and inference to obtain the opponent's future trajectory, maneuver intention, and tactical intention and perform feature encoding, and these information are jointly used as the input information of the actor network;

[0013] S6. Comprehensive analysis of trajectory and intention, through the comprehensive analysis of the trajectory prediction model and intention prediction model, the information modeled by the opponent is fused into the decision-making network to obtain a more advanced intelligent agent.

[0014] In step S2 of the present invention, in the data collection stage, different sampling frequencies and data cleaning methods are set for different types of data, and the trajectory data and intention data are cleaned according to the game scenario under incomplete information.

[0015] In step S3 of the present invention, during the training of the trajectory prediction model, the long sequence data processed by the sliding window technique is used, and the pre-trained model parameters are utilized to accelerate the model convergence, improving the prediction accuracy and the ability to predict long sequence data.

[0016] In step S4 of the present invention, the intention prediction model performs classification and annotation processing on both the upper-layer tactical intention and the middle-layer maneuver intention, and uses a long short-term memory network to train the intention prediction model to enhance the model's recognition and prediction ability for complex intentions.

[0017] In step S5 of the present invention, after model fusion, through multiple simulation verifications and actual tests, the performance and stability of the fusion model are continuously optimized.

[0018] In step S5 of the present invention, the method deploys and fuses the trained trajectory prediction model and intention prediction model, comprehensively considers the results of trajectory prediction and intention prediction, and according to the opponent's trajectory and intention information predicted by the model with trained parameters, these information are spliced as the input information of the actor network.

[0019] In step S6 of the present invention, through the comprehensive analysis of the trajectory and intention, the information modeled by the opponent is fused into the decision network. Specifically, in the actor network, feature fusion is performed on the information obtained in S5. Specifically, a convolutional neural network is added to extract features of the opponent's future long sequence trajectory, and then the maneuver intention and tactical intention information are spliced and input into a fully connected network for decision-making. Reinforcement learning training is carried out using the new actor network architecture containing opponent information to obtain a more advanced intelligent agent.

[0020] Another object of the present invention is to further provide a system for implementing the above-mentioned intelligent agent training method for complex game confrontation tasks. The system includes a data acquisition stage unit, a trajectory prediction model training unit, an intention prediction model training unit, and a trajectory and intention comprehensive analysis unit; the data acquisition stage unit is used to collect data and perform data cleaning by using the trained intelligent agents to conduct mutual game confrontations; the trajectory prediction model unit is used to predict the opponent's future trajectory by using the past detected opponent trajectories; the intention prediction model unit is used to predict the opponent's future intention according to the battlefield situation information; the trajectory and intention comprehensive analysis unit is used to fuse the opponent's future trajectory and intention into the decision network and then retrain to achieve performance improvement.

[0021] The data acquisition stage unit of the present invention includes steps S1 and S2, the trajectory prediction model training unit includes step S3, the intention prediction model training unit includes steps S4 and S5, and the trajectory and intention comprehensive analysis unit includes step S6.

[0022] The beneficial effects of adopting the above technical solutions are as follows: 1. With the unique sliding window data processing method and powerful Transformer model training, in actual air combat simulations, compared with traditional trajectory prediction methods, the present invention can more accurately capture the complex dynamic information in the opponent's long-order trajectory, significantly enhancing the adaptability to sudden maneuvers and complex tactical changes, and greatly improving the accuracy and reliability of trajectory prediction; 2. The present invention uses LSTM to train the tactical intention prediction model and the maneuver intention prediction model respectively, comprehensively considering various intentions of the opponent in air combat. Through learning and analyzing a large amount of air combat data, as shown in the embodiment Figure 4 the intention prediction accuracy of the model reaches more than 95%, providing comprehensive opponent information for one's own side and helping one's own side to layout in advance and gain the initiative in air combat; 3. By integrating the trajectory prediction model and the intention prediction model, the present invention constructs a comprehensive opponent modeling method, enabling one's own side to deeply understand the opponent's behavior and intention as a whole in the confrontation, and then formulating more scientific and effective confrontation strategies. After being verified by multiple simulated air combat experiments, the winning rate of one's own aircraft in air combat is improved after adopting this method, effectively enhancing the odds of winning in air combat confrontation; 4. In complex scenarios such as air combat confrontation, through in-depth research on multi-modal opponent modeling, the present invention enhances the ability to construct comprehensive and effective strategies based on opponent modeling, improves the performance of the intelligent agent in complex tasks, and its ability to adjust strategies in real time according to the opponent's dynamics. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of an intelligent agent training method for a complex game confrontation task of the present invention;

[0024] Figure 2 is a prediction effect diagram of the trajectory prediction part in air combat confrontation in the embodiment;

[0025] Figure 3 is a confrontation schematic diagram of a reinforcement learning opponent modeling intelligent agent after being trained by integrating the opponent modeling method in the embodiment;

[0026] Figure 4 is a test accuracy data diagram of the intention prediction model during the training process in the example;

[0027] Figure 5 is a comparison diagram of the training effects of the intelligent agent before and after opponent information fusion in the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0028] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in specific implementations. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application. It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.

[0029] A typical embodiment of the present invention is a multi-modal opponent modeling method for complex tasks in complex air combat scenarios, such as Figure 1 shown, including the following steps:

[0030] S1. Conduct intelligent agent confrontation training in a high-fidelity air combat simulation environment that can simulate various real air combat scenarios, including different meteorological conditions (such as sunny days, rainy days, heavy fog, etc.), geographical environments (such as over plains, mountains, oceans, etc.), and complex electromagnetic interference environments, etc. Based on the proximal policy optimization algorithm, a reinforcement learning intelligent agent with initial intelligence is trained.

[0031] S2. Use the intelligent agent to collect multi-field air combat confrontation data. Each confrontation includes detailed data of both sides' aircraft during the entire combat process. For the data of the opponent's aircraft, key data such as its flight trajectory (including coordinate information such as longitude, latitude, altitude, etc.), speed, acceleration, weapon launch status, and relative position relationship with one's own aircraft are collected. Tens of thousands of game confrontation data are collected in total, forming a rich air combat dataset.

[0032] S3. For the opponent's trajectory data, set the sliding window length to 60 time steps (each time step is 0.2 seconds), slide 1 step each time, extract the trajectory segments within the window from the starting position, and generate a time series sample dataset containing 1 million samples. Each sample includes information such as the position and speed of the opponent within 60 time steps. Preprocess the collected data, including cleaning (removing abnormal and incorrect data) and normalization (unifying the data scale) to improve the data quality and usability. Carefully annotate the air combat data, mark the opponent's maneuvering intention as turning left, etc., and classify the tactical intention into categories such as threat maintenance. When annotating, comprehensively consider various behavioral characteristics such as flight speed, and focus on the flight attitude and the changing trends of speed and acceleration.

[0033] S4. Training of the multi-modal opponent modeling model

[0034] (1) Training of the trajectory prediction model:

[0035] The Transformer model is selected for long-sequence trajectory prediction. Based on the self-attention mechanism, the Transformer model can effectively capture the dependencies between different time steps in long-sequence data, overcoming the limitations of traditional recurrent neural networks in processing long sequences. The calculation method of its self-attention mechanism is as follows:

[0036]

[0037] where Q (Query), K (Key), and V (Value) are the query vector, key vector, and value vector obtained by linearly transforming the input data respectively, and d k is the dimension of the key vector. The multi-head attention mechanism concatenates the results of multiple self-attention and then performs a linear transformation to obtain a richer feature representation.

[0038] Model training: Input the past 60-step data in the time series sample dataset into the Transformer model, and the output is the predicted result of the opponent's trajectory for the next 30 steps. During the training process, the mean squared error (MSE) loss function is used to measure the difference between the predicted trajectory and the real trajectory y, and the formula is as follows:

[0039]

[0040] where n is the number of samples, and y i are the predicted value and the real value of the i-th sample respectively. The parameters of the model are continuously adjusted through the backpropagation algorithm, so that the model can accurately learn the change rules and dynamic features of the opponent's trajectory. To prevent the model from overfitting, technical means such as L1 or L2 regularization are adopted to improve the generalization ability of the model.

[0041] (2) Training of the tactical intention prediction model:

[0042] Build an LSTM model for tactical intention prediction. The LSTM cell processes time series data through the input gate i t , forget gate f t , output gate o t and memory cell C t , and its calculation formula is as follows:

[0043] i t = σ(W ii x t + W hi h t―1 + b i )

[0044] f t = σ(W if x t+W hf h t―1 +b f )

[0045] o t = σ(W io x t +W ho h t―1 +b o )

[0046]

[0047] h t = o t ⊙tanh(C t )

[0048] where σ is the sigmoid function, tanh is the hyperbolic tangent function, ⊙ represents element-wise multiplication, W is the weight matrix, b is the bias vector, x t is the input at the current time step, and h t―1 is the hidden state at the previous time step. The model input is data such as the position, velocity, acceleration, weapon state of the opponent over a period of time (e.g., 30 time steps), as well as the relative position and relative angle with the own aircraft. After preprocessing, it is input into the LSTM network. During the training process, the cross-entropy loss function is used to measure the difference between the prediction result and the true label p, and the formula is:

[0049]

[0050] Optimization algorithms such as the Adam optimizer are used to adjust the model parameters to improve the prediction accuracy of the model.

[0051] (3) Maneuvering intention prediction model:

[0052] Similarly, another LSTM model is constructed for maneuvering intention prediction. The model structure is similar to that of the tactical intention prediction model, but the input data and model parameters are optimized according to the characteristics of maneuvering intention prediction. The model input is the position and velocity change information (including longitude, latitude, altitude, velocity, etc.) of the opponent detected by the radar in a short period of time (e.g., 10 time steps). By learning and analyzing this information, the model can accurately predict the maneuvering intention of the opponent. During the training process, the cross-entropy loss function and optimization algorithms are also used to train and optimize the model.

[0053] S5. Integrate the trained trajectory prediction model and intention prediction model. In practical applications, when the current trajectory data of the opponent is obtained, it is first input into the trajectory prediction model to obtain the trajectory prediction result of the opponent's next 30 steps. The prediction effect is as Figure 2As shown, at the same time, the current and recent flight state information is respectively input into the tactical intention prediction model and the maneuver intention prediction model to obtain the opponent's intention prediction result. The accuracy rate of the intention prediction evaluated in the test set is as Figure 4 As shown, the future maneuver trajectory, overall tactical intention, and maneuver intention of the opponent predicted by the above opponent prediction model are respectively encoded by a fully connected network for features, and finally these information are concatenated as the input information of the actor network.

[0054] S6. Add the multi-modal information of the opponent obtained in S5 to the input of the actor network of the proximal policy optimization algorithm in S1. Specifically, an additional convolutional neural network layer is added to the actor network to extract features from the opponent's future long sequence trajectory, and then the maneuver intention and tactical intention information are concatenated and input into the fully connected network for decision-making. The new actor network architecture containing the opponent information is used to re-train to obtain a more intelligent agent, as Figure 5 As shown, the training effect of the agent after fusing the opponent information is doubled compared with the stable reward value after fusing the opponent information. The trained agent with opponent modeling can comprehensively combine the results of trajectory prediction and intention prediction to formulate countermeasures for its own aircraft. For example, when the trajectory prediction shows that the opponent's aircraft will approach its own quickly and the tactical intention prediction is to attack with all its strength, and when the tactical intention prediction is to break away and avoid, its own aircraft can select an appropriate attitude and timing to launch a missile to shoot down the opponent according to the predicted escape trajectory of the opponent, as Figure 3 shown.

[0055] A system for an agent training method for complex game confrontation tasks, including a data collection stage unit (steps S1, S2), a trajectory prediction model training unit (step S3), an intention prediction model training unit (steps S4, S5), and a trajectory and intention comprehensive analysis unit (step S6); the data collection stage unit is used to collect data and clean data by using trained agents to conduct mutual game confrontation; the trajectory prediction model unit is used to predict the opponent's future trajectory by using the opponent's trajectory detected in the past; the intention prediction model unit is used to predict the opponent's future intention according to the battlefield situation information; the trajectory and intention comprehensive analysis unit is used to fuse the opponent's future trajectory and intention into the decision network and then re-train to achieve performance improvement.

[0056] Through the above specific implementation manners, the agent training method and system of the present invention for complex game confrontation tasks can effectively achieve accurate prediction of the opponent's trajectory and intention, providing strong technical support for opponent modeling in game confrontation tasks.

[0057] In summary, the present invention provides an agent training method and system for complex game confrontation tasks. Through innovative data processing and model training techniques, the method effectively improves the ability to predict the opponent's trajectory and intention, providing more powerful support for opponent modeling and intelligent decision-making in game confrontation tasks.

[0058] The above embodiments are only used to illustrate rather than limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the present invention can still be modified or equivalently replaced, and any modification or partial replacement without departing from the spirit and scope of the present invention shall be covered by the scope of the claims of the present invention.

Claims

1. A method for training an intelligent agent for complex game confrontation tasks, characterized in that: The method comprises: S1. In a high-fidelity aerial game confrontation simulation platform, a reinforcement learning agent with initial intelligence is trained based on the proximal strategy optimization algorithm; S2, data collection stage, using reinforcement learning agents with initial intelligence to play against each other, collect confrontation data, and conduct all-round data collection; S3, trajectory prediction model training, using the sliding window method to pre-process the collected trajectory data and convert it into time series samples suitable for model training; at the same time, the collected game data is finely annotated, and the opponent's maneuver intentions are divided into high action abstraction, speed action abstraction and direction action abstraction, and the tactical intentions are divided into threat maintenance, feint attack, full-force attack, and escape and evasion; S4, intention prediction model training, using Transformer to train the trajectory prediction model, and LSTM to train the tactical and maneuver intention prediction models respectively; S5: In the decision-making process, the trajectory prediction model and intention prediction model trained in S4 are deployed and inferred to obtain the opponent's future trajectory, maneuver intention and tactical intention, and perform feature encoding, and these information are used as input information of the actor network. S6. Comprehensive analysis of trajectory and intention. Through comprehensive analysis of trajectory prediction model and intention prediction model, the information modeled by the opponent is integrated into the decision network to obtain a more advanced intelligent agent.

2. The method for training an intelligent agent for complex game confrontation tasks according to claim 1, characterized in that: In the step S2, the data collection phase, different sampling frequencies and data cleaning methods are set for different types of data, and trajectory data and intention data are cleaned according to the game scenario under incomplete information.

3. The method for training an intelligent agent for complex game confrontation tasks according to claim 1, characterized in that: In the step S3, in the trajectory prediction model training, the long sequence data processed by the sliding window technology is used, and the pre-trained model parameters are used to accelerate the model convergence, thereby improving the prediction accuracy and the ability to predict long sequence data.

4. A method for training an intelligent agent for complex game confrontation tasks according to any one of claims 1 to 3, characterized in that: In the step S4, the intention prediction model classifies and labels the upper-level tactical intentions and the middle-level maneuvering intentions, and uses a long short-term memory network to train the intention prediction model to improve the model's ability to recognize and predict complex intentions.

5. An agent training method for complex game confrontation tasks according to any one of claims 1 to 3, characterized in that: In step S5, after the model is integrated, the performance and stability of the integrated model are continuously optimized through multiple simulation verifications and actual tests.

6. An agent training method for complex game confrontation tasks according to any one of claims 1 to 3, characterized in that: In the step S5, the method deploys and fuses the trained trajectory prediction model and intention prediction model, predicts the opponent's trajectory and intention information based on the model with trained parameters, and splices this information as input information of the actor network.

7. An agent training method for complex game confrontation tasks according to any one of claims 1 to 3, characterized in that: The step S6 integrates the information modeled by the opponent into the decision network through comprehensive analysis of the trajectory and intention. Specifically, the information obtained in S5 is feature fused in the actor network. Specifically, a layer of convolutional neural network is added to extract features of the opponent's future long sequence trajectory, and then the maneuver intention and tactical intention information are spliced ​​and input into the fully connected network for decision-making. The new actor network architecture containing the opponent information is used for reinforcement learning training to obtain a more advanced intelligent agent.

8. A system for implementing the agent training method for complex game confrontation tasks as described in any one of claims 1 to 7, characterized in that: The system includes a data collection phase unit, a trajectory prediction model training unit, an intention prediction model training unit, and a trajectory and intention comprehensive analysis unit; the data collection phase unit is used to use the trained intelligent agents to conduct mutual game confrontation to achieve data collection and data cleaning; the trajectory prediction model unit is used to use the opponent's trajectory detected in the past to predict the opponent's future trajectory; the intention prediction model unit is used to predict the opponent's future intention based on battlefield situation information; The trajectory and intention comprehensive analysis unit is used to integrate the opponent's future trajectory and intention into the decision network for retraining to achieve performance improvement.

9. The system of the agent training method for complex game confrontation tasks according to claim 8, characterized in that: The data collection phase unit includes steps S1 and S2, the trajectory prediction model training unit includes step S3, the intention prediction model training unit includes steps S4 and S5, and the trajectory and intention comprehensive analysis unit includes step S6.