Unmanned aerial vehicle countermeasure decision method based on experience and deduction model

CN118409503BActive Publication Date: 2026-10-09UNIV OF CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410378189.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-29
Publication Date
2026-10-09
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

[0008]为了克服上述问题,本发明人进行了深入研究,提出了一种基于经验和推演模型的无人机对抗决策方法,其不依赖于专家数据集并且能够较好考虑决策对未来影响,解决基于传统强化学习的无人机对抗决策算法不能很好的考虑长期规划,使用序列决策模型决策算法需要专家数据集且决策性能不稳定的问题

Benefits of technology

[0040] (1) Compared with traditional methods, the present invention has stronger ability to characterize the environment and make decision-making reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118409503B_ABST
    Figure CN118409503B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle confrontation decision-making methods based on experience and deductive model, and the optimal confrontation decision of world model is obtained by establishing initial confrontation trajectory data set training pre-decision model, second data set is constructed, and the optimal pre-decision model is obtained by using second data set training pre-decision model, and the optimal action of current time is output from optimization pre-decision model, and unmanned aerial vehicle is flown according to optimal action control.The method disclosed in the application can fully consider the long-term impact of decision on the environment by predicting the future state and reward, so that the unmanned aerial vehicle makes a decision that is more beneficial to the long term, and by considering long-term decision, a decision better than the decision of confrontation trajectory data set can be made, thereby achieving performance improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a UAV adversarial decision-making method based on experience and deductive models, belonging to the field of flight control technology. Background Technology

[0002] Currently, the mainstream method for combating drones is reinforcement learning, represented by SAC, PPO and their variants. Reinforcement learning is a decision-making method that allows an agent to interact with the environment and obtain the maximum cumulative reward in the environment through trial and error.

[0003] In the early stages of training drone adversarial agents based on reinforcement learning methods, the inaccurate valuation of state actions and future discounted rewards leads to almost random adversarial strategies, requiring extensive trial and error to arrive at an effective one. Furthermore, reinforcement learning-based adversarial algorithms necessitate playing numerous games from scratch, failing to fully utilize existing human experience.

[0004] Furthermore, in adversarial decision-making for drones, drone adversarial strategies based on reinforcement learning methods only consider the current state and future discounted rewards, without making long-term plans to achieve a certain goal. This makes the strategy prone to getting stuck in local optima or the current optimal solution rather than the global optimal solution.

[0005] Existing technologies also disclose long-term planning methods, such as the MUZERO method proposed by OpenAI. This method uses a Monte Carlo tree-based approach to enable agents to plan for the future and make more beneficial long-term decisions. However, Monte Carlo tree-based methods consume significant computational resources, and even more are needed for complex environments with high-dimensional action spaces, such as UAV combat, making it difficult to apply the MUZERO method in practice.

[0006] Existing technologies also include sequence decision-making methods, such as Decision Transformer. Based on a Transformer decoder architecture, it describes the reinforcement learning problem as a sequence problem. Unlike traditional reinforcement learning strategies that train from scratch, it uses expert trajectories as its dataset to fully leverage existing expert experience. This allows Decision Transformer to be trained on large amounts of offline data without interacting with the environment. Thanks to its decoder structure and sequence modeling of the reinforcement learning problem, Decision Transformer can handle high-dimensional and complex state and action spaces and has the advantage of being insensitive to reward functions. However, Decision Transformer only achieves satisfactory performance when trained on sufficiently large and high-quality trajectory datasets. Obtaining sufficiently large and high-quality UAV adversarial trajectory datasets is very difficult, limiting the potential of Decision Transformer in UAV adversarial environment decision-making. Furthermore, the performance of Decision Transformer is sensitive to the manually set hyperparameter Target_Reward; an inappropriate Target_Reward can cause significant variations in the model's decision performance, leading to unstable decision-making performance.

[0007] Therefore, it is necessary to further study existing UAV countermeasures decision-making methods to address the aforementioned issues. Summary of the Invention

[0008] To overcome the above problems, the inventors conducted in-depth research and proposed a UAV adversarial decision-making method based on experience and inference models. This method does not rely on expert datasets and can better consider the impact of decisions on the future. It solves the problems that traditional UAV adversarial decision-making algorithms based on reinforcement learning cannot well consider long-term planning, and that decision-making algorithms using sequential decision models require expert datasets and have unstable decision performance.

[0009] Specifically, the UAV adversarial decision-making method based on experience and deductive models includes the following steps:

[0010] S1. Establish the initial adversarial trajectory dataset;

[0011] S2. Construct a first model for predicting state-action pair sequences in the trajectory. Train the first model using the initial adversarial trajectory dataset to obtain a pre-decision model.

[0012] A second model is constructed to predict the sequence of states, actions, and rewards in adversarial trajectories. The second model is trained using the initial adversarial trajectory dataset to obtain the world model.

[0013] S3. Set simulation conditions and use a pre-decision model to perform online decision simulation to obtain past time rewards, the best action at the current time, and the second-best action at the current time.

[0014] S4

[0015] Input the past state, past action, past reward, current state, and current best action into the world model to obtain the reward value of the current best action and the next best state;

[0016] Input the past state, past action, past reward, current state, and current second-best action into the world model to obtain the reward value of the current second-best action and the next second-best state.

[0017] S5. The optimal state and the second-best state at the next moment are respectively used as inputs to the pre-decision model. The optimal action and the second-best action are obtained by the output of the pre-decision model. The obtained optimal action and the second-best action are used as inputs to the world model. The world model outputs the state at the next moment and the current reward.

[0018] S6. Repeat S4 and S5 multiple times to perform N-step predictions, obtaining 2 N A future path;

[0019] S7. Obtain the cumulative reward for each path, select the path with the largest cumulative reward, and use the initial action of that path as the decision for the current state.

[0020] S8. Repeat S4 to S7 at each time step to obtain the online adversarial decision for the complete time step;

[0021] S9. Change the simulation conditions and repeat S3 to S8 multiple times to obtain online adversarial decisions at multiple complete time points. Combine them into a second dataset. Use the second dataset to train the pre-decision model and the world model to obtain the optimized pre-decision model and the optimized world model.

[0022] The actual UAV combat conditions are input into the optimization pre-decision model, which outputs the optimal action at the current moment. The UAV then performs flight control based on the optimal action.

[0023] In a preferred embodiment, in S2, the first model includes a first input layer, a first fully connected layer, a first encoding layer, and a first decoder connected in sequence.

[0024] In a preferred embodiment, the first input layer is used to arrange the states and actions in the trajectory in chronological order to form a state-action sequence;

[0025] The first fully connected layer is used to embed the state-action sequence to reduce the dimensionality of the state-action sequence;

[0026] The first encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action sequence to represent the time corresponding to the state and action;

[0027] The first decoder includes multiple first decoder blocks connected end to end. Each first decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through the fully connected layer to obtain the reward and predicted action.

[0028] In a preferred embodiment, the second model includes a second input layer, a second fully connected layer, a second encoding layer, and a second decoder connected in sequence.

[0029] In a preferred embodiment, the second input layer is used to arrange the states, actions, and rewards in the trajectory in chronological order to form a state-action-reward sequence;

[0030] The second fully connected layer is used to embed the state-action-reward sequence to reduce the dimensionality of the state-action sequence;

[0031] The second encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action-reward sequence to represent the time corresponding to the state, action, and reward.

[0032] The second decoder includes multiple decoder blocks connected end to end. Each decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action-reward sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through a fully connected layer to obtain the state, reward, and predicted action.

[0033] In a preferred embodiment, when training the first model, the loss is calculated only for the actions and backpropagation is performed to update the parameters.

[0034] In a preferred embodiment, when training the second model, the loss is calculated only on the state and reward, and backpropagation is performed to update the parameters.

[0035] In a preferred embodiment, in S3, during the online decision simulation process, the state s from the previous time step is... t-2 Action a at the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 and the current state s tInput the pre-decision model, and the pre-decision model outputs the optimal action. and suboptimal actions

[0036] In a preferred embodiment, the state s at the previous time step t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the optimal action at the current moment Input the world model to obtain the optimal action at the current time. Reward value And the predicted optimal state at the next moment.

[0037] The state s from the previous time step t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the second-best action at the current moment Input the world model to obtain the second-best action at the current time step. Reward value And the predicted suboptimal state at the next moment.

[0038] In a preferred embodiment, in S9, if the configuration of our drone is the same as that of the target, the online adversarial decision of the winner is added to the second dataset; if the configuration of our drone is different from that of the target, the online adversarial decision of our drone is added to the second dataset.

[0039] The beneficial effects of this invention include:

[0040] (1) Compared with traditional methods, the present invention has stronger ability to characterize the environment and make decision-making reasoning.

[0041] (2) By predicting future states and rewards, the long-term impact of decisions on the environment can be fully considered, thereby enabling drones to make decisions that are more conducive to the long term. This improves, to some extent, the problem that existing drone-based decision-making methods do not adequately consider the long-term impact of decisions on the environment.

[0042] (3) By considering long-term decision-making, it can make decisions that are superior to those made on adversarial trajectory datasets, thereby improving performance. Therefore, it is not necessary to strictly use expert datasets for training.

[0043] (4) By training the pre-decision model, the available dimensions of the action space are greatly reduced, which significantly reduces the resources required for long-term planning of high-dimensional action space environments compared to long-term planning methods based on Monte Carlo tree search. Attached Figure Description

[0044] Figure 1 A schematic diagram of the process of a UAV adversarial decision-making method based on experience and deduction model according to a preferred embodiment of the present invention is shown.

[0045] Figure 2 The diagram shows a comparison between Example 1 and Comparative Example 1. Detailed Implementation

[0046] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Through these descriptions, the features and advantages of the present invention will become clearer and more apparent.

[0047] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments. Although various aspects of embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless specifically indicated otherwise.

[0048] According to the present invention, a UAV adversarial decision-making method based on experience and deductive models is provided, such as... Figure 1 As shown, it includes the following steps:

[0049] S1. Establish the initial adversarial trajectory dataset;

[0050] S2. Construct a first model for predicting state-action pair sequences in the trajectory. Train the first model using the initial adversarial trajectory dataset to obtain a pre-decision model.

[0051] A second model is constructed to predict the sequence of states, actions, and rewards in adversarial trajectories. The second model is trained using the initial adversarial trajectory dataset to obtain the world model.

[0052] S3. Set simulation conditions and use a pre-decision model to perform online decision simulation to obtain past time rewards, the best action at the current time, and the second-best action at the current time.

[0053] S4. Input the past state, past action, past reward, current state, and current best action into the world model to obtain the reward value of the current best action and the next best state.

[0054] Input the past state, past action, past reward, current state, and current second-best action into the world model to obtain the reward value of the current second-best action and the next second-best state.

[0055] S5. The optimal state and the second-best state at the next moment are respectively used as inputs to the pre-decision model. The optimal action and the second-best action are obtained by the output of the pre-decision model. The obtained optimal action and the second-best action are used as inputs to the world model. The world model outputs the state at the next moment and the current reward.

[0056] S6. Repeat S4 and S5 multiple times to perform N-step predictions, obtaining 2 N A future path;

[0057] S7. Obtain the cumulative reward for each path, select the path with the largest cumulative reward, and use the initial action of that path as the decision for the current state.

[0058] S8. Repeat S4 to S7 at each time step to obtain the online adversarial decision for the complete time step;

[0059] S9. Change the simulation conditions and repeat S3 to S8 multiple times to obtain online adversarial decisions at multiple complete time points. Combine them into a second dataset. Use the second dataset to train the pre-decision model and the world model to obtain the optimized pre-decision model and the optimized world model.

[0060] The actual UAV combat conditions are input into the optimization pre-decision model, which outputs the optimal action at the current moment. The UAV then performs flight control based on the optimal action.

[0061] In S1, there are no restrictions on the method for establishing the initial adversarial trajectory dataset. It can be used to make decisions in the corresponding adversarial environment through reinforcement learning or other decision-making methods and collect the decision trajectory; or an existing adversarial trajectory dataset can be used.

[0062] According to the present invention, the samples of the adversarial trajectory dataset are the trajectories of a complete adversarial exercise between our UAV and the target, including the state sequence, action sequence, reward sequence, next state sequence, and completion flag sequence of all time steps of the UAV.

[0063] In S2, the first model includes a first input layer, a first fully connected layer, a first encoding layer, and a first decoder connected in sequence.

[0064] The first input layer is used to arrange the states and actions in the trajectory in chronological order to form a state-action sequence, preferably forming {state}. 时刻1 ,action 时刻1 ,state时刻2 ,action 时刻2 The state-action sequence of ,...}.

[0065] The first fully connected layer is used to embed the state-action sequence to reduce the dimensionality of the state-action sequence. The fully connected layer is a common structure in neural networks, and its specific structure will not be described in detail in this invention.

[0066] The first encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action sequence to represent the time corresponding to the state and action.

[0067] Preferably, in the first coding layer, the state and action are segmented and coded to distinguish the state and action features in the sequence.

[0068] The first decoder includes multiple first decoder blocks connected end to end. Each first decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through the fully connected layer to obtain the reward and predicted action.

[0069] Preferably, the self-attention module is a masked multi-head self-attention module.

[0070] According to the present invention, the purpose of training the first model is to enable the first model to learn the decision-making experience in the trajectory dataset and to make preliminary decisions based on the state.

[0071] Specifically, the first model is trained using the initial adversarial trajectory dataset. The predicted actions are compared with the actual actions in the dataset to calculate the loss, and gradient backpropagation is performed to update the parameters, thus obtaining the pre-decision model f. θ The model parameter is θ. The pre-decision model can learn from the experience of the dataset and generate decision actions based on a sequence of states.

[0072] Preferably, in S2, when training the first model, only the loss of the action is calculated and backpropagation is performed to update the parameters, so as to simplify the training difficulty.

[0073] In S2, the second model includes a second input layer, a second fully connected layer, a second coding layer, and a second decoder connected in sequence.

[0074] The second input layer is used to arrange the states, actions, and rewards in the trajectory in chronological order to form a state-action-reward sequence, preferably forming a {state} sequence. 时刻1 ,action 时刻1 ,award 时刻1 ,state 时刻2 ,action 时刻2 ,award 时刻2 ,

[0075] ...} is a sequence of state-action rewards.

[0076] The second fully connected layer is used to embed the state-action-reward sequence to reduce the dimensionality of the state-action sequence.

[0077] The second encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action-reward sequence to represent the time corresponding to the state, action, and reward.

[0078] Preferably, in the coding layer, the state, action, and reward are segmented and encoded to facilitate the differentiation of state, action, and reward features in the sequence.

[0079] The second decoder includes multiple decoder blocks connected end to end. Each decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action-reward sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through the fully connected layer to obtain the state, reward, and predicted action.

[0080] Preferably, the self-attention module is a masked multi-head self-attention module.

[0081] According to the present invention, the purpose of training the second model is to enable the second model to learn the dynamics of the adversarial environment and the reward mechanism, and to be able to infer the upcoming next state and the resulting reward based on the current state and decision actions.

[0082] According to the present invention, for our UAV, the target can be regarded as part of the environment, so the behavior of the target is regarded as part of the dynamic process of resisting the environment, and the behavior of the target is predicted while training the dynamic process of the environment.

[0083] Specifically, the second model is trained using the states, actions, and rewards from the initial adversarial trajectory dataset. The predicted states and rewards are compared with the actual states and rewards in the dataset as a loss, and gradient backpropagation is performed to update the parameters, thus obtaining the world model. Model parameters are The world model can deduce the next state based on the current state and actions, and evaluate the reward obtained in the next state.

[0084] In a preferred embodiment, during S2, when training the second model, the loss is calculated only for the state and reward, and backpropagation is performed to update the parameters, thereby simplifying the training difficulty.

[0085] According to the present invention, by training a pre-decision model, the available dimensions of the action space are significantly reduced, thereby greatly reducing the resources required for long-term planning of high-dimensional action space environments compared to long-term planning methods based on Monte Carlo tree search.

[0086] In S3, the simulation conditions can be set by those skilled in the art according to actual needs, such as the position and speed of our drone, the position and speed of the target, etc.

[0087] It should be noted that in setting the simulation conditions in S3, the drones and targets need to be of the same type and configuration as the drones and targets in the initial adversarial trajectory dataset. That is, the parameters of the pre-decision model and world model are only applicable to one adversarial configuration. When the type and configuration of the adversarial drones and targets are changed, step S2 needs to be repeated.

[0088] According to the present invention, in S3, during the online decision-making simulation process, the past time state, past time action, and current time state s are... t The input is a pre-decision model, which outputs the optimal and suboptimal actions. Furthermore, at the initial moment of the online decision simulation, since there is no previous state or action, only the current state needs to be input into the pre-decision network.

[0089] Preferably, the state s at the previous time step t-2 Action a at the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 and the current state s t Input the pre-decision model, and the pre-decision model outputs the optimal action. and suboptimal actions

[0090] In S4, preferably, the state s from the previous time step is... t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the optimal action at the current moment Input the world model to obtain the optimal action at the current time. Reward value And the predicted optimal state at the next moment.

[0091] Similarly, the state s from the previous time step t-2 Action a at the previous momentt-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the second-best action at the current moment Input the world model to obtain the second-best action at the current time step. Reward value And the predicted suboptimal state at the next moment.

[0092] In S5, preferably, the optimal state at the previous moment, the optimal action at the previous moment, the optimal state at the current moment, the optimal action at the current moment, and the optimal state at the next moment are used as inputs to the pre-decision model, and the pre-decision model outputs the optimal action at the next moment.

[0093] Similarly, the optimal state at the previous moment, the optimal action at the previous moment, the optimal state at the current moment, the optimal action at the current moment, and the suboptimal state at the next moment are used as inputs to the pre-decision model, and the pre-decision model outputs the suboptimal action at the next moment.

[0094] The state s from the previous moment t- The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t The action a at the current moment t The reward r at the current moment t The state s at the next moment t+1 And the optimal action in the next moment Input the world model to obtain the reward value for the optimal action in the next time step. And the predicted optimal state at the next next time step.

[0095] Similarly, the state s from the previous moment... t- The action a at the previous moment r-1 The reward r from the previous moment t-1 The current state s t The action a at the current moment t The reward r at the current moment t The state s at the next moment t+1 And the next optimal action Input the world model to obtain the reward value for the second-best action in the next time step. And the predicted suboptimal state at the next next time step.

[0096] By repeating S4 to S5 multiple times, N-step prediction can be completed.

[0097] In S9, preferably, if the configuration of our drone and the target drone are the same, then the winning drone's online adversarial decision is added to the second dataset; if the configuration of our drone and the target drone are different, then the online adversarial decision of our drone is added to the second dataset. Since each step of the online decision predicts N future time steps and selects the action that maximizes the total future reward as the decision action, it can be considered that the decision trajectory obtained by the online decision is better than or no worse than the decision trajectory used for training. By predicting future states and rewards, the long-term impact of decisions on the environment can be fully considered, thus enabling adversarial drones to make decisions that are more beneficial in the long run, and to some extent improving the problem that existing drone adversarial decision-making methods do not adequately consider the long-term environmental impact of their decisions.

[0098] Furthermore, the optimized pre-decision model trained using the second dataset, by considering long-term decision-making, can make decisions superior to those made using the adversarial trajectory dataset, thus achieving performance improvement. Therefore, it is not necessary to strictly use the expert dataset for training.

[0099] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.

[0100] Example

[0101] Example 1

[0102] The simulation experiment includes the following steps:

[0103] S1. Establish the initial adversarial trajectory dataset;

[0104] S2. Construct a first model for predicting state-action pair sequences in the trajectory. Train the first model using the initial adversarial trajectory dataset to obtain a pre-decision model.

[0105] A second model is constructed to predict the sequence of states, actions, and rewards in adversarial trajectories. The second model is trained using the initial adversarial trajectory dataset to obtain the world model.

[0106] S3. Set simulation conditions and use a pre-decision model to perform online decision simulation to obtain past time rewards, the best action at the current time, and the second-best action at the current time.

[0107] S4. Input the past state, past action, past reward, current state, and current best action into the world model to obtain the reward value of the current best action and the next best state.

[0108] Input the past state, past action, past reward, current state, and current second-best action into the world model to obtain the reward value of the current second-best action and the next second-best state.

[0109] S5. Using the optimal state and the second-best state at the next time step as inputs to the pre-decision model, the pre-decision model outputs the optimal and second-best actions for subsequent time steps. These optimal and second-best actions are then used as inputs to the world model, which outputs the state and reward for the next time step. S6. Repeating S4 and S5 multiple times for N steps of prediction yields 2... N A future path;

[0110] S7. Obtain the cumulative reward for each path, select the path with the largest cumulative reward, and use the initial action of that path as the decision for the current state.

[0111] S8. Repeat S4 to S7 at each time step to obtain the online adversarial decision for the complete time step;

[0112] S9. Change the simulation conditions and repeat S3 to S8 multiple times to obtain online adversarial decisions at multiple complete time points. Combine them into a second dataset. Use the second dataset to train the pre-decision model and the world model to obtain the optimized pre-decision model and the optimized world model.

[0113] The actual UAV combat conditions are input into the optimization pre-decision model, which outputs the optimal action at the current moment. The UAV then performs flight control based on the optimal action.

[0114] In S2, the first model includes a first input layer, a first fully connected layer, a first encoder layer, and a first decoder connected in sequence; the first input layer is used to arrange the states and actions in the trajectory in chronological order to form a state-action sequence;

[0115] The first fully connected layer is used to embed the state-action sequence to reduce the dimensionality of the state-action sequence;

[0116] The first encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action sequence to represent the time corresponding to the state and action;

[0117] The first decoder includes multiple first decoder blocks connected end to end. Each first decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through the fully connected layer to obtain the reward and predicted action.

[0118] The second model includes a second input layer, a second fully connected layer, a second encoder layer, and a second decoder connected in sequence. The second input layer is used to arrange the state, action, and reward in the trajectory in chronological order to form a state-action-reward sequence.

[0119] The second fully connected layer is used to embed the state-action-reward sequence to reduce the dimensionality of the state-action sequence;

[0120] The second encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action-reward sequence to represent the time corresponding to the state, action, and reward.

[0121] The second decoder includes multiple decoder blocks connected end to end. Each decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action-reward sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through a fully connected layer to obtain the state, reward, and predicted action.

[0122] In S3, during the online decision-making simulation, the state s from the previous time step is... t-2 Action a at the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 and the current state s t Input the pre-decision model, and the pre-decision model outputs the optimal action. and suboptimal actions

[0123] The state s from the previous time step t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the optimal action at the current moment Input the world model to obtain the optimal action at the current time. Reward value And the predicted optimal state at the next moment.

[0124] The state s from the previous time step t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the second-best action at the current moment Input the world model to obtain the second-best action at the current time step. Reward value And the predicted suboptimal state at the next moment.

[0125] The initial adversarial trajectory dataset was obtained by making decisions in a simulation environment using the traditional reinforcement learning algorithm SAC (Soft-Actor-Critic).

[0126] Comparative Example 1

[0127] Under the same simulation experimental environment and evaluation function as in Example 1, the pre-decision model is directly trained using the trajectory dataset obtained by the SAC method, and the optimal action is output by the pre-decision model.

[0128] The results of Example 1 and Comparative Example 1 are as follows: Figure 2 As shown, the blue line represents the air combat evaluation score of the trajectory obtained by the SAC method as a function of the number of network updates, the orange line represents the air combat evaluation score of Comparative Example 1 as a function of the number of network updates, and the green line represents the air combat evaluation score of Example 1 as a function of the number of network updates. According to the evaluation rules, the greater the energy advantage of the UAV, the shorter the successful tail-chase time, and the higher the score.

[0129] from Figure 2 It can be seen that the method in Example 1 can achieve a higher evaluation score.

[0130] The present invention has been described above with reference to preferred embodiments; however, these embodiments are merely exemplary and illustrative. Various substitutions and modifications can be made to the present invention based on these embodiments, all of which fall within the scope of protection of the present invention.

Claims

1. A UAV adversarial decision-making method based on experience and inference models, characterized in that, Includes the following steps: S1. Establish the initial adversarial trajectory dataset; S2. Construct a first model for predicting state-action pair sequences in the trajectory. Train the first model using the initial adversarial trajectory dataset to obtain a pre-decision model. A second model is constructed to predict the sequence of states, actions, and rewards in adversarial trajectories. The second model is trained using the initial adversarial trajectory dataset to obtain the world model. S3. Set simulation conditions and use a pre-decision model to perform online decision simulation to obtain past time rewards, the best action at the current time, and the second-best action at the current time. S4. Input the past state, past action, past reward, current state, and current best action into the world model to obtain the reward value of the current best action and the next best state. Input the past state, past action, past reward, current state, and current second-best action into the world model to obtain the reward value of the current second-best action and the next second-best state. S5. The optimal state and the second-best state at the next time step are respectively used as one of the inputs to the pre-decision model. The optimal action and the second-best action at the next time step are obtained by the output of the pre-decision model. The obtained optimal action and the second-best action are used as one of the inputs to the world model. The world model outputs the state at the next time step and the reward at the next time step. S6. Repeat S4 and S5 multiple times to perform N-step predictions, obtaining 2 N A future path; S7. Obtain the cumulative reward for each path, select the path with the largest cumulative reward, and use the initial action of that path as the decision for the current state. S8. Repeat S4 to S7 at each time step to obtain the online adversarial decision for the complete time step; S9. Change the simulation conditions and repeat S3 to S8 multiple times to obtain online adversarial decisions at multiple complete time points. Combine them into a second dataset. Use the second dataset to train the pre-decision model and the world model to obtain the optimized pre-decision model and the optimized world model. The actual UAV combat conditions are input into the optimization pre-decision model, which outputs the optimal action at the current moment. The UAV then performs flight control based on the optimal action.

2. The UAV adversarial decision-making method based on experience and deductive models according to claim 1, characterized in that, In S2, the first model includes a first input layer, a first fully connected layer, a first coding layer, and a first decoder connected in sequence.

3. The UAV adversarial decision-making method based on experience and deductive models according to claim 2, characterized in that, The first input layer is used to arrange the states and actions in the trajectory in chronological order to form a state-action sequence; The first fully connected layer is used to embed the state-action sequence to reduce the dimensionality of the state-action sequence; The first encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action sequence to represent the time corresponding to the state and action; The first decoder includes multiple first decoder blocks connected end to end. Each first decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through the fully connected layer to obtain the reward and predicted action.

4. The UAV adversarial decision-making method based on experience and deductive models according to claim 1, characterized in that, The second model includes a second input layer, a second fully connected layer, a second encoding layer, and a second decoder connected in sequence.

5. The UAV adversarial decision-making method based on experience and deductive models according to claim 4, characterized in that, in, The second input layer is used to arrange the states, actions, and rewards in the trajectory in chronological order to form a state-action-reward sequence; The second fully connected layer is used to embed the state-action-reward sequence to reduce the dimensionality of the state-action sequence; The second encoding layer is used to perform triangular position encoding on the reduced-dimensional state-action-reward sequence to represent the time corresponding to the state, action, and reward. The second decoder includes multiple decoder blocks connected end to end. Each decoder block includes a self-attention module and a fully connected layer. The self-attention module extracts features from the state-action-reward sequence after triangular position encoding, performs residual connection and normalization processing, and then outputs through a fully connected layer to obtain the state, reward, and predicted action.

6. The UAV adversarial decision-making method based on experience and deductive models according to claim 1, characterized in that, When training the first model, only the loss of the action is calculated and backpropagation is performed to update the parameters.

7. The UAV adversarial decision-making method based on experience and deductive models according to claim 1, characterized in that, When training the second model, only the loss is calculated on the state and reward, and backpropagation is performed to update the parameters.

8. The UAV adversarial decision-making method based on experience and deductive models according to claim 1, characterized in that, In S3, during the online decision-making simulation, the state s from the previous time step is... t-2 Action a at the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 and the current state s t Input the pre-decision model, and the pre-decision model outputs the optimal action. and suboptimal actions 9. The UAV adversarial decision-making method based on experience and deductive models according to claim 8, characterized in that, The state s from the previous time step t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the optimal action at the current moment Input the world model to obtain the optimal action at the current time. Reward value And the predicted optimal state at the next moment. The state s from the previous time step t-2 Action a at the previous moment t-2 The reward r of the previous moment t-2 The state s of the previous moment t-1 The action a at the previous moment t-1 The reward r from the previous moment t-1 The current state s t and the second-best action at the current moment Input the world model to obtain the second-best action at the current time step. Reward value And the predicted suboptimal state at the next moment.

10. The UAV adversarial decision-making method based on experience and deductive models according to claim 1, characterized in that, In S9, if our drone has the same model configuration as the target, the winning side's online adversarial decision is added to the second dataset; if our drone has a different model configuration than the target, the winning side's online adversarial decision is added to the second dataset.

Citation Information

Patent Citations

  • Remote sensing object detection method based on few samples

    CN107657279A

  • Model-enhanced unmanned aerial vehicle flight path reinforcement learning optimization method

    CN114879738A