A three-dimensional trajectory planning method for anti-tank missiles based on deep reinforcement learning

Through the method based on deep reinforcement learning, the problem that existing aircraft trajectory planning algorithms cannot be efficiently solved under large-scale and wide-area conditions is solved, efficient three-dimensional trajectory planning is achieved, and the motion planning ability and robustness of the aircraft are improved.

CN118466207BActive Publication Date: 2025-06-06HARBIN INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410629598.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-06-06
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

The existing aircraft track planning algorithm cannot be efficiently solved under large-scale and wide-area conditions, and there are problems of randomness, high complexity and poor versatility.

Method used

The three-dimensional trajectory planning method of anti-tank missiles based on deep reinforcement learning is adopted. By establishing a nonlinear dynamic model, relative kinematic model and deep reinforcement learning framework, the state space and action space of agent training are designed, and reward functions and neural networks are constructed to realize the training and testing of ballistic planning models.

Benefits of technology

It realizes efficient generation of three-dimensional trajectories under large-scale and wide-area conditions, avoids complex solution processes in traditional methods, and improves the motion planning ability and robustness of the aircraft.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118466207B_ABST
    Figure CN118466207B_ABST
Patent Text Reader

Abstract

The invention relates to the field of aircraft guidance and control, and discloses a three-dimensional trajectory planning method for an anti-tank missile based on deep reinforcement learning. The method comprises the following steps: establishing a nonlinear dynamic model of the anti-tank missile; establishing a relative kinematic model of the anti-tank missile and a target; establishing a deep reinforcement learning framework for anti-tank missile training; establishing a state space and an action space for intelligent agent training; establishing a reward function and a neural network for striking a fixed target point; and designing a trajectory planning model training and testing method based on deep reinforcement learning. The deep reinforcement learning algorithm of the invention directly learns flight attack angle and sideslip angle instructions, thereby avoiding the complex solution process from overload instructions to attack angle / sideslip angle instructions in a traditional trajectory planning method. The intelligent trajectory online planning algorithm can give full play to the strike capability of the anti-tank missile and realize three-dimensional trajectory planning with full coverage of the anti-tank missile range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aircraft guidance and control, and in particular to a three-dimensional trajectory planning method for an anti-tank missile based on deep reinforcement learning. Background Art

[0002] Aircraft online trajectory planning is a key technology for realizing autonomous navigation of intelligent aircraft, and is one of the important research directions in the field of artificial intelligence and navigation and guidance control. When an intelligent aircraft performs a certain flight mission, planning a feasible trajectory is the most basic task requirement of the intelligent aircraft. According to different mission requirements, the requirements for trajectory planning will also be different, such as the shortest flight trajectory, the minimum flight time, the lowest energy consumption, etc. Higher mission requirements and complex flight environments pose new challenges to aircraft trajectory planning. Aircraft trajectory planning is essentially a multi-constrained optimization problem. At present, the widely used motion planning algorithms mainly include A* algorithm and D* algorithm based on environmental model, search-based random path map method (PRM) and rapid exploration random tree method (RRT), strategy-based fuzzy logic method and dynamic window method, as well as genetic algorithm, ant colony algorithm and bee colony algorithm based on bionic planning algorithm. Generally, under the conditions of known map, static obstacles and simple environment, these motion planning algorithms can complete simple tasks through environmental modeling or probabilistic search. However, the above algorithms have certain randomness and high complexity, and cannot be solved efficiently under large-scale and wide-area conditions. In addition, traditional planning methods are mostly customized algorithms, and there are many problems such as large program size, poor versatility, and high power consumption.

[0003] In order to solve the above problems, the intelligent aircraft trajectory planning strategy based on deep reinforcement learning theory can effectively avoid the defects of the trajectory planning algorithm in the prior art, such as randomness, high complexity, and inability to efficiently solve under large-scale and wide-area conditions. With the increase in the complexity of the aircraft's working environment, the increase in randomness, and the decrease in the amount of information, the aircraft's motion planning ability has been severely challenged. Researching efficient and autonomous motion planning theories and methods for intelligent aircraft so that they can always maintain good adaptability to complex environments during long-term missions is of great significance to ensuring flight safety and improving mission execution efficiency. Designing an intelligent aircraft motion planning method with autonomous decision-making capabilities, thereby making up for the defects of traditional motion planning methods and improving the robustness and generalization ability of aircraft motion planning methods, is one of the problems that need to be solved urgently. Based on this, the present invention proposes a three-dimensional trajectory planning method for anti-tank missiles based on deep reinforcement learning. Summary of the invention

[0004] The purpose of the present invention is to provide a three-dimensional trajectory planning method for an anti-tank missile based on deep reinforcement learning to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: a three-dimensional trajectory planning method for an anti-tank missile based on deep reinforcement learning, comprising the following steps:

[0006] Establish a nonlinear dynamic model of anti-tank missile;

[0007] Establish the relative kinematic model between anti-tank missile and target;

[0008] Establish a deep reinforcement learning framework for anti-tank missile training;

[0009] Establish state space and action space for agent training;

[0010] Establish a reward function and neural network for striking fixed target points;

[0011] Design a training and testing method for trajectory planning models based on deep reinforcement learning.

[0012] Preferably: the three-dimensional anti-tank missile dynamics model is as follows,

[0013]

[0014] Preferably, the relative kinematic model establishment process between the anti-tank missile and the target is as follows: in the line of sight coordinate system, R is defined as the relative distance between the target and the missile, q ε is the line of sight inclination, q β is the line of sight deflection angle, the relative distance between the missile and the target and its derivative are represented by R and Indicates that there is

[0015]

[0016] where x r =x t -x m ,y r =y t -y m ,z r =z t -z m .

[0017] Sight angle q ε ,q β and line of sight angular velocity The expression is

[0018]

[0019] Among them, q ε ,q β are the line of sight angles in pitch and yaw directions, respectively.

[0020] Preferably: In the deep reinforcement learning framework for anti-tank missile training, the Q value function Q π (s t ,a t ) is used to indicate that in state s t Next, take action a t Expected value of revenue that can be achieved

[0021] Q π (s t ,a t )=E π [R t |s t ,a t ]

[0022] Use a deep neural network to approximate the Q value in equation (11), assuming θ Q is the Q-value network hyperparameter, then

[0023] Q π (s t ,a t )≈Q π (s t ,a t |θ Q ) (12)

[0024] The anti-tank missile model is trained using a deep deterministic policy gradient algorithm, where the network Q value is expressed as

[0025] y t =r t +γQ(s t+1 ,π(s t+1 |φ′ π )|θ′ Q ) (13)

[0026] where θ′ Q and φ′ π is the hyperparameter of the target Q value network and V value network. The policy network parameter update method in the DDPG algorithm is

[0027]

[0028] Where N is the sample size.

[0029] The action network parameter update method in the DDPG algorithm is

[0030]

[0031] The target network is updated using the following soft update method:

[0032]

[0033] Where τ is the soft update factor.

[0034] Preferred: Based on the anti-tank missile of STT configuration, the angle of attack and sideslip angle are selected as the aircraft action space, as follows

[0035]

[0036] Taking into account the differences in the characteristics of the aircraft model and the order of magnitude of the variables, the state and action actually observed by the environment are

[0037]

[0038] Among them, R,q ε ,q β is the missile-target relative distance, the achieved inclination angle and the line of sight deviation angle, max_α and max_β are the maximum flight angle of attack and sideslip angle of the STT missile. Through the above transformation, the order of magnitude of the system state variables observed by the environment is between [-100, 100], and the order of magnitude of the action variables is between [-1, 1].

[0039] Preferably: the reward function includes the following three parts: 1) distance reward function

[0040]

[0041] Among them, R(t-1) and R(t) are the relative distances between the projectile and the target at time t-1 and t respectively. is the velocity vector magnitude w 1 is the reward weight coefficient, the default value is 1; 2) Heading angle reward function

[0042]

[0043] Where: and are respectively the ballistic deviation angle at time t and the expected ballistic inclination angle at time t, and the expression is

[0044]

[0045] where x t arg et ,y t arg et ,z t arg et is the position coordinate of the target point in the ground coordinate system; 3) Process constraint reward function

[0046]

[0047] Where: t max , R maxis the maximum training step length and the maximum allowed miss amount. The training end reward function is triggered when the missile successfully intercepts the target or other termination conditions (such as training step length timeout, aircraft state singularity, etc.) occur. The comprehensive reward function of the agent training process is

[0048]

[0049] Preferably: the neural network includes 4 neural networks, namely 2 action networks and 2 evaluation networks. The Actor action network (target action network) adopts a 4-layer fully connected layer BP neural network structure of 6-256-256-2, and the Critic evaluation network (target evaluation network) adopts a 4-layer BP neural network structure of 8-256-256-1. The Relu function is used as the hidden layer activation function, and the network parameters adopt the Adam adaptive gradient optimization algorithm.

[0050] Preferably, the steps of the trajectory planning model training method based on deep reinforcement learning are as follows:

[0051] 1) Set the physical properties and basic performance parameters of anti-tank missiles;

[0052] 2) Design the DDPG deep reinforcement learning framework and initialize the network parameters;

[0053] 3) Initialize the initial state and target point position of the anti-tank missile;

[0054] 4) Input the current state information into the DDPG algorithm to calculate the current guidance instructions;

[0055] 5) Update the aircraft status and missile-target relative information;

[0056] 6) Calculate the reward function and store the current state, action, reward value and next state in the experience pool;

[0057] 7) Randomly sample historical data from the experience pool and update DDPG network parameters through certain learning rules;

[0058] 8) Determine whether the stop condition of this round is met. If so, reinitialize the aircraft state, otherwise proceed to the next step;

[0059] 9) Determine whether the reward function converges. If so, stop training the agent. Otherwise, go to step 4).

[0060] Preferably, the steps of the trajectory planning model testing method based on deep reinforcement learning are as follows:

[0061] 1) Randomly initialize the initial state of the anti-tank missile;

[0062] 2) Randomly initialize the target point position;

[0063] 3) Importing Actor network in DDPG framework;

[0064] 4) Input the current missile-target relative information into the Actor network and calculate the guidance instructions;

[0065] 5) Update the flight status of anti-tank missiles;

[0066] 6) Determine whether the off-target condition is met, if so, stop the test, otherwise proceed to step 4).

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] 1) Randomly initialize the initial state and target point position of the anti-tank missile to generate a three-dimensional trajectory of the target online;

[0069] 2) The deep reinforcement learning algorithm directly learns the flight angle of attack and sideslip angle commands, avoiding the complex solution process from overload commands to angle of attack / sideslip angle commands in traditional trajectory planning methods;

[0070] 3) Intelligent trajectory online planning algorithm can give full play to the strike capability of anti-tank missiles and realize three-dimensional trajectory planning with full coverage of the anti-tank missile range. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is a flowchart of three-dimensional trajectory planning for anti-tank missiles based on deep reinforcement learning;

[0072] Figure 2 It is a schematic diagram of the three-dimensional trajectory planning of the anti-tank missile;

[0073] Figure 3 This is a schematic diagram of the DDPG deep reinforcement learning algorithm;

[0074] Figure 4 It is a graph showing the change of reward function during the agent training process;

[0075] Figure 5 It is the change diagram of the relative distance between the projectile and the target during the training process of the intelligent agent;

[0076] Figure 6 It is the three-dimensional trajectory diagram of the anti-tank missile planned by the DDPG algorithm;

[0077] Figure 7 It is the state variable diagram of the anti-tank missile;

[0078] Figure 8 It is the guidance instruction graph generated by the DDPG algorithm;

[0079] Fig. 9 It is the diagram of the change of the relative distance between the projectile and the target;

[0080] Fig.10 It is the trajectory planning diagram of 50 Monte Carlo simulations with random changes in the target point;

[0081] Fig.11 This is the state variable diagram of the anti-tank missile after 50 Monte Carlo simulations;

[0082] Fig.12 This is the guidance command graph generated by DDPG after 50 Monte Carlo simulations;

[0083] Fig.13 This is a graph showing the change in the relative distance between missile and target after 50 Monte Carlo simulations. DETAILED DESCRIPTION

[0084] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0085] Example

[0086] See also Figure 1-Figure 13 , a three-dimensional trajectory planning method for an anti-tank missile based on deep reinforcement learning is shown in the figure, comprising the following steps:

[0087] Establish a nonlinear dynamic model of anti-tank missile;

[0088] Establish the relative kinematic model between anti-tank missile and target;

[0089] Establish a deep reinforcement learning framework for anti-tank missile training;

[0090] Establish state space and action space for agent training;

[0091] Establish a reward function and neural network for striking fixed target points;

[0092] Design a training and testing method for trajectory planning models based on deep reinforcement learning.

[0093] In this embodiment, the establishment of the relative kinematic model between the anti-tank missile and the target first requires the definition of the coordinate system, including the ground coordinate system: the coordinate system Axyz fixed to the earth's surface is the ground coordinate system. The center of mass of the missile at the moment of launch is selected as the coordinate origin A, and the intersection of the ballistic plane and the horizontal plane is selected as the Ax axis, and the direction pointing to the target is positive; the Ay axis is perpendicular to the Ax axis and points upward, and the Az axis is given according to the right-hand rule, thereby determining the coordinate system Axyz; the projectile coordinate system: it is fixed to the projectile, moves and rotates in space with the projectile, and is a dynamic coordinate system. Usually its coordinate origin O is selected at the center of mass of the missile, Ox 1 The axis is generally parallel to the symmetry axis of the missile body or parallel to the average aerodynamic chord of the wing, and the direction to the head is positive; 1 The axis is located in the left-right symmetric plane of the projectile, Oy 1 Axis and Ox 1 The axes are vertical and contain the origin, Oy 1 Axis is positive; Oz 1 The axis is determined according to the right-hand rule; ballistic coordinate system: the origin O of the coordinate system is selected at the center of mass of the missile (target), which is a dynamic coordinate system that moves in space with the center of mass of the missile. 2 The axis is in the same direction as the velocity vector V of the missile's mass center. 2 Axis through Ox 2 The vertical plane of the axis is in contact with Ox 2 Vertical, the direction away from the center of the earth is positive; Oz 2 The axis is determined by the right-hand rule; the sight line coordinate system: the sight line of the projectile is selected as Ox 4 Axis, where the origin O is selected at the missile's center of mass. The direction in which the missile points to the target is positive. 4 Axis and Ox 4 The axis is vertical and located at the 4 On the vertical plane of the axis, Oy 4 Positive is pointing upward. 4 The axis is determined by the right-hand rule. The relative motion equation between the missile and the target in the line of sight direction is mainly established in the line of sight coordinate system. The missile and the target are fixed to the line of sight coordinate system, which is a moving coordinate system.

[0094] Furthermore, based on the basic principles of flight mechanics, dynamic equations describing the motion of anti-tank missiles and targets are established. The missile motion equations also include kinematic equations describing the relationship between various motion parameters. To determine the motion trajectory (trajectory) of the missile's center of mass relative to the ground coordinate system, it is necessary to establish the kinematic equations of the missile's center of mass relative to the ground coordinate system. When calculating aerodynamic forces and thrust, it is necessary to know the height of the missile at any instant, and determine the transient position through trajectory calculation. Therefore, it is necessary to establish the position equation of the missile's center of mass relative to the ground coordinate system Axyz.

[0095]

[0096] Among them, x, y, and z are the three-dimensional positions of the aircraft in the ground coordinate system.

[0097] According to the definition of the ballistic coordinate system, the velocity vector of the missile's center of mass coincides with the axis of the ballistic coordinate system, so

[0098]

[0099] Among them, V x2 ,V y2 ,V z2 is the component of the vehicle velocity vector in the ballistic coordinate system.

[0100] Using the conversion relationship between the ground coordinate system and the ballistic coordinate system, we can get

[0101]

[0102] Then there is

[0103]

[0104] in They are flight speed, ballistic inclination and ballistic deviation respectively.

[0105] The aircraft dynamics model is

[0106]

[0107] In the formula, F x2 ,F y2 ,F z2 All external forces of the missile except thrust (total aerodynamic force R, gravity G, etc.) are respectively 2 y 2 z 2 The algebraic sum of the components on the axis, P x2 ,P y2 ,P z2 The thrust is in Ox 2 y 2 z 2 Substituting the expressions of aerodynamic force R, gravity G and thrust P into equation (5) yields

[0108]

[0109] Where X, Y, Z are the drag, lift and lateral force on the aircraft, α, β, γ V are the angle of attack, sideslip angle and tilt angle. In summary, the following three-dimensional anti-tank missile dynamic model can be obtained:

[0110]

[0111] In this embodiment, the process of establishing the relative kinematic model between the anti-tank missile and the target is as follows: in the line of sight coordinate system, R is defined as the relative distance between the target and the missile, q ε is the line of sight inclination, q β is the line of sight deflection angle. The relative distance between the missile and the target and its derivative are represented by R and Indicates that there is

[0112]

[0113] where x r =x t -x m ,y r =y t -y m ,z r =z t -z m .

[0114] Sight angle q ε ,q β and line of sight angular velocity The expression is

[0115]

[0116] Among them, q ε ,q β are the line of sight angles in pitch and yaw directions, respectively.

[0117] In this embodiment, a deep reinforcement learning framework for anti-tank missile training is established. Reinforcement learning theory is an intelligent heuristic optimization algorithm that aims to maximize expected returns. By simulating human trial and error behavior, a large amount of historical experience data on the interaction between the intelligent agent and the external environment is obtained, and the optimal strategy is generated through certain optimization rules.

[0118] A typical reinforcement learning process can usually be represented by a 5-tuple <S,A,P,R,γ RL >, where S, A, P, and R are the system state space, action space, system state transition probability, and reward function space, respectively, and γ RL is the discount factor. The goal of reinforcement learning is to find the optimal strategy π * :S→A maximizes the following expected return

[0119]

[0120] Q-value function Q π (s t ,a t ) is used to indicate that in state s t Next, take action a tExpected value of revenue that can be achieved

[0121] Q π (s t ,a t )=E π [R t |s t ,a t ] (11)

[0122] In order to solve the dimensionality curse problem that occurs in classical reinforcement learning as the system state dimension or action dimension increases, a deep neural network is usually used to approximate the Q value in equation (11), thus forming the deep reinforcement learning theory. Let θ Q is the Q-value network hyperparameter, then

[0123] Q π (s t ,a t )≈Q π (s t ,a t |θ Q ) (12)

[0124] The Actor-Critic framework can effectively improve the interaction efficiency and sample utilization rate between the agent and the external environment, and is one of the current mainstream deep reinforcement learning frameworks. The Deep Deterministic Policy Gradient (DDPG) algorithm is one of the commonly used Actor-Critic algorithms. The training of the anti-tank missile model of the present invention adopts this algorithm, where the network Q value is expressed as

[0125] y t =r t +γQ(s t+1 ,π(s t+1 |φ′ π )|θ′ Q ) (13)

[0126] where θ′ Q and φ′ π are the hyperparameters of the target Q-value network and V-value network.

[0127] The policy network parameter update method in the DDPG algorithm is

[0128]

[0129] Where N is the sample size.

[0130] The action network parameter update method in the DDPG algorithm is

[0131]

[0132] The target network is updated using the following soft update method:

[0133]

[0134] Where τ is the soft update factor.

[0135] In this embodiment, the state space and action space of the agent training are established. In order to carry out the attack task on a fixed target point, the flight state and relative state information of the anti-tank missile are selected as the agent observation variable s. t Since the anti-tank missile under study is of STT configuration, the angle of attack and sideslip angle are selected as the aircraft action space, as follows

[0136]

[0137] Taking into account the differences in the characteristics of the aircraft model and the order of magnitude of the variables, the state and action actually observed by the environment are

[0138]

[0139] Among them, R,q ε ,q β is the missile-target relative distance, the achieved inclination angle and the line of sight deviation angle, and max_α, max_β are the maximum flight attack angle and sideslip angle of the STT missile. Through the above transformation, the order of magnitude of the system state variables observed by the environment is between [-100, 100], and the order of magnitude of the action variables is between [-1, 1].

[0140] In this embodiment, a reward function and a neural network for striking a fixed target point are established. In order to guide the anti-tank missile to reach the target point smoothly while satisfying certain process constraints and control variable constraints, the reward function consists of the following three parts: 1) distance reward function

[0141]

[0142] Among them, R(t-1) and R(t) are the relative distances between the projectile and the target at time t-1 and t respectively. is the velocity vector magnitude w 1 is the reward weight coefficient, the default value is 1; 2) Heading angle reward function

[0143]

[0144] Where: and are respectively the ballistic deviation angle at time t and the expected ballistic inclination angle at time t, and the expression is

[0145]

[0146] where x t arg et ,y t arg et ,z t arg et is the position coordinate of the target point in the ground coordinate system; 3) Process constraint reward function

[0147]

[0148] Where: t max , R max is the maximum training step length and the maximum allowed miss amount. The training end reward function is triggered when the missile successfully intercepts the target or other termination conditions (such as training step length timeout, aircraft state singularity, etc.) occur. The comprehensive reward function of the agent training process is

[0149]

[0150] The online trajectory planning method based on the DDPG framework needs to design 4 neural networks, including 2 action networks and 2 evaluation networks. The Actor action network (target action network) of the present invention adopts a 4-layer fully connected layer BP neural network structure of 6-256-256-2, and the Critic evaluation network (target evaluation network) adopts a 4-layer BP neural network structure of 8-256-256-1. The Relu function is used as the hidden layer activation function, and the network parameters adopt the Adam adaptive gradient optimization algorithm.

[0151] In this embodiment, the design of the trajectory planning model training and testing method based on deep reinforcement learning takes an anti-tank missile as the research object, with an initial mass of 60 kg, an initial flight speed of 180 m / s, a characteristic length of 2 m, a characteristic area of ​​0.015 m2, a flight angle of attack and a sideslip angle range of -20°-20°, and a maximum range of 10 km.

[0152] During the agent training phase, the initial state of the anti-tank missile is set to

[0153]

[0154] The target point position is set to

[0155] x t0 =5km,y y0 =0,z t0 =2km

[0156] In the test phase, the agent randomly initializes the target point position and the missile initial launch angle θ within the missile range. m0 ~[0,90°] and ballistic angle

[0157] Agent reward function weight coefficient w 1 =1,w 2 =0.05, the DDPG deep reinforcement learning algorithm parameters are set as follows:

[0158] Action network structure: 6-256-256-2

[0159] Evaluation network structure: 8-256-256-1

[0160] Action network learning rate: 3×10 -4

[0161] Evaluation network learning rate: 3×10 -4

[0162] Discount rate γ: 0.99

[0163] Soft update rate τ: 0.05

[0164] Sample number batch size: 4*256

[0165] Experience pool size: 1×10 7

[0166] Training rounds: 2000

[0167] The anti-tank missile trajectory planning method based on the DDPG algorithm includes two processes: trajectory planning method training and online use. The purpose of trajectory planning method training is to store historical experience data through a large number of interactions between anti-tank missiles and the external environment, and to complete offline learning and optimization of the Actor network and Critic network parameters in the DDPG framework through certain optimization rules. The specific steps are as follows:

[0168] 1) Set the physical properties and basic performance parameters of anti-tank missiles;

[0169] 2) Design the DDPG deep reinforcement learning framework and initialize the network parameters;

[0170] 3) Initialize the initial state and target point position of the anti-tank missile;

[0171] 4) Input the current state information into the DDPG algorithm to calculate the current guidance instructions;

[0172] 5) Update the aircraft status and missile-target relative information;

[0173] 6) Calculate the reward function and store the current state, action, reward value and next state in the experience pool;

[0174] 7) Randomly sample historical data from the experience pool and update DDPG network parameters through certain learning rules;

[0175] 8) Determine whether the stop condition of this round is met. If so, reinitialize the aircraft state, otherwise proceed to the next step;

[0176] 9) Determine whether the reward function converges. If so, stop training the agent. Otherwise, go to step 4).

[0177] In the online use stage of the trajectory online planning method, it is only necessary to save and extract the Actor network in the converged DDPG framework, generate guidance instructions online under the condition of inputting missile-target relative information, and realize the online trajectory planning of the anti-tank missile. The specific implementation steps are as follows:

[0178] 1) Randomly initialize the initial state of the anti-tank missile;

[0179] 2) Randomly initialize the target point position;

[0180] 3) Importing Actor network in DDPG framework;

[0181] 4) Input the current missile-target relative information into the Actor network and calculate the guidance instructions;

[0182] 5) Update the flight status of anti-tank missiles;

[0183] 6) Determine whether the off-target condition is met, if so, stop the test, otherwise proceed to step 4).

[0184] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.

[0185] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A three-dimensional trajectory planning method for anti-tank missiles based on deep reinforcement learning, characterized in that: The steps include: Establish a nonlinear dynamic model of anti-tank missile; Establish the relative kinematic model between anti-tank missile and target; Establish a deep reinforcement learning framework for anti-tank missile training; Establish state space and action space for agent training; Establish a reward function and neural network for striking fixed target points; Design a training and testing method for trajectory planning models based on deep reinforcement learning; The neural network includes 4 neural networks, 2 action networks and 2 evaluation networks. The Actor action network adopts a 4-layer fully connected layer BP neural network structure of 6-256-256-2, and the Critic evaluation network adopts a 4-layer BP neural network structure of 8-256-256-1. The Relu function is used as the hidden layer activation function, and the network parameters adopt the Adam adaptive gradient optimization algorithm. The reward function includes the following three parts: 1) distance reward function Among them, R(t-1) and R(t) are the relative distances between the projectile and the target at time t-1 and t respectively. is the velocity vector magnitude w1 is the reward weight coefficient, the default value is 1; 2) Heading angle reward function Where: and are respectively the ballistic deviation angle at time t and the expected ballistic inclination angle at time t, and the expression is where x target ,z target is the position coordinate of the target point in the ground coordinate system; 3) Process constraint reward function Where: t max , R max is the maximum training step length and the maximum allowed miss amount. The training end reward function is triggered when the missile successfully intercepts the target or other termination conditions occur. The comprehensive reward function of the agent training process is 2. The method for three-dimensional trajectory planning of an anti-tank missile based on deep reinforcement learning according to claim 1, characterized in that: The three-dimensional anti-tank missile dynamics model is as follows: Among them, V,θ, are respectively the flight speed, ballistic inclination and ballistic deviation; X, Y, Z are the drag, lift and lateral force on the aircraft, α, β, γ V are the angle of attack, sideslip angle and tilt angle; x, y, z are the positions of the aircraft in three directions in the ground coordinate system; m is the mass of the missile.

3. The method for three-dimensional trajectory planning of an anti-tank missile based on deep reinforcement learning according to claim 2, characterized in that: The process of establishing the relative kinematic model between the anti-tank missile and the target is as follows: In the line of sight coordinate system, R is defined as the relative distance between the target and the missile, q ε is the line of sight inclination, q β is the line of sight deflection angle, the relative distance between the missile and the target and its derivative are represented by R and Indicates that there is where x r = x t - x m , y r = y t - y m , z r = z t - z m ; Sight angle q ε ,q β and line of sight angular velocity The expression is Among them, q ε ,q β are the line of sight angles in pitch and yaw directions, respectively.

4. The method for three-dimensional trajectory planning of an anti-tank missile based on deep reinforcement learning according to claim 3 is characterized in that: In the deep reinforcement learning framework for anti-tank missile training, the Q-value function Q π (s t ,a t ) is used to indicate that in state s t Next, take action a t Expected value of revenue that can be achieved Q π (s t ,a t )=E π [R t |s t ,a t ](11) Use a deep neural network to approximate the Q value in equation (11), assuming θ Q is the Q-value network hyperparameter, Then there is Q π (s t ,a t )≈Q π (s t ,a t |θ Q )(12) The anti-tank missile model is trained using a deep deterministic policy gradient algorithm, where the network Q value is expressed as y t =r t +γQ(s t+1 ,π(s t+1 |f π ′)|θ′ Q (13) where θ Q ′Q and φ π ′ is the hyperparameter of the target Q value network and V value network. The policy network parameter update method in the DDPG algorithm is Where N is the sample size; The action network parameter update method in the DDPG algorithm is The target network is updated using the following soft update method: Where τ is the soft update factor.

5. The method for three-dimensional trajectory planning of an anti-tank missile based on deep reinforcement learning according to claim 4, characterized in that: For an anti-tank missile based on the STT configuration, the angle of attack and sideslip angle are selected as the aircraft action space, as follows Taking into account the differences in the characteristics of the aircraft model and the order of magnitude of the variables, the state and action actually observed by the environment are Among them, R,q ε ,q β are the missile-target relative distance, line of sight inclination and line of sight deviation, max_α, max_β are the maximum flight angle of attack and sideslip angle of the STT missile. Through the above transformation, the order of magnitude of the system state variables observed by the environment is between [-100, 100], and the order of magnitude of the action variables is between [-1, 1].

6. The method for three-dimensional trajectory planning of an anti-tank missile based on deep reinforcement learning according to claim 1, characterized in that: The steps of the trajectory planning model training method based on deep reinforcement learning are as follows: 1) Set the physical properties and basic performance parameters of anti-tank missiles; 2) Design the DDPG deep reinforcement learning framework and initialize the network parameters; 3) Initialize the initial state and target point position of the anti-tank missile; 4) Input the current state information into the DDPG algorithm to calculate the current guidance instructions; 5) Update the aircraft status and missile-target relative information; 6) Calculate the reward function and store the current state, action, reward value and next state in the experience pool; 7) Randomly sample historical data from the experience pool and update DDPG network parameters through certain learning rules; 8) Determine whether the stop condition of this round is met. If so, reinitialize the aircraft state, otherwise proceed to the next step; 9) Determine whether the reward function converges. If so, stop training the agent. Otherwise, go to step 4).

7. The method for three-dimensional trajectory planning of an anti-tank missile based on deep reinforcement learning according to claim 6, characterized in that: The steps of the trajectory planning model testing method based on deep reinforcement learning are as follows: 1) Randomly initialize the initial state of the anti-tank missile; 2) Randomly initialize the target point position; 3) Importing Actor network in DDPG framework; 4) Input the current missile-target relative information into the Actor network and calculate the guidance instructions; 5) Update the flight status of anti-tank missiles; 6) Determine whether the off-target condition is met, if so, stop the test, otherwise proceed to step 4).

Citation Information

Patent Citations

  • Multi-aircraft flight path planning method based on deep Q learning algorithm

    CN110928329A