An aircraft no-fly zone online avoidance system based on deep reinforcement learning
Patent Information
- Application Number
- CN202310346486.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-04-03
AI Technical Summary
[0006]有鉴于上述背景技术,本发明提供了一种基于深度强化学习的飞行器禁飞区在线规避系统;针对当前轨迹规划与制导方法实时性弱,智能自主性较差,无法规避实时探测到的禁飞区域的问题,首先采用基于改进A-star算法的航路点决策技术,决策出能够引导飞行器进行轨迹规划的飞行器航路点;其次采用基于深度强化学习的轨迹优化方法,生成飞行器实际参考轨迹以及倾侧角制导指令
[0030](1)相较于现有的路径规划方法,所述的飞行器航路点决策模块的算法实时性强。本发明中的航路点决策技术采用自适应搜索步长,相较于一般的固定步长搜索方法,本发明所采用的技术能够大大缩短路径搜索和航路点决策时间。
Smart Images

Figure CN116414149B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aircraft trajectory planning and guidance technology, and in particular to an online no-fly zone avoidance system for aircraft based on deep reinforcement learning. Background Technology
[0002] The aircraft trajectory planning and guidance system (TRASS) is the decision-making layer within the integrated aircraft navigation, guidance, and control system. Based on the aircraft's flight mission, it is responsible for generating reference flight trajectories and reference guidance commands, and issuing flight status control commands to the aircraft control system, thus holding an extremely important position. As the battlefield combat environment for aircraft becomes increasingly three-dimensional and complex, effectively avoiding no-fly zones such as detection and interception zones along the flight path has become a key issue in the field of aircraft trajectory planning and guidance. Traditional offline trajectory planning and online tracking guidance methods can effectively avoid static, known no-fly zones; however, with the increasing flexibility of enemy detection and interception defense systems, many real-time no-fly zones have emerged in the battlefield environment. Since aircraft cannot obtain this no-fly information before executing a flight mission, but can only obtain this information during flight through detection systems and ground command systems, traditional trajectory planning methods cannot achieve avoidance of these no-fly zones. Therefore, a real-time, highly flexible online trajectory planning system for aircraft is needed.
[0003] Intelligent graph search algorithms are a type of intelligent path planning algorithm that constructs a graph data structure using directed graphs formed by connecting nodes. In graph search algorithms, nodes represent locations in a map model, and edges represent the cost of location transitions. The basic unit of a graph search algorithm is a triple (child node, parent node, node cost). Intelligent graph search algorithms employ a priority-based search approach to plan the feasible path with the minimum cost. By introducing the concept of node cost during the search process, intelligent graph search algorithms significantly improve search efficiency, thus greatly enhancing real-time performance while ensuring the acquisition of the optimal path.
[0004] Deep reinforcement learning is an artificial intelligence technique belonging to machine learning. Deep reinforcement learning first requires constructing a Markov decision model based on the specific task. Its basic components include a state space, action space, agent, and reward function model. The reward function model guides the agent to output the desired action, making it crucial for designing the Markov decision model. After determining the Markov decision model, the reinforcement learning algorithm optimizes the agent's parameters based on extensive offline simulation results, thereby obtaining the optimal agent. The advantages of deep reinforcement learning technology lie in its ability to determine the complex mapping relationship between states and actions through extensive offline simulations. Furthermore, because deep reinforcement learning uses neural networks to output the agent's policy, its algorithm has strong real-time performance, which is beneficial for real-time online planning of aircraft.
[0005] Therefore, how to provide an online no-fly zone avoidance system for aircraft based on deep reinforcement learning has become a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] In view of the above background technology, the present invention provides an online no-fly zone avoidance system for aircraft based on deep reinforcement learning. Addressing the problems of weak real-time performance, poor intelligent autonomy, and inability to avoid real-time detected no-fly zones in current trajectory planning and guidance methods, the present invention first employs a waypoint decision-making technology based on an improved A-star algorithm to determine aircraft waypoints that can guide the aircraft in trajectory planning; secondly, it employs a trajectory optimization method based on deep reinforcement learning to generate the aircraft's actual reference trajectory and roll angle guidance commands.
[0007] The present invention solves the technical problem by adopting the following technical solution:
[0008] An online no-fly zone avoidance system for aircraft based on deep reinforcement learning includes an online waypoint decision-making module, an intelligent roll angle decision-making environment module, and an intelligent trajectory generation module; wherein:
[0009] The aircraft waypoint online decision-making module quickly generates a two-dimensional incremental real-time map model based on the aircraft's real-time position information, real-time situation information, and no-fly zone situation information; and through an intelligent search algorithm, it quickly and in real-time generates waypoints to guide the aircraft to avoid no-fly zones.
[0010] The intelligent decision-making environment module for tilt angle optimizes the aircraft tilt angle decision-making model through offline simulation.
[0011] The intelligent trajectory generation module for the aircraft uses the aircraft tilt angle decision model to output the aircraft's real-time optimal tilt angle based on the aircraft's real-time position information, real-time situation information, and no-fly zone situation information, thereby generating the aircraft's optimal avoidance trajectory.
[0012] Furthermore, it also includes an information acquisition and environmental map model conversion module; used to acquire real-time aircraft location information, real-time situation information, and no-fly zone situation information; and to convert the above information into environmental map information that can be processed by the path planning algorithm and store it in the airborne database.
[0013] Furthermore, the tilt angle intelligent decision-making environment module includes a training environment construction module, an agent network construction module, a reward model construction module, and an agent parameter optimization module; wherein:
[0014] The training environment construction module constructs environmental state, environmental actions, and environmental dynamics models based on the aircraft's flight status and no-fly zone avoidance mission.
[0015] The intelligent agent network construction module constructs a complex mapping between environmental states and environmental actions;
[0016] The reward model construction module generates aircraft decision feedback based on the interaction results between environmental actions and the training environment.
[0017] The intelligent agent parameter optimization module optimizes the model parameters of the aircraft tilt angle decision model based on the aircraft decision feedback.
[0018] Furthermore, the two-dimensional incremental real-time map model mainly includes the real-time position information of the aircraft, the real-time situation information, and the situation information of the no-fly zone. By processing the above information in real time and discretizing it, the two-dimensional incremental real-time map model is obtained, and the model is updated and maintained in real time.
[0019] Furthermore, based on the two-dimensional incremental real-time map model, and using the improved A-star intelligent search algorithm, a two-dimensional flight trajectory that can avoid no-fly zones is generated in real time. The trajectory is then filtered, and through a discretization method, waypoints that guide the aircraft to avoid no-fly zones are determined online.
[0020] Furthermore, the algorithm flow for improving the A-star intelligent search algorithm is as follows:
[0021] W1 initializes the graph model G, starting point S, target point T, untraversed node list Open, traversed node list Close, and stores the starting point S in Open;
[0022] W2 checks if there are any nodes in Open; if no nodes are found, the search ends.
[0023] W3, calculate the node cost in the Open list and sort all costs; select the node K with the lowest cost in the Open list and move K to the Close list;
[0024] W4 determines whether K is the target point. If so, it outputs the optimal path from the starting point to K.
[0025] W5. If K is not the target point, generate a child node of K. If the child node is in the Close list, delete the child node.
[0026] W6. If the child node is in the Open table, then determine the relationship between the cost from the parent node to the child node and the cost from K to the child node. If the cost from K to the child node is smaller, then update the parent node of the child node to K.
[0027] W7. If the child node is not in the Close or Open table, add it to the Open table and proceed to step W2.
[0028] Furthermore, the intelligent trajectory generation module for aircraft is divided into two parts: offline interactive training and online output. The offline interactive training part mainly uses a deep deterministic policy gradient algorithm to construct a reinforcement learning agent model training process and optimizes the agent model parameters through extensive offline simulation interaction. The online output part mainly uses the real-time position information, real-time situation information, and no-fly zone situation information of the aircraft, as well as the environmental map model conversion module and intelligent decision-making environment module, to obtain the agent state. Using the offline optimized agent model, the agent action signal is output and then converted into the aircraft roll angle guidance quantity. Through the above online output process, the optimal avoidance trajectory of the aircraft is obtained.
[0029] Beneficial effects:
[0030] (1) Compared with existing path planning methods, the algorithm of the aircraft waypoint decision module has strong real-time performance. The waypoint decision technology in this invention adopts an adaptive search step size. Compared with the general fixed step size search method, the technology adopted in this invention can greatly shorten the path search and waypoint decision time.
[0031] (2) Compared to existing path planning methods, the flight waypoint decision module described above offers stronger security in its decision-making results. Since general path planning algorithms only consider trajectory length as a cost, without taking into account the safe distance between the trajectory and the no-fly zone, this can lead to the actual trajectory of the aircraft falling within the no-fly zone due to factors such as accuracy errors and aircraft dynamics. The waypoint decision technology employed in this invention considers both trajectory length and the safe distance of the trajectory as costs, thus ensuring that the waypoint planning results maintain a certain safe distance from the no-fly zone, thereby guaranteeing the security of the trajectory planning results.
[0032] (3) Compared with existing trajectory optimization and guidance technologies, the intelligent trajectory generation module for aircraft has stronger real-time performance. Since general trajectory optimization techniques use linear recursion and full-trajectory integration to derive the aircraft reference trajectory, the trajectory calculation time is long and the real-time performance is poor. The intelligent trajectory generation technology proposed in this invention uses a deep neural network to output the aircraft's tilt angle guidance law and reference trajectory, thus the algorithm has a faster calculation speed, higher operating efficiency, and stronger real-time performance.
[0033] (4) Compared to general no-fly zone avoidance algorithms, the aircraft trajectory intelligent generation module described above exhibits stronger intelligent autonomy. Traditional no-fly zone avoidance algorithms can only plan trajectories for offline-known no-fly zones. However, for no-fly zones detected online in real time, this method cannot avoid them due to its weak intelligent autonomy. The trajectory generation technology based on deep reinforcement learning described in this invention, by using a large number of offline task scenarios to interactively train the agent, can adapt to multi-task scenarios and has strong intelligent autonomy.
[0034] (5) Compared with general trajectory optimization techniques, the modules of the aircraft no-fly zone online avoidance technology are lightweight. Since the present invention can parameterize each module and save it to the onboard computer, the computational resource consumption of the model is low and the model is highly lightweight.
[0035] As can be seen from the above technical solution, compared with the prior art, the present invention discloses an online no-fly zone avoidance method for aircraft based on deep reinforcement learning. Addressing the problems of low real-time performance and intelligent autonomy in existing trajectory optimization methods, and their inability to avoid real-time no-fly zones, this invention generates waypoint information through intelligent waypoint decision-making technology, thereby reducing the difficulty of trajectory optimization and improving the time efficiency of the algorithm. Furthermore, through intelligent aircraft trajectory generation technology based on deep reinforcement learning, the network parameters of the agent are trained using offline simulation, and the guidance law and reference trajectory of the aircraft are rapidly generated online through real-time interaction between the agent and the environment. Attached Figure Description
[0036] Figure 1 A schematic diagram of the structure of the online no-fly zone avoidance system for aircraft based on deep reinforcement learning provided by the present invention.
[0037] Figure 2 This is a schematic diagram of the information acquisition and environmental map model conversion module provided by the present invention.
[0038] Figure 3 The flowchart of the improved A-star algorithm provided by this invention.
[0039] Figure 4 A schematic diagram of the intelligent waypoint decision-making technology for aircraft based on the improved A-star algorithm provided by this invention.
[0040] Figure 5 The present invention provides a flow chart for an aircraft trajectory optimization algorithm based on a deep deterministic policy gradient algorithm.
[0041] Figure 6 A schematic diagram of the intelligent aircraft trajectory generation technology framework based on deep reinforcement learning provided by this invention. Detailed Implementation
[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0043] refer to Figure 1-6 This invention discloses an online no-fly zone avoidance system for aircraft based on deep reinforcement learning, comprising an online waypoint decision module, an intelligent decision environment module for roll angle, and an intelligent trajectory generation module for aircraft; wherein:
[0044] The aircraft waypoint online decision-making module quickly generates a two-dimensional incremental real-time map model based on the aircraft's real-time position information, real-time situation information, and no-fly zone situation information; and through an intelligent search algorithm, it quickly and in real-time generates waypoints to guide the aircraft to avoid no-fly zones.
[0045] The intelligent decision-making environment module for tilt angle optimizes the aircraft tilt angle decision-making model through offline simulation.
[0046] The intelligent trajectory generation module for the aircraft uses the aircraft tilt angle decision model to output the aircraft's real-time optimal tilt angle based on the aircraft's real-time position information, real-time situation information, and no-fly zone situation information, thereby generating the aircraft's optimal avoidance trajectory.
[0047] The present invention also includes an information acquisition and environmental map model conversion module; used to acquire real-time location information, real-time situation information and no-fly zone situation information of the aircraft; and to convert the above information into environmental map information that can be processed by the path planning algorithm and store it in the airborne database.
[0048] The intelligent decision-making environment module for the tilt angle includes a training environment construction module, an agent network construction module, a reward model construction module, and an agent parameter optimization module. Specifically: the training environment construction module constructs environmental state, environmental actions, and environmental dynamics models based on the aircraft's flight state and no-fly zone avoidance mission; the agent network construction module constructs a complex mapping between environmental states and environmental actions; the agent network is a multi-hidden-layer neural network; the reward model construction module generates aircraft decision feedback through the interaction results between environmental actions and the training environment; and the agent parameter optimization module optimizes the model parameters of the aircraft tilt angle decision-making model based on the aircraft decision feedback.
[0049] The intelligent decision-making environment module for tilt angle combines the aforementioned aircraft flight status information, real-time situational information, and waypoint decision information to construct a deep reinforcement learning training environment. Simultaneously, leveraging the strong fitting ability of deep neural networks, it constructs a complex mapping relationship between the environmental state and the agent's actions, i.e., using the neural network to determine the aircraft's tilt angle control quantity. The agent parameter optimization module performs flight simulation deduction based on the tilt angle decision results using digital simulation, and optimizes the network parameters of the agent's neural network based on the simulation results.
[0050] In this embodiment, the two-dimensional incremental real-time map model mainly includes the aircraft's real-time position information, real-time situational information, and no-fly zone situational information. Through real-time processing of the above information and discretization, a two-dimensional incremental real-time map model is obtained, and this model is updated and maintained in real-time. Based on the two-dimensional incremental real-time map model, an improved A-star intelligent search algorithm is used to generate a two-dimensional flight trajectory that can avoid the no-fly zone in real time. This trajectory is then filtered, and through discretization, waypoints guiding the aircraft to avoid the no-fly zone are determined online.
[0051] In this embodiment, the algorithm flow of the improved A-star intelligent search algorithm is as follows:
[0052] W1 initializes the graph model G, starting point S, target point T, untraversed node list Open, traversed node list Close, and stores the starting point S in Open;
[0053] W2 checks if there are any nodes in Open; if no nodes are found, the search ends.
[0054] W3, calculate the node cost in the Open list and sort all costs; select the node K with the lowest cost in the Open list and move K to the Close list;
[0055] W4 determines whether K is the target point. If so, it outputs the optimal path from the starting point to K.
[0056] W5. If K is not the target point, generate a child node of K. If the child node is in the Close list, delete the child node.
[0057] W6. If the child node is in the Open table, then determine the relationship between the cost from the parent node to the child node and the cost from K to the child node. If the cost from K to the child node is smaller, then update the parent node of the child node to K.
[0058] W7. If the child node is not in the Close or Open table, add it to the Open table and proceed to step W2.
[0059] The intelligent trajectory generation module for aircraft is divided into two parts: offline interactive training and online output. The offline interactive training part mainly uses a deep deterministic policy gradient algorithm to construct a reinforcement learning agent model training process and optimizes the agent model parameters through extensive offline simulation interaction. The online output part mainly uses the real-time position information, real-time situation information, and no-fly zone situation information of the aircraft, as well as the environmental map model conversion module and intelligent decision-making environment module to obtain the agent state. Using the offline optimized agent model, the agent action signal is output and then converted into the aircraft roll angle guidance quantity. Through the above online output process, the optimal avoidance trajectory of the aircraft is obtained.
[0060] The intelligent aircraft trajectory generation module can generate no-fly zone avoidance trajectories and provide battlefield command decision-making support for ground commanders. This module employs the aforementioned offline-optimized aircraft roll angle decision model. Based on the real-time flight status of the aircraft and the situational information of the no-fly zone, it outputs the optimal roll angle of the aircraft in real time. Through digital simulation performed by an onboard computer, the predicted flight trajectory of the aircraft is obtained. Ground command determines whether the aircraft can avoid the no-fly zone based on the predicted flight trajectory, thus deciding whether the aircraft should track the predicted trajectory and sending early warning information to the aircraft.
[0061] Example
[0062] This invention first uses an information acquisition and environmental map model conversion module to collect the current aircraft position information and detected no-fly zone information, mainly including the aircraft's latitude and longitude coordinates, flight speed, flight altitude, aircraft attitude angles, no-fly zone center latitude and longitude coordinates, and no-fly zone radius. This information is then integrated with a map environment model to generate a two-dimensional mesh model.
[0063] Secondly, an improved A-star algorithm is constructed and used to search for the optimal path on the aforementioned two-dimensional grid model. Compared to the traditional A-star algorithm, the improved A-star algorithm has the following advantages: First, it sets a safe distance around the no-fly zone, thus avoiding trajectory planning results that are too close to the no-fly zone, which could reduce aircraft safety. Second, the improved A-star algorithm uses an adaptive search step size; that is, when the current point is far from the target point, the algorithm uses a larger search step size, and when the current point is near the target point, the algorithm uses a smaller step size for refined searching. After finding feasible paths using the improved A-star algorithm, the search results are discretized to obtain the aircraft waypoint decision results. After obtaining the waypoint decision results, the trajectory planning task is decomposed into multiple sub-tasks, with the starting point and target point of each sub-task being two adjacent waypoints.
[0064] Based on the trajectory planning sub-task described above, a tilt angle intelligent decision-making environment module is constructed to facilitate the offline training process of the agent model. The tilt angle intelligent decision-making environment module mainly includes a training environment construction module, an agent network construction module, a reward model construction module, and an agent parameter optimization module. The agent state includes information such as aircraft speed, flight altitude, relative motion between the aircraft and the no-fly zone, and relative motion between the aircraft and the target point. The agent state model extracts features from the above information and transforms these features into agent states using a normalization method. Both the agent action neural network and the evaluation neural network are multi-hidden-layer deep neural networks, and their input is the agent state. The difference is that the action neural network outputs the agent's action selection strategy, while the evaluation neural network outputs an evaluation of the strategy. The agent reward function model is used to evaluate the agent's "performance"; the greater the total reward obtained by the agent, the more its expected strategy matches the expectations of the decision-making task.
[0065] Finally, based on the aforementioned intelligent decision-making environment module for tilt angle, an intelligent trajectory generation module for the aircraft is constructed. This module consists of two parts: offline interactive training and online output. The offline interactive training part primarily uses the Deep Deterministic Policy Gradient (DDPG) algorithm to construct a reinforcement learning agent model training process and optimizes the agent model parameters through extensive offline simulation interactions. The online output part mainly uses the aforementioned information acquisition and environmental map model conversion module and intelligent decision-making environment module to obtain the agent's state; using the offline-optimized agent model, it outputs the agent's action signals, which are then converted into aircraft tilt angle guidance quantities. Through this online output process, the aircraft trajectory is obtained.
[0066] To further optimize the above technical solution, the information acquisition and environmental map model conversion module includes online acquisition of aircraft and no-fly zone information and a raster-based map environment modeling part.
[0067] like Figure 2 As shown, the information acquisition section is mainly based on the aircraft's online information detection system, integrating information such as the aircraft's latitude and longitude coordinates, flight speed, flight altitude, aircraft attitude angles, latitude and longitude coordinates of the no-fly zone center, and the radius of the no-fly zone. The environmental map model conversion section mainly adopts the raster method, combining the above information to establish an incremental map grid model.
[0068] Grid-based modeling involves representing the missile's mission environment using a large number of grids. These grids are divided into obstacle grids and feasible free space grids. Obstacles can be further subdivided into terrain obstacles, enemy interception and detection threats, etc., all areas the missile's path is not permitted to traverse. By representing obstacles and free space in the environment using grids, a mathematical model is established, laying the foundation for subsequent path planning. The grid points are first selected to an appropriate size, and then the obstacle areas and feasible free space are processed separately according to the grid size to obtain the offline environment model. The detailed process is as follows:
[0069] (1) Determine the location of the missile launch point and multiple targets, as well as the grid size;
[0070] (2) Based on the known topographic map;
[0071] (3) Process the image data on the map according to the grid size.
[0072] To further optimize the above technical solution, the aircraft waypoint decision module mainly adopts the improved A-star algorithm for path search and discretizes the search results into aircraft waypoints.
[0073] The improved A-star algorithm structure consists of four parts: state space, action space, heuristic function, and algorithm search process.
[0074] This invention considers adding a suitable safety distance to the traditional state space to expand infeasible areas, ensuring that the planned waypoints have a certain safety margin, thereby facilitating the generation of subsequent aircraft avoidance trajectories. The graph search state space is set as follows:
[0075]
[0076] In the above formula, S represents the state space of the graph search algorithm; Δx and Δy represent the relative distances between the nodes and the no-fly zone boundary; r represents the radius of the no-fly zone; and dis_safe represents the safe distance.
[0077] This invention employs an adaptive search step size. A larger step size is used in the initial stage of the path search to reduce the search time; while the step size is reduced in the final stage, allowing the algorithm to perform a fine-grained search and thus ensuring the optimality of the path planning result. The search step size formula for the A-star algorithm is as follows:
[0078]
[0079] In the above formula, s k The search step size of the search algorithm is represented by s1 and s2, which represent the long search step size and the short search step size, respectively; (x k ,y k),(x T ,y T ) represent the grid coordinates of the current point and the target point, respectively; r_dis represents the separation distance between the initial search segment and the final search segment.
[0080] Based on the search step size mentioned above, the action space of the A-star algorithm is further constructed as follows:
[0081] A k+1 ={(x k+1 ,y k+1 )∈S|x k+1 =x k ±s k ,y k+1 =y k ±s k}
[0082] A k+1 Represents the action space of the search algorithm; (x k+1 ,y k+1 ) represents the grid coordinates of the expanded node at the next time step.
[0083] The heuristic function of the A-star algorithm can be expressed as:
[0084] f(n) = g(n) + h(n)
[0085] g(n) represents the movement cost function of the current search path. g(n) measures the cost incurred by the algorithm to move from the starting node to the current node; therefore, the value of g(n) is known for each node on the path. h(n) represents the heuristic function of the current node, which measures the estimated cost from the current node to the target node, i.e., the estimated distance traveled from the current node to the target node.
[0086] The algorithm flow of the improved A-star algorithm is as follows: Figure 3 As shown:
[0087] (1) Initialize the graph model G, starting point S, target point T, untraversed node list Open, traversed node list Close, and store the starting point S in Open;
[0088] (2) Determine if there are any nodes in Open. If there are no nodes, the search ends.
[0089] (3) Calculate the node cost in the Open table and sort all the costs; select the node K with the lowest cost in the Open table and move K to the Close table;
[0090] (4) Determine whether K is the target point. If so, output the optimal path from the starting point to K.
[0091] (5) If K is not the target point, generate a child node of K. If the child node is in the Close list, delete the child node.
[0092] (6) If the child node is in the Open table, then determine the relationship between the cost from the parent node to the child node and the cost from K to the child node. If the cost from K to the child node is smaller, then update the parent node of the child node to K.
[0093] (7) If the child node is not in the Close or Open table, add it to the Open table and go to step (2).
[0094] To further optimize the above technical solution, the aircraft waypoint decision-making steps are as follows: Figure 4 As shown, the specific process is as follows:
[0095] (1) During the flight mission, the aircraft will obtain the location information of the enemy's no-fly zone threat through various detection information sources. First, it is determined whether these no-fly zones will pose a threat to the aircraft's re-entry process; if the enemy's no-fly zone will pose a threat to the aircraft's flight process (such as the aircraft's pre-flying route passing through the no-fly zone), the aircraft waypoint decision module will first establish a two-dimensional grid map model using the grid method based on the location information of the no-fly zone and the estimated radius of the no-fly zone, combined with the aircraft's current position information.
[0096] (2) In order to improve the real-time performance of waypoint decision-making, the algorithm prioritizes a larger simulation search step size and uses an improved A* algorithm to perform path search on a two-dimensional grid map.
[0097] (3) If the algorithm cannot obtain a feasible path, or the latitude and longitude coordinates of the path nodes obtained after coordinate transformation are not within the feasible range, then reduce the simulation search step size and search the path again.
[0098] (4) By sampling the above path nodes, the waypoint decision results are finally obtained, and the waypoint location information is provided to the subsequent aircraft avoidance trajectory generation module.
[0099] This invention constructs an intelligent trajectory optimization algorithm for aircraft based on a deep deterministic policy gradient algorithm. The algorithm includes an intelligent decision-making environment model for the tilt angle and an offline optimization algorithm for the intelligent agent. For example... Figure 5 As shown, the tilt angle intelligent decision-making environment model mainly includes an agent state model, an agent action neural network, an agent evaluation neural network, and an agent reward function model, as detailed below:
[0100] (1) For the task of generating trajectories to avoid no-fly zones, the selection of state variables needs to consider three main factors simultaneously: the flight state of the aircraft itself, the relative situation between the aircraft and the no-fly zone, and the relative situation between the aircraft and the target point. This invention selects ten variables as the state variables of the trajectory generation agent: the line-of-sight angle between the aircraft and the center of the no-fly zone, the line-of-sight angle between the aircraft and the target point, the relative distance between the aircraft and the center of the no-fly zone, the relative distance between the aircraft and the target point, the aircraft's flight speed, flight altitude, longitude, latitude, ballistic inclination angle, and ballistic deflection angle.
[0101] (2) In this invention, both the agent action network and the agent evaluation network are deep neural networks with two hidden layers. The network inputs of both the agent action network and the evaluation network are the aforementioned agent state variables; the network output of the agent action network is the aircraft roll angle guidance quantity, and the network output of the agent evaluation network is the policy evaluation of the action network, i.e., the action value function. The training objective of the action network is to maximize the action evaluation of the evaluation network; the training objective of the evaluation network is to minimize the error between the policy evaluation output by the network and the true action value function.
[0102] (3) In the training process of reinforcement learning agents, the main role of the reward function is to guide the agent to complete the training task, to mathematically model the reinforcement learning task, and to transform it into an optimal planning problem that maximizes the expected total reward. For the task of generating no-fly zone avoidance trajectories, the reward function model in this invention mainly considers five aspects: reentry flight constraints, aircraft terminal error, aircraft position guidance, aircraft no-fly zone constraint violation, and aircraft energy loss.
[0103] The intelligent trajectory generation module of this invention, such as Figure 6 As shown above, the reinforcement learning agent action network outputs agent actions based on various normalized environmental state variables. This invention further constructs a tilt angle conversion module to convert the agent network output into an aircraft tilt angle guidance quantity; through digital simulation, the aircraft motion state for the next guidance cycle is obtained. The training environment outputs the agent state and action evaluation reward for the next moment based on the aircraft motion state, and stores the aforementioned state transition data in the experience pool R.
[0104] When the amount of data in the experience pool meets a certain value, the algorithm randomly extracts experience data from the experience pool and performs parameter optimization training on the action network and evaluation network.
[0105] The training objective of the action network is to improve the evaluation network's evaluation of the action network's actions. The training process for the action network is as follows: The evaluation network outputs a policy evaluation Q for the action network's policy based on the state information and the action information output by the action network. π(s,a); Based on this evaluation, the action network generates the total network loss function, and finally generates the gradients of all network parameters through backpropagation. The parameter gradients of the action network are expressed as:
[0106]
[0107] In the above formula, ▽ θ J(μ θ ) represents the gradient of the network parameters of the action network; ▽ θ μ θ (s) represents the derivative of the action network output with respect to the network parameters; ▽ a Q μ (s,μ θ (s) represents the derivative of the action value function with respect to the action; batch represents a set of state-action pairs (s, a) data sampled from the experience pool; n represents the number of data in the batch.
[0108] The training objective of the evaluation network is to ensure that the policy evaluation output by the network reflects the true action value function Q. π * The training process for the evaluation network (s, a) is as follows: Based on the reward information and action value information in the data, the training target value y of the evaluation network is generated, and the mean squared error function is used to construct the target function Loss of the evaluation network. MSE (Q π -y), and finally, through the backpropagation algorithm, the gradients of all network parameters are obtained. The loss function for evaluating the network can be expressed as:
[0109] y t =r(s t ,a t )+γQ(s t+1 ,μ(s t+1 )|θ Q )
[0110]
[0111] In the above formula, y t The truth value of an action; r(s) t ,a t Q(s) represents the reward obtained from the current action; t+1 ,μ(s t+1 )|θ Q L(θ) represents the estimated action value function at the next moment; γ represents the discount factor; L(θ) represents the value of the action at the next moment. Q ) represents the loss function for evaluating the network; Q(s) t ,a t |θ Q ) represents the evaluation network output at the current moment.
[0112] This allows us to obtain the parameter gradients of the evaluation network:
[0113]
[0114] In the above formula, ▽ θ L(θ Q ) represents the gradient of the network parameters of the action network; ▽ θ Q(s t ,a t |θ Q ) represents the derivative of the action network output with respect to the network parameters.
[0115] Based on the aforementioned network parameter gradients, the network parameters of the action network and evaluation network are optimized and updated using a step-size update method:
[0116]
[0117]
[0118] In the above formula, θ actor ,θ′ actor θ represents the network parameters of the action network before and after the update, respectively. critic ,θ′ critic These represent the network parameters of the evaluation network before and after the update, respectively; LR actor ,LR critic represents the learning rates of the action network and the evaluation network, respectively; These represent the gradients of the network parameters for the action network and the evaluation network, respectively.
[0119] The above process is repeated until the total reward of the agent converges, thus obtaining the aircraft tilt angle decision agent. Finally, the above agent is used to construct the aircraft trajectory online output module: the online output part is mainly based on the above information acquisition and environmental map model conversion module and intelligent decision environment module to obtain the agent state; using the offline optimized agent model, the agent action signal is output, which is then converted into the aircraft tilt angle guidance quantity and the aircraft trajectory is obtained.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An online no-fly zone avoidance system for aircraft based on deep reinforcement learning, characterized in that, This includes an online waypoint decision-making module, an intelligent decision-making environment module for roll angles, and an intelligent trajectory generation module for aircraft; among which: The aircraft waypoint online decision-making module quickly generates a two-dimensional incremental real-time map model based on the aircraft's real-time position information, real-time situation information, and no-fly zone situation information; and through an intelligent search algorithm, it quickly and in real-time generates waypoints to guide the aircraft to avoid no-fly zones. The intelligent decision-making environment module for tilt angle optimizes the aircraft tilt angle decision-making model through offline simulation. The intelligent trajectory generation module of the aircraft uses the aircraft tilt angle decision model to output the aircraft's real-time optimal tilt angle based on the aircraft's real-time position information, real-time situation information and no-fly zone situation information, thereby generating the aircraft's optimal avoidance trajectory. The intelligent trajectory generation module for aircraft is divided into two parts: offline interactive training and online output. The offline interactive training part mainly uses a deep deterministic policy gradient algorithm to construct a reinforcement learning agent model training process and optimizes the agent model parameters through extensive offline simulation interaction. The online output part mainly uses the real-time position information, real-time situation information, and no-fly zone situation information of the aircraft, as well as the environmental map model conversion module and intelligent decision-making environment module to obtain the agent state. Using the offline optimized agent model, the agent action signal is output and then converted into the aircraft roll angle guidance quantity. Through the above online output process, the optimal avoidance trajectory of the aircraft is obtained. A tilt angle conversion module is constructed to convert the agent network output into the aircraft tilt angle guidance quantity; through digital simulation, the aircraft motion state for the next guidance cycle is obtained; the training environment outputs the agent state and action evaluation reward for the next moment based on the aircraft motion state, and stores the above state transfer data in the experience pool. middle; When the amount of data in the experience pool meets a certain value, the algorithm randomly extracts experience data from the experience pool and performs parameter optimization training on the action network and evaluation network. The training objective of an action network is to improve the evaluation network's assessment of the actions taken by the action network. The training process for the action network is as follows: The evaluation network outputs a policy evaluation of the action network's policy based on the state information and the action information output by the action network. Based on this evaluation, the action network generates the total network loss function, and finally generates the gradients of all network parameters through backpropagation. The parameter gradients of the action network are expressed as: In the above formula, This represents the gradient of the network parameters of the action network; This represents the derivative of the action network output with respect to the network parameters. This represents the derivative of the action value function with respect to the action; This represents a set of state-action pairs sampled from the experience pool. data; express The amount of data in the data; The training objective of an evaluation network is to ensure that the policy evaluation output by the network reflects the true action value function. The training process for the evaluation network is as follows: Based on the reward information and action value information in the data, the training target value for the evaluation network is generated. The objective function for evaluating the network is constructed using the mean square error function. Finally, the gradients of all network parameters are obtained through backpropagation; the loss function for evaluating the network is expressed as: In the above formula, The truth value representing the value of an action; Indicates the reward obtained from the current action; This represents the estimated action value function for the next moment; Indicates the discount factor; This represents the loss function used to evaluate the network. This represents the network output at the current moment; This leads to the parameter gradients of the evaluation network: In the above formula, This represents the gradient of the network parameters of the action network; This represents the derivative of the action network output with respect to the network parameters; Based on the aforementioned network parameter gradients, the network parameters of the action network and evaluation network are optimized and updated using a step-size update method: In the above formula, These represent the network parameters of the action network before and after the update, respectively. These represent the network parameters of the evaluation network before and after the update, respectively. represents the learning rates of the action network and the evaluation network, respectively; These represent the gradients of the network parameters for the action network and the evaluation network, respectively.
2. The online no-fly zone avoidance system for aircraft based on deep reinforcement learning according to claim 1, characterized in that, It also includes an information acquisition and environmental map model conversion module; used to acquire real-time aircraft location information, real-time situation information and no-fly zone situation information; and convert the above information into environmental map information that can be processed by path planning algorithms and store it in the airborne database.
3. The online no-fly zone avoidance system for aircraft based on deep reinforcement learning according to claim 1, characterized in that, The tilt angle intelligent decision-making environment module includes a training environment construction module, an agent network construction module, a reward model construction module, and an agent parameter optimization module; wherein: The training environment construction module constructs environmental state, environmental actions, and environmental dynamics models based on the aircraft's flight status and no-fly zone avoidance mission. The intelligent agent network construction module constructs a complex mapping between environmental states and environmental actions; The reward model construction module generates aircraft decision feedback based on the interaction results between environmental actions and the training environment. The intelligent agent parameter optimization module optimizes the model parameters of the aircraft tilt angle decision model based on the aircraft decision feedback.
4. The online no-fly zone avoidance system for aircraft based on deep reinforcement learning according to claim 1, characterized in that, The two-dimensional incremental real-time map model mainly includes the real-time location information of the aircraft, the real-time situation information, and the situation information of the no-fly zone. By processing the above information in real time and discretizing it, the two-dimensional incremental real-time map model is obtained, and the model is updated and maintained in real time.
5. The online no-fly zone avoidance system for aircraft based on deep reinforcement learning according to claim 4, characterized in that, Based on the two-dimensional incremental real-time map model, and using the improved A-star intelligent search algorithm, a two-dimensional flight trajectory that can avoid no-fly zones is generated in real time. The trajectory is then filtered, and waypoints that guide the aircraft to avoid no-fly zones are determined online through a discretization method.
6. The online no-fly zone avoidance system for aircraft based on deep reinforcement learning according to claim 5, characterized in that, The algorithm flow of the improved A-star intelligent search algorithm is as follows: W1 initializes the graph model G, starting point S, target point T, untraversed node list Open, traversed node list Close, and stores the starting point S in Open; W2 checks if there are any nodes in Open; if no nodes are found, the search ends. W3, calculate the node cost in the Open list and sort all costs; select the node K with the lowest cost in the Open list and move K to the Close list; W4 determines whether K is the target point. If so, it outputs the optimal path from the starting point to K. W5. If K is not the target point, generate a child node of K. If the child node is in the Close list, delete the child node. W6. If the child node is in the Open table, then determine the relationship between the cost from the parent node to the child node and the cost from K to the child node. If the cost from K to the child node is smaller, then update the parent node of the child node to K. W7. If the child node is not in the Close or Open table, add it to the Open table and proceed to step W2.
Citation Information
Patent Citations
Multisensor fusion-based nmanned aerial vehicle SLAM (simultaneous localization and mapping) navigation method and system
CN108827306A
Reentry vehicle trajectory planning method based on reinforcement learning
CN112947592A