A method, device, equipment and medium for determining an avoidance strategy of an unmanned aerial vehicle based on reinforcement learning
By applying reinforcement learning technology in the models of drones and other aircraft, optimizing the evasion strategy of drones, solving the problem of relying on expert knowledge and complete models in the existing technology, and achieving efficient evasion and strategy scalability of drones in complex scenarios.
Patent Information
- Application Number
- CN202510344718.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-24
AI Technical Summary
Existing methods for drones to avoid other aircraft rely on expert prior knowledge and complete mathematical models, which are difficult to effectively solve in complex nonlinear systems and are not very scalable.
Using reinforcement learning-based methods, we build models of drones and other aircraft in the engineering computing simulation platform and build reinforcement learning environment in the programming language platform, and use reinforcement learning models to optimize drone evasion strategies, get rid of the dependence on expert knowledge, and avoid modeling error interference.
It has achieved a steady increase in the avoidance success rate in complex scenarios, enhanced the scalability of the avoidance strategy, and formed a virtuous cycle optimization avoidance strategy.
Smart Images

Figure CN119861744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous control of unmanned aerial vehicles, and particularly to a method, device, equipment and medium for determining an avoidance strategy of an unmanned aerial vehicle based on reinforcement learning. Background Art
[0002] In low-altitude logistics transportation, the flight of unmanned aerial vehicles (UAVs) in scenarios with dense and intersecting flight routes is greatly threatened by other aircraft. Other aircraft may have characteristics such as high speed, strong maneuverability, and variable flight routes, which pose many obstacles to the flight of UAVs. Currently, with the transformation of logistics transportation from manned operation to unmanned and from manual operation to full autonomy, decision-making autonomy has become an inevitable trend in the development of the current low-altitude economy.
[0003] Traditional UAV obstacle avoidance methods include the expert system method, differential game method, optimal control method, etc. The expert system method relies on the prior knowledge of experts. Once the subsystems of other aircraft or airplanes change, experts need to analyze the new subsystems and give new avoidance maneuver strategies again. Both the differential game method and the optimal control method rely on clear and complete mathematical models, and it is essential to solve complex differential equations. However, the problem of an aircraft avoiding other aircraft involves multiple complex non-linear systems, and it is inevitable that there will be errors in the modeling of each subsystem, which greatly increases the difficulty of solving this problem by the above methods. In addition, most of the UAV methods for avoiding other aircraft are completed under many restrictive conditions, and can only vaguely explain which maneuvers are more advantageous under a certain type of encounter situation, and the scalability is not strong. Therefore, the above problems urgently need to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method, device, equipment and medium for determining an avoidance strategy of an unmanned aerial vehicle based on reinforcement learning, which can get rid of the dependence on the prior knowledge of experts, avoid the interference of modeling errors on avoidance decisions, prompt the reinforcement learning model to optimize the avoidance strategy, form a virtuous cycle, and stably improve the avoidance success rate, greatly enhancing the scalability of the avoidance strategy. The specific solutions are as follows:
[0005] In the first aspect, the present application discloses a method for determining an avoidance strategy of an unmanned aerial vehicle based on reinforcement learning, including:
[0006] Build a six-degree-of-freedom model of the unmanned aerial vehicle in an engineering calculation and simulation platform, and establish an action library for the unmanned aerial vehicle;
[0007] Build a three-degree-of-freedom particle model of other aircraft currently in flight in the engineering calculation and simulation platform, constrain the maneuver characteristics of other aircraft based on constraint rules, and realize the tracking of the other aircraft on the unmanned aerial vehicle according to the relative motion situation; the relative motion situation is the relative motion situation between the other aircraft and the unmanned aerial vehicle;
[0008] Build a reinforcement learning environment in the programming language platform and establish a corresponding reinforcement learning model;
[0009] Input the position and attitude information of the drone and the position and attitude information of the other aircraft into the reinforcement learning model, so as to output the flight actions in the drone action library through the reinforcement learning model;
[0010] When the drone avoids the other aircraft based on the flight actions, judge whether the drone successfully avoids based on the maneuvering characteristics of the other aircraft. If successful, give the corresponding reward value to the reinforcement learning model based on the maneuvering area situation of the other aircraft to motivate the reinforcement learning model to optimize the avoidance strategy.
[0011] Optionally, the maneuvering characteristics of the other aircraft include the flight time of the other aircraft, the safety range of the other aircraft, and the overload that the other aircraft can withstand. Then, the constraint on the maneuvering characteristics of the other aircraft based on the constraint rules includes:
[0012] Set the simulation time of the other aircraft in the engineering calculation and simulation platform, so as to constrain the flight time of the other aircraft through the simulation time of the other aircraft;
[0013] Set the size of the simulation space in the engineering calculation and simulation platform, so as to constrain the safety range of the other aircraft through the size of the simulation space;
[0014] Control the overload that the other aircraft can withstand not to exceed the preset overload threshold, so as to constrain the overload that the other aircraft can withstand; wherein, the overload that the other aircraft can withstand includes tangential overload and normal overload.
[0015] Optionally, the judgment of whether the drone successfully avoids based on the maneuvering characteristics of the other aircraft includes:
[0016] Within the flight time of the other aircraft and within the overload that the other aircraft can withstand, judge whether the distance between the other aircraft and the drone is greater than a preset value. If so, it is determined that the drone successfully avoids.
[0017] Optionally, the maneuvering area situation of the other aircraft includes the distance between the other aircraft and the drone, the relative speed between the other aircraft and the drone, the collision probability of the other aircraft, and the situation level of the area where the other aircraft is located; wherein, the situation level represents the importance degree of the area where the other aircraft is located to the drone during avoidance.
[0018] Optionally, giving corresponding reward values to the reinforcement learning model based on the maneuvering area situation of other aircraft to encourage the reinforcement learning model to optimize the avoidance strategy includes:
[0019] Giving a first reward value to the reinforcement learning model based on the distance between the other aircraft and the UAV and through a first reward function to encourage the reinforcement learning model to optimize the avoidance strategy; wherein, in the first reward function, the distance between the other aircraft and the UAV is negatively correlated with the reward value.
[0020] Optionally, giving corresponding reward values to the reinforcement learning model based on the maneuvering area situation of other aircraft to encourage the reinforcement learning model to optimize the avoidance strategy includes:
[0021] Giving a second reward value to the reinforcement learning model based on the relative speed between the other aircraft and the UAV and through a second reward function to encourage the reinforcement learning model to optimize the avoidance strategy; wherein, in the second reward function, the relative speed between the other aircraft and the UAV is negatively correlated with the reward value.
[0022] Optionally, giving corresponding reward values to the reinforcement learning model based on the maneuvering area situation of other aircraft to encourage the reinforcement learning model to optimize the avoidance strategy includes:
[0023] If the situation level of the area where the other aircraft is located is the first priority and the collision probability of the other aircraft is greater than a preset probability threshold, then giving a third reward value to the reinforcement learning model to encourage the reinforcement learning model to optimize the avoidance strategy;
[0024] If the situation level of the area where the other aircraft is located is the second priority or the collision probability of the other aircraft is not greater than the preset probability threshold, then giving a fourth reward value to the reinforcement learning model to encourage the reinforcement learning model to optimize the avoidance strategy;
[0025] Wherein, the first priority is higher than the second priority, and the third reward value is greater than the fourth reward value.
[0026] In a second aspect, the present application discloses a UAV avoidance strategy determination device based on reinforcement learning, including:
[0027] A UAV model building module, configured to build a six-degree-of-freedom model of the UAV in an engineering calculation and simulation platform and establish a UAV action library;
[0028] An other aircraft model building module, configured to build a three-degree-of-freedom particle model of other aircraft currently in a flight state in the engineering calculation and simulation platform;
[0029] Other aircraft characteristic constraint module, which is used to constrain the maneuver characteristics of other aircraft based on constraint rules and enable other aircraft to track the UAV according to the relative motion situation; the relative motion situation is the relative motion situation between the other aircraft and the UAV.
[0030] Reinforcement learning model building module, which is used to build a reinforcement learning environment in a programming language platform and establish a corresponding reinforcement learning model.
[0031] Flight action output module, which is used to input the position and attitude information of the UAV and the position and attitude information of the other aircraft into the reinforcement learning model, so as to output the flight actions in the UAV action library through the reinforcement learning model.
[0032] Avoidance strategy optimization module, which is used to judge whether the UAV successfully avoids the other aircraft based on the maneuver characteristics of the other aircraft when the UAV avoids the other aircraft based on the flight actions. If successful, a corresponding reward value is given to the reinforcement learning model based on the maneuver area situation of the other aircraft to stimulate the reinforcement learning model to optimize the avoidance strategy.
[0033] In a third aspect, the present application discloses an electronic device, including:
[0034] A memory, which is used to store a computer program;
[0035] A processor, which is used to execute the computer program to implement the foregoing disclosed method for determining a UAV avoidance strategy based on reinforcement learning.
[0036] In a fourth aspect, the present application discloses a computer-readable storage medium, which is used to store a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed method for determining a UAV avoidance strategy based on reinforcement learning is implemented.
[0037] It can be seen that the present application proposes a method for determining an avoidance strategy for an unmanned aerial vehicle (UAV) based on reinforcement learning, including: building a six-degree-of-freedom model of the UAV in an engineering calculation and simulation platform and establishing a UAV action library; building a three-degree-of-freedom particle model of other flying vehicles currently in a flight state in the engineering calculation and simulation platform, constraining the maneuvering characteristics of other flying vehicles based on constraint rules, and realizing the tracking of the UAV by other flying vehicles according to the relative motion situation; building a reinforcement learning environment and establishing a corresponding reinforcement learning model in a programming language platform; inputting the position and attitude information of the UAV and the position and attitude information of other flying vehicles into the reinforcement learning model to output the flight actions in the UAV action library through the reinforcement learning model; when the UAV avoids other flying vehicles based on the flight actions, judging whether the UAV's avoidance is successful based on the maneuvering characteristics of other flying vehicles. If successful, a corresponding reward value is given to the reinforcement learning model based on the maneuvering area situation of the other flying vehicle to encourage the reinforcement learning model to optimize the avoidance strategy. It can be seen that the present application builds a six-degree-of-freedom model of the UAV and a three-degree-of-freedom particle model of other flying vehicles currently in a flight state in the engineering calculation and simulation platform, and combines the reinforcement learning environment and the reinforcement learning model, enabling the UAV to autonomously learn coping strategies. Even if other flying vehicles adopt a brand-new guidance subsystem, the UAV does not need to wait for expert analysis, but relies on the learning results accumulated in the simulation environment by itself, quickly selects a suitable flight action from the action library for avoidance, getting rid of the dependence on expert prior knowledge. Secondly, the reinforcement learning model in the present application is not based on the calculation results of a theoretical model that may have errors, but adjusts the flight actions of the UAV flexibly according to the coping patterns learned from past similar scenarios, effectively avoiding the interference of modeling errors on the avoidance decision. Further, after the UAV successfully avoids other flying vehicles, the present application gives a corresponding reward value to the reinforcement learning model according to the maneuvering area situation of the other flying vehicle, thus prompting the reinforcement learning model to optimize the avoidance strategy and forming a virtuous cycle. Finally, the present application can flexibly adjust parameters to simulate the characteristics of a variety of different environments with the help of the engineering calculation and simulation platform, enabling the UAV to stably improve the avoidance success rate with the avoidance ability obtained through training in various complex and changeable scenarios, greatly enhancing the scalability of the strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0039] Figure 1 It is a flowchart of a method for determining an avoidance strategy for an unmanned aerial vehicle based on reinforcement learning disclosed in the present application;
[0040] Figure 2 Schematic diagram of the basic framework of a reinforcement learning disclosed in this application;
[0041] Figure 3 Schematic diagram of the interaction simulation process between other aircraft and drones disclosed in this application;
[0042] Figure 4 Schematic diagram of the structure of a device for determining an avoidance strategy for drones based on reinforcement learning disclosed in this application;
[0043] Figure 5 Schematic diagram of the structure of an electronic device disclosed in this application. Specific implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] Traditional methods for drones to avoid other aircraft include the expert system method, differential game method, optimal control method, etc. The expert system method relies on the prior knowledge of experts. Once the subsystems of other aircraft or airplanes change, experts need to analyze the new subsystems and give new avoidance maneuver strategies again; both the differential game method and the optimal control method rely on clear and complete mathematical models, and it is essential to solve complex differential equations. However, the problem of an aircraft avoiding other aircraft involves multiple complex nonlinear systems, and it is inevitable to have errors in modeling each subsystem, which greatly increases the difficulty of solving this problem by the above methods. In addition, the methods for drones to avoid other aircraft usually only vaguely explain which maneuvers are more advantageous under a certain type of encounter situation, and the scalability is not strong.
[0046] Therefore, the embodiments of this application propose a solution for determining an avoidance strategy for drones based on reinforcement learning, which can get rid of the dependence on the prior knowledge of experts, avoid the interference of modeling errors on the avoidance decision, prompt the reinforcement learning model to optimize the avoidance strategy, form a virtuous cycle, and stably improve the avoidance success rate, greatly enhancing the scalability of the avoidance strategy.
[0047] The embodiments of this application disclose a method for determining an avoidance strategy for drones based on reinforcement learning. Refer to Figure 1 as shown, this method includes:
[0048] Step S11: Build a six-degree-of-freedom model of the drone in an engineering calculation and simulation platform, and establish a drone action library.
[0049] In this embodiment, first, an idealized setting is made for the drone. It is regarded as an ideal rigid body that is symmetric left and right, and it is clear that the control of the drone mainly relies on engine thrust and aerodynamic control surfaces. Next, before establishing the motion equations, the following two conditions need to be met: (1) The drone is regarded as a rigid body, and the mass of the drone remains constant; the ground coordinate axis system is selected as the inertial coordinate system. On this basis, the curvature of the earth is ignored, and it is assumed that the acceleration due to gravity does not change with the change of flight altitude; for a drone with a symmetric layout, it is clear that the plane of the drone is the symmetric plane of the drone. In this way, the geometric shape and mass distribution of the drone are symmetric about the plane; (2) The spatial motion of the drone is described using six degrees of freedom. The spatial motion of the drone specifically includes linear motion (the displacement of the center of mass) and angular motion (the rotation about the center of mass). Linear motion includes several motion forms such as the increase and decrease of flight speed, ascent and descent, and lateral movement. Angular motion includes pitch angular motion, yaw angular motion, and roll angular motion. Finally, a six-degree-of-freedom model of the drone is built on the engineering calculation and simulation platform (Matlab) to accurately simulate and analyze the motion state of the drone in space through the six-degree-of-freedom model of the drone.
[0050] Six-degree-of-freedom model of the drone:
[0051] ;
[0052] Among them, represents the speed of the drone, represents the track deviation angle, represents the track inclination angle, represents the climb rate, represents the velocity component of the drone in the x-axis direction, represents the velocity component of the drone in the y-axis direction, represents the velocity component of the drone in the z-axis direction.
[0053] ;
[0054] ;
[0055] ;
[0056] Among them, represents the derivative of the drone speed, represents the mass of the drone, represents the drag, represents the thrust, represents the angle of attack, represents the derivative of the angle of attack, represents the sideslip angle, represents the derivative of the sideslip angle, represents the acceleration due to gravity, represents the derivative of the track deviation angle, represents the lift force, represents the speed roll angle, represents the derivative of the speed roll angle, represents the lateral force, represents the derivative of the track inclination angle, represents the pitch rate, represents the derivative of the pitch rate, represents the roll rate, represents the derivative of the roll rate, represents the yaw rate, represents the derivative of the yaw rate, and and represent the components of the moment on the three axes of x, y, and z, represents the angular momentum.
[0057] For the established six-degree-of-freedom model of the UAV, based on the damper and stability augmentation design, the control law design for the longitudinal motion and lateral-directional motion of the UAV is carried out to obtain the outer-loop control command , represents the speed command, represents the altitude command, represents the aircraft roll command. By setting the command sequence, the UAV is made to complete the following seven basic maneuvering actions: uniform straight-and-level flight, maximum-force acceleration, maximum-force deceleration, climb, dive, horizontal left turn, and horizontal right turn. It should be noted that the design of the avoidance maneuver library is usually divided into two categories: one is the typical maneuver library based on classical maneuvers, and the other is the basic maneuver library based on common maneuvering methods. In this embodiment, the UAV action library is obtained by setting the command sequence, enabling the UAV to complete the above seven maneuvering actions. In this way, it can be executed and switched at any time, being flexible and simple, and not restricted by the completion time of the maneuvering actions.
[0058] In this embodiment, the update and switching of maneuver commands are implemented based on the following logic: The update of maneuver commands takes the step length of maneuver decision-making as the cycle to perform switching operations on maneuver action commands. Among them, the maneuver action determined by maneuver decision-making lasts for one maneuver decision step length, and the duration of the maneuver action command corresponding to this maneuver action is one command update cycle. Moreover, the update of maneuver commands has the characteristic of continuity. Each update of the command value is based on the speed, altitude, and roll angle command values when the previous maneuver action is completed, and corresponding increase or decrease operations are performed on this basis. Secondly, considering the characteristics of the constructed UAV model itself, its altitude and roll angle can change significantly within 5 s, and this characteristic can ensure the accurate execution of one maneuver action. Therefore, in this embodiment, the time interval of maneuver decision-making based on the reinforcement learning algorithm is set to 5 s. Finally, when a maneuver action needs to be switched, the speed, altitude, and roll angle command values should be adaptively changed synchronously. For example, when the aircraft completes the maximum afterburner climb maneuver and starts to execute the uniform linear horizontal flight maneuver, the speed command value should be set to the actual speed at the moment of switching from the climb maneuver to the uniform linear motion, and the altitude command value should also be the altitude corresponding to that moment. In this way, it can ensure the stable, continuous flight state of the UAV and conform to the flight logic, and then achieve the reasonable and smooth switching of maneuver commands.
[0059] Step S12: Build a three-degree-of-freedom particle model of other aircraft in the current flight state in the engineering calculation and simulation platform, constrain the maneuver characteristics of other aircraft based on the constraint rules, and realize the tracking of the UAV by other aircraft according to the relative motion situation; the relative motion situation is the relative motion situation between the other aircraft and the UAV.
[0060] Build a reinforcement learning environment in the programming language platform and establish a corresponding reinforcement learning model.
[0061] In this embodiment, a three-degree-of-freedom particle simplified model of other aircraft is established in the engineering calculation and simulation platform:
[0062] ;
[0063] Among them, represents the tangential overload, represents the normal overload, represents the track deviation angle, represents the derivative of the track deviation angle.
[0064] In this embodiment, the maneuvering characteristics of other aircraft include the flight time of other aircraft, the safety range of other aircraft, and the overload that other aircraft can withstand. Further, the maneuvering characteristics of other aircraft are constrained based on constraint rules, including: setting the simulation time of other aircraft in the engineering calculation and simulation platform to constrain the flight time of other aircraft through the simulation time of other aircraft; setting the size of the simulation space in the engineering calculation and simulation platform to constrain the safety range of other aircraft through the size of the simulation space; controlling the overload that other aircraft can withstand not to exceed a preset overload threshold to constrain the overload that other aircraft can withstand; where the overload that other aircraft can withstand includes tangential overload and normal overload. For example, for the flight time of other aircraft, setting the simulation time of other aircraft to 10 seconds in the engineering calculation and simulation platform means that the actual flight time of other aircraft also fluctuates around this magnitude, thereby achieving its constraint; for the safety range of other aircraft, setting the size of the simulation space to a cubic space with a length, width, and height of 10 kilometers in the engineering calculation and simulation platform, then the safety range of other aircraft is set within this space; for the overload that other aircraft can withstand, set , where represents the tangential overload, represents the normal overload, represents the preset tangential overload threshold, represents the preset normal overload threshold.
[0065] In this embodiment, other aircraft track the UAV according to the relative motion situation, and the relative motion situation is the relative motion situation between the other aircraft and the UAV. Among them, the relative motion model used to describe the relative motion situation is as follows:
[0066] ;
[0067] where represents the derivative of the distance between other aircraft and the UAV with respect to time, represents the relative course angle between other aircraft and the UAV, represents the derivative of the relative course angle with respect to time, represents the speed of other aircraft, is the guidance angle of other aircraft, is the guidance angle of the UAV, represents the derivative of the guidance angle with respect to time, represents the course angle of other aircraft, represents the course angle of the UAV, is the guidance coefficient, which reflects the response degree of the guidance system to the change of the course angle. The pitch control and yaw control of other aircraft are independent of each other. Thus, it can be obtained that:
[0068] ;
[0069] Among them, represents the pitch angle command of the pitch channel, represents the yaw angle command of the yaw channel; and respectively represent the derivatives of the components on the two channels of pitch and yaw, and respectively represent the guidance coefficients on the two channels of pitch and yaw. By adjusting the values of the guidance coefficients on the two channels, other aircraft can achieve better tracking effects.
[0070] Step S13: Build a reinforcement learning environment and establish a reinforcement learning model in a programming language platform.
[0071] Build a reinforcement learning environment and establish a reinforcement learning model in a programming language platform (python), and then establish UDP (User Datagram Protocol) communication between matlab and python to prepare for subsequent information transmission and model training. When building the reinforcement learning environment, in addition to using OpenAI Gym (a toolkit for developing and comparing reinforcement learning algorithms), TensorFlow (an open-source machine learning framework), or PyTorch (an open-source machine learning library), some auxiliary tools will also be used. For example, NumPy provides basic support for data processing, and Pandas is used for data cleaning and analysis. Especially when dealing with a large amount of data generated by the interaction between the agent and the environment, it can quickly organize and count the data.
[0072] Step S14: Input the position and attitude information of the drone and the position and attitude information of other aircraft into the reinforcement learning model to output the flight actions in the drone action library through the reinforcement learning model.
[0073] To effectively construct an intelligent decision-making mechanism, this embodiment introduces a Markov decision process. The Markov decision process is composed of a five-tuple constitutes. Among them, represents the state space, represents the action space, represents the state transition probability, represents the reward function, which is used to measure the return obtained by the agent when taking actions in different states, represents the discount factor, which is used to calculate the cumulative return. See Figure 2As shown, the agent continuously interacts with the environment, tries different actions, and adjusts its behavior strategy according to the obtained rewards, and finally learns the optimal strategy that can obtain the maximum cumulative reward in this environment. In this embodiment, the state space is constructed by selecting the position and attitude information of the UAV and the position and attitude information of other aircraft to comprehensively represent the three-dimensional state space. Among them, , , represent the position coordinates of the UAV, represents the roll angle of the UAV, represents the pitch angle of the UAV, represents the yaw angle of the UAV, , , represent the position coordinates of other aircraft, represents the roll angle of other aircraft, represents the pitch angle of other aircraft, represents the yaw angle of other aircraft. Through this information, the position and attitude of the UAV and other aircraft in space can be accurately described. The action space is established by numbering the basic maneuvering actions from 1 to 7 and representing them with the corresponding maneuvering action numbers. These basic maneuvering actions cover uniform linear level flight, maximum force acceleration, maximum force deceleration, climbing, diving, horizontal left turn, horizontal right turn, etc. In this way, a discrete action space is constructed, enabling the decision-making of the agent to correspond to specific and executable actions. The specific operation process of this step is as follows: input the position and attitude information of the UAV and the position and attitude information of other aircraft into the reinforcement learning model constructed based on the Markov decision process. This reinforcement learning model has strong analysis capabilities after learning a large amount of historical data and continuous optimization of the algorithm. It deeply analyzes the input state information, and through complex calculations and logical judgments, finally outputs appropriate flight actions from the pre-established UAV action library. This process realizes the efficient conversion from the current state information to specific action decisions.
[0074] Step S15: When the UAV avoids the other aircraft based on the flight action, judge whether the UAV successfully avoids based on the maneuvering characteristics of the other aircraft. If successful, give the corresponding reward value to the reinforcement learning model based on the maneuvering area situation of the other aircraft to encourage the reinforcement learning model to optimize the avoidance strategy.
[0075] In this embodiment, within the flight time of other aircraft and within the overload that other aircraft can withstand, it is determined whether the distance between the other aircraft and the UAV is greater than a preset value. If so, it is determined that the UAV has successfully evaded, and a corresponding reward value is given to the reinforcement learning model based on the situation of the maneuvering area of the other aircraft to encourage the reinforcement learning model to optimize the evasion strategy. The situation of the maneuvering area of the other aircraft includes the distance between the other aircraft and the UAV, the relative speed between the other aircraft and the UAV, the relative angle between the other aircraft and the UAV, the collision probability of the other aircraft, and the situation level of the area where the other aircraft is located. Among them, the situation level represents the importance degree of the area where the other aircraft is located when evading the UAV. Exemplarily, in a low-altitude route logistics transportation scenario, there is a logistics command and dispatch center, with dense UAV transportation bases deployed around and dense routes. This area is set as the first-priority area, while an open area at the edge of the city with only a few scattered routes is set as the second-priority area. In the first implementation manner, based on the distance between the other aircraft and the UAV, a first reward value is given to the reinforcement learning model through a first reward function to encourage the model to optimize the evasion strategy. For example, the first reward function can be: , represents the first reward value, is a constant representing the influence of distance on the reward. In this way, the distance between the other aircraft and the UAV is negatively correlated with the reward value, so that the other aircraft can obtain a higher reward when approaching the UAV, thus encouraging the other aircraft to approach the UAV.
[0076] In the second implementation manner, based on the relative speed between the other aircraft and the UAV, a second reward value is given to the reinforcement learning model through a second reward function to encourage the model to optimize the evasion strategy. For example, the first reward function can be: , represents the second reward value, is a constant representing the influence of relative speed on the reward. In this way, the relative speed between the other aircraft and the UAV is negatively correlated with the reward value. When the relative speed between the other aircraft and the UAV is large, the probability of collision increases. At this time, a negative reward value is set to prevent the UAV from choosing an evasion path with too large a relative speed. In addition, the situation of the maneuvering area of the other aircraft can also be the relative angle between the other aircraft and the UAV. The greater the change in the relative angle between the other aircraft and the UAV, the greater the difficulty of collision. At this time, a positive reward value is set to encourage the UAV to choose an evasion path with a larger change in relative angle.
[0077] In the third implementation, if the situation level of the area where other aircraft are located is the first priority and the collision probability of other aircraft is greater than the preset probability threshold, the reinforcement learning model is given a third reward value to encourage the model to optimize the avoidance strategy; if the situation level of the area where other aircraft are located is the second priority or the collision probability of other aircraft is not greater than the preset probability threshold, the reinforcement learning model is given a fourth reward value to encourage the model to optimize the avoidance strategy. The first priority is higher than the second priority, and the third reward value is greater than the fourth reward value. That is, when other aircraft are in a high-priority area and the probability of collision is high, a positive reward is given; if they are in a low-priority area or the success probability is low, a negative reward or a lower positive reward is given. It should be noted that for the above-mentioned collision probability, in this embodiment, a collision probability model can be established based on factors such as the flight state of other aircraft and the movement of the UAV, and this probability can be used as part of the situation quantification to reflect the collision probability of other aircraft in a certain state.
[0078] See Figure 3 as shown Figure 3 discloses a schematic diagram of the interaction simulation process between other aircraft and a UAV. First, the initial situation is set, including but not limited to relevant parameters such as the initial positions, speeds, and attitudes of other aircraft and the UAV, laying a foundation for the subsequent simulation process. After the initial situation is set, the iteration link of the position and attitude information of other aircraft is entered. In this way, during the simulation process, the position and attitude information of other aircraft will be continuously updated to simulate the dynamic changes of other aircraft during flight. Further, corresponding decision actions are input to avoid other aircraft. After the decision actions are input, the parameters of the UAV will be iteratively updated accordingly to reflect the actual response and dynamic changes of the UAV based on the decision actions. Further, it is judged whether the flight of other aircraft has ended. If the flight of other aircraft has not ended, the process will return to the iteration link of the position and attitude information of other aircraft, continue to update the information of other aircraft, and repeat the subsequent steps of inputting decision actions and iterating the UAV parameters, forming a loop until the flight of other aircraft ends. When the flight of other aircraft ends, the simulation end judgment link is entered. Here, it is judged whether the entire simulation process has ended. If the simulation has not ended, the process will return to the initial situation setting link to reset the initial situation and start a new round of simulation process. If the simulation ends, the simulation results are output. The simulation results include relevant information such as whether the UAV successfully avoids other aircraft, the flight trajectories of other aircraft, the flight trajectory of the UAV, and various performance indicators, and these results can be used for subsequent analysis, evaluation, and decision-making.
[0079] In summary, (1) through the tracking algorithm of other aircraft based on the proportional navigation law, the autonomous tracking maneuver of other aircraft to the UAV can be realized; (2) by establishing the basic maneuver action library of the UAV, the flight state of the UAV is stable, continuous and reasonable, and the logical switching of maneuver commands is realized; (3) through the method for the UAV to avoid incoming other aircraft based on reinforcement learning proposed in this application, the UAV can finally autonomously learn an effective maneuver avoidance strategy without relying on the prior knowledge of avoiding other aircraft.
[0080] In low-altitude logistics transportation, UAV A is responsible for parcel delivery and flies along the planned route at a speed of 8 m / s and an altitude of 120 m. During the flight, it encounters UAV B flying in the opposite direction (regarded as other aircraft), with a speed of 7 m / s, and there is a risk of collision between the two. Before performing the task, technicians built a six-degree-of-freedom model and action library of UAV A, and a three-degree-of-freedom particle model of UAV B on the engineering calculation and simulation platform. At the same time, the flight time, safety range and tolerable overload of UAV B are restricted. For example, the simulation flight time of UAV B is set to 15 s, the simulation space size (safety range) is a cube with a length, width and height of 10 km each, and its tolerable overload is controlled within a certain threshold. A reinforcement learning environment and model are built on the programming language platform, and a large amount of data is used for training to form a rich action library, including actions such as accelerating, decelerating, climbing, diving, and turning. UAV A obtains the position and attitude information of itself and UAV B in real time through on-board sensors and transmits this information to the reinforcement learning model. This model is constructed based on the Markov decision process, and according to the pre-designed reward function, the reward values of different actions are calculated by comprehensively considering factors such as distance and relative speed. According to the calculation results, the model selects the action of "quickly turning to the right and appropriately decelerating" from the action library and outputs it. After receiving the instruction, the flight control system of UAV A executes the corresponding action. During the execution process, the system monitors in real time. If the distance between the two aircraft is greater than 50 m within the flight time and tolerable overload of UAV B, it is determined that UAV A has successfully avoided. Then, according to the maneuver area situation of UAV B, such as the danger degree and collision probability of the area where it is located, the corresponding reward value is given to the reinforcement learning model. If UAV B is in a dangerous area and the collision probability is high, a high reward value is given; if it is in a safe area or the collision probability is low, a low reward value is given, so as to optimize the model decision-making strategy. Finally, UAV A successfully avoids UAV B and continues to complete the parcel delivery task. This decision-making method based on reinforcement learning can quickly respond to complex airspace situations compared with traditional methods and is applicable to the high-density flight scenarios of urban low-altitude logistics.
[0081] It can be seen that the present application proposes a method for determining an avoidance strategy for an unmanned aerial vehicle (UAV) based on reinforcement learning, including: building a six-degree-of-freedom model of the UAV in an engineering calculation and simulation platform and establishing a UAV action library; building a three-degree-of-freedom particle model of other flying vehicles currently in a flight state in the engineering calculation and simulation platform, constraining the maneuver characteristics of other flying vehicles based on constraint rules, and realizing the tracking of the UAV by other flying vehicles according to the relative motion situation; building a reinforcement learning environment and establishing a corresponding reinforcement learning model in a programming language platform; inputting the position and attitude information of the UAV and the position and attitude information of other flying vehicles into the reinforcement learning model to output the flight actions in the UAV action library through the reinforcement learning model; when the UAV avoids other flying vehicles based on the flight actions, judging whether the UAV has successfully avoided based on the maneuver characteristics of other flying vehicles, and if so, giving a corresponding reward value to the reinforcement learning model based on the maneuver area situation of the other flying vehicle to motivate the reinforcement learning model to optimize the avoidance strategy. It can be seen that the present application builds a six-degree-of-freedom model of the UAV and a three-degree-of-freedom particle model of other flying vehicles currently in a flight state in the engineering calculation and simulation platform, and combines the reinforcement learning environment and the reinforcement learning model, enabling the UAV to autonomously learn coping strategies. Even if other flying vehicles adopt a brand-new guidance subsystem, the UAV does not need to wait for expert analysis, but relies on its own learning results accumulated in the simulation environment, quickly selects a suitable flight action from the action library for avoidance, getting rid of the dependence on expert prior knowledge. Secondly, the reinforcement learning model in the present application is not based on the calculation results of a theoretical model that may have errors, but adjusts the flight actions of the UAV flexibly according to the coping patterns learned from past similar scenarios, effectively avoiding the interference of modeling errors on the avoidance decision. Further, after the UAV successfully avoids other flying vehicles, the present application gives a corresponding reward value to the reinforcement learning model according to the maneuver area situation of other flying vehicles, thus prompting the reinforcement learning model to optimize the avoidance strategy and forming a virtuous cycle. Finally, the present application can flexibly adjust parameters to simulate the characteristics of a variety of different environments with the help of the engineering calculation and simulation platform, enabling the UAV to stably improve the avoidance success rate with the avoidance ability obtained through training in various complex and changeable scenarios, greatly enhancing the scalability of the strategy.
[0082] Correspondingly, the embodiment of the present application also discloses a device for determining an avoidance strategy for an unmanned aerial vehicle based on reinforcement learning, as shown in Figure 4 The device includes:
[0083] A UAV model building module 11, configured to build a six-degree-of-freedom model of the UAV in an engineering calculation and simulation platform and establish a UAV action library;
[0084] An other flying vehicle model building module 12, configured to build a three-degree-of-freedom particle model of other flying vehicles currently in a flight state in the engineering calculation and simulation platform;
[0085] The other aircraft characteristic constraint module 13 is used to constrain the maneuvering characteristics of other aircraft based on constraint rules and enable other aircraft to track the UAV according to the relative motion situation; the relative motion situation is the relative motion situation between the other aircraft and the UAV.
[0086] The reinforcement learning model building module 14 is used to build a reinforcement learning environment in a programming language platform and establish a corresponding reinforcement learning model.
[0087] The flight action output module 15 is used to input the position and attitude information of the UAV and the position and attitude information of the other aircraft into the reinforcement learning model, so as to output the flight actions in the UAV action library through the reinforcement learning model.
[0088] The avoidance strategy optimization module 16 is used to, when the UAV avoids the other aircraft based on the flight actions, judge whether the UAV's avoidance is successful based on the maneuvering characteristics of the other aircraft. If successful, a corresponding reward value is given to the reinforcement learning model based on the maneuvering area situation of the other aircraft to motivate the reinforcement learning model to optimize the avoidance strategy.
[0089] Among them, for the more specific working processes of the above-mentioned various modules, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be elaborated here.
[0090] As can be seen, the present application proposes a method for determining an avoidance strategy for an unmanned aerial vehicle (UAV) based on reinforcement learning, including: building a six-degree-of-freedom model of the UAV in an engineering calculation and simulation platform, and establishing a UAV action library; building a three-degree-of-freedom particle model of other aircraft currently in a flight state in the engineering calculation and simulation platform, constraining the maneuvering characteristics of other aircraft based on constraint rules, and realizing the tracking of the UAV by other aircraft according to the relative motion situation; building a reinforcement learning environment and establishing a corresponding reinforcement learning model in a programming language platform; inputting the position and attitude information of the UAV and the position and attitude information of other aircraft into the reinforcement learning model, so as to output the flight actions in the UAV action library through the reinforcement learning model; when the UAV avoids other aircraft based on the flight actions, judging whether the UAV's avoidance is successful based on the maneuvering characteristics of other aircraft. If successful, a corresponding reward value is given to the reinforcement learning model based on the maneuvering area situation of the other aircraft to encourage the reinforcement learning model to optimize the avoidance strategy. As can be seen, the present application builds a six-degree-of-freedom model of the UAV and a three-degree-of-freedom particle model of other aircraft currently in a flight state in the engineering calculation and simulation platform, and combines the reinforcement learning environment and the reinforcement learning model, enabling the UAV to autonomously learn coping strategies. Even if other aircraft adopt a brand-new guidance subsystem, the UAV does not need to wait for expert analysis, but relies on the learning results accumulated in the simulation environment by itself, quickly selects a suitable flight action from the action library for avoidance, and gets rid of the dependence on expert prior knowledge. Secondly, the reinforcement learning model in the present application is not based on the calculation results of a theoretical model that may have errors, but adjusts the flight actions of the UAV flexibly according to the coping patterns learned from past similar scenarios, effectively avoiding the interference of modeling errors on the avoidance decision. Further, after the UAV successfully avoids other aircraft, the present application gives a corresponding reward value to the reinforcement learning model according to the maneuvering area situation of the other aircraft, so as to prompt the reinforcement learning model to optimize the avoidance strategy and form a virtuous cycle. Finally, the present application can flexibly adjust parameters to simulate the characteristics of a variety of different environments with the help of the engineering calculation and simulation platform, enabling the UAV to stably improve the avoidance success rate with the avoidance ability obtained through training in various complex and changeable scenarios, greatly enhancing the scalability of the strategy.
[0091] Furthermore, an embodiment of the present application also provides an electronic device. Figure 5 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of the present application.
[0092] Figure 5Schematic diagram of the structure of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the method for determining a UAV avoidance strategy based on reinforcement learning disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0093] In this embodiment, the power supply 26 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 25 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed on it here; the input / output interface 24 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are made here.
[0094] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon may include a computer program 221, and the storage method may be short-term storage or permanent storage. Among them, in addition to the computer program that can be used to complete the method for determining a UAV avoidance strategy based on reinforcement learning executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 221 may further include a computer program that can be used to complete other specific tasks.
[0095] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the method for determining a UAV avoidance strategy based on reinforcement learning disclosed above is implemented.
[0096] For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0097] The various embodiments in this application are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference may be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and reference may be made to the description in the method part for related parts.
[0098] Those skilled in the art may further realize that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0099] The steps of the methods or algorithms described in connection with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0100] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0101] The above has introduced in detail a method, device, equipment, and storage medium for determining an avoidance strategy for an unmanned aerial vehicle based on reinforcement learning provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A method for determining a drone avoidance strategy based on reinforcement learning, characterized in that: include: Build a six-degree-of-freedom model of a UAV in the engineering computing simulation platform and establish a UAV motion library; A three-degree-of-freedom particle model of other aircraft currently in flight is constructed in the engineering computing simulation platform, the maneuverability characteristics of other aircraft are constrained based on constraint rules, and the tracking of the UAV by other aircraft is achieved according to the relative motion situation; the relative motion situation is the relative motion situation between the other aircraft and the UAV; Build a reinforcement learning environment in the programming language platform and establish a corresponding reinforcement learning model; Inputting the position and attitude information of the UAV and the position and attitude information of the other aircraft into the reinforcement learning model, so as to output the flight actions in the UAV action library through the reinforcement learning model; When the UAV avoids the other aircraft based on the flight action, judging whether the UAV avoids the other aircraft successfully based on the maneuvering characteristics of the other aircraft, and if successful, giving the reinforcement learning model a corresponding reward value based on the situation of the maneuvering area of the other aircraft, so as to encourage the reinforcement learning model to optimize the avoidance strategy; Wherein, the maneuverability characteristics of other aircraft include the flight time of other aircraft, the safety range of other aircraft, and the tolerable overload of other aircraft. Then, constraining the maneuverability characteristics of other aircraft based on the constraint rules includes: Setting the simulation time of other aircraft in the engineering calculation simulation platform so as to constrain the flight time of the other aircraft through the simulation time of the other aircraft; The size of the simulation space is set in the engineering calculation simulation platform so as to constrain the safety range of the other aircraft through the size of the simulation space; The overload that the other aircraft can withstand is controlled to be no greater than a preset overload threshold, so as to constrain the overload that the other aircraft can withstand; wherein the overload that the other aircraft can withstand includes tangential overload and normal overload.
2. The method for determining a drone avoidance strategy based on reinforcement learning according to claim 1, characterized in that: The determining whether the UAV successfully avoids the maneuver based on the maneuver characteristics of the other aircraft includes: Within the flight time of the other aircraft and within the tolerable overload of the other aircraft, determine whether the distance between the other aircraft and the drone is greater than a preset value. If so, determine that the drone has successfully avoided the obstacle.
3. The method for determining a drone avoidance strategy based on reinforcement learning according to claim 1 or 2, characterized in that: The situation of the maneuvering area of other aircraft includes the distance between the other aircraft and the UAV, the relative speed between the other aircraft and the UAV, the collision probability of the other aircraft and the situation level of the area where the other aircraft is located; wherein the situation level represents the importance of the area where the other aircraft is located to the UAV when avoiding it.
4. The method for determining a drone avoidance strategy based on reinforcement learning according to claim 3 is characterized in that: Giving the reinforcement learning model a corresponding reward value based on the situation of other aircraft in the maneuvering area to encourage the reinforcement learning model to optimize the avoidance strategy includes: Based on the distance between the other aircraft and the UAV, a first reward value is given to the reinforcement learning model through a first reward function to encourage the reinforcement learning model to optimize the avoidance strategy; wherein, in the first reward function, the distance between the other aircraft and the UAV is negatively correlated with the reward value.
5. The method for determining a drone avoidance strategy based on reinforcement learning according to claim 3 is characterized in that: Giving the reinforcement learning model a corresponding reward value based on the situation of other aircraft in the maneuvering area to encourage the reinforcement learning model to optimize the avoidance strategy includes: Based on the relative speed between the other aircraft and the UAV, a second reward value is given to the reinforcement learning model through a second reward function to encourage the reinforcement learning model to optimize the avoidance strategy; wherein, in the second reward function, the relative speed between the other aircraft and the UAV is negatively correlated with the reward value.
6. The method for determining a drone avoidance strategy based on reinforcement learning according to claim 3 is characterized in that: Giving the reinforcement learning model a corresponding reward value based on the situation of other aircraft in the maneuvering area to encourage the reinforcement learning model to optimize the avoidance strategy includes: If the situation level of the area where the other aircraft is located is the first priority and the collision probability of the other aircraft is greater than the preset probability threshold, a third reward value is given to the reinforcement learning model to encourage the reinforcement learning model to optimize the avoidance strategy; If the situation level of the area where the other aircraft is located is the second priority or the collision probability of the other aircraft is not greater than the preset probability threshold, giving the reinforcement learning model a fourth reward value to encourage the reinforcement learning model to optimize the avoidance strategy; The first priority is higher than the second priority, and the third reward value is greater than the fourth reward value.
7. A device for determining a drone avoidance strategy based on reinforcement learning, characterized in that: include: The UAV model building module is used to build a six-degree-of-freedom model of the UAV in the engineering computing simulation platform and establish a UAV motion library; Other aircraft model building module, used to build three-degree-of-freedom particle models of other aircraft currently in flight in the engineering calculation simulation platform; Other aircraft characteristic constraint module, used for constraining the maneuvering characteristics of other aircraft based on constraint rules, and realizing tracking of the UAV by other aircraft according to relative motion situation; the relative motion situation is the relative motion situation between the other aircraft and the UAV; The reinforcement learning model building module is used to build a reinforcement learning environment in the programming language platform and establish a corresponding reinforcement learning model; A flight action output module, used for inputting the position and attitude information of the UAV and the position and attitude information of the other aircraft into the reinforcement learning model, so as to output the flight actions in the UAV action library through the reinforcement learning model; an avoidance strategy optimization module, configured to determine whether the UAV has successfully avoided the other aircraft based on the maneuvering characteristics of the other aircraft when the UAV avoids the other aircraft based on the flight action, and if successful, give the reinforcement learning model a corresponding reward value based on the situation of the maneuvering area of the other aircraft to encourage the reinforcement learning model to optimize the avoidance strategy; Wherein, the maneuverability characteristics of other aircraft include the flight time of other aircraft, the safety range of other aircraft, and the tolerable overload of other aircraft. Then, constraining the maneuverability characteristics of other aircraft based on the constraint rules includes: Setting the simulation time of other aircraft in the engineering calculation simulation platform so as to constrain the flight time of the other aircraft through the simulation time of the other aircraft; The size of the simulation space is set in the engineering calculation simulation platform so as to constrain the safety range of the other aircraft through the size of the simulation space; The overload that the other aircraft can withstand is controlled to be no greater than a preset overload threshold, so as to constrain the overload that the other aircraft can withstand; wherein the overload that the other aircraft can withstand includes tangential overload and normal overload.
8. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method for determining a drone avoidance strategy based on reinforcement learning as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, it implements the method for determining a drone avoidance strategy based on reinforcement learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Six-degree-of-freedom unmanned combat aerial vehicle short-range dogfight method based on simplified model machine game
CN105204512A
Agent dynamic autonomous obstacle avoidance motion method and device, server and storage medium
CN113589810A