UAV intelligent decision-making method, system and storage medium based on reinforcement learning
By generating a decision-making solution based on the detection result update model and using the boundary grid reward obtained by differential game, the problem that multi-agent reinforcement learning is difficult to accurately design the reward function in the DCA problem is solved, real-time optimization of the drone formation defense strategy and improving the defense effect.
Patent Information
- Application Number
- CN202510176233.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Multi-agent reinforcement learning is difficult to accurately design the reward function in defensive anti-aircraft operations (DCA) problems, making it difficult to effectively optimize the drone formation defense strategy.
By obtaining the drone formation information of the defensive party and the drone formation defense strategy, the detection results are used to generate a decision plan update model, including state space, decision plan, reward return and profit function, and the boundary grid reward obtained based on differential game is used to ensure that the reward mechanism can accurately reflect the goal achievement of both parties.
Real-time scheduling and optimization of drone formation defense strategies has been realized, and the defense effect has been improved, so that drones can gradually learn effective defense strategies in complex environments.
Smart Images

Figure CN119647631B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicles, and particularly relates to an intelligent decision-making method, system and storage medium for unmanned aerial vehicles based on reinforcement learning. Background Art
[0002] With the improvement of the intelligence level and the enhancement of the performance of unmanned aerial vehicles, using multi-unmanned aerial vehicle formations to execute aerial missions has become a research hotspot in the field. The attacker can use unmanned aerial vehicles, with their wide coverage and powerful firepower, to pose a serious threat to the core assets of the defender. Therefore, anti-air combat has become an indispensable key element in military strategic planning, which covers two core strategies: offensive counter-air combat (OCA) and defensive counter-air combat (DCA). The main goal of OCA is to maximize the weakening of the attacker's combat effectiveness near its source by destroying, disrupting or suppressing the attacker's forces; while DCA focuses on intercepting aerial targets to ensure the safety and stability of the defender's important assets by reducing their threat levels.
[0003] As a cutting-edge technology in the field of artificial intelligence, multi-agent reinforcement learning has achieved remarkable results in complex game scenarios such as StarCraft and Texas Hold'em with its self-learning and strong exploration capabilities. In recent years, this method has also been widely applied to solve the DCA problem. The reward function is an important part of the reinforcement learning algorithm. Due to the existence of the reward function, the algorithm can iterate network parameters and finally train an algorithm model that meets the research problem. Generally speaking, rewards are artificially created based on experience. In simple research problems, the reward function is easy to construct and understand. However, as the complexity of the problem increases, such as the DCA problem, it is more difficult to design the reward, especially the intermediate continuous reward is even more difficult.
[0004] In view of this, it is necessary to provide a new intelligent decision-making solution for unmanned aerial vehicles based on reinforcement learning to make up for the deficiency that the reward function in the multi-agent reinforcement learning for the DCA problem is difficult to accurately design. Summary of the Invention
[0005] (I) Technical Problems to be Solved
[0006] Aiming at the deficiencies of the prior art, the present invention provides an intelligent decision-making method, system and storage medium for unmanned aerial vehicles based on reinforcement learning, and solves the technical problem that the reward function in the multi-agent reinforcement learning for the DCA problem is difficult to accurately design.
[0007] (II) Technical Solutions
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0009] An intelligent decision-making method for unmanned aerial vehicles based on reinforcement learning includes:
[0010] Obtain information on the defending UAV formation and the UAV formation defense strategy; the defending UAV formation includes multiple UAVs; the UAV formation defense strategy is used to indicate the decision-making plan for the multiple UAVs, and the decision-making plan includes defense actions and flight directions; the defense actions include aerial patrol, calling for friendly aircraft, and the UAV's own retreat;
[0011] Control the defending UAV formation to execute the UAV formation defense strategy;
[0012] Obtain the patrolling UAVs in the patrolling state and obtain the detection results of the patrolling UAVs;
[0013] Generate a decision-making plan update model based on the detection results; the decision-making plan update model includes a state space, a decision-making plan, a reward return, and a benefit function; the state space is generated from the detection results; the reward return includes at least the boundary fence reward obtained based on differential game;
[0014] Solve the decision-making plan update model and update the decision-making plan of the patrolling UAVs according to the model solution;
[0015] Update the UAV formation defense strategy according to the updated decision-making plan and control the defending UAV formation to execute the updated UAV formation defense strategy.
[0016] Preferably, the solving of the decision-making plan update model includes:
[0017] Obtain the UAV state information of the attacking UAVs according to the state space;
[0018] Calculate the reward return including at least the boundary fence reward according to the UAV state information;
[0019] Solve the benefit function according to the reward return.
[0020] Preferably, the UAV state information includes the speed ratio between the patrolling UAVs and the attacking UAVs;
[0021] The process of obtaining the boundary fence reward includes:
[0022] Based on the speed ratio and the position of the defense position of the defending UAV formation, use differential game to construct the boundary fence of the patrolling UAVs to construct the pursuit area and the escape area of the patrolling UAVs;
[0023] Obtaining the boundary fence reward includes:
[0024]
[0025] where the subscript CCZ , C EZ respectively represent all attacking drones located in the pursuit area and the escape area of the patrol drone; represents the reduction in the distance between the patrol drone and all attacking drones in the pursuit area within a single cycle; represents the reduction in the distance between the patrol drone and all attacking drones in the escape area within a single cycle;
[0026] The drone status information further includes the number of the first attacking drones defeated by the patrol drone and the number of the second attacking drones that enter the defense position of the defense drone formation;
[0027] Calculating a reward return including at least the fence reward based on the drone status information further includes:
[0028] Obtaining a first reward return based on the drone status information includes:
[0029]
[0030] wherein, R target represents the first reward return; m 1 represents the number of the first attacking drones, m 2 represents the number of the second attacking drones;
[0031] represents the reward score for defeating an attacking drone. If the attacking drone j is defeated within the escape area of the patrol drone after the patrol drone calls for friendly aircraft, the reward score is upgraded to a preset multiple of the first basic reward score; otherwise, it remains the first basic reward score; j
[0032] j represents the reward score for an attacking drone j entering the defense position. If the attacking drone enters the defense position within the pursuit area of the patrol drone after the patrol drone calls for friendly aircraft, the reward score is upgraded to a preset multiple of the second basic reward score; otherwise, it remains the second basic reward score;
[0033] Obtaining a second reward return includes:
[0034]
[0035] R wherein, m time represents the second reward return; m 3 represents the number of patrol drones;
[0036] Calculate the reward return, including:
[0037]
[0038] Wherein, R represents the reward return.
[0039] Preferably, the revenue function is:
[0040]
[0041] Wherein, argmaxf( x ) represents the variable x for which the objective function f( x ) takes the maximum value; θ represents the flight direction; R ( t ) represents the reward return for the t th cycle; γ represents the discount factor, γ t represents the t th power of the discount factor.
[0042] Preferably, after obtaining the patrol UAV in the patrol state, it further includes:
[0043] Control the patrol UAV to time and detect the timing duration;
[0044] If the patrol UAV does not detect the attacking UAV within the preset patrol cycle, update the decision-making scheme of the patrol UAV to set the defense action of the patrol UAV to retreat by itself;
[0045] If the patrol UAV detects the attacking UAV within the preset patrol cycle, update the timing duration to 0 and control the patrol UAV to time.
[0046] Preferably, after obtaining the detection result of the patrol UAV, it further includes:
[0047] Obtain the patrol threat value of the patrol UAV;
[0048] If the patrol threat value is greater than or equal to the preset threat value threshold, update the decision-making scheme of the patrol UAV to set the defense action of the patrol UAV to call for friendly aircraft;
[0049] If the patrol threat value is less than the preset threat value threshold, perform the step of generating a decision-making scheme update model according to the detection result.
[0050] Preferably, obtaining the patrol threat value of the patrol drone includes:
[0051] According to the detection result of the patrol drone, the target attacking drone detected by the patrol drone is obtained;
[0052] Obtain the action strategy of the target attacking drone; the action strategy includes an advance strategy, a retreat strategy and other strategies; the advance strategy is the behavior of the target attacking drone advancing toward the defensive position of the defending drone formation, the retreat strategy is the behavior of the target attacking drone retreating, and the other strategies are strategic behaviors other than the advance strategy and the retreat strategy;
[0053] According to the action strategy, obtain the strategy threat value;
[0054] Get the location threat value, including:
[0055]
[0056] Among them, W distance indicates the location threat value, d max indicates the maximum patrol distance of the drones in the defending drone formation, Indicates the target attacking drone j * Distance from defensive positions;
[0057] Calculate the patrol threat value of the patrol drone according to the strategy threat value and the location threat value.
[0058] Preferably, the model solution includes probability distributions of multiple decision-making schemes;
[0059] Updating the decision plan of the patrol drone according to the model solution, including:
[0060] According to the model solution, the decision plan with the highest probability is obtained, and updated as the decision plan of the patrol drone in the next cycle;
[0061] Updating the drone formation defense strategy according to the updated decision plan, including:
[0062] According to the updated decision plan, the number of third drones whose defensive action is to call friendly drones is obtained, and the number of retreating drones whose defensive action is to retreat the drone itself is obtained;
[0063] Acquire the backup drones corresponding to the number of the third drones in the defense drone formation, and update the decision plan of the backup drones to set the defense action of the backup drones to air patrol;
[0064] Set the candidate UAV as a patrol UAV and cancel the setting of the retreat UAV as a patrol UAV;
[0065] Count the updated decision-making plans of multiple UAVs in the defender's UAV formation to generate the UAV formation defense strategy for the next cycle.
[0066] An intelligent UAV decision-making system based on reinforcement learning, the system includes:
[0067] An information acquisition module, configured to acquire defender UAV formation information and UAV formation defense strategies; the defender UAV formation includes multiple UAVs; the UAV formation defense strategy is used to indicate the decision-making plans of the multiple UAVs, and the decision-making plans include defense actions and flight directions; the defense actions include aerial patrol, calling for friendly aircraft, and the UAV's own retreat;
[0068] A strategy execution module, configured to control the defender UAV formation to execute the UAV formation defense strategy;
[0069] A result detection module, configured to acquire the patrol UAVs in the patrol state and acquire the detection results of the patrol UAVs;
[0070] A model generation module, configured to generate a decision-making plan update model according to the detection results; the decision-making plan update model includes a state space, a decision-making plan, a reward return, and a benefit function; the state space is generated from the detection results; the reward return includes at least the boundary reward obtained based on differential game;
[0071] A plan update module, configured to solve the decision-making plan update model and update the decision-making plans of the patrol UAVs according to the model solution;
[0072] A strategy update module, configured to update the UAV formation defense strategy according to the updated decision-making plans and control the defender UAV formation to execute the updated UAV formation defense strategy.
[0073] A storage medium that stores a computer program for intelligent UAV decision-making based on reinforcement learning, wherein the computer program causes a computer to execute the intelligent UAV decision-making method based on reinforcement learning as described above
[0074] (III) Beneficial effects
[0075] The present invention provides an intelligent UAV decision-making method, system, and storage medium based on reinforcement learning. Compared with the prior art, it has the following beneficial effects:
[0076] In the present invention, on the one hand, the detection results of the unmanned aerial vehicle (UAV) can be used to update the decision-making scheme of the UAV, so as to adjust the scale of the UAV formation in real time when performing tasks, thereby improving the defense effect. On the other hand, a boundary grid reward obtained based on differential game is designed to ensure that the reward mechanism can accurately reflect the achievement of the goals of both sides of the game. Through continuous iteration and optimization of reinforcement learning, the UAV can gradually learn an effective UAV formation defense strategy in a complex environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0078] Figure 1 Schematic diagram of the scenario of the UAV intelligent decision-making method based on reinforcement learning provided by the embodiment of the present invention;
[0079] Figure 2 Schematic diagram of the attack range and detection range of the UAV provided by the embodiment of the present invention;
[0080] Figure 3 Schematic diagram of the state space provided by the embodiment of the present invention;
[0081] Figure 4 Schematic diagram of the boundary grid structure provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0083] By providing a UAV intelligent decision-making method, system, and storage medium based on reinforcement learning, the embodiments of the present application solve the technical problem that it is difficult to accurately design the reward function in the multi-agent reinforcement learning for the DCA problem.
[0084] To better understand the above technical solutions, the following will describe the above technical solutions in detail in combination with the drawings in the specification and specific embodiments.
[0085] Multiple drone formations can perform air missions, and the attacker can use drone formations to attack the defender's defensive positions. For the defender, anti-air operations can be performed to prevent the attacker's drone formation from entering the defensive position.
[0086] The defender can implement two strategies: offensive counter-air operations (OCA) and defensive counter-air operations (DCA). Among them, the main goal of OCA is for the defender's drone to destroy, disrupt or suppress the attacker's forces, and minimize their combat effectiveness near their source. DCA is for the defender's drone to intercept the attacker's drone, and by reducing its threat level, ensure the safety and stability of the defender's important assets.
[0087] DCA, as a key strategy to ensure airspace security, covers two major aspects: active air defense and missile defense. It relies on a variety of asset types and system integration, including but not limited to fighter jets, surface-to-air missiles, anti-aircraft guns, electromagnetic warfare systems, and ballistic missile defense systems. These systems work together to destroy enemy forces or effectively weaken their offensive capabilities. When performing DCA missions, the fighter jets used may be traditional manned models or advanced unmanned aerial vehicles (UAVs). In particular, weaponized drones designed specifically for combat, namely unmanned combat aerial vehicles (UCAVs), can be remotely controlled and have the ability to perform tasks autonomously.
[0088] When using the DCA strategy, the defender usually fixes the size of the drones that perform the mission. For example, if the drone formation consists of 5 drones, the defender usually lets all 5 drones take off and perform the defense mission. In the embodiment of the present application, performing the defense mission means that the defender's drone patrols in the air to detect the attacker's drone and attack it.
[0089] However, the arrival of the attacking drones is often uncertain, and the number of the attacking drones is usually unknown. The attacking drones may advance to the defense position in batches and at multiple time periods. The defender's use of a preset scale of drones may result in a waste or shortage of resources, making it difficult to respond to emergencies and resulting in poor defense results.
[0090] In particular, multi-agent reinforcement learning methods are also widely used to solve DCA problems. The present application also recognizes that the reward function is an important part of the reinforcement learning algorithm. Due to the existence of the reward function, the algorithm can iterate the network parameters and finally train an algorithm model that meets the research problem. Different from simple research problems, the reward function is easy to construct and understand. Since the DCA problem is more complex, its rewards are also more difficult to design, especially the intermediate continuous rewards.
[0091] Example 1:
[0092] To solve the above problems, an embodiment of the present invention provides an intelligent decision-making method for drones based on reinforcement learning, which is executed by a computer and includes the following steps:
[0093] An intelligent decision-making method for drones based on reinforcement learning, characterized by including:
[0094] S1. Obtain the information of the defender's drone formation and the drone formation defense strategy; the defender's drone formation includes multiple drones; the drone formation defense strategy is used to indicate the decision-making scheme of the multiple drones, and the decision-making scheme includes defense actions and flight directions; the defense actions include aerial patrol, calling for friendly aircraft, and the retreat of the own aircraft;
[0095] S2. Control the defender's drone formation to execute the drone formation defense strategy;
[0096] S3. Obtain the patrolling drones in the patrolling state and obtain the detection results of the patrolling drones;
[0097] S4. Generate a decision-making scheme update model according to the detection results; the decision-making scheme update model includes a state space, a decision-making scheme, a reward return, and a revenue function; the state space is generated from the detection results; the reward return includes at least the boundary fence reward obtained based on differential game;
[0098] S5. Solve the decision-making scheme update model and update the decision-making scheme of the patrolling drones according to the model solution;
[0099] S6. Update the drone formation defense strategy according to the updated decision-making scheme and control the defender's drone formation to execute the updated drone formation defense strategy.
[0100] Figure 1 It is a schematic diagram of the scenario of the intelligent decision-making method for drones based on reinforcement learning provided by the embodiment of the present invention. The following is a specific analysis of each step:
[0101] In step S1, obtain the information of the defender's drone formation and the drone formation defense strategy.
[0102] Among them, the defender's drone formation may include multiple drones. The information of the defender's drone formation includes the information of each drone in the defender's drone formation. All drones in the defender's drone formation can adopt the same model, that is, parameters such as speed and maximum fuel power are the same. Each drone can be numbered in advance for distinction.
[0103] When the defense UAV formation is performing a defense mission, it can execute according to the UAV formation defense strategy. The UAV formation defense strategy can indicate the action strategy of each UAV, which is called the decision-making plan in the embodiments of the present application.
[0104] The decision-making plan can be represented by a two-dimensional vector ( θ , z ). Among them, θ represents the flight direction, that is, the maneuvering angle of the UAV. The flight direction is a continuous value, and the range is (0, 2 π ), representing the maneuvering angle of the UAV. z represents the defense action, that is, the action content that the UAV should execute. The defense actions include air patrol, calling for friendly aircraft, and the UAV's own retreat. The UAV can execute one of these defense actions at the same time.
[0105] Air patrol means that the UAV flies and patrols in the air and detects the UAVs of the attacking party. In the embodiments of the present application, the UAVs of the attacking party are called attacking UAVs. The UAV has a detection function and can detect the situation of other UAVs within a certain area in front. It can detect the attacking UAVs and also detect the friendly UAVs.
[0106] Calling for friendly aircraft means that the UAV detects attacking UAVs beyond its own processing capacity. For example, the number of attacking UAVs is large, so it calls the ground formation of the defense position to send friendly aircraft for reinforcement.
[0107] The UAV's own retreat means that the UAV retreats and flies back to the defense position, including the need to retreat due to factors such as insufficient battery power or fuel to continue patrolling, or the UAV being damaged itself.
[0108] All UAVs can be equipped with fire attack weapons and electronic jamming weapons, and the UAV can attack a certain enemy UAV. Affected by weapon factors, the weapons of the UAV have a certain effective area, that is, the attack of the UAV is limited within a certain range. In the embodiments of the present application, the weapon effective area is called the attack range.
[0109] The attack range of the UAV can be considered as a part of the detection range, that is, the attack range is smaller than the detection range. Figure 2 Shows a schematic diagram of the attack range and detection range of the UAV in some embodiments.
[0110] When the defensive drone detects the attacking drone, that is, when a certain attacking drone enters the detection range of the defensive drone, the defensive drone will fly towards the attacking drone until the attacking drone enters the attack range of the defensive drone, and the defensive drone will launch an attack on the attacking drone through its weapons. At the same time, in the embodiments of the present application, it is set that when the defensive drone encounters an attacking drone within its attack range, it will surely destroy the attacking drone.
[0111] When the defensive drone detects the attacking drone and flies towards it, it will lock on to the attacking drone and keep chasing it until it enters the attack range. At the same time, in the embodiments of the present application, it is set that if the defensive drone detects an attacking drone entering the defensive position of the defense side, it is considered that the attacking drone has completed the penetration mission, and the defense of the defensive drone against the attacking drone fails. At this time, the attacking drone is no longer considered, and the defensive drone will no longer pursue the attacking drone. The defensive drone resumes patrolling to detect the next attacking drone.
[0112] It should be noted that in order to cope with the variability of the attacking drones, in the embodiments of the present application, it is set that the number of drones performing the patrol mission at the same time is variable, that is, the scale of the drones flying for patrol can be adjusted. For the formation of defensive drones, only some of the drones may perform flight patrol, while the other drones stay in the defensive position for standby.
[0113] At the same time, in the embodiments of the present application, a scheduling period is preset in advance, for example, it can be 2 seconds. The formation of defensive drones can update the defense strategy of the drone formation according to the scheduling period. After determining a defense strategy for the drone formation, the drone formation can execute this strategy for the duration of a scheduling period, and then update the defense strategy of the drone formation to adjust the scale of the drones flying for patrol.
[0114] Before the defense mission starts, all the drones in the formation of defensive drones are located within the defensive position. At this time, an initial defense strategy for the drone formation can be preset, for example, the decision-making scheme of each drone can be set randomly.
[0115] In step S2, control the formation of defensive drones to execute the defense strategy of the drone formation.
[0116] At the start of the defense mission, the formation of defensive drones can execute the preset initial defense strategy for the drone formation. The execution duration of this initial strategy is the preset scheduling period. After the execution duration reaches this scheduling period, the defense strategy of the drone formation can be updated.
[0117] In step S3, obtain the patrolling UAVs in the patrolling state, and obtain the detection results of the patrolling UAVs.
[0118] To update the UAV formation defense strategy, the task execution situation of each UAV in the current scheduling period can be obtained first. In the embodiments of the present application, the UAVs that are currently flying and in the patrolling state are referred to as patrolling UAVs.
[0119] All the patrolling UAVs in the current period can be determined, and the detection results of each patrolling UAV can be obtained. The detection result refers to the detection situation of the UAVs within the detection range of the UAV, including the situation of the detected attacking UAVs. It should be noted that the UAV can detect in real time and can detect multiple times within a scheduling period. In the embodiments of the present application, it is set that the obtained detection result is the detection situation of the UAV at the last time point of the scheduling period.
[0120] In step S4, generate a decision scheme update model according to the detection results.
[0121] The decision scheme update model is used to update the decision scheme of the patrolling UAVs. The decision scheme update model may include a state space, the current UAV formation defense strategy, a reward return, and a revenue function.
[0122] The decision scheme update model can use Markov game (MG) to formulate the DCA problem. MG is an extension of Markov decision process (MDP) in multi-agent problems.
[0123] The MG of the DCA problem is defined as a tuple < N, S, A, R, T >, where N is n the set of S and A are the joint state space and the joint action space of the multi-UAVs respectively. Among them, S= ( S 1 ,S 2 ,…,S n ), A= ( A 1 ,A 2 ,…,A n ), S i and A i are the state space and the action space of the UAV i respectively. The reward function of the UAV i is Ri Indicates taking an action at state S i and the obtained reward. The state transition function A i is the probability of entering the next state T when taking an action at state S and taking action A . In the embodiments of the present application, the next state is the next scheduling cycle. Since all the drones in the environment are homogeneous, all drone parameters are shared and a single policy S’ is shared, which represents the probability distribution of the actions selected by the drones. The revenue function aims to maximize the cumulative reward. π w In some embodiments, the state space is generated from the detection results. The detection results can characterize the drone information detected by the patrol drone within its detection range, and the detection results can be visually represented in the form of a matrix, and the matrix can be used as the state space.
[0124] For the state space
[0125] i S which is used to represent the discrete space of the local position of the drone i , the positions of friendly drones, and the positions of attacking drones. In the embodiments of the present application, a double matrix is constructed to represent the state space of the drone i at the t th scheduling cycle . The first matrix is used to observe the positions of friendly drones, and the second matrix is used to observe the positions of attacking drones.
[0126] It should be noted that considering that the detection area of the drone is fan-shaped, not all elements in a matrix are within the detection range of the drone. In the embodiments of the present application, the matrix elements outside the detection range are set to -1, and the matrix elements within the detection range are set to 1 or 0, where 0 indicates that no drone is detected and 1 indicates that a drone is detected.
[0127] Figure 3 Figure 3 shows a schematic diagram of the state space in some embodiments. As shown, two initial matrices are set in the initial state, representing the friendly drone matrix and the attacking drone matrix respectively. After determining the patrol range of the patrol drone, the initial matrix is adjusted to determine the matrix elements corresponding to the detection area. Among them, the matrix elements outside the detection range, the drones not detected within the detection range, and the drones detected within the detection range are represented in three forms. After obtaining the detection results of the patrol drone, the matrix elements corresponding to the detection area in the matrix are assigned values accordingly, so as to obtain the state space of the patrol drone in the current cycle.
[0128] The matrix of the state space can be expressed as:
[0129]
[0130] Wherein, the matrix element of the friendly aircraft matrix of the UAV i at the k th row and the j th column. When this position is not within the detection range of the UAV, the matrix element is -1. If it is within the detection range and there is a friendly aircraft, the matrix element is 1. If there is no friendly aircraft, the matrix element is 0.
[0131] the matrix element of the attacking UAV matrix of the UAV i at the k th row and the j th column. When this position is not within the detection range of the UAV, the matrix element is -1. If it is within the detection range and there is an attacking UAV, the matrix element is 1. If there is no attacking UAV, the matrix element is 0.
[0132] The detection result can be input into the decision plan update model, and the decision plan update model can convert the detection result into a state space.
[0133] The action space is used to represent the flight actions of the patrol UAV to execute tasks, and the action space is the decision plan of the UAV. In the embodiments of the present application, it is set that the decision plan calculated in the previous scheduling period is used as the action space in the current period.
[0134] The defense actions in the decision plan can be expressed as:
[0135]
[0136] The reward return is used to characterize the execution effect of the defense task of the patrol UAV in the current scheduling period. In the embodiments of the present application, in order for the UAV to gradually learn effective strategies in complex environments, and at the same time to make up for the deficiency that it is difficult to accurately design the reward function in the DCA problem, it is set that the reward return at least includes the boundary grid reward obtained based on differential game. By using differential game for reward shaping, it is ensured that the reward mechanism can accurately reflect the goal achievement situation of both sides of the game (i.e., the defense UAV and the attacking UAV).
[0137] It should be noted that the boundary grid is an important concept in the theory of qualitative differential games and is widely used in fields such as pursuit-evasion and military confrontation. The boundary grid divides the confrontation area into an escape area and a pursuit area, and satisfies the following conclusion: for the attacking UAV target in the escape area, no matter what maneuvering strategy the defending party adopts, this attacking UAV target must have at least one strategy to successfully reach the destination; for the attacking UAV target in the pursuit area, no matter what penetration strategy the attacking UAV adopts, the defending party must have a maneuvering strategy to successfully intercept the attacking UAV.
[0138] In addition, the reward returns can be set to include event-triggered rewards (also called the first reward return) and time-continuous rewards (also called the second reward return).
[0139] Event-triggered rewards: Such rewards are obtained only once when the corresponding event is triggered. This type of reward is a goal-oriented reward and punishment item, that is, the capture and escape rewards. This reward represents the reward value generated when our UAV successfully captures the attacking UAV and the attacking UAV successfully enters our defense position. This type of reward value is relatively large. This reward mainly enables our UAV to learn the mission objectives, that is, to minimize the number of the enemy entering our defense position, and to learn to capture the attacking UAV.
[0140] Furthermore, the above-mentioned corresponding event can also be associated with the relevant concepts of the escape area and the pursuit area of the boundary grid reward mentioned above, and the reward upgrade method for the event-triggered reward is designed. This is not introduced in detail here. For specific details, please refer to the construction content of the first reward return in the following text.
[0141] Time-continuous rewards: Such rewards are continuously calculated and accumulated during the confrontation process. Time-step penalty, each UAV on the field will deduct a small reward value at each time step. This reward enables the UAV to capture the attacking UAV as soon as possible on the one hand, and on the other hand, it can limit the number of UAVs on the field through this reward.
[0142] In the embodiment of the present application, the benefit function is set as:
[0143]
[0144] where argmaxf( x ) represents the value of the variable that makes the objective function f( x ) take the maximum value x ; θ represents the flight direction; R ( t ) represents the reward return in the t th cycle; γ represents the discount factor, γ t represents the of the discount factort power
[0145] It should be noted that the process of the UAV performing air defense has relatively high requirements for the real-time update of the decision-making plan. Therefore, it can be set that the patrol UAV updates its own decision-making plan by itself. An update unit can be preset in the UAV, and a decision-making plan update model is stored in the update unit. After the patrol UAV executes the decision-making plan for one scheduling cycle, it can directly input the detection result into the decision-making plan update model in the UAV itself, so as to update and obtain the decision-making plan for the next scheduling cycle by using the decision-making plan update model.
[0146] In step S5, solve the decision-making plan update model and update the decision-making plan of the patrol UAV according to the model solution.
[0147] The model solution includes the following steps:
[0148] S501. Obtain the UAV state information of the attacking UAV according to the state space.
[0149] Among them, the UAV state information refers to the attack result of the patrol UAV on the detected attacking UAV, including the situation of the attacking UAV defeated by the patrol UAV and the situation of the attacking UAV that breaks through the defense position by the attacked UAV.
[0150] Specifically, the UAV state information may include the speed ratio between the patrol UAV and the attacking UAV.
[0151] Furthermore, the UAV state information may also include the number of attacking UAVs defeated by the patrol UAV, which is called the number of the first attacking UAVs in the embodiment of the present application. The UAV state information also includes the number of attacking UAVs that enter the defense position of the defense UAV formation, which is called the number of the second attacking UAVs in the embodiment of the present application.
[0152] S502. Calculate the reward return including at least the fence reward according to the UAV state information.
[0153] Referring to the boundary fence construction theory in the existing literature (Rui Yan, Zongying Shi, and Yisheng Zhong, "Reach-Avoid Games With Two Defenders and One Attacker: An Analytical Approach," IEEE Transactions on Cybernetics, Volume: 49, Issue: 3, March 2019), under the condition of knowing the speed ratio of both sides and the position of the defense position of the drone formation of our side's defense, the boundary fence of each patrol drone can be quickly solved, so as to obtain the escape area and the pursuit area. Based on this, a boundary fence is constructed for each patrol drone, and the action strategy of the patrol drone is guided according to the position of the attacking drone in the pursuit area and the escape area.
[0154] Figure 4 The schematic diagram of the boundary fence construction in some embodiments is shown. Taking the patrol drone on the leftmost side in the figure as an example: The dotted part is its boundary fence, the left side of the dotted line is its escape area, and the right side of the dotted line is its pursuit area.
[0155] On this basis, calculating the reward return including at least the boundary fence reward according to the drone state information includes the following steps:
[0156] Based on the speed ratio and the position of the defense position of the drone formation of the defense side, differential game is used to construct the boundary fence of the patrol drone, so as to construct the pursuit area and the escape area of the patrol drone:
[0157]
[0158] Among them, represents the pursuit area of the patrol drone i ; represents the escape area of the patrol drone i ; Barrier represents the relevant formula for constructing the boundary fence in the above-mentioned literature "Reach-Avoid Games With Two Defenders and One Attacker: An Analytical Approach", which will not be elaborated here.
[0159] When at least one attacking drone is observed by the patrol drone i , two sets are defined:
[0160] represents being located in the attacking drones within; represents being located in All attacking drones within
[0161] Obtain fence rewards, including:
[0162]
[0163] Subscript C CZ and C EZ respectively represent all attacking drones located in the pursuit area and the escape area of the patrol drone; represents the reduction in distance between the patrol drone and all attacking drones in the pursuit area within a single cycle among all attacking drones; represents the reduction in distance between the patrol drone and all attacking drones in the escape area within a single cycle, that is, the reduction in distance between the patrol drone i and the set among all attacking drones.
[0164] Calculate the reward return including at least the fence reward according to the drone status information, and also include:
[0165] Obtain the first reward return according to the drone status information, including:
[0166]
[0167] wherein, R target represents the first reward return; m 1 represents the number of the first attacking drones, m 2 represents the number of the second attacking drones.
[0168] represents the reward score for defeating an attacking drone j If the attacking drone j is defeated within the escape area of the patrol drone after the patrol drone calls for friendly aircraft, the reward score is upgraded to a preset multiple of the first basic reward score, otherwise it remains the first basic reward score; that is, when the attacking drone j ∈ , after the output z of the patrol drone i becomes 1, if the patrol drone i defeats the attacking drone j , then the reward is upgraded, otherwise the original reward score is maintained.
[0169] Exemplarily, the above first basic reward score is taken as 1, and the preset multiple is taken as a positive integer k .
[0170] Indicates the attacking UAV j The reward score for the attacking UAV to enter the defense position. If the attacking UAV j enters the defense position within the pursuit area of the patrol UAV after the patrol UAV calls for friendly aircraft, the reward score is upgraded to a preset multiple of the second basic reward score; otherwise, it remains the second basic reward score. That is, when the attacking UAV j ∈ , and after the output z of the patrol UAV i becomes 1, if the patrol UAV i defeats the attacking UAV j , then the reward is upgraded; otherwise, the original reward score is maintained.
[0171] Exemplarily, the above-mentioned second basic reward score is taken as -1, and the preset multiple is taken as a positive integer k .
[0172] It is not difficult to understand that in the embodiment of the present application, when there is an attacking UAV in the escape area of the patrol UAV i , it is hoped that the patrol UAV i can call for support at this time. However, for this situation, the action only needs to be done once. For example, in the action learning of a robot entering a door, it is necessary to press the doorbell first and then enter the door. Not pressing the doorbell or continuously pressing the doorbell is not the optimal result. Therefore, based on the escape area and pursuit area in the boundary grid reward, the embodiment of the present application designs the above-mentioned reward upgrade method for event-triggered rewards to make the original rewards and punishments more after the patrol UAV executes a specific action, further ensuring that the provided reward mechanism can accurately reflect the actual interaction situation and goal achievement situation of both sides of the game.
[0173] Obtain the second reward return, including:
[0174]
[0175] Among them, R time represents the second reward return; m 3 represents the number of patrol UAVs.
[0176] Calculate the reward return, including:
[0177]
[0178] Among them, R represents the reward return.
[0179] S503. Solve the revenue function according to the reward return. The reward return value can be substituted into the revenue function to obtain the model solution that maximizes the revenue function.
[0180] Among them, the model solution includes the probability distribution of multiple decision-making plans, that is, the probability situation of the decision-making plan of the patrol drone in the next cycle.
[0181] The decision-making plan of the patrol drone can be updated according to the model solution, including:
[0182] Obtain the decision-making plan with the highest probability according to the model solution, and update the decision-making plan with the highest probability as the decision-making plan of the patrol drone in the next cycle.
[0183] If the patrol drone updates its decision-making plan by itself, the updated decision-making plan can be returned to the ground command center.
[0184] In step S6, update the drone formation defense strategy according to the updated decision-making plan, and control the defense-side drone formation to execute the updated drone formation defense strategy.
[0185] Among them, updating the drone formation defense strategy includes the following steps:
[0186] S601. Obtain the number of the third drones whose defense action is to call for friendly aircraft according to the updated decision-making plan, and obtain the retreating drones whose defense action is to retreat by itself.
[0187] It should be noted that when there is one patrol drone calling for friendly aircraft, the ground command center needs to dispatch a new drone to perform flight patrol as a patrol drone. Therefore, the number of drones calling for friendly aircraft can be counted to dispatch the corresponding number of drones for support.
[0188] S602. Obtain the candidate drones corresponding to the number of the third drones in the defense-side drone formation, and update the decision-making plans of the candidate drones to set the defense actions of the candidate drones as aerial patrol.
[0189] Candidate drones can be selected from the drones currently staying in the defense position to perform flight patrol tasks.
[0190] S603. Set the candidate drones as patrol drones, and cancel the setting of the retreating drones as patrol drones.
[0191] S604. Statistically analyze the updated decision-making plans of multiple drones in the defense-side drone formation to generate the drone formation defense strategy for the next cycle.
[0192] In the next cycle, the defense-side drone formation can be controlled to execute the updated drone formation defense strategy.
[0193] In some embodiments, considering that there may be some emergencies during the patrol of the patrol drone, such as insufficient power or fuel, etc., a direct retreat is required.
[0194] For this reason, after obtaining the patrol drone in the patrol state, the situation of each patrol drone can be detected to determine whether a retreat is needed.
[0195] In the embodiments of the present application, a patrol cycle is set. When the continuous patrol duration of a drone reaches the patrol cycle and no attacking drones are detected during this period, the drone is ordered to retreat.
[0196] For this reason, the patrol drone can be controlled to start timing after taking off and performing patrol, and the timing duration is detected.
[0197] If within the preset patrol cycle, the patrol drone does not detect an attacking drone, the decision-making scheme of the patrol drone is updated to set the defense action of the patrol drone to retreat by itself.
[0198] If within the preset patrol cycle, the patrol drone detects an attacking drone, the timing duration is updated to 0, and the patrol drone is controlled to start timing again and continue to perform patrol.
[0199] In some embodiments, if the attacking drone poses a greater threat to the defense position, it may be necessary for the drone to directly call for friendly aircraft.
[0200] In the embodiments of the application, it is set that: after the patrol drone has detected for a scheduling cycle, the position degree of the current attacking drone can be obtained according to the detection result. If the threat degree is greater, friendly aircraft are directly called. If the threat degree is smaller, the decision-making scheme is updated through the decision-making scheme update model.
[0201] For this reason, after obtaining the detection result of the patrol drone, the patrol threat value of the patrol drone can be obtained. In the embodiments of the present application, two types of patrol threats are set: strategic threat and position threat.
[0202] The obtaining of the patrol threat value includes the following steps:
[0203] The target attacking drone detected by the patrol drone is obtained according to the detection result of the patrol drone.
[0204] Obtain the action strategy of the target attacking drone. The action strategy of the target attacking drone includes an advance strategy, a retreat strategy, and other strategies. The advance strategy is the behavior of the target attacking drone advancing toward the defense position of the defending drone formation, the retreat strategy is the behavior of the target attacking drone retreating, and the other strategies are strategic behaviors other than the advance strategy and the retreat strategy;
[0205] According to the action strategy, the strategy threat value is obtained. Among them, the threat value of the advance strategy can be set to 50, the threat value of the retracement strategy can be set to 10, and the threat value of other strategies can be set to 30.
[0206] Get the location threat value, including:
[0207]
[0208] Among them, W distance indicates the location threat value, d max indicates the maximum patrol distance of the drones in the defending drone formation, Indicates the target attacking drone j * Distance from defensive positions;
[0209] Calculate the patrol threat value of the patrol drone according to the strategy threat value and the location threat value.
[0210] For each target attacking drone, the strategic threat value and position threat value are calculated to obtain the threat value of each target attacking drone.
[0211] The patrol drone may detect multiple target attack drones, and the threat values of all target attack drones are summed up to obtain the patrol threat value of the patrol drone.
[0212] In some embodiments,
[0213] If the patrol threat value is greater than or equal to the preset threat value threshold, the decision plan of the patrol drone is updated to set the patrol drone's defense action to call a friendly aircraft. The threat value threshold can be set to 100.
[0214] If the patrol threat value is less than the preset threat value threshold, the step of generating a decision solution and updating the model according to the detection result is executed.
[0215] In some embodiments, when a patrol drone detects multiple attacking drones, it can attack the attacking drones in order of threat value from high to low.
[0216] Embodiment 2:
[0217] The embodiment of the present invention further provides an intelligent decision-making system for unmanned aerial vehicles based on reinforcement learning, and the system includes:
[0218] An information acquisition module, configured to acquire the information of the defense unmanned aerial vehicle formation and the defense strategy of the unmanned aerial vehicle formation; the defense unmanned aerial vehicle formation includes multiple unmanned aerial vehicles; the defense strategy of the unmanned aerial vehicle formation is used to indicate the decision-making plan of the multiple unmanned aerial vehicles, and the decision-making plan includes defense actions and flight directions; the defense actions include aerial patrol, calling for friendly aircraft, and the retreat of the local aircraft;
[0219] A strategy execution module, configured to control the defense unmanned aerial vehicle formation to execute the defense strategy of the unmanned aerial vehicle formation;
[0220] A result detection module, configured to acquire the patrolling unmanned aerial vehicles in the patrolling state and acquire the detection results of the patrolling unmanned aerial vehicles;
[0221] A model generation module, configured to generate a decision-making plan update model according to the detection results; the decision-making plan update model includes a state space, a decision-making plan, a reward return, and a revenue function; the state space is generated by the detection results; the reward return at least includes a boundary reward obtained based on differential game;
[0222] A plan update module, configured to solve the decision-making plan update model and update the decision-making plan of the patrolling unmanned aerial vehicle according to the model solution;
[0223] A strategy update module, configured to update the defense strategy of the unmanned aerial vehicle formation according to the updated decision-making plan and control the defense unmanned aerial vehicle formation to execute the updated defense strategy of the unmanned aerial vehicle formation.
[0224] Embodiment 3:
[0225] The embodiment of the present invention further provides a storage medium, which stores a computer program for intelligent decision-making of unmanned aerial vehicles based on reinforcement learning, wherein the computer program enables a computer to execute the intelligent decision-making method for unmanned aerial vehicles based on reinforcement learning as described in Embodiment 1.
[0226] It can be understood that the intelligent decision-making system for unmanned aerial vehicles based on reinforcement learning, the storage medium provided by the embodiment of the present invention correspond to the intelligent decision-making method for unmanned aerial vehicles. For the explanations, examples, beneficial effects, etc. of the relevant content, reference can be made to the corresponding content in the intelligent decision-making method for unmanned aerial vehicles based on reinforcement learning, and details are not described herein again.
[0227] In summary, compared with the prior art, the following beneficial effects are achieved:
[0228] 1. The embodiments of the present invention combine differential game with reinforcement learning to overcome the problem of solving the boundary grid in multi-aircraft cooperation in differential games. Specifically, a boundary grid reward obtained based on differential game is designed to ensure that the reward mechanism can accurately reflect the achievement of the goals of both sides of the game. Through continuous iterative optimization of reinforcement learning, the unmanned aerial vehicle (UAV) can gradually learn an effective UAV formation defense strategy in a complex environment.
[0229] 2. The embodiments of the present invention can obtain the information of the UAV formation of the defense side and the UAV formation defense strategy. Among them, the UAV formation defense strategy is used to indicate the decision-making scheme of multiple UAVs, and the decision-making scheme includes defense actions and flight directions; the defense actions include aerial patrol, calling for friendly aircraft, and the retreat of the own aircraft. Control the UAV formation of the defense side to execute the UAV formation defense strategy, and obtain the patrolling UAVs in the patrol state, and obtain the detection results of the patrolling UAVs. Generate a decision-making scheme update model according to the detection results and solve it, and update the decision-making scheme of the patrolling UAVs according to the solution of the model. Update the UAV formation defense strategy according to the updated decision-making scheme, and control the UAV formation of the defense side to execute the updated UAV formation defense strategy. Through the detection results of the UAVs, the decision-making scheme of the UAVs can be updated, so as to adjust the scale of the UAV formation in real time when performing tasks, and improve the defense effect.
[0230] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0231] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A UAV intelligent decision-making method based on reinforcement learning, characterized in that: include: Obtain the defender’s drone formation information and drone formation defense strategy; The defending drone formation includes multiple drones; the drone formation defense strategy is used to indicate the decision-making plan of the multiple drones, and the decision-making plan includes defense actions and flight directions; the defense actions include air patrol, calling friendly aircraft and the own aircraft retreating; Controlling the defending drone formation to execute the drone formation defense strategy; Acquire a patrol drone in a patrol state, and acquire a detection result of the patrol drone; A decision-making scheme update model is generated according to the detection results; the decision-making scheme update model includes a state space, a decision scheme, a reward return and a benefit function; the state space is generated by the detection results; the reward return at least includes a boundary fence reward obtained based on a differential game; Solving the decision-making scheme update model, and updating the decision-making scheme of the patrol drone according to the model solution; Updating the drone formation defense strategy according to the updated decision plan, and controlling the defending drone formation to execute the updated drone formation defense strategy; The method for intelligent decision-making of unmanned aerial vehicles solves the decision-making scheme update model, including: Acquire drone state information of the attacking drone according to the state space; Calculate a reward return including at least the boundary fence reward according to the drone state information; Solving the profit function according to the reward return; The drone state information of the drone intelligent decision-making method includes the speed ratio of the patrol drone to the attacking drone; The process of obtaining the boundary fence reward includes: Based on the speed ratio and the position of the defense position of the defense drone formation, differential game is used to construct the boundary fence of the patrol drone to construct the pursuit zone and escape zone of the patrol drone; Get Boundary Fence rewards, including: Among them, the subscript C CZ , C EZ Respectively represent all attacking drones located in the pursuit zone and escape zone of the patrol drone; It represents the reduction in the distance between the patrol drone and all attacking drones in the pursuit area in a single cycle; It represents the reduction in the distance between the patrol drone and all attacking drones in the escape zone in a single cycle; The drone status information also includes the number of first attacking drones defeated by the patrol drone and the number of second attacking drones that entered the defense position of the defending drone formation; Calculating a reward return including at least the boundary fence reward according to the drone state information also includes: Obtaining a first reward return according to the drone state information includes: in, R target Indicates the first reward return; m 1 indicates the number of drones of the first attacking party, m 2 represents the number of drones of the second attacking party; Defeated the attacking drone j Bonus score, if the attacking drone j If a patrol drone is defeated in its escape zone after it calls a friendly drone, the bonus score is upgraded to the preset multiple of the first basic bonus score, otherwise it remains at the first basic bonus score; Indicates the attacking drone j Bonus points for entering a defensive position, if the attacking drone j If you enter a defensive position in a patrol drone's pursuit zone after the patrol drone calls a friendly drone, the bonus score will be upgraded to the preset multiple of the second basic bonus score, otherwise it will remain at the second basic bonus score; Get the second reward, including: in, R time Indicates the second reward return; m 3 indicates the number of patrol drones; Calculate reward returns, including: in, R Indicates reward return.
2. The intelligent decision-making method for unmanned aerial vehicles according to claim 1, characterized in that: The profit function is: Among them, argmaxf( x ) means that the objective function f( x ) The variable with the maximum value x The value of θ Indicates the flight direction; R ( t ) indicates the t Reward return for each cycle; γ represents the discount factor, γ t The discount factor t Power.
3. The intelligent decision-making method for unmanned aerial vehicles according to claim 1, characterized in that: After executing the acquisition of the patrol drone in patrol state, it also includes: Control patrol drones to time and detect the timing duration; If the patrol drone does not detect the attacking drone within the preset patrol period, the decision plan of the patrol drone is updated to set the patrol drone's defense action to retreat; If the patrol drone detects the attacking drone within the preset patrol cycle, the timing duration is updated to 0, and the patrol drone is controlled to perform timing.
4. The UAV intelligent decision-making method according to claim 1, characterized in that: After obtaining the detection result of the patrol drone, the method further includes: Obtaining the patrol threat value of the patrol drone; If the patrol threat value is greater than or equal to a preset threat value threshold, updating the decision plan of the patrol drone to set the defense action of the patrol drone to call a friendly aircraft; If the patrol threat value is less than a preset threat value threshold, the step of generating a decision solution and updating the model according to the detection result is performed.
5. The intelligent decision-making method for unmanned aerial vehicles according to claim 4, characterized in that: Obtaining the patrol threat value of the patrol drone, including: Acquire the target attacking drone detected by the patrol drone according to the detection result of the patrol drone; Obtaining the action strategy of the target attacking drone; the action strategy includes an advance strategy, a retreat strategy, and other strategies; the advance strategy is the behavior of the target attacking drone advancing toward the defense position of the defending drone formation, the retreat strategy is the behavior of the target attacking drone retreating, and the other strategies are strategy behaviors other than the advance strategy and the retreat strategy; Obtaining a strategy threat value according to the action strategy; Get the location threat value, including: in, W distance Indicates the position threat value, d max Indicates the maximum patrol distance of the drones in the defending drone formation. Indicates the target attacking drone j * distance from defensive positions; The patrol threat value of the patrol drone is calculated according to the strategy threat value and the position threat value.
6. The intelligent decision-making method for unmanned aerial vehicles according to claim 1, characterized in that: The model solution includes probability distributions of multiple decision-making options; The decision plan of the patrol drone is updated according to the model solution, including: Obtaining a decision plan with the highest probability according to the model solution, and updating it as the decision plan of the patrol drone in the next cycle; The drone formation defense strategy is updated according to the updated decision plan, including: According to the updated decision plan, the number of third drones whose defensive action is to call friendly drones is obtained, and the number of retreating drones whose defensive action is to retreat. Acquire a number of standby drones corresponding to the third number of drones in the defending drone formation, and update the decision plan of the standby drones to set the defense action of the standby drones to air patrol; Setting the standby drone as a patrol drone and canceling the setting of the retreating drone as a patrol drone; Statistics are collected on the updated decision plans of multiple drones in the defending drone formation to generate the drone formation defense strategy for the next cycle.
7. A UAV intelligent decision-making system based on reinforcement learning, characterized in that: For executing the UAV intelligent decision-making method according to claim 1, the system comprises: The information acquisition module is configured to acquire the defender's drone formation information and the drone formation defense strategy; the defender's drone formation includes multiple drones; the drone formation defense strategy is used to indicate the decision-making plan of the multiple drones, the decision-making plan includes defense actions and flight directions; the defense actions include air patrol, calling friendly aircraft and the aircraft's own retreat; A strategy execution module, configured to control the defending drone formation to execute the drone formation defense strategy; A result detection module is configured to obtain a patrol drone in a patrol state and obtain a detection result of the patrol drone; A model generation module is configured to generate a decision solution update model according to the detection result; the decision solution update model includes a state space, a decision solution, a reward return and a benefit function; the state space is generated by the detection result; the reward return at least includes a boundary fence reward obtained based on a differential game; A solution updating module is configured to solve the decision solution updating model and update the decision solution of the patrol drone according to the model solution; The strategy update module is configured to update the drone formation defense strategy according to the updated decision plan, and control the defending drone formation to execute the updated drone formation defense strategy.
8. A storage medium, characterized in that: It stores a computer program for unmanned aerial vehicle intelligent decision-making based on reinforcement learning, wherein the computer program enables a computer to execute the unmanned aerial vehicle intelligent decision-making method based on reinforcement learning as described in any one of claims 1 to 6.