Reinforcement learning methods and systems for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs)

By using reinforcement learning to make adaptive intelligent decisions for multiple drones, the size of the drone formation is dynamically adjusted, which solves the problem of poor defense effectiveness of the defending drones and achieves effective interception of attacking drones and security of the defense position.

CN119690101BActive Publication Date: 2025-10-28HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411788978.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-10-28
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

When the defending drones are used in a preset number, the defense effect is poor, and it is difficult to cope with the uncertainty of the number of attacking drones and unexpected situations.

Method used

A reinforcement learning method for multi-UAV adaptive intelligent decision-making is adopted. By acquiring UAV formation information and detection results, a decision scheme update model is generated, which solves and updates the UAV formation defense strategy, and dynamically schedules the UAV scale to deal with attacks.

Benefits of technology

It improved the defensive effectiveness, adjusted the drone formation size in real time, enhanced the ability to intercept attacking drones, and ensured the security of the defensive positions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119690101B_ABST
    Figure CN119690101B_ABST
Patent Text Reader

Abstract

This invention provides a reinforcement learning method and system for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs). The invention acquires a UAV formation defense strategy to instruct decision-making schemes for multiple UAVs. The decision-making scheme includes defensive actions and flight directions; defensive actions include aerial patrol, calling on friendly aircraft, and self-retreat. The system controls the defending UAV formation to execute the UAV formation defense strategy and acquires the detection results of patrolling UAVs. Based on the detection results, a decision-making scheme update model is generated and solved, and the decision-making schemes of the patrolling UAVs are updated based on the model solution. The updated UAV formation defense strategy is then updated, and the defending UAV formation is controlled to execute the updated strategy. By updating the UAV decision-making schemes using the UAV detection results, the scale of the UAV formation during mission execution can be scheduled in real time, improving the defense effectiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and more specifically to a reinforcement learning method and system for adaptive intelligent decision-making among multiple UAVs. Background Technology

[0002] With the increasing intelligence and enhanced performance of unmanned aerial vehicles (UAVs), utilizing multi-UAV swarms for air combat missions has become a hot research topic in the field. Attackers can use UAVs, with their wide coverage and powerful firepower, to pose a serious threat to the core assets of defenders. Therefore, counter-air operations have become an indispensable key element in military strategic planning, encompassing two core strategies: Offensive Counter-Air Operations (OCA) and Defensive Counter-Air Operations (DCA). The primary objective of OCA is to weaken the attacking force near its source by destroying, disrupting, or suppressing it; while DCA focuses on intercepting aerial targets, ensuring the safety and stability of the defender's critical assets by reducing their threat level.

[0003] In the complex scenarios of DCA, the defending drones must intercept the constantly attacking drones to the greatest extent possible, while adopting continuous and flexible maneuvering strategies to cope with the rapid changes in the battlefield environment.

[0004] However, the arrival of attacking drones is often unpredictable, and they gradually approach the defender's predetermined targets over time. Furthermore, when the number of attacking drones is unknown, using a predetermined number of drones may result in wasted or insufficient resources, making it difficult to respond to unforeseen circumstances and leading to poor defensive effectiveness. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a reinforcement learning method and system for multi-UAV adaptive intelligent decision-making, which solves the technical problem that the defense effect is poor when the defending UAV uses a preset number of UAVs for defense in existing technologies.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions:

[0009] This invention addresses its technical problem by providing a reinforcement learning method for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs). The reinforcement learning method is executed by a computer and includes the following steps:

[0010] A reinforcement learning method for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs), wherein the reinforcement learning method is executed by a computer, is characterized by comprising the following steps:

[0011] Acquire information on the defending drone formation and its defense strategy; the defending drone formation includes multiple drones; the drone formation defense strategy is used to instruct the decision-making schemes of the multiple drones, the decision-making schemes include defensive actions and flight directions; the defensive actions include aerial patrol, calling on friendly aircraft, and self-retreat;

[0012] Control the defending drone formation to execute the drone formation defense strategy;

[0013] Acquire patrol drones that are in patrol mode, and acquire the detection results of the patrol drones;

[0014] A decision scheme update model is generated based on the detection results; the decision scheme update model includes a state space, decision schemes, reward rewards, and a payoff function; the state space is generated from the detection results.

[0015] Solve the decision-making scheme update model and update the decision-making scheme of the patrol drone based on the model solution;

[0016] The drone formation defense strategy is updated according to the updated decision scheme, and the defending drone formation is controlled to execute the updated drone formation defense strategy.

[0017] Preferably, solving the decision update model includes:

[0018] The drone state information of the attacker's drone is obtained based on the state space;

[0019] Calculate the reward based on the drone status information;

[0020] The revenue function is solved based on the stated reward.

[0021] Preferably, the drone status information includes the number of first attacking drones defeated by the patrol drone and the number of second attacking drones that entered the defensive drone formation's defensive position. The reward calculation based on the drone status information includes:

[0022] The first reward is obtained based on the drone's status information, including:

[0023] R target =m1×f1+m2×f2

[0024] in,

[0025] R target This indicates the first reward / reward.

[0026] m1 represents the number of drones of the first attacking party, and m2 represents the number of drones of the second attacking party;

[0027] f1 represents the bonus points for defeating the attacking drone, and f2 represents the bonus points for the attacking drone entering the defensive position.

[0028] To receive the second reward, you will receive:

[0029] R time = -0.01×m3

[0030] in,

[0031] R time This indicates the second reward; m3 represents the number of patrol drones.

[0032] Calculate reward returns, including:

[0033] R = R target +R time

[0034] in,

[0035] R represents the reward or incentive.

[0036] Preferably, the payoff function is:

[0037]

[0038] in,

[0039] R (t) This represents the reward for the t-th period;

[0040] γ represents the discount factor, γ t This represents the discount factor raised to the power of t.

[0041] Preferably, after acquiring a patrol drone in patrol mode, the process further includes:

[0042] Control the patrol drone to keep track of time and monitor the duration of the timer;

[0043] If the patrol drone does not detect the attacking drone within the preset patrol period, the patrol drone's decision-making scheme is updated to set the patrol drone's defensive action to retreat.

[0044] If the patrol drone detects an attacking drone within the preset patrol period, the timer duration is updated to 0, and the patrol drone is controlled to keep a timer running.

[0045] Preferably, after acquiring the detection results of the patrol drone, the method further includes:

[0046] Obtain the patrol threat value of the patrol drone;

[0047] If the patrol threat value is greater than or equal to a preset threat value threshold, the decision scheme of the patrol drone is updated to set the patrol drone's defensive action to call a friendly drone.

[0048] If the patrol threat value is less than a preset threat value threshold, then the step of generating a decision scheme and updating the model based on the detection results is executed.

[0049] Preferably, obtaining the patrol threat value of the patrol drone includes:

[0050] The target attacking drone detected by the patrol drone is obtained based on the detection results of the patrol drone;

[0051] The action strategy of the target attacking drone is obtained; the action strategy includes an advance strategy, a retreat strategy, and other strategies; the advance strategy is the action of the target attacking drone advancing towards the defensive position of the defending drone formation; the retreat strategy is the action of the target attacking drone retreating; and the other strategies are strategies other than the advance strategy and the retreat strategy.

[0052] Obtain the strategy threat value based on the described action strategy;

[0053] Obtain location threat value, including:

[0054]

[0055] in,

[0056] W distance Indicates the location threat value, d max This indicates the maximum patrol distance of the defending drone formation. This indicates the distance between the attacking drone i and the defensive position;

[0057] The patrol threat value of the patrol drone is calculated based on the strategy threat value and the location threat value.

[0058] Preferably, the model solution includes the probability distributions of multiple decision options;

[0059] The decision scheme for the patrol drone is updated based on the model solution, including:

[0060] The decision scheme with the highest probability is obtained from the model solution and updated as the decision scheme of the patrol drone in the next cycle;

[0061] The drone swarm defense strategy is updated according to the updated decision-making scheme, including:

[0062] The number of third drones whose defensive action is to call a friendly drone is obtained based on the updated decision scheme, and the number of retreating drones whose defensive action is to retreat.

[0063] Obtain the candidate drones corresponding to the number of the third drones in the defending drone formation, and update the decision scheme of the candidate drones so as to set the defensive action of the candidate drones to air patrol.

[0064] Set the standby drone as a patrol drone, and cancel the retreat drone from being set as a patrol drone;

[0065] The system analyzes the updated decision-making schemes of multiple drones in the defending drone formation to generate the drone formation defense strategy for the next cycle.

[0066] This invention provides a reinforcement learning system for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs) to solve its technical problem. The system includes:

[0067] The information acquisition module is configured to acquire information about the defending drone formation and its defense strategy. The defending drone formation includes multiple drones. The drone formation defense strategy is used to instruct the decision-making schemes of the multiple drones, including defensive actions and flight directions. The defensive actions include aerial patrol, calling on friendly aircraft, and self-retreat.

[0068] The execution module is configured to control the defending drone formation to execute the drone formation defense strategy;

[0069] The detection module is configured to acquire the detection results of the patrol drone in patrol mode;

[0070] The model generation module is configured to generate a decision scheme update model based on the detection results; the decision scheme update model includes a state space, decision schemes, reward rewards, and a payoff function; the state space is generated from the detection results.

[0071] The decision scheme update module is configured to solve the decision scheme update model and update the decision scheme of the patrol drone based on the model solution.

[0072] The defense strategy update module is configured to update the UAV formation defense strategy according to the updated decision scheme, and control the defending UAV formation to execute the updated UAV formation defense strategy.

[0073] The present invention provides a computer-readable storage medium that solves its technical problem by storing a computer program for reinforcement learning of multi-UAV adaptive intelligent decision-making, wherein the computer program causes a computer to execute the reinforcement learning method for multi-UAV adaptive intelligent decision-making as described above.

[0074] (3) Beneficial effects

[0075] This invention provides a reinforcement learning method and system for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs). Compared with existing technologies, it has the following advantages:

[0076] This invention can acquire information about the defending drone formation and its defense strategy. The defense strategy instructs multiple drones on decision-making methods, including defensive actions and flight directions. Defensive actions include aerial patrol, calling on friendly drones, and drone retreat. The system controls the defending drone formation to execute the defense strategy and acquires information about patrol drones in patrol mode, obtaining their detection results. Based on the detection results, a decision-making update model is generated and solved, and the decision-making strategies for the patrol drones are updated according to the model solution. The updated decision-making strategy is then used to update the defense strategy, and the defending drone formation is controlled to execute the updated strategy. By updating the drone decision-making strategies based on their detection results, the scale of the drone formation during mission execution can be scheduled in real time, improving the defense effectiveness. Attached Figure Description

[0077] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0078] Figure 1 This is a schematic diagram of a scenario for the reinforcement learning method for adaptive intelligent decision-making of multiple unmanned aerial vehicles provided in an embodiment of the present invention.

[0079] Figure 2 The diagram illustrates the attack range and detection range of the drone in some embodiments;

[0080] Figure 3 A schematic diagram of the state space is shown in some embodiments. Detailed Implementation

[0081] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0082] This application provides a reinforcement learning method and system for multi-UAV adaptive intelligent decision-making, which solves the problem of poor defense effect of UAV formations in the prior art and improves the quality of UAV formations performing defense.

[0083] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0084] Multiple drone swarms can perform air combat missions, allowing the attacker to use drone swarms to attack the defender's defensive positions. For the defender, this allows for counter-air operations, preventing the attacker's drone swarms from entering their defensive positions.

[0085] The defending side can employ two strategies: Offensive Counter-Air Operations (OCA) and Defensive Counter-Air Operations (DCA). The primary objective of OCA is for the defending side's drones to destroy, disrupt, or suppress the attacking force, maximizing the weakening of its combat effectiveness near its source. DCA, on the other hand, involves the defending side's drones intercepting the attacking drones, reducing their threat level to ensure the security and stability of the defending side's critical assets.

[0086] DCA (Distributed Combat Air Response) is a key strategy for ensuring airspace security, encompassing both active air defense and missile defense. It relies on a diverse range of asset types and system integrations, including but not limited to fighter jets, surface-to-air missiles, anti-aircraft artillery, electromagnetic warfare systems, and ballistic missile defense systems. These systems work together to destroy enemy forces or effectively weaken their offensive capabilities. In carrying out DCA missions, the fighter jets used can be either traditional manned aircraft or advanced unmanned aerial vehicles (UAVs). In particular, weaponized unmanned aerial vehicles (UCAVs) specifically designed for combat are capable of both remote control and autonomous mission execution.

[0087] When using the DCA (Distributed Calibration and Exploitation) strategy, the defender typically maintains a fixed number of drones performing the mission. For example, a drone swarm consisting of five drones would typically have all five drones take off and perform the defensive mission. In this embodiment, performing the defensive mission refers to the defender's drones patrolling the air to detect and attack the attacker's drones.

[0088] However, the arrival of attacking drones is often unpredictable, and their numbers are usually unknown. Attacking drones may advance on the defensive position in multiple waves over various time periods. Using a predetermined number of drones may result in wasted or insufficient resources for the defender, making it difficult to respond to unforeseen circumstances and leading to poor defensive effectiveness.

[0089] To address the above problems, this invention provides a reinforcement learning method for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs). This method is executed by a computer and includes the following steps:

[0090] S1. Obtain information on the defending drone formation and its defense strategy. The defending drone formation comprises multiple drones; the defense strategy instructs the drones on decision-making options, including defensive actions and flight directions; the defensive actions include aerial patrol, calling for friendly aircraft, and self-retreat.

[0091] S2. Control the defending drone formation to execute the drone formation defense strategy;

[0092] S3. Obtain the patrol drone in patrol mode and obtain the detection results of the patrol drone;

[0093] S4. Generate a decision scheme update model based on the detection results; the decision scheme update model includes a state space, decision schemes, reward rewards, and a payoff function; the state space is generated from the detection results.

[0094] S5. Solve the decision scheme update model and update the decision scheme of the patrol drone according to the model solution;

[0095] S6. Update the drone formation defense strategy according to the updated decision scheme, and control the defending drone formation to execute the updated drone formation defense strategy.

[0096] Figure 1 This is a schematic diagram illustrating a scenario of the reinforcement learning method for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs) provided in an embodiment of the present invention. The following is a detailed analysis of each step.

[0097] In step S1, information on the defending drone formation and the drone formation defense strategy are obtained.

[0098] The defending drone swarm can include multiple drones. The defending drone swarm information includes information about each drone in the swarm. All drones in the defending drone swarm can be of the same model, meaning they have identical parameters such as speed and maximum fuel capacity. Each drone can be pre-numbered for easy identification.

[0099] When the defending drone swarm performs a defense mission, it can follow a drone swarm defense strategy. This strategy can instruct the action plan for each drone, referred to as a decision scheme in this embodiment.

[0100] The decision-making scheme can be represented using a two-dimensional vector (θ, z). Here, θ represents the flight direction, i.e., the UAV's maneuver angle. The flight direction is a continuous value ranging from (0, 2π), representing the UAV's maneuver angle. z represents the defensive action, i.e., the action the UAV should perform. Defensive actions include aerial patrol, calling for friendly aircraft, and self-retreat. The UAV can perform one of these defensive actions simultaneously.

[0101] Aerial patrol refers to the aerial patrol by unmanned aerial vehicles (UAVs) to detect attacking UAVs. In this application embodiment, the attacking UAV is referred to as the attacking UAV. The UAV has detection capabilities and can detect the situation of other UAVs in a certain area ahead. It can detect both attacking and friendly UAVs.

[0102] Calling for friendly aircraft refers to a drone detecting an attacking drone that exceeds its own processing capacity, such as when there are many attacking drones, and thus calling on the ground formation of the defensive position to send friendly aircraft for reinforcement.

[0103] "Unmanned aerial vehicle (UAV) withdrawal" refers to the UAV's retreat back to its defensive position, including situations where the UAV itself needs to withdraw due to factors such as insufficient power or fuel to continue patrolling, or damage to itself.

[0104] All drones can be equipped with firepower and electronic jamming weapons, enabling them to attack enemy drones. Due to weapon limitations, drone weapons have a certain effective range, meaning their attacks are restricted to a specific area. In this embodiment, the effective range of the weapon is referred to as the attack range.

[0105] The attack range of a drone can be considered as part of its detection range, meaning the attack range is smaller than the detection range. Figure 2 The diagram illustrates the attack range and detection range of the drone in some embodiments.

[0106] When a defending drone detects an attacking drone—that is, when an attacking drone enters the defending drone's detection range—the defending drone will fly towards the attacking drone until the attacking drone enters the defending drone's attack range. At this point, the defending drone will launch an attack on the attacking drone using its weapons. Furthermore, this embodiment stipulates that when a defending drone encounters an attacking drone within its attack range, it will definitely destroy the attacking drone.

[0107] When a defending drone detects an attacking drone flying towards it, it locks onto the attacking drone and pursues it until it enters its attack range. Furthermore, in this embodiment, if the defending drone detects an attacking drone entering its defensive position, it considers the attacking drone to have completed its penetration mission, and the defending drone's defense against the attacking drone has failed. At this point, the defending drone is no longer considered and will cease pursuit. The defending drone then resumes its patrol to detect the next attacking drone.

[0108] It should be noted that, in order to cope with the variability of attacking drones, the number of drones performing patrol missions at the same time is set to be variable in this embodiment, that is, the scale of drones conducting patrols can be scheduled. For the defending drone formation, only some drones may perform patrols, while others remain on standby within the defensive position.

[0109] Meanwhile, in this embodiment, a scheduling period is pre-set, for example, 2 seconds. The defending drone formation can update its defense strategy according to the scheduling period. After determining a drone formation defense strategy, the drone formation can execute the strategy for one scheduling period, and then update the drone formation defense strategy to adjust the scale of the patrolling drones.

[0110] Before the defense mission begins, all drones in the defending drone formation are located within the defense position. At this time, an initial drone formation defense strategy can be pre-set, for example, the decision-making scheme for each drone can be randomly set.

[0111] In step S2, the defending drone formation is controlled to execute the drone formation defense strategy.

[0112] At the start of a defense mission, the defending drone swarm can execute a pre-set initial drone swarm defense strategy. The execution duration of this initial strategy is a pre-set scheduling period. After the execution duration reaches the scheduling period, the drone swarm defense strategy can be updated.

[0113] In step S3, the patrol drone in patrol mode is acquired, and the detection results of the patrol drone are obtained.

[0114] To update the drone swarm defense strategy, the mission execution status of each drone in the current scheduling period can be obtained first. In this embodiment, a drone currently flying and in patrol mode is referred to as a patrol drone.

[0115] It can identify all patrol drones within the current period and obtain the detection results for each patrol drone. The detection results refer to the drone's detection status within its detection range, including detected attacking drones. It should be noted that drones can perform detection in real time and can perform multiple detections within a single scheduling period. In this embodiment, the obtained detection results are defined as the drone's detection status at the last point in time within the scheduling period.

[0116] In step S4, a decision scheme update model is generated based on the detection results.

[0117] The decision scheme update model is used to update the decision scheme of patrol drones. The decision scheme update model can include a state space, the current drone formation defense strategy, reward reward, and payoff function.

[0118] Decision update models can use Markov games (MG) to formulate DCA problems. MG is an extension of Markov Decision Processes (MDPs) for multi-agent problems.

[0119] The MG in aerial combat is defined as a tuple.<N,S,A,R,T> Where N is a set of n UAVs; S and A are the joint state space and joint action space of the multiple UAVs, respectively, where S = (S1, S2, ..., S...). n ), A = (A1, A2, ..., A n ), S i and A i Let R be the state space and action space of drone i, respectively. The reward function for drone i is R. i Indicates that in state S i Take action A at that time i The reward obtained. The probability of entering the next state S' when taking action A in state S using the state transition function T. In this embodiment, the next state is the next scheduling cycle. Since all UAVs in the air combat environment are homogeneous, all UAVs share parameters and a shared policy π. w , representing the probability steps of the drone's selected action. The reward function aims to maximize the cumulative return.

[0120] In some embodiments, the state space is generated from the detection results. The detection results can characterize the drone information detected by the patrol drone within its detection range, and can be intuitively represented in the form of a matrix, which can serve as the state space.

[0121] For the state space S i Let be the discrete space describing the position of UAV i, the position of friendly UAVs, and the position of enemy UAVs. In this embodiment, a bimatrix is ​​constructed to represent the state space of UAV i in the t-th scheduling period. The first matrix is ​​used to observe the position of friendly aircraft, and the second matrix is ​​used to observe the position of enemy aircraft.

[0122] It should be noted that, considering that the detection area of ​​the UAV is fan-shaped, the entire matrix is ​​not within the detection range of the UAV. In this embodiment, the matrix elements outside the detection range are set to -1, and the matrix elements within the detection range are set to 1 or 0, where 0 indicates that no UAV was detected and 1 indicates that a UAV was detected.

[0123] Figure 3 A schematic diagram of the state space is shown in some embodiments. For example... Figure 3 As shown, two initial matrices are set initially, representing the friendly aircraft matrix and the enemy aircraft matrix, respectively. After determining the patrol range of the patrol drone, the initial matrices are adjusted to determine the matrix elements corresponding to the detection area. Matrix elements outside the detection range, drones not detected within the detection range, and drones detected within the detection range are represented in three forms. After obtaining the detection results of the patrol drone, the matrix elements corresponding to the detection area are assigned values ​​according to the detection results, thereby obtaining the state space of the patrol drone in the current period.

[0124] The matrix of the state space can be represented as:

[0125]

[0126]

[0127] in, The element in the k-th row and j-th column of the matrix of drone i is -1 when the position is not within the detection range of the drone, and 1 when there is a friendly drone and 0 when there is no friendly drone.

[0128] The element in the k-th row and j-th column of the enemy aircraft matrix for UAV i is -1 when the position is not within the detection range of the UAV, 1 when there is an enemy aircraft within the detection range, and 0 when there is no enemy aircraft.

[0129] The detection results can be input into the decision update model, which can then convert the detection results into a state space.

[0130] The action space is used to represent the flight actions of the patrol drone in performing its missions; the action space is essentially the drone's decision-making scheme. In this embodiment, the decision-making scheme calculated in the previous scheduling cycle is set as the action space for the current cycle.

[0131] The defensive actions in the decision-making scheme can be represented as:

[0132]

[0133] The reward is used to characterize the effectiveness of the patrol drone's defense mission execution within the current scheduling cycle. In this embodiment, the reward includes event-triggered rewards (also known as first reward rewards) and time-duration rewards (also known as second reward rewards).

[0134] Event-triggered rewards: These rewards are only given once when a specific event is triggered. These rewards are goal-oriented, specifically capture and escape rewards. The reward value is generated when our drone successfully captures an enemy aircraft or when the enemy aircraft successfully enters our defensive line. These reward values ​​are relatively large. This reward primarily helps our drones learn the mission objective: minimizing enemy entry into our defensive line and learning to capture enemy aircraft.

[0135] Time-based rewards: These rewards are continuously calculated and accumulated during the battle. A time-step penalty deducts a small reward value from each drone on the field at each time step. This reward serves two purposes: firstly, it encourages drones to capture enemy drones as quickly as possible; secondly, it limits the number of drones that can be deployed.

[0136] In this embodiment of the application, the revenue function is set as follows:

[0137]

[0138] in,

[0139] R (t) This represents the reward for the t-th period;

[0140] γ represents the discount factor, a constant in the range [0,1]. t This represents the discount factor raised to the power of t.

[0141] It should be noted that the process of UAVs performing air defense requires high real-time updates of decision-making schemes. Therefore, patrol UAVs can be configured to update their own decision-making schemes automatically. An update unit can be pre-configured within the UAV, storing the decision-making scheme update model. After executing the decision-making scheme for one scheduling cycle, the patrol UAV can directly input the detection results into its onboard decision-making scheme update model, thereby using the model to update and obtain the decision-making scheme for the next scheduling cycle.

[0142] In step S5, the decision scheme update model is solved, and the decision scheme of the patrol drone is updated according to the model solution.

[0143] The model solution includes the following steps:

[0144] S501. Obtain the drone status information of the attacking drone based on the state space. The drone status information refers to the attack results of the patrol drone against the detected attacking drone, including the attacking drone defeated by the patrol drone, and the attacking drone that breached the defensive position.

[0145] Specifically, the drone status information may include the number of attacking drones defeated by the patrol drone, referred to as the first number of attacking drones in this embodiment. The drone status information may also include the number of attacking drones that have entered the defensive position of the defending drone formation, referred to as the second number of attacking drones in this embodiment.

[0146] S502. Calculate the reward based on the drone status information. This includes the following steps:

[0147] The first reward is obtained based on the drone's status information, including:

[0148] R target =m1×f1+m2×f2

[0149] in,

[0150] R target This indicates the first reward / reward.

[0151] m1 represents the number of drones of the first attacking party, and m2 represents the number of drones of the second attacking party;

[0152] f1 represents the bonus points for defeating the attacking drone, and f2 represents the bonus points for the attacking drone entering the defensive position. f1 can be 1, and f2 can be -1.

[0153] To receive the second reward, you will receive:

[0154] R time = -0.01×m3

[0155] in,

[0156] R time m3 indicates the second reward; m3 indicates the number of patrol drones.

[0157] Calculate reward returns, including:

[0158] R = R target +R time

[0159] in,

[0160] R represents the reward or incentive.

[0161] S503. Solve the revenue function based on the reward return. The reward return value can be substituted into the revenue function to obtain the model solution that maximizes the revenue function.

[0162] The model solution includes the probability distribution of multiple decision options, i.e. the probability of the patrol drone's decision options in the next cycle.

[0163] The decision-making scheme for patrol drones can be updated based on the model solution, including:

[0164] The decision scheme with the highest probability is obtained from the model solution, and the decision scheme with the highest probability is updated as the decision scheme of the patrol drone in the next cycle.

[0165] If the patrol drone updates its decision-making plan on its own, it can send the updated plan back to the ground command center.

[0166] In step S6, the drone formation defense strategy is updated according to the updated decision scheme, and the defending drone formation is controlled to execute the updated drone formation defense strategy.

[0167] Updating the drone swarm defense strategy includes the following steps:

[0168] S601. Based on the updated decision scheme, obtain the number of third drones whose defensive action is to call friendly drones, and obtain the number of retreating drones whose defensive action is to retreat.

[0169] It should be noted that when a patrol drone calls for a friendly drone, the ground command center needs to dispatch a new drone to perform a patrol flight. To do this, the number of drones that have called for a friendly drone can be counted so that the corresponding number of drones can be dispatched for support.

[0170] S602. Obtain the candidate drone corresponding to the number of the third drone in the defending drone formation, and update the decision scheme of the candidate drone to set the defensive action of the candidate drone to air patrol.

[0171] Candidate drones can be selected from those currently stationed at the defensive positions to perform flight patrol missions.

[0172] S603. Set the backup drone as a patrol drone and cancel the retreating drone from being set as a patrol drone.

[0173] S604. Analyze the updated decision-making schemes of multiple drones in the defending drone formation to generate the drone formation defense strategy for the next cycle.

[0174] In the next cycle, the defender's drone formation can be controlled to execute the updated drone formation defense strategy.

[0175] In some embodiments, considering that patrol drones may encounter some unforeseen situations during patrols, such as insufficient power or fuel, they may need to retreat directly.

[0176] Therefore, after acquiring information on patrol drones in patrol mode, the status of each patrol drone can be monitored to determine whether it is necessary to withdraw them.

[0177] In this embodiment of the application, a patrol cycle is set. When a drone patrols continuously for a period of time until the patrol cycle is reached and no attacking drone is detected during this period, the drone is ordered to retreat.

[0178] Therefore, it is possible to control the patrol drone to start timing after takeoff and begin patrolling, and to detect the duration of the timing.

[0179] If the patrol drone does not detect the attacking drone within the preset patrol period, the patrol drone's decision-making scheme is updated to set the patrol drone's defensive action to retreat.

[0180] If the patrol drone detects an attacking drone within the preset patrol period, the timer duration is updated to 0, the patrol drone is controlled to restart the timer, and the patrol continues.

[0181] In some embodiments, if the attacking drone poses a significant threat to the defending position, it may be necessary for the drone to directly call upon a friendly drone.

[0182] In this embodiment, the following configuration is made: after the patrol drone has completed a scheduling cycle, the location severity of the attacking drone can be determined based on the detection results. If the threat level is high, a friendly drone is directly called. If the threat level is low, the decision-making scheme is updated through a decision-making scheme update model.

[0183] Therefore, after obtaining the detection results of the patrol drone, the patrol threat value of the patrol drone can be obtained. In this embodiment, the patrol threat is defined in two ways: strategic threat and location threat.

[0184] Obtaining patrol threat values ​​involves the following steps:

[0185] The target attacking drone detected by the patrol drone is obtained based on the detection results of the patrol drone.

[0186] The action strategy of the target attacking drone is obtained. The action strategy of the target attacking drone includes an advance strategy, a retreat strategy, and other strategies. The advance strategy is the action of the target attacking drone advancing towards the defensive position of the defending drone formation; the retreat strategy is the action of the target attacking drone withdrawing; and the other strategies are strategies and behaviors other than the advance strategy and the retreat strategy.

[0187] The threat value is obtained based on the action strategy. Specifically, the threat value for the advance strategy can be set to 50, the threat value for the retreat strategy can be set to 10, and the threat value for other strategies can be set to 30.

[0188] Obtain location threat value, including:

[0189]

[0190] in,

[0191] W distance Indicates the location threat value, d max This indicates the maximum patrol distance of the defending drone formation. This indicates the distance between the target attacking drone i and the defensive position.

[0192] The patrol threat value of the patrol drone is calculated based on the strategy threat value and the location threat value.

[0193] For each target attacking drone, a strategy threat value and a location threat value are calculated to obtain the threat value of each target attacking drone.

[0194] The patrol drone may detect multiple target attack drones, and the threat values ​​of all target attack drones are summed to obtain the patrol threat value of the patrol drone.

[0195] In some embodiments,

[0196] If the patrol threat value is greater than or equal to a preset threat value threshold, the decision scheme of the patrol drone is updated to set the patrol drone's defensive action to call a friendly drone. The threat value threshold can be set to 100.

[0197] If the patrol threat value is less than a preset threat value threshold, then the step of generating a decision scheme and updating the model based on the detection results is executed.

[0198] In some embodiments, when a patrol drone detects multiple attacking drones, it can attack them in order of threat value from highest to lowest.

[0199] This invention also provides a reinforcement learning system for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs), characterized in that the system includes:

[0200] The information acquisition module is configured to acquire information about the defending drone formation and the drone formation defense strategy; the defending drone formation includes multiple drones; the drone formation defense strategy is used to instruct the decision-making schemes of the multiple drones, the decision-making schemes include defensive actions and flight directions; the defensive actions include aerial patrol, calling friendly aircraft, and self-retreat;

[0201] The execution module is configured to control the defending drone formation to execute the drone formation defense strategy;

[0202] The detection module is configured to acquire the detection results of the patrol drone in patrol mode;

[0203] The model generation module is configured to generate a decision scheme update model based on the detection results; the decision scheme update model includes a state space, decision schemes, reward rewards, and a payoff function; the state space is generated from the detection results.

[0204] The decision scheme update module is configured to solve the decision scheme update model and update the decision scheme of the patrol drone based on the model solution.

[0205] The defense strategy update module is configured to update the UAV formation defense strategy according to the updated decision scheme, and control the defending UAV formation to execute the updated UAV formation defense strategy.

[0206] It is understood that the target allocation system provided in this embodiment of the invention corresponds to the reinforcement learning method. The explanation, examples, and beneficial effects of the relevant content can be referred to the corresponding content in the reinforcement learning method for multi-UAV adaptive intelligent decision-making, and will not be repeated here.

[0207] The invention also provides a computer-readable storage medium storing a computer program for reinforcement learning in multi-UAV adaptive intelligent decision-making, wherein the computer program causes a computer to execute the reinforcement learning method for multi-UAV adaptive intelligent decision-making as described above.

[0208] In summary, compared with existing technologies, it has the following beneficial effects:

[0209] This invention can acquire information about the defending drone formation and its defense strategy. The defense strategy instructs multiple drones on decision-making methods, including defensive actions and flight directions. Defensive actions include aerial patrol, calling on friendly drones, and drone withdrawal. The system controls the defending drone formation to execute the defense strategy and acquires information about patrol drones in patrol mode, obtaining their detection results. Based on the detection results, a decision-making update model is generated and solved, and the decision-making strategies for the patrol drones are updated according to the model solution. The updated decision-making strategy is then implemented, and the defending drone formation is controlled to execute the updated strategy. By updating the drone decision-making strategies based on their detection results, the scale of the drone formation during mission execution can be scheduled in real time, improving the defense effectiveness.

[0210] It should be noted that, through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms. The technical solutions described above, in essence or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or certain parts of embodiments. Numerous specific details are set forth in the specification provided herein. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0211] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0212] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A reinforcement learning method for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs), wherein the reinforcement learning method is executed by a computer, characterized in that, Includes the following steps: Acquire information on the defending drone formation and its defense strategy; the defending drone formation includes multiple drones; the drone formation defense strategy is used to instruct the decision-making schemes of the multiple drones, the decision-making schemes include defensive actions and flight directions; the defensive actions include aerial patrol, calling on friendly aircraft, and self-retreat; Control the defending drone formation to execute the drone formation defense strategy; Acquire patrol drones that are in patrol mode, and acquire the detection results of the patrol drones; A decision scheme update model is generated based on the detection results; the decision scheme update model includes a state space, decision schemes, reward rewards, and a payoff function; the state space is generated from the detection results. Solve the decision-making scheme update model and update the decision-making scheme of the patrol drone based on the model solution; The drone formation defense strategy is updated according to the updated decision scheme, and the defending drone formation is controlled to execute the updated drone formation defense strategy. Solving the decision update model includes: The drone state information of the attacker's drone is obtained based on the state space; Calculate the reward based on the drone status information; Solve the revenue function based on the reward return; The drone status information includes the number of first attacking drones defeated by the patrol drone and the number of second attacking drones that have entered the defensive position of the defending drone formation. The reward is calculated based on the drone's status information, including: The first reward is obtained based on the drone's status information, including: R target =m1×f1+m2×f2 in, m1 represents the number of drones of the first attacking party, and m2 represents the number of drones of the second attacking party; f1 represents the bonus points for defeating the attacking drone, and f2 represents the bonus points for the attacking drone entering the defensive position. To receive the second reward, you will receive: R time =-0.01×m3 in, m3 represents the number of patrol drones; Calculate reward returns, including: R=R target +R time 。 2. The reinforcement learning method according to claim 1, characterized in that, The profit function is: in, This represents the reward for the t-th period; Υ represents the discount factor, Υ t This represents the discount factor raised to the power of t.

3. The reinforcement learning method according to claim 1, characterized in that, After acquiring patrol drones in patrol mode, the process also includes: Control the patrol drone to keep track of time and monitor the duration of the timer; If the patrol drone does not detect the attacking drone within the preset patrol period, the patrol drone's decision-making scheme is updated to set the patrol drone's defensive action to retreat. If the patrol drone detects an attacking drone within the preset patrol period, the timer duration is updated to 0, and the patrol drone is controlled to keep a timer running.

4. The reinforcement learning method according to claim 1, characterized in that, After obtaining the detection results from the patrol drone, the process also includes: Obtain the patrol threat value of the patrol drone; If the patrol threat value is greater than or equal to a preset threat value threshold, the decision scheme of the patrol drone is updated to set the patrol drone's defensive action to call a friendly drone. If the patrol threat value is less than a preset threat value threshold, then the step of generating a decision scheme and updating the model based on the detection results is executed.

5. The reinforcement learning method according to claim 4, characterized in that, Obtaining the patrol threat value of the patrol drone includes: The target attacking drone detected by the patrol drone is obtained based on the detection results of the patrol drone; The action strategy of the target attacking drone is obtained; the action strategy includes an advance strategy, a retreat strategy, and other strategies; the advance strategy is the action of the target attacking drone advancing towards the defensive position of the defending drone formation; the retreat strategy is the action of the target attacking drone retreating; and the other strategies are strategies other than the advance strategy and the retreat strategy. Obtain the strategy threat value based on the described action strategy; Obtain location threat value, including: in, d max This indicates the maximum patrol distance of the defending drone formation. This indicates the distance between the attacking drone i and the defensive position; The patrol threat value of the patrol drone is calculated based on the strategy threat value and the location threat value.

6. The reinforcement learning method according to claim 1, characterized in that, The model solution includes the probability distribution of multiple decision options; The decision scheme for the patrol drone is updated based on the model solution, including: The decision scheme with the highest probability is obtained from the model solution and updated as the decision scheme of the patrol drone in the next cycle; The drone swarm defense strategy is updated according to the updated decision-making scheme, including: Based on the updated decision scheme, obtain the number of third drones whose defensive action is to call friendly drones, and obtain the number of retreating drones whose defensive action is to retreat. Obtain the candidate drones corresponding to the number of the third drones in the defending drone formation, and update the decision scheme of the candidate drones so as to set the defensive action of the candidate drones to air patrol. Set the standby drone as a patrol drone, and cancel the retreat drone from being set as a patrol drone; The system analyzes the updated decision-making schemes of multiple drones in the defending drone formation to generate the drone formation defense strategy for the next cycle.

7. A reinforcement learning system for adaptive intelligent decision-making among multiple unmanned aerial vehicles (UAVs), characterized in that, The system for performing the reinforcement learning method as described in claim 1 includes: The information acquisition module is configured to acquire information about the defending drone formation and its defense strategy. The defending drone formation includes multiple drones. The drone formation defense strategy is used to instruct the decision-making schemes of the multiple drones, including defensive actions and flight directions. The defensive actions include aerial patrol, calling on friendly aircraft, and self-retreat. The execution module is configured to control the defending drone formation to execute the drone formation defense strategy; The detection module is configured to acquire the detection results of the patrol drone in patrol mode; The model generation module is configured to generate a decision scheme update model based on the detection results; the decision scheme update model includes a state space, decision schemes, reward rewards, and a payoff function; the state space is generated from the detection results. The decision scheme update module is configured to solve the decision scheme update model and update the decision scheme of the patrol drone based on the model solution. The defense strategy update module is configured to update the UAV formation defense strategy according to the updated decision scheme, and control the defending UAV formation to execute the updated UAV formation defense strategy.

8. A computer-readable storage medium, characterized in that, It stores a computer program for reinforcement learning of multi-UAV adaptive intelligent decision-making, wherein the computer program causes a computer to execute the reinforcement learning method for multi-UAV adaptive intelligent decision-making as described in any one of claims 1 to 6.