Unmanned aerial vehicle swarm defense penetration decision model construction method based on reinforcement learning

By constructing a drone swarm penetration decision-making model based on reinforcement learning, combining dynamic characteristics and trajectory planning, and using a proximal strategy optimization algorithm PPO, the problem of insufficient autonomous combat capabilities of drone swarms in complex environments is solved, and the penetration effect and combat effectiveness of drone swarms are improved.

CN120578083AActive Publication Date: 2025-09-02JIANGNAN ELECTROMECHANICAL DESIGN INST

Patent Information

Application Number
CN202510492197.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-09-02
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing classic reinforcement learning algorithms have poor generalization in drone swarm penetration task scenarios, slow convergence speed, large state dimensions, and high training difficulty. They cannot effectively improve the autonomous combat capability of drone swarms in complex environments.

Method used

Build a drone swarm penetration decision-making model based on reinforcement learning. By building an offensive and defense confrontation training environment, establish a drone single kinematic model, generate trajectories and perform control quantity calculations, define Markov decision-making process, and use a proximity strategy optimization algorithm PPO for model optimization, combining dynamic characteristics and trajectory planning to improve the penetration effect of drone swarms.

Benefits of technology

It significantly improves the autonomous combat capability of drone bee swarms in complex environments, improves the penetration effect and combat effectiveness of drone bee swarms, and achieves more efficient execution of autonomous air combat missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120578083A_ABST
    Figure CN120578083A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle swarm defense penetration decision model construction method based on reinforcement learning. The method comprises the following steps: constructing an attack-defense confrontation-oriented training environment; establishing an unmanned aerial vehicle single body kinematic model, including defining a motion model of the unmanned aerial vehicle and main constraints of the unmanned aerial vehicle, and calculating a speed position integral at a next moment; the main constraints comprise regional constraints and dynamic constraints; the regional constraint limits the exploration region of the unmanned aerial vehicle, and the dynamic constraint mainly restrains the speed and acceleration of the unmanned aerial vehicle; generating an unmanned aerial vehicle track, associating the unmanned aerial vehicle track with each unmanned aerial vehicle, and determining a track starting point and a penetration target point; resolving the control quantity of the unmanned aerial vehicle; defining a Markov decision process model; and optimizing the Markov decision process model, and constructing an unmanned aerial vehicle swarm defense penetration decision model. According to the technical scheme, strategy optimization of the unmanned aerial vehicle can contain global environment information, the strategy can be adjusted for a specific air attack task, and the defense penetration effect and combat efficiency of the unmanned aerial vehicle are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence applications for unmanned aerial vehicles (UAVs), and in particular to a method for constructing a UAV swarm penetration decision-making model based on reinforcement learning. Background Art

[0002] In today's era, intelligence and autonomy have become key features of technological development, and drone technology is being widely used in fields such as military, logistics, and inspection. As demand for drone swarm operations continues to grow, the ability of drones to operate autonomously and penetrate complex environments has become increasingly crucial. However, faced with complex and ever-changing air combat environments and high-intensity interference, drone swarms often need to make rapid decisions and execute evasive or penetration missions in dynamic and uncertain environments.

[0003] Reinforcement learning (RL) is a machine learning method that aims to learn how to take optimal actions in different states to maximize long-term cumulative rewards through the interaction between an agent and its environment. In the RL framework, the agent selects an action based on the state of the environment, and the environment provides rewards and updates the state based on this action. The agent continuously adjusts its policy to learn how to take the best action in a given state. Reinforcement learning is widely used in fields such as robotic control, drone navigation, gaming, and autonomous driving, demonstrating strong adaptability and decision-making optimization capabilities, particularly in complex tasks with high-dimensional state and action spaces.

[0004] In this context, existing technologies often use reinforcement learning as a key tool for autonomous learning and decision-making in drones. Through reinforcement learning training, drone swarms can continuously optimize their strategies in complex adversarial environments, improving their evasion, defense, and penetration capabilities, thereby enabling more efficient autonomous air combat mission execution. However, existing classic reinforcement learning algorithms cannot be directly applied to multi-drone coordinated penetration mission scenarios, suffering from poor generalization, slow convergence, large state dimensions, and high training difficulty. Summary of the Invention

[0005] To solve the above problems, this application provides a method for constructing a UAV swarm penetration decision model based on reinforcement learning, which includes the following steps:

[0006] Establishing a training environment for offensive and defensive confrontation; the training environment includes: the movement area of ​​the UAV, the combat area for the UAV swarm penetration, and the parameters of the UAV;

[0007] Establish a UAV single-unit kinematic model based on the UAV single-unit kinematics. Establishing the UAV single-unit kinematic model includes defining the UAV's motion model and the main constraints of the UAV, and calculating the velocity and position integral at the next moment. The main constraints include area constraints and dynamic constraints. Area constraints limit the UAV's exploration area, and dynamic constraints mainly constrain the UAV's speed and acceleration.

[0008] Based on the UAV single-unit kinematic model, the polynomial coefficient method is used to generate trajectories. Based on the trajectories, a UAV trajectory consisting of multiple discrete points is generated. The UAV trajectory is associated with each UAV, with the UAV position as the trajectory starting point and the trajectory end point as the penetration target point.

[0009] Calculate the control amount of the drone based on the drone trajectory to obtain the linear acceleration and angular acceleration required by the drone when it moves from its current position to the penetration target point;

[0010] Define a Markov decision process model; the polynomial coefficients when generating the drone trajectory constitute the drone state set S, the polynomial coefficient increments constitute the action set A, define the reward function R, the transition probability P and the discount factor γ, and combine the drone state set S and the action set A to construct the Markov decision process quintuple (S, A, R, P, γ);

[0011] The proximal strategy optimization algorithm is used to optimize the Markov decision process model and construct a UAV swarm penetration decision model.

[0012] Among them, the establishment of a UAV single-unit kinematic model based on the UAV single-unit kinematics includes:

[0013] The linear acceleration and angular velocity of the UAV are defined by the motion model to reflect the UAV's behavior. The motion model is expressed as:

[0014] Among them, x and y are the current positions of the drone, ψ is the heading angle, and a x is the linear acceleration, a z is the angular acceleration, V is the linear velocity, ω is the angular velocity, is the projection of the drone speed on the x-axis, is the projection of the UAV speed on the y-axis;

[0015] The dynamic constraint is expressed as:

[0016]

[0017] Among them, x min 、x max 、y min 、y max is the UAV position range, V min 、Vmax is the speed range, a xmax is the maximum linear acceleration, a zmax is the maximum angular acceleration;

[0018] The method for calculating the velocity position integral at the next moment is:

[0019] Where: y n It is the speed or position at the current moment, y n+1 is the speed or position at the next moment, h is the step size, k1, k2, k3 and k4 are the estimated values ​​of the four slopes, and the calculation method is:

[0020] When the polynomial coefficient method is used to generate trajectories, a set of polynomials defines a trajectory, and the coefficients of the polynomials determine the shape of the trajectory. By adjusting the polynomial coefficients to form multiple sets of polynomials, different trajectories can be generated.

[0021] A set of polynomials defines a trajectory, which is generated by combining X(t) and Y(t) to generate a trajectory starting at (0, 0) and ending at (1, 0) in a two-dimensional plane. It is expressed as:

[0022]

[0023] Among them, d0, d1, …, d n 、c0,c1,…,c n are the coefficients of the polynomial, t is the normalized time variable, and t∈[0,1].

[0024] Furthermore, generating a UAV trajectory consisting of multiple discrete points based on the trajectory includes the following steps:

[0025] Define the number of sampling points N, perform uniform sampling based on the normalized time variable t, and calculate X(t) and Y(t) corresponding to each time variable t;

[0026] The points corresponding to X(t) and Y(t) are combined to form the UAV trajectory composed of discrete points, which can be expressed as:

[0027] trajectory={(x1,y1),(x2,y2),…,(x m ,y m )}.

[0028] Furthermore, associating the drone trajectory with each drone means:

[0029] The drone trajectory composed of discrete points is translated so that the starting point of the drone trajectory is translated to the position of the drone. At this time, the end point of the drone trajectory is the drone's penetration target point.

[0030] Specifically, associating the drone trajectory with each drone includes the following steps:

[0031] Use a normalized vector to represent the direction from the starting point to the end point in the UAV trajectory;

[0032] Calculate the angle θ between the normalized vector and the positive direction of the X-axis;

[0033] Construct a two-dimensional rotation matrix R to express the angle θ of the trajectory rotation around the origin, so that the normalized vector of each discrete point is aligned from the original direction to the direction of the penetration target point; the two-dimensional rotation matrix R is expressed as:

[0034] Scale the rotated discrete points so that the length of the trajectory is extended from unit length to the modulus length of the target vector;

[0035] Translate the drone trajectory so that the starting point of the drone trajectory coincides with the position of the translated drone.

[0036] Furthermore, calculating the control amount of the drone based on the drone trajectory includes:

[0037] Calculate the drone's angular acceleration, linear velocity, and target time;

[0038] The angular acceleration of the UAV is obtained by calculating the deviation between the current heading angle of the UAV and the target position when the UAV deviates from the track;

[0039] The target time is the time required for the UAV to navigate from the initial speed to the maximum acceleration. At the same time, the UAVs on other tracks navigate at an acceleration lower than the maximum acceleration.

[0040] Among them, the reward function R in the Markov decision process quintuple is: in the r time step of a drone breakout operation t If the requirement of maximizing the number of drone breakthroughs is met, the design reward is the number of drones that successfully break through.

[0041] Furthermore, an attack-defense simulation environment is defined during the proximal strategy optimization process. The core components of this environment include: a drone state set S, an action set A consisting of polynomial coefficient increments, and a reward function R. During the optimization process, the drone agent interacts with the environment through the reset() and step() methods to sample states, select actions, receive rewards, and determine whether the optimization is complete.

[0042] Among them, the reset() method returns to the initial state, and the step(action) method returns the current state, executes the reward, and determines whether it is finished.

[0043] According to the present invention, the attack and defense confrontation simulation method integrating reinforcement learning can be used to combine dynamic characteristics, trajectory planning and reinforcement learning, so that the strategy optimization of the UAV not only includes global environmental information, but also can adjust the strategy for specific air strike missions, thereby significantly improving the UAV's penetration effect and combat effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 Flowchart of a method for constructing a UAV swarm penetration decision model based on reinforcement learning according to an embodiment of the present invention;

[0045] Figure 2 2. Schematic diagram of a battlefield simulation environment for a method for constructing a UAV swarm penetration decision model based on reinforcement learning according to an embodiment of the present invention;

[0046] Figure 3 This is a drone control action diagram for a method of constructing a drone swarm penetration decision model based on reinforcement learning in this application. DETAILED DESCRIPTION

[0047] Based on deep reinforcement learning, the present invention adopts the Proximal Policy Optimization (PPO) algorithm to solve the problem of UAV swarm penetration in air combat environment, and proposes a UAV swarm penetration decision model construction method based on reinforcement learning.

[0048] The specific implementation of the present invention is described in detail below with reference to the accompanying drawings.

[0049] like Figure 1 As shown in FIG, the method for constructing a UAV swarm penetration decision model based on reinforcement learning includes the following steps:

[0050] A method for constructing a UAV swarm penetration decision model based on reinforcement learning, characterized by comprising the following steps:

[0051] Step S100: Building a training environment for attack and defense confrontation;

[0052] In the present invention, only the planar motion of the drone performing constant-altitude flight is considered. Therefore, the three-dimensional motion in reality is simplified to two dimensions (x and y axes). The training environment does not consider the influence of other external factors such as weather, network speed, temperature, etc., in order to simplify the simulation environment.

[0053] Building a training environment includes:

[0054] 1) Define the drone movement area: The combat zone for the drone swarm penetration is defined as 30 kilometers long and 30 kilometers wide. The blue team is the air attack team, the red team is the air defense team, the air attack team has N drones, and the air defense team has M weapons. The blue team's drones can take off from a specific area, and the red team's position is a circular area with a radius of 5 kilometers located in the center of the battlefield.

[0055] 2) Define the drone's parameters: Parameter settings include a breakout radius of 0 to 1000 meters, a maximum speed of 50 meters per second, and a minimum speed of 30 meters per second; a maximum linear acceleration of 2 meters per second squared in the X direction, and a maximum angular acceleration of 20 degrees per second squared in the Z direction. The defense system equipment in this invention uses an already developed equipment model, so agent modeling of the defense system is not required.

[0056] 3) Define the combat area for drone cluster penetration: In the present invention, it is assumed that multiple drone formations arrive at the mission area boundary at the same time. After arriving at the mission area, the drone formations attack along their respective routes. Therefore, the flight trajectory of the drones involved in the present invention is divided into two stages, such as Figure 2 As shown: In the first stage, the drones need to form a formation and arrive at the designated mission point. In the second stage, the drones take off from the mission point to conduct a breakthrough.

[0057] Step S110: establishing a drone unit kinematics model according to the drone unit kinematics;

[0058] Establishing a single UAV kinematic model includes defining the UAV's motion model and main constraints, and calculating the velocity and position integrals at the next moment; the main constraints include area constraints and dynamic constraints; the area constraints limit the UAV's exploration area, and the dynamic constraints mainly constrain the UAV's velocity and acceleration;

[0059] The kinematic modeling process of a UAV unit includes:

[0060] The linear acceleration and angular velocity of the drone are defined by the motion model to reflect the drone's behavior: Since the drone's behavior set includes straight flight and turning, which can be achieved by controlling the drone's acceleration, the drone's behavior is simplified to linear acceleration and angular acceleration. The drone's motion model is expressed as: Among them, x and y are the current positions of the drone, ψ is the heading angle, and a x is the linear acceleration, a z is the angular acceleration, V is the linear velocity, ω is the angular velocity, is the projection of the drone speed on the x-axis, is the projection of the UAV speed on the y-axis;

[0061] The dynamic constraint is expressed as:

[0062]

[0063] Among them, x min 、x max 、y min 、y max is the UAV position range, V min 、V max is the speed range, a xmax is the maximum linear acceleration, a zmax is the maximum angular acceleration.

[0064] In this invention, the Runge-Kutta method is used to calculate the velocity and position integral at the next moment, avoiding the divergence problem caused by an inappropriate time step. When the derivative and initial value of the equation are known, the computer simulation application is used to eliminate the complex process of solving the differential equation; wherein the differential equation is expressed as:

[0065] y′=f(t,y),y(t0)=y0;

[0066] The Runge-Kutta method is used to perform differential fitting, which is implemented by the following equation:

[0067]

[0068] Where: y n It is the speed or position at the current moment, y n+1 is the speed or position at the next moment, h is the step size, k1, k2, k3 and k4 are the estimated values ​​of the four slopes, and the calculation method is:

[0069] Step S120: Based on the UAV single body kinematic model, a polynomial coefficient method is used to generate a trajectory; based on the trajectory, a UAV trajectory consisting of multiple discrete points is generated; the UAV trajectory is associated with each UAV, with the UAV position being the trajectory starting point and the trajectory end point being the penetration target point;

[0070] 1) When using the polynomial coefficient method to generate trajectories, a set of polynomials defines a trajectory, and the coefficients of the polynomials determine the shape of the trajectory. By adjusting the polynomial coefficients to form multiple sets of polynomials, different trajectories can be generated.

[0071] In this step, using polynomial coefficients to generate trajectories is a method of modeling and generating trajectories using polynomial functions. The core idea of ​​this method is to fit the position change relationship of trajectory points (usually two-dimensional x and y coordinates) through polynomial fitting, such as Figure 1 As shown in step 120, expert-designed drone trajectories or existing empirical data can be introduced to significantly improve the decision accuracy of the decision model and reduce the training time of the intelligent decision model.

[0072] In trajectory generation, each trajectory is defined by a set of polynomials. The coefficients of the polynomials determine the shape of the trajectory. By adjusting the polynomial coefficients, different trajectories can be generated. Using X(t) and Y(t) together, a trajectory generated by the polynomial with a starting point at (0, 0) and an end point at (1, 0) can be expressed as:

[0073]

[0074] Among them, d0, d1, …, d n 、c0,c1,…,c n are the coefficients of the polynomial, t is the normalized time variable, and t∈[0,1].

[0075] In addition, for routes designed based on expert experience and rules, Opencv can be used to collect route data, and then the polynomial parameters can be obtained by data fitting.

[0076] 2) The process of generating a UAV trajectory consisting of multiple discrete points based on the trajectory includes:

[0077] Define the number of sampling points N and perform uniform sampling based on the normalized time variable t. That is, divide the trajectory generated in the above steps into multiple points and calculate the X(t) and Y(t) corresponding to each time variable t to represent the position in two-dimensional space.

[0078] The points corresponding to X(t) and Y(t) are combined to form the UAV trajectory composed of discrete points, which can be expressed as:

[0079] trajectory={(x1,y1),(x2,y2),…,(x m ,y m )};

[0080] The drone trajectories generated by this step can be associated with each drone.

[0081] Ultimately, each drone corresponds to a drone trajectory consisting of multiple discrete points. By adjusting the polynomial coefficients, the shape of the trajectory can be flexibly controlled to meet specific needs.

[0082] 3) Further, the drone trajectory generated in the above step is associated with the drone: the drone trajectory composed of discrete points is translated so that the starting point of the drone trajectory is translated to the position of the drone. At this time, the end point of the drone trajectory is the drone's penetration target point;

[0083] The implementation method of this step specifically includes:

[0084] Use a normalized vector to represent the direction from the starting point to the end point in the UAV trajectory;

[0085] Calculate the angle θ between the normalized vector and the positive direction of the X-axis. The angle θ can be used for each discrete point to represent the angle of rotation from the original direction (along the positive direction of the X-axis) to the new direction (pointing to the penetration target point).

[0086] Construct a two-dimensional rotation matrix R to express the angle θ of the trajectory rotation around the origin, so that the normalized vector of each discrete point is aligned from the original direction to the direction of the penetration target point; the two-dimensional rotation matrix R is expressed as:

[0087] Scale the rotated discrete points so that the length of the trajectory is extended from unit length to the modulus length of the target vector;

[0088] Translate the drone trajectory so that the starting point of the drone trajectory is consistent with the position of the translated drone;

[0089] At this time, the starting and ending points of the drone trajectory constitute the specific discrete trajectory from the current position of the drone to the penetration target point.

[0090] Step S130: After associating the drone trajectory with each drone, the drone control variable is calculated based on the drone trajectory to obtain the linear acceleration and angular acceleration required by the drone in the process of moving from the current position to the penetration target point;

[0091] The calculation of UAV control variables includes:

[0092] 1) Calculate the angular acceleration of the drone:

[0093] When the drone deviates from the track, it is necessary to correct the target direction of the drone. The principle of correction is as follows: Figure 3 As shown: With the current position of the drone as the center of the circle, extend the point on the track closest to the drone (the intersection of the green dotted line and the track) along the direction of the drone's movement. The extended point C is the point the drone should reach.

[0094] The angular acceleration can be obtained by calculating the deviation between the current heading angle of the drone and the target position.

[0095] 2) Calculate linear velocity:

[0096] The acceleration formula is expressed as: Among them, v0 represents the initial velocity, l represents the distance the current drone needs to travel, a and t represent acceleration and time respectively.

[0097] 3) Calculate target time:

[0098] The UAV has speed and acceleration limits, and the minimum and maximum UAV navigation speeds are recorded as v min and v max , the minimum and maximum UAV navigation acceleration are recorded as amin and a max , the longest track is recorded as l max , then the target time The calculation method is:

[0099]

[0100] Target time The time required for the drone to accelerate from its initial speed to its maximum acceleration is 2.5 times the time required for the drone to accelerate from its initial speed to its maximum acceleration. Meanwhile, drones on other trajectories are traveling at an acceleration lower than the maximum acceleration. In addition, to increase the probability of penetration, the drone must enter the equipment's detection range at its maximum speed.

[0101] 4) Calculate the acceleration a of other drones:

[0102] The system of equations defining the trajectory of the drone is expressed as:

[0103]

[0104] Among them, l1 and l2 are the flight distances of the drone, and v0 is the initial velocity;

[0105] At this time, the acceleration a required by other UAVs can be obtained by solving the equation group, and the UAVs can be guaranteed to navigate at the maximum speed when entering the detection range of the equipment.

[0106] At the same time, in order to avoid the situation where there is no analytical value for the equation group, the equation group is optimized and expressed as:

[0107]

[0108] st-2≤a≤2,

[0109] in,

[0110] Step S140: Define a Markov decision process model; the polynomial coefficients used in generating the drone trajectory constitute the drone state set S, and the polynomial coefficient increments constitute the action set A. Define the reward function R, the transition probability P, and the discount factor γ to combine the drone state set S and the action set A to construct a Markov decision process quintuple (S, A, R, P, γ);

[0111] 1) The state set S is the polynomial coefficient for generating the corresponding trajectory of each drone, s i represents the vector of polynomial coefficients (five coefficients corresponding to X(t) and five coefficients corresponding to Y(t)) that generate the trajectory of the i-th UAV in the formation;

[0112] 2) The state set A is the increment of each polynomial coefficient in the iterative optimization process; a iA vector representing the changes in the polynomial coefficients used to generate the trajectory of the i-th UAV in the formation; the polynomial coefficients are adjusted by adjusting the number of polynomial coefficients to change the UAV's route;

[0113] 3) The reward function R uses the rules that match the PPO (Proximal Policy Optimization) algorithm, and the reward value r is directly obtained through environmental interaction. t , at each time step (i.e., one round of internal drone breakout operation) r t If the requirement of maximizing the number of drone breakthroughs is met, the reward is the number of drones that successfully break through (survive and reach the breakthrough target point);

[0114] 4) The transition probability P adopts the rule that matches the PPO (Proximal Policy Optimization) algorithm. Since PPO directly generates trajectory data (s t ,a t ,r t ,s t+1 ), and optimizing the strategy based on this data does not require explicit modeling of the state transition function P(s'|s,a) of the environment.

[0115] Step S150: Using a proximal strategy optimization algorithm, the Markov decision process model is optimized to construct a UAV swarm penetration decision model.

[0116] In this paper, Proximal Policy Optimization (PPO) is used as the core reinforcement learning algorithm, and the policy gradient method is used to optimize the agent's policy. PPO constrains the pace of policy updates by clipping the objective function, balancing the algorithm's convergence speed and stability, and increasing training stability in high-dimensional state spaces and complex environments. PPO performs multiple stochastic gradient descents (SGD) after each round of sampling, avoiding complex second-order gradient calculations. The specific optimization process is as follows:

[0117] 1) At each policy update, PPO uses a clipping objective function to limit the range of change between the old and new policies, expressed as:

[0118] in, Represents the probability ratio of the new and old strategies, ∈ is the clipping threshold, which is used to limit r t The range of variation of (θ), A t is the advantage function, which measures the current action a t How good or bad the action is compared to the average;

[0119] 2) At the same time, the joint optimization of the strategy and value function, that is, PPO simultaneously optimizes the strategy network π θ Sum value function V φ (s t ), the goal is to minimize the following loss function:

[0120] L(θ,φ)=L CLIP (θ)-c1L VF (φ)+c2H(π θ ),

[0121] in, is the value function error, H(π θ ) is the entropy regularization term of the strategy, which is used to increase exploratory power. c1 and c2 are hyperparameters that control the weights of the value function loss and the entropy term, respectively.

[0122] 3) The PPO algorithm does not explicitly rely on a traditional experience replay pool (such as DQN's experience pool), but instead adopts an on-policy strategy to generate data directly based on the current policy.

[0123] In the present invention, the mechanism of experience sampling and processing is as follows:

[0124] Sampling data generation: Each round samples (s t ,a t ,r t ,s t+1 , done). According to the current strategy π θ Execute action a t , record the corresponding reward r t and the next state s t+1 ;

[0125] Advantage function calculation: The advantage function is calculated using the GeneralizedAdvantageEstimation (GAE) method:

[0126] A t =δ t +(γλ)δ t+1 +(γλ) 2 δ t+2 +…,

[0127] Where δt=rt+γV(s t+1 )-V(s t );

[0128] Data storage: The sampled data is stored in a memory queue (experience pool), and the size of each batch is set by the training batch;

[0129] Strategy update: After each round of sampling, the strategy and value function are optimized using mini-batch stochastic gradient descent (SGD). In SGD, the samples in the experience pool are traversed multiple times to fully utilize the samples.

[0130] 4) The present invention designs a custom environment (attack and defense simulation environment) for specific mission scenarios, and uses the RayRLlib framework for encapsulation and training. The environment design registers the custom environment as a Gym interface through the env_creator method, which supports direct algorithm calls. The core components of the custom environment design include: state S, action A, and reward R. During the optimization process, the reset() and step() methods are used to realize the interaction between the drone agent and the environment, sample the state, select the action, receive the reward, and determine whether it is over; the state S is the polynomial coefficient for generating the trajectory, the action A is the change in the polynomial coefficient, and the reward R is the number of drones that break through; the reset() method returns to the initial state, and the step(action) method returns the current state, executes the reward, determines whether it is over, and additional information;

[0131] 5) Use independent configuration files to flexibly define environmental parameters (such as map scale, number of agents, target point distribution, etc.) to support training needs in different scenarios.

[0132] 6) In this invention, efficient training and model management are achieved through the distributed computing framework provided by the Ray platform; PyTorch is used as a deep learning framework to facilitate model design and optimization.

[0133] 7) The specific configuration of the PPO algorithm is as follows:

[0134] Training batch: The total number of samples for each training is set to 2048;

[0135] Mini-batch: The mini-batch size for optimization using stochastic gradient descent (SGD) is 512;

[0136] Optimization steps: 10 optimization iterations per training batch;

[0137] Discount factor: set to 0.99 to balance short-term rewards with long-term gains;

[0138] Working process: multi-process training to improve training efficiency;

[0139] Resource allocation: Configure 1 GPU and multiple CPU cores for efficient parallel computing.

[0140] After each training session, the average reward and episode length are output, facilitating real-time monitoring of algorithm performance. Model checkpoints are saved at regular intervals (e.g., 50 iterations), allowing for subsequent loading and continued training. Furthermore, models can be loaded from checkpoints to restore training status.

[0141] The method for constructing a drone swarm penetration decision model provided by the present invention first constructs a reinforcement learning simulation environment for drone interaction; then constructs a single drone kinematic model to lay the foundation for subsequent strategy optimization; and generates drone trajectories by combining the polynomial coefficient method with drone characteristics, and solves the drone control quantity based on the discrete point trajectory; further combines the reinforcement learning method to model the Markov decision process, and adopts an intelligent optimization framework for drone penetration decision-making; finally, the drone and defense equipment are trained on an intelligent agent for attack and defense confrontation, and the drone swarm's air combat breakthrough capability is improved during the continuous optimization of the drone swarm penetration strategy, thereby realizing the construction of a drone swarm penetration decision model. The present invention integrates the attack and defense confrontation simulation method of reinforcement learning, combines dynamic characteristics, trajectory planning and reinforcement learning, so that the drone strategy optimization not only includes global environmental information, but also can adjust the strategy for specific air attack missions, significantly improving the drone's penetration effect and combat effectiveness.

[0142] The above disclosures are only a few specific embodiments of the present invention. However, the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present invention.

Claims

1. A method for constructing a UAV swarm penetration decision model based on reinforcement learning, characterized in that: The following steps are involved: Build a training environment for offensive and defensive confrontation; The training environment includes: the UAV's movement area, the UAV swarm penetration combat area, and the UAV's parameters; Establishing a single-unit UAV kinematic model based on the UAV's single-unit kinematics, wherein establishing the single-unit UAV kinematic model includes defining the UAV's motion model and the main constraints of the UAV, and calculating the velocity and position integral at the next moment; the main constraints include area constraints and dynamic constraints; the area constraints limit the UAV's exploration area, and the dynamic constraints mainly constrain the UAV's velocity and acceleration; Based on the UAV single body kinematic model, a polynomial coefficient method is used to generate a trajectory; based on the trajectory, a UAV trajectory consisting of multiple discrete points is generated; the UAV trajectory is associated with each UAV, with the position of the UAV being the starting point of the trajectory and the end point of the trajectory being the penetration target point; Calculating the control amount of the drone based on the drone trajectory to obtain the linear acceleration and angular acceleration required by the drone in the process of moving from the current position to the penetration target point; Define a Markov decision process model; the polynomial coefficients when generating the drone trajectory constitute the drone state set S, the polynomial coefficient increments constitute the action set A, define the reward function R, the transition probability P and the discount factor γ, and combine the drone state set S and the action set A to construct the Markov decision process quintuple (S, A, R, P, γ); The proximal strategy optimization algorithm is used to optimize the Markov decision process model and construct a UAV swarm penetration decision model.

2. The method for constructing a UAV swarm penetration decision model according to claim 1, characterized in that: The establishment of a UAV monomer kinematic model according to the UAV monomer kinematics includes: The linear acceleration and angular velocity of the UAV are defined by the motion model to reflect the UAV's behavior. The motion model is expressed as: Among them, x and y are the current positions of the drone, ψ is the heading angle, and a x is the linear acceleration, a z is the angular acceleration, V is the linear velocity, ω is the angular velocity, is the projection of the drone speed on the x-axis, is the projection of the UAV speed on the y-axis; The dynamic constraint is expressed as: Among them, x min 、x max 、y min 、y max is the UAV position range, V min 、V max is the speed range, a xmax is the maximum linear acceleration, a zmax is the maximum angular acceleration; The method for calculating the velocity position integral at the next moment is: Where: y n It is the speed or position at the current moment, y n+1 is the speed or position at the next moment, h is the step size, k1, k2, k3 and k4 are the estimated values ​​of the four slopes, and the calculation method is:

3. The method for constructing a UAV swarm penetration decision model according to claim 1, characterized in that: When the polynomial coefficient method is used to generate a trajectory, a set of polynomials defines a trajectory, and the coefficients of the polynomials determine the shape of the trajectory; by adjusting the polynomial coefficients to form multiple sets of polynomials, different trajectories are generated.

4. The method for constructing a UAV swarm penetration decision model according to claim 3, characterized in that: The set of polynomials defining a trajectory means that by jointly using X(t) and Y(t), a trajectory starting at (0, 0) and ending at (1, 0) is generated in a two-dimensional plane, which is expressed as: Among them, d0, d1, …, d n 、c0,c1,…,c n are the coefficients of the polynomial, t is the normalized time variable, and t∈[0,1].

5. The method for constructing a UAV swarm penetration decision model according to claim 4, characterized in that: Generating a UAV trajectory consisting of a plurality of discrete points based on the trajectory comprises the following steps: Define the number of sampling points N, perform uniform sampling based on the normalized time variable t, and calculate X(t) and Y(t) corresponding to each time variable t; The points corresponding to X(t) and Y(t) are combined to form the UAV trajectory composed of discrete points, which can be expressed as: trajectory={(x1,y1),(x2,y2),…,(x m ,y m )}。 6. The method for constructing a UAV swarm penetration decision model according to claim 5, characterized in that: Associating the drone trajectory with each drone refers to: The drone trajectory composed of discrete points is translated so that the starting point of the drone trajectory is translated to the position of the drone. At this time, the end point of the drone trajectory is the drone's penetration target point.

7. The method for constructing a UAV swarm penetration decision model according to claim 6, characterized in that: The process of associating the drone trajectory with each drone comprises the following steps: Use a normalized vector to represent the direction from the starting point to the end point in the UAV trajectory; Calculate the angle θ between the normalized vector and the positive direction of the X-axis; Construct a two-dimensional rotation matrix R to express the angle θ of the trajectory rotation around the origin, so that the normalized vector of each discrete point is aligned from the original direction to the direction of the penetration target point; the two-dimensional rotation matrix R is expressed as: Scale the rotated discrete points so that the length of the trajectory is extended from unit length to the modulus length of the target vector; Translate the drone trajectory so that the starting point of the drone trajectory coincides with the position of the translated drone.

8. The method for constructing a UAV swarm penetration decision model according to claim 1, characterized in that: The calculation of the drone control amount based on the drone trajectory includes: Calculate the drone's angular acceleration, linear velocity, and target time; The angular acceleration of the UAV is obtained by calculating the deviation between the current heading angle of the UAV and the target position when the UAV deviates from the track; The target time is the time required for the UAV to navigate from the initial speed to the maximum acceleration. At the same time, the UAVs on other tracks navigate at an acceleration lower than the maximum acceleration.

9. The method for constructing a UAV swarm penetration decision model according to claim 1, characterized in that: The reward function R in the Markov decision process quintuple is: t If the requirement of maximizing the number of drone breakthroughs is met, the design reward is the number of drones that successfully break through.

10. The method for constructing a UAV swarm penetration decision model according to claim 1, characterized in that: The proximal strategy optimization process defines an attack-defense simulation environment. The core components of the attack-defense simulation environment include: a drone state set S, an action set A consisting of polynomial coefficient increments, and a reward function R. During the optimization process, the reset() and step() methods are used to enable the drone agent to interact with the environment, sample states, select actions, receive rewards, and determine whether to end. Among them, the reset() method returns to the initial state, and the step(action) method returns the current state, executes the reward, and determines whether it is finished.

Citation Information

Patent Citations

  • Unexpected threat-oriented unmanned bee colony cooperative route planning method

    CN116448119A

  • Unmanned aerial vehicle cluster flight simulation system

    CN118034087A

Cited By

  • Training method based on dynamic adjustment reward mechanism

    CN120975268A