Rectangular region encirclement method for unmanned aerial vehicles based on differential game and reinforcement learning

By combining differential game theory and reinforcement learning, and introducing a rectangular flight area constraint loss function and batch temporal behavior constraints, the decision network of the UAV swarm is optimized, solving the problem of collaborative encirclement within a rectangular area and achieving efficient and stable collaborative interception of UAV swarms.

CN121325911BActive Publication Date: 2026-05-15GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511471869.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-15
Publication Date
2026-05-15
Estimated Expiration
2045-10-15

AI Technical Summary

Technical Problem

Existing collaborative interception technologies for drone swarms within a rectangular, limited flight area suffer from problems such as the decision network's difficulty in achieving stable convergence, the reward function's difficulty in effectively learning constraints, and poor swarm collaboration. In particular, under strong spatial constraints, it is difficult to achieve efficient and stable encirclement and capture.

Method used

By employing a method based on differential game theory and reinforcement learning, combined with a rectangular flight area constraint loss function and batch temporal behavior constraints, the decision network is optimized through a differential game value evaluation module, and spatial and temporal constraints are introduced to achieve stable collaborative capture of drone swarms.

Benefits of technology

It improves the decision-making stability and collaborative efficiency of drone swarms within rectangular areas, enhances environmental adaptability and interception success rate, and achieves rapid and stable collaborative encirclement and capture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121325911B_ABST
    Figure CN121325911B_ABST
Patent Text Reader

Abstract

The application discloses a rectangular area unmanned aerial vehicle encircling method based on differential game and reinforcement learning. The environment state information is obtained by the unmanned aerial vehicle cluster; the spatial constraint loss is calculated through the rectangular flight area restriction loss function module, and the time sequence consistency loss is calculated through the batch time sequence behavior constraint module, so that the state behavior joint hidden variable of the fusion space and time sequence constraint is generated. The hidden variable and the current environment state are fused and input into the differential game value evaluation module, the differential game between the unmanned aerial vehicle cluster and the black unmanned aerial vehicle is solved, the future expected value based on Nash equilibrium is obtained, and the decision network parameter is updated. The cluster executes the optimized encircling action, and the experience data is stored for continuous training. The application effectively improves the decision stability, cooperative efficiency and task success rate of the unmanned aerial vehicle cluster in the rectangular limited flight area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative control and intelligent decision-making technology for unmanned aerial vehicles (UAVs), and more specifically, to a method for UAV encirclement and capture in a rectangular area based on differential game theory and reinforcement learning. Background Technology

[0002] Cooperative interception technology for unauthorized drones flying within a rectangular, confined flight area has become a key means of countering threats from "low, slow, and small" targets. These areas have clearly defined spatial boundaries and environmental constraints; for example, military restricted zones need to prevent external drone intrusion, and tethered drones are limited by cable length, forming rectangular operational airspaces. These scenarios place special demands on the collaborative control and decision-making capabilities of drone swarms.

[0003] Traditional drone interception methods mainly include: interception methods based on manual control, interception methods based on expert experience, and interception methods based on optimization theories such as genetic algorithms. Although the above methods have certain effects in specific scenarios, they generally have the following problems in the strongly constrained environment of a rectangular finite flight area: (1) The decision network is difficult to converge stably: there are mutual influences between multiple agents, resulting in non-stationary environmental states, large oscillations in the training process, and difficulty in convergence. (2) The reward function is difficult to effectively learn constraints: traditional reinforcement learning often uses reward function approximation constraints, which easily leads to frequent violations of spatial constraints in the real environment. (3) The swarm collaboration effect is poor: there is a lack of effective constraints on the overall behavior of the swarm, and the abnormal decision-making of individual drones will seriously affect the overall interception efficiency.

[0004] With the increasing scale and complexity of drone swarms, there is a pressing need for an intelligent decision-making method capable of efficient, stable, and collaborative capture under strong spatial constraints. In recent years, the combination of reinforcement learning and game theory has provided new insights into multi-agent collaborative decision-making, demonstrating significant advantages, especially in dynamic and adversarial environments. However, existing methods still have significant limitations in dealing with high-dimensional state inputs, hard spatial constraints, and multi-drone collaborative game theory within a finite rectangular region. Therefore, it is necessary to develop a collaborative decision-making method that integrates differential game theory and reinforcement learning frameworks, enabling rapid, stable, and collaborative capture of unauthorized drones within a limited flight area, thereby improving the overall interception success rate and environmental adaptability of the system. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention proposes a rectangular area drone encirclement method based on differential game theory and reinforcement learning. By proposing a method that combines differential game theory principles, a rectangular flight constraint loss function, and batch temporal behavior constraints, this invention solves problems such as poor overall group strategy and susceptibility to decision deviation caused by minor environmental disturbances.

[0006] The first aspect of this invention provides a method for drone encirclement and capture in a rectangular area based on differential game theory and reinforcement learning, comprising the following steps:

[0007] S1: Our drone cluster acquires environmental status information, which includes the position and speed of both friendly and enemy drones, as well as their historical status and action sequence over the previous t time periods;

[0008] S2: Input the environmental state information into the rectangular flight area constraint loss function module to calculate the spatial constraint loss caused by the UAV deviating from the preset rectangular area; input the historical state and action sequence into the batch temporal behavior constraint module to calculate the temporal consistency loss used to ensure decision smoothness; combine the spatial constraint loss and the temporal consistency loss to generate a joint latent variable of state behavior that integrates spatial and temporal constraints;

[0009] S3: The joint latent variables of the state behavior are fused with the latest environmental state information and input into the differential game value evaluation module; by calculating the differential game between our drone cluster and the black drone, the expected future value based on Nash equilibrium is obtained, and the decision network parameters of our drone cluster are updated through reinforcement learning.

[0010] S4: Our drone swarm executes the updated decision network output of the capture action and stores the current experience data in the experience replay buffer for continuous model training and optimization.

[0011] In this solution, environmental status information perception and data preprocessing include:

[0012] Each of our drones collects environmental data through onboard sensors, identifies all detected drone targets as friend or foe, and classifies them according to pre-assigned electronic identification tags to distinguish between unauthorized drones and our drones.

[0013] For the identified unauthorized drones and our own drones, we extract their three-dimensional position coordinates, three-dimensional velocity vectors, and timestamps to form a real-time state vector.

[0014] During operation, the state-action tuple at each decision moment is stored in the database of the experience replay buffer. Each tuple contains: the state at time t, the action executed by the cluster at time t, the reward obtained after executing the action, and the new state after the action is executed.

[0015] At the current decision-making moment, extract all state-action tuples from the historical buffer in chronological order for the previous t moments, encapsulate all extracted state-action tuples to form the batch chronological historical state and action sequence required for decision-making;

[0016] The real-time state vector is integrated with the batch time-series historical state and action sequence to form environmental state information.

[0017] In this solution, the environmental state information is input into the rectangular flight area constraint loss function module to calculate the spatial constraint loss caused by the UAV deviating from the preset rectangular area, including:

[0018] The boundaries of the rectangular flight area are preset according to the mission requirements, and the coordinates of the geometric center point of the rectangular flight area are obtained as the position reference point.

[0019] For each of our drones at the current moment, read the real-time position coordinates and determine whether the real-time position coordinates fall within a preset rectangular flight area. If it is within the preset rectangular flight area, return the coordinates of the geometric center point of the rectangular flight area; otherwise, return the current real-time position coordinates of our drone.

[0020] The returned value is compared with the location reference point, and the difference is calculated using the mean squared error loss function to obtain the spatial constraint loss caused by the UAV deviating from the preset rectangular area. The spatial constraint loss is expressed as:

[0021] ,

[0022] in Let cross-entropy be the loss function. A function to determine whether the current real-time position of the drone is within the rectangular flight area. The desired position.

[0023] In this scheme, the historical states and action sequences are input into the batch temporal behavior constraint module to calculate the temporal consistency loss used to ensure decision smoothness, including:

[0024] Encode the batch of temporal historical states and action sequences from the historical buffer to generate joint latent variables of historical state behavior, and use the front-end network to generate preliminary joint latent variables of current state behavior at the current moment based on environmental state information.

[0025] The difference in behavioral trajectory between the joint latent variables of historical state behavior and the joint latent variables of current state behavior is calculated using a distance function, and the temporal consistency cost is calculated based on the difference in behavioral trajectory.

[0026] A stochastic constraint is introduced, and the stochastic constraint is added to the time-series consistency cost to obtain a comprehensive penalty term. This comprehensive penalty term is then transformed into a consistency score, and a time-series consistency loss is constructed based on the consistency score. The time-series consistency loss is expressed as:

[0027] ,

[0028] in For historical state behavior joint latent variables, For the current state behavior joint latent variables, To balance the weighting factors, The standard deviation parameter, This is a hyperparameter.

[0029] In this scheme, spatial constraint loss and temporal consistency loss are combined to generate joint latent variables of state behavior that integrate spatial and temporal constraints, including:

[0030] The spatial constraint loss and the temporal consistency loss are algebraically added to generate the total constraint loss. The total constraint loss is used as the optimization objective, and the parameters of the front-end network are updated through gradient backpropagation. The front-end network is iteratively trained with the goal of minimizing the total constraint loss.

[0031] Batch temporal historical states and action sequences are imported into the trained front-end network to generate joint latent variables of state and behavior that integrate spatial and temporal constraints.

[0032] In this scheme, the expected future value based on Nash equilibrium is obtained by calculating the differential game between our drone swarm and the unauthorized drones, including:

[0033] The state-behavior joint latent variables, which integrate spatial and temporal constraints, are weighted and fused with the latest environmental state information obtained from sensors to form an enhanced state vector. This vector is then loaded with the state-action pair from the previous iteration. Valuation of value ;

[0034] The differential game value evaluation module is used to model the current adversarial situation as a differential game, with participants including our drone swarm and unauthorized drones, and the payoff function of each participant is defined.

[0035] Based on the aforementioned payoff function, the future value calculated using Nash equilibrium is solved through gradient ascent in differential calculus, by solving the differential game equations. Received, among which For payment functions, Let i be the state of the i-th agent. To achieve the Nash equilibrium state of the i-th agent;

[0036] The immediate reward is added to the Nash equilibrium future value after discount factor decay to calculate the target Q value. A smooth update strategy is then used to update the value function by weighted averaging the previous Q value with the calculated target Q value, expressed as:

[0037] ,

[0038] in Let be the updated value function of the j-th agent taking action a in state s. To balance the weighting factors, Let be the instantaneous reward obtained by the j-th agent at time t. For future value calculated based on Nash equilibrium, This is the discount factor.

[0039] In this scheme, the updated value signal is used as a benchmark to feed back to the decision network of our drone swarm. Using the Bellman optimality equation principle in reinforcement learning, the internal policy parameters are adjusted through gradient escalation and other methods with the goal of maximizing value.

[0040] In this scheme, our drone swarm executes the updated encirclement actions output by the decision network and stores the current experience data in the experience replay buffer for continuous model training and optimization, including:

[0041] Each drone in our drone swarm receives action instructions from the updated decision network, interprets the received action instructions into specific control signals, and each drone executes flight actions according to the control signals.

[0042] After executing an action, the drone swarm acquires the environmental state information of the next moment through sensors, calculates the immediate reward obtained for this action according to the preset reward rules, encapsulates the complete cycle data of this decision and the execution of the flight action into a state-action tuple, and stores the newly generated state-action tuple in the experience replay buffer.

[0043] Randomly sample historical state-action tuples from the experience replay buffer, calculate spatial constraint loss and temporal consistency loss, update the Q-value function in the differential game value evaluation module, and use the updated Q-value to update the parameters of the decision network through gradient descent.

[0044] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0045] This invention proposes a multi-drone encirclement method based on differential game theory learning. By optimizing the overall workflow, this invention can better utilize game theory principles for decision-making, resulting in faster execution compared to traditional linear programming-based methods. The introduction of a differential game value evaluation module allows for better prediction of future rewards in dynamic game environments, enabling the decision network to more closely approximate global judgments and thus facilitating training. Furthermore, the introduction of batch temporal behavior constraints allows for state memorization over a fixed time period, resulting in more stable cluster decisions and superior performance. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments or examples of the present invention, the drawings used in the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained according to these drawings without creative effort.

[0047] Figure 1 A flowchart of a method for drone encirclement and capture in a rectangular area based on differential game theory and reinforcement learning is shown.

[0048] Figure 2 The flowchart of the differential game value evaluation module is shown;

[0049] Figure 3 A flowchart illustrating the closed-loop optimization of drone action execution and experience playback is shown. Detailed Implementation

[0050] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0051] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0052] like Figure 1 As shown, this embodiment provides a method for drone encirclement in a rectangular area based on differential game theory and reinforcement learning, including:

[0053] S1: Our drone cluster acquires environmental status information, which includes the position and speed of both friendly and enemy drones, as well as their historical status and action sequence over the previous t time periods;

[0054] S2: Input the environmental state information into the rectangular flight area constraint loss function module to calculate the spatial constraint loss caused by the UAV deviating from the preset rectangular area; input the historical state and action sequence into the batch temporal behavior constraint module to calculate the temporal consistency loss used to ensure decision smoothness; combine the spatial constraint loss and the temporal consistency loss to generate a joint latent variable of state behavior that integrates spatial and temporal constraints;

[0055] S3: The joint latent variables of the state behavior are fused with the latest environmental state information and input into the differential game value evaluation module; by calculating the differential game between our drone cluster and the black drone, the expected future value based on Nash equilibrium is obtained, and the decision network parameters of our drone cluster are updated through reinforcement learning.

[0056] S4: Our drone swarm executes the updated decision network output of the capture action and stores the current experience data in the experience replay buffer for continuous model training and optimization.

[0057] It should be noted that each of our drones collects environmental data through onboard sensors, including but not limited to GPS modules, inertial measurement units, visual cameras, lidar, and communication links. All detected drone targets undergo friend-or-foe identification, classifying their shapes based on pre-assigned electronic identification tags or visual recognition models to distinguish between unauthorized drones and our own. For identified unauthorized and our drones, three-dimensional position coordinates, three-dimensional velocity vectors, and timestamps are extracted to form a real-time state vector. During operation, the state-action tuples at each decision moment are stored in the database of the experience replay buffer. Each tuple contains: the state at time t, the action executed by the cluster at time t, the reward obtained after executing the action, and the new state after the action. At the current decision moment, all state-action tuples from the previous t moments are extracted from the historical buffer in chronological order. All extracted state-action tuples are encapsulated to form the batch-series historical state and action sequence required for decision-making. The real-time state vector and the batch-series historical state and action sequence are integrated to form environmental state information.

[0058] It should be noted that the rectangular flight area constraint loss function module calculates the spatial constraint loss caused by the UAV deviating from the preset rectangular area, thereby guiding the UAV swarm to move within the specified rectangular area. The boundary of the preset rectangular flight area is determined according to mission requirements, typically by defining the coordinates of the four vertices of the rectangular area or by providing the center point, length, width, and height of the area. The geometric center point coordinates of the rectangular flight area are obtained as a position reference point. For each friendly UAV at the current moment, its real-time position coordinates are read, and it is determined whether the real-time position coordinates fall within the preset rectangular flight area. If it is within the preset rectangular flight area, the geometric center point coordinates of the rectangular flight area are returned; otherwise, the current real-time position coordinates of the friendly UAV are returned. The returned value is compared with the position reference point, and the difference is calculated using the mean squared error loss function. The mean squared error loss amplifies large position deviations, resulting in a larger loss value as the UAV deviates further from the center point. The spatial constraint loss caused by the UAV deviating from the preset rectangular area is obtained, and the spatial constraint loss is expressed as:

[0059] ,

[0060] in Let cross-entropy be the loss function. A function to determine whether the current real-time position of the drone is within the rectangular flight area. The desired position.

[0061] When the drone is within the pre-defined rectangular flight area, the return value is the center point of the given area, with zero difference from the target center point, resulting in zero loss and no negative impact on training. When the drone is outside the pre-defined rectangular flight area, the return value is its own position, with a large difference from the target center point, resulting in negative loss. In gradient descent, the optimizer will try to avoid this huge penalty, thereby driving the drone's position to move within the area.

[0062] It should be noted that the batch temporal behavior constraint module calculates the temporal consistency loss used to ensure the smoothness of decision-making. By introducing historical decision information, it reduces the sudden changes in decision-making caused by instantaneous environmental disturbances, ensuring that the actions of our drone swarm are smooth, coherent and predictable in the time dimension, thereby improving the stability of swarm collaboration.

[0063] Batch time-series historical states and action sequences from the historical buffer are encoded to generate joint latent variables of historical state and behavior. These joint latent variables encapsulate the cluster's past decision-making style and situational evolution information. The front-end network generates preliminary joint latent variables of current state and behavior based on environmental state information. These current state and behavior joint latent variables represent the network's potential immediate decision-making tendency without historical constraints. A distance function is used to quantify the behavioral trajectory differences between the joint latent variables of historical state and behavior and the joint latent variables of current state and behavior. The time-series consistency cost is calculated based on these behavioral trajectory differences; a larger distance indicates a greater deviation between the current decision and past behavioral patterns. A stochastic constraint is introduced to allow the strategy to explore and innovate appropriately when necessary, preventing it from falling into overly conservative local optima. The stochastic constraint is added to the time-series consistency cost to obtain a comprehensive penalty term. The comprehensive penalty term is transformed into a consistency score. The larger the penalty, the lower the score. A temporal consistency loss is constructed based on the consistency score, and the temporal consistency loss is expressed as:

[0064] ,

[0065] in For historical state behavior joint latent variables, For the current state behavior joint latent variables, To balance the weighting factors, is the standard deviation parameter, and is the hyperparameter.

[0066] By minimizing the loss of temporal consistency, decisions that align with historical behavior patterns are encouraged, while decisions that deviate from historical trends are penalized. This effectively filters out transient noise in the environment, making the decisions and behaviors of the entire drone swarm more consistent and stable over time, and effectively reducing deviations in decision outcomes.

[0067] The rectangular flight area constraint loss function module and the batch temporal behavior constraint module operate independently, calculating the spatial constraint loss caused by deviations from the preset rectangular area and the temporal consistency loss used to ensure decision smoothness for the drones in the cluster. The front-end network receives raw, unconstrained batch temporal historical states and action sequences as input, processes the input data, and initially generates initial joint latent variables of state and behavior that do not yet contain explicit spatial and temporal constraint information. The spatial constraint loss and temporal consistency loss are algebraically added to generate the total constraint loss. Using this total constraint loss as the optimization objective, the parameters of the front-end network are updated through gradient backpropagation. The front-end network is iteratively trained with the goal of minimizing the total constraint loss. The parameters of the front-end network are continuously adjusted, forcing changes in the internal computational logic, enabling the model to learn to consider both spatial regularity and temporal smoothness when generating latent variables. The batch temporal historical states and action sequences are then imported into the trained front-end network to generate joint latent variables of state and behavior that integrate spatial and temporal constraints.

[0068] It should be noted that the differential game value evaluation module goes beyond traditional unilateral optimization. In a dynamic adversarial environment, it calculates a collaborative and balanced optimal strategy for our drone swarm. The internal working steps of the differential game value evaluation module are as follows: Figure 2 As shown, the joint latent variables of state-behavior with spatial and temporal constraints are weighted and fused with the latest environmental state information obtained from sensors to form an enhanced state vector. This enhanced state vector integrates the current environmental situation and the joint latent variable information refined from spatial and temporal constraints. The module loads the state-action pair based on the previous iteration. Valuation of value .

[0069] The differential game value evaluation module models the current adversarial situation as a differential game, with participants including our drone swarm and unauthorized drones, each possessing an independent strategy set and objective, and defining the payoff function for each participant. Nash equilibrium is a strategically stable point; in this state, neither party can gain additional benefit by unilaterally changing its strategy. Based on the payoff function, the future value calculated from Nash equilibrium is solved using gradient ascent in differential calculus, thereby solving the differential game equations. Received, among which For payment functions, Let i be the state of the i-th agent. To achieve the Nash equilibrium state for the i-th agent; in this state, no agent can obtain a higher payoff by unilaterally changing their own state.

[0070] The target Q value is calculated by adding the immediate reward to the future Nash equilibrium value after discount factor decay. A smooth update strategy is adopted, which calculates a weighted average of the Q-value from the previous time step and the calculated target Q-value, with the weights determined by a balancing weight factor. Control, to obtain the updated value function The update process allows value estimation to retain past experience while incorporating new insights based on game theory. The updated value function is expressed as:

[0071] ,

[0072] in Let be the updated value function of the j-th agent taking action a in state s. To balance the weighting factors, Let be the instantaneous reward obtained by the j-th agent at time t. Let be the future value calculated by the j-th agent based on Nash equilibrium, representing the optimal future gain that the agent can obtain. This is the discount factor.

[0073] The updated value signal is fed back to the decision network of our drone swarm as a benchmark. Using the Bellman optimality equation principle in reinforcement learning, the internal policy parameters are adjusted through gradient descent and other methods with the goal of maximizing value.

[0074] It should be noted that this involves translating intelligent decision-making into physical actions and collecting data to drive continuous learning and optimization of the system. For example... Figure 3As shown, each UAV in our UAV swarm receives action commands from the updated decision network and interprets these commands into specific control signals, including precise adjustments to throttle, control surfaces, and motor speed. Each UAV executes flight maneuvers based on these control signals, collectively achieving tactical intentions such as pursuit, interception, and encirclement within a rectangular flight area under strict constraints. After executing a maneuver, the UAV swarm acquires environmental state information for the next moment through sensors and calculates the immediate reward for the action according to preset reward rules. For example, a positive reward is given for successfully approaching the target, a negative reward for flying out of the area, and a small negative reward for consuming energy. The complete cycle data of this decision and flight maneuver is encapsulated into state-action tuples, and the newly generated state-action tuples are stored in the experience replay buffer. During the training phase of the decision network and the differential game value evaluation module, historical state-action tuples are randomly sampled from the experience replay buffer, spatial constraint loss and temporal consistency loss are calculated, the Q-value function in the differential game value evaluation module is updated, and the parameters of the decision network are updated using the updated Q-value through gradient descent. Through repeated practice within a rectangular area and continuous fine-tuning of decision-making strategies, the task of capturing unauthorized drones was eventually completed skillfully and efficiently.

[0075] The second embodiment of the present invention provides a computer-readable storage medium, which includes a program for a method of capturing drones in a rectangular area based on differential game theory and reinforcement learning. When the program for capturing drones in a rectangular area based on differential game theory and reinforcement learning is executed by a processor, it implements the steps of the method for capturing drones in a rectangular area based on differential game theory and reinforcement learning.

[0076] In the several embodiments provided in this application, it should be understood that the disclosed methods and systems can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, and can be electrical, mechanical, or other forms. Furthermore, in the various embodiments of the present invention, all functional units can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0077] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0078] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for drone encirclement and capture in a rectangular area based on differential game theory and reinforcement learning, characterized in that, Includes the following steps: S1: Our drone cluster acquires environmental status information, which includes the position and speed of both friendly and enemy drones, as well as their historical status and action sequence over the previous t time periods; S2: Input the environmental state information into the rectangular flight area constraint loss function module to calculate the spatial constraint loss caused by the UAV deviating from the preset rectangular area; input the historical state and action sequence into the batch temporal behavior constraint module to calculate the temporal consistency loss used to ensure decision smoothness; combine the spatial constraint loss and the temporal consistency loss to generate a joint latent variable of state behavior that integrates spatial and temporal constraints; S3: The joint latent variables of the state behavior are fused with the latest environmental state information and input into the differential game value evaluation module; by calculating the differential game between our drone cluster and the black drone, the expected future value based on Nash equilibrium is obtained, and the decision network parameters of our drone cluster are updated through reinforcement learning. S4: Our drone swarm executes the updated decision network output of the capture action and stores the current experience data in the experience replay buffer for continuous training and optimization of the model; By calculating the differential game between our drone swarm and the unauthorized drones, we obtain the expected future value based on Nash equilibrium, including: The state-behavior joint latent variables, which integrate spatial and temporal constraints, are weighted and fused with the latest environmental state information obtained from sensors to form an enhanced state vector. This vector is then loaded with the state-action pair from the previous iteration. Valuation of value ; The differential game value evaluation module is used to model the current adversarial situation as a differential game, with participants including our drone swarm and unauthorized drones, and the payoff function of each participant is defined. Based on the aforementioned payoff function, the future value calculated using Nash equilibrium is solved through gradient ascent in differential calculus, by solving the differential game equations. We obtain, where is the payment function. Let i be the state of the i-th agent. To achieve the Nash equilibrium state of the i-th agent; The immediate reward is added to the Nash equilibrium future value after discount factor decay to calculate the target Q value. A smooth update strategy is then used to update the value function by weighted averaging the previous Q value with the calculated target Q value, expressed as: , in Let be the updated value function of the j-th agent taking action a in state s. To balance the weighting factors, Let be the instantaneous reward obtained by the j-th agent at time t. For future value calculated based on Nash equilibrium, This is the discount factor.

2. The method for UAV encirclement and capture in a rectangular area based on differential game theory and reinforcement learning according to claim 1, characterized in that, Environmental status information perception and data preprocessing, including: Each of our drones collects environmental data through onboard sensors, identifies all detected drone targets as friend or foe, and classifies them according to pre-assigned electronic identification tags to distinguish between unauthorized drones and our drones. For the identified unauthorized drones and our own drones, we extract their three-dimensional position coordinates, three-dimensional velocity vectors, and timestamps to form a real-time state vector. During operation, the state-action tuple at each decision moment is stored in the database of the experience replay buffer. Each tuple contains: the state at time t, the action executed by the cluster at time t, the reward obtained after executing the action, and the new state after the action is executed. At the current decision-making moment, extract all state-action tuples from the historical buffer in chronological order for the previous t moments, encapsulate all extracted state-action tuples to form the batch chronological historical state and action sequence required for decision-making; The real-time state vector is integrated with the batch time-series historical state and action sequence to form environmental state information.

3. The rectangular area drone encirclement method based on differential game theory and reinforcement learning according to claim 1, characterized in that, The environmental state information is input into the rectangular flight area constraint loss function module to calculate the spatial constraint loss caused by the UAV deviating from the preset rectangular area, including: The boundaries of the rectangular flight area are preset according to the mission requirements, and the coordinates of the geometric center point of the rectangular flight area are obtained as the position reference point. For each of our drones at the current moment, read the real-time position coordinates and determine whether the real-time position coordinates fall within a preset rectangular flight area. If it is within the preset rectangular flight area, return the coordinates of the geometric center point of the rectangular flight area; otherwise, return the current real-time position coordinates of our drone. The returned value is compared with the location reference point, and the difference is calculated using the mean squared error loss function to obtain the spatial constraint loss caused by the UAV deviating from the preset rectangular area. The spatial constraint loss is expressed as: , in Let cross-entropy be the loss function. A function to determine whether the current real-time position of the drone is within the rectangular flight area. The desired position.

4. The method for UAV encirclement and capture in a rectangular area based on differential game theory and reinforcement learning according to claim 1, characterized in that, The historical states and action sequences are input into the batch temporal behavior constraint module to calculate the temporal consistency loss used to ensure decision smoothness, including: Encode the batch of temporal historical states and action sequences from the historical buffer to generate joint latent variables of historical state behavior, and use the front-end network to generate preliminary joint latent variables of current state behavior at the current moment based on environmental state information. The difference in behavioral trajectory between the joint latent variables of historical state behavior and the joint latent variables of current state behavior is calculated using a distance function, and the temporal consistency cost is calculated based on the difference in behavioral trajectory. A stochastic constraint is introduced, and the stochastic constraint is added to the time-series consistency cost to obtain a comprehensive penalty term. This comprehensive penalty term is then transformed into a consistency score, and a time-series consistency loss is constructed based on the consistency score. The time-series consistency loss is expressed as: , in For historical state behavior joint latent variables, For the current state behavior joint latent variables, To balance the weighting factors, The standard deviation parameter, This is a hyperparameter.

5. The method for UAV encirclement and capture in a rectangular area based on differential game theory and reinforcement learning according to claim 4, characterized in that, By combining spatial constraint loss with temporal consistency loss, joint latent variables of state behavior that integrate spatial and temporal constraints are generated, including: The spatial constraint loss and the temporal consistency loss are algebraically added to generate the total constraint loss. The total constraint loss is used as the optimization objective, and the parameters of the front-end network are updated through gradient backpropagation. The front-end network is iteratively trained with the goal of minimizing the total constraint loss. Batch temporal historical states and action sequences are imported into the trained front-end network to generate joint latent variables of state and behavior that integrate spatial and temporal constraints.

6. The method for UAV encirclement and capture in a rectangular area based on differential game theory and reinforcement learning according to claim 1, characterized in that, The updated value signal is fed back to the decision network of our drone swarm as a benchmark. Using the Bellman optimality equation principle in reinforcement learning, the internal policy parameters are adjusted through gradient descent and other methods with the goal of maximizing value.

7. The method for UAV encirclement and capture in a rectangular area based on differential game theory and reinforcement learning according to claim 1, characterized in that, Our drone swarm executes the updated decision network output of the encirclement action and stores the current experience data in the experience replay buffer for continuous model training and optimization, including: Each drone in our drone swarm receives action instructions from the updated decision network, interprets the received action instructions into specific control signals, and each drone executes flight actions according to the control signals. After executing an action, the drone swarm acquires the environmental state information of the next moment through sensors, calculates the immediate reward obtained for this action according to the preset reward rules, encapsulates the complete cycle data of this decision and the execution of the flight action into a state-action tuple, and stores the newly generated state-action tuple in the experience replay buffer. Randomly sample historical state-action tuples from the experience replay buffer, calculate spatial constraint loss and temporal consistency loss, update the Q-value function in the differential game value evaluation module, and use the updated Q-value to update the parameters of the decision network through gradient descent.