A method for unmanned aerial vehicle confrontation scheduling under incomplete observation information
By using online learning and batch processing techniques, combined with local observation data and feedback estimation, a UAV scheduling strategy is generated. This solves the problems of UAV scheduling methods' dependence on enemy behavior models and high resource consumption, and achieves efficient autonomous scheduling and resource conservation under incomplete observation information.
Patent Information
- Application Number
- CN202511847095.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-12-09
AI Technical Summary
Existing drone scheduling methods rely on strong assumptions about enemy behavior models and idealized observations of global battlefield information. This makes it difficult to effectively defend against enemy attacks and control scheduling frequency under incomplete observation information, resulting in high resource consumption.
By employing online learning and batch processing techniques, combined with local observation data and feedback estimation, a drone scheduling strategy is generated. Through online learning, the strategy adapts to changes in enemy strategies, reduces scheduling frequency, and optimizes defense strategies.
In situations where enemy behavior is unknown and observation information is incomplete, efficient autonomous scheduling can be achieved, increasing the cost of enemy attacks and reducing the amount of UAVs to be scheduled, thereby improving the system's robustness and resource utilization efficiency in dynamic environments.
Smart Images

Figure CN121277229B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unmanned aerial vehicle (UAV) defense and autonomous decision-making, and in particular to a UAV adversarial scheduling method under incomplete observation information. Background Technology
[0002] With the popularization of drone technology, its application in security fields such as border patrol and key area defense is becoming increasingly widespread. In typical drone attack and defense scenarios, we face a dual challenge: First, due to limitations in the number of drones and sensor energy efficiency, we can only observe local feedback information from the deployment area of defensive drones, leading to high decision-making uncertainty; second, drone endurance and communication resources are limited, and frequent scheduling updates will cause significant resource consumption. Against this backdrop, we need to effectively defend against enemy attacks and increase their attack costs based on incomplete observation information, while simultaneously controlling the resource consumption caused by frequent adjustments to scheduling strategies.
[0003] Most existing drone scheduling methods rely on static scheduling rules or model derivations based on certain prior information. However, these methods generally suffer from the following problems: on the one hand, they typically assume that our side can fully grasp the enemy's behavior or battlefield state, ignoring the adversarial nature of the enemy's strategies and the incompleteness of our observation information in the real environment; on the other hand, traditional methods often struggle to balance "defense effectiveness" and "scheduling frequency control," limiting the algorithm's adaptability and practicality in dynamic environments. Therefore, there is an urgent need for a drone scheduling method capable of making optimal decisions under unknown enemy strategies and incomplete observation information, effectively increasing the enemy's attack cost and reducing the rescheduling of our drones. Summary of the Invention
[0004] To address the problems of existing UAV adversarial scheduling methods, such as strong reliance on enemy behavior models, idealized observation requirements for global battlefield information, and low resource efficiency due to neglecting scheduling frequency control, this invention provides a UAV adversarial scheduling method under incomplete observation information. This method integrates online learning and batch processing techniques, dividing the adversarial process into multiple consecutive batches. Within each batch, the scheduling strategy for friendly UAVs is fixed to reduce scheduling frequency. A perturbation decision-making mechanism based on historical experience and a uniform exploration mechanism are combined to generate defense strategies for each batch, balancing utilization and exploration. Based on local observation data, feedback estimation is used to construct a global decision-making basis, achieving continuous strategy optimization. This invention enables efficient autonomous scheduling of friendly UAV swarms under conditions of unknown enemy behavior patterns and incomplete observation information, effectively increasing the enemy's attack cost and reducing the scheduling load of friendly UAVs.
[0005] The technical solution of this invention is: a UAV adversarial scheduling method under incomplete observation information, comprising the following steps:
[0006] Step S1: Initialize UAV countermeasures scheduling parameters; specifically including: system and environmental parameter settings: defense zone set. Number of our drones Number of enemy drones ,in Rescheduling cost coefficient Total number of rounds of confrontation Parameter settings: Exploration rate Disturbance parameters Geometric resampling times Switching time ;
[0007] Step S2: Based on the switching duration Calculate the total number of batches ,in Indicates rounding up, then the... Batch corresponding to round ;
[0008] Step S3: Generate from indivual Exploration strategy set composed of dimensional vectors And vector The Each component being 1 indicates a region. Deployed defensive drones;
[0009] Step S4: Initialize a cumulative strategy attack cost estimate for each defense zone. Combine all the initial cumulative strategy attack cost estimates into a single... Dimensional cumulative strategy attack cost estimation vector , where the superscript " " represents the transpose of a vector. This symbol will be used in the following text. Each batch will execute subsequent steps to generate the drone scheduling strategy for that batch.
[0010] Step S5: In the first Batch, based on parameters The exponential distribution is obtained by independent sampling. random variables ,composition Perturbation quantity ;
[0011] Step S6: Scheduling strategy for our drone swarm in this batch (i.e., in the round) Our drone swarm scheduling strategy Keep as The following is an example: based on the exploration rate. Explore the strategy set and randomly select one strategy from the exploration strategy set as the scheduling strategy for this batch of our drones; otherwise, use probability... Based on the cumulative strategy attack cost estimation vector With disturbance quantity The summation result is in our strategy space The strategy of selecting the one with the largest inner product of the summation results As the scheduling strategy for this batch of our drone fleet, namely ;
[0012] Step S7: According to the scheduling strategy The decision was made to deploy this batch of our drone fleet, among which The Each component If it is 1, then the first Deploy defensive drones in designated areas; if the value is 0, no defensive drones are deployed. Observe the cost vector of the enemy's strategic attack in this batch based on the enemy's actions. Partial components ( and This refers to the cost of our strategic attack strategy in the area where we deploy defense drones;
[0013] Step S8: Use geometric resampling estimation technique to estimate the current batch. The probability of selecting an area where defensive drones are actually deployed;
[0014] Step S9: Based on the actual scheduling strategy adopted by our drone swarm, the observed partial components of the enemy's strategic attack cost vector for this batch, and the simulated reciprocal of the probability. To update the cumulative strategy attack cost estimate Return to S5 to continue generating the next batch. The drone scheduling strategy, until The round of attack and defense has ended.
[0015] An electronic device includes: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method.
[0016] A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to implement the method described thereon.
[0017] A computer program product includes a computer program that, when executed by a processor, implements the method.
[0018] The advantages of this invention compared to the prior art are as follows:
[0019] (1) Based on an online learning framework, this invention can learn directly from the historical interactions with the enemy without relying on prior information such as enemy behavior models and utility functions, thus effectively dealing with intelligent opponents with unknown strategies and adversarial capabilities.
[0020] (2) The present invention integrates local observation and feedback estimation update mechanism, which can make effective decisions based on incomplete feedback information, breaks through the ideal dependence of traditional methods on the availability of global information, and significantly improves the robustness and practicality of the system in real environment.
[0021] (3) The present invention introduces batch processing technology, which reduces the scheduling frequency and saves communication and computing resources caused by frequent policy updates, thereby achieving a good balance between improving defense effectiveness and controlling scheduling overhead. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating the implementation of a UAV adversarial scheduling method under incomplete observation information as described in this invention.
[0023] Figure 2 This is a flowchart illustrating the implementation of the geometric resampling estimation algorithm described in this invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other. To achieve the above objectives, this invention adopts the following technical solution.
[0025] First, the UAV counter-scheduling problem solved by this invention is described. Both sides engage in... Round of confrontation, total We possess [number] defensive zones. The enemy can deploy drones in the same round. A drone, of which Only one attack and / or defense drone can be deployed in each area. Our drone swarm is deployed according to the scheduling strategy generated by this invention to intercept the intrusion actions of the attack drone swarm.
[0026] Our strategy is described as follows: Our drone swarm scheduling strategy space is denoted as... Each of these strategies It is A two-dimensional vector, that is, each component , used to indicate the deployment location of the drone; if and only if At that time, the first Defensive drones were deployed in several areas; this strategy satisfies... ,in The 1-norm of a vector is denoted as 1, i.e. (In the following text, symbols) This definition is used for all cases. In the... In this round, our actual drone swarm scheduling strategy is represented as follows: .
[0027] The enemy's strategy is described as follows: The enemy's drone swarm attack strategy space is denoted as... Each of its strategies It is also A two-dimensional vector that satisfies ; if and only if its weight At that time, the first Attack drones were deployed in several areas, in the... The enemy drone swarm attack strategy is represented as follows: And the enemy in the The generation of attack strategies does not depend on our strategies in previous rounds.
[0028] The cost of an enemy attack is described as follows: In each round of attack and defense, due to factors such as the varying distances of enemy drones from different defense zones, the cost of attacking enemy drones in different defense zones also varies. Let the inherent cost vector of the attack be denoted as... ,in This indicates that each component takes a value of of A set of dimensional vectors, whose components Indicates the enemy's attack area The inherent costs. In the first Round, given the enemy's attack strategy Then, the "strategy attack cost" vector generated by this strategy is defined as follows: ,in This represents the Hadamard product of vectors, i.e., when the enemy chooses to attack the defensive area. The inherent costs will only be incurred at this time. If defensive drones are deployed in the area at this time, enemy attacks will be ineffective and will incur attack costs. If no defensive drones are deployed in the area, the attack benefits offset the inherent costs, and the actual attack cost can be considered zero. For defensive areas not targeted by the enemy, regardless of whether defensive drones are deployed, the enemy incurs no attack cost. Therefore, for our side in the [missing information]... Wheel defense strategy The cost of the enemy's attack in this round can be described as ,in Represents the dot product of vectors.
[0029] There are two important features in the scene setting:
[0030] (1) Our side has no prior knowledge of the attack: Before each round of attack and defense begins, our side has no knowledge of the enemy's behavior pattern.
[0031] (2) Our observation capabilities are limited: After each round of attack and defense, we can only observe the enemy's strategic attack cost in the area where the defensive drones were deployed in that round.
[0032] In dynamic adversarial environments, frequent adjustments to defense strategies lead to significant resource consumption. To comprehensively evaluate strategy performance and scheduling efficiency, this invention introduces the "scheduling regret" metric, defined as the difference between the optimal fixed strategy with hindsight and the total cost (including attack cost and rescheduling amount) under the actual strategy sequence. This metric considers both the enemy's attack cost and our own scheduling cost, making it more in line with practical engineering needs. The mathematical expression for "scheduling regret" under round-robin attack and defense is as follows:
[0033] ;
[0034] in, This represents the rescheduling cost coefficient. This represents the amount of our drones rescheduled between two adjacent rounds of strategy, used to quantify scheduling overhead.
[0035] Based on the aforementioned scenario, this invention proposes a UAV adversarial scheduling method under incomplete observation information. The specific implementation process is described in [link to implementation details]. Figure 1 It includes the following steps:
[0036] Step S1: Initialize the basic parameters of the UAV countermeasures scheduling system; specifically including: system and environmental parameter settings: defense zone set. Number of our drones Number of enemy drones ,in Rescheduling cost coefficient Total number of rounds of confrontation Core parameter setting: Exploration rate Disturbance parameters Geometric resampling times Switching time Implementation details are as follows:
[0037] S11: Set the exploration rate ;
[0038] S12: Set disturbance parameters ;
[0039] S13: Set the geometric resampling count ;
[0040] S14: Set switch duration ,in Indicates rounding down;
[0041] Step S2: Based on the switching duration Calculate the total number of batches ,in Indicates rounding up, then the... Batch corresponding to round .
[0042] Step S3: Generate from indivual Exploration strategy set composed of dimensional vectors And vector The Each component being 1 indicates a region. Deployed defensive drones, with the remaining components being 0 or 1.
[0043] Step S4: Initialize a cumulative strategy attack cost estimate for each defense zone. Combine all the initial cumulative strategy attack cost estimates into a single... Dimensional cumulative strategy attack cost estimation vector Each batch of drones is processed through subsequent steps to generate the drone scheduling strategy for that batch.
[0044] Step S5: In the first Batch, based on parameters The exponential distribution is obtained by independent sampling. random variables ,composition Perturbation quantity .
[0045] Step S6: Scheduling strategy for our drone swarm in this batch (i.e., in the round) Our drone swarm scheduling strategy Keep as The following is an example: based on the exploration rate. Explore the strategy set and randomly select one strategy from the exploration strategy set as the scheduling strategy for this batch of our drones; otherwise, use probability... Based on the cumulative strategy attack cost estimation vector With disturbance quantity The summation result is in our strategy space The strategy of selecting the one with the largest inner product of the summation result. As the scheduling strategy for this batch of our drone fleet, namely .
[0046] Step S7: According to the scheduling strategy The decision was made to deploy this batch of our drone fleet, among which The Each component If it is 1, then the first Deploy defensive drones in designated areas; if the value is 0, no defensive drones are deployed. Observe the cost vector of the enemy's strategic attack in this batch based on the enemy's actions. Partial components ( and This refers to the cost of our strategy to attack areas where we deploy defense drones.
[0047] Step S8: Use geometric resampling estimation technique to estimate the current batch. The probability of selecting an area where defensive drones are actually deployed; see [link / reference]. Figure 2 The implementation details are as follows:
[0048] S81: For the areas where defensive drones were actually deployed in this batch. Initialize the hit counter Set the maximum number of resampling times ;
[0049] S82: Proceed arrive Resampling is performed in the next loop, and S83-S87 are executed in each loop.
[0050] S83: Based on the current cumulative strategy attack cost estimation vector, execute S5-S6 to generate a simulated scheduling strategy for our drone swarm. ;
[0051] S84: For each area where defensive drones are actually deployed in this batch. Execute S85-S86;
[0052] S85: If , The Each component and ,set up ;
[0053] S86: If , for those who still maintain Set the area ;
[0054] S87: If the areas where defensive drones are actually deployed in this batch All have been recorded as hits (i.e.) If the loop terminates prematurely, then the loop will exit early.
[0055] S88: From this, we obtain As a region In this batch Probability of deploying defensive drones The reciprocal of the estimate is truncated.
[0056] Step S9: Based on the actual scheduling strategy adopted by our drone swarm, the observed attack cost of some strategies in this batch, and the simulation results... To update the cumulative strategy attack cost estimate Its weight The specific update follows the formula below:
[0057] ;
[0058] in, Consistent with the definition above. Return to S5 to continue generating the next batch. The drone scheduling strategy, until The round of attack and defense has ended.
[0059] Theoretical analysis shows that, under the aforementioned scenario setting of no prior knowledge and incomplete observation information, the scheduling strategy generated by this invention has a strict mathematical upper bound guarantee on its performance. Specifically, the expected value of the scheduling regret is limited to a low order of magnitude:
[0060] ;
[0061] in, Represents the mathematical expectation. Denotes an asymptotic upper bound, that is, the existence of and irrelevant positive numbers Make That is, in The expected upper bound of scheduling regrets after each round increases sublinearly, which means that as... The increase in average scheduling regret The average performance of the strategy of this invention will approach 0, meaning that the average performance of the strategy converges to the optimal fixed strategy.
[0062] In summary, this invention provides a UAV adversarial scheduling method under incomplete observation information. The innovations of this method lie in: employing an online learning framework, which eliminates reliance on prior enemy knowledge and allows for autonomous learning and adaptation to adversarial attacks solely through historical interactions; combining local observation and feedback estimation mechanisms to overcome the dependence on global information in traditional methods, achieving effective decision-making under incomplete information; and introducing batch processing technology to reduce scheduling frequency, thereby conserving system communication and computing resources while maintaining defensive effectiveness. Ultimately, this method achieves a good balance between increasing the enemy's attack costs and reducing the number of UAVs scheduled by the user.
Claims
1. A method for UAV adversarial scheduling under incomplete observation information, characterized in that, Includes the following steps: Step S1: Initialize UAV countermeasures scheduling parameters; specifically including: system and environmental parameter settings: defense zone set. Number of our drones Number of enemy drones ,in Rescheduling cost coefficient Total number of rounds of confrontation Parameter settings: Exploration rate Disturbance parameters Geometric resampling times Switching time ; Step S2: Based on the switching duration Calculate the total number of batches ,in Indicates rounding up, then the... Batch corresponding to round ; Step S3: Generate from indivual Exploration strategy set composed of dimensional vectors And vector The Each component being 1 indicates a region. Deployed defensive drones; Step S4: Initialize a cumulative strategy attack cost estimate for each defense zone. Combine all the initial cumulative strategy attack cost estimates into a single... Dimensional cumulative strategy attack cost estimation vector The superscript " " represents the transpose of a vector, and each batch executes subsequent steps to generate the drone scheduling strategy for that batch; Step S5: In the first Batch, based on parameters The exponential distribution is obtained by independent sampling. random variables ,composition Perturbation quantity ; Step S6: Scheduling strategy for our drone swarm in this batch That is, in the round Our drone swarm scheduling strategy Keep as The following is based on the exploration rate: Explore the strategy set and randomly select one strategy from the exploration strategy set as the scheduling strategy for this batch of our drones; otherwise, use probability... Based on the cumulative strategy attack cost estimation vector With disturbance quantity The summation result is in our strategy space The strategy of selecting the one with the largest inner product of the summation results As the scheduling strategy for this batch of our drone fleet, namely ; Step S7: According to the scheduling strategy The decision was made to deploy this batch of our drone fleet, among which The Each component If it is 1, then the first Deploy defensive drones in designated areas; if the value is 0, no defensive drones are deployed. Observe the cost vector of the enemy's strategic attack in this batch based on the enemy's actions. Partial components ( and This refers to the cost of our strategy to attack areas where we deploy defense drones. Step S8: Use geometric resampling estimation technique to estimate the current batch. The probability of selecting an area where defensive drones are actually deployed; Step S9: Based on the actual scheduling strategy adopted by our drone swarm, the observed partial components of the enemy's strategic attack cost vector for this batch, and the simulated reciprocal of the probability. To update the cumulative strategy attack cost estimate Return to S5 to continue generating the next batch. The drone scheduling strategy, until The round of attack and defense has ended.
2. The UAV adversarial scheduling method under incomplete observation information according to claim 1, characterized in that, The specific problem of drone combat scheduling is as follows: both sides conduct... Round of confrontation, total We possess [number] defensive zones. The enemy deployed drones in the same round. A drone, of which Only one attack or defense drone can be deployed in each area. Our drone swarm will be deployed according to the generated scheduling strategy to intercept the intrusion of the attack drone swarm. Our strategy is described as follows: Our drone swarm scheduling strategy space is denoted as... Each of these strategies It is A two-dimensional vector, that is, each component , used to indicate the deployment location of the drone; if and only if At that time, the first Defensive drones were deployed in several areas; this strategy satisfies... ,in The 1-norm of a vector is denoted as 1, i.e. In the In this round, our drone swarm scheduling strategy is expressed as follows: ; The enemy's strategy is described as follows: The enemy's drone swarm attack strategy space is denoted as... Each of its strategies It is also A two-dimensional vector that satisfies ; if and only if its weight At that time, the first Attack drones were deployed in several areas; In the The enemy drone swarm attack strategy is represented as follows: And the enemy in the The generation of attack strategies does not depend on our strategies in previous rounds.
3. The UAV adversarial scheduling method under incomplete observation information according to claim 2, characterized in that, Step S1 includes: S11: Set the exploration rate ; S12: Set disturbance parameters ; S13: Set the geometric resampling count ; S14: Set switch duration ,in This indicates rounding down to the nearest integer.
4. The UAV adversarial scheduling method under incomplete observation information according to claim 3, characterized in that, Step S8 includes: S81: For the areas where defensive drones were actually deployed in this batch. Initialize the hit counter Set the maximum number of resampling times ; S82: Proceed arrive Resampling is performed in the next loop, and S83-S87 are executed in each loop. S83: Based on the current cumulative strategy attack cost estimation vector, execute S5-S6 to generate a simulated scheduling strategy for our drone swarm. ; S84: For each area where defensive drones are actually deployed in this batch. Execute S85-S86; S85: If , The Each component and ,set up ; S86: If , for those who still maintain Set the area ; S87: If the areas where defensive drones are actually deployed in this batch All have been recorded in the hit. If so, the loop will exit prematurely; S88: From this, we obtain As a region In this batch Probability of deploying defensive drones The reciprocal of the estimate is truncated.
5. The UAV adversarial scheduling method under incomplete observation information according to claim 4, characterized in that, In step S9, the cumulative strategy attack cost estimate is updated. Its weight The specific update follows the formula below: 。 6. The UAV adversarial scheduling method under incomplete observation information according to claim 2, characterized in that, Introduce a scheduling regret index. The mathematical expression for the scheduling regrets under round-robin attack and defense is as follows: ; This represents the amount of our drones rescheduled between two adjacent rounds of strategy.
7. The UAV adversarial scheduling method under incomplete observation information according to claim 6, characterized in that, The expected value of policy scheduling regret is limited to: ; in, Represents the mathematical expectation. Denotes an asymptotic upper bound, which exists and irrelevant positive numbers Make .
8. An electronic device, characterized in that, include: One or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method of any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, It stores executable instructions that, when executed by a processor, cause the processor to perform the method described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative confrontation method and system, and storage medium
CN114721424A
Multi-batch attack-oriented incomplete information dynamic game modeling method
CN118228490A