A method, device and storage medium for cooperative confrontation of a UAV cluster

CN117647994BActive Publication Date: 2026-09-29NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311635464.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2026-09-29
Estimated Expiration
2043-12-01

AI Technical Summary

Technical Problem

[0003]目前的无人机集群协同主要是通过传统的做法是假设任务区域的态势已知,在地图上标注防空火力阵地等威胁范围、地形、任务目标等态势信息,将地图输入到无人机中基于规则的方式进行协同,但面对复杂任务和多变环境时仍存在一定的局限性,导致协同效果不佳

Benefits of technology

[0034]本申请实施例第四方面,提供了一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现本申请实施例第一方面中的无人机集群的协同对抗方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117647994B_ABST
    Figure CN117647994B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for cooperative confrontation of a UAV cluster, an equipment and a storage medium, relates to the technical field of UAVs, and can improve the cooperative effect of the UAV cluster. The specific scheme comprises the following steps: acquiring global situation information of a confrontation scene, acquiring role information of the confrontation scene, the roles including reconnaissance roles, interference roles and attack roles; acquiring a state space of each role and a behavior space of each role; according to the role information, calling a first sub-model corresponding to a behavior type, inputting the state space of the role and the behavior space of the role into the corresponding first sub-model, and obtaining joint behavior information of each role; calling a second sub-model corresponding to the role according to the joint behavior information of the role, inputting the acquired state space information into the second sub-model, obtaining behavior information of each UAV in the role, and enabling each UAV to perform cooperative confrontation based on the behavior information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a method, apparatus, device, and storage medium for cooperative combat of UAV swarms. Background Technology

[0002] With the continuous development of technologies such as computers, automatic control, and robotics, unmanned systems, primarily new types of drones and unmanned vehicles, are playing an increasingly important role and are widely used in both military and civilian fields. However, individual unmanned systems are limited by their own power, functionality, and performance, making it impossible to complete complex tasks independently. Therefore, drone swarms are used to address this issue. Consequently, how drone swarms can cooperate effectively to complete tasks has become a crucial aspect.

[0003] Current drone swarm collaboration mainly relies on the traditional approach of assuming the situation in the mission area is known. This involves marking threat ranges such as air defense positions, terrain, and mission objectives on a map, and then inputting the map into the drones for rule-based collaboration. However, this approach still has limitations when facing complex missions and changing environments, resulting in poor collaboration performance. Summary of the Invention

[0004] This application provides a method, apparatus, device, and storage medium for cooperative countermeasures of drone swarms, which can improve the cooperative effect of drone swarms.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] In a first aspect, this application provides a method for cooperative countermeasures by a swarm of unmanned aerial vehicles (UAVs), the method comprising:

[0007] The global situation information of the adversarial scenario is obtained, which includes: the real-time positions of multiple first UAVs in the first party, the real-time positions of multiple second UAVs in the second party, the position of the radar in the first party, the position of the radar in the second party, and the position of the target in the second party. The first party and the second party are adversaries, and the target position of the second party is the attack position of the first party.

[0008] Obtain the role information of the adversarial scenario, including the number of roles and the drone set of each role, and the roles include reconnaissance roles, jamming roles and attack roles;

[0009] Obtain the state space and behavior space of each of the aforementioned roles. The behavior space includes empty behavior, movement behavior, and payload behavior. The payload behavior includes reconnaissance behavior, interference behavior, and attack behavior.

[0010] Based on the role information, the first sub-model corresponding to the behavior type is invoked, and the state space and behavior space of the role are input into the corresponding first sub-model to obtain the joint behavior information of each role;

[0011] The corresponding second sub-model is invoked based on the joint behavior information of the role, and the obtained state space information is input into the second sub-model to obtain the behavior information of each drone in the role. Each drone performs cooperative combat based on the behavior information. The state space information includes: the state information of the drone itself, the state information of the drone cluster of the role, environmental information, and task information.

[0012] In one embodiment, before acquiring the global situational information of the adversarial scenario, the method further includes:

[0013] Construct an initial adversarial model, which includes multiple initial first sub-models and multiple initial second sub-models.

[0014] In one embodiment, after constructing the initial adversarial model, the method further includes:

[0015] Acquire sample information of multiple movement behaviors, the sample information of the movement behaviors includes multiple movement sample state spaces and movement sample behavior spaces, the movement sample state space is a stacked first grid map, the first grid map includes the UAV's own position information, cluster position information, target position information, unmanned system domain information and global situation information, the movement sample behavior space includes empty behavior and movement behavior;

[0016] The initial second sub-model of the mobile behavior is iteratively trained using the mobile sample state space, the mobile sample behavior space, and the behavior information of the mobile role drone. Each iteration yields mobile joint behavior information, which is then evaluated using a preset first reward function to obtain a first evaluation value. This process continues until the first evaluation value meets a second preset threshold, at which point the first sub-model of the mobile behavior is obtained.

[0017] In one embodiment, after constructing the initial adversarial model, the method further includes:

[0018] Acquire sample information of multiple reconnaissance behaviors. The sample information of the reconnaissance behaviors includes multiple reconnaissance sample state spaces and reconnaissance sample behavior spaces. The reconnaissance sample state space includes a stacked second grid map, which includes the UAV's own position information, reconnaissance role cluster position information, unmanned system domain information, global situation information, and reconnaissance situation information. The reconnaissance sample behavior space includes aerial behavior and movement behavior.

[0019] The initial second sub-model of the reconnaissance behavior is iteratively trained using the reconnaissance sample state space, the reconnaissance sample behavior space, and the behavior information of the reconnaissance role UAV. Each iteration yields the corresponding reconnaissance joint behavior information. The mobile joint behavior information is evaluated using a preset second reward function to obtain a second evaluation value. The process continues until the second evaluation value meets a second preset threshold, at which point the first sub-model of the reconnaissance behavior is obtained.

[0020] In one embodiment, after constructing the initial adversarial model, the method further includes:

[0021] The system acquires sample information of multiple interference behaviors, which includes multiple interference sample state spaces and interference sample behavior spaces. The interference sample state space includes a stacked third grid diagram, which includes the UAV system's own state information, interference role cluster state information, UAV system domain information, global situation information, interference situation information, and attack role cluster location information. The interference sample behavior space includes empty behavior, movement behavior, and interference behavior.

[0022] The initial second sub-model of the interference behavior is iteratively trained using the state space of the interference sample, the behavior space of the interference sample, and the behavior information of the UAV of the interference role. Each iteration yields the corresponding joint interference behavior information. The joint interference behavior information is evaluated using a preset third reward function to obtain a third evaluation value. The process continues until the third evaluation value meets a third preset threshold, at which point the first sub-model of the interference behavior is obtained.

[0023] In one embodiment, after constructing the initial adversarial model, the method further includes:

[0024] Acquire sample information of multiple attack behaviors. The sample information of the attack behaviors includes multiple attack sample state spaces and attack sample behavior spaces. The attack sample state space includes a stacked fourth grid diagram. The fourth grid diagram includes the state information of the UAV system, the state information of the attack role cluster, the domain information of the UAV system, the global situation information, and the interference situation information. The attack sample behavior space includes empty behavior, movement behavior, and attack behavior.

[0025] The initial second sub-model of the attack behavior is iteratively trained using the attack sample space, the attack sample behavior space, and the behavior information of the drone of the attacking role. Each iteration yields the corresponding joint attack behavior information. The joint attack behavior information is evaluated using a preset fourth reward function to obtain a fourth evaluation value. The process continues until the fourth evaluation value meets a fifth preset threshold, at which point the first sub-model of the attack behavior is obtained.

[0026] In one embodiment, the first sub-model is a model based on the MADDPG algorithm, and the second sub-model is a model based on the PPO algorithm.

[0027] A second aspect of this application provides a cooperative countermeasure device for a drone swarm, the device comprising:

[0028] The first acquisition module is used to acquire global situational information of the adversarial scenario. The global situational information includes: the real-time positions of multiple first UAVs in the first party, the real-time positions of multiple second UAVs in the second party, the position of the radar in the first party, the position of the radar in the second party, and the position of the target in the second party. The first party and the second party are adversaries to each other, and the target position of the second party is the attack position of the first party.

[0029] The second acquisition module is used to acquire the role information of the adversarial scenario. The role information includes the number of roles and the drone set of each role. The roles include reconnaissance roles, jamming roles and attack roles.

[0030] The third acquisition module is used to acquire the state space and behavior space of each of the aforementioned roles. The behavior space includes empty behavior, movement behavior, and payload behavior. The payload behavior includes reconnaissance behavior, interference behavior, and attack behavior.

[0031] The first processing module is used to call the first sub-model corresponding to the behavior type according to the role information, input the state space and behavior space of the role into the corresponding first sub-model, and obtain the joint behavior information of each role;

[0032] The second processing module is used to call the corresponding second sub-model according to the joint behavior information of the role, and input the obtained state space information into the second sub-model to obtain the behavior information of each drone in the role. Each drone performs cooperative combat based on the behavior information. The state space information includes: the state information of the drone itself, the state information of the drone cluster of the role, environmental information, and task information.

[0033] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the cooperative anti-drone swarm method of the first aspect of this application.

[0034] In a fourth aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, it implements the cooperative anti-drone swarm method of the first aspect of this application.

[0035] The beneficial effects of the technical solutions provided in this application include at least the following:

[0036] The collaborative adversarial method for drone swarms provided in this application embodiment acquires global situational information of the adversarial scenario. This global situational information includes: the real-time positions of multiple first drones in the first party, the real-time positions of multiple second drones in the second party, the position of the radar in the first party, the position of the radar in the second party, and the target position in the second party. The first party and the second party are adversaries, and the target position of the second party is the attack position of the first party. The method also acquires role information of the adversarial scenario, including the number of roles and the drone set for each role. The roles include reconnaissance roles, jamming roles, and attack roles. Finally, the method acquires the state space and behavior space of each role. The system includes spatial behavior, movement behavior, and payload behavior, with payload behavior encompassing reconnaissance, interference, and attack behaviors. Based on the role information, a first sub-model corresponding to the behavior type is invoked, and the role's state space and behavior space are input into the corresponding first sub-model to obtain joint behavior information for each role. Based on the joint behavior information, a corresponding second sub-model is invoked, and the acquired state space information is input into the second sub-model to obtain the behavior information of each UAV within the role. Each UAV then performs cooperative combat based on this behavior information. The state space information includes: the UAV's own state information, the state information of the UAV cluster within the role, environmental information, and task information. Compared to traditional collaborative methods, the role-driven hierarchical collaborative strategy proposed in this invention significantly improves reusability and scalability. The upper-level strategy is based on role modeling; by proposing the concept of a role, the coordination strategy relationships between roles are learned and can be reused in other scenarios, improving the reusability of the strategy.

[0037] Furthermore, unmanned swarms consist of multiple unmanned systems that move at high speeds and may join or leave at any time, resulting in dynamic and volatile swarms. This makes it impossible to model the execution strategy of this behavior as multi-agent reinforcement learning, as multi-agent reinforcement learning requires a fixed number of nodes, thus failing to effectively support dynamic changes in nodes. While single-agent reinforcement learning is difficult to apply in multi-agent environments because each agent might consider other agents as part of the environment, leading to instability and hindering policy convergence, multiple agents with the same action space and objective exhibit high policy similarity. Therefore, model sharing can be used, allowing multiple agents to share a single policy model. In this invention, after the agents are layered at the upper level, lower-level agents of the same type have the same state and action spaces. Therefore, lower-level policies can use a joint state representation method to improve policy reusability and node scalability, enabling cooperative policies to have good adaptability when solving complex problems. The role-driven unmanned swarm cooperative adversarial strategy proposed in this invention uses reinforcement learning for modeling, enabling the learned strategy to solve cooperative problems in complex and ever-changing environments. Furthermore, this invention designs different state spaces, action spaces, and reward functions for different strategies, guiding each strategy to learn effectively and improving the intelligence level of the unmanned swarm. Attached Figure Description

[0038] Figure 1 A flowchart illustrating a cooperative countermeasure method for a drone swarm provided in this application embodiment;

[0039] Figure 2 This application provides a schematic diagram of a cooperative combat scenario involving a drone swarm.

[0040] Figure 3 This is a schematic diagram of a collaborative framework for a drone swarm provided in an embodiment of this application;

[0041] Figure 4 A schematic diagram of the state space of a role provided in an embodiment of this application;

[0042] Figure 5 A schematic diagram of the behavioral space of a character provided in an embodiment of this application;

[0043] Figure 6 This is a state space diagram of a movement strategy provided in an embodiment of this application;

[0044] Figure 7 This is a state space diagram of a reconnaissance strategy provided in an embodiment of this application;

[0045] Figure 8 This is a state space diagram of an interference strategy provided in an embodiment of this application;

[0046] Figure 9 This is a state space diagram of an attack strategy provided in an embodiment of this application;

[0047] Figure 10 This application provides an embodiment of a method for cooperative countermeasures against unmanned swarms, illustrating the situation of the unmanned swarm during the countermeasure process.

[0048] Figure 11 This is a structural diagram of a cooperative combat device for a drone swarm provided in an embodiment of this application. Detailed Implementation

[0049] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0050] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more.

[0051] In addition, the use of “based on” or “according to” implies openness and inclusivity, because processes, steps, calculations or other actions “based on” or “according to” one or more conditions or values ​​can in practice be based on additional conditions or values ​​beyond those conditions.

[0052] With the continuous development of technologies such as computers, automatic control, and robotics, unmanned systems, primarily new types of drones and unmanned vehicles, are playing an increasingly important role and are widely used in both military and civilian fields. However, individual unmanned systems are limited by their own power, functionality, and performance, making it impossible to complete complex tasks independently. Therefore, drone swarms are used to address this issue. Consequently, how drone swarms can cooperate effectively to complete tasks has become a crucial aspect.

[0053] Current drone swarm collaboration mainly relies on the traditional approach of assuming the situation in the mission area is known. This involves marking threat ranges such as air defense positions, terrain, and mission objectives on a map, and then inputting the map into the drones for rule-based collaboration. However, this approach still has limitations when facing complex missions and changing environments, resulting in poor collaboration performance.

[0054] Furthermore, most traditional reinforcement learning research is based on learning strategies in fixed simulation environments, which leads to excessive coupling between the strategy and the environment. If the environment is changed to a similar one, the effect may be very poor, and a complete retraining is required. However, logically speaking, cooperative strategy is an abstract and high-level concept that has a certain degree of reusability in similar tasks. For example, in a cooperative search and rescue scenario, a cooperative search should be conducted first, and then a rescue operation should be carried out on the searched target. This kind of cooperative strategy with temporal relationship can be reused.

[0055] To address the aforementioned issues, this application proposes a cooperative adversarial method for drone swarms. This method significantly improves reusability and scalability through a role-driven hierarchical cooperative strategy. The upper-level strategy is based on role modeling. By proposing the concept of roles, the coordination strategy relationships between roles are learned and can be replicated in other scenarios, thus enhancing the reusability of the strategy.

[0056] This application provides a cooperative countermeasure method for unmanned aerial vehicle (UAV) swarms, such as... Figure 1 As shown, the method includes the following steps:

[0057] Step 101: Obtain global situational information of the adversarial scenario.

[0058] The global situation information includes: the real-time positions of multiple first UAVs in the first party, the real-time positions of multiple second UAVs in the second party, the position of the radar in the first party, the position of the radar in the second party, and the position of the target in the second party. The first party and the second party are adversaries, and the target position of the second party is the attack position of the first party.

[0059] Optionally, information included in the global situational awareness can be extracted from image information collected by multiple drones and input information.

[0060] Figure 2 In this implementation scenario, within a rectangular area, the first party intends to attack the island, employing a swarm of multiple drones carrying different payloads to break through the second party's defense system and destroy their command post. The second party intends to defend the island, relying on ground-based air defense and radar to protect their command post. The victory rules for this scenario are: the first party wins if they destroy the second party's command post; the second party wins if they destroy all of the first party's threat units or the timeout period expires. The first party, acting as the algorithm carrier, uses the drone swarm cooperative combat method provided in this application, while the second party uses fixed rules to make decisions. Through continuous interaction and confrontation between the first and second parties, the first party ultimately learns a cooperative combat strategy to defeat the second party.

[0061] Step 102: Obtain the character information of the confrontation scenario.

[0062] The role information includes the number of roles and the drone set for each role, which includes reconnaissance roles, jamming roles, and attack roles.

[0063] The character information is determined based on the current combat mission.

[0064] Step 103: Obtain the state space and behavior space of each of the aforementioned roles.

[0065] The behavior space includes empty behavior, movement behavior, and payload behavior, and the payload behavior includes reconnaissance behavior, interference behavior, and attack behavior.

[0066] Step 104: Based on the role information, call the first sub-model corresponding to the behavior type, input the state space and behavior space of the role into the corresponding first sub-model, and obtain the joint behavior information of each role.

[0067] Step 105: Call the corresponding second sub-model according to the joint behavior information of the role, and input the obtained state space information into the second sub-model to obtain the behavior information of each drone in the role. Each drone performs cooperative combat based on the behavior information.

[0068] The state space information includes: the state information of the drone itself, the state information of the drone cluster of the role, environmental information, and task information.

[0069] In actual implementation, the above steps require the pre-design of a role-driven hierarchical collaborative framework for unmanned swarms, consisting of an upper-level role collaboration layer and a lower-level behavior execution layer. Based on this framework, the upper-level role collaboration strategy for collaborative adversarial tactics is designed, including the design of the role state space, role behavior space, and reward function. Based on this framework, the lower-level behavior execution strategy for collaborative adversarial tactics is designed, with a total of four behavior execution strategies: movement, reconnaissance, interference, and attack, including the design of the state space, action space, and reward function. Finally, the upper and lower-level strategies are jointly updated to guide each unmanned system in selecting the optimal action.

[0070] Figure 3The overall architecture of this framework is shown in the diagram. It consists of two layers: an upper-layer role coordination layer and a lower-layer behavior execution layer. The upper layer uses multi-agent reinforcement learning to model role coordination strategies, making decisions every few time steps and outputting joint role behaviors. These behaviors represent a higher-level, abstract form of coordination. Instead of directly interacting with the environment, these behaviors activate corresponding behavior execution strategies. All unmanned systems select their corresponding behavior execution strategies based on their assigned roles, deciding their actions to interact with the environment. Considering the reusability of strategies and the dynamic adaptability of cluster size, the lower-layer behavior execution strategies are modeled using single-agent reinforcement learning. By jointly representing the state of each unmanned system, model sharing and strategy reuse are achieved.

[0071] Optionally, before performing step 101, the method further includes:

[0072] Construct an initial adversarial model, which includes multiple initial first sub-models and multiple initial second sub-models.

[0073] Specifically, the character state space is as follows: Figure 4 As shown, S r =[S r 1 ,S r 2 ,…,S r m ], where the state space S of the i-th character r i =[S env, S r 1 ,S r 2 ,…,S r m ,S extra i ], where S env For environmental information, in adversarial scenarios, it is represented as the enemy's situation, S r S represents the situational information for all roles. extra i This represents supplementary information for the i-th role's decision-making. In adversarial scenarios, the supplementary information for the reconnaissance role is reconnaissance situational information, including reconnaissance range and interference range, while the supplementary information for the interference and attack roles is interference situational information. The role state space is formally represented using a grid graph to characterize the situational information of unmanned swarms of different sizes; the role behavior space is defined as A. r =[A r 1 A r 2 ,…,A r m], where the behavior space A of the i-th character r i =[A no-op A move A load ],like Figure 5 As shown, the first two behaviors are general behaviors, which do not have the concept of roles and include waiting in the air, hovering, and A. no-op The empty behavior represents the implementation of role-time coordination, A. move This represents movement behavior, which is further subdivided into task movement behavior A. task-move Cooperative movement behavior A role-move It is the realization of role-space collaboration, that is, which two roles are close to each other, and the load behavior A. load These are behaviors unique to each role, with the payload behaviors for the three roles being reconnaissance, interference, and attack. The first reward function is determined to be R. r =[R r 1 ,R r 2 ,…,R r m ], where the reward function R for the i-th character is r i =R acc +R time-penalty +R trans-penalty , where R acc The cumulative reward at the lower level is represented by R, which is the average reward obtained by all unmanned clusters of role i within the time unit (k time steps) of the upper-level decision-making process. time-penalty R represents the time penalty. trans-penalty This indicates a penalty for changes in decision-making.

[0074] Optionally, after constructing the initial adversarial model, the method further includes:

[0075] Multiple sample information of mobile behaviors are acquired. The sample information of mobile behaviors includes multiple mobile sample state spaces and mobile sample behavior spaces. The mobile sample state space is a stacked first grid map. The first grid map includes the UAV's own position information, cluster position information, target position information, unmanned system domain information, and global situational information. The mobile sample behavior space includes empty behavior and mobile behavior. The preset initial first sub-model of the mobile behavior is iteratively trained using the mobile sample state space and the mobile sample behavior space. Each iteration obtains mobile joint behavior information. The mobile joint behavior information is evaluated using a preset first reward function to obtain a first evaluation value. The first evaluation value is obtained until it meets a second preset threshold, then the first sub-model of the mobile behavior is obtained.

[0076] In actual implementation, such as Figure 6 As shown, the state space, action space, and reward function of the mobile behavior execution strategy are determined.

[0077] The state space is designed as a joint representation of the unmanned system's own position information, cluster position information, target point position information, unmanned system's domain information, and global situational information. It is represented using a grid method, achieved through the stacking of multiple grid maps. The action space is designed as empty actions and movement actions.

[0078] The first reward function is designed as R. move =R target +R group +R alive +R time-penalty , where R target R represents the reward for moving towards the target point; the closer to the target location, the greater the reward. group This indicates that the more clustered the group, the greater the reward for movement. R alive R represents the survival reward. time-penalty This indicates a time penalty.

[0079] Optionally, after constructing the initial adversarial model, the method further includes: acquiring sample information of multiple reconnaissance behaviors, wherein the sample information of the reconnaissance behaviors includes multiple reconnaissance sample state spaces and reconnaissance sample behavior spaces, wherein the reconnaissance sample state space includes a stacked second grid map, wherein the second grid map includes the UAV's own position information, the reconnaissance role cluster position information, the unmanned system's domain information, global situation information, and reconnaissance situation information; wherein the reconnaissance sample behavior space includes air behavior and movement behavior; and iteratively training a preset initial first sub-model of the reconnaissance behavior using the reconnaissance sample state space and the reconnaissance sample behavior space, wherein each iteration obtains corresponding reconnaissance joint behavior information, and the movement joint behavior information is evaluated using a preset second reward function to obtain a second evaluation value, until the second evaluation value meets a second preset threshold, thereby obtaining the first sub-model of the reconnaissance behavior.

[0080] Among them, such as Figure 7 As shown, the first sub-model of reconnaissance behavior includes the state space, action space, and reward function for the execution strategy of reconnaissance behavior. The state space is designed as a joint representation of the unmanned system's own position information, the reconnaissance role cluster's position information, the unmanned system's domain information, global situational information, and reconnaissance situational information, and is represented using a grid method, consisting of multiple stacked grid diagrams. The action space is designed as no-action actions and movement actions. The second reward function is designed as R = R explore +R disperse +R alive +R time-penalty , where R exploreThis represents exploration rewards, such as bonus points awarded for exploring new areas or finding new objectives. R disperse This indicates a reward for dispersed reconnaissance; the more dispersed the reconnaissance, the higher the reward (R). alive R represents the survival reward. time-penalty This indicates a time penalty.

[0081] Optionally, after constructing the initial adversarial model, the method further includes: acquiring sample information of multiple interference behaviors, wherein the sample information of interference behaviors includes multiple interference sample state spaces and interference sample behavior spaces, wherein the interference sample state space includes a stacked third grid diagram, wherein the third grid diagram includes the unmanned aerial vehicle system's own state information, interference role cluster state information, unmanned system domain information, global situation information, interference situation information, and attack role cluster location information; and the interference sample behavior space includes empty behavior, movement behavior, and interference behavior.

[0082] The initial first sub-model of the interference behavior is iteratively trained using the interference sample state space and the interference sample behavior space. Each iteration yields the corresponding interference joint behavior information. The interference joint behavior information is evaluated using a preset third reward function to obtain a third evaluation value. The process continues until the third evaluation value meets a third preset threshold, at which point the first sub-model of the interference behavior is obtained.

[0083] Specifically, such as Figure 8 As shown, the first sub-model of the interference behavior includes the state space, action space, and reward function of the interference behavior execution strategy. The state space is designed as a joint representation of the unmanned system's own state information, the state information of the interference character cluster, the unmanned system's domain information, global situational information, interference situational information, and the positional information of the attacking character cluster. It adopts a grid method, consisting of multiple stacked grid diagrams. The action space is designed as empty actions, movement actions, and interference actions. The reward function is designed as R... dist =R valid +R overlap +R alive +R time-penalty , where R valid Reward for effective interference; the interference range covers both friendly and enemy forces. (R) overlap R represents the penalty for overlapping interference. The larger the overlap area, the larger the penalty score. alive R represents the survival reward. time-penalty This indicates a time penalty.

[0084] Optionally, after obtaining the first sub-model, the method further includes: acquiring sample information of multiple attack behaviors, wherein the sample information of the attack behaviors includes multiple attack sample state spaces and attack sample behavior spaces, wherein the attack sample state space includes a stacked fourth grid diagram, wherein the fourth grid diagram includes state information of the unmanned aerial vehicle system, state information of the attacking role cluster, domain information of the unmanned system, global situation information, and interference situation information, and the attack sample behavior space includes empty behavior, movement behavior, and attack behavior;

[0085] The initial first sub-model of the preset attack behavior is iteratively trained using the attack sample space and the attack sample behavior space. Each iteration yields the corresponding joint attack behavior information. The joint attack behavior information is evaluated using a preset fourth reward function to obtain a fourth evaluation value. The process continues until the fourth evaluation value meets a fourth preset threshold, at which point the first sub-model of the attack behavior is obtained.

[0086] Among them, such as Figure 9 As shown, the first sub-model of the attack behavior includes the state space, action space, and reward function for the attack behavior execution strategy. The state space is designed as a joint representation of the unmanned system's state information, the attacker's cluster state information, the unmanned system's domain information, global situational information, and interference situational information, and is formally constructed using stacked grid graphics. The action space is designed as empty actions, movement actions, and attack actions. The reward function is designed as R... atk =R pos +R effic +R alive +R time-penalty , where R pos This indicates the attack position bonus; the closer to the enemy's priority, the higher the bonus value. (R) effic R represents the attack efficiency bonus. alive R represents the survival reward. time-penalty This indicates a time penalty.

[0087] The first sub-model is based on the MADDPG algorithm, and the second sub-model is based on the PPO algorithm.

[0088] In practical implementation, within an unmanned swarm, roles can be distinguished based on load. Roles are classifications of the unmanned systems and are closely related to the actions they can perform. This invention uses the MADDPG algorithm for decision-making. Unlike traditional methods, it does not model MADDPG on individuals, but rather on roles. For a specific task scenario, regardless of the size of the unmanned swarm, the roles required to complete the task are always determined, and the number of roles is relatively small and does not increase with the number of unmanned systems. The role state space S... rAs input to the MADDPG algorithm, a decision is made every K time steps to determine the joint behavior A for each role, i.e., the behavior that each role should perform within the current K time steps. This behavior is then used by the lower layer to select a specific model for invocation, and is determined based on the reward function R. r The reward is calculated and the model parameters are updated via backpropagation.

[0089] As the lower layer, the behavior execution layer needs to execute the behavioral instructions issued by the upper-level role coordination layer, while also considering the dynamic changes of the unmanned swarm. The PPO algorithm is introduced into the behavior execution strategy, based on model sharing modeling, enabling it to learn an effective behavior execution strategy.

[0090] Each unit follows the role-based joint behavior A output from the upper layer. r This determines whether the current lower layer should use an empty behavior, movement behavior, interference behavior, or attack behavior. Each behavior model corresponds to a PPO model. The lower layer's input state space... The following formula represents the state information of unmanned system j when it performs behavior i. It includes four parts: the state of unmanned system j itself, the state information of unmanned cluster k, the environmental information and task information that are of concern during the execution of the behavior, where unmanned cluster k refers to unmanned system clusters with the same role.

[0091]

[0092] Each time step will record the current state of the individual unit. The PPO algorithm determines the specific action (Arj) of the current unit and calculates the reward based on the current behavior model using a reward function. The model parameters are then updated via backpropagation. This process is repeated until the model or the number of training iterations reaches the preset maximum number of training rounds.

[0093] The cooperative adversarial method for UAV swarms provided in this application uses role-based modeling as its upper-layer strategy. By proposing the concept of roles, the coordination strategy relationships between roles are learned and can be replicated in other scenarios, improving the reusability of the strategy. Furthermore, UAV swarms consist of multiple unmanned systems that move at high speeds and may join or leave at any time, resulting in dynamic variability within the swarm. This makes it impossible to model the execution strategy of this behavior as multi-agent reinforcement learning, as multi-agent reinforcement learning requires a fixed number of nodes, thus failing to effectively support dynamic changes in nodes. While single-agent reinforcement learning is difficult to apply in multi-agent environments because any agent might consider other agents as part of the environment, leading to instability and hindering strategy convergence, multiple agents with the same action space and objective exhibit high policy similarity. Therefore, model sharing can be used to allow multiple agents to share a single policy model. In this invention, after the agents are layered at the upper layer, lower-layer agents of the same type have the same state and action spaces. Therefore, the lower-layer strategy can use a joint state representation method to improve policy reusability and node scalability, enabling the cooperative strategy to have good adaptability when solving complex problems. The role-driven unmanned swarm cooperative adversarial strategy proposed in this invention uses reinforcement learning for modeling, enabling the learned strategy to solve cooperative problems in complex and ever-changing environments. Furthermore, this invention designs different state spaces, action spaces, and reward functions for different strategies, guiding each strategy to learn effectively and improving the intelligence level of the unmanned swarm.

[0094] Figure 10 This is a schematic diagram of the unmanned aerial vehicle (UAV) swarm cooperative combat method of the present invention during the combat process. For example... Figure 10 As shown in the simulation results, the present invention can effectively achieve adaptive cooperative adversarial decision-making for unmanned swarms during the constantly changing offensive and defensive situation. Furthermore, during the adversarial process, unmanned systems with different roles effectively cooperate to complete the mission of striking enemy targets. Therefore, the role-driven unmanned swarm cooperative adversarial strategy adopted in this invention can solve complex coordination problems and demonstrates that this cooperative adversarial strategy has good reusability and dynamic adaptability to swarm size.

[0095] This application also provides a cooperative countermeasure device for drone swarms, such as... Figure 11 As shown, the device includes:

[0096] The first acquisition module 11 is used to acquire global situational information of the adversarial scenario. The global situational information includes: the real-time positions of multiple first UAVs in the first party, the real-time positions of multiple second UAVs in the second party, the position of the radar in the first party, the position of the radar in the second party, and the position of the target in the second party. The first party and the second party are adversaries to each other, and the target position of the second party is the attack position of the first party.

[0097] The second acquisition module 12 is used to acquire the role information of the adversarial scenario. The role information includes the number of roles and the drone set of each role. The roles include reconnaissance roles, jamming roles and attack roles.

[0098] The third acquisition module 13 is used to acquire the state space and behavior space of each of the roles. The behavior space includes empty behavior, movement behavior and load behavior. The load behavior includes reconnaissance behavior, interference behavior and attack behavior.

[0099] The first processing module 14 is used to call the first sub-model corresponding to the behavior type according to the role information, input the state space and behavior space of the role into the corresponding first sub-model, and obtain the joint behavior information of each role;

[0100] The second processing module 15 is used to call the corresponding second sub-model according to the joint behavior information of the role, and input the obtained state space information into the second sub-model to obtain the behavior information of each drone in the role. Each drone performs cooperative combat based on the behavior information. The state space information includes: the state information of the drone itself, the state information of the drone cluster of the role, environmental information, and task information.

[0101] In one embodiment, the device further includes a training module 16, for:

[0102] Construct an initial adversarial model, which includes multiple initial first sub-models and multiple initial second sub-models.

[0103] In one embodiment, the training module 16 is further configured to:

[0104] Acquire sample information of multiple movement behaviors, the sample information of the movement behaviors includes multiple movement sample state spaces and movement sample behavior spaces, the movement sample state space is a stacked first grid map, the first grid map includes the UAV's own position information, cluster position information, target position information, unmanned system domain information and global situation information, the movement sample behavior space includes empty behavior and movement behavior;

[0105] The initial second sub-model of the mobile behavior is iteratively trained using the mobile sample state space, the mobile sample behavior space, and the behavior information of the mobile role drone. Each iteration yields mobile joint behavior information, which is then evaluated using a preset first reward function to obtain a first evaluation value. This process continues until the first evaluation value meets a second preset threshold, at which point the first sub-model of the mobile behavior is obtained.

[0106] In one embodiment, the training module 16 is further configured to:

[0107] Acquire sample information of multiple reconnaissance behaviors. The sample information of the reconnaissance behaviors includes multiple reconnaissance sample state spaces and reconnaissance sample behavior spaces. The reconnaissance sample state space includes a stacked second grid map, which includes the UAV's own position information, reconnaissance role cluster position information, unmanned system domain information, global situation information, and reconnaissance situation information. The reconnaissance sample behavior space includes aerial behavior and movement behavior.

[0108] The initial second sub-model of the reconnaissance behavior is iteratively trained using the reconnaissance sample state space, the reconnaissance sample behavior space, and the behavior information of the reconnaissance role UAV. Each iteration yields the corresponding reconnaissance joint behavior information. The mobile joint behavior information is evaluated using a preset second reward function to obtain a second evaluation value. The process continues until the second evaluation value meets a second preset threshold, at which point the first sub-model of the reconnaissance behavior is obtained.

[0109] In one embodiment, the training module 16 is further configured to:

[0110] The system acquires sample information of multiple interference behaviors, which includes multiple interference sample state spaces and interference sample behavior spaces. The interference sample state space includes a stacked third grid diagram, which includes the UAV system's own state information, interference role cluster state information, UAV system domain information, global situation information, interference situation information, and attack role cluster location information. The interference sample behavior space includes empty behavior, movement behavior, and interference behavior.

[0111] The initial second sub-model of the interference behavior is iteratively trained using the state space of the interference sample, the behavior space of the interference sample, and the behavior information of the UAV of the interference role. Each iteration yields the corresponding joint interference behavior information. The joint interference behavior information is evaluated using a preset third reward function to obtain a third evaluation value. The process continues until the third evaluation value meets a third preset threshold, at which point the first sub-model of the interference behavior is obtained.

[0112] In one embodiment, the training module 16 is further configured to:

[0113] Acquire sample information of multiple attack behaviors. The sample information of the attack behaviors includes multiple attack sample state spaces and attack sample behavior spaces. The attack sample state space includes a stacked fourth grid diagram. The fourth grid diagram includes the state information of the UAV system, the state information of the attack role cluster, the domain information of the UAV system, the global situation information, and the interference situation information. The attack sample behavior space includes empty behavior, movement behavior, and attack behavior.

[0114] The initial second sub-model of the attack behavior is iteratively trained using the attack sample space, the attack sample behavior space, and the behavior information of the drone of the attacking role. Each iteration yields the corresponding joint attack behavior information. The joint attack behavior information is evaluated using a preset fourth reward function to obtain a fourth evaluation value. The process continues until the fourth evaluation value meets a fifth preset threshold, at which point the first sub-model of the attack behavior is obtained.

[0115] In one embodiment, the first sub-model is a model based on the MADDPG algorithm, and the second sub-model is a model based on the PPO algorithm.

[0116] The collaborative countermeasure device for drone swarms provided in this embodiment can execute the above-described collaborative countermeasure method embodiment for drone swarms. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0117] Specific limitations regarding the cooperative countermeasure device for drone swarms can be found in the limitations on cooperative countermeasure methods for drone swarms mentioned above, and will not be repeated here. Each module in the aforementioned cooperative countermeasure device for drone swarms can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor of the electronic device in hardware form, or stored in the memory of the electronic device in software form, so that the processor can call and execute the corresponding operations of each module.

[0118] In another embodiment of this application, an electronic device is also provided, including a memory and a processor. The memory stores a computer program, which, when executed by the processor, implements the steps of the cooperative anti-drone swarm method as described in the embodiments of this application.

[0119] In another embodiment of this application, a computer-readable storage medium is also provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of the cooperative anti-drone swarm method as described in the embodiments of this application are implemented.

[0120] In another embodiment of this application, a computer program product is also provided, which includes computer instructions that, when executed on a collaborative countermeasure device for a drone swarm, cause the collaborative countermeasure device for the drone swarm to perform each step of the collaborative countermeasure method for a drone swarm in the method flow shown in the above method embodiment.

[0121] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software programs, implementation can be, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device containing one or more servers, data centers, etc., that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks, SSDs).

[0122] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0123] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A cooperative countermeasure method for unmanned aerial vehicle (UAV) swarms, characterized in that, The method includes: The global situation information of the adversarial scenario is obtained, which includes: the real-time positions of multiple first UAVs in the first party, the real-time positions of multiple second UAVs in the second party, the position of the radar in the first party, the position of the radar in the second party, and the position of the target in the second party. The first party and the second party are adversaries, and the target position of the second party is the attack position of the first party. Obtain the role information of the adversarial scenario, which includes the number of roles and the drone set of each role. The roles include reconnaissance roles, jamming roles, and attack roles. The state space and behavior space of each of the aforementioned roles are obtained. The behavior space includes empty behavior, movement behavior, and payload behavior. The payload behavior includes reconnaissance behavior, interference behavior, and attack behavior. The state space information includes: the state information of the UAV itself, the state information of the UAV cluster of the aforementioned roles, environmental information, and task information. Based on the role information, the first sub-model corresponding to the behavior type is invoked, and the state space and behavior space of the role are input into the corresponding first sub-model to obtain the joint behavior information of each role; the first sub-model is a role collaboration model based on the multi-agent deep deterministic policy gradient (MADDPG) algorithm, which is used to output the joint behavior information of the role based on the role's state space and behavior space; The corresponding second sub-model is invoked based on the joint behavior information of the role, and the obtained state space information is input into the second sub-model to obtain the behavior information of each drone in the role. Each drone performs cooperative combat based on the behavior information. The second sub-model is a behavior execution model based on the near-end policy optimization PPO algorithm, which is used to output the behavior information of a drone based on the state space information of a single drone.

2. The method according to claim 1, characterized in that, Before acquiring the global situational information of the adversarial scenario, the method further includes: Construct an initial adversarial model, which includes multiple initial first sub-models and multiple initial second sub-models.

3. The method according to claim 2, characterized in that, After constructing the initial adversarial model, the method further includes: Acquire sample information of multiple movement behaviors, the sample information of the movement behaviors includes multiple movement sample state spaces and movement sample behavior spaces, the movement sample state space is a stacked first grid map, the first grid map includes the UAV's own position information, cluster position information, target position information, unmanned system domain information and global situation information, the movement sample behavior space includes empty behavior and movement behavior; The initial second sub-model of the mobile behavior is iteratively trained using the mobile sample state space, the mobile sample behavior space, and the behavior information of the mobile role drone. Each iteration yields mobile joint behavior information, which is then evaluated using a preset first reward function to obtain a first evaluation value. This process continues until the first evaluation value meets a second preset threshold, at which point the first sub-model of the mobile behavior is obtained.

4. The method according to claim 2, characterized in that, After constructing the initial adversarial model, the method further includes: Acquire sample information of multiple reconnaissance behaviors. The sample information of the reconnaissance behaviors includes multiple reconnaissance sample state spaces and reconnaissance sample behavior spaces. The reconnaissance sample state space includes a stacked second grid map, which includes the UAV's own position information, reconnaissance role cluster position information, unmanned system domain information, global situation information, and reconnaissance situation information. The reconnaissance sample behavior space includes aerial behavior and movement behavior. The initial second sub-model of the reconnaissance behavior is iteratively trained using the reconnaissance sample state space, the reconnaissance sample behavior space, and the behavior information of the reconnaissance role UAV. Each iteration yields the corresponding reconnaissance joint behavior information. The mobile joint behavior information is evaluated using a preset second reward function to obtain a second evaluation value. The process continues until the second evaluation value meets a second preset threshold, at which point the first sub-model of the reconnaissance behavior is obtained.

5. The method according to claim 2, characterized in that, After constructing the initial adversarial model, the method further includes: The system acquires sample information of multiple interference behaviors, which includes multiple interference sample state spaces and interference sample behavior spaces. The interference sample state space includes a stacked third grid diagram, which includes the UAV system's own state information, interference role cluster state information, UAV system domain information, global situation information, interference situation information, and attack role cluster location information. The interference sample behavior space includes empty behavior, movement behavior, and interference behavior. The initial second sub-model of the interference behavior is iteratively trained using the state space of the interference sample, the behavior space of the interference sample, and the behavior information of the UAV of the interference role. Each iteration yields the corresponding joint interference behavior information. The joint interference behavior information is evaluated using a preset third reward function to obtain a third evaluation value. The process continues until the third evaluation value meets a third preset threshold, at which point the first sub-model of the interference behavior is obtained.

6. The method according to claim 2, characterized in that, After constructing the initial adversarial model, the method further includes: Acquire sample information of multiple attack behaviors. The sample information of the attack behaviors includes multiple attack sample state spaces and attack sample behavior spaces. The attack sample state space includes a stacked fourth grid diagram. The fourth grid diagram includes the state information of the UAV system, the state information of the attack role cluster, the domain information of the UAV system, the global situation information, and the interference situation information. The attack sample behavior space includes empty behavior, movement behavior, and attack behavior. The initial second sub-model of the attack behavior is iteratively trained using the attack sample state space, the attack sample behavior space, and the behavior information of the drone of the attacking role. Each iteration yields the corresponding joint attack behavior information. The joint attack behavior information is evaluated using a preset fourth reward function to obtain a fourth evaluation value. The process continues until the fourth evaluation value meets a fifth preset threshold, at which point the first sub-model of the attack behavior is obtained.

7. The method according to claim 1, characterized in that, The first sub-model is based on the MADDPG algorithm, and the second sub-model is based on the PPO algorithm.

8. A cooperative combat device for a swarm of unmanned aerial vehicles (UAVs), characterized in that, The device includes: The first acquisition module is used to acquire global situational information of the adversarial scenario. The global situational information includes: the real-time positions of multiple first UAVs in the first party, the real-time positions of multiple second UAVs in the second party, the position of the radar in the first party, the position of the radar in the second party, and the position of the target in the second party. The first party and the second party are adversaries to each other, and the target position of the second party is the attack position of the first party. The second acquisition module is used to acquire the role information of the adversarial scenario. The role information includes the number of roles and the drone set of each role. The roles include reconnaissance roles, jamming roles and attack roles. The third acquisition module is used to acquire the state space and behavior space of each of the aforementioned roles. The behavior space includes empty behavior, movement behavior, and payload behavior. The payload behavior includes reconnaissance behavior, interference behavior, and attack behavior. The state space information includes: the state information of the UAV itself, the state information of the UAV cluster of the aforementioned roles, environmental information, and task information. The first processing module is used to call the first sub-model corresponding to the behavior type according to the role information, input the state space and behavior space of the role into the corresponding first sub-model, and obtain the joint behavior information of each role; the first sub-model is a role collaboration model based on the multi-agent deep deterministic policy gradient (MADDPG) algorithm, and is used to output the joint behavior information of the role according to the state space and behavior space of the role. The second processing module is used to call the corresponding second sub-model according to the joint behavior information of the role, and input the obtained state space information into the second sub-model to obtain the behavior information of each drone in the role. Each drone performs cooperative combat based on the behavior information. The second sub-model is a behavior execution model based on the near-end strategy optimization PPO algorithm, which is used to output the behavior information of a drone according to the state space information of a single drone.

9. An electronic device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, which, when executed by the processor, implements the cooperative countermeasure method for a drone swarm as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the cooperative countermeasure method for a drone swarm as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-unmanned-aerial-vehicle intelligent cooperative defense penetration confrontation method

    CN112198892A

  • Aircraft soldier system intelligent behavior modeling method based on global situation information

    CN112560332A