Large-scale unmanned cluster game strategy construction system and method

By designing a large-scale unmanned cluster game strategy construction system, using MAPPO algorithm and distributed execution module, the problem of difficult expansion of existing technology to large-scale unmanned clusters and central nodes is solved, and efficient simulation and verification and task reliability and flexibility are improved.

CN120068641APending Publication Date: 2025-05-30POLIXIR TECH LTD

Patent Information

Application Number
CN202510204026.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing unmanned cluster simulation and control methods are difficult to expand to large-scale clusters, the central node's computational burden is too heavy, the communication delay is increased, and single-point failure is lacking in meticulous modeling of real task scenarios.

Method used

A large-scale unmanned cluster game strategy construction system was designed, including simulation environment module, centralized training module and distributed execution module. Centralized training is carried out through the MAPPO algorithm combined with the policy network and the value network, and each agent can independently execute the strategy in the distributed execution module, and independently select the action distribution based on local observation information.

Benefits of technology

It realizes efficient simulation and verification of large-scale unmanned clusters, overcomes the computing bottlenecks and communication delay problems in traditional methods, enhances the reliability and flexibility of tasks, and makes up for the shortcomings of existing simulation systems in real task scenario modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068641A_ABST
    Figure CN120068641A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale unmanned cluster game strategy construction system and method. The system comprises a simulation environment module, a centralized training module and a distributed execution module, wherein the centralized training module and the distributed execution module are based on the simulation environment module. The simulation environment module is used for simulating a dynamic interaction environment of a large-scale unmanned cluster; the centralized training module adopts a cluster game confrontation reinforcement learning algorithm MAPPO and combines a strategy network and a value network to realize strategy optimization of a large-scale unmanned cluster through centralized training; and the distributed execution module applies the strategy generated by the centralized training module to the large-scale unmanned cluster, so that each agent in the large-scale unmanned cluster independently executes the strategy, and independently selects action distribution based on own local observation information. The method can cope with dimension disasters caused by too high state space dimension in a large-scale scene, and makes up for the defects of an existing simulation system in the aspect of real task scene modeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned cluster game confrontation, and particularly relates to a system and method for constructing a large-scale unmanned cluster game strategy. Background Art

[0002] In recent years, unmanned cluster game confrontation technology has shown great application potential in many fields. Especially in the military and security fields, the research on cluster operations of unmanned aerial vehicles and unmanned vehicles has received much attention. However, the existing unmanned cluster simulation and control methods have significant limitations and need to be further improved and optimized.

[0003] First of all, the current simulation systems are mostly designed for small-scale unmanned clusters, usually with the number of unmanned devices not exceeding 5. The dynamic models of these systems are complex and usually focus on the precise simulation of the control of unmanned devices, sensor characteristics, and environmental interaction. However, the high-fidelity modeling method has a huge computational cost and is difficult to scale to large-scale cluster scenarios. The existing simulation systems have significant deficiencies in terms of scene complexity and cluster scale scalability, and it is difficult to meet the actual needs of cluster cooperation and confrontation with a scale of more than a hundred sorties.

[0004] Secondly, traditional methods mostly adopt the way of centralized training and centralized execution, that is, the strategies of all unmanned devices are uniformly trained and executed by the central node. With the increase in the number of unmanned devices, the computational requirements of the central node increase exponentially. It is necessary to process the observation data of all devices simultaneously and generate strategy instructions, which puts extremely high requirements on computational resources and performance, resulting in excessive computational pressure. At the same time, the central node needs to communicate with each unmanned device frequently to achieve strategy distribution and status feedback, and the communication burden increases significantly in large-scale clusters, resulting in high latency problems, making it difficult for unmanned devices to respond to changes in the dynamic environment in a timely manner. In addition, the centralized architecture relies on a single central node. Once the node fails, the entire unmanned cluster will lose control. This risk of single-point failure severely limits the reliability and adaptability of unmanned clusters in task execution.

[0005] In addition, the existing simulation systems generally lack a detailed modeling of real task scenarios. For example, in the task of unmanned aerial vehicles, key elements such as the target distribution, hostile behavior, and the attack method of unmanned aerial vehicles in typical task scenarios such as offense and defense and reconnaissance are not fully reflected in the existing simulation systems. This deficiency directly affects the practical applicability and robustness of strategy training and is difficult to meet the application requirements in complex task environments. Summary of the Invention

[0006] To solve the above problems, the present invention discloses a system and method for constructing a large-scale unmanned cluster game strategy, which solves the limitation that traditional simulation systems are difficult to scale to large-scale unmanned clusters, overcomes problems such as excessive computational burden on the central node, sharp increase in communication delay, and single-point failure in the centralized strategy execution mode, can cope with the dimensionality disaster caused by the excessively high state space dimension in large-scale scenarios, and makes up for the deficiencies of existing simulation systems in modeling real task scenarios.

[0007] The specific solutions are as follows:

[0008] A large-scale unmanned cluster game strategy construction system, characterized by including a simulation environment module, a centralized training module based on the simulation environment module, and a distributed execution module; the simulation environment module is used to simulate the dynamic interaction environment of a large-scale unmanned cluster; the centralized training module adopts the cluster game adversarial reinforcement learning algorithm MAPPO, combines a policy network and a value network, and realizes the strategy optimization of a large-scale unmanned cluster through centralized training; the distributed execution module applies the strategy generated by the centralized training module to the large-scale unmanned cluster, enables each agent in the large-scale unmanned cluster to independently execute the strategy, and independently selects an action distribution based on its own local observation information.

[0009] Furthermore, the steps for the simulation environment module to simulate the dynamic interaction environment include:

[0010] (1) Initialize the simulation environment according to the set parameters, including the typical unmanned cluster game adversarial scenario, the initial state of the agents, the task objectives, and the ability parameters of the agents;

[0011] (2) Receive the actions of each agent output by the policy network and execute them;

[0012] (3) Update the environment state and the new local observation information of each agent according to the execution results of the agents' actions;

[0013] (4) Generate a reward value for each agent according to the task objective completion situation, resource consumption, and conflict resolution factors to guide strategy optimization;

[0014] (5) Render the simulation scenario in real time to show the interaction and dynamic adversarial effects of the large-scale unmanned cluster.

[0015] Furthermore, in the centralized training module, the policy network is responsible for generating an action distribution, the value network is used to evaluate the state value, and the two jointly drive the policy update; each agent uses local observation information to generate a specific policy distribution, and at the same time the value network uses global state information to optimize the policy evaluation; through the interaction results of the simulation environment, calculate the reward value and optimize the policy network and the value network.

[0016] Furthermore, the centralized training module enables each agent to better cooperate to complete tasks through a shared experience pool and joint optimization, while using parallel computing technology to improve the training speed and ensure the training efficiency in large-scale cluster scenarios.

[0017] Furthermore, in the distributed decision-making stage of the distributed execution module, the risk of single-point failure is avoided and the dependence on the central node is eliminated through the autonomous decision-making of the agents; in the distributed execution stage, it supports the rapid configuration and expansion of different task scenarios, improving the flexibility and applicability of the simulation system.

[0018] Furthermore, the large-scale unmanned cluster includes a large-scale unmanned aerial vehicle cluster, a large-scale unmanned ground vehicle cluster or a large-scale unmanned ship cluster.

[0019] A method for constructing a large-scale unmanned cluster game strategy includes the following steps:

[0020] S1. Select typical unmanned cluster game confrontation scenarios, such as offense and defense or reconnaissance, and configure the ability parameters of the unmanned aerial vehicles, such as speed, attack payload, etc.;

[0021] S2. Select the corresponding unmanned cluster game confrontation strategy model, and set the number of test rounds and other relevant parameters;

[0022] S3. Start the multi-round simulation test process. During the simulation operation, the system records and prints the key indicators of the unmanned cluster game confrontation in real time, and stores the simulation data and rendering playback in the specified path;

[0023] S4. After the test is completed, the user observes and evaluates the behavior and strategy effect of the unmanned cluster in depth by analyzing the simulation data and playback video.

[0024] The beneficial effects of the present invention are as follows:

[0025] 1. Support efficient simulation of multi-task scenarios: The present invention constructs a simulation system that supports two typical task scenarios of offense and defense and reconnaissance, covering the main task types in the current application of unmanned aerial vehicle clusters, significantly improving the practicality and robustness of the simulation system and the cluster game confrontation strategy.

[0026] 2. Achieve efficient simulation and verification of large-scale clusters: For unmanned clusters with a scale of 100 vs. 100, the present invention breaks through the scale limitation of the existing simulation system and can efficiently complete simulation verification. Through the optimized two-dimensional simulation model and distributed architecture design, while ensuring the simulation accuracy, the computational overhead is greatly reduced, meeting the simulation requirements of large-scale clusters.

[0027] 3. Solving the bottleneck of the centralized decision-making mode: The present invention adopts a distributed decision-making and execution architecture, effectively overcoming the computational bottleneck and communication delay problems of traditional centralized methods. The distributed training and execution mode not only improves the task reliability and avoids task failure caused by single-point failures, but also significantly enhances the real-time performance and flexibility of task completion.

[0028] 4. Supporting real-time rendering of large-scale scenarios: The simulation system has a real-time rendering function, which can intuitively display the dynamic behaviors of a large-scale unmanned cluster during task execution. This visualization ability not only helps with the debugging and optimization of strategies, but also provides important support for in-depth research on complex adversarial environments. Brief Description of the Drawings

[0029] Figure 1 It is a system framework diagram of the present invention.

[0030] Figure 2 It is a communication model diagram of the unmanned aerial vehicle in the present invention.

[0031] Figure 3 It is a usage method diagram of the present invention in the scenario of an unmanned aerial vehicle cluster. Detailed Embodiment

[0032] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0033] As Figure 1 shown, the present invention provides a large-scale unmanned cluster game strategy construction system, which mainly includes two core contents: constructing an efficient unmanned cluster simulation system and developing a cluster game confrontation strategy based on the paradigm of centralized training and distributed execution. Specifically, it includes a simulation environment module, a centralized training module, and a distributed execution module based on the simulation environment module; the simulation environment module is used to simulate the dynamic interaction environment of a large-scale unmanned cluster; the centralized training module adopts the cluster game confrontation reinforcement learning algorithm MAPPO, combines a policy network and a value network, and realizes the strategy optimization of a large-scale unmanned cluster through centralized training; the distributed execution module applies the strategy generated by the centralized training module to the large-scale unmanned cluster, enables each intelligent agent in the large-scale unmanned cluster to independently execute the strategy, and independently selects an action distribution based on its own local observation information.

[0034] In this embodiment, a large-scale unmanned aerial vehicle cluster is taken as an example for detailed introduction. The present application is equally applicable to the scenario of unmanned vehicles.

[0035] 1. Simulation Environment Module: The simulation environment module is the core foundation of the entire system and is used to simulate the dynamic interaction environment of a large-scale UAV cluster. First, it initializes the simulation environment according to the set parameters, including typical cluster game confrontation scenarios (such as offense and defense and reconnaissance), the initial states of UAVs, mission objectives, and the capability parameters of UAVs (such as speed, communication radius, and attack methods). Subsequently, it receives the actions of each agent output by the policy network (divided into heading actions, acceleration actions, and attack actions) and executes them. Based on the execution results of the UAV actions, it calculates the new environmental state and the new local observation information of each UAV. At the same time, it generates reward values for each agent according to factors such as mission objective completion, resource consumption, and conflict resolution to guide policy optimization. In addition, it can also render the simulation scene in real time to show the interaction and dynamic confrontation effects of the UAV cluster.

[0036] To simplify the modeling, the present invention sets the simulation system as a two-dimensional plane system. The flight of the UAV only needs to control the heading angle and speed, and its dynamic model is as follows:

[0037]

[0038] v t+1 = v t + a t ·d t

[0039]

[0040] Where, and are the heading angles of the agent at the current time step and the next time step respectively, v t and v t+1 are the speeds of the agent at the current time step and the next time step respectively, (x t , y t ) and (x t+1 , y t+1 ) are the positions of the agent at the current time step and the next time step respectively, p t and a t represent the heading angle change value and the acceleration value at the current time step respectively. In this model, p t and a t directly affect the motion state of the UAV and are output by the policy network. By adjusting these two parameters, the UAV can achieve basic control actions such as acceleration, deceleration, and turning, so as to simulate its motion behavior in various tasks. In implementation, p t and a t can be discretized to reduce the training difficulty of the policy network. Specifically, assume that the maximum acceleration value is 5m / s 2, the acceleration can be divided into 5 discrete actions, as shown in the following table:

[0041]

[0042] Similarly, by setting the maximum steering angle within a range, such as ±17 degrees, the steering angle can be divided into several discrete actions, for example, divided into 7 discrete actions, as shown in the following table:

[0043]

[0044] To solve the curse of dimensionality problem in large-scale scenarios, the present invention designs an efficient communication model (as Figure 2 shown). The communication range of the UAV is a circular radar scanning area with a radius of D. Only the information of the 5 nearest enemy aircraft and 5 friendly aircraft within the communication range is retained as local observations, significantly reducing the state space dimension. The specific observation information is as follows:

[0045]

[0046] In addition, we design different reward mechanisms for the attack-defense scenario and the reconnaissance scenario respectively. The design of the reward mechanism is crucial for the reinforcement learning of multi-agent systems because it directly affects the behavior and strategy optimization direction of the agents. In the attack-defense scenario, we use the multi-agent reinforcement learning method to train the strategy of the red UAVs. Its task objective is to protect the base, and all UAVs share the team reward. This team reward mechanism can promote the cooperation between UAVs and improve the overall task completion efficiency and effect. The specific reward design is shown in the following table:

[0047]

[0048] Among them, the time penalty prompts the red UAVs to complete the task as soon as possible and alleviates the problem of sparse rewards to a certain extent; the agent damage penalty encourages the red UAVs to avoid being attacked by the enemy as much as possible during the mission, keep themselves alive, and increase the probability of mission success; the penalty for the core area being attacked aims to protect the core area, prompting the agents to be more active in defensive actions during defense, preventing the enemy from breaking through the defense line and attacking the core area; the penalty for being far from the base aims to prompt the agents to stay within the base defense range and not leave the defense area due to chasing the enemy or other reasons, thus ensuring the safety of the base; the reward for damaging the blue agents aims to encourage the red UAVs to actively attack, destroy the enemy forces, and reduce the combat effectiveness of the enemy; the reward for approaching the base encourages the agents to stay close to the base during defense, form an effective defense formation, and improve the defense intensity of the base.

[0049] In the reconnaissance scenario, we adopt a multi-agent reinforcement learning method to train the strategy of the reconnaissance side (red side), whose goal is to complete the reconnaissance of all high-value areas and 70% of the ordinary areas of the blue side. In implementation, the area to be reconnoitered by the blue side is divided into grids of 50m x 50m, and the grid where the current position of the aircraft is located represents successful reconnaissance. The specific reward design is shown in the following table:

[0050]

[0051]

[0052] Among them, the reward for survival time encourages the agent to stay alive during the mission execution, increasing the mission's persistence and the probability of ultimate success; the reward for reconnoitering high-value areas aims to guide the agent to prioritize the reconnaissance of strategic areas, ensuring the collection of information with the optimal value and improving the efficiency and effectiveness of the reconnaissance mission; the reward for reconnoitering ordinary areas encourages the agent to comprehensively cover the reconnaissance mission area, ensuring the integrity and comprehensiveness of information collection within the mission area; the reward for damaging the blue side's agents encourages the agent to actively engage in confrontation, reducing the enemy's strength and enhancing the survival ability and mission success rate of itself and its teammates; the reward for approaching the reconnaissance area guides the agent to approach the mission target area, improving the execution effect and efficiency of the reconnaissance mission; the penalty for staying away from the reconnaissance area prompts the agent to stay within the mission scope, ensuring the reconnaissance coverage rate of the mission target area and the smooth completion of the mission; the penalty for entering the threat area prevents the agent from entering the dangerous area, thereby reducing the risk of being attacked or damaged and ensuring the smooth progress of the mission; the penalty for the agent being damaged prompts the agent to pay attention to avoiding the enemy's attack during the mission execution, reducing losses and increasing the survival rate and mission success rate; the penalty for repeated reconnaissance encourages the agent to explore new unknown areas, improving the reconnaissance coverage rate and the comprehensiveness of information collection; the winning reward encourages and ensures that the agent always takes the completion of the mission as the ultimate goal during the mission execution, motivating it to actively execute the mission and improving the quality and efficiency of mission completion.

[0053] 2. Centralized training module: The centralized training module adopts the game adversarial reinforcement learning algorithm (MAPPO), combines the policy network and the value network, and realizes the optimization of the cluster policy through centralized training. Among them, the policy network is responsible for generating the action distribution, and the value network is used to evaluate the state value. The two jointly drive the policy update. Each agent generates a specific policy distribution using local observation information, while the value network optimizes the policy evaluation using global state information. Through the interaction results of the simulation environment, the reward value is calculated and the policy network and the value network are optimized. In addition, through the way of sharing the experience pool and joint optimization, each UAV agent can better cooperate to complete the mission. At the same time, parallel computing technology is used to improve the training speed and ensure the training efficiency in large-scale cluster scenarios.

[0054] 3. Distributed Execution Module: The distributed execution module applies the strategies generated by centralized training to a large-scale cluster. Each agent independently executes the strategy, independently selects the action distribution based on its own local observation information, enhancing the reliability and flexibility of the task. During the distributed decision-making phase, the risk of single-point failure is avoided and the dependence on the central node is eliminated through the autonomous decision-making of the agents. During the distributed execution phase, it supports the rapid configuration and expansion of different task scenarios, improving the flexibility and applicability of the simulation system.

[0055] As Figure 3 shown, taking the UAV mission as an example, the implementation of the present invention mainly includes the following steps: First, select a typical cluster game confrontation scenario, such as offense and defense or reconnaissance, and configure the ability parameters of the UAV, such as speed and attack payload; then, select the corresponding cluster game confrontation strategy model, set the number of test rounds and other relevant parameters; then, start the multi-round simulation test process. During the simulation operation, the system will record and print the key indicators of the cluster game confrontation in real time, and at the same time store the simulation data and rendering playback in the specified path; after the test is completed, the user can deeply observe and evaluate the cluster behavior and strategy effect by analyzing the simulation data and playback video.

[0056] Aiming at the limitation that traditional simulation systems are difficult to be extended to large-scale unmanned clusters, the present invention designs and constructs an efficient simulation system that supports the coordination and confrontation of unmanned clusters with a scale of more than a hundred sorties; to overcome the problems such as the excessive computing burden of the central node, the sharp increase in communication delay, and single-point failure in the centralized strategy execution mode, a new strategy optimization method based on distributed training and execution is proposed, thus significantly improving the task reliability and execution efficiency of unmanned clusters; to cope with the curse of dimensionality caused by the excessively high state space dimension in large-scale scenarios, an improved reinforcement learning algorithm framework is developed, effectively reducing the state space dimension and improving the training efficiency and the actual effect of confrontation strategies; at the same time, to make up for the deficiencies of existing simulation systems in modeling real task scenarios, typical scenarios such as offense and defense and reconnaissance of UAV missions are introduced into the system, enhancing the practicality and robustness of the trained strategies in complex environments.

[0057] The technical means disclosed in the solution of the present invention are not limited to the technical means disclosed in the above embodiments, but also include technical solutions composed of any combination of the above technical features. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A large-scale unmanned cluster game strategy construction system, characterized in that: It includes a simulation environment module and a centralized training module and a distributed execution module based on the simulation environment module; the simulation environment module is used to simulate the dynamic interactive environment of a large-scale unmanned cluster; the centralized training module adopts the cluster game adversarial reinforcement learning algorithm MAPPO, and combines the strategy network and the value network to achieve strategy optimization of the large-scale unmanned cluster through centralized training; the distributed execution module applies the strategy generated by the centralized training module to the large-scale unmanned cluster, allowing each intelligent agent in the large-scale unmanned cluster to independently execute the strategy and independently select the action distribution based on its own local observation information.

2. A large-scale unmanned cluster game strategy construction system according to claim 1, characterized in that: The steps of simulating the dynamic interactive environment by the simulation environment module include: (1) Initialize the simulation environment according to the set parameters; (2) Receive the actions of each agent output by the policy network and execute them; (3) Update the state of the environment and the new local observation information of each agent based on the results of the agent's actions; (4) Generate reward values ​​for each agent to guide strategy optimization; (5) Real-time rendering of simulation scenes to demonstrate the interactive and dynamic confrontation effects of large-scale unmanned swarms.

3. A large-scale unmanned cluster game strategy construction system according to claim 1, characterized in that: In the centralized training module, the policy network is responsible for generating action distribution, and the value network is used to evaluate state value, and the two jointly drive policy updates; each agent generates a specific policy distribution using local observation information, while the value network optimizes policy evaluation using global state information; through the interaction results of the simulation environment, the reward value is calculated and the policy network and value network are optimized.

4. A large-scale unmanned cluster game strategy construction system according to claim 3, characterized in that: The centralized training module enables each intelligent agent to better collaborate to complete tasks by sharing the experience pool and jointly optimizing, while using parallel computing technology to increase the training speed and ensure training efficiency in large-scale cluster scenarios.

5. A large-scale unmanned cluster game strategy construction system according to claim 1, characterized in that: In the distributed decision-making stage, the distributed execution module avoids the risk of single point failure and eliminates the dependence on central nodes through the autonomous decision-making of the intelligent agent; in the distributed execution stage, it supports the rapid configuration and expansion of different task scenarios, and improves the flexibility and applicability of the simulation system.

6. A large-scale unmanned cluster game strategy construction system according to claim 1, characterized in that: The large-scale unmanned swarm includes a large-scale drone swarm, a large-scale unmanned vehicle swarm or a large-scale unmanned ship swarm.

7. A large-scale unmanned cluster game strategy construction method, characterized in that: The system according to any one of claims 1 to 6 comprises the following steps: S1. Select a typical unmanned cluster game confrontation scenario; S2. Select the corresponding unmanned swarm game confrontation strategy model and set the number of test rounds; S3, start multiple rounds of simulation test processes. During the simulation process, the system records and prints the key indicators of unmanned cluster game confrontation in real time, and stores the simulation data and rendering playback in the specified path; S4. After the test is completed, users can conduct in-depth observation and evaluation of the unmanned swarm behavior and strategy effects by analyzing simulation data and replaying videos.

Citation Information

Patent Citations

  • Distributed parallel multi-agent cooperative training system and method

    CN114707404A

  • Fixed-wing unmanned aerial vehicle cluster maneuvering cooperation method and system in simulation environment

    CN118655916A

  • Unmanned aerial vehicle cluster near-end strategy optimization collaborative confrontation method based on anti-fact baseline

    CN118778678A

  • Cross-domain heterogeneous unmanned cluster game confrontation strategy generation method and system

    CN118885000A

  • Unmanned aerial vehicle cluster collaborative confrontation decision-making method based on reinforcement learning

    CN119002521A

Cited By

  • Multi-agent game construction method and multi-agent system

    CN120560133A

  • Unmanned ship formation path planning training system based on MAPPO

    CN120972937A

  • Unmanned ship confrontation strategy optimization method and system based on reinforcement learning

    CN121723842A

  • Heterogeneous unmanned cluster game decision-making method and device

    CN122114184A