A control strategy optimization method and device for multi-agent games

By building a high-reality three-dimensional simulation environment and introducing Actor-Critic network, combining deep reinforcement learning algorithms, multi-agent game control strategies are optimized, and the problem of insufficient realism in the multi-agent training environment in the existing technology is solved, and the multi-agent training effect with high generalization and practicality is achieved, and VR interactive experience is supported.

CN116224799BActive Publication Date: 2025-07-18CHINA ACADEMY OF ELECTRONICS AND INFORMATION TECHNOLOGY OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310256330.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-16
Publication Date
2025-07-18
Estimated Expiration
2043-03-16

AI Technical Summary

Technical Problem

The existing multi-agent game training environment has low authenticity, insufficient integration of the simulation platform, and cannot effectively train the multi-agent synergy and game confrontation capabilities, and lacks VR-based interactive experience.

Method used

Build a high-reality three-dimensional visual simulation environment, introduce Actor and Critic networks, combine generalized advantage function estimation method for gradient training, optimize multi-agent game control strategies, and train drones and unmanned vehicle clusters through deep reinforcement learning algorithms to support the interactive experience of VR functions.

Benefits of technology

It realizes the high generalization and practicality of multi-agent games, improves the training effect and interactive experience, and can effectively train the coordination and game confrontation capabilities of drones and unmanned vehicle clusters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116224799B_ABST
    Figure CN116224799B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for optimizing a control strategy for multi-agent games, including: obtaining the observation of an unmanned aerial vehicle (UAV) and the joint observation of a UAV cluster; inputting the observation of UAV i at time t into the Actor network, obtaining the probability distribution of actions according to the policy, and outputting the control quantity of a fixed-wing UAV according to the Gaussian distribution; inputting the joint observation of the UAV cluster into the Critic network, obtaining the evaluation value of UAV i according to the policy, and determining the joint value of the UAV cluster; executing the action of UAV i obtained according to the policy; constructing a training set of a target size based on the joint observation, joint value, and reward of the UAV cluster; calculating the cumulative return after the current moment according to Generalized Advantage Estimation (GAE); performing gradient training based on the calculated cumulative return after the current moment; and outputting the control strategy. The embodiments of the present application introduce different learning algorithms and flexibly change the initial scenario, and have high generalization and practicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of unmanned aerial vehicles, and particularly to a method and device for optimizing control strategies for multi-agent games. Background Art

[0002] There are various multi-agent systems in nature, such as bird flocks, wolf packs, etc. Agents have obtained the ability to survive in nature through interacting with nature and learning from each other within and between populations. Multi-agent learning algorithms draw on the mechanisms by which individuals and groups in nature interact with the environment and evolve, and adapt to the environment by enabling agents to learn and improve from trial and error, thereby obtaining optimal group benefits.

[0003] In recent years, in-depth research on deep reinforcement learning has enabled the rapid development of multi-agent game training algorithms, which have also been widely applied in other fields. In virtual environments with a high degree of authenticity, both sides in the game confrontation face many problems, such as both sides being complex multi-agent systems with continuous action spaces, one side may have means such as radar / air defense that the other side cannot know, and the weather and lighting are constantly changing, which greatly increases the difficulty of learning.

[0004] Most of the current multi-agent game training environments on the market are based on real-time strategy (RTS) games and self-conceived scenarios. There are also some GIS-based simulation platforms that incorporate deep reinforcement learning algorithms for intelligent deduction and simulation.

[0005] Most of the current multi-agent game training environments on the market are based on real-time strategy (RTS) games and self-conceived scenarios. However, if one wants to apply intelligent algorithms to real-world environments, real-time strategy games are not of reference significance, and self-conceived scenarios are usually relatively single and lack elements.

[0006] GIS-based simulation platforms are usually used for the deduction of large-scale scenarios, focusing on the overall deduction results, but the description of the details of the environmental scenarios is not clear enough, and they do not pay attention to the specific behaviors and controls of small groups of agents, and cannot train the collaborative and game confrontation capabilities of multi-agents.

[0007] Existing multi-agent training simulation platforms based on intelligent algorithms are limited by the low authenticity of the simulation environment and the insufficient algorithm integration, and the application scenarios are very limited. The training effect and the generalization ability of the model are difficult to meet the application requirements. Moreover, there is no multi-agent training method based on VR function, resulting in insufficient interaction experience of the platform. Summary of the Invention

[0008] The embodiments of this application provide a method and device for optimizing control strategies for multi-agent games, which introduce different learning algorithms and flexibly change the initial scenario, and have high generalization and practicability.

[0009] An embodiment of the present application provides an optimization method for the control strategy of multi-agent games. The agents at least include unmanned aerial vehicles (UAVs), which is applied to optimize the control strategies of homogeneous and heterogeneous multi-agent games, and includes the following steps:

[0010] Pre-construct the required terrain model, environment model, and agent model;

[0011] Establish an Actor network for each UAV and a Critic network for the UAV cluster;

[0012] Obtain the observations of the UAVs and the joint observations of the UAV cluster;

[0013] Input the observation of UAV i at time t into the Actor network, and obtain the probability distribution of actions according to the policy π i ; Output the control quantity u of the fixed-wing UAV according to the Gaussian distribution t, i, and map the control quantity u t,i to the dynamic control range;

[0014] Input the joint observation O of the UAV cluster t into the Critic network, and obtain the evaluation value of UAV i according to the policy W and determine the joint value V of the UAV cluster;

[0015] Execute the action of UAV i according to the policy π i to obtain the reward obtained by UAV i for executing this action and the joint reward Rt of the UAV cluster, as well as the observation of UAV i at the (t + 1)-th moment and the joint observation O of the UAV cluster t+1 ;

[0016] Construct a training set of the target size based on the joint observations, joint values, and rewards of the UAV cluster;

[0017] Calculate the cumulative return after the current moment according to the Generalized Advantage Estimator (GAE) method;

[0018] Based on the calculated cumulative return after the current moment and the constructed training set, perform gradient ascent training to update the policy of the Actor network, and perform gradient descent training to update the evaluation index of the Critic network;

[0019] Output the optimized control strategy.

[0020] Optionally, it further includes the following initialization steps:

[0021] Initialize the parameters θ of the Actor network and the parameters of the Critic network such that θ(0) and satisfy the orthogonal initialization of the neural network;

[0022] Set the learning rate α;

[0023] Set the total number of steps Step of the deep reinforcement learning max and the required size batch_size of the training set;

[0024] Initialize the buffer D.

[0025] Optionally, the dynamic control range to which the control quantity u is mapped includes: roll angle (Roll), pitch angle (Pitch), yaw angle (Yaw), and thrust (Throttle). t,i

[0026] Optionally, it satisfies according to the Generalized Advantage Estimator (GAE) method:

[0027]

[0028] where γ is the discount factor and λ is the parameter of the GAE method;

[0029] Calculate that the cumulative return after the current moment satisfies:

[0030]

[0031] where T is the total number of steps Step max .

[0032] Optionally, based on the calculated cumulative return after the current moment and the constructed training set, perform gradient ascent training to update the policy of the Actor network, including: based on the calculated cumulative return after the current moment and the constructed training set, through the Adam optimizer, perform gradient ascent training using the following formula:

[0033]

[0034] where is the result calculated by the GAE method, S is the policy entropy, σ is the hyperparameter of the policy entropy, and n is the number of drones;

[0035] Perform gradient descent training using the following formula to update the evaluation index of the Critic network:

[0036] ​

[0037] The parameter ε of the Adam optimizer is 1e-5.

[0038] Optionally, the reward obtained by the drone executing the action satisfies:

[0039]

[0040] The embodiment of the present application also proposes a control strategy optimization device for multi-agent game, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the control strategy optimization method for multi-agent game as described above are implemented.

[0041] The embodiment of the present application also proposes a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the control strategy optimization method for multi-agent game as described above are implemented.

[0042] The embodiment of the present application can achieve flexible change of the initial scenario by introducing different learning algorithms, and has high generalization and practicability.

[0043] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0045] Figure 1 is the overall process schematic of the embodiment of the present application;

[0046] Figure 2 is the schematic diagram of the drone being controlled by the "brain" in the embodiment of the present application;

[0047] Figure 3 is the schematic diagram of the communication component architecture in the embodiment of the present application;

[0048] Figure 4 is the structure of the control center in the embodiment of the present application;

[0049] Figure 5 is the schematic diagram of the MAPPO algorithm training the drone collaborative reconnaissance process in the embodiment of the present application. Detailed implementation manners

[0050] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully communicated to those skilled in the art.

[0051] An embodiment of the present application provides a method for optimizing a control strategy for multi-agent games. The agents at least include unmanned aerial vehicles and may also include unmanned vehicles, and are applied to realize the optimization of control strategies for homogeneous and heterogeneous multi-agent games, such as Figure 1 shown, and includes the following steps:

[0052] In step S101, a required terrain model, environment model, and agent model are pre-constructed. For example, it can be implemented based on the Unreal Engine 4. Specifically, the terrain and environment can be constructed in the Blender software, and then in the Unreal Engine 4 (UE4), the terrain and environment models are imported. According to the map image data and terrain data, it is converted into a DEM (Digital Elevation Model) through KQGIS DESKTOP, and then the WorldCreator is used to generate an initial white model that can be stored in the UE. Then the white model is imported into UE4 to construct a simulation environment. Water systems are added through the components of UE4, an image map is loaded, and the addition of scene water systems and surface textures is completed through functions such as baking, edging, and normal blending. The residential area features are made in 3DMAX, then imported into the UE4 scene, the coordinates are corrected through the image base map, offset to the specified position, and then multi-object binding is performed to generate a new blueprint class. The angles and positions of the buildings are finely adjusted through the UE4 editor and the lighting function is constructed. The lighting effects of fixed light sources and dynamic light sources are set in the scene, the fineness of the static lighting is adjusted using the world lighting settings, a weather system plug-in is added to the scene, and the meteorological function is configured. The construction of the overall environment is completed through the above steps.

[0053] In some specific examples, the models to be simulated include quadrotor unmanned aerial vehicles, fixed-wing unmanned aerial vehicles, and cars. The above models can be made in 3DMAX, then converted into FBX format files and imported into UE4. In UE4, the components of the models are spliced, and the corresponding blueprint events are set. The flight control of the unmanned aerial vehicle of this model, the control simulation model of the unmanned vehicle, and the actions in different flight states are implemented through C++ and blueprints.

[0054] In the embodiments of the present application, a specialized communication component for the environment, the agent, and the intelligent learning algorithm is further described. This component is a plug-in for the UE4 platform. It maps the control functions of unmanned devices into the Python language through interfaces such as drones and unmanned vehicles. The user establishes an agent model on the Python side and trains the agent using deep reinforcement learning algorithms. The trained agent calls these control functions to control and query the status of drones and unmanned vehicles (the system interface uses the RPC protocol (Remote Procedure Call Protocol) based on TCP / IP, and develops the function of the program interface through the RPCLIB (RPC protocol library). When the scenario simulation model is started, the interface opens port 41451 to listen for incoming requests. The agent connects to this port through the system background, sends RPC calls using the MsgPack serialization format, and realizes data interaction through the interface function). The specific functions of this component are as follows: The images, infrared images, lidar, and other position, direction, and distance information obtained from the high-fidelity three-dimensional visualization simulation environment captured by the cameras mounted on drones and unmanned vehicles are transmitted to the input end of the intelligent algorithm through this component. After being processed and learned by the intelligent algorithm, control information is output, the control function is called to control the decision-making of drones and unmanned vehicles, and the decision is mapped into the system scenario environment in the high-fidelity three-dimensional visualization simulation environment. The system scenario environment mainly includes scenario elements such as meteorology, lighting, and volumetric clouds. The meteorological element is a separate subsystem, and lighting and volumetric clouds are plug-ins developed based on UE4. The interaction of the system scenario environment mainly uses the UE4 blueprint scripting system and interface functions. Through blueprint events, the changes of scenario meteorology, lighting, and volumetric clouds are driven, and function parameters are passed by the background call to realize the interaction of scenario environment information.

[0055] As Figure 2 shown, the behaviors of drones and unmanned vehicles are controlled by software called "brains", and drones and unmanned vehicles with similar functions are controlled by one "brain". There are three modes of the "brain". The first is the "external communication" mode, which can control the behaviors of drones and unmanned vehicles through external learning algorithms and continuously learn to give new action plans. The second is the "internal communication" mode, which can load the strategies trained by the intelligent algorithm into the "brain" for decision-making and no longer conduct training. The third is the "script" mode, which controls drones and unmanned vehicles through fixed scripts or strategies, does not directly involve training, but can be part of the game training.

[0056] When using the deep reinforcement learning algorithm for training, select the control mode of the "brain" as "external communication". Transmit the data required by the deep reinforcement learning algorithm from the high-fidelity three-dimensional visualization environment to the Python side in real time. Train with the deep reinforcement learning algorithm on the Python side, output control information, and transmit this control information back to the three-dimensional environment in real time to control the actions of drones and unmanned vehicles. After the training is completed, the network model based on the TensorFlow / Pytorch framework generated during training can be loaded into the "brain", change the control mode of the "brain" to "internal communication", and use the deep reinforcement learning model to control drones and unmanned vehicles for intelligent decision-making and task execution.

[0057] The specific working process of this component is as follows:

[0058] Figure 3 This is an example of the communication component architecture for the embodiments of this application. When using the deep reinforcement learning algorithm, during training in the "external communication" mode, in the high-fidelity three-dimensional visualization environment, drones and unmanned vehicles (i.e., agents) transmit their observations (such as images captured by cameras, distances detected by radars to targets or obstacles, data such as wind power affecting flight, etc.) to the corresponding "brains". Each "brain" transmits all the information collected and the rewards obtained after the drones and unmanned vehicles execute actions to the "control center". The "control center" uses the MsgPack-RPC protocol to transmit the data to the Python side through SocketIO in the external communication component. After training with the deep reinforcement learning algorithm on the Python side, control information is generated and returned to the "control center". The "control center" distributes the action information of each "brain". After obtaining the action information, the "brain" uses it to control the behavior of the corresponding drones and unmanned vehicles (agents), and saves the data of this round of training to the network model. In this way, the data in the model is continuously iteratively updated until the training ends. In the "internal communication" mode, when using the generated neural network model to execute tasks, the "brain" calculates the model according to the current states of drones and unmanned vehicles through the TensorFlow C++ API or Pytorch C++ API, makes decisions that conform to the model strategy, and then controls drones and unmanned vehicles to take actions to achieve intelligent decision-making. When using fixed strategies, scripts, or heuristic algorithms to control drones and unmanned vehicles to execute tasks or actions in the "script" mode, the strategies of drones and unmanned vehicles need to be set in advance, and during the task execution process, the strategies of drones and unmanned vehicles will not change due to unexpected events. The "internal communication" mode and the "script" mode can also be used during the training process of the "external communication" mode, that is, the trained network model and the fixed Heuristic strategy can both be used as part of the training to train the new strategies of agents, enhance the intelligence of decision-making, improve the autonomy of the unmanned system, and the control center architecture of the communication component is asFigure 4 as shown

[0059] The embodiments of the present application also propose to train the UAV and unmanned vehicle clusters through the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm, so that they can successfully perform reconnaissance tasks. As Figure 5 shown, in step S102, an Actor network with parameters θ is established for each UAV, and a Critic network with parameters In some embodiments, the following initialization steps are further included: initializing the parameters θ of the Actor network and the parameters of the Critic network such that θ(0) and satisfy the orthogonal initialization of the neural network;

[0060] Set the learning rate α;

[0061] Set the total number of steps Step of the deep reinforcement learning max and the required training set size batch_size. For example, the total number of steps Step max = 10e6, batch_size = 1024.

[0062] Initialize the buffer D.

[0063] In step S103, obtain the observations of the UAV and the joint observations of the UAV cluster. In some specific examples, the partial observations of the i-th UAV at the t-th moment can be defined as a matrix formed by concatenating the array converted from the photo taken by the camera and the parameters such as the distance, direction, altitude, speed, and attitude of the target detected by the radar, as well as the distances, directions, altitudes, speeds, and attitudes of the other friendly UAVs, then is the joint observation of the UAV cluster at the t-th moment.

[0064] In step S104, input the observations of the i-th UAV at the t-th moment into the Actor network, and obtain the probability distribution of the action according to the policy π i Output the control amount of the fixed-wing UAV according to the Gaussian distribution and map the control amount u to the dynamic control range. t,i In step S105, input the joint observation O

[0065] of the UAV cluster into the Critic network, and obtain the evaluation value t of the i-th UAV according to the policy W and determine the joint value of the UAV cluster

[0066] In step S106, perform according to policy π i to obtain the action of drone i, and obtain the reward obtained by drone i executing this action as well as the joint reward of the drone swarm and the observation of drone i at the (t + 1)-th moment and the joint observation of the drone swarm

[0067] In step S107, based on the joint observation, joint value, and reward of the drone swarm, construct a training set of the target size. Specifically, it can be according to the formula τ+ = [O t , V, u t , O t+1 , r], where r is the reward, and save the data into the temporary array τ. After completing the collection of experiences of a batch_size size, in step S108, according to the Generalized Advantage Estimator (GAE) method, calculate the cumulative return after the current moment.

[0068] In step S109, based on the calculated cumulative return after the current moment and the constructed training set, perform gradient ascent training and update the policy of the Actor network, and perform gradient descent training to update the evaluation index of the Critic network.

[0069] In step S110, output the optimized control strategy.

[0070] The embodiments of the present application can achieve flexible change of the initial scenario by introducing different learning algorithms, and have high generalization and practicability.

[0071] In some embodiments, the dynamic control range mapped by the control quantity u t,i includes: roll angle (Roll), pitch angle (Pitch), yaw angle (Yaw), and thrust (Throttle).

[0072] In some embodiments, the Generalized Advantage Estimator (GAE) method satisfies:

[0073]

[0074] where γ is the discount factor and λ is the parameter of the GAE method. In some specific examples, the value of γ can be 0.99, and the value of λ can be 0.95.

[0075] The cumulative return after the current moment is calculated to satisfy:

[0076]

[0077] where T is the total number of steps Step max .

[0078] In some embodiments, based on the calculated cumulative return after the current moment and the constructed training set, gradient ascent training is performed, and the strategy for updating the Actor network includes: saving the temporary array τ to the buffer D, and fetching a batch _ of data with a size of size. Based on the calculated cumulative return after the current moment and the constructed training set, through the Adam optimizer, gradient ascent training is performed using the following formula:

[0079]

[0080] where is the result calculated by the GAE method, S is the policy entropy, σ is the hyperparameter of the policy entropy, and n is the number of drones.

[0081] Gradient descent training is performed using the following formula to update the evaluation index of the Critic network:

[0082]

[0083] The parameter ε of the Adam optimizer is 1e-5.

[0084] In some embodiments, the reward obtained by the drone for performing an action satisfies:

[0085]

[0086] In some specific application scenarios, a multi-agent training method based on VR functions can also be developed to improve the interactive experience of the system. Users can switch to the VR mode to enter the first / third-person VR perspective and the scene roaming perspective during the agent training process, and realize the real-time monitoring of the agent training process. Specifically, the following steps can be adopted to achieve it:

[0087] Step 1: Load the SteamVR plugin in UE4, configure the SteamVR mirror mode, create a mode blueprint for the VR mode, set the world properties and mode properties, and configure the input and output of the VIVE device control.

[0088] Step 2: Use the motion component, add and configure the controller, and configure the handle control buttons and the interaction API.

[0089] Step 3: Configure the camera following model and use blueprint scripts in combination with scene elements for function development.

[0090] Step 4: Introduce the VRExpansionPlugin plugin, which covers some basic interaction logics and operations of VR, facilitating the extension and combination of complex VR functions.

[0091] Through the above steps, the development of the intelligent agent training method based on VR functions is achieved. Users can switch to the VR mode to enter the first / third-person perspective and the scene roaming perspective during the intelligent agent training process, realizing real-time monitoring of the intelligent agent training process, thereby enhancing the interaction experience of the system.

[0092] In some application examples, first, some VR-related environment setups are carried out on the basis of the existing project, install the Steam software, install SteamVR and connect the VR device, and perform some basic VR-related settings on the project. For example, enable forward shading, enable instantiated stereoscopic function, enable the SteamVR plugin in the project, recompile the lighting shader, etc., to complete the development of the entire system function.

[0093] The embodiment of this application is based on languages such as Python commonly used in the UE4 platform and intelligent algorithms. A scene simulating the real environment is built on the UE4 platform, and simulation models and control models of drones and unmanned vehicles are constructed. On this basis, the embodiment of this application designs a communication plugin based on the UE4 platform, which is used to generate intelligent agents using intelligent algorithms to control drones and unmanned vehicles to perform multi-agent game training in a specific scene, and give corresponding multi-agent game plans to guide the intelligent agents to complete tasks, realize the optimization of intelligent control strategies, and thus complete the specified tasks. Through the multi-agent proximal policy optimization algorithm, the drone and unmanned vehicle clusters are trained to enable them to successfully execute reconnaissance tasks. The designed system has excellent training effects and outstanding generalization performance.

[0094] The embodiment of this application also proposes a control strategy optimization device for multi-agent games, including a processor and a memory. A computer program is stored on the memory, and when the computer program is executed by the processor, the steps of the control strategy optimization method for multi-agent games as described above are implemented.

[0095] The embodiment of this application also proposes a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the control strategy optimization method for multi-agent games as described above are implemented.

[0096] It should be noted that, in this document, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device that includes a series of elements not only includes those elements but also other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device that includes such element.

[0097] The serial numbers of the embodiments of the present application described above are for description only and do not represent the superiority or inferiority of the embodiments.

[0098] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0099] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the purpose of the present application and the scope protected by the claims, can also make many forms, and these all fall within the protection scope of the present application.

Claims

1. A method for optimizing a control strategy of multi-agent games, characterized in that The agent at least includes an unmanned aerial vehicle (UAV), and is applied to optimize the control strategies of homogeneous and heterogeneous multi-agent games, including the following steps: Pre-construct the required terrain model, environment model, and agent model; Establish an Actor network for each UAV and a Critic network for the UAV cluster; Obtain the observations of the UAVs and the joint observations of the UAV cluster; The observed quantity of the UAV i at time t is input into the Actor network, and according to the policy the probability distribution of the action is obtained , and the control quantity of the fixed-wing UAV is output according to the Gaussian distribution , and the control quantity is mapped to the dynamic control range; Input the joint observation of the UAV swarm into the Critic network, and obtain the evaluation value of the UAV according to the policy , and determine the joint value of the UAV swarm ; Execute according to the policy Obtain the action of UAV i, obtain the UAV Reward obtained by executing this action And the joint reward of the UAV cluster , and at the Time UAV Observation value And the joint observation value of the UAV cluster ; Construct a training set of the target size based on the joint observations, joint values, and rewards of the UAV cluster; Calculate the cumulative return after the current moment according to the Generalized Advantage Estimation (GAE) method; Based on the calculated cumulative return after the current moment and the constructed training set, perform gradient ascent training, update the policy of the Actor network, and perform gradient descent training to update the evaluation index of the Critic network; Output the optimized control strategy.

2. The control strategy optimization method for multi-agent games according to claim 1, characterized in that It further includes the following initialization steps: Initialize the parameters of the Actor network and the parameters of the Critic network , such that and satisfy the orthogonal initialization of the neural network; Set the learning rate ; Set the total number of steps for deep reinforcement learning and the required size of the training set ; Initialize the buffer .

3. The control strategy optimization method for multi-agent game according to claim 1, characterized in that, The control quantity The mapped dynamic control range includes: roll angle, pitch angle, yaw angle, and propulsion force.

4. The control strategy optimization method for multi-agent games according to claim 2, characterized in that, The Generalized Advantage Estimation (GAE) method satisfies: Among them, is the discount factor, is a parameter of the GAE method; The calculation of the cumulative return after the current moment satisfies: Among them, is the total step length .

5. The control strategy optimization method for multi-agent game according to claim 4, characterized in that Based on the calculated cumulative return after the current moment and the constructed training set, perform gradient ascent training to update the policy of the Actor network, including: based on the calculated cumulative return after the current moment and the constructed training set, through the Adam optimizer, perform gradient ascent training using the following formula: Among them, , is the result calculated by the GAE method, is the policy entropy, is the hyperparameter of the policy entropy, is the number of UAVs; Perform gradient descent training using the following formula to update the evaluation index of the Critic network: Parameters of the Adam optimizer 。 6. The control strategy optimization method for multi-agent game according to claim 5, characterized in that, The reward obtained by the UAV executing the action satisfies:

7. A control strategy optimization device for multi-agent games, characterized in that, It includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the steps of the control strategy optimization method for multi-agent games described in any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the control strategy optimization method for multi-agent games described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Unmanned aerial vehicle cluster collaborative learning method based on multi-agent reinforcement learning

    CN112131660A

  • Multi-agent game training method and system in virtual environment

    CN114444716A