Intelligent Decision-making Method, Device and Storage Medium for Wargaming
By adopting a multi-agent reinforcement learning method with stratified training in wargame deduction scenarios, the problem of inefficiency of centralized training algorithms is solved, and efficient decision-making and training effects are achieved.
Patent Information
- Application Number
- CN202410282120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-03-12
AI Technical Summary
In wargame deduction scenarios, the centralized training algorithm directly based on the joint value function decomposition of all actions is poor in efficiency, and it is difficult to effectively cope with the complexity of high-dimensional state space, observation space and action space.
The multi-agent reinforcement learning intelligent decision-making method for war chess deduction is adopted, and the upper and lower-level decision-making networks are trained in layered. The upper-level decision network adopts centralized training, while the lower-level decision network adopts independent training, and uses RNN network to realize the interaction between the upper and lower-level decision networks.
The training efficiency in wargame deduction scenarios has been improved, effective decision-making in large-scale decision-making space has been achieved, and the application effect of multi-agent reinforcement learning in wargame deduction has been significantly improved.
Smart Images

Figure CN118001744B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of multi-agent reinforcement learning and wargaming, and particularly to a multi-agent reinforcement learning intelligent decision-making method, device, and storage medium for wargaming. Background Art
[0002] The problem of multi-agent intelligent decision-making is the core research content in the field of multi-agent systems. As an important example of the multi-agent intelligent decision-making problem, wargaming has received extensive attention in the research of the multi-agent system field. For such multi-agent game problems, they have a large-scale discrete decision space and a flexible and changeable environmental situation. How to make the reinforcement learning algorithm effectively cope with these challenges and thus be applied to such problems is an important research topic.
[0003] In recent years, many algorithms for multi-agent reinforcement learning based on a centralized training and distributed execution framework have attempted to solve the multi-agent game problem, such as a series of value function approximation algorithms based on joint action value function decomposition (QMIX). However, in the wargaming environment, due to the following challenges faced by the wargaming scenario, it is very difficult to directly represent the coordination relationship of each agent's actions through value function decomposition, and the centralized training algorithm based on the joint value function decomposition of all actions has poor efficiency; among them, the challenges faced by the wargaming scenario include:
[0004] Large-scale state space: The wargaming scenario generally includes nearly 5,000 hexagonal grids, and each of the two players has 6 operators. The state information of each operator and the capture and control points in the map are constantly changing. Roughly estimated, the state space of each operator is 5000 6×2 ·3 6×2+2 = 1.1677e51, and both the state space and the observation space of the operator are high-dimensional;
[0005] Complex action space: A general wargaming scenario includes 11 actions: null action (null), move, hide, occupy, shoot, guide-shoot, indirect-shoot, get-on, get-off, decompress, and stop the ongoing action (stop); some actions have many parameters and the parameter sizes are variable, and the entire action space is too large;
[0006] Long-term decision-making: Generally, each game in the wargaming scenario lasts at least 1,600 decision steps, and the scenarios caused by many non-shooting actions may occur long after the action is executed. Therefore, the long time series and action delay pose difficulties for the training of the reinforcement learning model. Summary of the Invention
[0007] An embodiment of the present invention provides a multi-agent reinforcement learning intelligent decision-making method, device and storage medium for wargaming, so as to solve the technical problem of poor efficiency of a centralized training algorithm directly based on the joint value function decomposition of all actions in a wargaming scenario with a high-dimensional state space, observation space and action space.
[0008] To achieve the above object, on the one hand, a multi-agent reinforcement learning intelligent decision-making method for wargaming is provided, including:
[0009] Step S1, model the wargaming scenario, including defining the set of agents in the wargaming scenario and modeling the state space, observation space and action space;
[0010] Step S2, according to the modeling of the wargaming scenario, construct an upper and lower layer hierarchical decision-making network for the wargaming scenario. Among them, the upper and lower layer hierarchical decisions are respectively regarded as Markov decision processes, and the decision results of the upper and lower layer hierarchical decision-making networks are used together to form the composite operations required by the environment; among them, the upper layer decision-making network is used to select available tasks for the agents from the task set; the lower layer decision-making network is used to select the actions to be executed by the agents according to the tasks selected by the upper layer decision-making network;
[0011] Step S3, train the hierarchical network of the upper and lower layer hierarchical decision-making network through reinforcement learning; among them, the upper layer decision-making network is trained in a centralized training manner for all multi-agents; the lower layer decision-making network is trained in an independent training manner for each agent;
[0012] Step S4, use the trained multi-agents to make battle decisions.
[0013] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, the modeling of the action space in step S1 includes:
[0014] Redefine the actions in the wargaming scenario based on the hierarchical actions of the upper and lower layers of tasks and behaviors; among them, the upper layer actions are tasks, and the tasks include: tasks based on hexagonal grids and tasks based on enemy operators; the lower layer actions are behaviors, and the behaviors are discrete actions, showing the moving direction of the agent at the current moment, including: six directions representing the surrounding hexagonal grids and stop.
[0015] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, the tasks based on the hexagonal grid include: an agent selects a grid from the candidate hexagonal grid set, and then executes the tasks related to the selected grid; among them, the tasks related to the selected grid include: getting on the vehicle, getting off the vehicle, seizing control or hiding at the selected grid; the tasks based on the enemy operator include: moving to a grid within a predetermined distance range from the enemy operator, and performing stopping, shooting or hiding.
[0016] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, a hierarchical decision-making network of upper and lower layers is constructed by interacting with the environment, where:
[0017] After the environment outputs the global system state s at the current moment t t , the control party obtains the visible original observation information from s t and, after structurally extracting the original observation information, transmits the observation information and the optional task set of each agent to the upper-layer decision-making network of each agent; then, the upper-layer decision-making network transmits the observation information of each agent and the tasks selected by the upper-layer decision-making network of each agent to the lower-layer decision-making network together; finally, the final action of the corresponding agent is obtained according to the tasks selected by the upper-layer decision-making network of the agent and the actions selected by the lower-layer decision-making network; the control party transmits the joint actions of all its agents back to the environment together, so that the environment advances according to the actions of both parties and gives the global system state s at the next moment t + 1 t+1 and transmits the joint return r of the control party at the current step t back to the hierarchical decision-making network of upper and lower layers.
[0018] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, the hierarchical decision-making network of upper and lower layers is implemented by an RNN network.
[0019] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, in step S3, the upper-layer decision-making network is trained by using a value decomposition method.
[0020] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, when the upper-layer decision-making network is trained by using a value decomposition method, the following first loss function is used:
[0021]
[0022] Among them, the first loss function updates the upper-layer decision-making network according to the samples in the upper-layer experience pool buffer B1 ∑ , and the content of each sample b1 ∑ is <s, o, g, r Σ , s′, o′, G′>, where o = {o 1 , o 2,…,o N} is the observation vector of all agents at the current moment, o′ is the observation vector of all agents at the next moment, N is the number of agents, g = {g 1 , g 2 ,…, g N} is the upper-layer joint action of all agents at the current moment, g′ is the upper-layer joint action of all agents at the next moment, s is the global system state at the current moment, s′ is the global system state at the next moment, r ∑ is the joint reward, G is the joint optional task set of all agents at the current moment, G′ is the joint optional task set of all agents at the next moment, θ in the first loss function is the parameter to be trained for the estimation network of the upper-layer joint value function, and in the first loss function is the parameter of the corresponding target network used to update the estimation network; γ is the discount factor; among them, the upper-layer decision network is updated by the Bellman update method.
[0023] Preferably, in the multi-agent reinforcement learning intelligent decision-making method, in step S3, the deep Q network is used to train the lower-layer decision network; among them, the following second loss function is used to train the lower-layer decision network:
[0024]
[0025] Among them, the second loss function updates the network according to the samples in the lower-layer experience pool buffer B2 L . Each sample b2 L has the content of <o, g, d, r L , o′>. Among them, o in the sample b2 L is the observation information of a single agent at the current moment, o′ is the observation information of a single agent at the next moment, g is the task required to be completed at the current moment selected by the upper-layer decision network of the corresponding single agent, g′ is the task required to be completed at the next moment by the upper-layer decision network of the corresponding single agent, d represents the action at the current moment corresponding to the lower-layer decision, d′ represents the action at the next moment corresponding to the lower-layer decision, and r L is the lower-layer reward for the corresponding single agent. Among them, r L adds an evaluation of the task completion degree of the upper-layer decision network. Different weight rewards are given according to the different task completion degrees of the upper-layer decision network; θ in the second loss function is the parameter of the estimation network, and in the second loss function is the parameter of the target network; γ is the discount factor.
[0026] On the other hand, a device for intelligent decision-making of multi-agent reinforcement learning for wargaming is provided, including a memory and a processor. The memory stores at least one segment of program, and the at least one segment of program is executed by the processor to implement the method described in any one of the above.
[0027] In another aspect, a computer-readable storage medium is provided, in which at least one segment of program is stored, and the at least one segment of program is executed by the processor to implement the method described in any one of the above.
[0028] The above technical solutions have the following technical effects:
[0029] The technical solution of the embodiment of the present invention proposes a hierarchical training method for reinforcement learning based on task-behavior (Task-Behavior Hierarchical Reinforcement Learning, TBHRL). Through centralized training of the upper-layer policy based on joint action value decomposition and independent training of the lower layer combined with expert knowledge, for the large-scale decision space scenario of wargaming, state features can be selectively extracted, and the overall decision can be explicitly divided into two layers of decision-making: task and behavior; using the solution of the embodiment of the present invention, for the decision-making in a certain scenario, the network framework first decides what task should be executed, and then decides the specific action required to execute a certain task in this scenario; the decision-making of the upper and lower layers is regarded as a Markov decision-making process respectively; thus, for complex training scenarios such as wargaming with high-dimensional state space, observation space and action space, the technical solution of the embodiment of the present invention improves the overall training efficiency and can achieve effective decision-making under a specific wargaming scenario. Description of the Drawings
[0030] Figure 1 It is a schematic flowchart of the intelligent decision-making method for multi-agent reinforcement learning for wargaming according to an embodiment of the present invention;
[0031] Figure 2 It is an exemplary reward function adopted in the method according to an embodiment of the present invention;
[0032] Figure 3 It is an overall framework diagram of the intelligent decision-making method for multi-agent reinforcement learning for wargaming according to an embodiment of the present invention;
[0033] Figure 4 It is a schematic diagram of the value decomposition scheme adopted in the upper-layer task decision network of the hierarchical network in the method according to an embodiment of the present invention;
[0034] Figure 5 It is a schematic diagram of the initial situation in the method according to an embodiment of the present invention;
[0035] Figure 6Schematic diagram of the structure of a multi-agent reinforcement learning intelligent decision-making device for wargaming according to an embodiment of the present invention. Detailed implementation manners
[0036] To further illustrate the embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, mainly used to illustrate the embodiments, and can be used to explain the operating principle of the embodiments in conjunction with the relevant descriptions in the specification. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are usually used to represent similar components.
[0037] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners.
[0038] Embodiment 1:
[0039] Figure 1 Schematic diagram of the process of a multi-agent reinforcement learning intelligent decision-making method for wargaming according to an embodiment of the present invention. As Figure 1 , the multi-agent reinforcement learning intelligent decision-making method in this embodiment includes:
[0040] Step S1, model the wargaming scenario, including defining the set of agents in the wargaming scenario and modeling the state space, observation space, and action space;
[0041] Step S2, according to the modeling of the wargaming scenario, construct an upper and lower layer hierarchical decision-making network for the wargaming scenario. Among them, the upper and lower layer hierarchical decisions are regarded as Markov decision processes respectively, and the decision results of the upper and lower layer hierarchical decision-making networks are used together to form the composite operations required by the environment; among them, the upper layer decision-making network is a task decision-making network, used to select available tasks for the agents from the task set; the lower layer decision-making network is a behavior decision-making network, used to select the behaviors to be executed by the agents according to the tasks selected by the upper layer decision-making network;
[0042] Step S3, train the hierarchical network of the upper and lower layer hierarchical decision-making networks through reinforcement learning; among them, the upper layer decision-making network is trained in a centralized training manner for all multi-agents; the lower layer decision-making network is trained in an independent training manner for each agent;
[0043] Step S4, use the trained multi-agents to make battle decisions.
[0044] In one implementation, modeling the wargaming scenario in step S1 includes: letting the set of all N agents controlled by the decision-making model, that is, the decision-making network, be I = {1, 2,..., N}, and the set of all N agents controlled by the opponent be I -={-1, -2, ..., -N}. From the perspective of the overall environment, i.e., the war game simulation platform environment, the entire decision-making form is still a five-tuple<S,A,T,R,γ> The MDP is a Markov decision process. In each decision step of the deduction process: the environment gives the current global system state s; consider each agent i∈I on the other side to obtain its observation o i , and give an action a to be performed according to the decision model i , the joint action transmitted to the environment is a = {a 1 ,a 2 ,…,a N}∈A, the joint action passed by the opponent to the environment is a - ={a -1 ,a -2 ,…,a -N}∈A - , the environment executes the actions passed to the environment by both parties and follows the transfer matrix T:S×A×A - →S gives the next global system state s′ of the environment, and the environment calculates the reward value r by comparing the two global system states s and s′, and transmits the reward value r back to its decision model. Among them, γ is the discount factor, γ∈[0,1].
[0045] In one implementation, modeling the state space and observation space includes: setting the global state space s = {σ E ,σ 1 ,σ 2 ,…,σ N ,σ -1 ,…,σ -N}, where σ E Refers to all information on the field except the chess pieces of both sides, also known as map information; i Represents all the information contained in the chess piece itself, also known as chess piece information. Among them, map information includes: attribute information of all grids on the map; the specific location of the control point grid; and the subordinate information belongs to which camp. Chess piece information includes: the original dynamic attributes and intrinsic attribute information of the chess piece; among them, for the controlling agent i∈I, its observation information can be expressed as o i ={σ i ,σ E ,σ 1 ,σ 2 ,…,σ N ,σ -1 ,…,σ -N}, when the enemy operator is invisible, the enemy information in the observation information is based on the enemy's last appearance information, using all observation information o = {o 1 ,o 2 ,…,o N} to replace the joint action - observation history τ = {τ 1 , τ 2 , …, τ N}.
[0046] In one implementation, modeling the action space includes: re - defining the actions in the wargame scenario based on a hierarchical action with two layers of task and behavior; where the upper - layer action is a task, and the tasks include two categories: tasks based on hexagons and tasks based on enemy operators. The lower - layer action is a behavior, and the behavior is a discrete action, showing the moving direction of the agent at the current moment, including: six directions representing the surrounding hexagons and stop. Among them, the tasks based on hexagons include: the agent selects a grid from the set of candidate hexagons, and then executes the tasks related to the selected grid; where the tasks related to the selected grid include: getting on the vehicle, getting off the vehicle, seizing and controlling, or hiding at the selected grid; the tasks based on enemy operators include: moving to a grid within a predetermined distance range from the enemy operator, and performing stop, shooting, or hiding.
[0047] In one implementation, modeling the wargame scenario also includes: modeling the reward function. Among them, Figure 2 Exemplarily shows multiple factors that make up the reward function. Such as Figure 2 , the multiple factors that make up the reward function include: events and the corresponding rewards for the events. For example, for the event of "being closer to the multi - control point than the previous state", the corresponding reward is +0.1, etc.
[0048] Figure 3 This is the overall framework diagram of the multi - agent reinforcement learning intelligent decision - making method for wargame in an embodiment of the present invention. Figure 3 Among them: on the right is the wargame platform environment, i.e., the engine, and on the left is the framework of the entire upper - and lower - layer hierarchical decision - making network; the middle part on the left is divided into two areas, the upper and lower columns, which are respectively the upper - layer decision - making network architecture and the lower - layer decision - making network architecture of the hierarchical decision - making. The entire upper - layer decision - making network architecture of the model includes the upper - layer decision - making networks corresponding to each agent and a mixing network for mixing the values of the actions output by the upper - layer decision - making networks of each agent. For details, see Figure 4 . The entire lower - layer decision - making network architecture of the model also includes the lower - layer decision - making networks corresponding to each agent. The area of the upper - left, i.e., the dotted - line box part at the top left of the left side, and the area of the lower - left, i.e., the dotted - line box part at the bottom left of the left side, are the information adaptation and transformation parts between the network and the environment. The information adaptation and transformation part in the upper - left is used to process the original information given by the environment, such as the original observation information and the reward information, into a data format that the hierarchical decision - making network can process; the information adaptation and transformation part in the lower - left is used to process and modify the information output by the hierarchical decision - making network, such as the joint actions of the agents, into a data format that the platform environment can process or execute.
[0049] Such asFigure 3 In the multi-agent reinforcement learning intelligent decision-making method of one embodiment of the present invention, the upper and lower hierarchical decision networks are constructed to interact with the platform environment as follows, wherein the interaction in step t is taken as an example for explanation.
[0050] At step t: the environment gives the global system state s at this moment t ; The controlling party from s t Obtain its visible observation information, and extract the information structuredly through the content required by the hierarchical network, and convert the observation information of each agent i∈I after structured extraction and optional task collection Then, the observation information is passed to the upper decision network of each agent. and the task selected by agent i in its upper decision network Finally, according to the task selected by agent i in its upper decision network and the behavior chosen by the decision network below it The final action of the agent is obtained according to the preset rules in the action modeling The controller then takes the combined action of all its agents The environment is sent back to the environment; the environment advances according to the rules based on the actions of both parties and gives the global system state s at the next moment t+1 , and the joint reward r of the controller of this step, i.e. the current step t Pass back to the hierarchical decision network.
[0051] In one implementation, both the upper and lower hierarchical decision networks are implemented through RNN networks.
[0052] In one implementation, performing centralized training on the upper-level task decision network includes: training the upper-level decision network using a value decomposition method. Figure 4 FIG. 1 shows a schematic diagram of a value decomposition scheme used in the upper task decision network of the hierarchical network. Figure 4 , the value corresponding to the task actually selected by all agents through their corresponding upper decision network, that is, the upper strategy network They are input together into the mixing network for mixing to output the value Q of the joint action Σ Among them, the parameters of the hybrid network are obtained after the global state s is input into the super network. The strategy training of the upper decision network only needs to minimize the first loss function:
[0053]
[0054] The first loss function is based on the upper experience pool buffer B1 Σ The samples in update the upper network, and each sample b1 ΣThe content is <s,o,g,r Σ ,s′,o′,G′>, where o = {o 1 ,o 2 ,…,o N} is the observation vector of all agents at the current moment, o′ is the observation vector of all agents at the next moment, g = {g 1 ,g 2 ,…,g N} is the joint action of the upper layer of all agents at the current moment, s is the global system state, r ∑ is the joint reward, g′ is the joint action of the upper layer of all agents at the next moment, G is the joint optional task set of all agents at the current moment, and G′ is the joint optional task set of all agents at the next moment; θ refers to the parameters to be trained for the estimation network of the upper layer joint value function. The estimation network includes the entire upper layer decision network and the mixing network for mixing; refers to the parameters of the corresponding target network used to update the estimation network. The entire update method adopts the Bellman update form, that is, it is obtained by adding the reward corresponding to a certain state-action and the maximum value among all state-action values at the next moment. The addition after conversion is to add after multiplying by the corresponding discount factor. γ represents the discount factor. In a specific implementation, both the mixing network and the super network can be implemented by an RNN network.
[0055] In one implementation, the lower layer decision network is independently trained for each agent by using a deep Q network for training. Among them, the task that the lower layer decision network needs to decide is a task targeted at hexagonal grids. In one example, it is the moving direction of the agent at the current moment. The behaviors or tasks that can be selected include: six directions representing the surrounding hexagonal grids and stopping, which can actually be considered as solving the path planning problem. The bottom reward function r L adds an evaluation of the completion degree of the upper layer task, and gives rewards with different weights according to the different completion degrees of the upper layer task. Among them, the training loss function of the entire bottom layer network is the following second loss function:
[0056]
[0057] The loss function updates the lower layer decision network according to the samples in the lower layer experience pool buffer B2 L . Among them, the content of each sample b2 L is <o,g,d,r L, o′>, where: o is the current observation information of a certain agent, and o′ is the observation information of the agent at the next moment; g is the currently required task selected by the upper-level decision network of the agent, and g′ is the required task of the upper-level decision network of the agent at the next moment; among them, the required tasks are consistent in the previous and next moments, and samples of task switching between the previous and next moments will not be saved in the lower-level experience pool to ensure that sub-problems for different purposes will not interfere with each other; d represents the current action corresponding to the next decision, and d′ represents the action at the next moment corresponding to the lower-level decision, r L is the bottom-layer reward for the agent; θ is the parameter of the estimation network that needs to be updated, and is the parameter corresponding to the target network.
[0058] Embodiment 2:
[0059] This embodiment takes a certain military wargame event as an example to introduce the specific application form of the present invention in military wargames.
[0060] Introduction to the wargame scenario
[0061] Name: Encounter battle scenario for company-level water network paddy field terrain
[0062] The Red Force's troop deployment is shown in Table 1:
[0063]
[0064]
[0065] Table 1
[0066] The Blue Force's troop deployment is shown in Table 2:
[0067]
[0068] Table 2
[0069] The capture and control points are shown in Table 3:
[0070] Capture control point Score Position Main capture control point 80 3636 Secondary capture control point 50 4039 Total 130 ——
[0071] Table 3
[0072] The Red Force's equipment is shown in Table 4:
[0073]
[0074]
[0075] Table 4 The Blue Force's equipment is shown in Table 5:
[0076]
[0077] Table 5
[0078] The initial situation is Figure 5 shown.
[0079] Scenario Introduction:
[0080] The Red Army's combined battle group was ordered to conduct a roundabout penetration into the enemy's depth, occupy favorable terrain in the depth and deploy key points for defense. The Blue Army assembled its forces to conduct offensive operations, seize key points, intercept the Red Army's combined battle group's roundabout penetration, and improve the overall situation.
[0081] Training process:
[0082] 1. Modeling war game scenarios based on current assumptions
[0083] (1) State space modeling
[0084] ① Map information: Add the hexagonal terrain, height difference and other map information of the two sides' chess pieces in the battle area to the state space;
[0085] ② Our status: add our status such as the current health percentage, current position, chess piece type, current mobility status, relative distance from the capture point, weapon cooling time, etc. of each chess piece to the status space;
[0086] ③Enemy status: add enemy status such as the relative position of each enemy piece from the control point, whether our side is visible, the current health percentage of the enemy piece, the current position, the piece type, the current mobility status, etc. to the status space.
[0087] (2) Observation space modeling
[0088] ① Self-observation: add the current chess piece health percentage, current position, chess piece type, current mobility status, relative distance to the control point, weapon cooling time and other self-status to the observation space;
[0089] ② Friendly observation: Add friendly status such as health percentage, current position, piece type, current mobility status, relative distance to the current piece, etc. of other friendly pieces in our camp except our own to the observation space;
[0090] ③Enemy observation: add enemy states such as the relative distance of the enemy pieces that the current piece can observe, the percentage of health points of the pieces, the current position, the type of pieces, the current mobility state, and the relative distance from the current piece to the observation space.
[0091] (3) Action Space Modeling
[0092] ①Upper level action behavior:
[0093] 1) Hexagon - based actions: The agent selects a cell from the set of candidate hexagons and then executes a task related to that cell. The tasks are usually maneuvering to the selected cell, boarding or alighting from a vehicle at the selected cell, seizing control, or taking cover.
[0094] 2) Enemy - operator - based actions: The tasks are usually moving to a cell at an appropriate distance from the operator and then stopping, shooting, or taking cover. In a specific implementation, the above - mentioned appropriate distance can be determined by comparing the distance to the enemy operator with a pre - set distance threshold.
[0095] ② Lower - layer action behavior: It represents the moving direction of the agent at the current moment. It is a discrete action, representing the six directions of the surrounding hexagons and stopping respectively.
[0096] (4) Modeling of the reward function
[0097] The reward function consists of multiple factors as Figure 2 shown.
[0098] 2. Constructing the upper - layer and lower - layer hierarchical decision - making network
[0099] The architecture of the constructed upper - layer and lower - layer hierarchical decision - making network is specifically as Figure 3 shown. For its specific composition and operation, please refer to the written description in Embodiment 1 for Figure 3 this, and it will not be elaborated here.
[0100] 3. Centralized training of the upper - layer, i.e., the high - level decision - making network
[0101] As Figure 4 shown, a value - decomposition scheme is adopted in the upper - layer decision - making network of the hierarchical network. For the specific value - decomposition scheme used, please refer to the written description in Embodiment 1 for Figure 4 this, and it will not be elaborated here.
[0102] 4. Independent training of the lower - layer, i.e., the bottom - layer behavior decision - making network
[0103] The lower - layer behavior decision - making network is trained using a deep Q - network. For the specific training scheme and loss function, please refer to the written description of the training of the lower - layer decision - making network in Embodiment 1, and it will not be elaborated here.
[0104] 5. Model training
[0105] Train until the model converges, and use the trained multi - agent to make battle decisions.
[0106] In specific practice, applying the agent obtained by training using the multi - agent reinforcement learning hierarchical training method based on the embodiments of the present invention to a military chess competition has achieved good results, thus verifying the effectiveness of the method of the present invention.
[0107] Embodiment 3:
[0108] The present invention also provides a multi-agent reinforcement learning intelligent decision-making device for wargaming, such as Figure 6 shown. The device includes a processor 601, a memory 602, a bus 603, and a computer program stored in the memory 602 and executable on the processor 601. The processor 601 includes one or more processing cores. The memory 602 is connected to the processor 601 through the bus 603. The memory 602 is used to store program instructions. When the processor executes the computer program, it implements the steps in the above method embodiment of Embodiment 1 of the present invention.
[0109] Furthermore, as an executable solution, the multi-agent reinforcement learning intelligent decision-making device may be a computer unit, and this computer unit may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer unit may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the above composition structure of the computer unit is only an example of the computer unit and does not constitute a limitation on the computer unit. It may include more or fewer components than the above, or combine some components, or different components. For example, the computer unit may further include input / output devices, network access devices, a bus, etc. The embodiments of the present invention do not make limitations in this regard.
[0110] Furthermore, as an executable solution, the so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer unit and connects various parts of the entire computer unit using various interfaces and lines.
[0111] The memory can be used to store the computer program and / or modules. By running or executing the computer program and / or modules stored in the memory, and invoking the data stored in the memory, the processor realizes various functions of the computer unit. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash device, or other volatile solid-state storage devices.
[0112] Embodiment 4:
[0113] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method in the embodiments of the present invention are realized.
[0114] If the modules / units integrated in the computer unit are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0115] Although the present invention is specifically shown and described in conjunction with the preferred embodiments, those skilled in the art should understand that various changes can be made to the present invention in form and detail without departing from the spirit and scope of the present invention defined by the appended claims, and all are within the protection scope of the present invention.
Claims
1. A multi-agent reinforcement learning intelligent decision-making method for war game simulation, characterized in that: include: Step S1, modeling the war game scenario, including defining a set of intelligent agents in the war game scenario and modeling the state space, observation space and action space; Among them, modeling the action space includes: The actions in the war game scenario are redefined based on hierarchical actions of tasks and behaviors. The upper-level actions are tasks, which include: hexagonal grid-based tasks and enemy operator-based tasks; the lower-level actions are behaviors, which are discrete actions that show the moving direction of the agent at the current moment, including: six directions representing the surrounding hexagonal grids and stop; Step S2, based on the modeling of the war game scenario, constructing an upper and lower hierarchical decision network of the war game scenario, wherein the upper and lower hierarchical decisions are respectively regarded as Markov decision processes, and the decision results of the upper and lower hierarchical decision networks are used together to form the composite operation required by the environment; wherein the upper decision network is used to select available tasks for the intelligent agent from the task set; and the lower decision network is used to select an action to be performed by the intelligent agent according to the task selected by the upper decision network; The upper and lower hierarchical decision network is constructed by interacting with the environment, wherein: Output the global system state s at the current time t in the environment t After that, the controller t The visible raw observation information is obtained from the control, and after structured extraction of the raw observation information, the observation information and optional task set of each agent are transmitted to the upper decision network of each agent; then, the upper decision network transmits the observation information of each agent together with the task selected by the upper decision network of each agent to the lower decision network; finally, the final action of the corresponding agent is obtained according to the task selected by the upper decision network of the agent and the behavior selected by the lower decision network; the controller transmits the joint action of all its agents back to the environment, so that the environment can advance according to the actions of both parties and give the global system state s at the next time t+1 t+1 , and the joint reward r of the current step controller t Returning the upper and lower hierarchical decision networks; Step S3, performing hierarchical network training on the upper and lower hierarchical decision networks by reinforcement learning; wherein the upper decision network is trained for all multi-agents in a centralized training manner; and the lower decision network is trained for each agent in an independent training manner; Step S4, using the trained multi-agent to make combat decisions.
2. The multi-agent reinforcement learning intelligent decision-making method according to claim 1, characterized in that: The hexagonal grid-based tasks include: the intelligent agent selects a grid from a set of candidate hexagonal grids, and then performs tasks related to the selected grid; wherein the tasks related to the selected grid include: getting on or off a vehicle, taking control or hiding at the selected grid; the enemy operator-based tasks include: moving to a grid within a predetermined distance range from the enemy operator, and stopping, shooting or hiding.
3. The multi-agent reinforcement learning intelligent decision-making method according to claim 1, characterized in that: The upper and lower layer hierarchical decision network is implemented through an RNN network.
4. The multi-agent reinforcement learning intelligent decision-making method according to claim 1, characterized in that: In the step S3, the upper layer decision network is trained by using a value decomposition method.
5. The multi-agent reinforcement learning intelligent decision-making method according to claim 4, characterized in that: When the upper decision network is trained by value decomposition, the following first loss function is used: Among them, the first loss function is based on the upper experience pool buffer B1 ∑ The samples in the upper decision network are updated, and each sample b1 ∑ The content is <s,o,g,r Σ ,s′,o′,G′>, where o={o1,o2,…,o N } is the observation vector of all agents at the current moment, o′ is the observation vector of all agents at the next moment, N is the number of agents, g={g1,g2,…,g N } is the upper-level joint action of all agents at the current moment, g′ is the upper-level joint action of all agents at the next moment, s is the global system state at the current moment, s′ is the global system state at the next moment, r ∑ is the joint reward, G is the joint optional task set of all agents at the current moment, G′ is the joint optional task set of all agents at the next moment, θ in the first loss function is the parameter to be trained for the estimation network of the upper joint value function, and is a parameter of the target network corresponding to the updated estimation network; γ is a discount factor; wherein the upper decision network is updated by the Bellman update method.
6. The multi-agent reinforcement learning intelligent decision-making method according to claim 1, characterized in that: In step S3, the lower decision network is trained using a deep Q network; wherein the lower decision network is trained using the following second loss function: Among them, the second loss function is based on the lower layer experience pool buffer B2 L The samples in update the network, where each sample b2 L The content is <o,g,d,r L ,o′>, where sample b2 L Where o is the observation information of a single agent at the current moment, o′ is the observation information of a single agent at the next moment, g is the task required to be completed at the current moment selected by the upper decision network of the corresponding single agent, g′ is the task required to be completed at the next moment by the upper decision network of the corresponding single agent, d represents the action of the lower decision at the current moment, d′ represents the action of the lower decision at the next moment, r L is the lower layer reward for the corresponding single agent, where r L The evaluation of the completion of the upper decision network task is added, and different weight rewards are given according to the completion of the uploaded decision network task; θ in the second loss function is the parameter of the estimated network, and are the parameters of the target network; γ is the discount factor.
7. A multi-agent reinforcement learning intelligent decision-making device for war game simulation, characterized in that: The method comprises a memory and a processor, wherein the memory stores at least one program, and the at least one program is executed by the processor to implement the method according to any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that: The storage medium stores at least one program, and the at least one program is executed by a processor to implement the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent war game deduction method based on distributed reinforcement learning
CN113222106A
Layered multi-agent reinforcement learning method for multi-element joint command and control
CN114330651A