A Group Adversarial System Based on Hierarchical Reinforcement Learning

Through a group adversarial system with layered reinforcement learning, the upper-level macro strategy network and the lower-level micro action network are used to solve the environmental instability and locality of observation information in the multi-agent system, and more accurate decision-making action generation is achieved, which is suitable for the multi-agent collaborative game adversarial environment.

CN115068953BActive Publication Date: 2025-07-22ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210657186.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-07-22
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

The problems of environmental instability, locality of observation information and individual task consistency of multi-agent systems are difficult to solve in the multi-agent collaborative game confrontation environment, resulting in difficult and inaccurate decision-making.

Method used

A group adversarial system based on hierarchical reinforcement learning, including the upper-level macro strategy network and the lower-level micro action network, is adopted, and the improved Q-mix algorithm and DQN algorithm are used to process global and local decisions respectively, and more accurate decision actions are generated through attention mechanism and state information decomposition.

Benefits of technology

In the multi-agent collaborative game confrontation environment, the agent can quickly and accurately generate decision-making actions, improving the accuracy and efficiency of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115068953B_ABST
    Figure CN115068953B_ABST
Patent Text Reader

Abstract

The present invention discloses a group confrontation system based on hierarchical reinforcement learning, which includes an upper-layer macro policy network and a lower-layer micro action network; the upper-layer macro policy network includes multiple policy networks and a hybrid network adopted by multiple agents, and each policy network is used to calculate and output a predicted sub-goal at the current moment according to the observed state at the current moment and the sub-goals of the previous multiple time steps; the hybrid network is used to calculate and output a macro total goal as the sub-goal of each agent at the next moment according to the full environmental state information and the predicted sub-goals output by the policy networks adopted by each sub-agent; the lower-layer micro action network includes multiple DQNs adopted by multiple agents, and each DQN is used to calculate and output a decision-making action according to the observed state at the current moment and the sub-goal at the current moment. In this system, agents can generate more accurate decisions while taking into account the macro total goal and individual sub-goals, and it is applicable to the game environment of multi-agent collaborative game confrontation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence and intelligent game confrontation, and particularly relates to a group confrontation system based on hierarchical reinforcement learning. Background Art

[0002] The theory of reinforcement learning originated from the cross-research of cybernetics, statistics, computer science and other disciplines. Reinforcement learning mainly uses reward signals or penalty signals to motivate decision-making individuals without the process of obtaining specific experience or data, so that the decisions of decision-making individuals are iteratively improved in the correct direction. Reinforcement learning explores the environment and uses existing strategies to continuously iteratively improve the strategies. Due to its characteristics of directly interacting with the environment for learning and not requiring existing data, reinforcement learning has been widely used in many fields, especially in the field of game agents. The performance of agents using reinforcement learning has far exceeded that of humans and other agents based on deep learning algorithms in many scenarios.

[0003] Currently, whether it is a game environment or a real-world problem, the characteristics of decentralized collaboration are prominent. Compared with reinforcement learning on single-agent systems, multi-agent systems have more extensive application value and more profound theoretical value. Reinforcement learning on multi-agent systems applied to game environments will encounter the following problems that have never been faced in the research of reinforcement learning on single-agent systems:

[0004] Non-stationarity of multi-agent environment: The evolution of the multi-agent system environment is not affected by a single agent, but by the combined actions of all agents. Therefore, the evolution of the environment is not determined by a single agent alone. For a single agent, the actions of other agents may be unpredictable. Therefore, the multi-agent system environment is unstable for a single agent.

[0005] Locality of observation information: In many agent research environments, a single agent cannot directly detect all environmental states, and cannot directly obtain the observations and reward signals of non-self agents.

[0006] Individual task consistency: Each agent may have different goals and tasks in the environment. The optimal reward for each agent may be the global optimal reward or the optimal individual reward of each agent itself.

[0007] Environmental complexity: The complexity of the environment in a multi-agent system increases with the number of agents.

[0008] To address the above problems and challenges, the research on multi-agent reinforcement learning has shifted from the early deployment of single-agent reinforcement learning algorithms in multi-agent systems to the direct study of collaborative decision-making and game confrontation in multi-agent systems. Among them, the collaborative decision-making problem has received extensive attention and applications due to its important position in game agent control and robot control. Moreover, the problem of sparse rewards in multi-agent environment systems is also worthy of attention. Summary of the Invention

[0009] In view of the above problems, the object of the present invention is to provide a group confrontation system based on hierarchical reinforcement learning, in which each agent can generate more accurate decisions while taking into account the macroscopic overall goal and individual sub-goals, so as to be applicable to the game environment of multi-agent collaborative game confrontation.

[0010] To achieve the above object of the invention, an embodiment provides a group confrontation system based on hierarchical reinforcement learning, comprising:

[0011] It includes an upper-layer macroscopic policy network and a lower-layer microscopic action network;

[0012] The upper-layer macroscopic policy network includes multiple policy networks and a mixing network adopted by multiple agents. Each policy network is used to calculate and output the decomposed value function at the current moment based on the observed state at the current moment and the sub-goals in the previous multiple time steps, and determine the predicted sub-goal from the decomposed value function; the mixing network is used to calculate and output the joint value function according to the full environment state information and the decomposed value functions output by the policy networks adopted by each sub-agent, and determine the macroscopic overall goal as the sub-goal of each agent at the next moment from the joint value function;

[0013] The lower-layer microscopic action network includes multiple DQNs adopted by multiple agents. Each DQN is used to calculate and output a decision action according to the observed state at the current moment and the sub-goal at the current moment.

[0014] In one embodiment, the policy network includes an MLP, a GRU, and an MLP connected in sequence, which are used to calculate the observed state at the current moment and the sub-goals in the previous multiple time steps as input, so as to output the decomposed value function at the current moment.

[0015] In one embodiment, the mixing network adopts an improved Q-mix algorithm based on sub-environment state attention information, including an attention mechanism layer and a mixing layer;

[0016] The attention mechanism layer is used to implement the mapping from a query to a series of key-value pairs. The input of each policy network is used as the key, the sub-environment state information extracted from the full environment state information is used as the query, and the decomposed value function output by the policy network is used as the value. After the query and the key are weighted and mapped, the mapping result is used as the weight to perform weighted summation with the value to obtain a fused value function considering all decomposed value functions. Here, the input of each policy network is the observation state of each agent at the current moment and the sub-goals of the previous multiple time steps.

[0017] The hybrid layer is used to perform hybrid calculation on the full environment state information and the fused value function to obtain a joint value function considering the full environment state information, and determine and output the macroscopic total goal from the joint value function.

[0018] In one embodiment, the full state information refers to the observation situations of all agents, which is represented by the observation matrix V. In the observation matrix V, the element Vij in the observation matrix V indicates that agent i and agent j can observe each other in the environment, and the value is the concatenation result of the observation vectors of agent i and agent j.

[0019] In one embodiment, the sub-environment state information extracted from the full environment state information includes: analyzing the observation matrix to obtain the connection relationship between agents. For agent i, calculate the one-hop neighbor agents that have a direct connection relationship with agent i according to the observation matrix V to obtain a direct relationship subset composed of the one-hop neighbor agents and the sub-goal of agent i. Then, calculate the two-hop neighbor agents that have an indirect connection relationship with agent i according to the one-hop neighbor agents to obtain a potential relationship subset composed of the two-hop neighbor agents. The union of the direct relationship subset and the potential relationship subset is used as the sub-environment state information of agent i.

[0020] In one embodiment, during the learning process of the upper-layer macroscopic policy network, the update goal is to maximize the external reward value actually obtained by each policy network.

[0021] In one embodiment, during the learning process of the lower-layer microscopic action network, when the decision actions output by the decision networks of each DQN reach the sub-goal, the evaluation networks of each DQN generate and output the artificially specified inherent reward value, and the update goal of each DQN is to maximize the cumulative inherent reward value.

[0022] In one embodiment, the system further includes 2 experience replay buffers for the upper-layer macroscopic policy network and the lower-layer microscopic action network to use.

[0023] Compared with the prior art, the beneficial effects of the present invention at least include:

[0024] By combining the upper - layer macro - strategy network and the lower - layer micro - action network, the upper layer is responsible for assigning sub - goals to each agent, and the lower layer performs micro - operations according to the sub - goals assigned by the upper layer to generate decision actions more accurately. In a multi - agent collaborative game environment, each agent can quickly and accurately generate decision actions to attack the target. Brief Description of the Drawings

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 FIG. [FIG number] is a schematic structural diagram of a group confrontation system based on hierarchical reinforcement learning provided by the embodiment;

[0027] Figure 2 FIG. [FIG number] is a schematic structural diagram of the upper - layer macro - strategy network provided by the embodiment;

[0028] Figure 3 FIG. [FIG number] is a schematic diagram of the definition and information aggregation of two - hop neighbors provided by the embodiment;

[0029] Figure 4 FIG. [FIG number] is a schematic structural diagram of the lower - layer micro - action network provided by the embodiment. Detailed Embodiments

[0030] To make the purpose, technical solutions and advantages of the present invention more clear, the following further details the present invention with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not limit the protection scope of the present invention.

[0031] Note: The "FIG. [FIG number]" in the translation is a placeholder. You need to fill in the actual figure numbers according to the specific content of the original text.In a game environment of multi-agent collaborative game confrontation, due to the complex environment and the large number of agents, a huge strategy space is formed, making it difficult and inaccurate for agents to make decision actions. Based on this, aiming at the problems of complex multi-agent environment, large decision space and easy ineffective exploration, the embodiment provides a group confrontation system based on hierarchical reinforcement learning for a general collaborative confrontation hierarchical structure in a battlefield confrontation environment. The strategy of the whole system is defined as two layers and uses different state information. The upper layer is the macro strategy responsible for assigning sub-goals to each agent, and the lower layer makes micro decision operations according to the sub-goals assigned by the upper layer. Different reinforcement learning algorithms are used in the upper and lower layers of the whole system. In order to master the global information and effectively perform value decomposition in the upper layer, an improved Q-mix algorithm with sub-environment state Attention information is used; in the lower layer, in order to efficiently execute single-agent tactics and avoid deadlocks with the actions of other agents on one's own side, a DQN algorithm with an LSTM network, teammate action communication and the macro total goal (output by the upper strategy) is adopted. The hierarchical strategy links the upper and lower layers together for training and use through the given goals and rewards, and its performance in the game environment is better than other current methods that perform well in this type of environment, that is, it can output accurate decision actions.

[0032] In a multi-agent battlefield confrontation environment, if controlled by humans, the thinking angle of humans is divided into several layers. Abstractly speaking, it is roughly divided into two layers, that is, first consider the group macro strategy, and then consider the specific execution process of each agent. If not considered hierarchically, it is very likely to be biased between the macro strategy and tactics and the micro execution tactics. Based on this, the present invention constructs a two-tier system. This is very consistent with the mode of humans to complete a complex task. When encountering a complex task, it will be disassembled into a series of small goals, and then these small goals will be achieved one by one. Each agent has a clear battle goal, that is, who the agent needs to attack currently; after receiving the instruction of the attack goal from the upper layer, the lower layer algorithm also has a clear execution logic, that is, how to walk and attack to attack the attack goal given by the upper layer.

[0033] Figure 1 It is a schematic structural diagram of the group confrontation system based on hierarchical reinforcement learning provided by the embodiment. As Figure 1As shown in the figure, the multi-agent adversarial system based on hierarchical reinforcement learning provided by the embodiment includes an upper-layer macro policy network and a lower-layer micro action network. Among them, the upper-layer macro policy network is a multi-agent reinforcement learning algorithm, which is mainly responsible for generating the macro total goal (GOAL) according to the observation state, and obtaining the reward (Reward). At the same time, the macro total goal of the upper layer is used as the input of the lower layer, so that the two layers are interconnected through the goal. The lower layer is mainly responsible for generating the micro decision-making actions (ACTION) of each agent after receiving the macro total goal, and executing the micro decision-making actions to feedback to the environment (Enviroment). In the embodiment, the entire environment is a multi-agent battlefield collaborative confrontation environment.

[0034] Upper-layer macro policy network

[0035] In the embodiment, in the multi-agent battlefield collaborative confrontation environment, the dimension of the joint behavior value function of multiple agents will increase exponentially with the increase in the number of agents. Therefore, the idea of training multiple agents as a single agent is not feasible. In the value function decomposition, each agent has its own behavior value function, and the centralized behavior value function can be decomposed into a certain combination of each behavior value function. The value function decomposition abandons the global optimality and makes a strong decomposition assumption. On the basis of value decomposition, it is more noteworthy that: in a multi-agent scenario, in many cases, each agent only needs to pay attention to specific entities in the environment. If there are teammates, opponents, and entities irrelevant to the confrontation in the environment, each agent should assign different attention levels to each entity in the environment, and the attention level should change dynamically over time. According to human experience, intuitively, the entities in the environment can be divided and grouped, and the influence of visible and invisible objects on their own Q values can be calculated according to different groups. Based on the above assumptions, each agent can ignore the irrelevant entities in the environment and will learn better and faster. Based on this, the embodiment improves the Q-MIX algorithm, introduces the idea of the multi-head attention mechanism, and on this basis, proposes the concept of sub-environment state information (Sub-State). The sub-environment state information describes the environmental state related to the agent, and the sub-environment state information is used for attention calculation, and the full-environment state information (Full-State) is added for calculation in the mixing network. Intuitively, considering the complexity of the global state in the multi-agent system, in order to reduce the computational complexity and increase the accuracy and robustness, an adaptive state is used for calculation.

[0036] Figure 2 is a schematic structural diagram of the upper-layer macro policy network provided by the embodiment. As Figure 2As shown, the upper-layer macro policy network provided by the embodiment includes multiple policy networks and a hybrid network adopted by multiple agents. Among them, each policy network calculates and outputs the decomposed value function at the current moment according to the observation state at the current moment and the sub-goals of the previous multiple time steps and determines the predicted sub-goal from the decomposed value function Q where, n denotes the observation state of the nth agent at the current moment t, and denotes the sub-goal of the nth agent at the previous c time steps from the current moment t, denotes the joint action, observation history. Specifically, each policy network includes an MLP, a GRU, and an MLP connected in sequence, that is, it calculates the observation state at the current moment input and the sub-goals of the previous multiple time steps and outputs the decomposed value function Q at the current moment with a greedy policy π with parameter ε n .

[0037] In the embodiment, the hybrid network is used to calculate according to the full environmental state information S full , the decomposed value functions output by each sub-agent using the policy network to output the joint value function Q total (S full , GOAL), and determines the maximum value from the joint value function Q total as GOAL as the macro total goal, and this macro total goal is used as the sub-goal of each agent at the next moment.

[0038] In the embodiment, an agent may have no relationship with other agents, and the reinforcement learning algorithm should not consider the interaction between this agent. The advantage of doing this is that it can reduce the complexity of the environment and pay more attention to useful information. Based on such an assumption, two types of environmental state information are adopted: full environmental state information (Full-State) S full and sub-environmental state information (Sub-State) S sub , and the two types of environmental state information are organically combined with the hybrid network.

[0039] There are two types of connections between agents in the battlefield collaborative game confrontation environment: Cooperation Relationship (CR): CR reflects the relationship between teammate agents in this environment, such as jointly attacking the same target, etc.; Antagonistic Relationship (AR): AR reflects the relationship between opponents in the environment, such as attacking the other agent or being attacked by the other agent, etc. Based on the above two relationships and environmental information, an observation vector is proposed, that is, the observation vector represents the situation of enemies or teammates existing within the line of sight. By splicing the observation vectors, the observation matrix V can be obtained. In the observation matrix V, each element Vij indicates that agent i and agent j can observe each other in the environment, and the element value is the splicing result of the observation vectors of agent i and agent j.

[0040] In the embodiment, by analyzing the observation matrix, the one-hop neighbor agents of agent i can be obtained. Make an assumption about the target of the attack, believing that the target also has the same association with a neighbor of its own agent. According to this relationship, the direct relationship subset Mi of agent i can be obtained. The direct relationship subset Mi is defined as a sub-goal goal that contains agent i i And the subset of one-hop neighbor agents with direct connection relationships.

[0041] One-hop neighbors represent the direct relationship between agents, but they still cannot represent the relationship between agent groups. Aggregating two-hop neighbors is also of great significance. On this basis, the calculation of two-hop neighbor agents is introduced. Two-hop neighbor agents are agents that have an indirect connection relationship with agent i calculated based on one-hop neighbor agents. The definition and information aggregation of two-hop neighbor agents are as Figure 3 shown, where k is the hop index. When k is 1, it represents one-hop, and when k is 2, it represents two-hop. The set of agents calculated through two-hop neighbors is called the potential relationship subset. Finally, the intersection of the direct neighbor subset and the potential relationship subset is used as the sub-environment state information S of agent i sub .

[0042] In the embodiment, an attention mechanism is further introduced into the hybrid network, that is, the hybrid network includes an attention mechanism layer and a hybrid layer. Among them, the attention mechanism layer is used to implement the mapping from a query to a series of key-value pairs. Imagine that the constituent elements in the Source are composed of a series of <Key, Value> data pairs. At this time, given a certain element Query in the Target, by calculating the similarity or correlation between the Query and each Key, the weight coefficient corresponding to each Value of the Key is obtained. After softmax normalization, the weighted sum of the weights and the corresponding Values is performed, that is, the final attention value is obtained. Therefore, in essence, the attention mechanism is to perform a weighted sum on the Value values of the elements in the Source, and the Query and Key are used to calculate the weight coefficients corresponding to the Values.

[0043] In the embodiment, the input of each policy network (the observation state of each agent at the current moment and the sub-goals of the previous multiple time steps) is used as the key Key, the sub-environment state information extracted from the full environment state information is used as the query Query, and the decomposed value function output by the policy network is used as the value Value. After the query Query and the key Key are weighted and mapped, the mapping result is used as the weight and weighted sum with the value Value to obtain a fused value function that considers all the decomposed value functions containing the predicted sub-goals. Since the full environment state information contains a lot of irrelevant information and noise information, the structure using the sub-environment state information instead of the full environment state information further precisifies the relationship between the single-agent observation and the overall information.

[0044] It should be noted that for each agent, after the query Query and the key Key are weighted and calculated by the scaled dot product, and then linearly transformed through SOFTMAX, the mapping result λ is obtained. n,h The mapping results of all agents and the corresponding values Value (decomposed value functions) of all agents are weighted and calculated by the scaled dot product to obtain the fused value function Q. h .

[0045] It should be noted that the embodiment adopts the multi-head attention mechanism. Each head performs an attention calculation for all agents once, and h heads perform h times of attention calculations. Each time, the fused value function Q is obtained. h Then, the concatenation of the fused value functions Q obtained h times h is input into the hybrid layer, and at the same time, the full environment state information is input into the hybrid layer for hybrid calculation to obtain a joint value function that considers the full environment state information and all the decomposed value functions, and outputs the macroscopic total goal included in the joint value function as the sub-goal of each agent at the next moment. The formula is expressed as Q total (S full , GOAL):

[0046]

[0047] In the embodiment, the mixing layer adopts a fully connected layer, and the output dimension is the action dimension, which is equal to the agent dimension, that is, an output action is selected therefrom as the macroscopic total goal.

[0048] In the embodiment, during the learning process of the upper-layer macroscopic policy network, the update goal is to maximize the external reward value actually obtained by each policy network, that is, the update goal is expressed as:

[0049]

[0050] where Q tot is the joint Q-value, u is the joint action, and τ is the joint action-observation history.

[0051] The lower-layer microscopic action network

[0052] Based on the upper-layer macroscopic policy network, how to perform microscopic decision-making operations to achieve the macroscopic policy (macroscopic total goal) is also a problem that each agent needs to consider. To address this problem, a bottom-layer reinforcement learning execution algorithm based on temporal information and the upper-layer policy is proposed. Since the cooperation and competition issues have been inferred and predicted at the macroscopic tactical level, the lower-layer microscopic action network uses a single-agent algorithm to reduce the complexity of the problem and accelerate the training efficiency to improve the success rate of the execution strategy without the need to consider the cooperation and competition issues between agents again.

[0053] In the embodiment, the lower-layer microscopic action network envelopes multiple DQNs adopted by multiple agents, adds the macroscopic total goal output by the upper-layer macroscopic policy network for action value function prediction, and achieves the purpose of selecting microscopic operations through value function prediction.

[0054] Specifically, on the basis of DQN, the macroscopic total goal (GOAL) calculated by the upper-layer macroscopic policy network is added as a part of the input state for use, and the decision action is selected according to the observation state of each agent itself. The specific structure is as Figure 4 shown. The observation state o at the current moment is obtained from the environment, and the sub-goal goal at the current moment is obtained from the buffer. Based on the sub-goal goal and the observation state o at the current moment, the decision action that maximizes the expected intrinsic reward is selected. Here, it is the same as the traditional DQN except that the sub-goal goal is added. It should be noted that inside the agent, this intrinsic reward is generated by the Critic network, and the intrinsic reward will be generated only when the decision action output by the decision network at the current moment reaches the sub-goal. During the learning process, the update goal of each DQN is to maximize the cumulative intrinsic reward value, which is expressed as:

[0055]

[0056] Among them, γ t′-t represents the intrinsic reward discount factor from time t' to time t, o t represents the observation state at time t, a t represents the action vector of the agent at time t, g t represents the reward obtained by the agent at time t, π ag is the policy, r t represents the intrinsic reward value.

[0057] In the embodiment, the decision network in the DQN adopted by each agent in the lower - layer microscopic action network uses a neural network with 3 fully - connected layers. The sizes of the three layers are: 128, 256, 128 respectively, and the output dimension is the action dimension of each agent itself. ReLU is used as the activation function, and Adam is used as the optimizer for optimization.

[0058] In the embodiment, since the system is composed of two - layer reinforcement learning algorithms and a complex policy hierarchical structure, the learning algorithm is also designed in detail. First, at the data level, based on the original game data, two experience replay buffers (Replay Buffer) are constructed. One is for the upper layer and the other is for the lower layer. Since the upper - layer instructions are not updated in every time slice. Considering this factor, the data structure of the buffer for the upper layer is constructed as:

[0059] <S t ,O t ,A t ,G t ,S -t ,O -t >

[0060] Among them, S t represents the environmental state at time t, including the full - environmental state information and the sub - environmental state information, O t represents the observation state of each agent at time t, A t represents the action vectors of all agents at time t, that is, the predicted sub - goals of all agents, that is, the macro - total goal, G t represents the reward obtained by the agent at time t, S -t represents the impact of generating the macro - total goal on the environment, O -t represents the impact of the predicted sub - goals output by all agents on the observation state.

[0061] Since the lower-level buffer needs to consider the sub-goal goal given by the upper level, and the lower level is actually a single-agent algorithm, considering this factor, the buffer data for the lower level is constructed as follows:

[0062] <o t ,a t ,g t ,r t ,o -t >

[0063] Among them, o t represents the observation state of each agent at time t, a t represents the decision-making action of each agent at time t, g t represents the sub-goal of each agent at time t, r t represents the intrinsic reward obtained by each agent at time t, o -t represents the impact of each agent's execution of the decision-making action on the observation state.

[0064] In the embodiment, in the selection of exploration and exploitation actions, both the upper and lower layers use the Epsilon-Greedy strategy. At the training structure level, the upper and lower layer networks are updated using a synchronous update method. First, for the upper and lower layer networks, the parameters are initialized. When the experience pool meets the requirements, a batch of experiences from the experience pool is extracted for algorithm update. The learning phase and the acquisition of experiences are logically separated, and based on randomly sampling from the experience pool, by repeatedly presenting these data to the upper and lower layer reinforcement learning networks, rapid learning is achieved from a limited amount of data. Specifically, when optimizing, the loss function used is:

[0065] L(θ) = E[(TargetQ - Q(o,a;θ)) 2

[0066] Among them, TargetQ is the target Q value, and Q(o,a;θ) is the Q value output by DQN.

[0067] The above-mentioned group confrontation system based on hierarchical reinforcement learning is carried out in the intelligent agent collaborative battle environment. Two types of verification indicators are considered: (1) Reward: The method is to calculate the average reward score. When testing the performance, a certain number of steps are executed, and all the obtained reward values are recorded and averaged. In this experiment, two evaluation methods are considered for the reward. One is to conduct 1000 game tests after convergence, and the average reward of each game is tested. The second is during the training process, and a curve graph is drawn based on the average reward of the past 100 episodes for quantitative evaluation of the training process.

[0068] ​(2) Win Rate: In the experiment, two evaluation methods are considered for the win rate. One is to conduct 1000-game tests after the algorithm converges and calculate the average win rate of these 1000 games. The second is during the algorithm training process, to draw a curve graph based on the average win rate of the past 100 games for quantitative evaluation of the training process. All experimental results were conducted 8 times and the average value was taken for evaluation and analysis.

[0069] The specific implementation manners described above have detailed the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment, characterized in that, Applied to a game environment of collaborative game confrontation composed of opponent agents and teammate agents, each agent in the upper layer has a clear combat goal, that is, the object to be attacked by the agent. After receiving the attack target instruction from the upper layer, the agents in the lower layer have a clear execution logic, that is, how to move to attack the attack target given by the upper layer; the multi-agents in the upper layer adopt an upper-layer macro strategy network in terms of tactics, and the multi-agents in the lower layer adopt a lower-layer micro action network in terms of tactics; The upper-layer macro strategy network includes multiple strategy networks and a mixing network adopted by multiple agents. Each strategy network is used to calculate and output the decomposed value function at the current moment based on the observation state at the current moment and the sub-goals in the previous multiple time steps, and determine the predicted sub-goal from the decomposed value function; the mixing network is used to calculate and output the joint value function based on the full environment state information and the decomposed value functions output by the strategy networks adopted by each sub-agent, and determine the macro total goal as the sub-goal of each agent at the next moment; among them, the full state information refers to the observation situation of all agents, including the situation of enemies or teammates in the visual memory; The lower-layer micro action network contains multiple DQNs adopted by multiple agents, and each DQN is used to calculate and output a decision action based on the observation state at the current moment and the sub-goal at the current moment.

2. The group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment according to claim 1, characterized in that, The strategy network includes an MLP, a GRU, and an MLP connected in sequence, and is used to calculate the observation state at the current moment and the sub-goals in the previous multiple time steps of the input, so as to output the decomposed value function at the current moment.

3. The group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment according to claim 1, characterized in that, The mixing network adopts an improved Q-mix algorithm based on the sub-environment state attention information, including an attention mechanism layer and a mixing layer; The attention mechanism layer is used to realize the mapping from a query to a series of key-value pairs. The input of each strategy network is used as the key Key, the sub-environment state information extracted from the full environment state information is used as the query Query, and the decomposed value function output by the strategy network is used as the value Value. After the query Query and the key Key are weighted and mapped, the mapping result is used as the weight to perform weighted summation with the value Value to obtain a fused value function considering all decomposed value functions. Among them, the input of each strategy network is the observation state of each agent at the current moment and the sub-goals in the previous multiple time steps; The mixing layer is used to mix and calculate the full environment state information and the fused value function to obtain a joint value function considering the full environment state information, and determine and output the macro total goal from the joint value function.

4. The group confrontation system based on hierarchical reinforcement learning for the battlefield confrontation environment according to claim 3, characterized in that, The full state information is represented by the observation matrix V. In the observation matrix V, the element Vij in the observation matrix V indicates that agent i and agent j can observe each other in the environment, and the value is the concatenation result of the observation vectors of agent i and agent j.

5. The group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment according to claim 4, characterized in that, The sub-environment state information extracted from the full-environment state information includes: analyzing the observation matrix to obtain the connection relationship between agents. For agent i, the one-hop neighbor agents directly connected to agent i are calculated according to the observation matrix V, and a direct relationship subset composed of the one-hop neighbor agents and the sub-goals of agent i is obtained. Then, the two-hop neighbor agents indirectly connected to agent i are calculated based on the one-hop neighbor agents, and a potential relationship subset composed of the two-hop neighbor agents is obtained. The union of the direct relationship subset and the potential relationship subset is used as the sub-environment state information of agent i.

6. The group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment according to claim 1, characterized in that During the learning process of the upper-layer macro policy network, the update goal is to maximize the external reward value actually obtained by each policy network.

7. The group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment according to claim 1, characterized in that During the learning process of the lower-layer micro action network, when the decision actions output by the decision-making network of each DQN reach the sub-goals, the evaluation network of each DQN generates and outputs the inherent reward value stipulated artificially, and the update goal of each DQN is to maximize the cumulative inherent reward value.

8. The group confrontation system based on hierarchical reinforcement learning for battlefield confrontation environment according to claim 3, characterized in that The system also includes 2 experience replay buffers for the upper-layer macro policy network and the lower-layer micro action network to use; Among them, the data structure of the experience replay buffer used by the upper-layer macro policy network is: <S t ,O t ,A t ,G t ,S -t ,O -t > Among them, S t represents the environmental state at time t, including the full environmental state information and the sub-environmental state information, O t represents the observation state of each agent at time t, A t represents the action vector of all agents at time t, that is, the predicted sub-goals of all agents, that is, the macro total goal, G t represents the reward obtained by the agent at time t, S -t represents the impact of the generation of the macro total goal on the environment, O -t represents the impact of the output of the predicted sub-goals of all agents on the observation state; The data structure of the experience replay buffer used by the lower-layer micro action network is: <o t ,a t ,g t ,r t ,o -t > Among them, o t represents the observation state of each agent at time t, a t represents the decision-making action of each agent at time t, g t represents the sub-goal of each agent at time t, r t represents the intrinsic reward obtained by each agent at time t, o -t represents the impact of each agent's execution of the decision-making action on the observation state.

Citation Information

Patent Citations

  • Setting method and device of pass game, storage medium and electronic device

    CN107970608A

  • Robust, scalable and generalizable machine learning paradigm for multi-agent applications

    CN113396428A