Value decomposition difference-based multi-subject comparison exploration method

By utilizing the multi-subject comparison exploration method of value decomposition differences in multi-agent collaboration, and using self-attention mechanism and confined hybrid network to generate joint value functions, the problems of inefficient sample efficiency and extended training convergence time in the existing technology are solved, and faster cooperative strategy learning and better joint action recognition are achieved.

CN119990245APending Publication Date: 2025-05-13UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510144491.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art In multi-agent collaboration, value decomposition methods rely on neural networks to lead to inefficiency of sample and prolong training convergence time.

Method used

A multi-subject comparison exploration method based on the difference in value decomposition is proposed. Using the estimation differences between different types of value decomposition methods, a joint value function is generated through self-attention mechanisms and restricted hybrid networks, promoting the exploration and guidance of the update of joint strategies.

Benefits of technology

By using value decomposition differences, we promote the exploration and guidance of the update of joint strategies, we improve the optimal joint action recognition ability of agents in the broad action space, and accelerate the learning process of cooperative strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990245A_ABST
    Figure CN119990245A_ABST
Patent Text Reader

Abstract

The invention discloses a value decomposition difference-based multi-agent comparison exploration method, which comprises the following steps of: determining an updating weight according to the difference between different value decomposition estimations by utilizing the value decomposition difference and a comparison principle, setting the updating weight and taking the difference as an internal target in an updating process. The MACE architecture comprises two value function estimators, each value function estimator is responsible for estimating joint state action value functions Qjt and Qtot corresponding to the two VD methods, and an implicit reward function and a weighting mechanism are created through the difference between the Qjt and the Qtot to guide exploration and used for updating the two internal function estimators. By means of the method, it is ensured that the action with the high Q value is preferentially sampled, the action with the small Q value still has the opportunity to be sampled, the exploration behavior is enhanced, the learning speed and the final performance are obviously superior to those of a baseline, and the complete expression capacity is effectively kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of coordinated robotics, autonomous unmanned aerial vehicles (UAVs), and traffic signal control. The correct estimation of the joint value function with full representation ability can guide the appropriate update of the joint value function with limited representation ability. This difference implicitly points out the direction that the joint cooperation strategy should explore. Appropriate exploration enables the model to quickly identify the optimal joint action in a wide action space, correct the misvaluation of the policy network, and accelerate the learning process of the entire cooperation strategy. Background Art

[0002] (1) Dec-POMDP: Consider the collaboration among multiple agents, since the system consists of n agents, represented as n∈N≡{1,…,n}, which operate in a restricted observation area and make decisions using their respective observation features. This decision-making process is usually embodied in the concept of decentralized partially observable Markov decision processes (DecentralizedPartiallyObservableMarkovDecisionProcesses, Dec-POMDPs), represented and modeled by the tuple G = (S,U,P,r,Z,O,n,γ), where at each time step t, the environment transitions according to the environment state probability function The actual state s∈S conveys the collective knowledge gained by all agents through sharing information and collaborative decision-making, as well as the supplementary environmental knowledge that the environment itself may contain in addition to the information observed by the agents. n Represents collective individual action u i ∈U, where U n represents the Cartesian product of the collective individual actions of n agents, that is, the combination of actions that each agent can choose. All agents share a common reward function r(s,u):S×U→R. Each agent is rewarded by observing the function O(z i |s,u i ):S×U→Z gets a partial observation z i ∈Z, the goal of all agents is to collaborate to maximize the cumulative team reward This is the joint value function defined.

[0003] (2) Task and value decomposition: In order to achieve better coordination, it is usually necessary to clearly define the specific tasks for each agent to cooperate, which allows the parameters of the policy network to be updated based on the assigned task information. This process is called task decomposition and ensures that each agent can effectively coordinate its actions and optimize the performance of the entire system. In the reinforcement learning (RL) paradigm, the completion of a task usually depends on the joint actions of the agents, which are represented by the joint action value function Qtot Therefore, task decomposition essentially corresponds to decomposing the joint action value function. The essence of VD is to form the joint state action value function Q during training tot and the utility function of individual agents Direct connection between and meets:

[0004]

[0005] Among them, τ represents the motion trajectory, which means that we want to find an action u such that under the given trajectory τ, the joint action value function Q tot Reached maximum value.

[0006] when Following the joint value function Q tot When the IGM criterion is satisfied, the related cooperative tasks are said to be decomposable. Therefore, Q tot (τ,u) can be decomposed into or Constitutes Q tot Two commonly used value function decomposition methods are VDN and QMIX, which satisfy the following sufficient conditions:

[0007]

[0008] This type of approach initially assumes a decomposable centralized value function, which is then decomposed into a compositional structure containing the utility functions of each agent. tot Expressed as the monomer utility function Q through the hybrid network i A monotone nonlinear combination of , where the parameter weights are all positive, can be described as The VDN can be expressed as a special case of QMIX. where w i ≡1 and satisfy the IGM principle. Usually, these positive parameters can be established by imposing structural constraints on the neural network, where the input utility function Q i and output Q tot This optimization effectively reduces arrive The time complexity of |U| usually represents the number of possible actions taken by a single agent, while n represents the total number of agents, which helps the agents learn effective joint actions faster.

[0009] The restricted positive correlation between input and output in these VD methods may lead to the totThe inability to fully capture the reward distribution associated with the task hinders the ability of the agents to discover the optimal joint action. Another type of VD method re-estimates the true joint action value function, uses QMIX for action selection, and reduces the update weights of non-optimal action samples during training, which encourages cooperative agents to pay more attention to samples with optimal joint actions in their updates. The WQMIX type of VD method exemplifies this, establishing an updated weight w(s,u)<1 for samples associated with the optimal action, and assigning weights to samples associated with non-optimal actions, expressed as follows:

[0010]

[0011] in is the estimated value of the target network. The joint value function Q jt The WQMIX fitting method shows better expressiveness than QMIX and VDN without being restricted by the neural network structure. However, since these methods rely on neural networks to estimate the centralized value function Q for selecting joint actions, jt and Q tot , thus encountering the problems of low sample efficiency and prolonged training convergence time.

[0012] (3) Policy Gradient Method: The core concept of the policy gradient method is to create a policy π controlled by the neural network parameters φ, and the update of these parameters is driven by the feedback received from the environment. This method is based on the policy gradient theorem, and the gradient update equation formula is:

[0013]

[0014] in, is the derivative of the loss function with respect to the parameter φ, the environmental state distribution P, is the gradient of the logarithmic probability of the policy π with respect to the parameter φ, that is, in state s t Take action a t The gradient of the probability, Q π (s t ,a t ) represents the Q function estimated by strategy π for fitting the cumulative reward of interaction with the environment. The accuracy of Q estimation determines the direction and accuracy of strategy update. The commonly used method is to evaluate the relative importance of actions by subtracting the state value, that is: A π (s t ,a t )=Q π (s t ,a t )-V π (s t ), where A πIt is called the advantage function. In order to enhance exploration and avoid getting stuck in local optima, a common practice is to include the entropy of the policy in the objective function, thereby modifying the gradient update form:

[0015]

[0016] The Shannon entropy representing the policy π and α is a temperature parameter that controls the randomness of the policy π and can be set to a constant or optimized with another objective function.

[0017] (4) Qualitatively describe the advantages and disadvantages of two different types of value decomposition methods

[0018] Unrestricted centralized value decomposition of value functions: As a core component of deep learning, neural networks are often used as feature extractors to express high-level features and abstract information in observed data. The proficiency of neural networks in completing the above tasks is attributed to their use of multi-layer network architecture and multiple activation functions for nonlinear transformation of input features. This design enables the network to effectively handle nonlinear relationships in the data, giving it the ability to approximate arbitrary functions. Typically, under the CTDE framework, MARL algorithms integrate the utility functions of all agents to consider joint decision-making in cooperative scenarios. Figure 1 As shown in the left part, the utility function Q of the agent i Fitted through the DRQN ​​module and entered into the hybrid network to generate the centralized value function Q jt The hybrid network for this type of VD method usually takes the form of a regular multilayer feedforward network and does not impose any constraints on the weight parameters. Each utility function selects an action from the action space of a single agent, so Q jt The joint action space is (|U| n ). The basic idea of ​​this type of algorithm is to assume that the centralized value function can be decomposed so as to maintain full representation capability throughout the process. However, the lack of constraints on the input actions when estimating the centralized value function and the utility function makes it impossible to clearly determine the relationship between joint actions and individual actions, which leads to slow optimization of the collaborative decision-making process of intelligent agents.

[0019] Value Decomposition of Concentrated Value Functions with Explicit Restrictions: The expressive power of a neural network is closely related to its weight parameters. In the CTDE framework, VD methods with restricted expressive power face challenges in fitting arbitrary functions. The fundamental problem stems from the absolute value normalization of the weight parameters generated by the hypernetwork, which aims to ensure the consistency of joint actions and individual actions. Forcing all weights to be positive causes the network to lose the ability to express negative correlations, making it unable to adapt to certain complex functional relationships. For example, Figure 1 As shown in the right part, the hybrid network integrates the utility function Q jtUse the weight parameter |w after absolute value normalization i |>0. This enforces a uniform positive correlation (or uniform monotonic relationship) between the concentration value function and the utility function, thus guaranteeing the IGM condition satisfaction.

[0020] When all weights are constrained to be positive, the neural network loses the ability to fit certain functions because it cannot represent the negative correlation between the input utility functions. This defect is reflected in the limiting ability of the concentrated value function. Summary of the invention

[0021] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a new multi-agent comparative exploration method based on value decomposition differences, which utilizes the estimation differences between different types of VD methods to promote exploration and guide the updating of joint strategies.

[0022] The object of the present invention is achieved through the following technical solution: a multi-agent comparative exploration method based on value decomposition difference, which utilizes the estimation differences between different types of VD methods to promote exploration and guide the update of joint strategies, comprising the following steps:

[0023] It utilizes the estimation differences between different types of VD methods to promote exploration and guide the update of joint strategies, and is characterized by including the following steps:

[0024] (1) The agent strategy network uses a GRU-based recurrent neural network to make decisions: the agent’s observation value o and the previous action u are combined t-1 As the input of the agent strategy network; after the input data is encoded by the feedforward neural network composed of MLP, GRU and MLP in sequence, the output state action value function Q i (τ i ,); then the action selection mechanism output by the action preference network is used to select the collaborative action u corresponding to the highest value of the state action value function i , and calculate the state action value function Q under the collaborative action i (τ i ,u i ), τ i is the trajectory of the i-th agent; finally, the state action value function Q i (τ i ,u i ) is input to a first joint value function estimation part having full expression capability and a second joint value function estimation part having limited expression capability;

[0025] (2) The first joint value function estimation part with full expressive power uses the self-attention mechanism to generate a centralized value function estimate Q jt(τ,u); the specific process is:

[0026] (2-1) Statistics of the state action value function Q of all agents i (τ i ,u i ), denoted as n is the total number of agents;

[0027] (2-2) The global state information S obtained by the agent through collecting surrounding environment information t The input is processed in a module consisting of MLP, attention mechanism module, soft attention module and Score function to obtain the weight ratio coefficient of the Attention mechanism;

[0028] (2-3) The weight ratio coefficient of the Attention mechanism output by the Score function is Perform matrix multiplication to obtain the enhanced utility function

[0029] (2-4) The enhanced utility function and global state information S t Input them into MLP for encoding processing respectively, and sum the data processed by the two MLPs to get the centralized value function estimate Q jt (τ,u);

[0030] (3) The second joint value function estimation part with restricted expressiveness combines the utility functions of the agents through a restricted hybrid network to produce a centralized value function that can decompose the cooperative task. The specific process is:

[0031] (3-1) Statistics of the state-action value function Q of all agents i (τ i ,u i ), denoted as

[0032] (3-2) Decomposed into the state function of each agent and advantage function It is expressed as:

[0033]

[0034] (3-3) At time t Each Q i (τ i ,u i ) to perform argmaxQ i (τ i ,u i) operation to obtain the joint optimal action consisting of the actions with the maximum action Q value of each agent

[0035] (3-4) The joint optimal action and global state information S t Input the MLP and attention mechanism modules in turn, use the attention mechanism to merge the global state information; and combine the data obtained by the attention mechanism module with the advantage function Perform a dot product;

[0036] (3-5) All the state functions of the agents Add them together, and then add them to the dot product result obtained in (3-4) to get the final concentrated value function

[0037] (4) According to Q jt (τ,u) and The difference between them is used to adaptively calculate the intrinsic target; then the optimization objective function is calculated, which is expressed as:

[0038]

[0039] in and Q jt and Q tot The target estimate of is calculated as maxQ tot (τ′,u′,s′); ω is the scaling factor; intrinT represents the intrinsic goal, which is Q jt (τ,u) and The difference between them; r is the immediate reward, γ is the discount factor; the update weight is set to α<1; maxQ tot (τ′,u′,s′) represents the maximum joint action value function Q of all possible actions u′ in the next state s′ tot The value of

[0040] By optimizing the above intrinsic objective function, the action preference network parameters, the first joint value function estimation part and the network parameters of the second joint value function estimation part with restricted expression ability are obtained; then the network is deployed to each intelligent agent.

[0041] The action preference network has two output components: Q function estimation and action-value distribution; Q function estimation is used to calculate the state action value function; action-value distribution is used to select a greedy action with probability ∈ according to the value of the state action value function, or to sample actions according to action preference with probability 1-∈, and execute the action to trigger the next state transition.

[0042] The beneficial effects of the present invention are: (1) The present invention proposes a new multi-agent comparative exploration method based on value decomposition difference, which utilizes the estimation differences between different types of VD methods to promote exploration and guide the updating of joint strategies; (2) MACE includes an action preference network based on the average policy entropy of all agents, which enables more effective action sampling, thereby enhancing exploration, helping agents to escape from local optimality and more easily find the optimal joint action. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Schematic diagram of two different types of value decomposition methods; (a) is a centralized state-action value function with complete representation capability, (b) is a centralized state-action value function with only positive representation;

[0044] Figure 2 Schematic diagram of the MACE model, A is the agent strategy network containing action preferences, B is the centralized value function part with infinite representation ability; C is the centralized value function part with limited representation ability;

[0045] Figure 3 Schematic diagram of the interaction between the MACE agent and the environment;

[0046] Figure 4 The training results of MACE and other baseline algorithms in different SMAC scenarios;

[0047] Figure 5 for the effect of the additional optimistic component on different SMAC scenarios during MACE training;

[0048] Figure 6 The role of action preference network in MACE and its training process;

[0049] Figure 7 The influence of different action preference network update frequencies on the MACE training process. DETAILED DESCRIPTION

[0050] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0051] The present invention can be applied in the fields of coordinated robots, autonomous unmanned aerial vehicles (UAVs) and traffic signal control. Intelligent agents usually refer to systems or devices that can autonomously perform tasks and make decisions. Specifically in these fields, intelligent agents can be the following devices: 1. Coordination robots: Intelligent agents are industrial robots with sensors and actuators. They can collect data about the surrounding environment through sensors, such as visual, tactile, position and other information, to form the global state information S of the intelligent agent. t2. Autonomous Unmanned Aerial Vehicles (UAVs): The intelligent agent is a drone equipped with sensors such as cameras, GPS, and inertial measurement units (IMUs). They can collect image data, position information, speed, acceleration, and other data during flight to form the global state information S of the intelligent agent. t 3. Traffic signal control: The agent may be a traffic signal control system equipped with cameras, traffic flow sensors, radars and other equipment, which can monitor traffic flow, vehicle speed, road conditions and other information in real time to form the global state information S of the agent. t .

[0052] In the context of multi-agent reinforcement learning (MARL), the data collected by agents through these sensors and devices will be used to train and update their decision-making strategies to achieve effective collaboration and task execution. For example, in traffic signal control, agents need to collect real-time traffic flow data to optimize the switching of traffic lights to reduce traffic congestion and improve road use efficiency.

[0053] A Multi-Agent Contrastive Exploration (MACE) method based on value decomposition differences, which uses the estimation differences between different types of VD (Value Decomposition) methods to promote exploration and guide the update of joint strategies. MACE consists of three parts: the agent policy network, the first joint value function Q with full expressive power, and the agent policy network. jt Estimation part, the second joint value function Q with limited expressive power tot The estimation part, the specific structure is as follows Figure 2 As shown; specifically including the following steps:

[0054] (1) The agent strategy network uses a GRU-based recurrent neural network to make decisions: the agent’s observation value o and the previous action u are combined t-1 As the input of the agent strategy network; after the input data is encoded by the feedforward neural network composed of MLP, GRU and MLP in sequence, the output state action value function Q i (τ i ,); then the action selection mechanism output by the action preference network is used to select the collaborative action u corresponding to the highest value of the state action value function i , and calculate the state action value function Q under the collaborative action i (τ i ,u i ), τ i is the trajectory of the ith agent, u iis the action of the i-th agent; finally, the state action value function Q i (τ i ,u i ) is input into a first joint value function estimation part with full expressive power and a second joint value function estimation part with limited expressive power.

[0055] (2) The first joint value function estimation part with full expressive power ( Figure 2 Module A in ) uses the self-attention mechanism to incorporate global state information and uses soft attention to improve the learning ability of the effective decision-making process to supplement the utility function and generate Once we get this enhanced utility function They are fed into the feedforward neural network for information encoding, and a state function V(s) encoded with global state information is added to finally generate a centralized value function estimate Q jt (τ,u); the specific process is:

[0056] (2-1) Statistics of the state action value function Q of all agents i (τ i ,u i ), denoted as n is the total number of agents;

[0057] (2-2) The global state information S obtained by the agent through collecting surrounding environment information t The input is processed in a module consisting of MLP, Attention module, Softmax module and Score function.

[0058] Get the weight ratio coefficient of the Attention mechanism;

[0059] (2-3) The weight ratio coefficient of the Attention mechanism output by the Score function is Perform matrix multiplication (Matmul) to obtain the enhanced utility function

[0060] (2-4) The enhanced utility function and global state information S t Input them into MLP for encoding processing respectively, and sum the data processed by the two MLPs to get the centralized value function estimate Q jt (τ,u);

[0061] The self-attention mechanism includes MLP, attention module, soft attention (Softmax) module and Score function, matrix multiplication and MLP module. The output of the entire self-attention mechanism is represented by Attention().

[0062] Q jt The calculation process of (τ,u) can be described by the following formula:

[0063]

[0064] Where W Q and W K represents the parameters of the Q network and K network in the self-attention mechanism module, d k Refers to the dimension of the encoding vector; V(s) represents the global state information S t The state function obtained after inputting into the MLP for processing.

[0065] (3) The second joint value function estimation part with restricted expressiveness combines the utility functions of the agents through a restricted hybrid network to produce a centralized value function that can decompose the cooperative task. The specific process is:

[0066] (3-1) Statistics of the state-action value function Q of all agents i (τ i ,u i ), denoted as

[0067] (3-2) Decomposed into the state function of each agent and advantage function It is expressed as:

[0068]

[0069] (3-3) At time t Each Q i (τ i ,u i ) to perform argmaxQ i (τ i ,u i ) operation to obtain the joint optimal action consisting of the actions with the maximum action Q value of each agent

[0070] (3-4) The joint optimal action and global state information S tInput the MLP and attention modules in turn, use the attention mechanism to merge the global state information; and compare the data obtained by the attention module with the advantage function Perform a dot product.

[0071] (3-5) All the state functions of the agents Add them together, and then add them to the dot product result obtained in (3-4) to get the final concentrated value function

[0072] This module focuses on the value function The expression can be normalized as:

[0073]

[0074] W Q ', W K ' represents the parameters of the Q network and K network in the attention mechanism, d' k It refers to the dimension of the encoding vector, ABS represents the absolute value function; Attention() refers to the output of the entire self-attention mechanism (including the network composed of MLP, attention mechanism (Attention) module and dot product module).

[0075] The action preference network of the present invention has two components: Q function estimation (Q Estimation) and action value-distribution (Action Distribution), such as Figure 3 As shown in Figure 2, these two components share a feed-forward encoding network consisting of an MLP, a GRU, and an MLP in sequence. During cooperative training, once the agent infers its action preferences, the network encourages action sampling based on these preferences, rather than following a greedy approach, promoting prior cooperative exploration.

[0076] Q function estimation (Q Estimation) is used to calculate the state action value function Q i (τ i ,), the action-value distribution is used to select a greedy action with probability ∈ according to the state action value function value, or to sample actions according to action preferences with probability 1-∈, and execute the action to trigger the next state transition; the customized ∈-greedy strategy is expressed as follows:

[0077]

[0078] Among them, π custom-∈ (u i |o i ) represents the policy function of the action preference network output action; η(u i|o i ) represents stochastic action u i The preference function of the corresponding action preference network is used to determine whether the output action is the local optimal action, and to influence the final action selection probability of the action preference network; Represents the best action chosen by the agent under a certain strategy or algorithm.

[0079] MACE uses the policy entropy of a single agent to determine the update direction of its action preference network and describes the objective function of the action preference network as:

[0080]

[0081] Where φ represents the parameter value of the action preference network; Indicates state o i Satisfy the state transfer matrix u i By the action preference network η(u i |o i ) for sampling; represents the output action u of the action preference network i In state o i The corresponding function value at V η (o i ) represents the action preference network output action in o i The state value at φ (u i |o i ) represents the action preference network preference function when the parameter is φ; Represents the action preference network output action u i In state o i Advantage function at ; Represents a preference network η φ (u i |o i ) is observing o i MACE calculates the average policy entropy of all individual agents and uses it as the final objective function; this allows the agent to adjust the sampling of priority actions according to the interaction trajectory in each round of interaction and optimize the action preference by maximizing this objective.

[0082] According to the objective function of the action preference network, the gradient of the action preference network and the new parameters calculated are written as:

[0083]

[0084] The temperature parameter β is automatically adjusted using the minimax optimization method:

[0085]

[0086] Among them, ξ entropy is the expected minimum entropy.

[0087] (4) Similar to traditional collaborative MARL, MACE also uses the team rewards obtained from the interaction between the module and the environment to update the parameters of the centralized value function network. However, MACE is different in that it introduces a new action preference network and uses its output to adjust the ∈-greedy mechanism to enhance the action preference of the agent. In addition, MACE jt (τ,u) and The difference between them is used to adaptively compute the intrinsic goal intrinT. The centralized value function objective of the cooperative action is implicitly modified to promote a wider range of action exploration.

[0088] First, we initialize the agent policy network, action preference network, and two types of VD parts (including the parameters of their self-attention modules) as well as the parameters of the target network. Next, we set up the experience replay buffer The learning rate η, batch size, and target network update interval are parameters. We use individual observations o i And the last action To evaluate the utility function Then the Q of the two parts of VD jt and Q tot The calculation is performed and the difference between these two centered value function estimates is calculated to obtain the intrinsic target.

[0089] During the training process, MACE Extract a small batch of interaction trajectories and estimate the utility function And these two concentrated value functions {Q jt , Q tot}. In this process, the action preference network is updated according to each interaction trajectory between the model and the environment. At this time, MACE calculates the average policy entropy of all agents and updates the parameters of the action preference network using formulas (5) and (6), adjusts the agent's ∈-greedy mechanism and obtains the action preference strategy based on this average entropy to ensure the maximum entropy of the collaborative strategy of the entire team.

[0090] By adding Q jt and Q tot The difference between them is scaled by a factor ω, which can be used to obtain an intrinsic goal to encourage the model to jump out of the local optimum and find the optimal joint action more easily, which is summarized as ω·(Q jt -Q tot). The intrinsic goal in MACE uses the contrastive learning principle to transform Q jt and Q tot The difference between Q and MACE is interpreted as a trainable metric to minimize. Ideally, this difference becomes zero when each MACE achieves the joint optimal action; for non-optimal actions, Q jt The full expressiveness makes it more tot is more suitable for capturing the value function of the optimal joint action. Reducing this gap implicitly encourages Q tot Jump out of local optima, thereby indirectly promoting exploration (this approach is consistent with the method of maximizing the separation of positive and negative samples in contrastive learning strategies).

[0091] On the other hand, MACE will also be based on the Q inspired by WQMIX jt and Q tot The difference between is used to adjust the weight of parameter update, so MACE will Q jt The update weight is set to α = 1 when its estimated value is less than Q in the same batch. jt Q tot The update weight of is set to α<1, which emphasizes the importance of optimal joint action. The conclusions are as follows:

[0092]

[0093] Then calculate the optimization objective function, expressed as:

[0094]

[0095] in and Q jt and Q tot The target estimate of is calculated as and maxQ tot (τ′,u′,s′). ω is the scaling factor used to adjust Q jt (τ,u) and Q tot The difference between (τ,u) is used to encourage the model to jump out of the local optimal solution and find the optimal joint action more easily. intrinT represents the intrinsic target, which is Q jt (τ,u) and The purpose is to encourage the model to explore and avoid falling into the local optimum. r is the immediate reward, which means the reward obtained immediately after taking action in the current state. θ is the discount factor used to calculate the current value of future rewards, which determines the importance of future rewards relative to immediate rewards. maxQ tot(τ′,u′,s′) represents the maximum joint action value function Q of all possible actions u′ in the next state s′ tot The value of .

[0096] By optimizing the above intrinsic objective function, the action preference network parameters, the first joint value function estimation part and the network parameters of the second joint value function estimation part with restricted expression ability are obtained; then the network is deployed to each intelligent agent.

[0097]

[0098]

[0099]

[0100] Figure 4 We show that in SMAC, MACE outperforms all other baselines in all scenarios as model training progresses. Notably, MACE achieves unexpected results in two super-hard scenarios. In MMM2, MACE consistently becomes the best method and outperforms ResQ, which also uses the Q(λ) method to smooth the reward information obtained, a technique that has been widely proven to be very effective in RL model training. MACE also shows a clear advantage in the comparison between 3s_5z and 3s_6z. After 2 million training steps (in most other methods, the model's win rate can only be observed and distinguished after 5 million training steps), other methods can hardly achieve any win rate, while MACE steadily improves and finally achieves a median win rate of nearly 21%, demonstrating the advantage of the proposed method. Even in some relatively simple scenarios (2m_1z, 2s_1sc, bane_bane), MACE converges faster than other methods with minimal fluctuations, demonstrating the advantage of the method in training collaborative strategies after integrating different types of VD estimates. This result highlights the effectiveness of MACE in SMAC scenarios and its potential as a valuable tool in the development of more complex RL algorithms.

[0101] Figure 5It is shown that the MACE model with optimism outperforms its opponent, which is mainly reflected in the acceleration of policy convergence during training, enabling the collaborative model to achieve task success faster. Another possible explanation is that the optimistic component is constructed from the information difference between two different types of value functions, representing the difference between the value function in the restricted set and the value function in the unrestricted set. The additional optimism enables the restricted value function component to continuously reduce this difference during the training update process. It enables the model to break free from the local optimal solution and fit the true value function distribution faster, thereby improving the overall performance of the model. In other words, MACE is able to perceive the differences between the two VD estimates, and these differences are more conducive to the update of the model during training, thereby enhancing the overall training effect of the model.

[0102] Figure 6 It is shown that in the 5m_6m scenario, the action preference network improves the training efficiency of the model and eventually converges to a collaborative strategy with an almost 100% winning rate, which is not achievable by many other baseline methods. In the super difficult scenario MMM2, the action preference network can improve the final performance of the model by about 8%. The potential explanation is that in more difficult situations, the agent needs to adapt its strategy faster for collaborative training. The action preference network continuously adjusts the action selection mechanism according to the principle of maximizing the average policy entropy, enabling the agent to break through the limitations of ordinary ∈-greedy and make more specific collaborative action choices, thereby improving the final performance of the model.

[0103] Figure 7 We show that the update frequency of the action preference network has little effect on the performance in easy and difficult scenarios, affecting only the convergence speed of some policies, but in more complex collaborative scenarios, the update frequency has a significant impact on the final performance, probably because the agent encounters new observation and action features in each interaction with the environment, which are subject to considerable variation. When the action preference network is updated too quickly, it can cause instability in the agent policy entropy, leading to overly random action selection, which can negatively impact the training of the collaborative policy, especially when the uncertainty of action selection is high in the early stages of training. On the other hand, when the action preference network is updated slowly, it cannot keep up with the pace of the policy learning process, which can cause the agent to make incorrect action choices, thus hindering the collaborative process. Therefore, we set the update frequency of the action preference network between 50 and 100 in our model, as this setting strikes a balance and performs well in different training scenarios.

[0104] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A multi-agent comparative exploration method based on value decomposition difference, which exploits the estimation differences between different types of VD methods to promote exploration and guide the update of joint strategies, characterized by: The following steps are involved: (1) The agent strategy network uses a GRU-based recurrent neural network to make decisions: the agent’s observation value o and the previous action u are combined t-1 As the input of the agent strategy network; after the input data is encoded by the feedforward neural network composed of MLP, GRU and MLP in sequence, the output state action value function Q i (τ i ,); then the action selection mechanism output by the action preference network is used to select the collaborative action u corresponding to the highest value of the state action value function i , and calculate the state action value function Q under the collaborative action i (τ i ,u i ), τ i is the trajectory of the i-th agent; finally, the state action value function Q i (τ i ,u i ) is input to a first joint value function estimation part having full expression capability and a second joint value function estimation part having limited expression capability; (2) The first joint value function estimation part with full expressive power uses the self-attention mechanism to generate a centralized value function estimate Q jt (τ,u); the specific process is: (2-1) Statistics of the state action value function Q of all agents i (τ i ,u i ), denoted as n is the total number of agents; (2-2) The global state information S obtained by the agent through collecting surrounding environment information t The input is processed in a module consisting of MLP, attention mechanism module, soft attention module and Score function to obtain the weight ratio coefficient of the Attention mechanism; (2-3) The weight ratio coefficient of the Attention mechanism output by the Score function is Perform matrix multiplication to obtain the enhanced utility function (2-4) The enhanced utility function and global state information S t Input them into MLP for encoding processing respectively, and sum the data processed by the two MLPs to get the centralized value function estimate Q jt (τ,u); (3) The second joint value function estimation part with restricted expressiveness combines the utility functions of the agents through a restricted hybrid network to produce a centralized value function that can decompose the cooperative task. The specific process is: (3-1) Statistics of the state-action value function Q of all agents i (τ i ,u i ), denoted as (3-2) Decomposed into the state function of each agent and advantage function It is expressed as: (3-3) At time t Each Q i (τ i ,u i ) to perform argmaxQ i (τ i ,u i ) operation to obtain the joint optimal action consisting of the actions with the maximum action Q value of each agent (3-4) The joint optimal action and global state information S t Input the MLP and attention mechanism modules in turn, and use the attention mechanism to merge the global state information; And the data obtained by the attention mechanism module is combined with the advantage function Perform a dot product; (3-5) All the state functions of the agents Add them together, and then add them to the dot product result obtained in (3-4) to get the final concentrated value function (4) According to Q jt (τ,u) and The difference between them is used to adaptively calculate the intrinsic target; then the optimization objective function is calculated, which is expressed as: in and Q jt and Q tot The target estimate of is calculated as maxQ tot (τ′,u′,s′); ω is the scaling factor; intrinT represents the intrinsic goal, which is Q jt (τ,u) and The difference between them; r is the immediate reward, γ is the discount factor; the update weight is set to α<1; maxQ tot (τ′,u′,s′) represents the maximum joint action value function Q of all possible actions u′ in the next state s′ tot The value of By optimizing the above intrinsic objective function, the action preference network parameters, the first joint value function estimation part and the network parameters of the second joint value function estimation part with restricted expression ability are obtained; then the network is deployed to each intelligent agent.

2. The multi-agent comparative exploration method based on value decomposition difference according to claim 1 is characterized in that: The action preference network has two output components: Q function estimation and action-value distribution; Q function estimation is used to calculate the state action value function; action-value distribution is used to select a greedy action with probability ∈ according to the value of the state action value function, or to sample actions according to action preference with probability 1-∈, and execute the action to trigger the next state transition.