Robust reinforcement learning method for heterogeneous multi-agent system
Through the robust reinforcement learning method of hierarchical decomposition and dual regularization mechanism, the environmental adaptability and collaboration efficiency of heterogeneous multi-agent systems are improved, and the problems of robustness and stability in heterogeneous agent systems are solved.
Patent Information
- Application Number
- CN202510327117.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
AI Technical Summary
The existing multi-agent reinforcement learning methods are difficult to effectively deal with environmental noise and action interference in heterogeneous agent systems, resulting in a degradation of system performance and insufficient robustness, especially in dynamic environments, which are difficult to maintain stability and collaboration efficiency.
The hierarchical decomposition strategy is adopted, type-level value functions and hyper-strategy networks are introduced, combined with the dual regularization mechanism, and the coordination and cooperation capabilities of heterogeneous agents are improved through type-level hybrid networks and global hybrid networks, and robustness is enhanced at the observation and action levels.
It improves the robustness and collaboration efficiency of heterogeneous multi-agent systems in dynamic environments, can effectively deal with environmental noise and action interference, and ensures that the system maintains stable operation in complex environments.
Smart Images

Figure CN120278225A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of heterogeneous multi-agent reinforcement learning, and particularly relates to a robust reinforcement learning method for heterogeneous multi-agent systems. Background Art
[0002] Multi-agent Reinforcement Learning (MARL) is a technology involving multiple agents interacting and learning in a shared environment, aiming to achieve an optimal solution for the entire system through cooperation or competition among agents. Each agent has the ability to perceive the environment, select actions, and obtain rewards. MARL has a wide range of applications in the real world, including fields such as cooperative robot systems, traffic management, smart grid scheduling, and network routing optimization. Through MARL technology, collaborative learning and optimization among multiple agents can be achieved, improving the overall performance and efficiency of the system.
[0003] In the field of multi-agent reinforcement learning, the existence of heterogeneous agents is a challenge. Since different agents may have different capabilities, goals, and learning preferences, this makes algorithm design and system stability more complex. For example, in an intelligent transportation system, there are both autonomous vehicles that need to comply with traffic rules and traffic lights that need to be flexibly scheduled. Their respective action spaces, perception capabilities, and learning strategies are very different. How to make them cooperate efficiently and optimize the overall traffic flow poses higher requirements for the compatibility and adaptability of the algorithm. If the heterogeneity problem is not effectively solved, the performance of the multi-agent system may be greatly reduced, or even the established goals cannot be achieved. In addition, the robustness problem has always been a crucial research direction in the field of multi-agent reinforcement learning because multi-agent systems in the real world often face various uncertainties and changes, such as environmental disturbances, agent failures, policy differences, and malicious behaviors of teammates. In the context of heterogeneous agents, on the one hand, robustness requires the algorithm to maintain stable performance under uncertain factors such as observation noise, joint action disturbances, or environmental dynamic changes; on the other hand, the diversity of heterogeneous agents leads to differences in the sensitivity and response mechanisms of different agents when facing disturbances, making it difficult for a unified robustness design to meet the needs of each agent. Therefore, studying robust MARL algorithms in the context of heterogeneous agents is the key to ensuring the reliable operation of multi-agent systems in complex and dynamic environments.
[0004] Currently, the mainstream method in the MARL field is to adopt the centralized training and decentralized executing (CTDE) paradigm. The main idea is that each agent learns its policy based on local observations, and at the same time, a centralized mechanism is designed to guide the training of all agents. In most such methods, the agents share their neural network parameters to calculate actions, which is called parameter sharing (PS). This method enables single-agent reinforcement learning algorithms to be directly applied to multi-agent reinforcement learning without introducing excessive computational and sample complexity burdens as the number of agents increases. Therefore, in the field of multi-agent reinforcement learning, PS technology is a common practice to improve sample efficiency and enhance algorithm performance. However, since the PS technology shares the network parameters among agents, it will make the agents show similar behavioral characteristics, that is, the homogenization of agents, which will hinder the emergence of agent diversity, limit the exploration ability of agents, reduce the final cooperation performance, and cause many algorithms to perform poorly in heterogeneous agent settings. For the methods starting from the PS technology, in order to encourage agent diversity to better handle heterogeneous agent tasks, either the way of not sharing parameters at all or the way of selective parameter sharing is adopted. The former often has non-stationary problems because the agents have independent decision-making networks, especially in tasks with a large number of agents. The latter, although to a certain extent balances the efficiency and diversity of multi-agent algorithms in heterogeneous agent tasks by using the learned classification method for selective parameter sharing, it is difficult to ensure the rationality of the classification scheme to better handle complex heterogeneous agent tasks with dynamic changes. Moreover, the above heterogeneous multi-agent methods often lack consideration of the robustness issues of the methods. The current robust multi-agent research focuses on enhancing the robustness of the algorithm by adding perturbations, adding interference and noise in multiple aspects such as actions, observations, and communication content, and using high-complexity mathematical modeling and optimization methods to improve the robustness of the algorithm in various aspects. However, these methods often have high computational complexity, strict theoretical assumptions are required and it is difficult to apply, and usually do not consider the heterogeneous agent task environment. Summary of the Invention
[0005] The object of the present invention is to provide a robust reinforcement learning method for heterogeneous multi-agent systems, which addresses the policy learning challenges in heterogeneous multi-agent systems due to differences in agent capabilities, state spaces or action spaces, improves the robustness of the model in the face of environmental noise and action interference, and enables it to effectively adapt to dynamic environments where the number or type of agents changes, thereby improving the stability and cooperation efficiency of heterogeneous multi-agent systems in complex dynamic scenarios.
[0006] To achieve the above functions, the present invention designs a robust reinforcement learning method for heterogeneous multi-agent systems, which performs the following steps A - E to complete the robust reinforcement of the multi-agent system facing environmental noise and action interference:
[0007] Step A: There are multiple types of agents and a buffer in the task environment. The types of agents are discretely defined, and each type of agent has a preset clear function set and ability limit for the task in the task environment. The task environment provides the global state at the current time step, the type information of each agent under the current task, and the local observations for each agent respectively. Each agent has its own local utility network. Each agent obtains its own local observation and makes an action according to its local utility network. The joint action is composed of the actions of all agents. The task environment receives the joint action feedback reward and updates the global state at the next time step. Add random noise to the local observations for each agent and mark them as perturbed observations, perform random perturbation processing on the joint action and mark it as the perturbed joint action, store the interaction information between each agent and the task environment in the buffer. If step A is executed for the first time, initialize the time step, otherwise increment the current time step by one, and then enter step B;
[0008] Step B: Select the interaction information between the agents and the task environment from the buffer, input the local observations and perturbed observations of each agent in the interaction information into the local utility network of each agent, respectively obtain the local action values under normal observations and the local action values under perturbed observations for each agent, calculate the loss according to the observation robust regularization constraint formula and update the local utility network parameters to complete the training of the local utility network, and then enter step C;
[0009] Step C: Based on the number of types of agents under the current task, construct the same number of type-based hybrid networks based on the type attention network. According to the type of agent, input the global state, the local observation of each agent, and the local action values corresponding to the joint action before perturbation and the joint action after perturbation under normal observations into the type-based hybrid network corresponding to its type respectively. After receiving the local observations, the joint action before perturbation and the local action values corresponding to the joint action after perturbation, and the global state of all agents of its corresponding type, the type-based hybrid network generates the action values of the joint action before perturbation and the joint action after perturbation based on their respective types, and then enter step D;
[0010] Step D: Input the global state and the type-based action values of all pre-disturbance joint actions and post-disturbance joint actions into the global mixing network respectively. The global mixing network will generate the global action values of the pre-disturbance joint action and the post-disturbance joint action respectively. Calculate the total loss according to the TD loss formula and the joint action robust regularization constraint formula and update all network parameters according to the chain rule of gradient calculation. If the current time step is greater than the preset maximum time step T, complete the training of the global mixing network and enter Step E; otherwise, return to continue executing Step A.
[0011] Step E: Based on the trained local utility network, type-based mixing network, and global mixing network of the agents, complete the robustness enhancement of the multi-agent system facing environmental noise and action interference. Deploy the trained agents in the task environment and make decisions according to the information provided by the task environment until the task is completed.
[0012] Advantageous effects: Compared with the prior art, the advantages of the present invention include:
[0013] Compared with the traditional value decomposition framework that directly decomposes the global action value function into local action values, this method introduces a type-level value function. The joint action value function is decomposed into type-level value functions using the global state, and then the type-level value functions are further decomposed into local action values using the global state, local observations, and other information. This hierarchical decomposition strategy makes full use of information at different levels, taking into account both the global cooperation requirements and the specific roles and capabilities of different types of agents, further enhancing the coordination and cooperation between heterogeneous agents from the perspective of value decomposition, which is an advantage not possessed by existing methods. To further improve the adaptability of the method in a dynamic environment and its robustness against environmental interference, the hyper-policy network is used as the local utility function. With the help of the hyper-parameter network, when the observed content changes in arrangement due to environmental dynamic changes, it can ensure that the internal data calculation of the network remains unchanged, while being able to adjust the arrangement order of the final output action values according to the sensitivity of different actions to the arrangement order, minimizing the adjustment amplitude while ensuring that the output action values contain order information, greatly reducing the observed action space, so that the local utility function can explore with higher efficiency and better cope with the dynamic changes of the environment. On this basis, to enhance the robustness against environmental interference, this method introduces a dual regularization mechanism. In the joint action value function and the local action value function, a joint action perturbation regularization constraint and an individual observation noise regularization constraint are respectively introduced. Different from the single-dimensional consideration of most methods, this dual regularization mechanism optimizes the robustness for the perturbation at the joint action level and the noise at the individual observation level respectively, aiming to achieve a more comprehensive improvement in robustness. Description of the Drawings
[0014] Figure 1 It is the overall structural framework diagram of a robust reinforcement learning method for heterogeneous multi-agent systems provided according to an embodiment of the present invention;
[0015] Figure 2 It is the structural diagram of the local utility network of the agent provided according to an embodiment of the present invention;
[0016] Figure 3 It is the structural diagram of the type-based hybrid network provided according to an embodiment of the present invention. Detailed implementation manners
[0017] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0018] A robust reinforcement learning method for heterogeneous multi-agent systems provided according to an embodiment of the present invention, whose application scenarios include: intelligent warehousing tasks, urban traffic control, etc. In the intelligent warehousing task, the robots are the agents. The available robots are respectively the handling robot, the picking robot and the packing robot according to their functions. The type-based hybrid network is used to independently consider the decisions of different types of robots, and at the same time consider the warehouse environment noise and the possible abnormalities in the robot actions. The trained handling robot optimizes the path planning through deep reinforcement learning. Even when there is noise in the sensor data, it can avoid obstacles through historical experience and alternative path planning to ensure the delivery of goods. The picking robot combines the reinforcement learning model of visual recognition and robotic arm control. Even in the case of light changes or slight deviations in the placement position of the goods, it can accurately grasp the target goods. The packing robot optimizes the packing process through reinforcement learning. Even when encountering goods of different sizes and shapes, it can adjust the packing strategy to ensure the firmness and efficiency of the packaging. Through the design of the global reward function, the three types of robots are encouraged to assist and cooperate in behavior. For example, the handling robot gives priority to responding to the needs of the picking robot, and the picking robot adjusts the picking order according to the processing ability of the packing robot, so as to achieve the robust and efficient operation of the entire warehousing system.
[0019] In an urban traffic control system, agents can include intelligent traffic lights and autonomous vehicles. As control nodes, intelligent traffic lights use a deep reinforcement learning algorithm to optimize signal timing through an independent type-mixed network. Even if some cameras have inaccurate traffic flow data due to weather or occlusion, the traffic lights can still make robust timing adjustments by leveraging information feedback from adjacent intersections, historical traffic data, and prediction models due to the robustness to environmental noise during policy learning, effectively alleviating congestion. As road users, autonomous vehicles train their driving strategies through deep reinforcement learning using an independent type-mixed network. Even when there are foreign object occlusions or poor signal feedback in the vehicle's body radar sensors, autonomous vehicles can ensure safe driving with their robust decision-making models. By designing a global reward function to encourage cooperation between different agents, traffic lights can optimize signal timing based on the average driving speed and route planning of autonomous vehicles; autonomous vehicles can adjust their route planning according to the traffic flow data and congestion conditions observed at each traffic light in the traffic network, avoid congested areas, and be guided to a smoother route. This collaborative mechanism enables the entire urban traffic system to remain efficient, stable, and safe in the face of various noises and disturbances.
[0020] A robust reinforcement learning method for a heterogeneous multi-agent system provided by an embodiment of the present invention, referring to Figure 1 , perform the following steps A - step E to complete the robustness enhancement of the multi-agent system facing environmental noise and action interference:
[0021] Step A: There are multiple types of agents and buffers in the task environment. The types of agents are discretely defined, and each type of agent has a preset clear function set and ability limit for the task in the task environment. Taking the task environment in the embodiment as an intelligent warehouse, for the intelligent warehouse task, the types of agents include, but are not limited to, handling, picking, and packing robots, and their respective function sets and ability limits include path planning, handling, and conveying goods; grasping and sorting target goods; packaging goods of different shapes and sizes, etc. In reality, the capabilities of agents may be distributed in a continuous space, and discrete types cannot accurately express the subtle differences in capabilities. The characteristics of agents may change dynamically over time, task environment, or task content, and it is difficult for fixed agent type classification to capture this dynamicity. The above situations are not within the scope of consideration in the embodiments of the present invention.
[0022] The task environment provides the global state at the current time step, the type information of each agent under the current task, and the local observations for each agent respectively. Each agent has its own local utility network. Each agent obtains its own local observation and makes an action based on its own local utility network. The joint action is composed of the actions of all agents. The task environment receives the feedback reward of the joint action and updates the global state at the next time step. Random noise is added to the local observations for each agent respectively and marked as the perturbed observation, and random perturbation processing is performed on the joint action and marked as the perturbed joint action. The interaction information between each agent and the task environment is stored in the buffer. If step A is executed for the first time, the time step is initialized, otherwise the current time step is incremented by one, and then step B is entered.
[0023] The specific steps of step A are as follows:
[0024] Step A.1: The task environment provides the global state s at the current time step t t , the type information M of each agent, and the local observations for each agent respectively The type information M includes the type information type to which each agent belongs i and the total number m of agent categories under the current task. Subsequently, each agent executes an action based on its received local observation according to its own local utility network and obtains the local action value corresponding to the action where represents the local utility network and θ local represents the local utility network parameters;
[0025] Each agent selects the action with the maximum value in according to the ε-greedy strategy and executes it. All actions form the joint action After the task environment receives the joint action A t , it gives the feedback reward r of the current task environment under this joint action A t , and updates the global state to s t , where 1 ≤ i ≤ I and I represents the total number of agents, t+1 represents the action of agent i among all agents at time step t;
[0026] If step A is executed for the first time, the current time step t = 0 is initialized and the current time step is recorded. Otherwise, t = t + 1 is executed, and the recorded time step t is incremented by one as the current time step and the current time step is recorded;
[0027] Step A.2: Add random noise to the local observations for each agent respectively For the local observation Do the range of the value domain o Preset the ratio as λ o Add Gaussian noise with the above preset ratio λ, mark it as the perturbed observation, and express it as the following formula:
[0028]
[0029] Among them, represents the perturbed observation of agent i, and δ o is the observation perturbation function;
[0030] Perform interference processing on the joint actions of agents. Randomly select a preset ratio λ a of agents among all agents, and change their actions to Mark it as the perturbed joint action and express it as the following formula:
[0031]
[0032] Among them, represents the perturbed joint action, and δ a is the action perturbation function;
[0033] It should be noted that for the noise addition process on the local observations of each agent and the interference processing on the joint actions of agents, the λ parameter in the formula only represents that the added interference randomness is fixed to a certain extent statistically, but the interference itself (i.e., Gaussian noise and the probability of modifying joint actions) is random.
[0034] Step A.3: Record the interaction information between each agent and the task environment in the current time step and store it in buffer B. The interaction information includes the information provided by the task environment to each agent and the actions taken by the agents accordingly The information provided by the task environment includes the type information M of each agent, the local observation at the current time step the global state s t , the feedback reward r t , the global state s at the next time step t+1 , the global state s at the next time step t+1 , as well as the perturbed observations of each agent and the perturbed joint actions
[0035] Step A.4: Judge whether the current content volume |B| of buffer B meets the capacity condition |B|≥batchsize. When the capacity condition is met, judge whether the current time step t meets the training interval condition t - t last train≥interval. If the above two conditions are met, select interaction information with a quantity of batchsize from buffer B, update the time step t when entering step B last time last train =t, and enter step B. If the above two conditions are not met, repeat step A; where batchsize is the preset minimum trainable buffer capacity, interval is the preset shortest training time interval, and t last train is the time step when successfully entering step B from step A last time. If it is the first time to execute step A, initialize t last train =0 and record t last train .
[0036] Step B: Select the interaction information between the agent and the task environment from the buffer, input the local observations and perturbed observations of each agent in the interaction information into the local utility network of each agent, respectively obtain the local action values under normal observations and the local action values under perturbed observations of each agent, calculate the loss according to the observation robustness regularization constraint formula, and update the local utility network parameters to complete the training of the local utility network, and then enter step C;
[0037] The specific method of step B is as follows:
[0038] Randomly extract interaction information from the buffer, and input the local observations and perturbed observations of each agent in each piece of interaction information into the local utility network of the agent, and execute and where represents the local utility network, and θ local represents the local utility network parameters; obtain the local action value under normal observations and the local action value under perturbed observations. Then calculate the observation robustness loss value where R o represents the observation robustness regularization constraint, which is calculated using the KL divergence, and its calculation formula is as follows:
[0039]
[0040] where P and Q are two distributions about the variable x, and p(x) and q(x) represent the probabilities of the variable x in distributions P and Q, respectively;
[0041] The observation robustness loss value loss o Calculate the gradient of the local utility network parameter θ local as follows:
[0042]
[0043] Update the local utility network parameter θ by means of stochastic gradient descent, and its update formula is as follows: local As follows:
[0044]
[0045] where α is a preset learning rate, is the gradient of the local utility network parameter θ local .
[0046] The local utility network described is a hyper-policy network structure, and its structure refers to Figure 2 , including a permutation-invariant input layer layer PI , a decision network processing layer layer policy , a fully connected layer layer FC and a permutation-equivariant output layer layer PE . The specific steps for the local utility network to calculate the local action value q according to the local observation o are as follows:
[0047] Step B.1: Input the local observation o into the permutation-invariant input layer layer of the local utility network PI . In the permutation-invariant input layer layer PI , for each entity feature component o[j] of the local observation o, execute the following formula:
[0048] W in [j] = hp in (o[j]; θ in )
[0049] where hp in represents an input-invariant hypernetwork. The input-invariant hypernetwork hp in consists of a multi-layer perceptron MLP, and its network parameter is θ in , which belongs to the local utility network parameter θ local ; W in [j] represents the weight matrix generated by the input-invariant hypernetwork hp in for each entity feature component o[j] of the local observation o; where 1 ≤ j ≤ m o , m o represents the number of entities to which the entity features included in the content of the local observation o belong;
[0050] Based on the weight matrix W in[j], execute the following formula to obtain the output z of the permutation-invariant input layer layer PI :
[0051]
[0052] where is a new local observation generated by any permutation combination with each entity feature component o[j] of the local observation o as the basic unit
[0053] Step B.2: Input the output z of the permutation-invariant input layer layer PI into the decision network processing layer layer policy , execute the following formula to obtain the output h of the decision network processing layer layer policy :
[0054] h = layer policy (z; θ policy )
[0055] where the decision network processing layer layer policy consists of a recurrent neural network RNN with its network parameters being θ policy , belonging to the local utility network parameters θ local ;
[0056] Step B.3: Input the output h of the decision network processing layer layer policy into the permutation-equivariant output layer layer PE and the fully connected layer layer FC respectively. The permutation-equivariant output layer layer PE outputs the permutation-equivariant action value q equiv , and the fully connected layer layer FC outputs the permutation-invariant action value q inv . Concatenate q equiv and q inv to generate the local action value q as the output of the agent's local utility network .
[0057] In Step B.3, the process of the permutation-equivariant output layer layer PE outputting the permutation-equivariant action value q equiv is as follows:
[0058] In the permutation-equivariant output layer layer PE , execute the following formula:
[0059] W out [j] = hp out (o[j]; θ out )
[0060] Among them, hp out represents the output-invariant hypernetwork. The output-invariant hypernetwork hp out is composed of a multi-layer perceptron MLP, and its network parameters are θ out , belonging to the local utility network parameters θ local , W out [j] represents the weight matrix generated by the output-invariant hypernetwork hp out for each entity feature component o[j] of the local observation o; where 1 ≤ j ≤ m o , m o represents the number of entities to which the entity features included in the content of the local observation o belong;
[0061] For W out [j], execute the following formula to obtain the permutation-equivariant action value q equiv for each action value part q equiv [j]:
[0062]
[0063] Among them, h represents the output of the decision network processing layer layer policy ;
[0064] Concatenate all action value parts q equiv [j] to obtain the permutation-equivariant action value q equiv ;
[0065] In step B.3, the fully connected layer layer FC outputs the permutation-invariant action value q inv The process is as follows:
[0066] In the fully connected layer layer FC execute the following formula to obtain the permutation-invariant action value q inv :
[0067] q inv = layer FC (h; θ FC )
[0068] Among them, θ FC is the fully connected layer network parameter, belonging to the local utility network parameter θ local , h represents the output of the decision network processing layer layer policy .
[0069] The local utility network obtains the local action value under normal observation according to the agent's local observation and according to the perturbed observation Obtain the local action value under perturbed observations All adopt the above calculation process.
[0070] Step C: According to the number of agent types in the current task, construct the same number of type-based hybrid networks based on the type attention network. According to the type of the agent, input the global state, the local observation of each agent, and the local action values corresponding to the pre-perturbation joint action and the post-perturbation joint action under normal observations into the type-based hybrid network corresponding to its type respectively. After receiving the local observations of all agents of its corresponding type, the local action values corresponding to the pre-perturbation joint action and the post-perturbation joint action, and the global state, the type-based hybrid network generates the action values of the pre-perturbation joint action and the post-perturbation joint action based on their respective types, and then enters Step D;
[0071] The specific steps of Step C are as follows:
[0072] Step C.1: Classify the agents according to the agent type information M provided by the current task environment, and construct a type-based hybrid network for each type to which the agent belongs Where represents the type-based hybrid network of the 1st...m types, and m is the total number of agent types;
[0073] Step C.2: From the perspective of each agent, according to the type to which each agent belongs, input the local observation the global state s t and the local action value corresponding to the pre-perturbation joint action under normal observations and the local action value corresponding to the post-perturbation joint action into the type-based hybrid network of the type type i to which agent i belongs, and execute the following formula:
[0074]
[0075] Where, represents the action value of the type to which agent i belongs under pre-perturbation under normal observations, represents the action value of the type to which agent i belongs after perturbation, represents the network parameters of the type-based hybrid network of the type to which agent i belongs; represents the type-based hybrid network of the type to which agent i belongs, represents the local action value corresponding to the pre-perturbation joint action of the type to which agent i belongs under normal observations, Denote the local action value corresponding to the joint action after perturbation of the type to which agent \(i\) belongs. Denote the local observation of the type to which agent \(i\) belongs.
[0076] Step C.3: From the perspective of agent types, according to the agent type, obtain the type-based action value of the joint action before perturbation under normal observation and the type-based action value of the joint action after perturbation Specifically, as shown in the following formula:
[0077]
[0078] where Denote the action value of type \(j\) of the joint action before perturbation under normal observation. Denote the action value of type \(j\) of the joint action after perturbation. \(m\) is the total number of agent types.
[0079] The expression of is as follows:
[0080]
[0081] where Denote the local action value corresponding to the joint action before perturbation of all agents belonging to type \(j\) under normal observation. Denote the local action value corresponding to the joint action after perturbation of all agents belonging to type \(j\). Denote the type-based hybrid network belonging to type \(j\), \(\theta\) type j Denote the network parameters of the type-based hybrid network belonging to type \(j\), \(\theta\) type Denote the set of network parameters \(\theta\) containing all type-based hybrid networks type \(=\{\theta\) type _1,...,\theta\) type m \}\), \(\{o\) i \}\) type j is the local observation of the agent belonging to type \(j\).
[0082] For convenience of representation, denote as \(\{q\) i \}\) type j , denote as \(Q\) type j In step C.3, taking the local action value \(\{q\) i \}\) type j as an example, the local action value \(\{q\) i \}\)type j , the global state s, and the local observation {o i} type j Input the type-based hybrid network Obtain the type-based action value Q belonging to type j type j The specific process is as follows:
[0083] Refer to Figure 3 , in the type-based hybrid network , the local action value {q i} type j the global state s, and the local observation {o i} type j will first be input into the type attention network, and the structure of the type attention network refers to the right gray part in Figure 3 , and execute the following formula:
[0084]
[0085] where, is the type attention network in the type-based hybrid network belonging to type j, and its network parameters are is the type-based local action value belonging to type j;
[0086] In the type attention network , the observation content {o i} type j will first perform the attention scaled dot product operation with the global state s, then normalize the operation result through softmax, and finally perform the matrix multiplication operation with the local action value {q i} type j to obtain the type-based local action value belonging to type j The specific process is as follows:
[0087] For the global state s and the local observation {o i} type j Execute the following formula to obtain the query content Q and the key content K:
[0088] Q = s × W q
[0089] K = o × W k
[0090] Among them, s represents the global state, o represents the local observation, and W q and W k are attention weight matrices;
[0091] Perform matrix multiplication on the query content Q and the key content K, scale the operation result by the square root of the matrix dimension size, and finally perform softmax probability normalization on the result and multiply it with the local action value {q i} type j Execute matrix multiplication to obtain the type-based local action value belonging to type j
[0092]
[0093] where d k is the second dimension size of the query content Q and the key content K matrices;
[0094] For Execute the following formula to obtain the type-based action value Q belonging to type j type j :
[0095]
[0096] where W type j and b type j respectively represent the weight matrix and the bias vector in the weighted bias process in the type-based hybrid network belonging to type j, and Q type j is the type-based action value belonging to type j;
[0097] The weight matrix W type j and the bias vector b type j are respectively composed of two hypernetworks hp w type j and hp b type j , execute W type j = hp w type j (s; θ w type j ) and b type j = hp b type j(s; θ b type j ) are generated. The inputs of both hyper-networks are the global state s, and the structure of both is composed of a multi-layer perceptron MLP, where θ w type j and θ b type j are the network parameters of hyper-networks hp w type j and hp b type j respectively. The network parameter θ of the type-based hybrid network belonging to type j type j contains the network parameters of the type attention network in the type-based hybrid network belonging to type j and is the two hyper-network parameters θ of the type-based hybrid network belonging to type j w type j and θ b type j .
[0098] The perturbed local action value the global state s, and the local observations {o i} type j are input into the type-based hybrid network to obtain the perturbed type-based action value belonging to type j The process is the same as the above process.
[0099] Step D: The global state and the type-based action values of all pre-perturbation joint actions and post-perturbation joint actions will be fed into the global hybrid network respectively. The global hybrid network will generate the global action values of the pre-perturbation joint action and the post-perturbation joint action respectively. The total loss will be calculated according to the TD loss formula and the joint action robust regularization constraint formula, and all network parameters will be updated according to the chain rule of gradient calculation. If the current time step is greater than the preset maximum time step T, the training of the global hybrid network is completed, and step E is entered; otherwise, return to step A and continue to execute;
[0100] The schematic diagram of step D refers to Figure 1 the uppermost part in, and the specific steps are as follows:
[0101] Step D.1: The global state s t and the type-based action values of all pre-perturbation joint actions and post-perturbation joint actions under all normal observations and are respectively sent into the global mixing network to execute the following formula:
[0102]
[0103] where is the global action value of the joint action before perturbation, is the global action value of the joint action after perturbation; θ total is the network parameter of the global mixing network ;
[0104] The global mixing network adopts dual weighted bias processing, in which, based on the type of action value Q type two weighted bias calculations will be executed in sequence:
[0105] Q total = [Q type × W1 total + b1 total × W2 total + b2 total
[0106] where Q total represents the global action value, W1 total and W2 total are respectively the weight matrices in the first weighted bias and the second weighted bias operations of the global mixing network , b1 total and b2 total are respectively the bias vectors in the first weighted bias and the second weighted bias operations;
[0107] Each weight matrix and bias vector are generated by four different hypernetworks. Executing W1 total = hp w1 total (s; θ w1 total ) and W2 total = hp w2 total (s; θ w2 total ) respectively obtain the weight matrices W1 total and W2 total in the first and second weighted bias operations. Executing b1 total = hp b1 total (s; θ b1 total ) and b2 total = hp b2 total (s; θb2 total ) respectively obtain the bias vectors b1 total and b2 total , all four hypernetwork inputs are the global state s, and the structure is composed of a multi-layer perceptron MLP, where hp w1 total and hp w2 total are respectively the hypernetworks used to generate the weight matrices in the first and second weighted bias operations, and θ w1 total and θ w2 total are respectively the network parameters of the hypernetworks hp w1 total and hp w2 total ; hp b1 total and kp b2 total are respectively the hypernetworks used to generate the weight matrices in the first and second weighted bias operations, and θ b1 total and θ b2 total are respectively the network parameters of the hypernetworks hp b1 total and hp b2 total ; the network parameter θ of the global mixing network total contains the network parameters θ of all the hypernetworks used to generate the weight matrices and bias vectors in the first and second weighted bias operations in the global mixing network w1 total and θ w2 total , and θ b1 total and θ b2 total ;
[0108] After updating the global network parameter θ Total , determine whether the current time step t satisfies the training termination condition t ≥ T, where T is the preset maximum training time step. If the condition is satisfied, enter step E; otherwise, return to execute step A.
[0109] Step D.2: According to the feedback reward r t given by the task environment at the same time step taken from the buffer B , combined with the global action value of the joint action before perturbation Calculate the TD loss value and the joint action robust loss value respectively according to the TD loss formula and the joint action robust regularization constraint formula and the joint action robust loss value where the joint action robust regularization constraint formula R a () is calculated using the root mean square error RMSE, as shown in the following formula:
[0110]
[0111] where y and y ′ represent the true vector and the predicted vector respectively, ‖‖2 represents the L2 norm of the vector, and n is the dimension size of the vector;
[0112] TD loss formula adopts the classical formula of value-based reinforcement learning, and its calculation formula is as follows:
[0113]
[0114] where γ is the time discount factor, is the network parameter θ of the globally mixed network that is periodically replicated and A total is the global joint action at time t+1. In practical applications, since the globally mixed network t+1 does not directly receive the global joint action A and the global state s, but receives the type-based action value generated by the type-based mixed network in step C so the above formula will adopt the following form in the application process: Therefore, in the application process, the above formula will adopt the following form:
[0115]
[0116] where, is the type-based action value at time step t+1, which is generated by the type-based mixed network in step C receiving the observations of each agent belonging to time step t+1 in the buffer B local action value and the global state s t+1 for processing and generation;
[0117] Step D.3: The TD loss value loss TD and the joint action robust loss value loss a calculated according to the TD loss formula and the joint action robust regularization constraint formula are subjected to an addition process to obtain the global loss value loss Total :
[0118] loss Total = loss TD+loss a
[0119] Utilize the global loss value loss total to calculate the global loss gradient
[0120]
[0121] where θ Total is the global network parameter, including the network parameter θ of the local utility network local , the network parameter θ of all types of type-based hybrid networks type and the network parameter θ of the global hybrid network total ;
[0122] After calculating the global loss gradient , update the global network parameter θ Total by means of stochastic gradient descent according to the preset learning rate. The stochastic gradient descent formula is as follows:
[0123]
[0124] where α is the preset learning rate.
[0125] Step E: Based on the local utility network, type-based hybrid network, and global hybrid network of the trained agents, complete the robustness enhancement of the multi-agent system facing environmental noise and action interference, deploy the trained agents in the task environment, and execute decisions according to the information provided by the task environment until the task is completed.
[0126] In Step E, deploy all the trained agents in the task environment. The local utility network of all agents will i execute and select the action a with the maximum value from the local action value q i after obtaining the local observation o i given by the task environment, and execute it until the task completion degree reaches the required expectation.
[0127] Compared with traditional value decomposition multi-agent reinforcement learning techniques, the method designed in the present invention addresses the differences among different types of agents in heterogeneous multi-agent systems by introducing a type-level hybrid network. Using the local observations and global states of each agent, a type-level value function is constructed in addition to the global-layer value function and local-layer value function, thereby capturing the inherent characteristics and diverse capabilities of different types of agents. This hierarchical decomposition strategy fully utilizes information at different levels, taking into account both the global cooperation requirements and the specific roles and capabilities of different types of agents. To further improve the adaptability of the method in a dynamic heterogeneous multi-agent task environment, the hyper-policy network is used as the local utility network. With the help of the internal hyper-parameter network, different types of agents can adaptively adjust the permutation relationship between the observed content and the local action through the local utility network, and while ensuring decision consistency, reduce the local decision space dimension and improve the exploration efficiency of different types of agents. On this basis, to enhance the robustness of the method against disturbances in the real task environment, the present method enhances the robustness of the system synchronously from two dimensions: perturbations at the joint action level and noise at the individual observation level by designing a dual regularization mechanism, ensuring that the system has a more robust adaptability in complex and dynamic environments.
[0128] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A robust reinforcement learning method for heterogeneous multi-agent systems, characterized in that, Perform the following steps A - E to complete the robustness enhancement of the multi - agent system in the face of environmental noise and action interference: Step A: There are multiple types of agents and a buffer in the task environment. The types of agents are discretely defined, and each type of agent has a preset and clear set of functions and ability limitations for the task in the task environment. The task environment provides the global state at the current time step, the type information of each agent under the current task, and the local observations for each agent respectively. Each agent has its own local utility network. Each agent obtains its own local observation and makes an action based on its local utility network. The joint action is composed of the actions of all agents. The task environment receives the joint action feedback reward and updates the global state at the next time step. Add random noise to the local observations for each agent and mark them as perturbed observations, perform random perturbation processing on the joint action and mark it as the perturbed joint action, store the interaction information between each agent and the task environment in the buffer. If step A is executed for the first time, initialize the time step, otherwise increment the current time step by one, and then enter step B; Step B: Select the interaction information between the agents and the task environment from the buffer. Input the local observations and perturbed observations of each agent in the interaction information into the local utility network of each agent to obtain the local action values under normal observations and the local action values under perturbed observations for each agent respectively. Calculate the loss according to the observation robustness regularization constraint formula and update the local utility network parameters to complete the training of the local utility network, and then enter step C; Step C: Based on the number of agent types under the current task, construct the same number of type - based hybrid networks based on the type attention network. According to the type of agent, input the global state, the local observation of each agent, and the local action values corresponding to the joint action before perturbation and the joint action after perturbation under normal observations into the type - based hybrid network corresponding to its type respectively. After receiving the local observations, the local action values corresponding to the joint action before perturbation and the joint action after perturbation, and the global state of all agents of its corresponding type, the type - based hybrid network generates the action values of the joint action before perturbation and the joint action after perturbation based on their respective types, and then enter step D; Step D: Input the global state and the type - based action values of all joint actions before perturbation and joint actions after perturbation into the global hybrid network. The global hybrid network generates the global action values of the joint action before perturbation and the joint action after perturbation respectively. Calculate the total loss according to the TD loss formula and the joint action robustness regularization constraint formula and update all network parameters according to the chain rule of gradient calculation. When the current time step is greater than the preset maximum time step T, complete the training of the global hybrid network and enter step E, otherwise return to continue executing step A; Step E: Based on the local utility network, type-based hybrid network, and global hybrid network of the trained agent, complete the robustness enhancement of the multi-agent system facing environmental noise and action interference. Deploy the trained agents in the task environment and execute decisions according to the information provided by the task environment until the task is completed.
2. The robust reinforcement learning method for a heterogeneous multi-agent system according to claim 1, characterized in that The specific steps of Step A are as follows: Step A.1: The task environment provides the global state s at the current time step t t , the type information M of each agent, and the local observations for each agent respectively The type information M includes the type information type to which each agent belongs i and the total number m of agent categories in the current task. Subsequently, each agent executes an action according to its received local observation and obtains the local action value corresponding to the action based on its respective local utility network where represents the local utility network, and θ local represents the local utility network parameters; Each agent is selected according to the ε-greedy strategy The action with the highest value And execute, all actions constitute a joint action The task environment receives the joint action A t After that, given the current task environment in this joint action A t The feedback reward r t , and update the global state to s t+1 , where 1≤i≤I, I represents the total number of agents, represents the action of agent i among all agents at time step t; If Step A is executed for the first time, initialize the current time step t = 0 and record the current time step. Otherwise, execute t = t + 1, increment the recorded time step t by one as the current time step, and record the current time step; Step A.2: Respectively for the local observations of each agent Add random noise to the local observations and perform a Gaussian noise superposition with a preset ratio of λ o in the value range of range o The perturbed observations are marked and expressed by the following formula: Among them, represents the perturbed observation of agent i, and δ o is the observation perturbation function; Perform interference processing on the joint actions of agents, and randomly select agents with a preset ratio λ a among all agents, and change their actions to which is marked as the perturbed joint action and expressed by the following formula: Among them, represents the combined action after perturbation, and δ a is the action perturbation function; Step A.3: Record the interaction information between each agent and the task environment in the current time step and store it in buffer B, where the interaction information includes the information provided by the task environment to each agent and the actions taken by the agents accordingly The information provided by the task environment includes the type information M of each agent, the local observation at the current time step the global state s t , the feedback reward r t the global state s of the next time step t+1 the global state s of the next time step t+1 , and the perturbed observations of each agent and the perturbed joint action Step A.4: Determine whether the content volume |B| of the current buffer B satisfies the capacity condition |B| ≥ batchsize. When the capacity condition is satisfied, determine whether the current time step t satisfies the training interval condition t - t lasttrain ≥ interval. If both of the above conditions are satisfied, select interaction information with a quantity of batchsize from buffer B, update the time step t when entering step B last time lasttrain = t, and enter step B. If the above two conditions are not satisfied, repeat step A; where batchsize is the preset minimum trainable buffer capacity, interval is the preset shortest training time interval, and t lasttrain is the time step when successfully entering step B from step A last time. If it is the first time to execute step A, initialize t lasttrain = 0 and record t lasttrain .
3. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 1, characterized in that, The specific method of Step B is as follows: For the interaction information randomly sampled from the buffer, the local observation of the agent in each piece of interaction information and the perturbed observation are fed into the local utility network of the agent to execute and where represents the local utility network, and θ local represents the local utility network parameters; the local action value under normal observation and the local action value under perturbed observation are obtained. Then, the observation robustness loss value is calculated, where R o represents the observation robustness regularization constraint, which is calculated using the KL divergence, and its calculation formula is as follows: Among them, P and Q are two distributions of the variable x, and p(x) and q(x) represent the probabilities of the variable x in the distributions P and Q, respectively; Observed robust loss value loss o Calculate the local utility network parameter θ local of the gradient, and its formula is as follows: Update the local utility network parameter θ by using the stochastic gradient descent method local The update formula is as follows: where α is a preset learning rate, is the gradient of the local utility network parameter θ local .
4. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 3, characterized in that The local utility network described in step B is a hyper-policy network structure, including a permutation-invariant input layer layer PI , a decision network processing layer layer policy , a fully connected layer layer FC and a permutation-equivariant output layer layer PE . The specific steps for the local utility network to calculate the local action value q based on the local observation o are as follows: Step B.1: Input the local observation o into the local utility network into the permutation-invariant input layer layer PI In the permutation-invariant input layer layer PI for each entity feature component o[j] of the local observation o, execute the following formula: W in [j] = hp in (o[j]; θ in ) Among them, hp in represents an input-invariant hypernetwork, the input-invariant hypernetwork hp in is composed of a multi-layer perceptron MLP, and its network parameters are θ in , belonging to the local utility network parameters θ local ; W in [j] represents the weight matrix generated by the input-invariant hypernetwork hp in for each entity feature component o[j] of the local observation o; where 1 ≤ j ≤ m o , m o represents the number of entities to which the entity features included in the content of the local observation o belong; Based on the weight matrix W in [j], execute the following formula to obtain the output z of the permutation-invariant input layer layer PI : Among them, is a new local observation generated by any permutation and combination with each entity feature component o[j] of the local observation o as the basic unit Step B.2: Input the output z of the permutation-invariant input layer layer PI into the decision network processing layer layer policy , and perform the following formula to obtain the output h of the decision network processing layer layer policy : h = layer policy (z; θ policy ) Among them, the decision network processing layer layer policy is composed of a recurrent neural network RNN, and its network parameters are θ policy , belonging to the local utility network parameter θ local ; Step B.3: Input the output h of the decision network processing layer layer policy into the permutation-equivariant output layer layer PE and the fully-connected layer layer PE respectively. The permutation-equivariant output layer layer PE outputs the permutation-equivariant action value q equiv , and the fully-connected layer layer FC outputs the permutation-invariant action value q inv . Concatenate q equiv and q inv to generate the local action value q, which serves as the output of the local utility network of the agent .
5. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 4, characterized in that, The permutation-equivariant output layer layer in step B.3 PE Output the permutation-equivariant action value q equiv The process is as follows: In the permutation-equivariant output layer layer PE the following equation is executed: W out [j] = hp out (o[j]; θ out ) Among them, hp out represents an output-invariant hypernetwork, and the output-invariant hypernetwork hp out is composed of a multi-layer perceptron MLP, and its network parameters are θ out , belonging to the local utility network parameter θ local , W out [j] represents the weight matrix generated by the output-invariant hypernetwork hp out for each entity feature component o[j] of the local observation o; where 1 ≤ j ≤ m o , m o represents the number of entities to which the entity features included in the content of the local observation o belong; For W out [j] Execute the following formula to obtain the permutation-equivariant action value q equiv for each action value part q equiv [j]: where h represents the output of the processing layer layer of the decision network policy of; Concatenate all action value parts q equiv [j] to obtain the permutation-equivariant action value q equiv ; The fully connected layer layer in step B.3 FC Outputs the permutation-invariant action value q inv The process is as follows: In the fully connected layer layer FC Execute the following formula to obtain the permutation-invariant action value q inv : q inv = layer FC (h; θ FC ) where, θ FC is the fully connected layer network parameter, belonging to the local utility network parameter θ local , and h represents the output of the processing layer layer policy of the decision network.
6. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 1, characterized in that The specific steps of Step C are as follows: Step C.1: Classify the agents according to the agent type information M provided by the current task environment. For each type to which the agents belong, construct a type-based hybrid network where represents the type-based hybrid networks of the 1st to mth types, and m is the total number of agent types; Step C.2: From the perspective of each agent, according to the respective type of each agent, the local observation of agent i global state s t and the local action value corresponding to the pre-disturbance joint action under normal observation and the local action value corresponding to the post-disturbance joint action are input into the type-based hybrid network of the type type i of agent i, and the following formula is executed: Among them, represents the action value of the type to which agent i belongs before perturbation under normal observation, represents the action value of the type to which agent i belongs after perturbation, represents the network parameters of the type-based hybrid network of the type to which agent i belongs; represents the type-based hybrid network of the type to which agent i belongs, represents the local action value corresponding to the joint action before perturbation of the type to which agent i belongs under normal observation, represents the local action value corresponding to the joint action after perturbation of the type to which agent i belongs, represents the local observation of the type to which agent i belongs; Step C.3: From the perspective of agent type, obtain the type-based action values of the joint actions before perturbation under normal observations according to the agent type and the type-based action values of the joint actions after perturbation Specifically, as shown in the following formula: Among them, represents the action value of the joint action of type j before perturbation under normal observation, represents the action value of the joint action of type j after perturbation, and m is the total number of agent types; The expression is as follows: Among them, represents the local action value corresponding to the joint action of all agents of type j before perturbation under normal observation, represents the local action value corresponding to the joint action of all agents of type j after perturbation, represents the type-based hybrid network of type j, θ typej represents the network parameters of the type-based hybrid network of type j, θ type represents the set of network parameters θ that includes all type-based hybrid networks type ={θ type1 ,...,θ typem}, {o i} typej is the local observation of the agent of type j.
7. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 6, characterized in that The local action value {q i} typej , the global state s, and the local observation {o i} typej are input into the type-based hybrid network to obtain the type-based action value Q belonging to type j typej The specific process is as follows: In a type-based hybrid network the local action value {q i} typej the global state s, and the local observation {o i} typej will first be input into a type attention network and perform the following formula: Among them, is a type-based hybrid network belonging to type j in the type attention network, whose network parameters are is the type-based local action value belonging to type j; In the type attention network for the global state s and the local observation {o i} typej perform the following formula to obtain the query content Q and the key content K: Q = s × W q K = o × W k where s represents the global state, o represents the local observation, and W q and W k are attention weight matrices; Based on the query content Q and the key content K, for the local action value {q i} typej Execute the following formula to obtain the type-based local action value belonging to type j where d k is the second dimension size of the query content Q and the key content K matrix; For execute the following formula to obtain the type-based action value Q of type j typej : where, W typej and b typej represent the weight matrix and the bias vector in the weighted bias process of the type-based hybrid network of type j, respectively, and Q typej is the type-based action value of type j; Weight matrix W typej and bias vector b typej are respectively generated by two hypernetworks hp wtypej and hp btypej . Executing W typej = hp wtypej (s; θ wtypej ) and b typej = hp btypej (s; θ btypej ). The inputs of the two hypernetworks are both the global state s, and the structures are both composed of a multi-layer perceptron MLP. Among them, θ wtypej and θ btypej are respectively the network parameters of hypernetworks hp wtypej and hp btypej . The network parameter θ typej of the type-based hybrid network belonging to type j contains the network parameters of the type attention network in the type-based hybrid network belonging to type j as the two hypernetwork parameters θ wtypej and θ btypej in the type-based hybrid network belonging to type j.
8. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 7, characterized in that The specific steps of Step D are as follows: Step D.1: Take the global state s t and the type-based action values of the joint actions before and after the perturbation under all normal observations and and feed them into the global mixing network respectively to execute the following formula: Among them, is the global action value of the combined action before perturbation, is the global action value of the combined action after perturbation; θ total is the global hybrid network of the network parameters; Step D.2: Based on the feedback reward r given by the task environment at the same time step taken from buffer B t , combined with the global action value of the joint action before the disturbance and the global action value of the joint action after disturbance Calculate the TD loss value according to the TD loss formula and the joint action robust regularization constraint formula and joint action Lupin loss value The joint action robust regularization constraint formula R a () is calculated using the root mean square error RMSE, as follows: where y and y ′ represent the true vector and the predicted vector respectively, ‖‖2 represents the L2 norm of the vector, and n is the dimension size of the vector; The TD loss value is specifically as follows: where γ is the time discount factor, is the globally mixed network that is periodically replicated with network parameters θ total , is the type-based action value at time step t + 1; Step D.3: Calculate the global loss value loss according to the following formula Total :[[]]END]] loss Total = loss TD + loss a Using the global loss value loss total to calculate the global loss gradient Among them, θ Total is the global network parameter, including the network parameter θ of the local utility network local , the network parameter θ of all types of type-based hybrid networks type and the network parameter θ of the global hybrid network total ; When calculating the global loss gradient After completion, according to the preset learning rate, the global network parameters θ Total are updated by using stochastic gradient descent. The stochastic gradient descent formula is as follows: Among them, α is a preset learning rate.
9. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 8, characterized in that Global mixing network in step D.1 For the input type-based action value Q type Execute the following formula: Q total = [Q type × W 1total + b 1total × W 2total + b 2total Among them, Q total represents the global action value, W 1total and W 2total are the weight matrices in the first weighted bias and the second weighted bias operations of the global hybrid network respectively, and b 1total and b 2total are the bias vectors in the first weighted bias and the second weighted bias operations respectively; Each weight matrix and bias vector are generated by four different hyper-networks, performing W 1total = hp w1total (s; θ w1total ) and W 2total = hp w2total (s; θ w2total ) respectively obtain the weight matrices W 1total and W 2total in the first and second weighted bias operations. Performing b 1total = hp b1total (s; θ b1total ) and b 2total = hp b2total (s; θ b2total ) respectively obtain the bias vectors b 1total and b 2total in the first and second weighted bias operations. The inputs of the four hyper-networks are all the global state s, and the structures are all composed of a multi-layer perceptron MLP. Among them, hp w1total and hp w2total are the hyper-networks for generating the weight matrices in the first and second weighted bias operations respectively. θ w1total and θ w2total are the network parameters of the hyper-networks hp w1total and hp w2total respectively. hp b1total and hp b2total are the hyper-networks for generating the weight matrices in the first and second weighted bias operations respectively. θ b1total and θ b2total are the network parameters of the hyper-networks hp b1total and hp b2total respectively; the network parameters θ of the global mixing network total include the network parameters θ of all the hyper-networks for generating the weight matrices and bias vectors in the first and second weighted bias operations in the global mixing network w1total and θ w2total , as well as θ b1total and θ b2total ; After updating the global network parameters θ Total it is determined whether the current time step t satisfies the training termination condition t≥T, where T is the preset maximum training time step. If the condition is satisfied, step E is entered; otherwise, step A is returned for execution.
10. A robust reinforcement learning method for heterogeneous multi-agent systems according to claim 9, characterized in that In step E, all the trained agents are deployed in the task environment, and the local utility networks of all the agents will execute after obtaining the local observation o i given by the task environment and select the action a with the maximum value from the local action value q i and execute it until the task completion degree reaches the required expectation. i
Citation Information
Cited By
Heterogeneous multi-agent cooperation method based on hierarchical value decomposition of super network
CN120874951A