A Human-Guided Multi-Agent Deep Reinforcement Learning Method
By constructing a multi-agent value decomposition framework and a human real-time intervention mechanism, combining hybrid experience replay and adaptive entropy regularization, the problem of inefficient credit allocation and exploration in multi-agent collaboration tasks is solved, and efficient and secure multi-agent collaboration training is achieved.
Patent Information
- Application Number
- CN202510466976.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing multi-agent reinforcement learning methods have shortcomings in credit allocation, distribution offset, exploration efficiency and human intervention mechanisms, and it is difficult to effectively solve the strategy optimization problems in multi-agent collaboration tasks.
A multi-agent value decomposition framework is constructed, combining human real-time intervention and mixed experience replay, and through nonlinear monotonic hybrid networks and adaptive entropy regularization, the policy network is optimized to achieve efficient and secure multi-agent collaborative training.
Provide high-quality samples through human intervention, combining priority mixed sampling and entropy adjustment, significantly accelerate strategy convergence, improve training security and adaptability, solve the problem of inefficient credit allocation and exploration, and ensure global optimal strategy generation.
Smart Images

Figure CN119990246B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of reinforcement learning training strategies, and particularly relates to a multi-agent deep reinforcement learning method based on human guidance. Background Art
[0002] Multi-agent reinforcement learning can solve complex tasks (such as cooperative hunting and joint path planning) that cannot be handled by single agents through the interaction and cooperation of multiple agents in a shared environment. However, the existing methods still face the following key problems:
[0003] Credit assignment problem: In fully cooperative tasks, it is difficult to fairly decompose the global reward into individual contributions. Traditional methods such as VDN decompose the Q-value through linear summation, but its expressive ability is limited; although QMIX introduces a non-linear monotonic mixing network, it still cannot handle non-monotonic dependencies, resulting in sub-optimal policies;
[0004] Distribution shift and insufficient generalization: The dynamic differences between the training environment and the real environment (such as sensor noise and sudden obstacles) can easily lead to policy failure, and existing methods (such as MADDPG and COMA) lack a robust design for distribution shift;
[0005] Low exploration efficiency: Traditional MARL relies on random exploration, with low sampling efficiency in high-dimensional state-action spaces and is prone to falling into local optima.
[0006] Existing human-in-the-loop reinforcement learning methods optimize policies through human demonstrations, interventions, or evaluations, but have the following deficiencies in multi-agent scenarios:
[0007] Single interaction mechanism: Existing methods (such as TAMER and Deep TAMER) mostly adopt static imitation learning and cannot dynamically integrate human feedback and autonomous exploration, resulting in insufficient policy flexibility;
[0008] Insufficient multi-agent collaboration support: Human intervention usually targets a single agent and lacks overall guidance for group cooperation behaviors, making it difficult to solve the credit assignment problem.
[0009] Summary of the comparison and deficiencies of existing technologies:
[0010] QMIX / VDN: Although it solves some credit assignment problems, it relies on environmental reward signals and cannot utilize human prior knowledge to accelerate training;
[0011] MADDPG: A simple centralized training and decentralized execution (CTDE) framework has poor adaptability to non-stationary environments;
[0012] COMA: The calculation of the counterfactual baseline is complex and difficult to scale to large-scale agent systems;
[0013] MASAC: In multi-agent reinforcement learning, the state space to be explored is large. If the exploration range is too large, the exploration efficiency will be low.
[0014] Traditional HITL method: It lacks a dynamic intervention mechanism for multi-agent collaboration and has a low sample utilization rate. Summary of the Invention
[0015] In view of the above problems, the present invention provides a human-guided multi-agent deep reinforcement learning method, including the following processes:
[0016] S1. Construct a multi-agent value decomposition framework, initialize the multi-agent environment and the state of each agent. At the same time, initialize the value function decomposition network, obtain the global action value function through the value function network and the policy network, and synthesize the global value through the non-linear monotonic mixing network to ensure that the individual-global maximization principle is satisfied;
[0017] S2. Human real-time monitoring and intervention; humans monitor the local observations, action selections, and environmental states of agents in real time through a visual interaction interface; when humans believe that the action efficiency of an agent is lower than a preset value or there is a risk, they input alternative actions through the interaction interface to overwrite the original actions of the agent and store the intervention experience in an independent replay pool;
[0018] S3. Store the high-quality experience samples generated by human intervention in an independent replay pool and mix and sample them with the samples generated by the agents' autonomous exploration according to priorities; the priorities are dynamically allocated by the temporal difference error;
[0019] S4. Maximize the entropy in the policy optimization objective to encourage exploration, and dynamically adjust the entropy coefficient through the target entropy; update the policy network by optimizing the loss function to ensure that the agents find a balance between exploration and exploitation;
[0020] S5. Recursively execute the steps of the interaction between the agents and the environment, experience saving, value network, and policy network update until the preset termination condition is met; finally, obtain the greedy policy of the agents according to the multi-agent soft actor-critic architecture combined with the value function decomposition technology as the output policy.
[0021] Preferably, the global action value function in S1 is synthesized through a non-linear monotonic mixing network, and the specific form is:
[0022] ;
[0023] Or expressed as:
[0024] ;
[0025] represents the global action value function is a hybrid network, is the global state, is the local observation history, is the joint action, is the hybrid weight matrix, is the bias vector, dynamically generated by the hypernetwork according to the global state; represents the agent 's local value function, is any number from 1 to N, with the input being the local observation history and the action .
[0026] Preferably, the hypernetwork is composed of a three-layer fully connected neural network, with the input being the global state , and the output layer uses the Softplus activation function to ensure the non-negativity of the hybrid weights; its calculation process is:
[0027] ;
[0028] ;
[0029] FC 1, FC 2, FC 3 are linear layers, ReLU is the activation function, and the Softplus layer is used to ensure non-negative weights.
[0030] Preferably, the human intervention mechanism in S2 includes:
[0031] Humans input alternative actions through the interaction interface , overriding the agent's original action , generating intervention experience and storing it in an independent replay pool ;
[0032] The policy update loss function introduces an imitation learning term, in the form of:
[0033] ;
[0034] Among them, is the entropy coefficient, is the imitation learning weight, is the action distribution output by the policy network, is the joint action value, is the current state, is the agent's current local observation, is the agent's action, is the human action, , To calculate the mathematical expectation of the content within [ ] within the experience pool D , range.
[0035] Preferably, during the processes of S2 and S3, the agent interacts with the environment using an exploration strategy, and humans observe and are ready to intervene at any time to replace the agent's actions and save the experience. Specifically:
[0036] S21: Calculate the value function of all actions at the current moment;
[0037] Each agent uses the current observation and the action at the previous moment as inputs and feeds them into its value function network for calculation to obtain the value functions of all possible actions at the current moment;
[0038] The calculation method of the value function network is as follows: ; where is the hidden state obtained through the RNN, representing the state information at the current moment;
[0039] S22: Exploration and exploitation decision-making;
[0040] Based on the current value function, the agent adopts -greedy strategy to select an action: with probability, randomly select an action, probability to select the action with the maximum value function. As the training progresses, will gradually decrease;
[0041] S23: Save the experience;
[0042] The agent stores the current observation, the selected action, the reward, and the new observation in the experience buffer pool, and this process continues until the experience buffer pool reaches the specified capacity.
[0043] Preferably, during the processes of S2 and S3, the training of the agent network is performed, including the following processes:
[0044] S31: Sample from the experience buffer pool;
[0045] During each training, sample N pieces of experience data from the experience buffer pool. Each piece of data contains the current observation, action, return, and the next observation, and then use these data to calculate the action value function of each agent;
[0046] S32: Calculate the global state encoding;
[0047] Perform one-hot encoding on the state of each agent through the global state information to ensure the effective integration of the global information shared by multiple agents;
[0048] S33: Use the hypernetwork to calculate the value decomposition weights and biases;
[0049] Use the QMIX network to perform weighted synthesis on the local value functions of each agent, and calculate the global joint action value function:
[0050] ;
[0051] Among them, is the joint action value function, is the local value function of agent , is the current state, is the agent action, is any number from 1 to N, is the weight matrix under the global state , is the bias term;
[0052] Step 3-4: Calculate and optimize the loss of the joint action value function;
[0053] Calculate and optimize the loss of the joint action value function, and use the standard mean squared error MSE loss function:
[0054] ;
[0055] Among them, is the expected return.
[0056] Preferably, the sampling method in the hybrid experience replay pool in S3 is priority sampling from the storing human experience and the storing non-human experience in the experience replay pool.
[0057] Preferably, the entropy coefficient in S4 is dynamically updated through the optimization objective, and the loss function is:
[0058] ;
[0059] Among them, is the entropy coefficient, is the current policy, is the target entropy, is the loss function, is the mathematical expectation operator, is the mathematical operator calculated by exponential moving average.
[0060] Preferably, the policy network update in S5 adopts the double Q network mechanism, and the update target is:
[0061] ;
[0062] Among them, the target value is calculated as:
[0063] ;
[0064] is the discount factor, is the output of the target network, and the target network parameters are kept stable through soft update where is the soft update weight coefficient, is the current state, is the local observation of the current agent, is the agent action, is the human intervention action, is the update target, is the mathematical expectation operator.
[0065] Preferably, both the local value function network and the policy network include a recurrent neural network (RNN) layer for encoding the historical observation sequence.
[0066] Preferably, the imitation learning weight is a preset hyperparameter and can be dynamically adjusted through a hidden layer neural network.
[0067] Preferably, the target network parameter update rate is a preset hyperparameter.
[0068] Compared with the prior art, the present invention has the following beneficial effects:
[0069] By integrating human active intervention and multi-agent value decomposition technology, combining hybrid experience replay, adaptive entropy regularization, and elastic policy update mechanism, the present invention realizes efficient and safe training of multi-agent collaborative tasks. Specifically, it is reflected in:
[0070] Efficient training and policy optimization: By providing high-quality demonstration samples through human intervention and combining priority hybrid sampling, the policy convergence is significantly accelerated, and the ineffective exploration time is reduced;
[0071] Improved safety and robustness: The human real-time monitoring and action coverage mechanism can avoid high-risk behaviors and ensure the safety of the training process; the value decomposition framework enhances the adaptability to dynamic changes in the environment and reduces the policy failure caused by distribution shift;
[0072] Optimized credit assignment: Based on the value decomposition architecture of the non-linear monotonic hybrid network, while satisfying the individual-global maximization principle (IGM), the present invention solves the credit assignment problem in multi-agent collaboration and ensures the generation of the global optimal policy;
[0073] Exploration-Exploitation Balance: By maximizing the policy entropy and dynamically adjusting the entropy coefficient, the exploration ability of the agent is enhanced, avoiding being trapped in local optima while maintaining the stability of the policy.
[0074] Multi-Agent Value Decomposition Framework Combined with Human Active Intervention:
[0075] Nonlinear Hybrid Network: Dynamically generate the weight matrix and bias vector through a hypernetwork to synthesize the global action value function, solving the problem of insufficient expressive power of traditional linear decomposition (such as VDN) and supporting the modeling of non-monotonic dependencies in complex tasks;
[0076] Independent Generation of Local Value Functions: Each agent independently generates local value functions, and combines with the hybrid network to synthesize the global value, ensuring the fairness of credit assignment and the interpretability of the policy.
[0077] Human Real-Time Intervention Mechanism:
[0078] High-Quality Experience Injection: Humans directly correct inefficient or dangerous behaviors through substitute actions, and the generated high-quality samples are stored in an independent replay pool, providing reliable prior knowledge for policy optimization;
[0079] Introduction of Imitation Learning Term: Add an imitation learning term of human actions to the policy loss function, forcing the policy network to approximate human demonstrations, accelerating policy optimization and improving action safety;
[0080] Hybrid Prioritized Experience Replay:
[0081] Dynamic Priority Sampling: Allocate sample priorities based on the temporal difference error (TD-error), and preferentially use high-value human intervention samples and self-exploration samples for mixed training, improving sample utilization and learning efficiency;
[0082] Independent Replay Pool Design: Separate human experience and non-human experience, avoiding sample distribution conflicts while retaining the diversity of self-exploration;
[0083] Soft Update of Parameters: Update the target network parameters through exponential moving average, avoiding policy mutations and ensuring the smoothness of the training process.
[0084] Hypernetwork Structure Optimization:
[0085] Non-Negative Weight Constraint: Ensure the non-negativity of the mixed weights through the Softplus activation function, meeting the monotonicity requirements of the IGM principle while supporting flexible global value synthesis. Description of the Drawings
[0086] Figure 1 It is a schematic flowchart of the method of the present invention.
[0087] Figure 2It is a schematic diagram of the interaction between modules in the algorithm of the present invention.
[0088] Figure 3 It is the specific update process of the algorithm of the present invention.
[0089] Figure 4 It is a schematic diagram of the network architecture in the present invention. Specific implementation manner
[0090] The present invention provides a multi-agent value decomposition deep reinforcement learning method based on human guidance. The overall process is as Figure 1 shown: It includes the following steps:
[0091] Step 1: Construct a multi-agent value decomposition framework;
[0092] Initialize the environment and agent states: Initialize the multi-agent environment and the state of each agent, including the local observation, action selection of the agent, and the global state of the environment; at the same time, initialize the value function decomposition network to generate the local values of each agent, and synthesize the global value through a non-linear monotonic mixing network to ensure that the individual-global maximization principle (IGM) is satisfied;
[0093] Step 2: Flexible policy update by integrating human intervention;
[0094] Human real-time monitoring and intervention: Humans monitor the local observations, action selections, and environmental states of agents in real time through a visual interaction interface; when humans believe that the action efficiency of the agent is low or there are risks, they input alternative actions through the interaction interface to overwrite the original actions of the agent, and store the intervention experience in an independent replay pool;
[0095] Step 3: Hybrid experience replay pool;
[0096] Experience sample storage and sampling: Store the high-quality experience samples generated by human intervention in an independent replay pool, and sample them by mixing with the samples generated by the agent's autonomous exploration according to priorities; the priorities are dynamically allocated by the temporal difference error (TD-error) to ensure that high-quality samples are fully sampled during training;
[0097] Step 4: Adaptive entropy regularization constraint;
[0098] Policy optimization and entropy adjustment: Maximize the entropy in the policy optimization objective to encourage exploration, and dynamically adjust the entropy coefficient through the target entropy; update the policy network by optimizing the loss function to ensure that the agent finds a balance between exploration and exploitation;
[0099] Step 5: Loop through training and evaluation
[0100] Iterative training and evaluation: The steps of interaction between the agent and the environment, experience preservation, and policy update are executed in a loop until a preset termination condition is met; finally, the greedy policy of the agent is obtained according to the agent value function as the output policy to complete the multi-agent reinforcement learning training.
[0101] The following further illustrates the present invention in conjunction with embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0102] Embodiment 1.
[0103] Step 1: Initialize the environment and the states of all agents;
[0104] Initialize the agent network:
[0105] Before training, the weights and biases of all agent value function networks and target networks are initialized;
[0106] Initialize the environment:
[0107] Initialize the environment of the task. For example, assume the scenario is a path planning task where multiple agents cooperate. The environmental state includes the current position information, target position, etc. of each agent;
[0108] Initialize the experience cache pool:
[0109] Set up the experience cache pool for storing the interaction experiences of each agent, including observations, actions, rewards, and new observations. The size of the pool is limited, and when it is full, the earliest experience will be discarded.
[0110] Step 2: The agent uses the exploration policy to interact with the environment and save the experience; during this process, humans can observe the agent's actions and the global state in real time through the visualization interface. If it is found that the agent has inappropriate behaviors (such as dangerous exploration behaviors or inefficient exploration behaviors), humans can directly interrupt the interaction between the agent and the environment through the interaction interface and input human intervention actions to replace the agent's actions, guide the agent's next behavior, and generate human intervention experiences (the current agent's local observation, global state, human intervention action);
[0111] Step 2-1: Calculate the value functions of all actions at the current moment;
[0112] Each agent uses the current observation and the action at the previous moment as inputs and sends them into its value function network for calculation to obtain the value functions of all possible actions at the current moment;
[0113] The calculation method of the value function network is as follows: Among them, is the hidden state obtained through the RNN, representing the state information at the current moment;
[0114] Step 2-2: Exploration and exploitation decision-making;
[0115] Based on the current value function, the agent adopts - The greedy policy to select an action: with probability, randomly select an action, probability to select the action with the maximum value function. As the training progresses, will gradually decrease; during this process, humans can observe the agent's actions and the global state in real time through the visualization interface. If it is found that the agent has dangerous or inefficient exploration behaviors, humans can interrupt the training through the interaction interface and input human intervention actions to replace the current agent's actions. For example, in a path planning task, agent A selects a circuitous path (30 meters away from the target and there are dynamic obstacles). The human operator judges that this path takes a long time and is risky, and inputs the alternative action "accelerate and move to the right" through the interaction interface, directly overriding the original action "go straight" of agent A;
[0116] Step 2-3: Save experience;
[0117] The agent stores the current observation, the selected action, the reward, and the new observation in the experience buffer pool. If there is human intervention in Step 2-2, the selected action is changed to the human intervention action. This process continues until the experience buffer pool reaches the specified capacity.
[0118] Step 3: Perform the training of the agent network;
[0119] Step 3-1: Sample from the experience buffer pool
[0120] Each time during training, sample N pieces of experience data from the experience buffer pool (each piece of data contains the current observation, action, return, and the next observation). Then use these data to calculate the action value function of each agent;
[0121] Step 3-2: Calculate the global state encoding;
[0122] Perform one-hot encoding on the state of each agent through the global state information to ensure the effective integration of the global information shared by multiple agents;
[0123] Step 3-3: Use the hypernetwork to calculate the value decomposition weights and biases;
[0124] Use the QMIX network to perform weighted synthesis on the local value functions of each agent to calculate the global joint action value function:
[0125] ;
[0126] Among them, is the joint action value function, is the local value function of the agent , is the global state under the weight matrix, is the bias term;
[0127] Step 3-4: Calculate and optimize the loss of the joint action value function;
[0128] Calculate and optimize the loss of the joint action value function, using the standard mean squared error (MSE) loss function:
[0129] ;
[0130] Among them, is the expected return.
[0131] Step 4: Loop through Steps 2 and 3 until the training terminates;
[0132] Steps 2 and 3 will be continuously looped until the training termination condition is reached, such as convergence or the maximum number of training epochs.
[0133] In this example, the value function decomposition method is for each independent agent to learn the local value function, and then combine these local values with a learnable hybrid neural network to generate the joint value , In this algorithm, the value function decomposition network in QMIX is used. It is a non-linear monotonic decomposition structure that can realize a richer function class while satisfying the Individual-Global Maximization (IGM) principle: the result of performing the global argmax operation on is the same as the result of performing the argmax operation on each local separately. The value network update method is similar to the Critic update method used in the traditional SAC algorithm, but now the Critic network consists of "local network + hybrid network ", where the network architecture of the hybrid network is modified from the network in the QMIX paper. It is composed of the local network of the agent and the hybrid network is composed, where the weights and biases of the hybrid network are generated by independent hyper-networks. Figure 4 shows the detailed structure of the value decomposition hybrid network. For each agent , there is a local network to represent its local value function . The hybrid network linearly mixes the outputs of the agent's local networks as inputs, and then generates values through an absolute value activation function.
[0134] During the update process of the policy network, two soft value networks are utilized. For j ∈ {1, 2}, the minimum value is taken as the target:
[0135] ;
[0136] Meanwhile, it can also be expressed in the following form:
[0137] ;
[0138] is the experience replay pool, which contains the previous sampled state transition processes (state s , local observation , action , reward , next state , next local observation ), is the target network, and its parameters are obtained through the exponentially moving average of the current network weights , which has been proven in the original SAC paper to be able to stabilize the training process.
[0139] The structure of the policy network is the same as the local network of the agent , but a softmax layer is added to output the probability distribution.
[0140] Note that in the formula, is from the current policy i of the agent Sampled, rather than sampled from the replay buffer. Compared with the original QMIX algorithm, an additional policy network is introduced, which outputs a probabilistic policy (probability mass function in the discrete domain), accurately representing the probability value of each agent choosing each discrete action. Therefore, the expected value can be accurately calculated. Recent research has theoretically proven that soft (or Boltzmann) policy iteration can guarantee improvement and convergence to the optimal policy. Based on the soft policy iteration process, the goal of policy update is as follows:
[0141] ;
[0142] Meanwhile, it can also be expressed in the following formula form:
[0143] ;
[0144] is a hyperparameter that controls the trade-off between maximizing the policy entropy and the expected discounted return, and is updated in the same way as in the original SAC paper. The overall value decomposition network architecture of the present invention is as Figure 4 shown.
[0145] Update method with human-in-the-loop:
[0146] As Figure 3 shown, in order to incorporate human intervention samples, an imitation learning loss term is added to encourage the policy to be close to the guiding behavior of humans. Let the human intervention sample be , the policy update becomes:
[0147] ;
[0148] where is the standard sample in the experience replay pool, is the sample generated by human intervention, and 𝛽 is a hyperparameter that controls the weight of the imitation learning loss term.
[0149] In the above way, human intervention directly affects the policy network update through imitation learning, driving the agent's policy closer to the high-quality behavior provided by humans.
[0150] After adding human intervention samples, human intervention affects the generation method of the target value. For the intervention samples provided by humans, assuming that the actions of humans are considered to be of higher quality, the value network update method becomes:
[0151] ;
[0152] where:
[0153] ;
[0154] After introducing human intervention samples, the joint Value Calculated from human actions Without using the action distribution of the policy network, it becomes:
[0155]
[0156] Example 2
[0157] The value decomposition multi-agent reinforcement learning training method using human guidance in the present invention can be applied to multiple fields such as robot control, traffic coordination, and manufacturing control.
[0158] The multi-agent reinforcement learning training method described in the present invention is used for the cooperative control of a multi-robot system. Among them, each agent represents a robot. Each robot includes multiple modules, such as a movement module, a grasping module, a perception module, a navigation module, etc. The policy of the agent is the currently executed action. Among them, the movement module controls the movement direction and speed of the robot, the grasping module controls the actions of grasping and placing objects, the perception module is used to process and understand environmental information, and the navigation module is responsible for planning the travel path of the robot. All these actions are controlled by the agent value function network. The reward function represents the efficiency of multiple robots in collaborating to complete tasks, such as task completion time, path length, cooperation effect, etc. Each robot can only perceive and control its own working state and cannot directly observe the internal states and behaviors of other robots. The global state of the environment contains the state information of each robot (such as position, task progress, power situation, etc.) and the cooperation state of the entire system. The robots coordinate and cooperate by observing local environmental information to achieve a common goal.
[0159] This embodiment relates to a multi-robot cooperation task. In this scenario, each robot is regarded as an agent and collaborates to complete a common task.
[0160] Step 1: Initialize the entire multi-robot system and the state of each robot;
[0161] 1.1 Initialize the agent network: Each robot (agent) initializes its value function network, which has the same structure as in the single-agent scenario and consists of three layers: a linear layer (MLP), a recurrent neural network (RNN) layer, and a linear output layer;
[0162] 1.2 Initialize the environment: The state of the environment includes the current positions, task objectives, and specific requirements of all robots. The state of each robot contains its local information and collaborates through sharing the global state.
[0163] Step 2: Each robot interacts with the environment using an exploration strategy and saves the experience in the experience cache pool; during this process, humans can observe the agent's actions and the global state in real time through the visualization interface. If it is found that the agent has improper behavior (such as dangerous exploration behavior or inefficient exploration behavior), humans can directly interrupt the agent's interaction with the environment through the controller and input human intervention actions to replace the agent's actions (current movement direction, speed, grasping data), guide the agent's next behavior, and generate human intervention experience (current local observation of the agent, global state, human intervention actions). For example, if Agent B (responsible for grasping) and Agent C (responsible for transportation) may collide due to path conflicts, humans can observe through the visualization interface that their predicted paths overlap and immediately intervene:
[0164] 1. Input the alternative action "pause grasping" for Agent B;
[0165] 2. Input the alternative action "turn left to avoid" for Agent C;
[0166] 2.1 Step 2-1: Each robot calculates the action value function;
[0167] Each robot calculates the action value function through the value function network based on the current observation (such as position, task status, etc.) and the action at the previous moment;
[0168] 2.2 Step 2-2: Exploration and exploitation decision-making;
[0169] Each robot decides whether to execute a random action or an optimal action based on the current value function. If the executed action is dangerous or inefficient, the human observer can directly interrupt the action execution and input human intervention actions through the controller to replace the agent's current action;
[0170] 2.3 Step 2-3: Save the experience;
[0171] Each robot saves the current state, action, reward, and new state in the experience pool. If there is human intervention in Step 2-2, the agent's action in the saved experience is the human intervention action.
[0172] Step 3: Execute the training of the agent network;
[0173] 3.1 Step 3-1: Sample from the experience cache pool;
[0174] Sample N pieces of experience data from the experience cache pool, calculate the action value function of each robot, and update its policy;
[0175] 3.2 Step 3-2: Calculate the global state encoding;
[0176] Encode the state of each robot using the global state to ensure information sharing in the multi-robot system;
[0177] 3.3 Step 3-3: Use the QMIX network for value decomposition;
[0178] Use the QMIX network to calculate the joint value function of all robots for effective credit assignment;
[0179] 3.4 Step 3-4: Calculation of the joint action value function;
[0180] Calculate the joint action value function through the decomposition structure and update the policy;
[0181] 3.5 Step 3-5: Optimize the target network;
[0182] Use the target network to calculate the loss of the joint action value function and optimize it, updating the networks of all robots.
[0183] Step 4: Loop through Step 2 and Step 3 until the training terminates
[0184] Steps 2 and 3 are looped until the termination condition is reached.
[0185] Value decomposition network architecture:
[0186] Local Network: Each agent independently trains the local function , with the input being the local observation history , and the output being the action value;
[0187] Hypernetwork design:
[0188] Input: Global state s;
[0189] Structure: A 3-layer fully connected network with the ReLU activation function for the hidden layer;
[0190] Output: Mixing weights and biases , guaranteed by the softplus activation function ;
[0191] Global Value synthesis:
[0192] Determine the global value :
[0193] .
[0194] Human active intervention mechanism:
[0195] Intervention process:
[0196] 1. Humans can view the local observations of each agent, , action probability distribution , and the global environmental state s in real time through the visualization interface;
[0197] 2. When it is observed that the agent's action is inefficient (such as a circuitous path) or there is a risk, humans intervene in the agent's action through the interaction interface and input an alternative action ;
[0198] 3. The system immediately executes , and transfers the states before and after the intervention and stores them .
[0199] Experience priority:
[0200] Human intervention samples are defaultly given the highest priority in priority sampling to ensure sufficient sampling during training.
[0201] Policy optimization:
[0202] Policy loss function:
[0203] ;
[0204] Where:
[0205] SAC target;
[0206] is the imitation learning term;
[0207] Critic double network and soft update:
[0208] Two soft value networks are adopted, and the minimum value is taken as the target. For j ∈{1,2}:
[0209] ;
[0210] Soft update mechanism: The parameters of the target network , to improve the training stability, where is a preset hyperparameter, such as .
[0211] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
[0212] Although the specific implementation manners of the present invention have been described above, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A multi-agent deep reinforcement learning method based on human guidance, characterized in that It includes the following processes: S1. Construct a multi-agent value decomposition framework, initialize the multi-agent environment and the state of each agent. Meanwhile, initialize the value function decomposition network, obtain the global action value function through the value function network and the policy network, and synthesize the global value through the non-linear monotonic mixing network to ensure compliance with the individual-global maximization principle; S2. Human real-time monitoring and intervention; humans real-time monitor the local observations, action selections, and environmental states of the agents through the visual interaction interface; when humans believe that the action efficiency of the agent is lower than the preset value or there are risks, they input alternative actions through the interaction interface to overwrite the original actions of the agent and store the intervention experience in an independent replay pool; S3. Store the high-quality experience samples generated by human intervention in the independent replay pool and mix and sample them with the samples generated by the agent's autonomous exploration according to the priority; the priority is dynamically allocated by the temporal difference error; S4. Maximize the entropy in the policy optimization objective to encourage exploration and dynamically adjust the entropy coefficient through the target entropy; update the policy network by optimizing the loss function to ensure that the agent finds a balance between exploration and exploitation; S5. Recursively execute the steps of the interaction between the agent and the environment, experience preservation, value network, and policy network update until the preset termination condition is met; finally, obtain the greedy policy of the agent according to the multi-agent soft actor-critic architecture combined with the value function decomposition technology as the output policy.
2. The multi-agent deep reinforcement learning method based on human guidance according to claim 1, characterized in that: In the above S1, the global action value function is synthesized through the non-linear monotonic mixing network, and the specific form is: ; Or expressed as: ; represents the global action value function, is a hybrid network, is the global state, is the local observation history, is the joint action, is the hybrid weight matrix, is the bias vector, dynamically generated by the hypernetwork according to the global state; represents the agent 's local value function, is any number from 1 to N, with the input being the local observation history and the action .
3. The multi-agent deep reinforcement learning method based on human guidance according to claim 2, wherein: The super network consists of three layers of fully connected neural networks, with the input being the global state , and the output layer ensures the non-negativity of the mixing weights through the Softplus activation function; its calculation process is as follows: ; ; FC 1, FC 2, FC 3 is a linear layer, ReLU is an activation function, and the Softplus layer is used to ensure non - negative weights.
4. A human-guided multi-agent deep reinforcement learning method according to claim 1, characterized in that: The human intervention mechanism in the above S2 includes: Humans input alternative actions through the interaction interface , overriding the original actions of the agent , generating intervention experience and storing it in an independent replay pool ; Policy update loss function Introduce an imitation learning term in the form of: ; Among them, is the entropy coefficient, is the imitation learning weight, is the action distribution output by the policy network, is the joint action value, is the current state, is the current local observation of the agent, is the agent's action, is the human action, 、 are in the experience pool 、 to calculate the mathematical expectation within the range.
5. A human-guided multi-agent deep reinforcement learning method according to claim 1, characterized in that: During the processes of S2 and S3, the agent uses the exploration strategy to interact with the environment, and humans observe and are ready to intervene at any time to replace the agent's actions and save the experience. Specifically: S21: Calculate the value functions of all actions at the current moment; Each agent uses the current observation and the action at the previous moment as inputs and sends them into its value function network for calculation to obtain the value functions of all actions at the current moment; The calculation method of the value function network is as follows: ; where is the hidden state obtained through the RNN, representing the state information at the current moment; S22: Exploration and exploitation decision-making; Based on the current value function, the agent adopts a greedy policy to select an action: with a probability of randomly select an action, and with a probability of select the action with the maximum value function. As the training progresses, will gradually decrease; S23: Save the experience; The agent stores the current observation, selected action, reward, and new observation in the experience cache pool, and this process continues until the experience cache pool reaches the specified capacity.
6. The multi-agent deep reinforcement learning method based on human guidance according to claim 1, wherein: During the processes of S2 and S3, the training of the agent network is performed, including the following processes: S31: Sample from the experience cache pool; During each training, sample N pieces of experience data from the experience cache pool. Each piece of data contains the current observation, action, return, and next observation, and then use these data to calculate the action value functions of each agent; S32: Calculate the global state encoding; Perform one-hot encoding on the state of each agent through the global state information to ensure the effective integration of the global information shared by multiple agents; S33: Use the hypernetwork to calculate the value decomposition weights and biases; Use the QMIX network to perform weighted synthesis on the local value functions of each agent to calculate the global joint action value function: ; Among them, is the joint action value function, is the local value function of the agent , is the current state, is the agent action, is any number from 1 to N, is the global state under the weight matrix, is the bias term; Step 3-4: Calculate the loss of the joint action value function and optimize it; Calculate the loss of the joint action value function and optimize it, using the standard mean squared error MSE loss function: ; Among them, is the expected return.
7. A human-guided multi-agent deep reinforcement learning method according to claim 1, characterized in that: The sampling method of the hybrid experience replay pool in S3 is to perform priority sampling from the part storing human experience and the part storing non-human experience in the experience replay pool. and the part storing non-human experience in the experience replay pool.
8. A multi-agent deep reinforcement learning method based on human guidance according to claim 1, characterized in that: The entropy coefficient in S4 is dynamically updated through an optimization objective, and the loss function is: ; Among them, is the entropy coefficient, is the current policy, is the target entropy, is the loss function, is the mathematical expectation operator, is the mathematical operator calculated by exponential moving average.
9. A human-guided multi-agent deep reinforcement learning method according to claim 1, characterized in that: The policy network update in S5 adopts a double Q-network mechanism, and the update objective is: ; Wherein, the target value is calculated as: ; is the discount factor, is the output of the target network, the target network parameters Maintain stability through soft update maintain stability, is the soft update weight coefficient, is the current state, is the local observation of the current agent, is the agent action, is the human intervention action, is the update target, is the mathematical expectation operator.
10. A multi-agent deep reinforcement learning method based on human guidance according to claim 1, characterized in that: Both the local value function network and the policy network include a recurrent neural network (RNN) layer for encoding historical observation sequences.
Citation Information
Patent Citations
Equipment optimal maintenance strategy searching method and system based on reinforcement learning
CN118941275A
Multi-agent cooperative control method and system based on deep reinforcement learning, and medium
CN119717508A