Multi-agent deep reinforcement learning method based on human guidance

By constructing a multi-agent value decomposition framework and a human real-time intervention mechanism, combining hybrid experience replay and adaptive entropy regularization, the problems of credit allocation, distribution offset and exploration efficiency in multi-agent reinforcement learning are solved, and efficient and secure multi-agent collaborative training is achieved.

CN119990246AActive Publication Date: 2025-05-13OCEAN UNIV OF CHINA

Patent Information

Application Number
CN202510466976.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing multi-agent reinforcement learning methods have shortcomings in credit allocation, distribution offset and generalization, exploration efficiency, etc., especially in multi-agent collaboration tasks, it is difficult to effectively utilize human prior knowledge and dynamic intervention mechanisms.

Method used

Using a multi-agent deep reinforcement learning method based on human guidance, efficient and secure multi-agent collaborative training is achieved by building a multi-agent value decomposition framework, human real-time monitoring and intervention, a hybrid experience playback pool, adaptive entropy regularization and elastic strategy update mechanism.

Benefits of technology

Significantly accelerate strategy convergence, reduce ineffective exploration time, improve training security and adaptability, optimize credit allocation, and find a balance between exploration and utilization to avoid local optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990246A_ABST
    Figure CN119990246A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent deep reinforcement learning method based on human guidance, and the method comprises the steps: constructing a multi-agent value decomposition framework, initializing a multi-agent environment and the state of each agent, innovatively fusing the elastic strategy update of human intervention, storing a high-quality experience sample generated by human intervention into an independent playback pool, and carrying out the recognition of a multi-agent value decomposition framework; mixing and sampling with samples generated by autonomous exploration of the intelligent agent according to priorities; maximizing the entropy in the strategy optimization target to encourage exploration, and dynamically adjusting the entropy coefficient through the target entropy; and circularly executing the steps of interaction between the intelligent agent and the environment, experience storage and strategy network and value network updating, and finally obtaining a greedy strategy of the intelligent agent according to a multi-intelligent agent soft actor commentator architecture combined with a value function decomposition technology as an output strategy. Through a human active intervention and value decomposition technology, efficient and safe training of a multi-agent cooperation task is realized, a human intervention sample and an autonomous exploration sample are trained in a mixed manner, and strategy optimization is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of reinforcement learning training strategies, and in particular relates to a multi-agent deep reinforcement learning method based on human guidance. Background Art

[0002] Multi-agent reinforcement learning can solve complex tasks that a single agent cannot handle (such as collaborative roundups and joint path planning) through the interaction and collaboration of multiple agents in a shared environment. However, existing methods still face the following key problems: Credit allocation problem: In fully collaborative tasks, it is difficult to fairly decompose global rewards into individual contributions. Traditional methods such as VDN decompose Q values ​​through linear addition, but its expression ability is limited; although QMIX introduces a nonlinear monotonic mixed network, it still cannot handle non-monotonic dependencies, resulting in suboptimal strategies; Distribution shift and lack of generalization: The dynamic differences between the training environment and the real environment (such as sensor noise and sudden obstacles) can easily lead to strategy failure. Existing methods (such as MADDPG and COMA) lack robustness design against distribution shift. Low exploration efficiency: Traditional MARL relies on random exploration, which has low sampling efficiency in high-dimensional state-action space and is prone to falling into local optimality.

[0003] Existing human-in-the-loop reinforcement learning methods optimize strategies through human demonstration, intervention, or evaluation, but have the following shortcomings in multi-agent scenario applications: Single interaction mechanism: Existing methods (such as TAMER and Deep TAMER) mostly use static imitation learning and cannot dynamically integrate human feedback and autonomous exploration, resulting in insufficient strategy flexibility; Insufficient support for multi-agent collaboration: Human intervention is usually targeted at a single agent, lacking overall guidance for group collaborative behavior, making it difficult to solve the credit allocation problem.

[0004] Comparison of existing technologies and summary of defects: QMIX / VDN: Although it solves part of the credit allocation problem, it relies on environmental reward signals and cannot use human prior knowledge to accelerate training; MADDPG: The simple centralized training decentralized execution (CTDE) framework has poor adaptability to non-stationary environments; COMA: Counterfactual Baseline is computationally complex and difficult to scale to large-scale intelligent systems. MASAC: The state space to be explored in multi-agent reinforcement learning is large, and too large an exploration range will lead to low exploration efficiency; Traditional HITL methods: lack of dynamic intervention mechanism for multi-agent collaboration and low sample utilization. Summary of the invention

[0005] In view of the above problems, the present invention provides a multi-agent deep reinforcement learning method based on human guidance, which includes the following process: S1, build a multi-agent value decomposition framework, initialize the multi-agent environment and the state of each agent, and at the same time, initialize the value function decomposition network, obtain the global action value function through the value function network and the strategy network, and synthesize the global value through a nonlinear monotone hybrid network to ensure that the individual-global maximization principle is met; S2, real-time monitoring and intervention by humans; humans monitor the local observation, action selection and environmental status of the agent in real time through a visual interactive interface; when humans believe that the action efficiency of the agent is lower than the preset value or there is a risk, they input alternative actions through the interactive interface to cover the original action of the agent, and store the intervention experience in an independent playback pool; S3, stores high-quality experience samples generated by human intervention into an independent replay pool, and mixes them with samples generated by autonomous exploration by the agent according to priority; the priority is dynamically assigned by the temporal difference error; S4, maximizes entropy in the policy optimization objective to encourage exploration, and dynamically adjusts the entropy coefficient through the target entropy; updates the policy network by optimizing the loss function to ensure that the agent finds a balance between exploration and exploitation; S5, loops through the steps of agent-environment interaction, experience preservation, value network, and policy network update until the preset termination condition is met; finally, the agent’s greedy strategy is obtained as the output strategy based on the multi-agent soft actor-critic architecture combined with value function decomposition technology.

[0006] Preferably, the global action value function in S1 is synthesized by a nonlinear monotone hybrid network, and the specific form is: ; Or as: ; represents the global action-value function, For hybrid networks, is the global state, is the local observation history, For joint action, is the mixing weight matrix, is the bias vector, which is dynamically generated by the hypernetwork according to the global state; Representing an Agent The local value function of Any number between 1 and N, input is the local observation history and actions .

[0007] Preferably, the hypernetwork is composed of a three-layer fully connected neural network, with the input being the global state , the output layer uses the Softplus activation function to ensure the non-negativity of the mixed weight; its calculation process is: ; ; FC 1, FC 2, FC 3 is a linear layer, ReLU is the activation function, and the Softplus layer is used to ensure that the weights are non-negative.

[0008] Preferably, the human intervention mechanism in S2 includes: Human input replaces action through interactive interface , covering the original action of the agent , generating intervention experience Deposit into independent replay pool ; Policy update loss function Introduce imitation learning terms in the form of: ; in, is the entropy coefficient, To imitate the learning weights, is the action distribution output by the policy network, is the joint action value, is the current state, is the current local observation of the agent, is the agent action, Action for humans, , For the experience pool D , Find the mathematical expectation of the contents in [ ] within the range.

[0009] Preferably, during S2 and S3, the agent uses an exploration strategy to interact with the environment, humans observe and are ready to intervene at any time, replace the agent's actions, and save the experience. Specifically: S21: Calculate the value function of all actions at the current moment; Each agent uses the current observation and the action at the previous moment as input, sends it to its value function network for calculation, and obtains the value function of all possible actions at the current moment; The value function network is calculated as follows: ;in, is the hidden state obtained by RNN, which represents the state information at the current moment; S22: exploration and exploitation decisions; Based on the current value function, the agent adopts - Greedy strategy selects actions: The probability of randomly choosing an action is The probability of selecting the action with the largest value function is as the training progresses. will gradually decrease; S23: Save experience; The agent stores the current observation, selected action, reward, and new observation in the experience buffer pool. This process continues until the experience buffer pool reaches the specified capacity.

[0010] Preferably, the training of the agent network is performed in the processes S2 and S3, including the following processes: S31: sampling from the experience buffer pool; During each training, N pieces of experience data are sampled from the experience buffer pool. Each piece of data contains the current observation, action, reward, and next observation. Then these data are used to calculate the action value function of each agent. S32: Calculate the global state code; One-hot encode the state of each agent through the global state information to ensure that the global information shared by multiple agents is effectively integrated; S33: Compute value decomposition weights and biases using hypernetwork; The QMIX network is used to weight the local value function of each agent and calculate the global joint action value function: ; in, is the joint action-value function, For intelligent agents The local value function of is the current state, is the agent action, is any number from 1 to N, For global status The weight matrix under is the bias term; Step 3-4: Calculate the loss of the joint action value function and optimize it; Calculate the loss of the joint action value function and optimize it, using the standard mean square error MSE loss function: ; in, For the expected return.

[0011] Preferably, the mixed experience replay pool sampling method in S3 is to store human experience from the experience replay pool. and storage of non-human experience Priority sampling is performed in .

[0012] Preferably, the entropy coefficient in S4 is dynamically updated by optimizing the target, and the loss function is: ; in, is the entropy coefficient, For the current strategy, is the target entropy, is the loss function, is the mathematical expectation operator symbol, The mathematical operator symbol is used to calculate the exponential moving average.

[0013] Preferably, the policy network update in S5 adopts a dual-Q network mechanism, and the update target is: ; The target value is calculated as: ; is the discount factor, is the target network output, target network parameters Through soft update Maintain stability, is the soft update weight coefficient, is the current state, is the local observation of the current agent, is the agent action, For human intervention, To update the target, is the mathematical expectation operator.

[0014] Preferably, the local value function network and the policy network both include a recurrent neural network (RNN) layer for encoding historical observation sequences.

[0015] Preferably, imitation learning weights To preset hyperparameters, they can be dynamically adjusted through hidden layer neural networks.

[0016] Preferably, the target network parameter update rate Preset hyperparameters.

[0017] Compared with the prior art, the present invention has the following beneficial effects: The present invention achieves efficient and safe training of multi-agent collaborative tasks by integrating human active intervention and multi-agent value decomposition technology, combined with hybrid experience replay, adaptive entropy regularization and elastic strategy update mechanism. Specifically, it is embodied in: Efficient training and strategy optimization: Providing high-quality demonstration samples through human intervention and combining priority mixed sampling can significantly accelerate strategy convergence and reduce invalid exploration time; Improved safety and robustness: Real-time human monitoring and action coverage mechanisms can avoid high-risk behaviors and ensure the safety of the training process; the value decomposition framework enhances adaptability to dynamic changes in the environment and reduces policy failures caused by distribution shifts; Credit allocation optimization: Based on the value decomposition architecture of nonlinear monotone hybrid networks, it solves the credit allocation problem in multi-agent collaboration while satisfying the individual-global maximization principle (IGM) and ensures the generation of the global optimal strategy; Balance between exploration and utilization: By maximizing the policy entropy and dynamically adjusting the entropy coefficient, the agent’s exploration ability is enhanced to avoid falling into local optimality while maintaining the stability of the policy.

[0018] Multi-agent value decomposition framework combined with active human intervention: Nonlinear hybrid network: Dynamically generates weight matrices and bias vectors through hypernetworks to synthesize global action value functions, solving the problem of insufficient expression of traditional linear decomposition (such as VDN) and supporting non-monotonic dependency modeling in complex tasks; Independent generation of local value functions: Each agent independently generates a local value function and combines it with the hybrid network to synthesize the global value, ensuring the fairness of credit allocation and the interpretability of the strategy.

[0019] Real-time human intervention mechanism: High-quality experience injection: Humans directly correct inefficient or dangerous behaviors through alternative actions. The generated high-quality samples are stored in an independent replay pool, providing reliable prior knowledge for strategy optimization. Introduction of imitation learning items: Add imitation learning items of human actions to the policy loss function, forcing the policy network to approach human demonstrations, accelerating policy optimization and improving action safety;

[0020] Mixed Priority Experience Replay: Dynamic priority sampling: Assign sample priorities based on temporal difference error (TD-error), give priority to mixed training with high-value human intervention samples and autonomous exploration samples, and improve sample utilization and learning efficiency; Independent playback pool design: Separate human experience from non-human experience to avoid sample distribution conflicts while retaining the diversity of autonomous exploration; Parameter soft update: Update the target network parameters through exponential moving average to avoid sudden changes in strategy and ensure the smoothness of the training process.

[0021] Hypernetwork structure optimization: Non-negative weight constraint: The Softplus activation function is used to ensure the non-negativity of the mixed weights, satisfying the monotonicity requirement of the IGM principle while supporting flexible global value synthesis. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic flow diagram of the method of the present invention.

[0023] Figure 2 It is a schematic diagram of the interaction of each module in the algorithm of the present invention.

[0024] Figure 3 This is the specific updating process of the algorithm of the present invention.

[0025] Figure 4 It is a schematic diagram of the network architecture in the present invention. DETAILED DESCRIPTION

[0026] The present invention provides a multi-agent value decomposition deep reinforcement learning method based on human guidance. The overall process is as follows: Figure 1 As shown: the following steps are included: Step 1: Construct a multi-agent value decomposition framework; Initialize the environment and agent states: Initialize the multi-agent environment and the states of each agent, including the local observations, action selections, and global states of the environment. At the same time, initialize the value function decomposition network, generate local values ​​for each agent, and synthesize the global value through a nonlinear monotonic hybrid network to ensure that the individual-global maximization principle (IGM) is met. Step 2: Update the resilience strategy by integrating human intervention; Real-time human monitoring and intervention: Humans monitor the agent's local observations, action selections, and environmental status in real time through a visual interactive interface. When humans believe that the agent's actions are inefficient or risky, they input alternative actions through the interactive interface to override the agent's original actions, and store the intervention experience in an independent replay pool. Step 3: Mix the experience replay pool; Experience sample storage and sampling: High-quality experience samples generated by human intervention are stored in an independent playback pool and mixed with samples generated by autonomous exploration by the agent according to priority. Priority is dynamically allocated by the temporal difference error (TD-error) to ensure that high-quality samples are fully sampled during training. Step 4: Adaptive entropy regularization constraint; Policy optimization and entropy regulation: Maximize entropy in the policy optimization objective to encourage exploration, and dynamically adjust the entropy coefficient through the target entropy; update the policy network by optimizing the loss function to ensure that the agent finds a balance between exploration and exploitation; Step 5: Loop through training and evaluation Iterative training and evaluation: The steps of interaction between the agent and the environment, experience preservation, and strategy update are executed cyclically until the preset termination condition is met; finally, the greedy strategy of the agent is obtained according to the agent value function and used as the output strategy to complete the multi-agent reinforcement learning training.

[0027] The present invention is further described below in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0028] Example 1.

[0029] Step 1: Initialize the environment and all agent states; Initialize the agent network: Before training, all weights and biases of the agent value function network and the target network are initialized; Initialize the environment: Initialize the task environment. For example, assume that the scenario is a multi-agent collaborative path planning task. The environment state includes the current location information of each agent, the target location, etc. Initialize the experience cache pool: Set up an experience cache pool to store each agent's interaction experience, including observations, actions, rewards, and new observations. The size of the pool is limited, and when the storage is full, the earliest experience will be discarded.

[0030] Step 2: The agent uses the exploration strategy to interact with the environment and saves experience. During this process, humans observe the agent's actions and global status in real time through a visual interface. If the agent is found to have inappropriate behavior (such as dangerous exploration behavior or inefficient exploration behavior), humans directly interrupt the interaction between the agent and the environment through the interactive interface, and input human intervention actions to replace the agent's actions, guide the agent's next behavior, and generate human intervention experience (current agent local observation, global status, human intervention actions). Step 2-1: Calculate the value function of all actions at the current moment; Each agent uses the current observation and the action at the previous moment as input, sends it to its value function network for calculation, and obtains the value function of all possible actions at the current moment; The value function network is calculated as follows: in, is the hidden state obtained by RNN, which represents the state information at the current moment; Step 2-2: Exploration and exploitation decision; Based on the current value function, the agent adopts - Greedy strategy selects actions: The probability of randomly choosing an action is The probability of selecting the action with the largest value function is . As the training progresses, will gradually decrease; in this process, humans observe the agent's actions and global status in real time through a visual interface. If the agent is found to have dangerous or inefficient exploration behavior, humans interrupt the training through the interactive interface and input human intervention actions to replace the current agent's actions. For example, in the path planning task, agent A chooses a circuitous path (30 meters away from the target, with dynamic obstacles). The human operator determines that this path is time-consuming and risky, and inputs the replacement action "accelerate to the right" through the interactive interface, directly overwriting the original action of agent A "go straight"; Step 2-3: Save experience; The agent stores the current observation, selected action, reward, and new observation in the experience buffer pool. If there is human intervention in step 2-2, the selected action is changed to the human intervention action. This process continues until the experience buffer pool reaches the specified capacity.

[0031] Step 3: Perform training of the agent network; Step 3-1: Sampling from the experience buffer pool During each training, N pieces of experience data are sampled from the experience buffer pool (each piece of data contains the current observation, action, reward, and next observation). These data are then used to calculate the action value function of each agent; Step 3-2: Calculate the global state code; One-hot encode the state of each agent through the global state information to ensure that the global information shared by multiple agents is effectively integrated; Step 3-3: Use the hypernetwork to calculate the value decomposition weights and biases; The QMIX network is used to weight the local value function of each agent and calculate the global joint action value function: ; in, is the joint action-value function, For intelligent agents The local value function of For global status The weight matrix under is the bias term; Step 3-4: Calculate the loss of the joint action value function and optimize it; Calculate the loss of the joint action value function and optimize it, using the standard mean squared error (MSE) loss function: ; in, For the expected return.

[0032] Step 4: Loop through steps 2 and 3 until the training is terminated; Steps 2 and 3 will be executed repeatedly until the training termination condition is reached, such as convergence or the maximum number of training rounds.

[0033] In this example, the value function decomposition method learns local value function, and then these local The values ​​are combined with a learnable hybrid neural network to generate the joint value , This algorithm uses the value function decomposition network in QMIX, which is a nonlinear monotone decomposition structure that can achieve a richer function class while satisfying the Individual-Global Maximization (IGM) principle: The result of performing the global argmax operation on each local The results of performing argmax operations on the value network are the same. The value network update method is similar to the Critic update method used in the traditional SAC algorithm, but now the Critic network consists of “local Network + Hybrid Network ”, where the hybrid network The network architecture is based on the QMIX paper The network is modified from the intelligent agent Part of network and hybrid networks The weights and biases of the hybrid network are generated by independent hyper-networks. Figure 4 Demonstrates value decomposition mixing The detailed structure of the network, for each agent , there is a local Network to represent its local Value Function The hybrid network localizes the agent The output of the network is linearly mixed as input and then generated by the absolute value activation function The value of .

[0034] In the process of updating the policy network, two soft Value Network ,for j ∈{1,2}, and take the minimum value as the target: ; At the same time, it can also be expressed as follows: ; is the experience replay pool, which contains the state transition process (state s , local observations ,action ,award , next state , the next local observation ), is the target The network, its parameters It is through the current Network weights This is obtained by taking an exponentially moving average of , which has been shown in the original SAC paper to be useful for stabilizing the training process.

[0035] The structure and agent of the policy network Part of The network is the same, but with an additional softmax layer to output a probability distribution.

[0036] Please note that in the formula, From the agent i Current strategy The algorithm is obtained by sampling instead of sampling from the replay buffer. Compared with the original QMIX algorithm, an additional policy network is introduced, which outputs a probabilistic policy (a probability mass function in the discrete domain) that accurately represents the probability value of each agent choosing each discrete action. Therefore, the expected value can be accurately calculated. Recent studies have theoretically proved that soft (or Boltzmann) policy iteration can guarantee improvement and convergence to the optimal policy. Based on the soft policy iteration process, the goals of policy update are as follows: ; At the same time, it can also be expressed as follows: ; is a hyperparameter that controls the trade-off between maximizing the entropy of the policy and the expected discounted return, and is updated in the same way as in the original SAC paper. The overall value decomposition network architecture of the present invention is as follows Figure 4 shown.

[0037] How to update in the loop: like Figure 3As shown, in order to incorporate human intervention samples, an imitation learning loss term is added to encourage the strategy to be close to human guidance behavior. Suppose the human intervention sample is , the strategy update becomes: ; in is a standard sample in the experience replay pool, are samples generated by human intervention, and 𝛽 is a hyperparameter that controls the weight of the imitation learning loss term.

[0038] In the above way, human intervention directly affects the update of the policy network through imitation learning, pushing the agent's strategy closer to the high-quality behavior provided by humans.

[0039] After adding human intervention samples, human intervention affects the way the target value is generated. For the intervention samples provided by humans, assuming that human actions are considered to have higher quality, the value network update method becomes: ; in: ; After introducing human intervention samples, the joint value By human action The action distribution without using the policy network is calculated to be:

[0040] Example 2.

[0041] The value decomposition multi-agent reinforcement learning training method using human guidance of the present invention can be applied to multiple fields such as robot control, traffic coordination, manufacturing control, etc.

[0042] The multi-agent reinforcement learning training method described in the present invention is used for the collaborative control of a multi-robot system, wherein each agent represents a robot. Each robot includes multiple modules, such as a mobile module, a grasping module, a perception module, a navigation module, etc. The strategy of the agent is the currently executed action, wherein the mobile module controls the movement direction and speed of the robot, the grasping module controls the action of grasping and placing objects, the perception module is used to process and understand environmental information, and the navigation module is responsible for planning the robot's travel path. All these actions are controlled by the agent value function network, and the reward function represents the efficiency of multiple robots in collaboratively completing tasks, such as task completion time, path length, collaborative effect, etc. Each robot can only perceive and control its own working state, and cannot directly observe the internal state and behavior of other robots. The global state of the environment contains the state information of each robot (such as position, task progress, power status, etc.) and the collaborative state of the entire system. The robots coordinate and cooperate by observing local environmental information to achieve common goals.

[0043] This embodiment involves a multi-robot collaborative task. In this scenario, each robot is regarded as an intelligent agent, and completes a common task through collaboration.

[0044] Step 1: Initialize the entire multi-robot system and the state of each robot; 1.1 Initialize the agent network: Each robot (agent) initializes its value function network, which has the same structure as the single agent scenario and consists of three layers: a linear layer (MLP), a recurrent neural network (RNN) layer, and a linear output layer; 1.2 Initialize the environment: The state of the environment includes the current position, mission objectives, and specific requirements of all robots. The state of each robot contains its local information, and they collaborate by sharing the global state.

[0045] Step 2: Each robot uses an exploration strategy to interact with the environment and saves the experience to the experience cache pool; during this process, humans observe the agent's actions and global status in real time through a visual interface. If the agent is found to have inappropriate behavior (such as dangerous exploration behavior or inefficient exploration behavior), humans directly interrupt the interaction between the agent and the environment through the controller, and input human intervention actions to replace the agent's actions (current movement direction, speed, grasping data), guide the agent's next behavior, and generate human intervention experience (current agent local observation, global status, human intervention action). For example, if agent B (responsible for grasping) and agent C (responsible for transportation) may collide due to path conflict, humans observe through the visual interface that the expected paths of the two overlap and intervene immediately: 1. Input the alternative action “pause grasping” to agent B; 2. Input the alternative action “turn left to avoid” for agent C; 2.1 Step 2-1: Each robot calculates the action value function; Each robot calculates the value function of the action through the value function network based on the current observation (such as position, task status, etc.) and the action at the previous moment; 2.2 Step 2-2: Exploration and Exploitation Decisions; Each robot chooses whether to perform a random action or an optimal action based on the current value function. If the action performed is dangerous or inefficient, the human observer can directly interrupt the action execution and input human intervention actions through the controller to replace the current action of the intelligent agent. 2.3 Step 2-3: Save experience; Each robot saves the current state, action, reward and new state to the experience pool. If there is human intervention in step 2-2, the agent action in the saved experience is the human intervention action.

[0046] Step 3: Perform training of the agent network; 3.1 Step 3-1: Sampling from the experience buffer pool; Sample N pieces of experience data from the experience cache pool, calculate the action value function of each robot, and update its strategy; 3.2 Step 3-2: Calculate the global state encoding; Use global state to encode the state of each robot and ensure information sharing in multi-robot systems; 3.3 Step 3-3: Value decomposition using QMIX network; Use the QMIX network to calculate the joint value function of all robots and perform effective credit allocation; 3.4 Step 3-4: Calculation of joint action value function; Calculate the joint action-value function through the decomposition structure and update the strategy; 3.5 Step 3-5: Optimize the target network; The target network is used to calculate and optimize the loss of the joint action value function, and the networks of all robots are updated.

[0047] Step 4: Repeat steps 2 and 3 until the training is terminated. Steps 2 and 3 are executed repeatedly until the termination condition is met.

[0048] Value decomposition network architecture: Local Network: Each agent trains its own local function , the input is the local observation history ,The output is the action value; Super Network Design: Input: global state s; Structure: 3-layer fully connected network, hidden layer activation function is ReLU; Output: Mixed weights and bias , through the softplus activation function to ensure ; Global Value synthesis: Determine the global by linear weighting value : .

[0049] Human active intervention mechanism: Intervention process: 1. Humans view local observations of each agent in real time through a visual interface , Action Probability Distribution and the environment global state s; 2. When the agent's actions are observed to be inefficient (e.g., detours) or risky, humans intervene in the agent's actions through the interactive interface and input alternative actions. ; 3. System executes immediately , and transfer the state before and after the intervention Deposit .

[0050] Experience Priority: Human intervention samples are given the highest priority by default in priority sampling to ensure that they are fully sampled during training.

[0051] Strategy Optimization: Policy loss function: ; in: SAC goals; is the imitation learning item; Critic Network and software updates: Using two soft Value Network , and take the minimum value as the target, for j ∈{1,2}: ; Soft update mechanism: target network parameters , improve training stability, including To preset hyperparameters, such as .

[0052] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0053] Although the above describes the specific implementation methods of the present invention, it is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art on the basis of the technical solution of the present invention without creative work are still within the scope of protection of the present invention.

Claims

1. A multi-agent deep reinforcement learning method based on human guidance, characterized in that: The process includes: S1, build a multi-agent value decomposition framework, initialize the multi-agent environment and the state of each agent, and at the same time, initialize the value function decomposition network, obtain the global action value function through the value function network and the strategy network, and synthesize the global value through a nonlinear monotone hybrid network to ensure that the individual-global maximization principle is met; S2, real-time monitoring and intervention by humans; humans monitor the local observation, action selection and environmental status of the agent in real time through a visual interactive interface; when humans believe that the action efficiency of the agent is lower than the preset value or there is a risk, they input alternative actions through the interactive interface to cover the original action of the agent, and store the intervention experience in an independent playback pool; S3, stores high-quality experience samples generated by human intervention into an independent replay pool, and mixes them with samples generated by autonomous exploration by the agent according to priority; the priority is dynamically assigned by the temporal difference error; S4, maximizes entropy in the policy optimization objective to encourage exploration, and dynamically adjusts the entropy coefficient through the target entropy; updates the policy network by optimizing the loss function to ensure that the agent finds a balance between exploration and exploitation; S5, loops through the steps of agent-environment interaction, experience preservation, value network, and policy network update until the preset termination condition is met; finally, the agent’s greedy strategy is obtained as the output strategy based on the multi-agent soft actor-critic architecture combined with value function decomposition technology.

2. A multi-agent deep reinforcement learning method based on human guidance as claimed in claim 1, characterized in that: The global action value function in S1 is synthesized by a nonlinear monotone hybrid network, and its specific form is: ; Or as: ; represents the global action-value function, For hybrid networks, is the global state, is the local observation history, For joint action, is the mixing weight matrix, is the bias vector, which is dynamically generated by the hypernetwork according to the global state; Representing an Agent The local value function of Any number between 1 and N, input is the local observation history and actions .

3. A multi-agent deep reinforcement learning method based on human guidance as claimed in claim 2, characterized in that: The hypernetwork consists of three layers of fully connected neural networks, with the input being the global state , the output layer uses the Softplus activation function to ensure the non-negativity of the mixed weight; its calculation process is: ; ; FC 1, FC 2, FC 3 is a linear layer, ReLU is the activation function, and the Softplus layer is used to ensure that the weights are non-negative.

4. A multi-agent deep reinforcement learning method based on human guidance as claimed in claim 1, characterized in that: The human intervention mechanisms in S2 include: Human input replaces action through interactive interface , covering the original action of the agent , generating intervention experience Deposit into independent replay pool ; Policy update loss function Introduce imitation learning terms in the form of: ; in, is the entropy coefficient, To imitate the learning weights, is the action distribution output by the policy network, is the joint action value, is the current state, is the agent’s current local observation, is the agent action, Action for humans, , For the experience pool , Find the mathematical expectation within the range.

5. The multi-agent deep reinforcement learning method based on human guidance as claimed in claim 1, characterized in that: In the process of S2 and S3, the agent uses the exploration strategy to interact with the environment, humans observe and are ready to intervene at any time, replace the agent's actions, and save the experience. Specifically: S21: Calculate the value function of all actions at the current moment; Each agent uses the current observation and the action at the previous moment as input, sends it to its value function network for calculation, and obtains the value function of all actions at the current moment; The value function network is calculated as follows: ;in, is the hidden state obtained by RNN, which represents the state information at the current moment; S22: exploration and exploitation decisions; Based on the current value function, the agent adopts Greedy strategy selects actions: The probability of randomly choosing an action is The probability of selecting the action with the largest value function is as the training progresses. will gradually decrease; S23: Save experience; The agent stores the current observation, selected action, reward, and new observation in the experience buffer pool. This process continues until the experience buffer pool reaches the specified capacity.

6. A multi-agent deep reinforcement learning method based on human guidance as claimed in claim 1, characterized in that: The training of the agent network is performed in S2 and S3, including the following steps: S31: sampling from the experience buffer pool; During each training, N pieces of experience data are sampled from the experience buffer pool. Each piece of data contains the current observation, action, reward, and next observation. Then these data are used to calculate the action value function of each agent. S32: Calculate the global state code; One-hot encode the state of each agent through the global state information to ensure that the global information shared by multiple agents is effectively integrated; S33: Compute value decomposition weights and biases using hypernetwork; The QMIX network is used to weight the local value function of each agent and calculate the global joint action value function: ; in, is the joint action-value function, For intelligent agents The local value function of is the current state, is the agent action, is any number from 1 to N, For global status The weight matrix under is the bias term; Step 3-4: Calculate the loss of the joint action value function and optimize it; Calculate the loss of the joint action value function and optimize it, using the standard mean square error MSE loss function: ; in, For the expected return.

7. The multi-agent deep reinforcement learning method based on human guidance according to claim 1, characterized in that: The mixed experience replay pool in S3 is sampled by storing human experience in the experience replay pool. and storage of non-human experience Priority sampling is performed in .

8. The multi-agent deep reinforcement learning method based on human guidance as claimed in claim 1, characterized in that: The entropy coefficient in S4 is dynamically updated by optimizing the target, and the loss function is: ; in, is the entropy coefficient, For the current strategy, is the target entropy, is the loss function, is the mathematical expectation operator symbol, The mathematical operator symbol is used to calculate the exponential moving average.

9. The multi-agent deep reinforcement learning method based on human guidance as claimed in claim 1, characterized in that: The policy network update in S5 adopts a dual-Q network mechanism, and the update target is: ; The target value is calculated as: ; is the discount factor, is the target network output, target network parameters Through soft update Maintain stability, is the soft update weight coefficient, is the current state, is the local observation of the current agent, is the agent action, For human intervention, To update the target, is the mathematical expectation operator.

10. The multi-agent deep reinforcement learning method based on human guidance according to claim 1, characterized in that: Both the local value function network and the policy network contain recurrent neural network (RNN) layers to encode historical observation sequences.

Citation Information

Patent Citations

  • Value decomposition multi-agent reinforcement learning training method using attention network

    CN114861932A

  • Multi-agent reinforcement learning method for optimizing experience storage and experience reutilization

    CN116205273A

  • Equipment optimal maintenance strategy searching method and system based on reinforcement learning

    CN118941275A

  • Multi-agent cooperative control method and system based on deep reinforcement learning, and medium

    CN119717508A

  • Sample-efficient reinforcement learning

    WO2023171102A1

Cited By

  • Decomposition line balancing method based on deep reinforcement learning

    CN120297153A

  • Comprehensive energy system optimization method and system in multiple uncertain environments

    CN120822667A

  • Space-time enhanced multi-agent reinforcement learning-based empty box allocation scheme optimization method

    CN121937051A

  • Empty container distribution optimization method based on space-time reinforcement multi-agent reinforcement learning

    CN121937051B