Collaborative Exploration Method for Multi-Agent Sparse Reward Environments Based on Intrinsic Motivation
By constructing artificial potential fields and calculating advantageous functions, the problem of low efficiency of multi-agent collaborative exploration in reward sparse environments is solved, and more efficient exploration strategy updates and full exploration of environmental spaces are achieved.
Patent Information
- Application Number
- CN202111455606.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-01
AI Technical Summary
In the reward sparse environment, the collaborative exploration efficiency of multiple agents is low, and the existing technology is difficult to effectively adapt to the strong correlation between behaviors between agents and the correlation between task completion and some elements of the environment, resulting in low exploration efficiency.
A multi-agent sparse reward environment collaborative exploration method based on intrinsic motivation is adopted to guide exploration strategies for exploration of underexplored areas in the environment by constructing artificial potential field functions, and a counterfactual baseline method is used to calculate the dominant function of the agent, allocate potential energy impact to update the exploration strategy.
It improves the exploration efficiency of multiple agents in sparse reward environments, is suitable for a variety of reinforcement learning algorithms, reduces dependence on prior information, and is suitable for most sparse reward exploration environments.
Smart Images

Figure CN114169421B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-agent deep reinforcement learning, and particularly to a multi-agent collaborative exploration method in a sparse reward environment. Background Art
[0002] The related prior arts of the present invention include:
[0003] I. Distributed Partially Observable Markov Decision Process (Dec-POMDP) is defined as follows:
[0004] <S, U, P, r, O, Z, n, γ>
[0005] Among them, n represents the number of agents, S is the state set, and U is the joint action of the agents.
[0006] II. COMA (Counterfactual Multi-Agent Policy Gradients) is an algorithm proposed for the belief assignment problem in multi-agent reinforcement learning. The belief assignment problem is one of the problems widely existing in multi-agent cooperation tasks. The difficulty of the problem lies in how to distinguish the contribution degree of each agent to the global reward when the agents share a global reward. COMA calculates an advantage function for each agent a:
[0007]
[0008] This advantage function can reflect the quality of the agent's current action selection compared with the unselected actions. Using this advantage function to update the agent's policy can solve the problem that the multi-agent algorithm cannot achieve good results due to the belief assignment problem.
[0009] Currently, the methods for exploration in sparse reward environments are limited by some special settings, such as strong correlations between the behaviors of agents, and the completion of tasks is only related to some elements in the environment. How to adapt to exploration in a wider range of sparse reward environments remains an open problem. Summary of the Invention
[0010] The present invention aims to solve the problem of multi-agent collaborative exploration in a sparse reward environment, and proposes a multi-agent collaborative exploration method in a sparse reward environment based on intrinsic motivation, which realizes multi-agent collaborative exploration in a sparse reward environment based on an artificial potential field.
[0011] The present invention is realized by the following technical solutions:
[0012] A multi-agent collaborative exploration method in a sparse reward environment based on intrinsic motivation specifically includes the following steps:
[0013] Step 1, initialize the target policy This strategy is used to learn to complete the target task while initializing the exploration strategy This strategy is used for sufficient exploration in the environment; where π represents the current strategy of the agent, and n is the number of agents;
[0014] Step 2: Construct an artificial potential field function. By constructing an artificial potential field in the environment, guide the exploration strategy to explore in the environment according to the potential energy in the artificial potential field, strengthen the exploration of the under-explored area, so as to obtain successful experience and guide the target strategy to learn;
[0015] Step 3: Potential energy influence distribution, specifically processed as follows:
[0016] Using the counterfactual baseline method, the advantage function of agent a is calculated using the following formula, as shown below:
[0017]
[0018] where u a represents the action of agent a, u -a represents the joint action of other agents, π represents the current strategy of agent a, and A a represents the influence of agent a taking action u a compared with taking other actions on the potential energy, the larger A a , the greater the influence of the current action u a of agent a on the potential energy compared with other actions, and vice versa. Then, for each agent i, its corresponding A i is calculated, and the proportion of the agent's internal influence by the potential energy is obtained through the softmax operation:
[0019]
[0020] Let the reward of agent i at each decision step t be as shown below:
[0021]
[0022] Step 4: Use the influence of the artificial potential field to update the exploration strategy, that is, use the intensity of the artificial potential field after belief assignment to guide the exploration strategy to explore, accelerate the exploration of the environmental space, and use the successful experience signal to guide the learning of the target strategy.
[0023] Compared with the prior art, the present invention has a relatively high improvement in the exploration efficiency of agents in a sparse reward environment and can be combined with a variety of reinforcement learning algorithms; it uses less prior information and can be applied to most exploration environments with sparse rewards. Brief Description of the Drawings
[0024] Figure 1 This is the overall flowchart of a multi-agent sparse reward environment collaborative exploration method based on intrinsic motivation of the present invention;
[0025] Figure 2 This is the algorithm architecture diagram of the present invention;
[0026] Figure 3 This is an example of the application scenario of the present invention. Specific implementation manners
[0027] The technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0028] As Figure 1 shown, this is a multi-agent sparse reward environment collaborative exploration method based on intrinsic motivation of the present invention. As Figure 2 shown, this is the algorithm architecture diagram of the present invention. The method specifically includes the following steps:
[0029] Step 1, initialize the target policy and exploration policy:
[0030] Initialize the target policy This policy is used to learn to complete the target task, and at the same time initialize the exploration policy This policy is used to fully explore in the environment; where, π represents the current policy of the agent, and n is the number of agents;
[0031] Step 2, construct an artificial potential field function. By constructing an artificial potential field in the environment, guide the exploration policy to explore in the environment according to the potential energy in the artificial potential field, strengthen the exploration for the insufficiently explored area, so as to obtain successful experience and guide the target policy to learn; for example, in the field of robot path planning, the configuration space is usually regarded as an area with undulating terrain, where the starting point and obstacle points are located in the higher areas, the end point is located in the lower area, and the robot is regarded as a sphere. Then the robot will slide along a certain trajectory from the higher starting point to the lower end point under the action of gravity and avoid the lower obstacles. This step specifically includes the following processing:
[0032] Step 2.1, perform exploration sufficiency measurement, and the specific processing is as follows:
[0033] Model the collaborative multi-agent exploration task as a distributed partially observable Markov decision process (Dec-POMDP), as shown in the following formula:
[0034] <S, U, P, r, O, Z, n, γ>
[0035] Among them, S is the set of global states of the agents, U is the set of joint actions of the agents, P is the transition function, r is the global reward function, O is the set of local observations of the agents, Z is the initial global state distribution, n is the number of agents, and γ is the reward discount factor in reinforcement learning
[0036] At time step t, there exists the joint state S of the agents t , and each agent selects an action to execute, obtains a global reward r, and transfers to the next joint state S according to the transition function P t+1 . In the training paradigm of centralized training and distributed execution, at each time step t, each agent i selects an action u i according to its own observation function o i , forms the joint action u t , interacts with the environment, and transfers to the next joint state s t+1 according to the environmental state transition function P(s t , u t ), and at the same time obtains the environmental feedback r t+1 (s, u). The Count-based exploration algorithm is one of the simplest algorithms for exploration by designing intrinsic rewards. The counter Counter C(S t represents the number of times the multi-agent system takes the joint action u t when the joint state is S t during the entire training process; t under the joint action of u t ;
[0037] Step 2.2, Distance measurement network training
[0038] In a multi-agent environment, since the positions of multiple agents are involved, a distance measurement network is used to measure the distance between two states. The input of the distance measurement network is the global state S t and the joint action u t , and the output is a value used to measure the distance between two states;
[0039] The distance measurement formula is as follows:
[0040] dis = ||f(s t+1 , u t+1 ) - f(s t , u t )|| 2
[0041] Among them, f() represents the fitting function, and dis represents the distance between two states.
[0042] Step 2.3, Construct an artificial potential field, and the specific processing is as follows:
[0043] Sample a batch of data from the data pool, and use the state-action pair with the least Counter as the target state (s, u). goal ,
[0044] To prevent the problem of excessive gravitational force at positions far from the target, sample the piecewise gravitational potential energy, so that there is gravitational potential energy, as shown in the following formula:
[0045]
[0046] where d((s, u), (s, u) goal ) represents the distance between the current state and the target, is a hyperparameter. When the distance between the two is less than or equal to , the gravitational potential energy shows a square form; otherwise, the magnitude of the gravitational potential energy is reduced;
[0047] Step 3: Potential energy influence distribution, the specific processing is as follows:
[0048] Using the counterfactual baseline method, calculate the advantage function of agent a with the following formula, as shown in the following formula:
[0049]
[0050] where u a is the action of agent a, u -a is the joint action of other agents, π represents the current policy of agent a, and A a represents the magnitude of the influence of potential energy when agent a takes action u a compared with taking other actions. The larger A a is, the greater the degree of influence of the current action u a of agent a by potential energy compared with other actions, and vice versa. Then calculate the corresponding A i for each agent i, and obtain the proportion of the internal influence of potential energy within the agent through the softmax operation:
[0051]
[0052] Let the reward of agent i at each time step t be as shown in the following formula:
[0053]
[0054] Step 4: Use the artificial potential field influence to update the exploration strategy, that is, use the artificial potential field strength influence after belief distribution to guide the exploration strategy to explore, accelerate the exploration of the environmental space, and use the successful experience signal to guide the target strategy learning.
[0055] The present invention aims to solve the problem of low efficiency of multi-agent collaborative exploration in a sparse reward environment. The present invention is not limited to the requirements of the relationship between agents or the particularity of tasks, and can be applied to most multi-agent collaborative exploration scenarios with sparse rewards.
[0056] As Figure 3 shown, it is an application scenario of the present invention. In this scenario, two robots need to cooperate to explore a room and will receive a reward after cleaning up all the banana peels in the room. The robots do not know the location of the banana peels in advance, nor do they know the location of their teammates. In the task as Figure 3 shown, n is 2.
[0057] The above is an exemplary description of the present invention. It should be noted that any simple deformation, modification, or equivalent replacement that can be made by those skilled in the art without creative efforts falls within the protection scope of the present invention without departing from the core of the present invention.
Claims
1. A collaborative robot path exploration method for a multi-agent sparse reward environment based on intrinsic motivation, characterized in that The method specifically includes the following steps: Step 1: Initialize the target policy This policy is used to learn to complete the target task; at the same time, initialize the exploration policy This policy is used to fully explore the environment; where n is the number of agents; Step 2: Construct an artificial potential field function for robot path planning. By constructing an artificial potential field in the environment, guide the exploration strategy to explore in the environment according to the potential energy in the artificial potential field, strengthen the exploration of the insufficiently explored area, so as to obtain successful experience and guide the target strategy to learn. Specifically, obtain a configuration space as an area with undulating terrain, where the starting point and obstacle points are located in the high area, and the end point is located in the low area. The robot is regarded as a sphere, then the robot will slide from the high starting point to the low end point along a certain trajectory under the action of gravity and avoid obstacles. The said Step 2 further includes the following processing: Step 2.1: Conduct exploration sufficiency measurement, and the specific processing is as follows: Model the collaborative multi-agent exploration task as a distributed partially observable Markov decision process (Dec-POMDP), as shown in the following formula: <S, U, P, r, O, Z, n, γ> where S represents the set of global states of the agents, U represents the set of joint actions of the agents, P represents the transition function, r is the global reward function, O represents the set of local observations of the agents, Z represents the initial global state distribution, n represents the number of agents, and γ represents the reward discount factor in reinforcement learning; Use the counter Counter C(S t , u t ) to represent the number of times the multi-agent system takes the joint action u t when the joint state is S t during the entire training process; Step 2.2: Training of the distance measurement network In a multi-agent environment, a distance measurement network is used to measure the distance between two states. The input of the distance measurement network is the global state S t and the joint action u t , and the output is a value used to measure the distance between two states. The distance measurement formula is as follows: dis = ||f(s t+1 , u t+1 ) - f(s t , u t )|| 2 where f() represents the fitting function, and dis represents the distance between two states; Step 2.3: Construct an artificial potential field, and the specific processing is as follows: Sample a batch of data from the data pool, and use the state-action pair with the fewest Counter as the target state (s, u) goal , Sample the segmented gravitational potential energy, and the gravitational potential energy is as shown in the following formula: where d((s, u), (s, u) goal ) represents the distance between the current state and the target state, represents a hyperparameter. When the distance between the two is less than or equal to , the gravitational potential energy exhibits a quadratic form; otherwise, the magnitude of the gravitational potential energy is reduced. Step 3: Conduct the distribution of potential energy influence, and the specific processing is as follows: Use the counterfactual baseline method to calculate the advantage function of agent a with the following formula, as shown in the following formula: A a = U att (S, U)-∑ u′ aπ(u′ a |o a )U att (S, U -a , u′ a ) Among them, U -a is the joint action of other agent-a, A a represents that agent a takes action u' under the current policy a compared with the magnitude of the potential energy affected by taking other actions, A a The larger it is, the greater the degree to which the current action of agent a is affected by potential energy compared with other actions, and vice versa; then calculate the corresponding A for each agent i i , and obtain the proportion of the potential energy influence within the agent through the softmax operation: Let the reward of agent \(i\) at each time step \(t\) be \(r\). i t As shown in the following formula: Step 4: Use the influence of the artificial potential field to update the exploration strategy, accelerate the exploration of the environmental space, and use the successful experience signal to guide the learning of the target strategy.
Citation Information
Patent Citations
Reinforced learning path planning algorithm based on potential field
CN110794842A
Mixed-experience multi-agent reinforcement learning motion planning method
CN113341958A