Multi-agent collaborative exploration method for constructing intrinsic individual rewards based on information gain
By using information gain and exploration frequency factors in the multi-agent reinforcement learning environment to construct intrinsic individual rewards, the problem of low exploration efficiency of agents in sparse reward environments is solved, and more efficient exploration and learning is achieved.
Patent Information
- Application Number
- CN202510167422.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-16
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-16
AI Technical Summary
In a multiagent reinforcement learning environment, it is difficult for an agent to effectively explore and attribute the impact of its behavior on the environment in a sparse reward environment, especially in the case of global intrinsic rewards.
By training multiple agents from both global and local environments, using information gain and exploration frequency factors to construct intrinsic individual rewards, guiding agents to prioritize exploring unknown or high predictive uncertainty.
It significantly improves the exploration efficiency of multi-agent systems, helps agents learn and make decisions more effectively in complex environments and unseen areas, and improves learning efficiency and decision-making capabilities.
Smart Images

Figure CN120087410A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent reinforcement learning, and particularly relates to a multi-agent collaborative exploration method for constructing an intrinsic individual reward based on information gain. Background Art
[0002] In multi-agent reinforcement learning, there are usually multiple agents learning and making decisions in a shared environment. The goal of reinforcement learning is to obtain the highest possible cumulative reward by interacting with the environment. However, in some tasks, the rewards in the environment may be sparse, that is, the agents do not receive reward feedback for a long time, which makes it difficult for them to judge which behaviors are beneficial, thus affecting the learning process. To help the agents explore these environments more effectively, researchers often introduce intrinsic rewards, that is, reward signals that do not rely on external environmental feedback, to encourage the agents to explore different state spaces.
[0003] However, in a multi-agent environment, although intrinsic rewards can promote exploration, they also bring another problem: the credit assignment problem. That is, when multiple agents participate in a task together, it is difficult to clearly attribute the contribution to each agent because their behaviors are highly coupled. Especially in the case of using global intrinsic rewards, it may be even more difficult to know which agent's behavior led to a certain change in the environment. Therefore, even if a reward signal is given, how to correctly attribute the reward to each agent's behavior is a complex problem. In traditional single-agent reinforcement learning, the agent's behavior directly affects its reward, so the credit assignment problem is relatively simple. However, in a multi-agent system, the behaviors of the agents are often interdependent, so multiple factors must be considered, such as the role of each agent, the task assignment, the cooperation mode, etc. The individual contribution as an intrinsic exploration scaffold guides the agent's exploration by evaluating each agent's contribution to the change of the global system state. In this way, the agent not only pays attention to its immediate reward, but also understands how its behavior affects the entire system in the long run, thereby significantly improving the agent's exploration ability in a sparse reward environment and accelerating the learning process in multi-agent collaborative tasks. Summary of the Invention
[0004] To address the problems existing in the above-mentioned prior art, the present invention trains multi-agent systems in both global and local environments. By calculating the information content obtained by the prediction model of the state after the agent explores a new state, the agent is guided to preferentially explore states with high uncertainty or low predictability. The information gain is combined with the exploration frequency factor as the intrinsic individual reward to measure the new information content obtained by the agent through exploring a state-action pair. If the prediction uncertainty of the model is significantly reduced due to exploration, then the state-action pair is worth exploring. This can further improve the exploration efficiency of the multi-agent system and enable the agent to learn in complex environments and unseen regions.
[0005] To achieve the above technical objectives, the present invention provides the following technical solutions:
[0006] A multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards, which specifically includes the following steps:
[0007] S1. Obtain the state s of the agent at the current moment t , the action a at the current moment t and the state s at the next moment t+1 , create a multi-agent reinforcement learning environment, and initialize the parameters of the dynamic model;
[0008] S2. Combine the generative adversarial network GAN to construct a conditional variational autoencoder CVAE. Use the CVAE as the dynamic model to compress the state space into the latent space and predict the state distribution of the agent at the next moment;
[0009] S3. Optimize the model parameters of the CVAE constructed in step S2 using the reconstruction loss and KL divergence loss to ensure the representation ability and distribution regularization of the latent space, and obtain the trained CVAE model;
[0010] S4. Use the trained CVAE model to estimate the prior entropy and posterior entropy of the state transition from the current moment to the next moment of the agent, and calculate the information gain;
[0011] S5. Design the exploration frequency factor, and combine the information gain obtained in step S4 to calculate the intrinsic individual reward, guiding the agent to preferentially select actions that reduce uncertainty;
[0012] S6. Generate the exploitation strategy and exploration strategy of the agent according to the intrinsic individual reward and the extrinsic reward. The exploitation strategy and the exploration strategy work together to guide the agent to select the action to be executed and complete the task objective.
[0013] Further, step S1 specifically includes:
[0014] S11. Create and configure the reinforcement learning environment, and obtain the shape of the observation space;
[0015] S12. Create the global baseline network and the local baseline network of the dynamic model; the global baseline network takes the global state in the multi-agent environment and the actions of all agents as inputs, and learns the potential laws of the overall state transition of the agents; while the local baseline network is trained only based on the local state and actions of the current agent, so as to independently predict the state transition of this agent;
[0016] S13. Collect the multi-step experience data of the agents, including the state s of each agent at the current moment t and the action a at the current moment t and the state s at the next moment t+1 ;
[0017] S14. Create and initialize the dynamic model, embed and normalize the high-dimensional state information and action information, and prepare for the training of the dynamic model.
[0018] Further, step S2 specifically includes:
[0019] S21. Input the concatenated vector of s t , a t , s t+1 . Use the multi-layer perceptron MLP as the encoder to gradually extract the features of the input data, and output the mean μ and standard deviation σ of the distribution parameters in the latent space; and then define that the posterior distribution in the latent space satisfies
[0020] S22. Sample the latent state z from the latent space generated by the encoder. To ensure that the sampling process is differentiable, use the reparameterization trick to sample the latent state which represents the potential law of the state transition from the current moment to the next moment;
[0021] S23. Use GAN to construct the decoder; GAN includes a generator and a discriminator; input the state s at the current moment t and the action a at the current moment t into the generator. The generator processes the input data through multiple fully connected layers and activation functions, extracts features from it, and learns the potential laws of state transition. The generator finally outputs the predicted state distribution at the next moment: where μ' is the mean of the predicted state distribution at the next moment, and σ' is the standard deviation of the predicted state distribution at the next moment; the discriminator then performs feedback optimization on the obtained predicted state distribution at the next moment, so as to finally realize the prediction of the decoder outputting a state distribution at the next moment that is close to the real situation.
[0022] Further, step S3 specifically includes:
[0023] S31. Design the reconstruction loss L reconstruction, so that the state distribution P(s t+1 |z, s t , a t ) predicted by the decoder is as close as possible to the true state s t+1 at the next moment, and the training process of the generator is optimized through the feedback of the discriminator to strengthen the simulation ability of the generator for the true state distribution; the reconstruction loss is defined as negative log-likelihood, and the formula is expressed as:
[0024]
[0025] where Q(z|s t , a t , s t+1 ) is the posterior distribution of the latent state z; E represents taking the expectation;
[0026] S32. Design the KL divergence loss L KL to make the posterior distribution Q(z|s t , a t , s t+1 ) of the latent state z close to its prior distribution P(z|s t , a t ), thereby regularizing the latent space distribution, preventing the model from overfitting, and ensuring that the structure of the latent space can fully capture the latent laws of state transitions; the formula expression of the KL divergence loss is:
[0027] L KL = D KL [Q(z|s t , a t , s t+1 ) || P(s t |a t )];
[0028] where D KL (.) is the KL divergence calculation formula;
[0029] Substituting into the normal distribution again, the formula expression of the KL divergence loss is obtained as:
[0030]
[0031] where μ i , σ i are the mean and standard deviation of the posterior distribution of the latent state z;
[0032] S33. Define the total loss L CVAE of the CVAE model as the weighted sum of the reconstruction loss and the KL divergence, and the formula is expressed as:
[0033] L CVAE = Lreconstruction +γL KL ;
[0034] where γ is a balance parameter used to adjust the weight between the reconstruction loss and the KL divergence;
[0035] S34. Minimize the total loss L by gradient descent CVAE to optimize the model parameters of the CVAE, and finally obtain a trained CVAE model.
[0036] Further, step S4 specifically includes:
[0037] S41. Use the global baseline network to estimate the overall state transition of the agent according to the state s at the current moment t ; the global limit network considers the dynamics of the entire multi-agent reinforcement learning environment and generates a global state transition prediction that does not depend on the actions of specific agents; then the prior network in the CVAE infers the prior distribution P(s t |s t+1 ) of the state transition of the dynamic model further based on the state s at the current moment t and the output of the global baseline network;
[0038] S42. Combine the local baseline network and the posterior network to estimate a more accurate posterior distribution of the state transition; the local baseline network combines the current state s t and the action a of the agent t to evaluate the specific impact of this action on the state transition, and the posterior network updates the estimation of the state transition based on the output of the local baseline network and the current state; finally, obtain the posterior distribution P(s t+1 |s t ,a t ) of the state transition of the dynamic model, that is, the state change that occurs after the agent takes the action a at the current moment t ;
[0039] S43. Calculate the prior entropy H(s t+1 |s t ) and the posterior entropy H(s t+1 |s t ,a t ) respectively according to P(s t+1 |s t ) and P(s t+1 |s t ,a t ) obtained in steps S41 and S42. The formula is expressed as:
[0040]
[0041] where \(e\) is the natural constant; \(d\) is the total number of dimensions of the agent state, and each dimension represents a state in the latent space; \(\sigma\) prior,i represents the standard deviation of the prior distribution of the state transition in the \(j\)-th dimension; \(\sigma\) posterior,j represents the standard deviation of the posterior distribution of the state transition in the \(j\)-th dimension;
[0042] S44. Obtain the information gain definition based on the prior entropy and posterior entropy calculated in step S43, and the formula is expressed as:
[0043] IG(s t ,a t ) = H(s t+1 |s t ) - H(s t+1 |s t ,a t );
[0044] Use the information gain IG(s t ,a t ) to represent the contribution of the current action \(a\) t in reducing the uncertainty of the state at the next moment. The larger the value, the more conducive \(a\) t is to exploring unknown states.
[0045] Furthermore, step S5 specifically includes:
[0046] S51. Based on the exploration frequency of the state-action pair \((s t ,a t ) during the training process, introduce the exploration frequency factor \(\omega(s t ,a t ) to promote the agent to explore less visited state-action pairs; the calculation formula of the exploration frequency factor is:
[0047]
[0048] where \(N(s t ,a t ) is the historical access frequency of the state-action pair \((s t ,a t ), and \(\delta\) is a hyperparameter that controls the degree of penalty for the historical access frequency;
[0049] S52. Combine the information gain and the exploration frequency factor to define the intrinsic individual reward, and the formula is expressed as:
[0050] r int (s t ,a t ) = \(\omega(s t ,a t ) · IG(s t ,at );
[0051] where r int (s t , a t ) is the intrinsic individual reward, and IG(s t , a t ) is the information gain;
[0052] S53. Normalize the intrinsic individual reward and map it to the interval [0, 1], and use the normalized intrinsic individual reward as the selection index for the agent's exploration strategy; the normalization is expressed by the formula:
[0053]
[0054] where max(r int ) and min(r int ) are the maximum and minimum values of the intrinsic individual rewards corresponding to all state-action pairs of the agent respectively; r int '(s t , a t ) is the normalized intrinsic individual reward.
[0055] Furthermore, step S6 specifically includes:
[0056] S61. Set the exploitation strategy The exploitation strategy is the strategy used by the agent to complete the task. It selects the action to be executed by evaluating the extrinsic rewards corresponding to each known action, that is, it selects the action that maximizes the extrinsic reward. π i represents the exploitation strategy of the i-th agent;
[0057] S62. Set the exploration strategy The exploration strategy selects the action to be executed by evaluating the intrinsic individual rewards corresponding to each unknown action, and preferentially selects the action with a large intrinsic individual reward, that is, the action that can significantly reduce the environmental uncertainty; v i represents the exploration strategy of the i-th agent;
[0058] S63. The action selection process of the agent's behavior strategy b i is That is, the final action u i is sampled from the behavior strategy, and the behavior strategy will balance between the exploration strategy and the exploitation strategy. The formula is expressed as:
[0059]
[0060] where represents sampling through the exploration strategy v i and preferentially selecting the action with a large intrinsic individual reward, It is represented by the target policy π i Sampling, preferentially selecting actions with high extrinsic rewards, where α is a hyperparameter that controls the balance between the exploration strategy and the exploitation strategy;
[0061] S64. After each exploration action is executed, the exploration strategy is further optimized based on the state and the motion trajectory; the optimization objective of the exploration strategy is set to maximize the intrinsic individual reward and the policy entropy, and the formula is expressed as:
[0062]
[0063] Among them, J i (ξ) is the objective function; ξ is the parameter of the exploration strategy, is the intrinsic individual reward under the exploration strategy v i ; E represents taking the expectation; H(·|τ i , s) represents the entropy of the exploration strategy, τ i and s represent the given motion trajectory and state, which are used to encourage the diversity of the exploration strategy, and β is a hyperparameter that controls the weight of entropy regularization.
[0064] Based on the above technical solutions, the method proposed by the present invention has at least the following beneficial effects:
[0065] The method proposed by the present invention introduces an intrinsic individual reward based on information gain and exploration frequency factor to help the intelligent agent better understand the impact of its actions on the environment, focus on valuable action selection during its learning process, and avoid ineffective exploration. At the same time, by combining extrinsic rewards and intrinsic individual rewards, it balances the exploitation strategy and the exploration strategy, thereby significantly improving the learning efficiency and decision-making ability of the intelligent agent.
[0066] The method proposed by the present invention shows good adaptability in different environments; the collaborative exploration mechanism set according to the intrinsic individual reward and the external reward enables the intelligent agent to quickly adapt and achieve effective strategy learning and optimization in different complex environments, improving the generality and robustness of the model; by setting the intrinsic individual reward, it encourages the intelligent agent to conduct more exploration during the training process, optimizes the reward sparsity problem in complex environments, enables the intelligent agent to obtain effective strategies faster, and avoids falling into local optimal solutions, thereby improving the training efficiency and the performance of the final model.
[0067] The method proposed by the present invention effectively solves the problems of low exploration efficiency, long training time, and poor adaptability in traditional methods, provides an efficient, flexible, and highly adaptable exploration method, and provides strong support for the application of intelligent agents in complex environments. Brief Description of the Drawings
[0068] Figure 1 is the overall flowchart of the method proposed by the present invention;
[0069] Figure 2 Schematic diagram of the experimental environment 2s3z in the embodiments of the present invention;
[0070] Figure 3 Schematic diagram of the experimental environment 3-vs-1 with Keeper in the embodiments of the present invention;
[0071] Figure 4 Test winning rate curve graph under the experimental environment 2s3z;
[0072] Figure 5 Training average score curve graph under the experimental environment 3-vs-1 with Keeper. Specific implementation manners
[0073] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying Figures 1-5 drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0074] Although the steps in the present invention are arranged with reference numerals, they are not used to limit the order of the steps. Unless the order of the steps is clearly stated or the execution of a certain step requires other steps as a basis, the relative order of the steps can be adjusted. It can be understood that the term "and / or" used herein relates to and encompasses any and all possible combinations of one or more of the associated listed items.
[0075] As Figure 1 shown, the multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards proposed by the present invention learns the environmental state transition law through a dynamic model (such as the conditional variational autoencoder CVAE), calculates the contribution of an action to reducing the uncertainty of the environmental state, so as to achieve efficient exploration. The present invention first trains the CVAE to learn the prior distribution and posterior distribution of the environmental state, and calculates the information gain through the difference between the two to measure the exploration value of each action. An intrinsic individual reward is generated by combining the information gain with the historical access frequency, and the behavior decision of the agent is guided by combining the extrinsic reward, thereby greatly improving the exploration efficiency.
[0076] The method proposed by the present invention specifically includes the following steps:
[0077] S1. Obtain the state s of the agent at the current moment t , the action a at the current moment t and the state s at the next moment t+1 , create a multi-agent reinforcement learning environment, and initialize the parameters of the dynamic model;
[0078] As a preferred embodiment, step S1 specifically includes:
[0079] S11. Create and configure a reinforcement learning environment, and obtain the shape of the observation space;
[0080] S12. Create a global baseline network and a local baseline network for the dynamic model; the global baseline network takes the global state in the multi-agent environment and the actions of all agents as inputs, and learns the potential laws of the overall state transition of the agents; while the local baseline network is trained only based on the local state and actions of the current agent, so as to independently predict the state transition of the agent;
[0081] S13. Collect multi-step experience data of the agents, including the state s of each agent at the current moment t and the action a at the current moment t and the state s at the next moment t+1 ;
[0082] S14. Create and initialize the dynamic model, embed and normalize the high-dimensional state information and action information, and prepare for the training of the dynamic model.
[0083] S2. Combine the generative adversarial network GAN to construct a conditional variational autoencoder CVAE, use the CVAE as the dynamic model, compress the state space into the latent space, and predict the state distribution of the agent at the next moment;
[0084] As a preferred embodiment, step S2 specifically includes:
[0085] S21. Input the concatenated vector of s t , a t , s t+1 , use a multi-layer perceptron MLP as the encoder, gradually extract the features of the input data, and output the mean μ and standard deviation σ of the distribution parameters in the latent space; furthermore, define the posterior distribution in the latent space to satisfy The existence of the latent space maps the high-dimensional input data to a low-dimensional space, making the input data more compact while retaining important features;
[0086] S22. Sample the latent state z from the latent space generated by the encoder. To ensure that the sampling process is differentiable, use the reparameterization trick to sample the latent state which represents the potential law of the state transition from the current moment to the next moment;
[0087] S23. Use GAN to construct a decoder; GAN includes a generator and a discriminator; input the state s at the current moment t and the action a at the current moment tInput into the generator, which processes the input data through multiple fully connected layers and activation functions, extracts features from it, and learns the potential laws of state transition. The generator finally outputs the predicted state distribution at the next moment: where μ' is the mean of the predicted state distribution at the next moment, and σ' is the standard deviation of the predicted state distribution at the next moment; the discriminator then provides feedback optimization on the obtained predicted state distribution at the next moment to ultimately achieve the prediction of the state distribution at the next moment where the decoder output is close to the true situation.
[0088] S3. Use the reconstruction loss and KL divergence loss to optimize the model parameters of the CVAE constructed in step S2, ensure the representation ability of the latent space and distribution regularization, and obtain the trained CVAE model;
[0089] In this embodiment, the optimization process is carried out in the discriminator of the GAN model. The discriminator will evaluate the difference between the state distribution generated by the generator and the true state, and give a feedback signal indicating the direction that the generator needs to optimize; the generator will adjust the parameters according to the output of the discriminator, gradually generate a prediction result closer to the true state, and strengthen the simulation ability of the generator for the true state distribution; in this application, the reconstruction loss and KL divergence loss are used as the evaluation criteria for optimization; thus, as a preferred implementation manner, step S3 specifically includes:
[0090] S31. Design the reconstruction loss L reconstruction , such that the state distribution P(s t+1 |z, s t , a t ) predicted by the decoder at the next moment is as close as possible to the true state s t+1 at the next moment, and optimize the training process of the generator through the feedback of the discriminator to strengthen the simulation ability of the generator for the true state distribution; the reconstruction loss is defined as negative log-likelihood, and the formula is expressed as:
[0091]
[0092] where Q(z|s t , a t , s t+1 ) is the posterior distribution of the latent state z; E represents taking the expectation;
[0093] S32. Design the KL divergence loss L KL , so that the posterior distribution Q(z|s t , a t , s t+1 ) of the latent state z is close to its prior distribution P(z|s t , a t), thus regularizing the potential space distribution, preventing the model from overfitting, and ensuring that the structure of the potential space can fully capture the potential laws of state transitions; the formula of the KL divergence loss is expressed as:
[0094] L KL = D KL [Q(z|s t , a t , s t+1 )||P(s t |a t )];
[0095] Among them, D KL (.) is the KL divergence calculation formula;
[0096] Substituting the normal distribution again, the formula representation of the KL divergence loss is:
[0097]
[0098] Among them, μ i , σ i are the mean and standard deviation of the posterior distribution of the latent state z;
[0099] S33. Define the total loss L CVAE of the CVAE model as the weighted sum of the reconstruction loss and the KL divergence, and the formula is expressed as:
[0100] I CVAE = L reconstruction + γL KL ;
[0101] Among them, γ is a balance parameter used to adjust the weights between the reconstruction loss and the KL divergence; flexibly adjust the balance between the generation task and the potential space regularization, so that the model can not only generate accurate samples but also have a good generation space structure;
[0102] S34. Minimize the total loss L CVAE by the gradient descent method to optimize the model parameters of the CVAE, and finally obtain the trained CVAE model.
[0103] S4. Use the trained CVAE model. The trained CVAE model not only has good representation ability but also has regularized distribution characteristics; use it to estimate the prior entropy and posterior entropy of the state transition of the agent from the current moment to the next moment, and calculate the information gain;
[0104] As a preferred implementation manner, step S4 specifically includes:
[0105] S41. Use the global baseline network according to the state s tto estimate the state transition of the agent as a whole; the global limit network considers the dynamics of the entire multi-agent reinforcement learning environment and generates global state transition predictions that are independent of the actions of specific agents; and then the prior network in CVAE is based on the current state s t and the output of the global baseline network, and further infer the prior distribution P(s) of the state transition of the dynamic model t+1 |s t ); This prior distribution reflects the dynamic changes inherent in the environment and helps the agent understand how the environment evolves without specific actions;
[0106] S42, combine the local baseline network and the posterior network to estimate a more accurate posterior distribution of state transitions; the local baseline network will combine the current state s t and the agent's action a t To evaluate the specific impact of the action on the state transition, the posterior network updates the estimate of the state transition based on the output of the local baseline network and the current state; finally, the posterior distribution P(s t+1 |s t ,a t ), that is, the agent takes action a at the current moment t The state change that occurs later;
[0107] S43, P(s) obtained according to steps S41 and S42 t+1 |s t )、P(s t+1 |s t ,a t ) respectively calculate the prior entropy H(s t+1 |s t ) and posterior entropy H(s t+1 |s t ,a t ), the formula is:
[0108]
[0109] Where e is a natural constant; d is the total number of dimensions of the agent state, each dimension represents a state in the latent space; σ prior,i represents the standard deviation of the prior distribution of the j-dimensional state transition; σ posterior,j Represents the standard deviation of the posterior distribution of the j-dimensional state transition;
[0110] S44, according to the priori entropy and posterior entropy calculated in step S43, the definition of information gain is obtained, and the formula is expressed as:
[0111] IG(s t ,a t )=H(s t+1 |st ) - H(s t+1 |s t , a t );
[0112] Use the information gain IG(s t , a t ) to represent the contribution of the action a at the current moment t in reducing the uncertainty of the state at the next moment. The larger the value, the more beneficial a t is for exploring unknown states, thus guiding the agent to conduct effective exploration.
[0113] S5. Design an exploration frequency factor, and calculate the intrinsic individual reward by combining the information gain obtained in step S4 to guide the agent to preferentially select actions that reduce uncertainty;
[0114] As a preferred implementation method, step S5 specifically includes:
[0115] S51. Based on the exploration frequency of the state - action pair (s t , a t ) during the training process, introduce the exploration frequency factor ω(s t , a t ) to promote the agent to explore state - action pairs that are less frequently visited; The calculation formula of the exploration frequency factor is:
[0116]
[0117] where N(s t , a t ) is the historical access frequency of the state - action pair (s t , a t ), and δ is a hyperparameter that controls the degree of penalty for the historical access frequency;
[0118] S52. Define the intrinsic individual reward by combining the information gain and the exploration frequency factor, and the formula is expressed as:
[0119] r int (s t , a t ) = ω(s t , a t ) · IG(s t , a t );
[0120] where r int (s t , a t ) is the intrinsic individual reward, and IG(s t , a t) is the information gain; in this way, the state-action pairs with high-frequency access will obtain smaller rewards, while the state-action pairs with low-frequency access will obtain larger rewards
[0121] S53. Since the range of information gain fluctuates greatly under different tasks and different scales, it is necessary to normalize the intrinsic individual rewards and map them to the [0,1] interval, and use the normalized intrinsic individual rewards as the selection index for the agent's exploration strategy; the normalization is expressed by the formula:
[0122]
[0123] where max(r int ), min(r int ) are respectively the maximum and minimum values of the intrinsic individual rewards corresponding to all state-action pairs of the agent; r int '(s t ,a t ) is the normalized intrinsic individual reward
[0124] S6. Generate the exploitation strategy and exploration strategy of the agent according to the intrinsic individual reward and the extrinsic reward. The exploitation strategy and the exploration strategy act synergistically to guide the agent to select the action to be executed and complete the task objective;
[0125] As a preferred embodiment, step S6 specifically includes:
[0126] S61. Set the exploitation strategy The exploitation strategy is the strategy for the agent to complete the task. It selects the action to be executed by evaluating the extrinsic rewards corresponding to each known action, that is, selects the action that maximizes the extrinsic reward, and π i represents the exploitation strategy of the i-th agent;
[0127] In this embodiment, the QMIX algorithm is used to calculate the extrinsic reward, specifically:
[0128] S611: Each agent i calculates its expected return after selecting the action a t in the current state s t , that is, the dynamic-action value function Q i (s t ,a t ). This value is usually estimated through historical experience or the current strategy;
[0129] S612: Combine the state-action value function Q i (s t ,a t ) of each agent into a global state-action value function Q tot (s t ,at ), and use it as an extrinsic reward; Q tot (s t , a t ) is obtained by weighted summation, and the formula is expressed as:
[0130]
[0131] where m i is the weight of each agent, which is used to reflect the importance or contribution of each agent in the system;
[0132] 613: Each agent i selects the optimal action in the exploitation strategy according to the extrinsic reward to maximize the overall return, and the formula is expressed as:
[0133]
[0134] S62. Set the exploration strategy The exploration strategy selects the action to be executed by evaluating the intrinsic individual reward corresponding to each unknown action, and preferentially selects the action with a large intrinsic individual reward, that is, the action that can significantly reduce the environmental uncertainty; v i represents the exploration strategy of the i-th agent;
[0135] S63. The action selection process of the agent's behavior strategy b i is That is, the final action u i is sampled from the behavior strategy, and the behavior strategy will trade off between the exploration strategy and the exploitation strategy. The formula is expressed as:
[0136]
[0137] where, represents sampling through the exploration strategy v i and preferentially selects the action with a large intrinsic individual reward, represents sampling through the target strategy π i and preferentially selects the action with a high extrinsic reward. α is a hyperparameter that controls the balance between the exploration strategy and the exploitation strategy;
[0138] S64. After each exploration action is executed, optimize the exploration strategy according to the state and motion trajectory; set the optimization goal of the exploration strategy to maximize the intrinsic individual reward and the policy entropy. The formula is expressed as:
[0139]
[0140] where J i (ξ) is the objective function; ξ is the parameter of the exploration strategy, and its randomness or diversity of the exploration strategy is adjusted through entropy regularization; To explore the strategy v i The intrinsic individual reward under the condition; E represents the expectation; H(·|τ i ,s) represents the entropy of the exploration strategy, τ i and s represent the given motion trajectory and state, which are used to encourage exploration strategy diversity. β is a hyperparameter that controls the weight of entropy regularization.
[0141] So far, the multi-agent collaborative exploration method proposed in the present invention constructs intrinsic individual rewards based on information gain, and constructs a collaborative exploration mechanism of agents by combining intrinsic individual rewards with extrinsic rewards (that is, considering both utilization strategies and exploration strategies, and taking into account both known environments and unknown environments), which is more conducive to the exploration of agents; the effect of the method proposed in the present invention is verified in combination with specific experimental examples.
[0142] Experimental Examples
[0143] This experimental example sets up two different experimental environments, such as Figure 2 , 3 shown. Figure 2 The StarCraft Multi-Agent Challenge (SMAC) environment is a multi-agent reinforcement learning environment developed by Blizzard and DeepMind based on the real-time strategy game StarCraft II. The SMAC environment contains a total of 23 maps, such as Figure 2 Shown is one of the Figure 2 Actual screen of s3z. This map is a heterogeneous symmetrical map. Both the enemy and our side are composed of two types of soldiers, two Stalkers and three Zealots. Each combat unit is controlled by an independent agent, which can only be trained according to the limited vision of the combat unit, while the opponent's units are uniformly controlled by hand-coded built-in AI. There are ten levels of built-in AI, consisting of gradually enhanced levels 1 to 9 and the strongest level A. This is a confrontation between two troops, and the model needs to learn one or more strategies to defeat the enemy.
[0144] Figure 3The Google Research Football (GRF) shown provides a football game environment released by Google. It offers a physics-based 3D football simulation scenario where agents control players, learn how to pass and cooperate among players, and learn how to break through the opponent's defense to score. This is a very challenging reinforcement learning task because in the football game scenario, agents need to maintain the connection and balance between immediate control (such as learned concepts like passing and dribbling) and high-level strategies. These scenarios are diverse. Some scenarios may require the agent to score against an empty goal, while in some scenarios, the agent must cooperate with teammates to break through the opponent's defensive line to score. The 3-vs-1 with Keeper environment simulates a 3-on-1 football game, with 3 attacking agents (attacking players) and 1 defending agent (goalkeeper). The goal of the attacking side is to attack through passing, cooperation, and shooting, while the goal of the defending side is to prevent the attacking side from shooting as much as possible and defend their own goal.
[0145] In this experimental example, a sparse reward setting is used in the SMAC and GRF environments. When all enemy forces in 2s3z are defeated, the reward value is +200, which is the ultimate reward for completing the task, encouraging agents to cooperate to eliminate all enemy forces; when one enemy is defeated, the reward value is +10, used to motivate the behavior of gradually killing enemy forces; when one friendly force is defeated, the reward value is -5, a punishment mechanism to encourage protecting friendly forces. In 3-vs-1 with Keeper, when our side scores, the reward value is +100, which is the target reward for the task, motivating agents to cooperate to achieve the final goal; when the opponent scores, the reward value is -1, and when the ball or our player returns to our own half, the reward value is -1, punishment signals to prevent agents from ignoring defensive strategies and choosing to stall or engage in ineffective exploration behaviors. Here, rewards are only given when key events occur, and there is no reward signal in the intermediate process, increasing the difficulty of learning and requiring agents to explore and collaborate more efficiently. Set the parameters of the deep reinforcement learning model as follows: the action embedding dimension is 4; the learning rate of the intrinsic reward framework is 0.0001; the Adam optimizer is used as the optimizer; the gradient clipping value is 0.1; the learning rate of the exploration network for the 2s3z task is set to 0.01, and for the 3-vs-1 with Keeper task is set to 0.001; the exploration-exploitation balance coefficient α is set to 0.1 in the 2s3z task, and set to 0.2 at the beginning of the 3-vs-1 with Keeper task training and gradually reduced to 0.05 as the training progresses; the intrinsic reward entropy weight β is set to 0.1 in the 2s3z task and 0.05 in the 3-vs-1 with Keeper task.
[0146] The traditional QMIX algorithm and the method proposed in the present invention are respectively used to train the agents, and the training results obtained after training are asFigure 4 , Figure 5 as shown in the figure. The abscissa in the figure represents the number of training steps, and the ordinate represents the test winning rate and the average training score respectively. The dashed line is the method proposed by the present invention, and the solid line is the QMIX algorithm. From Figure 4 , it can be seen that when the number of steps reaches 0.5M (500,000) steps, the test winning rate of the method proposed by the present invention begins to increase significantly, and starts to converge at 1M steps, and maintains the highest team winning rate until the training ends. The curve of the traditional QMIX algorithm starts to rise at 1M steps and is relatively slow. Figure 5 It shows that the average score of the present invention begins to rise at about 1M steps and begins to gradually converge at 2M steps. From the perspective of the animation effect of the environment rendering, when the method proposed by the present invention reaches convergence, the paths taken by the agents are all the shortest paths, while the task completion rate of the traditional QMIX algorithm is very low. In summary, the present invention greatly improves the convergence speed, enhances the exploration rate of the environment, and has a good effect on exploration.
[0147] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0148] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards, characterized by: The specific steps include: S1. Get the current state s of the agent t , the action at the current moment a t and the state s at the next moment t+1 , create a multi-agent reinforcement learning environment and initialize the parameters of the dynamic model; S2. Combine the generative adversarial network GAN to construct the conditional variational autoencoder CVAE. Use CVAE as a dynamic model to compress the state space into the latent space and predict the state distribution of the agent at the next moment. S3, using reconstruction loss and KL divergence loss to optimize the model parameters of the CVAE constructed in step S2, ensuring the representation ability and distribution regularization of the latent space, and obtaining a trained CVAE model; S4. Use the trained CVAE model to estimate the prior entropy and posterior entropy of the state transition of the agent from the current moment to the next moment, and calculate the information gain; S5. Design an exploration frequency factor and calculate the intrinsic individual reward based on the information gain obtained in step S4 to guide the agent to prioritize actions that reduce uncertainty. S6. Generate the agent’s utilization strategy and exploration strategy based on intrinsic individual rewards and extrinsic rewards. The utilization strategy and exploration strategy work together to guide the agent to select actions to perform and complete the task objectives.
2. The multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards according to claim 1, characterized in that: Step S1 specifically includes: S11. Create and configure a reinforcement learning environment to obtain the shape of the observation space; S12. Create a global baseline network and a local baseline network of the dynamic model; the global baseline network takes the global state and the actions of all agents in the multi-agent environment as input to learn the underlying laws of the state transition of the agent as a whole; while the local baseline network is trained only based on the local state and action of the current agent, so as to independently predict the state transition of the agent; S13, collect multi-step experience data of the agent, including the current state s of each agent t , the action at the current moment a t , the state s at the next moment t+1 ; S14. Create and initialize a dynamic model, embed and normalize high-dimensional state information and action information, and prepare for training the dynamic model.
3. The multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards according to claim 1, characterized in that: Step S2 specifically includes: S21, input s t 、a t 、s t+1 The concatenated vector is used as the encoder to gradually extract the features of the input data and output the mean μ and standard deviation σ of the distribution parameters of the latent space; then the posterior distribution of the latent space is defined to satisfy S22, sample the potential state z from the latent space generated by the encoder. To ensure that the sampling process is differentiable, a reparameterization technique is used to sample the potential state z = μ + σ · ∈, It represents the potential law of state transition from the current moment to the next moment; S23. Use GAN to build a decoder. GAN includes a generator and a discriminator. Input the current state s t and the current action a t The data is input to the generator, which processes the input data through multiple fully connected layers and activation functions, extracts features from it, and learns the potential rules of state transition. The generator finally outputs the predicted state distribution at the next moment: Where μ' is the mean of the predicted state distribution at the next moment, and σ' is the standard deviation of the predicted state distribution at the next moment; the discriminator then performs feedback optimization on the predicted state distribution at the next moment, so as to ultimately achieve a prediction of the state distribution at the next moment that is close to the actual situation when the decoder outputs the output.
4. The multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards according to claim 3 is characterized in that: Step S3 specifically includes: S31, design reconstruction loss L reconstruction , so that the decoder predicts the state distribution P(s) at the next moment t+1 |z,s t ,a t ) is as close as possible to the actual state s of the next moment t+1 , the training process of the generator is optimized through the feedback of the discriminator, and the generator's simulation ability for the real state distribution is strengthened; the reconstruction loss is defined as the negative log-likelihood, and the formula is expressed as: Among them, Q(z|s t ,a t ,s t+1 ) is the posterior distribution of the potential state z; E represents the expectation; S32. Design KL divergence loss L KL , so that the posterior distribution Q(z|s t ,a t ,s t+1 ) is close to its prior distribution P(z|s t ,a t ), thereby regularizing the potential space distribution, preventing the model from overfitting, and ensuring that the structure of the potential space can fully capture the potential laws of state transition; the formula for KL divergence loss is expressed as: L KL =D KL [Q(z|s t ,a t ,s t+1 )||P(s t |a t )]; Among them, D KL (.) is the KL divergence calculation formula; Then bring in the normal distribution, and the formula for KL divergence loss is expressed as: Among them, μ i , σ i is the mean and standard deviation of the posterior distribution of the latent state z; S33. Define the total loss L of the CVAE model CVAE It is the weighted sum of reconstruction loss and KL divergence, and the formula is expressed as: THE CVAE =L reconstruction +γL KL ; Among them, γ is a balance parameter used to adjust the weight between reconstruction loss and KL divergence; S34. Minimize the total loss L by gradient descent method CVAE , optimize the model parameters of CVAE, and finally obtain the trained CVAE model.
5. The multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards according to claim 2, characterized in that: Step S4 specifically includes: S41, use the global baseline network according to the current state s t to estimate the state transition of the agent as a whole; the global limit network considers the dynamics of the entire multi-agent reinforcement learning environment and generates global state transition predictions that are independent of the actions of specific agents; and then the prior network in CVAE is based on the current state s t and the output of the global baseline network, and further infer the prior distribution P(s) of the state transition of the dynamic model t+1 |s t ); S42, combine the local baseline network and the posterior network to estimate a more accurate posterior distribution of state transitions; the local baseline network will combine the current state s t and the agent's action a t To evaluate the specific impact of the action on the state transition, the posterior network updates the estimate of the state transition based on the output of the local baseline network and the current state; finally, the posterior distribution P(s t+1 |s t ,a t ), that is, the agent takes action a at the current moment t The state change that occurs later; S43, P(s) obtained according to steps S41 and S42 t+1 |s t )、P(s t+1 |s t ,a t ) respectively calculate the prior entropy H(s t+1 |s t ) and posterior entropy H(s t+1 |s t ,a t ), the formula is: Where e is a natural constant; d is the total number of dimensions of the agent state, each dimension represents a state in the latent space; σ prior,i represents the standard deviation of the prior distribution of the j-dimensional state transition; σ posterior,j Represents the standard deviation of the posterior distribution of the j-dimensional state transition; S44, according to the priori entropy and posterior entropy calculated in step S43, the definition of information gain is obtained, and the formula is expressed as: IG(s t ,a t )=H(s t+1 |s t )-H(s t+1 |s t ,a t ); Using information gain IG(s t ,a t ) represents the action a at the current moment t The contribution to reducing the uncertainty of the state at the next moment, the larger the value, the greater the t Helpful for exploring unknown states.
6. The multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards according to claim 1, characterized in that: Step S5 specifically includes: S51, based on state-action pair (s t ,a t ) is the exploration frequency during the training process, introducing the exploration frequency factor ω(s t ,a t ) to encourage the agent to explore less visited state-action pairs; the calculation formula of the exploration frequency factor is: Among them, N(s t ,a t ) is a state-action pair (s t ,a t )’s historical access frequency, δ is a hyperparameter that controls the degree of penalty for historical access frequency; S52. Define the intrinsic individual reward by combining information gain and exploration frequency factor. The formula is: r int (and t ,the t )=ω(s t ,the t )·IG(s t ,the t ); Among them, r int (s t ,a t ) is the intrinsic individual reward, IG(s t ,a t ) is the information gain; S53, normalize the intrinsic individual reward and map it to the interval [0,1], and use the normalized intrinsic individual reward as the selection index of the intelligent agent's exploration strategy; the normalization formula is expressed as: Among them, max(r int )、min(r int ) are the maximum and minimum values of the corresponding intrinsic individual rewards of all state-action pairs of the agent; r int '(s t ,a t ) is the normalized intrinsic individual reward.
7. The multi-agent collaborative exploration method based on information gain to construct intrinsic individual rewards according to claim 6, characterized in that: Step S6 is specifically as follows: S61. Set up utilization strategy As the strategy used by the agent to complete the task, it selects the action to be performed by evaluating the extrinsic reward corresponding to each known action, that is, it selects the action that maximizes the extrinsic reward, π i represents the utilization strategy of the i-th agent; S62. Set exploration strategy The exploration strategy selects the action to be performed by evaluating the intrinsic individual reward corresponding to each unknown action, giving priority to actions with large intrinsic individual rewards, that is, actions that can significantly reduce environmental uncertainty; i represents the exploration strategy of the i-th agent; S63, Agent's Behavior Strategy b i The action selection process is That is, the final action u i It is sampled from the behavior strategy, which makes a trade-off between the exploration strategy and the exploitation strategy. The formula is: in, Represents that through the exploration strategy v i Sampling, giving priority to actions with large intrinsic individual rewards, Represents the target strategy π i Sampling, prioritizes actions with high extrinsic rewards. α is a hyperparameter that controls the balance between exploration strategy and exploitation strategy. S64, after each exploration action is performed, the exploration strategy is optimized according to the state and motion trajectory; the optimization goal of the exploration strategy is set to maximize the intrinsic individual reward and strategy entropy, and the formula is expressed as: Among them, J i (ξ) is the objective function; ξ is the parameter of the exploration strategy, To explore the strategy v i The intrinsic individual reward under the condition; E represents the expectation; H(·|τ i ,s) represents the entropy of the exploration strategy, τ i and s represent the given motion trajectory and state, which are used to encourage exploration strategy diversity. β is a hyperparameter that controls the weight of entropy regularization.
Citation Information
Patent Citations
Defense resource allocation method based on multi-agent collaborative behavior prediction
CN118153995A
Behavior modal division method, and training method and reasoning method of multi-modal trajectory prediction model
CN118569382A
Intelligent agent collaborative exploration method
CN118707973A
Multi-agent experience exploration cooperation method based on curiosity mechanism
CN119150914A
Multi-agent collaborative data collection learning method based on alliance formation game
CN119421202A
Cited By
Multi-agent collaborative decision-making method based on teammate representation and incentive communication
CN120337980A
Multi-unmanned aerial vehicle cooperative search and rescue intelligent decision-making method and device based on hierarchical intention
CN120672085A
Hierarchical intention-based multi-uav cooperative search and rescue intelligent decision method and device
CN120672085B