Multi-agent reinforcement learning method and system based on evolutionary curriculum learning

Through the strategies and value networks of evolutionary course learning and self-attention mechanism optimization, the problem of mismatch in knowledge transfer in multi-agent systems is solved, and the training efficiency and decision-making performance and adaptability of the agent are improved.

CN120031100BActive Publication Date: 2025-09-02ROCKET FORCE UNIV OF ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510495200.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-09-02
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing multi-agent reinforcement learning methods have target mismatch in the process of knowledge transfer between new and old agents, resulting in unstable performance of agents in large-scale systems and large calculation volume, which affects training efficiency.

Method used

Using an evolutionary course learning method, we perform evolutionary operations on some initial populations with the best adaptability in the multiagent system, combining the strategies and value network of the self-attention mechanism, dynamically handle the changes in the number of agents, and reload and reuse of models and experiences.

Benefits of technology

It significantly improves the training efficiency of multi-agent systems, improves the decision-making performance and adaptability of agents in different environments, simplifies the training process, reduces redundant competition, and improves the efficiency of knowledge transfer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031100B_ABST
    Figure CN120031100B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-agent reinforcement learning method and system based on evolutionary course learning, relating to the field of multi-agent decision-making technology. The method comprises the following steps: First, during each course learning phase, multiple populations of agents are trained in parallel to generate an initial multi-agent population for each role; then, an evolutionary population selection process is performed on the initial population to select the optimal population for training the next course learning phase. This step is repeated until the set number of agents is reached, at which point the training process ends. This method effectively addresses the poor adaptability of knowledge transfer in traditional course learning and improves the performance of traditional course learning. Furthermore, during the optimal population selection process, a balance is achieved between algorithm training efficiency and performance by rationally simplifying the population evolution process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multi-agent decision-making technology, and specifically relates to a multi-agent reinforcement learning method and system based on evolutionary curriculum learning. Background Art

[0002] With the deepening development of artificial intelligence, AI is gradually evolving from perception, cognition, and intelligence to decision-making intelligence. Since most real-world problems involve interactions between multiple agents, particularly in areas like multi-robot control, autonomous driving, and traffic signal control, collaborative decision-making intelligence based on multiple agents has emerged. However, the increasing number of agents complicates the interaction environment and inter-agent collaboration, making it particularly difficult to learn optimal strategies in multi-agent game-playing tasks.

[0003] To fully leverage the collaborative performance of multiple agents and improve the scalability of multi-agent reinforcement learning (MARL), existing research has adopted a curriculum-based training approach, gradually increasing the number of agents to enhance the algorithm's learning capabilities. However, this approach can lead to a mismatch in the knowledge transfer process between new and old agents, causing the current top-performing agent to experience unstable performance in a larger-scale agent system and fail to maintain its peak performance. This means that the current top-performing agent may not necessarily maintain its peak performance in a larger multi-agent system (MAS).

[0004] To effectively select agents with higher fitness and equip new agents with better initial decision-making capabilities, some researchers have introduced evolutionary thinking to address the goal mismatch problem. These methods incorporate selection, crossover, and mutation into course learning, selecting the seed with the highest fitness as the knowledge transfer target for the new agent in the next course. However, this approach inevitably requires a large amount of computation, severely impacting the algorithm's training efficiency. Summary of the Invention

[0005] The embodiments of the present invention provide a multi-agent reinforcement learning method and an agent system based on evolutionary curriculum learning, which can solve the problem of low efficiency of reinforcement learning in current multi-agent systems.

[0006] In a first aspect, an embodiment of the present invention provides a multi-agent reinforcement learning method based on evolutionary curriculum learning. The method is applied to a multi-agent reinforcement learning system based on evolutionary curriculum learning, wherein the system includes agents with multiple roles and the number of agents is not constant. The method includes:

[0007] For the first The agents in the system are trained multiple times in the learning phase to generate multiple The initial population of the learning phase, where the first The number of agents in the initial population of each learning stage is The number of agents set in each learning phase;

[0008] From each role The best fitness is selected from the initial population trained in the learning phase. The population performs evolution operations to obtain multiple The optimal population of the learning phase, where the first The optimal population in the first learning phase is used for +1 training session;

[0009] Determine the Whether the number of agents in the system in each learning stage is less than the set maximum number of agents;

[0010] If the If the number of agents in the system during the learning phase is not less than the maximum number of agents, the training is terminated and a trained agent is obtained.

[0011] In a second aspect, an embodiment of the present invention provides a multi-agent reinforcement learning system based on evolutionary curriculum learning, wherein the system includes agents with multiple roles and the number of agents is not constant; the system is used to perform evolutionary curriculum learning according to the method provided in the first aspect above.

[0012] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: according to the multi-agent reinforcement learning method provided by the present invention, by selecting only part of the trained initial population with the best fitness to perform evolutionary operations, the optimal population is obtained after completing this round of evolutionary course learning. Compared with the traditional method of allowing all populations to participate in the optimal population selection process, the present invention can significantly improve the training efficiency of the multi-agent system. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 A schematic diagram of the structure of a multi-agent reinforcement learning system based on evolutionary curriculum learning provided by an embodiment of the present invention;

[0014] Figure 2 A schematic diagram of the structure of a policy network and a value network based on a self-attention mechanism provided by an embodiment of the present invention;

[0015] Figure 3 A schematic diagram of the specific structure of a value network provided by an embodiment of the present invention;

[0016] Figure 4 A schematic diagram of a scenario in which a multi-agent reinforcement learning system based on evolutionary curriculum learning performs reinforcement learning based on an evolutionary curriculum learning algorithm provided by an embodiment of the present invention;

[0017] Figure 5 A flowchart for implementing a multi-agent reinforcement learning method based on evolutionary curriculum learning provided by an embodiment of the present invention;

[0018] Figure 6 A flowchart for implementing an evolution operation provided by an embodiment of the present invention;

[0019] Figure 7 A schematic diagram of another scenario of a multi-agent reinforcement learning system based on evolutionary curriculum learning and performing reinforcement learning based on an evolutionary curriculum learning algorithm provided by an embodiment of the present invention;

[0020] Figure 8 A schematic diagram of a scenario in which an intelligent agent saves experience replay data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0021] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.

[0022] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0023] It will also be understood that the term "and / or" used in the present description and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0024] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0025] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0026] References to "one embodiment" or "some embodiments" in the present specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present invention. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0027] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.

[0028] Example 1

[0029] Figure 1 Shown is a structural diagram of a multi-agent reinforcement learning system based on evolutionary curriculum learning provided by an embodiment of the present invention.

[0030] In one possible implementation, see Figure 1 The multi-agent reinforcement learning system based on evolutionary curriculum learning (hereinafter referred to as the system) 100 may include Different roles, each role includes The intelligent agent, is less than or equal to A positive integer.

[0031] For example, the system 100 may include Heterogeneous roles, one of which is a food delivery robot and the other is a patrol drone. Role 1 can include =5 food delivery robots, role 2 can include =6 patrol drones.

[0032] For example, the number of agents included in the system 100 is not constant, and users can add or delete agents in the system 100 according to their needs.

[0033] For example, if the system 100 is used for hotel services, the manager can increase the number of agents that play the role of food delivery robots due to the increase in the number of guests. After the hotel's peak season is over, the manager can also reduce the number of agents that play the role of food delivery robots according to demand.

[0034] In one example, each intelligent agent in the system 100 may be provided with a policy network and a value network. The system 100 may perform reinforcement learning on the policy networks and value networks within all intelligent agents based on the evolutionary curriculum learning algorithm proposed in the present invention to obtain trained intelligent agents.

[0035] Example 2

[0036] Since curriculum learning steadily improves the decision-making performance of MARL by gradually increasing the number of agents, the dimension of the agent observation vector continues to increase. However, traditional neural networks have difficulty processing input and output of dynamic dimensions, which affects the knowledge transfer of agents between courses. In view of this, the present invention proposes a network structure with a self-attention mechanism to optimize the policy network and value network in the traditional evolutionary curriculum learning algorithm, so that it can dynamically respond to changes in input and output. By dynamically calculating the weight of the input information of each agent, the information is flexibly aggregated and the input information of varying dimensions is converted into information of fixed dimensions, thereby effectively processing variable-length input sequences.

[0037] Figure 2 The diagram below shows a schematic diagram of a policy network and value network based on a self-attention mechanism according to an embodiment of the present invention. By way of example and not limitation, the policy network 210 may include a policy sub-network 212 and a first encoding network 211 based on a dynamic self-attention mechanism; the value network 220 may include a value sub-network 222 and a second encoding network 221.

[0038] For example, the first coding network 211 and the second coding network 221 have the same structure. The first coding network 211 and the second coding network 221 can be used to identify the intelligent agent (e.g., the first coding network 211 and the second coding network 221) in which the coding network is located. The observation vector and action vector of an agent are encoded to obtain an embedding vector of a preset dimension.

[0039] The strategy sub-network 212 is used to generate the first The action to be performed by the agent at the next moment (i.e. The value sub-network 222 is used to evaluate the strategy and output its value according to the embedding vector of the preset dimension of the agent at the current moment.

[0040] Optionally, the structures of the policy sub-network 212 and the value sub-network 222 may be respectively the same as the structures of the policy network and the value network in the current Multi-Agent Soft Actor-Critic (MASAC) algorithm.

[0041] In a possible implementation, taking the second encoding network 221 in the value network 220 as an example, the second encoding network 221 may include multiple first multi-layer perceptrons, first attention layers, second attention layers, and second multi-layer perceptrons.

[0042] For example, see Figure 3 , No. The first multilayer perceptron in an agent can The observation vector of the agent and its own motion vector Perform dimension transformation to uniformly transform the input observation vector and action vector into the first preset dimension, and obtain the action vector of the first preset dimension and the observation sub-vector of the first preset dimension. The observation vector obtained by an agent observing other agents - and - Input to the first attention layer for encoding embedding to obtain the first encoding vector, where is the total number of agents in the agent system 100. The observation vector obtained by an agent observing an obstacle - Input to the second attention layer for encoding embedding to obtain the second encoding vector, where is the dimension of the observation vector, will vary with the number of agents in the agent system 100. The observation vector obtained by the agent observing itself , the first encoding vector, the second encoding vector and its own action vector are embedded and connected to obtain an embedding vector of the second preset dimension = , Indicates the The second encoding network of the agent.

[0043] Specifically, the observation sub-vector is a row vector or a column vector in the observation vector.

[0044] In one example, see Figure 3, the value sub-network 222 may include multiple third attention layers, an attention concatenation layer and a third multi-layer perceptron.

[0045] For example, see Figure 3 , No. The third attention layer can The embedding vectors of other agents are linearly transformed to obtain key vectors Sum value vector , for The query vector is obtained by linearly transforming the embedding vector of the agent ; First, query vector and key vector Perform inner product operation and then use Softmax The function activates the result of the inner product operation to obtain the first Other agents relative to the The attention weight vector of the agent , and finally through Pair value vector The weighted average is obtained Other agents relative to the The weighted attention vector of each agent The attention concatenation layer can concatenate all The third multi-layer perceptron can obtain the spliced ​​multi-head attention. Evaluate the output of the policy network The value of the action to be performed by the agent at the next moment.

[0046] Specifically, Less than or equal to and is not equal to .

[0047] Exemplarily, the evaluation process of the value network 220 can be expressed as follows:

[0048] ,

[0049] in, For the The value network 220 in the intelligent body is based on The observation vector of the agent and motion vector The value of the output, Indicates Figure 3 The last two fully connected layers in the third multilayer perceptron shown.

[0050] For example, the spliced ​​multi-head attention can satisfy the following formula:

[0051] ,

[0052] in, represents a nonlinear activation function (e.g. Figure 3 in Softmax function), For the embedding vectors of other agents, Indicates the A second encoding network of other agents, 、 Respectively observation vectors and action vectors input by other agents.

[0053] For example, the attention weight vector can satisfy the following formula:

[0054] ,

[0055] in, and Respectively The query vector , key vector Sum value vector The three trainable attention parameter matrices used for linear transformation, is the transpose symbol.

[0056] According to the strategy network and value network provided by the present invention, by introducing an encoding network based on attention mechanism encoding, it is possible to flexibly respond to the dynamic addition of intelligent agents and effectively solve the problem of processing information with variable length dimensions, so that the evolutionary curriculum learning method can be more effectively used to train intelligent agents and gradually improve the collaborative decision-making performance of the network.

[0057] Example 3

[0058] As an example, see Figure 4 The diagram in FIG shows a scenario of a multi-agent reinforcement learning system based on evolutionary curriculum learning that performs reinforcement learning based on an evolutionary curriculum learning algorithm. The solid-line circle within the rounded rectangle represents the agent belonging to role 1, and the dotted-line oval represents the agent belonging to role 2. The two roles compete for resources in the scene environment (see Figure 4 The arrows connecting agents indicate that two agents with different roles are competing, and the arrows connecting an agent and a resource indicate that the agent is picking up the resource.

[0059] In one example, each time a new agent is added, the system 100 enters a new learning phase, retraining the policy network and value network within each agent. Alternatively, if the system 100 initially includes a large number of agents, the system 100 may be trained in multiple learning phases.

[0060] For example, see Figure 4 When the system 100 is used for the first time (i.e., the initial stage), both character 1 and character 2 include 8 agents, and the system 100 can be divided into three learning stages for training. In the first learning stage, it is set that the system 100 includes 2 agents belonging to character 1 and 2 agents belonging to character 2. These two agents with different roles compete and pick up resources. After completing the learning of the first stage, it enters the second learning stage and sets the system 100 to include 4 agents for both character 1 and character 2 for training. After that, the system 100 is set to include 8 agents for both character 1 and character 2 for training.

[0061] Figure 5 The flowchart shown is an implementation flow of a multi-agent reinforcement learning method based on evolutionary curriculum learning provided by an embodiment of the present invention. By way of example and not limitation, the method can be applied to the aforementioned system 100. The method can include steps S501-S504, each of which is described below.

[0062] S501, for The agents in the system are trained multiple times in the learning phase to generate multiple The initial population for the first learning phase.

[0063] In one example, the system 100 can train the strategy network and value network of each agent multiple times based on the MASAC algorithm to obtain the An initial population.

[0064] For example, is a positive integer greater than or equal to 2.

[0065] For example, each role The number of agents in the initial population of each learning stage is The number of agents set in each learning phase.

[0066] It should be understood that the intelligent agents included in various groups in the present invention are simulations of actual intelligent agents (such as drones, etc.).

[0067] S502, from each role The best fitness is selected from the initial population trained in the learning phase. The population performs evolutionary operations and obtains the first The optimal population for each learning phase.

[0068] For example, is a positive integer greater than or equal to 2.

[0069] S503, determine the Whether the number of agents in the system in each learning stage is less than the set maximum number of agents.

[0070] In one example, if If the number of agents in the system 100 in the learning stage is not less than the set maximum number of agents, step S504 can be performed to end the training and obtain the trained agents.

[0071] For example, referring to the above example, the system 100 is conducting evolutionary course learning in the third learning stage. At this time, the number of intelligent agents in the system 100 is 16, which is equal to the maximum number of intelligent agents set, and the training can be terminated.

[0072] In another example, if In the learning phase, the number of agents in the system 100 is less than the set maximum number of agents. = +1, increase the number of agents and continue the next training from the above step S501.

[0073] For example, referring to the first learning stage and the second learning stage in the above example, the process may return to step S501 to continue the next training.

[0074] According to the multi-agent reinforcement learning method provided by the present invention, by selecting only part of the trained initial population with the best fitness to perform evolutionary operations, the optimal population is obtained after completing this round of evolutionary course learning. Compared with the traditional method of allowing all populations to participate in the optimal population selection process, the present invention can significantly improve the training efficiency of the multi-agent system.

[0075] Example 4

[0076] Figure 6 The flowchart of an implementation of an evolution operation provided by an embodiment of the present invention is shown. As an example and not a limitation, the method can be a possible specific implementation of the above step S502. The method can include steps S601-S603, each of which is described below.

[0077] S601, based on the selection probability, select the first The best fitness is selected from the initial population trained in the learning phase. The populations compete to get the first The population after competition in the learning phase.

[0078] In some embodiments, see Figure 7 (a) in the above example can be found from the first The initial populations in the learning phase are selected based on the selection probability. the population with the best fitness; see Figure 7 (b) in the figure, then the different roles can be selected to compete with each other and get the The population after the competition of the learning phase is then selected from the population after the competition. The population with the best fitness performs the cross-fusion operation in the subsequent step S602.

[0079] For example, if the selection probability of a population is 0.5, then there is a 50% probability that the population will be selected for the subsequent competition operation.

[0080] For example, see Figure 7 In (b), referring to the above example, we can make each population in role 1 compete with each population in role 2. The competition results of the selected populations, before selection The populations are then cross-fused.

[0081] Compared with the current traditional course learning method that requires each population to participate in the competition, the present invention selects some populations with better fitness to compete based on probability, which can effectively shorten the training time of the model while ensuring the performance and diversity of the model (i.e., the strategy network and the value network).

[0082] For example, , It is also a positive integer greater than or equal to 2.

[0083] For example, the selection probability of a population can be determined according to the fitness of the population. The better the fitness of the population, the greater the selection probability.

[0084] In one possible implementation, while maintaining a certain degree of randomness, we can focus on populations with higher fitness to further improve training effectiveness. Therefore, we can first scale the fitness using exponential scaling, then determine the probability of selecting a population based on the scaled fitness, and finally select the population with higher fitness based on the selection probability.

[0085] In one example, the selection probability of a population may satisfy the following formula:

[0086] ,

[0087] in, For the The selection probability of a population, For the The first role The total number of initial populations for each learning phase.

[0088] For example, the fitness of the population after scaling can satisfy the following formula:

[0089] ,

[0090] in, For the The fitness of the population after scaling, is the exponential scaling factor, For the The fitness of a population.

[0091] In one example, each role can be The first stage of learning The performance and stability of each agent in the initial population are calculated by calculating the standardized average reward and the standardized average variance of the population; then according to the The standardized average reward and the standardized average variance of the population determine the The fitness of a population.

[0092] For example, The fitness of a population can satisfy the following formula:

[0093] ,

[0094] in, 、 are performance weight and stability weight respectively, For the The normalized average reward of the population, For the The standardized mean variance of the population.

[0095] This probabilistic sampling population selection method based on fitness scaling can simplify the population selection process and accelerate the training progress of evolutionary curriculum learning.

[0096] S602, from The first The population with the best fitness is cross-fused to obtain the first The population after cross-fusion of the learning stages.

[0097] In one example, see Figure 7 In (c), you can first select the same role After the competition of the first learning stage, the population is cross-grouped, and then the knowledge transfer operations such as model reloading and experience reuse are performed on it to obtain the first The population after the intra-group crossover in the first learning phase. After the intra-group crossover in the first learning phase, the population is crossovered between groups, and then the knowledge transfer operations such as model reloading and experience reuse are performed on it to obtain the first The population after cross-fusion of the learning stages.

[0098] Through fusion operations such as intra-group crossover and inter-group crossover, more promising offspring can be produced, improving the performance and adaptability of intelligent agents in different environments.

[0099] For example, after performing intra-group crossover, each role can obtain populations; after crossover between groups, each role can obtain A population.

[0100] It should be understood that the present invention does not limit the order of performing intra-group crossover and inter-group crossover.

[0101] S603, from After the cross-fusion of the learning phase, some populations are selected for mutation operation to obtain the first The optimal population for each learning phase.

[0102] In one possible implementation, each role can No. Random sampling from the population after cross-fusion of learning stages Perform multiple mutation operations on each population to obtain the first The population after the mutation of the learning phase.

[0103] In one example, see Figure 7 In (d), each time a mutation occurs, zero-mean Gaussian noise can be added to the population after the previous mutation to obtain the population after this mutation.

[0104] For example, the standard deviation of the Gaussian noise added to each mutation operation can be dynamically adjusted according to the evolutionary stage. For example, a larger standard deviation is used in the early stage (i.e., the first few mutations) to enhance the exploration ability, and a smaller standard deviation is used in the later stage (the last few mutations) to fine-tune the model.

[0105] This mutation method can help intelligent agents adapt to a wider range of changes in different environments by increasing the diversity of the population, and improve the global search ability and convergence efficiency during model training.

[0106] According to the multi-agent reinforcement learning method provided by the present invention, by selecting only part of the trained initial population with the best fitness to perform evolutionary operations, the optimal population is obtained after completing this round of evolutionary course learning. Compared with the traditional method of allowing all populations to participate in the optimal population selection process, the present invention can significantly improve the training efficiency of the multi-agent system.

[0107] In some embodiments, the model reloading operation mentioned in step S602 may include: copying the previous The best fitness The model parameters of the agents in the population after the competition of the learning phase; the experience reuse operation can include: indivual The best fitness The previous population is extracted and saved from the competition after the learning phase. -Experience replay data for 1 learning phase.

[0108] Specifically, see Figure 4 and Figure 8 , an experience replay pool can be set up inside each agent to save the training data of the agents in the population as experience replay data; When performing the cross-fusion operation (i.e., the above step S602) in the first learning stage, the selected The network parameters of the agents in the population after the competition of the first learning phase are copied to the new population generated by the intra-group crossover or component crossover, and the previous parameters of these old populations are extracted. -1 The experience replay data of the learning phase is saved to the experience replay pool of the new population of agents.

[0109] For example, The model reloading process of the learning stage can be expressed as: , ;in, 、 Respectively The first role The network parameters of the policy network and value network in the new population generated by the cross-fusion operation in the learning stage, 、 Respectively before indivual The best fitness After the first learning phase of competition The network parameters of the policy network and value network of the agents in a population.

[0110] By migrating the network parameters of agents from the same population with the highest fitness during the previous learning phase to the new population generated by the cross-fusion operation, the agents in the new population are given optimal decision-making capabilities. Furthermore, compared to traditional knowledge transfer methods that only reload models, this simultaneous knowledge transfer method of reloading models and reusing experience can enhance the fitness of agents and improve their performance.

[0111] In one possible implementation, During the first learning stage, if - The experience playback data of each stage in a learning phase is used in turn Represents the number of experience replay data extracted from each learning stage Can satisfy: ,in, The maximum capacity of the experience replay pool.

[0112] For example, see Figure 4 ,As the learning phase progresses, each agent can gradually update ,the experience replay data in chronological order, replacing old ,experience samples with new ones, to ensure that the agent can maintain ,the effectiveness and adaptability of learning in a constantly changing ,environment.

[0113] Optionally, due to the changing number of agents at each learning stage, the state space at different learning stages may be different. The number of agents in the first learning stage is greater than that in the -1 learning stage, then the state space of the corresponding agent is also larger than the first -1 learning stage state space. Due to the different state dimensions The first learning stage cannot be used directly -1 learning stage experience data. Therefore, the experience replay data can be dimensionally adjusted by zero-filling the samples with smaller state dimensions to make them consistent with the stage with larger state dimensions, thereby achieving experience sharing and reuse between different stages.

[0114] In one example, after introducing experience reuse into the evolutionary curriculum learning course based on the MASAC algorithm, the ultimate goal of the evolutionary curriculum learning can be to maximize the objective function of the policy network and minimize the loss of the value network.

[0115] For example, the objective function of the policy network can satisfy the following formula:

[0116] ,

[0117] in, Indicates the Agent-based policy network The objective function, Express expectations, Indicates the The state of an agent It is from Experience replay pool for each learning phase Extracted from For the The action vector of each agent, Indicates the The value network of an intelligent agent, For the value of its output, Indicates the The policy network of each agent is based on The obtained strategy (i.e. The action to be performed by the agent at the next moment), represents the entropy of the policy, is the temperature coefficient, which is used to adjust the weight of entropy in the objective function. express is generated.

[0118] For example, the loss of the value network can satisfy the following formula:

[0119] ,

[0120] in, Indicates the The loss of the value network of each agent, Represents experience replay data It is from Experience replay pool for each learning phase Extracted from For the The reward received by the agent, For the The state of the agent at the next moment, For the The next action of an agent, For the The value network of each agent at the next moment, express is generated; is the discount factor, reflecting the The degree to which an agent cares about future rewards.

[0121] This hybrid transfer method based on model reloading and experience reuse can further improve the efficiency of knowledge transfer between courses and give full play to the potential of course learning.

[0122] According to the multi-agent reinforcement learning method provided by the present invention, by selecting only part of the trained initial population with the best fitness to perform evolutionary operations, the optimal population is obtained after completing this round of evolutionary course learning. Compared with the traditional method of allowing all populations to participate in the optimal population selection process, the present invention can significantly improve the training efficiency of the multi-agent system.

[0123] Furthermore, (1) the probability sampling population selection method based on fitness scaling effectively reduces redundant competition between populations while ensuring population performance and diversity, greatly simplifies the selection process of the optimal population, and accelerates the training speed of evolutionary curriculum learning. (2) In the cross-integration process of evolutionary curriculum learning, knowledge transfer between new and old populations is carried out based on model reloading and experience reuse; it can give the intelligent agents in the new population the optimal decision-making ability, further improve the efficiency of knowledge transfer between courses, give full play to the potential of course learning, and thus improve the performance of intelligent agents. (3) By introducing an encoding network based on attention mechanism encoding in the value network and policy network, it can flexibly respond to the dynamic addition of intelligent agents and effectively solve the problem of processing variable-length dimensional information, so that the evolutionary curriculum learning method can be more effectively used to train intelligent agents and gradually improve the collaborative decision-making performance of the network.

[0124] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

Claims

1. A multi-agent reinforcement learning method based on evolutionary curriculum learning, characterized in that: The method is applied to a multi-agent reinforcement learning system based on evolutionary curriculum learning, wherein the system includes agents with multiple roles and the number of agents is not constant. The agents have two roles, namely, a patrol drone and a food delivery robot. The method includes: For the first The agents in the system are trained multiple times in the learning phase to generate multiple The initial population of the learning phase, where the first The number of agents in the initial population of each learning stage is The number of agents set in each learning phase; From each role The best fitness is selected from the initial population trained in the learning phase. The population performs evolution operations to obtain multiple The optimal population of the learning phase, where the first The optimal population in the first learning phase is used for +1 training session; Determine the Whether the number of agents in the system in each learning stage is less than the set maximum number of agents; If the If the number of agents in the system during the learning phase is not less than the maximum number of agents, the training is terminated to obtain trained agents, wherein the trained patrol drone is used for patrolling, and the trained food delivery robot is used for food delivery; Among them, the The best fitness is selected from the initial population trained in the learning phase. The population performs evolution operations to obtain multiple The optimal population for each learning phase includes: Based on the probability of selection from each role The best fitness is selected from the initial population trained in the learning phase. The populations compete to get the first The population after the competition of learning phases, where the fitness of the population is determined by the performance and stability of all agents in the population; From the said The first The population with the best fitness is cross-fused to obtain the first The population after cross-fusion of the learning stages; Regarding the The population after the cross-fusion of the learning stages is mutated to obtain the The optimal population for each learning phase; Among them, the The first The population with the best fitness is cross-fused to obtain the first The population after cross-fusion of the learning stages includes: From the said The first The population with the best fitness performs intra-group crossover, and performs model reloading and experience reuse operations to obtain the first The population after intra-group crossover in the learning phase; Regarding the After the intra-group crossover in the first learning phase, the population undergoes inter-group crossover, and performs model reloading and experience reuse operations to obtain the first The population after cross-fusion of the learning stages; The model reloading operation includes: copying the The best fitness Model parameters of the agents in the population after competition in the learning phase; Wherein, the experience reuse operation includes: The best fitness The first -1 learning phase experience playback data and save; Each intelligent body is provided with a strategy network and a value network. The strategy network includes a first coding network and a strategy sub-network. The value network includes a second coding network and a value sub-network. The first coding network and the second coding network have the same structure. The second encoding network includes: a plurality of first multi-layer perceptrons, a first attention layer, a second attention layer, and a second multi-layer perceptron; No. The first multi-layer perceptron in the intelligent body is used to The observation vector and action vector of each agent are transformed into a dimension to obtain an action vector of a first preset dimension and an observation sub-vector of a first preset dimension, wherein the observation sub-vector is the first preset dimension. The row vector or column vector of the agent's own observation vector; The said The first attention layer in the agent is used to An agent observes other agents to obtain an observation sub-vector of a first preset dimension, and performs encoding embedding to obtain a first encoding vector; The said The second attention layer in the agent is used to The observation sub-vector of the first preset dimension obtained by the intelligent agent observing the obstacle is encoded and embedded to obtain a second encoding vector; The said The second multi-layer perceptron in the intelligent body is used to The observation sub-vector of the first preset dimension obtained by the agent observing itself, the first encoding vector, the second encoding vector and the action vector of the first preset dimension are embedded and connected to obtain an embedding vector of the second preset dimension.

2. The method according to claim 1, characterized in that The fitness of the population satisfies the following formula: , in, For the The fitness of a population, 、 are performance weight and stability weight respectively, For the The normalized average reward of the population, For the The standardized mean variance of the population; The said The normalized average reward of the population is calculated based on The performance of all agents in the population is determined by the The standardized mean variance of the population is based on the The stability of all agents in a population is determined.

3. The method according to claim 2, characterized in that The selection probability satisfies the following formula: , in, For the said The selection probability of a population, For the said The fitness of the population after scaling, is the exponential scaling factor, For the The first role The total number of initial populations for each learning phase.

4. The method according to claim 1, wherein The number of experience replay data extracted from each learning stage satisfies the following formula: , in, is the number of experience replay data extracted from each learning stage, The maximum capacity of the experience replay pool.

5. A multi-agent reinforcement learning system based on evolutionary curriculum learning, characterized by: The system includes intelligent agents of multiple roles and the number of intelligent agents is not constant. The system is used to perform evolutionary curriculum learning according to the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multidirectional course reinforcement learning method and device for multi-agent decision

    CN116523076A

  • Multidirectional course learning training method and device for cluster robot cooperative navigation

    CN117540203A