Multi-agent decision-making method and device based on CTDE
By adopting the CTDE framework and centralized evaluation network method in the multi-agent system, the problem of environmental instability in the multi-agent system is solved, and a more stable multi-agent decision-making strategy is achieved.
Patent Information
- Application Number
- CN202411925552.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
In multi-agent systems, strategies and actions between agents influence each other, resulting in unstable environments, making it difficult to make effective decisions for single agents.
The multi-agent decision-making method based on CTDE is adopted, and the multi-agent decision-making strategy is initialized and trained through a centralized evaluation network to obtain the trained multi-agent decision-making strategy. The method includes obtaining the multiagent state, training the policy network and evaluation network, and determining it as a decision strategy if the loss value is less than the preset value.
Through centralized training of decentralized execution architecture, each agent adds the state and actions of other agents to consider, improves the search ability for global optimal solutions, improves the data format of the experience pool, and enhances the decision stability of the multi-agent system.
Smart Images

Figure CN120046644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a decision-making method and device for multi-agent based on CTDE. Background Art
[0002] In a multi-agent system, the strategies and actions among agents influence each other. The action decision of each agent not only needs to consider its own state but also the mutual relationship with other agents. If a single-agent reinforcement learning method is simply adopted for each agent, for each agent, the states and actions of other agents are part of the environment, resulting in environmental instability, and the application effect in the scenario of two-agent game confrontation is not ideal. Therefore, how to provide a decision-making method and device for multi-agent based on CTDE has become a technical problem urgently to be solved in this field. Summary of the Invention
[0003] The object of the present invention is to provide a decision-making method and device for multi-agent based on CTDE.
[0004] According to the first aspect of the present invention, a decision-making method for multi-agent based on CTDE is provided, including:
[0005] Obtain multi-agents, and use a centralized evaluation network to initialize the multi-agents to obtain a multi-agent state, a policy network, and an evaluation network corresponding to each multi-agent;
[0006] Train the policy network and evaluation network of the multi-agents according to the preset number of training cycles, training cycle length, experience replay buffer pool, and multi-agent state to obtain the trained multi-agents; the multi-agent state is a combination of the states of all agents in the environment at the same moment and a combination of the actions of all agents in the environment;
[0007] If the loss value of the trained multi-agents is less than a preset value, then determine the trained multi-agents as the decision-making strategy of the multi-agents, and the decision-making strategy of the multi-agents is used to adjust the strategy of the multi-agents.
[0008] Optionally, the method further includes:
[0009] For each agent, select the execution action of each agent according to the current policy and exploration noise;
[0010] Obtain the current return value according to the execution action of each agent, and convert the return value into a new state;
[0011] Store the execution action, the corresponding return value, and the new state into the experience replay buffer pool.
[0012] Optionally, the method further includes:
[0013] During the training process, if the amount of experience data in the experience replay buffer pool is greater than the preset start training data amount, training samples are obtained from the experience replay buffer pool.
[0014] Optionally, if the loss value of the trained multi-agent is less than the preset value, determining the trained multi-agent as the decision-making strategy of the multi-agent includes:
[0015] Calculating a first loss value of the policy network of the multi-agent, and updating the policy network according to the first loss value;
[0016] Calculating a second loss value of the evaluation network of the multi-agent, and updating the evaluation network according to the second loss value.
[0017] Optionally, calculating a first loss value of the policy network of the multi-agent, and updating the policy network according to the first loss value includes:
[0018] Obtaining the observation state of the agent from the experience replay buffer pool;
[0019] Inputting the observation state of the agent into the policy network, and the output is the action executed by the agent;
[0020] Replacing the action executed by the multi-agent stored in the experience replay buffer pool with the action executed by the agent;
[0021] Invoking the evaluation network to obtain the evaluation value of the overall next state and action of all agents.
[0022] Optionally, calculating a second loss value of the evaluation network of the multi-agent, and updating the evaluation network according to the second loss value includes:
[0023] Training the evaluation network through forward propagation, with the input being the state set and action set of all agents, and the output being the evaluation value;
[0024] Updating the network using TD mean squared error backpropagation, sequentially extracting the next observation state of each agent from the experience replay buffer pool, and inputting it into their respective target policy networks to obtain the next actions of all agents;
[0025] Invoking the target evaluation network to obtain the evaluation value of the overall next state and action of all agents.
[0026] According to the second aspect of the present invention, there is provided a decision-making device for a multi-agent based on CTDE, including:
[0027] An acquisition module, configured to acquire multiple agents, and initialize the multiple agents by using a centralized evaluation network to obtain a multi-agent state, a policy network, and an evaluation network corresponding to each multi-agent;
[0028] A training module, configured to train the policy network and the evaluation network of the multiple agents according to a preset number of training cycles, a training cycle length, an experience replay buffer pool, and the multi-agent state to obtain trained multiple agents; the multi-agent state is a combination of the states of all agents in the environment at the same moment, and a combination of the actions of all agents in the environment;
[0029] A determination module, configured to, if a loss value of the trained multiple agents is less than a preset value, determine the trained multiple agents as a decision-making strategy of the multiple agents, and the decision-making strategy of the multiple agents is used to adjust the strategy of the multiple agents.
[0030] Optionally, the acquisition module is configured to:
[0031] For each agent, select an execution action of each agent according to the current policy and exploration noise;
[0032] Obtain a current return value according to the execution action of each agent, and convert the return value into a new state;
[0033] Store the execution action, the corresponding return value, and the new state into the experience replay buffer pool.
[0034] Optionally, the acquisition module is configured to:
[0035] During the training process, if the amount of experience data in the experience replay buffer pool is greater than a preset start training data amount, obtain training samples from the experience replay buffer pool.
[0036] Optionally, the determination module is configured to:
[0037] Calculate a first loss value of the policy network of the multiple agents, and update the policy network according to the first loss value;
[0038] Calculate a second loss value of the evaluation network of the multiple agents, and update the evaluation network according to the second loss value.
[0039] Optionally, the determination module is configured to:
[0040] Obtain the observation state of the agent from the experience replay buffer pool;
[0041] Input the observation state of the agent into the policy network, and the output is the action executed by the agent;
[0042] Replace the actions executed by the multi-agent stored in the experience replay buffer pool with the actions executed by the agent;
[0043] Call the evaluation network to obtain the evaluation value of the next state and actions of all agents as a whole.
[0044] Optionally, the determination module is configured to:
[0045] Train the evaluation network through forward propagation, with the input being the state set and action set of all agents, and the output being the evaluation value;
[0046] Use TD mean squared error backpropagation to update the network. Sequentially extract the next observation state of each agent from the experience replay buffer pool and input it into its respective target policy network to obtain the next actions of all agents;
[0047] Call the target evaluation network to obtain the evaluation value of the next state and actions of all agents as a whole.
[0048] In a third aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the method described in any of the above aspects.
[0049] In a fourth aspect, the present application discloses a non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the method described in any of the above aspects.
[0050] In a fifth aspect, the present application discloses a computer program product, when the instructions in the computer program product are executed by the processor of the electronic device, enabling the electronic device to execute the method described in any of the above aspects.
[0051] The beneficial effects brought by the present invention are as follows:
[0052] As can be seen from the above solution, the embodiments of the present invention provide a decision-making method and device for multi-agent based on CTDE, which obtain multi-agents and use a centralized evaluation network to initialize the multi-agents to obtain a multi-agent state, a policy network, and an evaluation network corresponding to each multi-agent; according to the preset number of training cycles, training cycle length, experience replay buffer pool, and multi-agent state, train the policy network and evaluation network of the multi-agents to obtain the trained multi-agents; the multi-agent state is a combination of the states of all agents in the environment at the same time, and a combination of the actions of all agents in the environment; if the loss value of the trained multi-agents is less than a preset value, then determine the trained multi-agents as the decision-making strategy of the multi-agents, and the decision-making strategy of the multi-agents is used to adjust the strategy of the multi-agents. A multi-agent training architecture of centralized training and decentralized execution is adopted, and each agent has a separate evaluation network and policy network; the decision-making of each agent additionally considers the states and actions of other agents, and realizes the observation of the other party by increasing the field of view of each agent; improve the data format in the experience pool, and additionally consider the states and actions of all agents; introduce a policy set method to improve the search ability for the global optimal solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 FIG. is a schematic flow chart of a decision-making method for multi-agent based on CTDE provided by an embodiment;
[0054] Figure 2 FIG. is a schematic structural diagram of a policy network provided by an embodiment;
[0055] Figure 3 FIG. is a schematic structural diagram of an evaluation network provided by an embodiment;
[0056] Figure 4 FIG. is a schematic diagram of an experience data generation process provided by an embodiment;
[0057] Figure 5 FIG. is a schematic diagram of a training and updating process of an evaluation network provided by an embodiment;
[0058] Figure 6 is a structural block diagram of a decision-making device for multi-agent based on CTDE of the present application.
[0059] Figure 7 is a block diagram of an electronic device of the present application.
[0060] Figure 8 is a block diagram of a computer-readable storage medium of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0062] In a multi-agent system, each agent not only needs to consider its own strategy but also the impact of other agents' strategies on itself. Multi-agent reinforcement learning methods include three architectures: centralized learning, independent learning, and centralized training with decentralized execution (CTDE). To improve the training effect while ensuring the training efficiency, a multi-agent reinforcement learning method with centralized training and centralized execution is selected.
[0063] To solve the environmental instability in the scenario of two-agent game confrontation, a multi-agent deep reinforcement learning method based on the CTDE framework is studied: 1) Adopt a multi-agent training architecture with centralized training and decentralized execution, where each agent has a separate evaluation network and policy network; 2) Each agent's decision-making additionally considers the states and actions of other agents, and observes the other party by increasing the vision range of each agent; 3) Improve the data format in the experience pool, and additionally consider the states and actions of all agents; 4) Introduce a policy set method to improve the search ability for the global optimal solution.
[0064] Refer to Figure 1 , which shows a flowchart of the steps of a decision-making method for a multi-agent based on CTDE in the present application. This method can be applied to an electronic device. Specifically, the method may include the following steps:
[0065] S101. Obtain multi-agents, and initialize the multi-agents using a centralized evaluation network to obtain a multi-agent state, a policy network, and an evaluation network corresponding to each multi-agent.
[0066] S102. Train the policy network and the evaluation network of the multi-agents according to the preset number of training cycles, training cycle length, experience replay buffer pool, and multi-agent state to obtain the trained multi-agents. The multi-agent state is a combination of the states of all agents in the environment at the same moment and a combination of the actions of all agents in the environment.
[0067] S103. If the loss value of the trained multi-agent is less than the preset value, the trained multi-agent is determined as the decision-making strategy of the multi-agent, and the decision-making strategy of the multi-agent is used to adjust the strategy of the multi-agent.
[0068] In an underwater confrontation multi-platform system, different types of platforms have different functions, and these functions correspond to different actions and goals in the algorithm. Therefore, each agent needs to have a separate policy network. In addition, the tasks of the agents on both sides in the underwater DK scenario are not exactly the same, and each agent needs to have a separate reward function and evaluation network. Through comprehensive analysis, the single-agent reinforcement learning algorithm is applied to the multi-agent scenario.
[0069] During the training process of the multi-agent, each agent makes action decisions based on its own local observation state, and the environment returns the reward value of each agent in the global context according to the actions of the agents in the overall system. Therefore, in the design of the experience memory bank, the experience data should include global information, and when training different networks, the required data can be selected from the experience pool.
[0070] Another embodiment of this application further supplements the decision-making method of the multi-agent based on CTDE provided in the above embodiment.
[0071] The multi-agent deep reinforcement learning algorithm based on CTDE applies the deep deterministic gradient policy to the multi-agent system and makes improvements according to the characteristics of the multi-agent system. The algorithm adopts a centralized training and decentralized execution framework, and each agent has a separate policy network and evaluation network. During the training process, a centralized evaluation network is adopted. This network regards all agents in the environment as part of the policy, and provides information such as the observations and potential behaviors of other agents to the agent, thus converting an unpredictable environment into a predictable environment, that is, the evaluation network guides the update of the policy network according to the global information; during the testing process, the agents do not need to obtain this centralized evaluation network, and they can take actions based on their own observations and predictions of the behaviors of other agents, that is, the policy network makes action decisions only based on the local observation information of the agent. Based on this framework, during the training phase, the agent can use more information, but the learned policy can only use local information (i.e., the observations of the agent itself) during the testing phase.
[0072] The multi-agent deep reinforcement learning algorithm based on CTDE applies the deep deterministic policy gradient algorithm to the multi-agent system, and adopts a multi-agent training architecture of centralized training and decentralized execution. It can not only learn the strategies of other agents according to the global information during the training phase, but also make quick responses according to the local observations during the testing process. In addition, each agent has a separate policy network and evaluation network, that is, each agent can make independent decisions and can learn separate task goals.
[0073] Each agent has four networks: a policy network, a value network, a target policy network, and a target value network. The differences are as follows: 1) The composition of data in the experience pool is different. In the multi-agent decision-making algorithm, each set of experience data includes the information of all agents at the same moment. 2) The input of the value network is different. In the multi-agent decision-making algorithm, the value network scores the impact of a certain action on the entire system. Even for the value network of an agent, its input includes the states and actions of all agents. 3) The update method of the target network is different. The target network is updated by copying. The parameters of the policy network and the value network are periodically copied as the parameters of the target policy network and the target value network respectively. In the multi-agent decision-making algorithm, the target network adopts a soft-copy method, partially copying the parameters of the policy network and the value network, while partially retaining the parameters of the target network itself. This method is more conducive to ensuring the stability of training.
[0074] Optionally, the method further includes:
[0075] For each agent, according to the current policy and exploration noise, select the execution action of each agent;
[0076] According to the execution actions of each agent, obtain the current return value and convert the return value into a new state;
[0077] Store the execution actions, the corresponding return values, and the new states into the experience replay buffer pool.
[0078] The network structure of the multi-agent decision-making algorithm is similar to that of the deep deterministic gradient algorithm. As shown in Figure 4, a fully connected manner is still adopted between each network layer. Compared with the network structure of the deep deterministic gradient algorithm, the main differences are as follows:
[0079] 1) The policy networks and value networks of each agent are independent of each other and no longer share weights.
[0080] 2) During the training process, the action part of the input of the value network also includes the actions output by the policy networks of other agents, and the output is still the corresponding evaluation value. As shown in Figure 2 and Figure 3 shown.
[0081] In a multi-agent system, each agent has two networks: a policy network and a value network. The inputs, outputs, and network structures of the two networks are all different. Training the policy network requires the individual information of the agent, while training the value network requires the information of all agents. Therefore, the training data of the multi-agent decision-making algorithm should contain the information of all agents.
[0082] In the experience pool of the Deep Deterministic Gradient Policy Algorithm, each set of experience data consists of the information of a single agent, including five sets of data: state, action, next state, reward value, and whether the task is completed. In the experience pool of the multi-agent deep reinforcement learning algorithm based on CTDE, each set of experience data consists of the information of all agents, including five sets: the states, actions, next states, reward values, and whether the task is completed of all agents.
[0083] In the multi-agent decision-making algorithm, the process of generating experience data is to interact with the environment through the policy network of the agents, extract experience data, and construct an experience pool. The difference is that the experience data in the multi-agent decision-making algorithm includes the information of all agents at the same moment. In a multi-agent system composed of N agents, assume that the policy network of agent i is represented as Take the observed state s of agent i i as the input of the policy network , and the output is the action decision a i , and add Gaussian noise as the executed action. The environment returns the reward value r i for the action. When all agents have completed interacting with the environment, store the experience data (X, a 1 , …, a N , X ′ , r 1 , …, r N ) in the experience pool, where X represents the set of all agents' states and X' represents the set of all agents' next states. The process of generating experience data in the multi-agent decision-making algorithm is as shown in Figure 4 .
[0084] Optionally, the method further includes:
[0085] During the training process, if the amount of experience data in the experience replay buffer pool is greater than the preset start training data amount, obtain training samples from the experience replay buffer pool.
[0086] Optionally, if the loss value of the trained multi-agent is less than the preset value, determine the trained multi-agent as the decision-making strategy of the multi-agent, including:
[0087] Calculate the first loss value of the policy network of the multi-agent and update the policy network according to the first loss value;
[0088] Calculate the second loss value of the evaluation network of the multi-agent and update the evaluation network according to the second loss value.
[0089] Optionally, calculating the first loss value of the policy network of the multi-agent and updating the policy network according to the first loss value includes:
[0090] Obtain the observation state of the agent from the experience replay buffer pool;
[0091] Input the observation state of the agent into the policy network, and the output is the action executed by the agent;
[0092] Replace the actions executed by multiple agents stored in the experience replay buffer pool with the actions executed by the agent;
[0093] Call the evaluation network to obtain the evaluation value of the overall next state and actions of all agents.
[0094] In the embodiments of the present application, the policy network can be updated, and the specific steps are as follows:
[0095] The policy network of agent i is represented as Taking a set of experience data (X, a 1 , …, a N , X ′ , r 1 , …, r N ) as an example, the training and update process of the policy network of agent i is shown in the following figure. Extract the observation state s of the agent from the experience data X ′ and input it into the policy network i , and the output is the action a executed by agent i. Replace a predict in the experience data with this action, call the evaluation network Q i (X, a i , …, a 1 , …, a N |θ i Q ) to obtain the evaluation value Q(X, a 1 , …, a N |θ i Q ). The update process of the evaluation network follows the following formula:
[0096]
[0097] Optionally, calculate the second loss value of the evaluation network of multiple agents, and update the evaluation network according to the second loss value, including:
[0098] Train the evaluation network through forward propagation. The input is the state set and action set of all agents, and the output is the evaluation value;
[0099] Adopt TD mean square error backpropagation to update the network. Sequentially extract the next observation state of each agent from the experience replay buffer pool and input it into their respective target policy networks to obtain the next actions of all agents;
[0100] Call the target evaluation network to obtain the evaluation values for the overall next state and actions of all agents.
[0101] The evaluation network of agent i is denoted as Q i (X, a 1 , …, a N | θ i Q ) and, taking a set of experience data (X, a 1 , …, a N , X ′ , r 1 , …, r N ) as an example, the training and updating process of the evaluation network of agent i is as Figure 5 shown. The evaluation network is trained through forward propagation, with the input being the state set and action set of all agents (X, a 1 , …, a N ), and the output being the evaluation value Q(X, a 1 , …, a N | θ i Q ). The network is updated using TD mean squared error backpropagation. The next observed state of each agent is sequentially extracted from the experience data X' and input into their respective target policy networks to obtain the next actions of all agents. Then, call the target evaluation network to obtain the evaluation values for the overall next state and actions of all agents The parameter update process of the evaluation network follows the following formula:
[0102]
[0103] On the above basis, the embodiments of the present application further include updating the target evaluation network and the target policy network, that is: the target evaluation network and the target policy network adopt the soft copy method, and the parameters of the evaluation network and the policy network are respectively partially copied at regular intervals for updating.
[0104]
[0105] On the basis of the above embodiments, the present application further includes updating the policy set:
[0106] An important problem in multi-agent reinforcement learning is the environmental non-stationarity caused by the constantly changing policies of agents. In a multi-agent adversarial game system, agents may overfit in policy learning in order to adapt to the behaviors of competitors. However, when platform D changes its policy, platform W cannot adapt to the new policy, resulting in failure. To obtain a multi-agent policy that is more robust to the policy changes of competing agents, a strategy set approach is adopted to aggregate K different sub-strategy sets. In each round of the game, this model randomly selects a specific sub-strategy for each agent to execute. Assume the policy μ i is the k-th sub-strategy of the K different sub-strategy sets, denoted as Then the goal of agent i is to maximize the set goal:
[0107]
[0108] Since different sub-strategies will be executed in different game rounds, a buffer is constructed for each sub-strategy of agent i Therefore, it can be deduced that the gradient of the set goal value with respect to is:
[0109]
[0110] The process of the multi-agent deep reinforcement learning algorithm based on CTDE is also roughly the same as that of the deep deterministic gradient policy algorithm. The main difference is that: since the multi-agent decision algorithm uses a centralized evaluation network, the state of the agent is no longer the state of the agent itself, but the combination of the states of all agents in the environment at the same time, and the action is no longer the action of itself, but the combination of the actions of all agents in the environment. The reward still uses the reward of the current agent.
[0111] The process of the multi-agent deep reinforcement learning algorithm based on CTDE is as follows:
[0112] Initialize the state, policy network and evaluation network of the agent, and set the number of training epochs M, the length of the training epoch T, and the experience replay buffer D;
[0113] FOR epoch = 1:M:
[0114] Initialize random noise, and initialize the environment, the policy network and the evaluation network of the agent;
[0115] FOR time step = 1:T:
[0116] For each agent:
[0117] Select the execution action ai according to the current policy and detection noise; each agent executes the corresponding action, obtains the current return value ri, and transitions to the new state si'; store the experience data (X, a 1 , …, a N , X ′ , r 1 , …, r N ) into the experience pool D. IF the amount of experience data > the preset start training data volume.
[0118] For each agent: randomly sample batch size groups of training samples from D; calculate the evaluation network loss value according to the formula; update the evaluation network according to the formula; calculate the policy network loss value according to the formula; update the policy network according to the formula; update the target evaluation network and target policy network according to the formula; end this training; end this cycle of training; end the training.
[0119] It should be noted that each implementable manner in this embodiment can be implemented separately or in any combined manner without conflict. This application does not make a limitation.
[0120] Another embodiment of this application provides a decision-making device for multi-agents based on CTDE, which is used to execute the decision-making method for multi-agents based on CTDE provided in the above embodiment.
[0121] As Figure 6 shown, it is a schematic structural diagram of the decision-making device for multi-agents based on CTDE provided in the embodiment of this application. The decision-making device for multi-agents based on CTDE includes an acquisition module 601, a training module 602, and a determination module 603, where:
[0122] The acquisition module 601 is used to acquire multi-agents and initialize the multi-agents by using a centralized evaluation network to obtain the multi-agent state, policy network, and evaluation network corresponding to each multi-agent;
[0123] The training module 602 is used to train the policy network and evaluation network of the multi-agents according to the preset number of training cycles, training cycle length, experience replay buffer pool, and multi-agent state to obtain the trained multi-agents; the multi-agent state is the combination of the states of all agents in the environment at the same moment, as well as the combination of the actions of all agents in the environment;
[0124] The determination module 603 is used to determine the trained multi-agents as the decision-making strategy of the multi-agents if the loss value of the trained multi-agents is less than the preset value, and the decision-making strategy of the multi-agents is used to adjust the strategy of the multi-agents.
[0125] Regarding the device in this embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0126] Another embodiment of the present application further supplements the decision-making device for multi-agent based on CTDE provided in the above embodiment.
[0127] Optionally, the acquisition module is used for:
[0128] For each agent, select the execution action of each agent according to the current policy and detection noise;
[0129] Obtain the current return value according to the execution actions of each agent, and convert the return value into a new state;
[0130] Store the execution action, the corresponding return value, and the new state in the experience replay cache pool.
[0131] Optionally, the acquisition module is used for:
[0132] During the training process, if the amount of experience data in the experience replay cache pool is greater than the preset start training data amount, obtain training samples from the experience replay cache pool.
[0133] Optionally, the determination module is used for:
[0134] Calculate the first loss value of the policy network of the multi-agent, and update the policy network according to the first loss value;
[0135] Calculate the second loss value of the evaluation network of the multi-agent, and update the evaluation network according to the second loss value.
[0136] Optionally, the determination module is used for:
[0137] Obtain the observation state of the agent from the experience replay cache pool;
[0138] Input the observation state of the agent into the policy network, and the output is the action executed by the agent;
[0139] Replace the action executed by the multi-agent stored in the experience replay cache pool with the action executed by the agent;
[0140] Call the evaluation network to obtain the evaluation value of the overall next state and action of all agents.
[0141] Optionally, the determination module is used for:
[0142] Train the evaluation network through forward propagation, with the input being the state set and action set of all agents, and the output being the evaluation value;
[0143] Update the network using TD mean squared error backpropagation. Sequentially extract the next observation state of each agent from the experience replay buffer pool and input it into their respective target policy networks to obtain the next actions of all agents.
[0144] Call the target evaluation network to obtain the evaluation values for the overall next state and actions of all agents.
[0145] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.
[0146] Optionally, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0147] An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0148] Figure 7 It is a block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0149] Refer to Figure 7 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0150] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0151] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, and the like. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0152] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0153] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0154] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0155] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0156] The sensor component 814 includes one or more sensors for providing an assessment of the status of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and a change in the temperature of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0157] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0158] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0159] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by a processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0160] Figure 8 FIG. is a block diagram of a computer-readable storage medium 1900 shown in the present application. For example, the computer-readable storage medium 1900 may be provided as a server.
[0161] Referring to Figure 8 , the computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0162] The computer-readable storage medium 1900 may further include a power supply component 1926 configured to perform power management of the computer-readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer-readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer-readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0163] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.
[0164] From the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0165] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
[0166] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.
[0167] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0168] In the embodiments provided in the present application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0169] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0170] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0171] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0172] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0173] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A multi-agent decision-making method based on CTDE, characterized in that: include: Acquire a multi-agent, and use a centralized evaluation network to initialize the multi-agent, and obtain a multi-agent state, a strategy network, and an evaluation network corresponding to each multi-agent; According to the preset number of training cycles, training cycle length, experience replay buffer pool and multi-agent state, the strategy network and evaluation network of the multi-agent are trained to obtain the trained multi-agent; the multi-agent state is a combination of the states of all agents in the environment at the same time, and a combination of the actions of all agents in the environment; If the loss value of the trained multi-agent is less than a preset value, the trained multi-agent is determined as the decision-making strategy of the multi-agent, and the decision-making strategy of the multi-agent is used to adjust the strategy of the multi-agent.
2. The CTDE-based multi-agent decision-making method according to claim 1, characterized in that: The method further comprises: For each agent, select the execution action of each agent according to the current strategy and detection noise; According to the execution actions of each agent, a current reward value is obtained, and the reward value is converted into a new state; The execution action, the corresponding reward value and the new state are stored in the experience replay buffer pool.
3. The CTDE-based multi-agent decision-making method according to claim 2, characterized in that: The method further comprises: During the training process, if the amount of experience data in the experience replay buffer pool is greater than the preset starting training data amount, a training sample is obtained from the experience replay buffer pool.
4. The CTDE-based multi-agent decision-making method according to claim 1, characterized in that: If the loss value of the trained multi-agent is less than a preset value, determining the trained multi-agent as a decision-making strategy of the multi-agent includes: Calculating a first loss value of a multi-agent policy network, and updating the policy network according to the first loss value; A second loss value of the multi-agent evaluation network is calculated, and the evaluation network is updated according to the second loss value.
5. The CTDE-based multi-agent decision-making method according to claim 4, characterized in that: The step of calculating a first loss value of a multi-agent policy network and updating the policy network according to the first loss value includes: Get the observed state of the agent from the experience replay buffer pool; Inputting the observed state of the agent into the policy network, and outputting the action performed by the agent; Replace the actions performed by the agent with the actions performed by the multi-agent stored in the experience replay buffer pool; Call the evaluation network to obtain the evaluation value of the next state and action of all agents as a whole.
6. The CTDE-based multi-agent decision-making method according to claim 4, characterized in that: The step of calculating a second loss value of the evaluation network of the multi-agent and updating the evaluation network according to the second loss value includes: The evaluation network is trained through forward propagation, with the input being the state set and action set of all agents, and the output being the evaluation value; The TD mean square error back propagation is used to update the network, and the next observation state of each agent is extracted from the experience replay buffer pool in turn, and input into the respective target policy network to obtain the next action of all agents; Call the target evaluation network to obtain the evaluation value of the next state and action of all agents as a whole.
7. A multi-agent decision-making device based on CTDE, characterized in that: include: An acquisition module is used to acquire multi-agents and initialize the multi-agents using a centralized evaluation network to obtain a multi-agent state, a strategy network, and an evaluation network corresponding to each multi-agent; A training module is used to train the strategy network and evaluation network of the multi-agent according to the preset number of training cycles, training cycle length, experience replay buffer pool and multi-agent state to obtain the trained multi-agent; the multi-agent state is a combination of the states of all agents in the environment at the same time, and a combination of the actions of all agents in the environment; A determination module is used to determine the trained multi-agent as the decision-making strategy of the multi-agent if the loss value of the trained multi-agent is less than a preset value, and the decision-making strategy of the multi-agent is used to adjust the strategy of the multi-agent.
8. The CTDE-based multi-agent decision-making device according to claim 7, characterized in that: The acquisition module is used to: For each agent, select the execution action of each agent according to the current strategy and detection noise; According to the execution actions of each agent, a current reward value is obtained, and the reward value is converted into a new state; The execution action, the corresponding reward value and the new state are stored in the experience replay buffer pool.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 6 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.