A reinforcement learning model training method and system for intelligent agents

By receiving and mixing empirical data of multiple agents in the central training server, training reinforcement learning models and sending predictive operation strategies, the problem of only single agent training in the existing technology is solved, and efficient training of multi-agent and multi-simulation environments is achieved.

CN114117752BActive Publication Date: 2025-06-06HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111326221.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-10
Publication Date
2025-06-06
Estimated Expiration
2041-11-10

AI Technical Summary

Technical Problem

The prior art is difficult to support the reinforcement learning model training of multiple agents, and only single agent training can be achieved.

Method used

By receiving empirical data from multiple agents in the central training server, mixing and storing them in a preset experience pool, when the preset data amount is reached, the reinforcement learning model is trained and the predicted operational strategy is sent to the environment server to enable it to execute the corresponding strategy.

Benefits of technology

Reinforcement learning model training of multiple agents is realized, the accuracy of the reinforcement learning model obtained by training is improved, and efficient training in multi-simulation environments is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114117752B_ABST
    Figure CN114117752B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method and system for training a reinforcement learning model of an intelligent agent, the method comprising: receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server; mixing the experience data of the associated intelligent agents and storing them in a preset experience pool; obtaining the mixed experience data as sample data, and triggering the training of the reinforcement learning model to be trained based on the sample data to obtain output prediction operation strategy information; sending the prediction operation strategy information to the environment server so that the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy; if the preset model training end condition is reached, the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training. That is, the embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple intelligent agents and multiple simulation environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular to a method and system for training a reinforcement learning model of an intelligent agent. Background Art

[0002] Reinforcement learning is one of the paradigms and methodologies of machine learning. It can be used to describe and solve the problem of how an intelligent agent can maximize rewards or achieve specific goals through learning strategies during its interaction with the environment. An intelligent agent refers to a software program or an entity (such as a person, vehicle, or robot) with basic characteristics such as autonomy, sociality, responsiveness, and proactiveness. An intelligent agent can be embedded in the environment, perceive the environment through sensors, and then act on the environment autonomously through effectors.

[0003] The traditional reinforcement learning model training method is: a single agent collects multiple sets of environmental cases in the simulation environment instance database through multiple sets of distributed samplers, exchanges information with the server based on the collected multiple sets of environmental cases, and outputs the trajectory data of the corresponding environmental cases. The server then initializes the agent through the reinforcement learning algorithm module.

[0004] However, traditional reinforcement learning model training methods can only realize single-agent reinforcement learning model training, and do not provide any training methods that support multiple agents. Summary of the invention

[0005] The purpose of the embodiments of the present invention is to provide a method and system for training a reinforcement learning model of an intelligent agent, so as to realize the reinforcement learning model training of multiple intelligent agents.

[0006] In a first aspect, an embodiment of the present invention provides a method for training a reinforcement learning model of an agent, which is applied to a central training server in a reinforcement learning model training system, wherein the system further comprises at least one environment server, each of which runs at least one simulation environment, each of which comprises at least one agent, and the total number of agents is greater than 1, wherein the method comprises:

[0007] Receive the experience data of each agent included in any simulation environment sent by the environment server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located;

[0008] When the amount of the experience data is not less than a first preset amount of data, the experience data of the associated agents are mixed, and the mixed experience data is stored in a preset experience pool;

[0009] When the amount of data in the preset experience pool reaches a second preset amount of data, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0010] Sending the prediction operation strategy information to the environment server so that: the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy, and after executing the prediction operation strategy, sends the status information of each simulation environment to the central training server;

[0011] Receive status information of each simulation environment sent by the environment server, and determine whether a preset model training end condition is met based on the status information of each simulation environment;

[0012] If the preset model training end condition is reached, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training;

[0013] If the preset model training end condition is not met, return to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server.

[0014] Optionally, determining whether a preset model training end condition is met based on the status information of each simulation environment includes:

[0015] Based on the status information of each simulation environment, determining whether each simulation environment in the environment server has completed a preset number of operations;

[0016] If each simulation environment in the environment server has completed the preset number of runs, it is determined that the preset model training end condition has been reached.

[0017] Optionally, when the data amount of the experience data is not less than a first preset data amount, mixing the experience data of the associated agents and storing the mixed experience data in a preset experience pool includes:

[0018] Acquire the association relationship between each intelligent agent from the environment server;

[0019] When the amount of the experience data is not less than the first preset data amount, for each agent, the experience data of the agent associated with the agent and the experience data of the agent are mixed according to the association relationship to obtain mixed experience data, and stored in the preset experience pool corresponding to the agent.

[0020] Optionally, before receiving the experience data of each agent included in any simulation environment sent by the environment server, the method further includes:

[0021] Obtaining configuration information of each of the environment servers;

[0022] Selecting a server for the environment to be configured based on the configuration information;

[0023] Based on the configuration information of the environment server to be configured, an SSH connection is created between the central training server and the environment server to be configured;

[0024] Sending a simulation environment startup instruction to the server of the environment to be configured through an SSH connection, so that the server of the environment to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the central training server returns the transmission port information corresponding to the simulation environment;

[0025] Based on the transmission port information, an information transmission channel is created between the central training server and the simulation environment, and the number of simulation environments running in the environment server to be configured is updated;

[0026] If the number of simulation environments running in the environment server to be configured does not reach the limited number of environments corresponding to the environment server to be configured, return to execute the step of sending a simulation environment startup instruction to the environment server to be configured through the SSH connection; otherwise, stop creating a simulation environment for the environment server to be configured, and return to execute the step of selecting the environment server to be configured based on the configuration information for the remaining environment servers until the number of simulation environments running in each environment server reaches the limited number of environments corresponding to the environment server.

[0027] Optionally, the receiving of the experience data of each agent included in any simulation environment sent by the environment server includes:

[0028] The experience data of each intelligent agent included in any simulation environment sent by the environment server is received through the information transmission channel between each simulation environment in the environment server and the central training server.

[0029] Optionally, the prediction operation strategy corresponding to each simulation environment carries the environment identifier of the simulation environment;

[0030] The sending the prediction operation strategy information to the environment server includes:

[0031] Determine the simulation environment corresponding to each prediction operation strategy based on the environment identifier carried by each prediction operation strategy in the prediction operation strategy information;

[0032] The prediction operation strategy is distributed to the simulation environment corresponding to the environment identifier in the environment server through the information transmission channel between the simulation environment and the central training server, so that the simulation environment in the environment server executes the prediction operation strategy.

[0033] Optionally, before sending the prediction operation strategy information to the environment server, the method further includes:

[0034] When the amount of data in the preset experience pool does not reach the second preset data amount, preset operation strategy information is sent to the environment server, so that the corresponding prediction operation strategy in the preset operation strategy information is executed in the corresponding simulation environment in the environment server, and after executing the prediction operation strategy, the status information of each simulation environment is sent to the central training server, and the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned.

[0035] In a second aspect, an embodiment of the present invention provides a method for training a reinforcement learning model of an intelligent agent, which is applied to any environment server in a reinforcement learning model training system, wherein the system includes a central training server and at least one environment server, each of which runs at least one simulation environment, each of which includes at least one intelligent agent, and the total number of intelligent agents is greater than 1, and the method includes:

[0036] Sending the experience data of each agent included in any simulation environment to the central training server so that the central training server performs the following steps:

[0037] When the amount of data in the preset experience pool reaches a second preset amount of data, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server; the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0038] Receive the predicted operation strategy information sent by the central training server, and make the corresponding simulation environment in the environment server execute the predicted operation strategy corresponding to the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy; wherein the experience data of each intelligent agent includes: the status information of the intelligent agent, the reward information determined by the environment server based on the status information of the intelligent agent, and the operation strategy of the simulation environment where the intelligent agent is located.

[0039] Optionally, before sending the experience data of each agent included in any simulation environment to the central training server, the method further includes:

[0040] Receiving a simulation environment startup instruction sent by the central training server through an SSH connection; wherein the SSH connection is a connection between the central training server and the environment server created based on the configuration information of the environment server;

[0041] Start a simulation environment according to the simulation environment start instruction;

[0042] After the simulation environment is started, the transmission port information corresponding to the simulation environment is returned to the central training server, so that the central training server performs the following steps: based on the transmission port information, an information transmission channel is created between the central training server and the simulation environment, and the number of simulation environments running in the environment server is updated; if the number of simulation environments running in the environment server does not reach the limited number of environments corresponding to the environment server, a step of sending a simulation environment startup instruction to the environment server through an SSH connection is performed; otherwise, the creation of the simulation environment for the environment server is stopped.

[0043] Optionally, the sending of the experience data of each agent included in any simulation environment to the central training server includes:

[0044] Through the information transmission channel between each simulation environment in the environment server and the central training server, the experience data of each intelligent agent included in any simulation environment is sent to the central training server.

[0045] Optionally, the prediction operation strategy corresponding to each simulation environment carries the environment identifier of the simulation environment;

[0046] The receiving the prediction operation strategy information sent by the central training server, and causing the corresponding simulation environment in the environment server to execute the prediction operation strategy corresponding to the prediction operation strategy information, includes:

[0047] Based on the environment identifier carried by each predicted operation strategy in the predicted operation strategy information, the predicted operation strategy carrying the environment identifier in the predicted operation strategy information is received through the information transmission channel between the simulation environment corresponding to the environment identifier in the environment server and the central training server, so that the simulation environment executes the predicted operation strategy.

[0048] Optionally, after any simulation environment is started, it also includes:

[0049] Initialize the simulation environment;

[0050] After issuing the preset operation strategy, obtaining the experience data of the intelligent agent included in the simulation environment;

[0051] Determining whether the simulation environment controls a plurality of agents based on experience data of agents included in the simulation environment;

[0052] If the simulation environment controls multiple agents, convert the experience data of the agents included in the simulation environment into experience data in the form of multiple agents, and send the experience data in the form of multiple agents to the central training server through the information transmission channel between the simulation environment and the central training server; if not, send the experience data of the agents included in the simulation environment to the central training server through the information transmission channel between the simulation environment and the central training server;

[0053] If the simulation environment ends, the information transmission channel between the simulation environment and the central training server is closed;

[0054] If the simulation environment has not yet finished running, receive the predicted operation strategy information sent by the central training server so that the simulation environment executes the predicted operation strategy corresponding to the predicted operation strategy information, and return to execute the step of obtaining the experience data of the intelligent agent included in the simulation environment.

[0055] Optionally, if the simulation environment controls multiple agents, converting the experience data of the agents included in the simulation environment into experience data in a multi-agent form includes:

[0056] If the simulation environment controls multiple agents, determining whether the multiple agents controlled by the simulation environment are asynchronous agents;

[0057] If the multiple agents controlled by the simulation environment are asynchronous agents, the experience data of the agents included in the simulation environment are converted into experience data in a multi-agent form.

[0058] Optionally, after any simulation environment is started, it also includes:

[0059] Determine whether the simulation environment has ended actively;

[0060] If the simulation environment ends its active operation, the simulation environment is deleted from the environment server where the simulation environment is located;

[0061] If the simulation environment does not automatically end, obtain the operation information of the simulation environment;

[0062] Based on the operation information, determining whether the information transmission channel between the simulation environment and the central training server is closed;

[0063] If the information transmission channel between the simulation environment and the central training server is closed, deleting the simulation environment from the environment server where the simulation environment is located;

[0064] If the information transmission channel between the simulation environment and the central training server is not closed, sending an environment destruction request to the central training server through the information transmission channel between the simulation environment and the central training server;

[0065] Receive an environment destruction instruction sent by the central training server, and delete the simulation environment from the environment server where the simulation environment is located according to the environment destruction instruction; wherein the environment destruction instruction is determined by the central training server according to the environment destruction request.

[0066] In a third aspect, an embodiment of the present invention provides a reinforcement learning model training system, including: a central training server and at least one environment server, each of the environment servers runs at least one simulation environment, each simulation environment includes at least one intelligent agent, and the total number of intelligent agents is greater than 1;

[0067] The environment server is used to send the experience data of each agent included in any simulation environment to the central training server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located;

[0068] The central training server is used to receive the experience data of each intelligent agent included in any simulation environment sent by the environment server; when the data volume of the experience data is not less than the first preset data volume, the experience data of the associated intelligent agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data volume in the preset experience pool reaches a second preset data volume, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain the output prediction operation strategy information; the prediction operation strategy information is sent to the environment server, the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met, if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0069] The environment server is also used to receive the predicted operation strategy information sent by the central training server, and enable the corresponding simulation environment in the environment server to execute the corresponding predicted operation strategy in the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy.

[0070] In a fourth aspect, an embodiment of the present invention provides a reinforcement learning model training device for an intelligent agent, which is applied to a central training server in a reinforcement learning model training system, wherein the system further includes at least one environment server, each of which runs at least one simulation environment, each of which includes at least one intelligent agent, and the total number of intelligent agents is greater than 1, and the method includes:

[0071] An experience data receiving module, used to receive the experience data of each agent included in any simulation environment sent by the environment server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located;

[0072] An experience data mixing module, configured to mix the experience data of the associated agents and store the mixed experience data in a preset experience pool when the data amount of the experience data is not less than a first preset data amount;

[0073] A model training module, used for obtaining mixed experience data from the preset experience pool as sample data when the amount of data in the preset experience pool reaches a second preset data amount, and triggering the training of the reinforcement learning model to be trained based on the sample data to obtain output prediction operation strategy information; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0074] A strategy information sending module is used to send the predicted operation strategy information to the environment server, so that: the corresponding simulation environment in the environment server executes the corresponding predicted operation strategy, and sends the status information of each simulation environment to the central training server after executing the predicted operation strategy;

[0075] The end condition determination module is used to receive the status information of each simulation environment sent by the environment server, and determine whether the preset model training end condition is met based on the status information of each simulation environment; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server.

[0076] In a fifth aspect, an embodiment of the present invention provides a reinforcement learning model training device for an agent, which is applied to any environment server in a reinforcement learning model training system, wherein the system includes a central training server and at least one environment server, each of which runs at least one simulation environment, each of which includes at least one agent, and the total number of agents is greater than 1, and the method includes:

[0077] An experience data sending module is used to send the experience data of each intelligent agent included in any simulation environment to the central training server, so that the central training server performs the following steps: when the amount of data in the preset experience pool reaches a second preset data amount, obtain the mixed experience data from the preset experience pool as sample data, and trigger the training of the reinforcement learning model to be trained based on the sample data to obtain the output prediction operation strategy information; send the prediction operation strategy information to the environment server; receive the status information of each simulation environment sent by the environment server, and determine whether the preset model training end condition is met based on the status information of each simulation environment; if the preset model training end condition is met, determine the current reinforcement learning model to be trained as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0078] A strategy information receiving module is used to receive the predicted operation strategy information sent by the central training server, and enable the corresponding simulation environment in the environment server to execute the corresponding predicted operation strategy in the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy; wherein the experience data of each intelligent agent includes: the status information of the intelligent agent, the reward information determined by the environment server based on the status information of the intelligent agent, and the operation strategy of the simulation environment where the intelligent agent is located.

[0079] In a sixth aspect, an embodiment of the present invention provides a server, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0080] Memory, used to store computer programs;

[0081] The processor is used to implement the method steps described in any one of the first aspect or the second aspect when executing the program stored in the memory.

[0082] In a seventh aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of the first aspect or the second aspect are implemented.

[0083] Beneficial effects of the embodiments of the present invention:

[0084] The method provided by the embodiment of the present invention is adopted, by receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server; when the data amount of the experience data is not less than the first preset data amount, the experience data of the associated intelligent agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data amount in the preset experience pool reaches a second preset data amount, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server, so that: the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy, and after executing the prediction operation strategy, the status information of each simulation environment is sent to the central training server; the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned. That is, the embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of related agents, the experience data can be shared between agents, thereby improving the accuracy of the trained reinforcement learning model.

[0085] Of course, it is not necessary to achieve all of the advantages described above at the same time to implement any product or method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.

[0087] Figure 1 A schematic diagram of the structure of a reinforcement learning model training system provided by an embodiment of the present invention;

[0088] Figure 2 Another structural schematic diagram of the reinforcement learning model training system provided by an embodiment of the present invention;

[0089] Figure 3 A schematic diagram of another structure of the reinforcement learning model training system provided by an embodiment of the present invention;

[0090] Figure 4 A flowchart of a reinforcement learning model training method for an intelligent agent provided by an embodiment of the present invention;

[0091] Figure 5 A flow chart of a simulation environment startup method provided by an embodiment of the present invention;

[0092] Figure 6 A flow chart of a reinforcement learning model training method for an intelligent agent of a central training server in a reinforcement learning model training system provided by an embodiment of the present invention;

[0093] Figure 7 A flow chart of the environment operation provided by the embodiment of the present invention;

[0094] Figure 8 A flow chart of environment destruction provided by an embodiment of the present invention;

[0095] Fig. 9 A structural diagram of a distributed multi-agent reinforcement learning model training provided by an embodiment of the present invention;

[0096] Fig.10 A structural schematic diagram of a reinforcement learning model training device for an intelligent agent of a central training server in a reinforcement learning model training system provided by an embodiment of the present invention;

[0097] Fig.11 A structural schematic diagram of a reinforcement learning model training device for an intelligent agent applied to any environment server in a reinforcement learning model training system provided by an embodiment of the present invention;

[0098] Fig.12 A schematic diagram of the structure of a server provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0099] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field based on this application belong to the scope of protection of the present invention.

[0100] In order to realize the reinforcement learning model training of multiple intelligent agents, the embodiments of the present invention provide a reinforcement learning model training method, system, server, computer-readable storage medium and computer program product for an intelligent agent.

[0101] The concepts involved in the embodiments of the present invention are described below:

[0102] Distributed: Multiple parallel tasks are distributed on multiple servers. Multiple server nodes can be used to execute multiple tasks in parallel to improve task execution efficiency.

[0103] Reinforcement learning: also known as reinforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning, which is used to describe and solve the problem of how intelligent agents can maximize rewards or achieve specific goals through learning strategies during their interaction with the environment.

[0104] Environment: The environment in which the reinforcement learning model interacts, which is different from the real environment and can be a game or a simulation software. Specifically, the environment can be a simulation environment including multiple intersections.

[0105] Simulation: Simulation is software used to simulate the real environment, such as traffic simulation Sumo, which is a simulation software that can simulate large-scale road networks and the operation of a large number of vehicles. It can obtain status information of entities such as vehicles, roads, and intersections on the road.

[0106] Intelligent agent: An intelligent agent refers to an entity with the basic characteristics of autonomy, sociality, responsiveness and proactivity. It can be regarded as a corresponding software program or an entity (such as a person, vehicle, robot, etc.). The intelligent agent is embedded in the environment, perceives the environment through sensors, acts on the environment autonomously through effectors and meets the design requirements.

[0107] Multi-agent: that is, multiple agents, which may compete or cooperate with each other.

[0108] Multi-agent reinforcement learning: Each agent in the multi-agent system obtains rewards by interacting with the environment to improve its own learning behavior, thereby obtaining the optimal operating strategy in the environment.

[0109] Simulation environment: any simulation environment. If there is no multi-agent, it is a single-agent simulation environment. If there is multi-agent, it is a multi-agent simulation environment. For example, in the traffic signal control simulation, if there is only one intersection in the road network, it is a single-agent environment, and no distinction is needed. If there are multiple intersections in the road network, it is a multi-agent environment. If only one intersection is controlled, it is still a single-agent environment, and other intersections operate according to the default signal timing without additional control.

[0110] Observation: includes an observation of the environment, including the vehicle's position, speed, whether the intersection allows release, simulation time, etc.

[0111] Environmental state: The effective information obtained after processing the observation according to specific requirements, generally a set of data required for the training of the reinforcement learning model. The observation includes an observation of the simulation environment, which can be the vehicle's position, speed, whether the intersection is allowed to be released, and the simulation time.

[0112] Reward: The reward calculated by observation or environment state, which is a scalar value. For example, the reward for a congested intersection is -5 points, the reward for a smooth intersection is 20 points, and the reward for a normal intersection is 1 point.

[0113] The operation strategy of the simulation environment: assign a behavior to the simulation environment, such as the moving direction and distance of up, down, left, and right in the game. For example, if the simulation environment is an intersection, and the actions given to the simulation environment are east-west and north-south travel times [30, 30], then the green light duration in the east-west direction of the intersection is 30 seconds, while the red light duration in the north-south direction is 30 seconds. In addition, the green light duration in the north-south direction of the intersection is 30 seconds, while the red light duration in the east-west direction is also 30 seconds.

[0114] The following first introduces the reinforcement learning model training system provided by an embodiment of the present invention. Figure 1 A structural diagram of a reinforcement learning model training system provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the reinforcement learning model training system includes: a central training server 110 and at least one environment server 120, each of the environment servers 120 runs at least one simulation environment, and each simulation environment includes at least one intelligent agent;

[0115] The environment server 120 is used to send the experience data of each agent included in any simulation environment to the central training server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located;

[0116] The central training server 110 is used to receive the experience data of each agent included in any simulation environment sent by the environment server; when the data volume of the experience data is not less than the first preset data volume, the experience data of the associated agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data volume in the preset experience pool reaches a second preset data volume, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server, and the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met, and if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, the step of receiving the experience data of each agent included in any simulation environment sent by the environment server is returned; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0117] The environment server 120 is also used to receive the predicted operation strategy information sent by the central training server, and enable the corresponding simulation environment in the environment server to execute the corresponding predicted operation strategy in the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy.

[0118] In the system provided by the embodiment of the present invention, the central training server receives the experience data of each intelligent agent included in any simulation environment sent by the environment server; when the data volume of the experience data is not less than the first preset data volume, the experience data of the associated intelligent agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data volume in the preset experience pool reaches a second preset data volume, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server The device is configured to: execute the corresponding prediction operation strategy in the simulation environment of the environment server, and send the status information of each simulation environment to the central training server after executing the prediction operation strategy; receive the status information of each simulation environment sent by the environment server, and determine whether the preset model training end condition is met based on the status information of each simulation environment, and if the preset model training end condition is met, determine the current reinforcement learning model to be trained as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server. That is, the embodiment of the present invention proposes a new efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of related agents, the experience data can be shared between agents, thereby improving the accuracy of the reinforcement learning model obtained by training.

[0119] In a possible implementation, the central training server is further configured to receive experience data of each agent included in any simulation environment sent by the environment server through an information transmission channel between each simulation environment in the environment server and the central training server.

[0120] In one possible implementation, Figure 2 Another structural diagram of the reinforcement learning model training system provided by the embodiment of the present invention is as follows Figure 2 As shown, the central training server 110 includes a multi-agent reinforcement learning module 210;

[0121] The multi-agent reinforcement learning module 210 is used to determine whether each simulation environment in the environment server has completed a preset number of operations based on the status information of each simulation environment; if each simulation environment in the environment server has completed a preset number of operations, it is determined that the preset model training end condition is reached, and the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training.

[0122] The preset number of times can be set to 100 times or 50 times etc. according to actual application.

[0123] In this embodiment, after the environment server receives the prediction operation strategy information sent by the central training server, the corresponding simulation environment in the environment server can execute the corresponding prediction operation strategy, and after the corresponding simulation environment in the environment server completes the execution of the prediction operation strategy, the environment server can send the status information of each simulation environment to the central training server; after receiving the status information of each simulation environment sent by the environment server, the central training server can determine whether the preset model training end condition is met based on the status information of each simulation environment. Specifically, the central training server can determine whether each simulation environment in the environment server has completed the preset number of operations based on the status information of each simulation environment; if each simulation environment in the environment server has completed the preset number of operations, it is determined that the preset model training end condition is met.

[0124] In another possible implementation, the multi-agent reinforcement learning module 210 is also used to obtain the association relationship between each agent from the environment server; when the data volume of the experience data is not less than the first preset data volume, for each agent, according to the association relationship, the experience data of the agent associated with the agent and the experience data of the agent are mixed to obtain mixed experience data, and stored in the preset experience pool corresponding to the agent.

[0125] The first preset data volume can be flexibly set according to actual applications and is not specifically limited here.

[0126] In another possible implementation, Figure 2 As shown, the central training server also includes a distributed simulation environment management module 220;

[0127] The distributed simulation environment management module 220 is used to obtain configuration information of each of the environment servers; select an environment server to be configured based on the configuration information; create an SSH connection between the central training server and the environment server to be configured based on the configuration information of the environment server to be configured; send a simulation environment startup instruction to the environment server to be configured through the SSH connection, so that the environment server to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the central training server returns the transmission port information corresponding to the simulation environment; based on the transmission port information, create an information transmission channel between the central training server and the simulation environment, and update the number of simulation environments running in the environment server to be configured; if the number of simulation environments running in the environment server to be configured does not reach the limited number of environments corresponding to the environment server to be configured, return to the step of sending the simulation environment startup instruction to the environment server to be configured through the SSH connection; otherwise, stop creating the simulation environment for the environment server to be configured, and return to the step of selecting the environment server to be configured based on the configuration information for the remaining environment servers until the number of simulation environments running in each environment server reaches the limited number of environments corresponding to the environment server.

[0128] In the embodiment of the present invention, a main process of reinforcement learning model training is run in the central training server, and the central training server may include a multi-agent reinforcement learning module and a distributed simulation environment management module. Among them, the multi-agent reinforcement learning module includes any multi-agent reinforcement learning model.

[0129] Each simulation environment can output the current state information S (State). The state information of each simulation environment can specifically include the experience data of each agent in the simulation environment. For example, the intersection of Road 1 and Road 2 is intersection a. If simulation environment 1 is an environment including intersection 1, Road 1 and Road 2, where intersection 1 is an agent, then the current state information S of simulation environment 1 can include: the traffic flow, vehicle queue data and vehicle speed of intersection 1, etc., where if the traffic flow of intersection 1 is expressed as [23,41,4 ,22], which means that the traffic volume going south through intersection 1 is 23, the traffic volume going north through intersection 1 is 41, the traffic volume going east through intersection 1 is 4, and the traffic volume going west through intersection 1 is 22; if the vehicle queue data of intersection 1 is expressed as [12,13,1,3], it means that the number of vehicles queuing south through intersection 1 is 12, the number of vehicles queuing north through intersection 1 is 13, the number of vehicles queuing east through intersection 1 is 1, and the number of vehicles queuing west through intersection 1 is 3.

[0130] The environment server can use the preset function rules to calculate the corresponding reward R (Reward) according to the state data of the simulation environment. The reward is a scalar value, and the higher the value, the higher the reward. For example, for intersection 1, according to the vehicle queue data, if the number of queued vehicles is > 10, the reward is -10 (a negative reward can be regarded as a penalty); if the number of queued vehicles is between 4-10, the reward is 3; if the number of queued vehicles is between 0-4, the reward is 20. Of course, the specific reward rules can be determined according to the specific environment. This is just an example without specific limitations.

[0131] The central training server can send prediction operation strategy information to the simulation environment. The prediction operation strategy information can be action information A (Action), that is, sending actions to the simulation environment. For example, in traffic signal control, first north-south traffic for 20 seconds and then east-west traffic for 40 seconds, that is, the central training server can send the action [20,40] to the simulation environment, and the simulation environment can execute the action of first north-south traffic for 20 seconds and then east-west traffic for 40 seconds as the input for the next environment operation.

[0132] In the embodiment of the present invention, the experience data of each agent may include: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located. Specifically, the experience data of each agent may include [S, R, A], where S is the current state information output by the simulation environment where the agent is located, R is the corresponding reward R calculated by the environment server for the agent based on the state data of the simulation environment, and A is the operation strategy of the simulation environment where the agent is located.

[0133] The central training server runs the main process of reinforcement learning model training and can obtain the association relationship between each intelligent agent from the environment server; when the data amount of the experience data is not less than the first preset data amount, for each intelligent agent, according to the association relationship, the experience data of the intelligent agent associated with the intelligent agent and the experience data of the intelligent agent are mixed to obtain mixed experience data, and stored in the preset experience pool corresponding to the intelligent agent.

[0134] Specifically, the central training server can obtain the association relationship between each intelligent agent according to the initial configuration by initializing the association relationship between each intelligent agent, or obtain the association relationship between each intelligent agent through the complex network relationship provided by the simulation environment. After receiving the experience data of the intelligent agent from the distributed process, the central training server can store the experience data of the intelligent agent in the specified variable and record the most recent experience data. Then, for each intelligent agent, the central training server can determine whether the intelligent agent is related to other intelligent agents according to the association relationship between each intelligent agent. If it is related, it will mix the state with other intelligent agents according to the predetermined rules, and store the mixed experience data in the preset experience pool corresponding to the intelligent agent; if it is not related, the experience data of the intelligent agent is directly stored in the preset experience pool corresponding to the intelligent agent. Specifically, the experience data of each intelligent agent carries the intelligent agent identification of the intelligent agent. The central training server can store the experience data of each intelligent agent in the preset experience pool corresponding to the intelligent agent according to the intelligent agent identification, or store the experience data of the intelligent agent mixed in the preset experience pool corresponding to the intelligent agent.

[0135] Then, the central training server can determine whether the number of preset experience pools corresponding to each agent has reached the limit. If the limit is reached, it will be deleted in order, or deleted according to the reward of the experience data, and low-quality experience data with low rewards will be eliminated. Among them, the number limit of the preset experience pool corresponding to each agent can be the same or different, and is not specifically limited here.

[0136] For example, if the number of preset experience pools corresponding to agent a is limited to 100,000, when the number of preset experience pools corresponding to agent a reaches the limit, the central training server can delete some experience data from it according to the weight, such as deleting experience data with too low rewards and leaving experience data with higher rewards; it can also delete older experience data in order, and it can be assumed that the old experience data has been used or the rewards are generally low.

[0137] In a possible implementation, the central training server may also include a multi-agent experience state mixing module.

[0138] There are relationships between multiple agents, so the experience data of the agents can be designed according to the state of reinforcement learning and combined on demand. For example, array-type experience data can be spliced, and digital experience data can be summed or weighted summed. The specific experience data mixing method can be determined according to the actual application needs. For example, the agent is intersection 1. The traffic flow of intersection 1 will be affected by the traffic flow of the previous intersection 2, that is, intersection 1 is associated with its previous intersection. Then, the mixed experience data can be formed by mixing the traffic flow data of the two intersections. In the case of a single agent, intersection 1 only contains its own traffic flow [23,12,31,51], but due to the influence of intersection 2, the central training server can mix the experience data of intersection 1 and intersection 2 to obtain the mixed experience data of intersection 1 [23,12,31,51,12,11,33,21]. When the reinforcement learning model is trained, intersection 1 includes the traffic flow information of intersection 2. The optimization process of reinforcement learning model training will adjust the actions of intersection 1 and intersection 2, so that intersection 1 can achieve the best under various intersection 1 and intersection 2 conditions. Table 1 below is a schematic diagram of the traffic flow of intersection 1 and intersection 2.

[0139] Table 1: Traffic flow diagram

[0140] East South West north Traffic flow at intersection 1 23 12 31 51 Traffic flow at intersection 2 12 11 33 21

[0141] In the embodiment of the present invention, when the amount of data in the preset experience pool reaches the second preset data amount, the central training server can obtain mixed experience data from the preset experience pool as sample data, and trigger the training of the reinforcement learning model to be trained based on the sample data to obtain output prediction operation strategy information. The second preset data amount can be specifically set according to the actual application situation, for example, set to 64.

[0142] Specifically, the central training server may obtain mixed experience data from a preset experience pool as sample data by random sampling. For example, 64 sample data may be obtained by sampling from any preset experience pool to trigger the training of the reinforcement learning model to be trained.

[0143] In one possible implementation, an environment server can run multiple simulation environments, wherein the premise for the operation of a distributed simulation environment is that the software packages that need to be run must be distributed to all environment servers that need to run the simulation environment, and the number of simulation environments running on the environment server must be kept consistent as much as possible.

[0144] There can be multiple simulation environment processes in each environment server, each simulation environment process can run a simulation environment, and each simulation environment can include one or more intelligent agents.

[0145] The simulation environment process can include two parts: a multi-agent adaptation module and a simulation environment. Among them, the simulation environment is any simulation environment. If there is no multi-agent, it is a single-agent simulation environment. If there is multi-agent, it is a single-agent simulation environment. For example, in the traffic signal control simulation: if there is only one intersection on the road network, it is a single-agent and no distinction is needed. If it is a road network with multiple intersections and multiple intersections need to be controlled, it is a multi-agent. If only one intersection is controlled, it is still a single-agent, and other intersections operate according to the default signal timing without additional control. Several key processes of the simulation environment are as follows: environment creation, environment operation, and environment destruction.

[0146] The process of creating a simulation environment may include:

[0147] The distributed simulation environment management module obtains configuration information of each of the environment servers; selects an environment server to be configured based on the configuration information; creates an SSH connection between the central training server and the environment server to be configured based on the configuration information of the environment server to be configured; sends a simulation environment startup instruction to the environment server to be configured through the SSH connection, so that the environment server to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the central training server returns the transmission port information corresponding to the simulation environment; based on the transmission port information, creates an information transmission channel between the central training server and the simulation environment, and updates the number of simulation environments running in the environment server to be configured; if the number of simulation environments running in the environment server to be configured does not reach the limited number of environments corresponding to the environment server to be configured, returns to the step of sending the simulation environment startup instruction to the environment server to be configured through the SSH connection; otherwise, stops creating the simulation environment for the environment server to be configured, and returns to the step of selecting the environment server to be configured based on the configuration information for the remaining environment servers until the number of simulation environments running in each environment server reaches the limited number of environments corresponding to the environment server.

[0148] Specifically, the distributed simulation environment management module can initialize the configuration and obtain the configuration information of all server nodes in the cluster, including server IP, user name, password, Python interpreter path, package path and server node resources, etc. The distributed simulation environment management module can maintain an environment server list, and initialize the number of simulation environments of all environment servers to 0. The distributed simulation environment management module can select an environment server in the environment server list as the environment server to be configured according to the configuration information. Specifically, if there is an environment server with the smallest number of simulation environments, the environment server with the smallest number of simulation environments can be selected as the environment server to be configured. If the number of simulation environments of all environment servers is the same, an environment server can be selected in order as the environment server to be configured.

[0149] The distributed simulation environment management module can create an SSH connection between the central training server and the environment server to be configured according to the configuration information of the environment server to be configured, such as IP, user name and password, connect to the remote central training server node through SSH, run the Python interpreter and program, and then send a simulation environment startup instruction to the environment server to be configured through the SSH connection, so that the environment server to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the transmission port information corresponding to the simulation environment can be returned to the central training server. Then, the distributed simulation environment management module can create an information transmission channel between the central training server and the simulation environment based on the transmission port information, and update the number of simulation environments running in the environment server to be configured; if the number of simulation environments running in the environment server to be configured does not reach the number of limited environments corresponding to the environment server to be configured, the distributed simulation environment management module continues to send the simulation environment startup instruction to the environment server to be configured through the SSH connection to create a new simulation environment; if the number of simulation environments running in the environment server to be configured reaches the number of limited environments corresponding to the environment server to be configured, then stop creating a simulation environment for the environment server to be configured. Then, the distributed simulation environment management module can select a new environment server to be configured from the remaining environment servers until the number of simulation environments running in each environment server reaches the number of restricted environments corresponding to the environment server. The number of restricted environments corresponding to the environment servers can be the same or different, and can be set according to actual applications, which is not specifically limited here.

[0150] Distributed simulation environments can be created in batches. For example, if a total of 1,000 simulation environments need to be run, 10 simulation environments are created in each batch. After all 10 simulation environments have finished running, 10 new simulation environments are created. This ensures that there is always only one batch of simulation environments running, which can effectively control the number of resources.

[0151] In one possible implementation, Figure 3 Another structural diagram of the reinforcement learning model training system provided by the embodiment of the present invention is as follows: Figure 3 As shown, the distributed simulation environment management module 220 includes:

[0152] The experience collection submodule 310 is used to receive the experience data of each agent included in any simulation environment sent by the environment server through the information transmission channel between each simulation environment in the environment server and the central training server;

[0153] The strategy distribution submodule 320 is used to determine the simulation environment corresponding to the prediction operation strategy based on the environment identifier carried by each prediction operation strategy in the prediction operation strategy information; and distribute the prediction operation strategy to the simulation environment corresponding to the environment identifier in the environment server through the information transmission channel between the simulation environment and the central training server, so that the simulation environment in the environment server executes the prediction operation strategy.

[0154] Specifically, each simulation environment communicates bidirectionally with the central training process created by the central training server. The simulation environment management module in the central training process is responsible for integrating the information of all processes, storing the environment status information, and distributing the predicted operation strategy information to each simulation process in the environment server.

[0155] In a possible implementation, the environment server is further used to initialize the simulation environment after any simulation environment is started; after issuing a preset operation strategy, obtain the experience data of the intelligent agent included in the simulation environment; based on the experience data of the intelligent agent included in the simulation environment, determine whether the simulation environment controls multiple intelligent agents; if the simulation environment controls multiple intelligent agents, convert the experience data of the intelligent agent included in the simulation environment into experience data in the form of multiple intelligent agents, and send the experience data in the form of multiple intelligent agents to the central training server through the information transmission channel between the simulation environment and the central training server; if not, send the experience data of the intelligent agent included in the simulation environment to the central training server through the information transmission channel between the simulation environment and the central training server; if the simulation environment ends, close the information transmission channel between the simulation environment and the central training server; if the simulation environment does not end, receive the predicted operation strategy information sent by the central training server, so that the simulation environment executes the corresponding predicted operation strategy in the predicted operation strategy information, and return to execute the step of obtaining the experience data of the intelligent agent included in the simulation environment.

[0156] Specifically, if the simulation environment controls multiple agents, the experience data of the agents included in the simulation environment are converted into experience data in the form of multi-agents. For example, for each agent, the experience data of the agent associated with it is mixed with its own experience data as the experience data of the agent in the form of multi-agents.

[0157] Specifically, each simulation environment runs independently in its own process, and the running process of any simulation environment may include the following process steps A1 to A6:

[0158] Step A1: The simulation environment runs and initializes the configuration information required by itself.

[0159] For example, the configuration information may include the traffic simulation software SUMO loading the road network and the simulation software configuration;

[0160] Step A2: Start the environment simulation software and send initialization actions to itself.

[0161] The purpose of this step is to run the simulation environment and collect enough initial information. For example, let the traffic simulation run for 5 cycles according to a fixed timing plan [30,30] to warm up the road network and fill the road with enough vehicles. Otherwise, the initialization state may be the information on an empty road, which has no valid value. Usually, the above steps A1 and A2 are generally combined into a reset, which is the simulation environment initialization and reset operation. Reset will return the initial information of the simulation environment to facilitate judgment and decision-making for reinforcement learning model training.

[0162] Step A3: If it is multi-agent, encapsulate the agent's experience data into a multi-agent form according to the multi-agent adaptation logic.

[0163] For example, encapsulation can be accomplished through a dictionary data structure, through which multiple agents can be filled in, or only the information of one agent can be filled in.

[0164] Step A4: Send the simulation status information to the information transmission channel between the simulation environment and the central training server.

[0165] The simulation status information is sent to the established socket pipeline, and the central training server can receive the environment status information.

[0166] Step A5: Determine whether the simulation environment has finished running.

[0167] The conditions for determining the end are provided by the simulation initialization parameters. For example, in traffic simulation, the end condition is the running time limit. For example, if the road traffic simulation is limited to 2 hours, the simulation environment running process will be ended after the simulation environment running time reaches 2 hours.

[0168] Step A6: If the simulation environment running process has not ended, continue to monitor the information transmission channel between the simulation environment and the central training server, and receive the prediction running strategy information sent by the central training server.

[0169] Specifically, after receiving the environment status information, the central training server will predict the predicted operation strategy information based on the environment status information, and send the predicted operation strategy information to the simulation environment. After receiving the predicted operation strategy information, the simulation environment can continue to execute the process of steps A2-A6 until the simulation environment runs to the end.

[0170] In one possible implementation mode, the environment server is specifically used to determine whether the multiple agents controlled by the simulation environment are asynchronous agents if the simulation environment controls multiple agents; if the multiple agents controlled by the simulation environment are asynchronous agents, convert the experience data of the agents included in the simulation environment into experience data in a multi-agent form.

[0171] Specifically, when the simulation environment is running, multiple agents may be synchronous or asynchronous. For example, in a table tennis game, the states of two players can be obtained and returned at the same time, which is a synchronous situation. In signal control, multiple agents are asynchronous. Intersection A runs for 60 seconds, and intersection B may run for 70 seconds. They will return their respective states at 60 seconds and 70 seconds respectively.

[0172] The environment server may include a multi-agent adaptation module. Specifically, if it is a single agent, the environment server will not load the multi-agent adaptation module when creating the environment; if multiple agents are specified when creating the simulation environment, the environment server will automatically load the multi-agent adaptation module to complete the adaptation of dynamic functions.

[0173] The multi-agent adaptation module includes the following functions: default action delivery, multi-agent action conversion, and multi-agent state conversion.

[0174] Specifically, the default action delivery function is: when the simulation environment fails to train or needs to be explored, the default action delivery module can be used to generate a fixed, random or expert algorithm-based, such as SQP linear programming method.

[0175] The multi-agent action conversion function is: for synchronous agents, no conversion is required, and the operation strategies of multiple agents can be received at the same time; for asynchronous agents, only one agent's operation strategy can be received at a time, so conversion is required and then sent to the simulation environment.

[0176] The multi-agent state transition function is to package the agent into a dictionary, which will be parsed at the upper layer and then placed into their respective multi-agent queues.

[0177] In a possible implementation, the process of destroying the simulation environment may be: the environment server is specifically used to determine whether the simulation environment has actively ended its operation after any simulation environment is started; if the simulation environment has actively ended its operation, the simulation environment is deleted from the environment server where the simulation environment is located; if the simulation environment has not actively ended its operation, the operation information of the simulation environment is obtained; based on the operation information, it is determined whether the information transmission channel between the simulation environment and the central training server is closed; if the information transmission channel between the simulation environment and the central training server is closed, the simulation environment is deleted from the environment server where the simulation environment is located; if the information transmission channel between the simulation environment and the central training server is not closed, an environment destruction request is sent to the central training server through the information transmission channel between the simulation environment and the central training server; an environment destruction instruction sent by the central training server is received, and the simulation environment is deleted from the environment server where the simulation environment is located according to the environment destruction instruction; wherein the environment destruction instruction is determined by the central training server according to the environment destruction request.

[0178] Specifically, the destruction of the simulation environment may include two situations.

[0179] Case 1: The process ends actively, that is, the process exits after the simulation environment finishes running on its own. This case also includes the process ending due to other abnormal situations.

[0180] Case 2: The central training process ends, and other distributed simulation environment processes need to be terminated to prevent process resource leakage. For example, if external factors affect the central training process and require it to end, other distributed simulation environment processes must also be terminated.

[0181] Specifically, in the embodiment of the present invention, the result of Done=True sent by the simulation environment running process can be used to determine whether the simulation environment running process has actively ended. If the simulation environment process sends the result of Done=True, it can be determined that the simulation environment running process has actively ended.

[0182] The embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of the associated agents, the experience data can be shared between the agents, thereby improving the accuracy of the reinforcement learning model obtained by training. In addition, by using the system provided by the embodiment of the present invention, by decoupling the experience pool, the reinforcement learning model and the simulation environment, any existing offline strategy reinforcement learning algorithm can be replaced at any time, and the experience pool can collect sample data of each step in real time, so that the use of sample data is more timely. In addition, the training environment of the reinforcement learning model can support multi-process and distributed coexistence, which can fully utilize the performance of a single server and also perform horizontal performance expansion, thereby improving the utilization rate of server resources. In addition, the training of the reinforcement learning model is not performed in parallel, and the number of samples can be greatly increased while ensuring the quality of the sample data, thereby improving the convergence speed of the algorithm. The central training server can use the CPU or the GPU for reinforcement learning model training, which can give full play to the hardware performance and improve the training speed.

[0183] Corresponding to the above reinforcement learning model training system, the embodiment of the present invention further provides a reinforcement learning model training method for an intelligent agent. The reinforcement learning model training method for an intelligent agent provided by the embodiment of the present invention is introduced below. Figure 4 A flowchart of a method for training a reinforcement learning model of an agent provided in an embodiment of the present invention. The method can be applied to a central training server in a reinforcement learning model training system. The system further includes at least one environment server, each of which runs at least one simulation environment, each of which includes at least one agent, and the total number of agents is greater than 1. Figure 4 As shown, the method includes:

[0184] S401, receiving experience data of each agent included in any simulation environment sent by the environment server.

[0185] The experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located.

[0186] S402, when the data amount of the experience data is not less than a first preset data amount, the experience data of the associated agents are mixed, and the mixed experience data is stored in a preset experience pool.

[0187] S403, when the amount of data in the preset experience pool reaches a second preset data amount, obtain the mixed experience data from the preset experience pool as sample data, and trigger the training of the reinforcement learning model to be trained based on the sample data to obtain the output prediction operation strategy information.

[0188] The predicted operation strategy information includes the predicted operation strategy of the corresponding simulation environment in the environment server.

[0189] S404, sending the prediction operation strategy information to the environment server, so that: the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy, and sends the status information of each simulation environment to the central training server after executing the prediction operation strategy.

[0190] S405, receiving status information of each simulation environment sent by the environment server, and determining whether a preset model training end condition is met based on the status information of each simulation environment.

[0191] S406: If the preset model training end condition is reached, the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training.

[0192] S407, if the preset model training end condition is not met, return to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server.

[0193] The method provided by the embodiment of the present invention is adopted, by receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server; when the data amount of the experience data is not less than the first preset data amount, the experience data of the associated intelligent agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data amount in the preset experience pool reaches a second preset data amount, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server, so that: the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy, and after executing the prediction operation strategy, the status information of each simulation environment is sent to the central training server; the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned. That is, the embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of related agents, the experience data can be shared between agents, thereby improving the accuracy of the trained reinforcement learning model.

[0194] In a possible implementation, determining whether a preset model training termination condition is met based on the status information of each simulation environment includes: determining whether each simulation environment in the environment server has completed a preset number of operations based on the status information of each simulation environment; if each simulation environment in the environment server has completed a preset number of operations, determining that a preset model training termination condition is met.

[0195] In a possible implementation, when the data amount of the experience data is not less than the first preset data amount, mixing the experience data of the associated agents and storing the mixed experience data in a preset experience pool includes the following steps B1-B2:

[0196] Step B1: Obtaining the association relationship between each intelligent agent from the environment server;

[0197] Step B2: When the amount of the experience data is not less than the first preset data amount, for each agent, according to the association relationship, the experience data of the agent associated with the agent and the experience data of the agent are mixed to obtain mixed experience data, and stored in the preset experience pool corresponding to the agent.

[0198] In another possible implementation, Figure 5 A flow chart of a simulation environment startup method provided by an embodiment of the present invention, such as Figure 5 As shown, before receiving the experience data of each agent included in any simulation environment sent by the environment server, the method further includes:

[0199] S501: Obtain configuration information of each of the environment servers.

[0200] S502: Select an environment server to be configured based on the configuration information.

[0201] S503: Based on the configuration information of the environment server to be configured, create an SSH connection between the central training server and the environment server to be configured.

[0202] S504, sending a simulation environment startup instruction to the environment server to be configured through an SSH connection, so that the environment server to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the central training server returns the transmission port information corresponding to the simulation environment.

[0203] S505: Based on the transmission port information, an information transmission channel is created between the central training server and the simulation environment, and the number of simulation environments running in the environment server to be configured is updated.

[0204] S506: If the number of simulation environments running in the environment server to be configured does not reach the limited number of environments corresponding to the environment server to be configured, return to execute the step of sending the simulation environment startup instruction to the environment server to be configured through the SSH connection; otherwise, stop creating the simulation environment for the environment server to be configured, and return to execute the step of selecting the environment server to be configured based on the configuration information for the remaining environment servers until the number of simulation environments running in each environment server reaches the limited number of environments corresponding to the environment server.

[0205] In a possible implementation, the receiving of the experience data of each intelligent agent included in any simulation environment sent by the environment server includes: receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server through an information transmission channel between each simulation environment in the environment server and the central training server.

[0206] In one possible implementation, the prediction operation strategy corresponding to each simulation environment carries the environment identifier of the simulation environment; sending the prediction operation strategy information to the environment server includes: determining the simulation environment corresponding to the prediction operation strategy based on the environment identifier carried by each prediction operation strategy in the prediction operation strategy information; distributing the prediction operation strategy to the simulation environment corresponding to the environment identifier in the environment server through the information transmission channel between the simulation environment and the central training server, so that the simulation environment in the environment server executes the prediction operation strategy.

[0207] In a possible implementation, before sending the predicted operation strategy information to the environment server, it also includes: when the amount of data in the preset experience pool does not reach the second preset data amount, sending the preset operation strategy information to the environment server, so that the corresponding prediction operation strategy in the preset operation strategy information is executed in the corresponding simulation environment in the environment server, and after executing the predicted operation strategy, the status information of each simulation environment is sent to the central training server, and the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned.

[0208] The embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of the associated agents, the experience data can be shared between the agents, thereby improving the accuracy of the reinforcement learning model obtained by training. In addition, by adopting the method provided by the embodiment of the present invention, by decoupling the experience pool, the reinforcement learning model and the simulation environment, any existing offline strategy reinforcement learning algorithm can be replaced at any time, and the experience pool can collect sample data of each step in real time, so that the use of sample data is more timely. In addition, the training environment of the reinforcement learning model can support multi-process and distributed coexistence, which can fully utilize the performance of a single server and perform horizontal performance expansion, thereby improving the utilization rate of server resources. In addition, the training of the reinforcement learning model is not performed in parallel, and the number of samples can be greatly increased while ensuring the quality of the sample data, thereby improving the convergence speed of the algorithm. The central training server can use the CPU or the GPU for reinforcement learning model training, which can give full play to the hardware performance and improve the training speed.

[0209] Corresponding to the above reinforcement learning model training system, the embodiment of the present invention further provides a reinforcement learning model training method for an intelligent agent. The reinforcement learning model training method for an intelligent agent provided by the embodiment of the present invention is introduced below. Figure 6 A flow chart of a reinforcement learning model training method for an agent of a central training server in a reinforcement learning model training system provided by an embodiment of the present invention, wherein the system further comprises at least one environment server, each of which runs at least one simulation environment, each of which comprises at least one agent, and the total number of agents is greater than 1. Figure 6 As shown, the method includes:

[0210] S601, sending the experience data of each agent included in any simulation environment to the central training server, so that the central training server executes step C1:

[0211] Step C1: When the amount of data in the preset experience pool reaches a second preset data amount, obtain mixed experience data from the preset experience pool as sample data, and trigger the training of the reinforcement learning model to be trained based on the sample data to obtain output prediction operation strategy information; send the prediction operation strategy information to the environment server; receive the status information of each simulation environment sent by the environment server, and determine whether the preset model training end condition is met based on the status information of each simulation environment; if the preset model training end condition is met, determine the current reinforcement learning model to be trained as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server.

[0212] S602, receiving the predicted operation strategy information sent by the central training server, and making the corresponding simulation environment in the environment server execute the predicted operation strategy corresponding to the predicted operation strategy information, and sending the status information of each simulation environment to the central training server after executing the predicted operation strategy; wherein the experience data of each intelligent agent includes: the status information of the intelligent agent, the reward information determined by the environment server based on the status information of the intelligent agent, and the operation strategy of the simulation environment where the intelligent agent is located.

[0213] The embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of related agents, the experience data can be shared between agents, thereby improving the accuracy of the trained reinforcement learning model.

[0214] In a possible implementation, before sending the experience data of each intelligent agent included in any simulation environment to the central training server, the method further includes: receiving a simulation environment startup instruction sent by the central training server through an SSH connection; wherein the SSH connection is a connection between the central training server and the environment server created based on the configuration information of the environment server; starting a simulation environment according to the simulation environment startup instruction; and returning the transmission port information corresponding to the simulation environment to the central training server after the simulation environment is started, so that the central training server performs the following steps: creating an information transmission channel between the central training server and the simulation environment based on the transmission port information, and updating the number of simulation environments running in the environment server, if the number of simulation environments running in the environment server does not reach the limited number of environments corresponding to the environment server, executing the step of sending the simulation environment startup instruction to the environment server through the SSH connection, otherwise, stopping the creation of the simulation environment for the environment server.

[0215] In a possible implementation, sending the experience data of each intelligent agent included in any simulation environment to the central training server includes: sending the experience data of each intelligent agent included in any simulation environment to the central training server through an information transmission channel between each simulation environment in the environment server and the central training server.

[0216] In one possible implementation, the prediction operation strategy corresponding to each simulation environment carries the environment identifier of the simulation environment; the receiving of the prediction operation strategy information sent by the central training server, and causing the corresponding simulation environment in the environment server to execute the prediction operation strategy corresponding to the prediction operation strategy information, includes: based on the environment identifier carried by each prediction operation strategy in the prediction operation strategy information, receiving the prediction operation strategy carrying the environment identifier in the prediction operation strategy information through the information transmission channel between the simulation environment corresponding to the environment identifier in the environment server and the central training server, so that the simulation environment executes the prediction operation strategy.

[0217] In one possible implementation, Figure 7 A flow chart of the environment operation provided by the embodiment of the present invention, such as Figure 7 As shown, after any simulation environment is started, the operation steps of the simulation environment include:

[0218] S701, initializing the simulation environment.

[0219] S702, after issuing the preset operation strategy, obtaining the experience data of the intelligent agent included in the simulation environment.

[0220] S703, based on the experience data of the intelligent agent included in the simulation environment, determine whether the simulation environment controls multiple intelligent agents.

[0221] S704, if the simulation environment controls multiple agents, convert the experience data of the agents included in the simulation environment into experience data in the form of multi-agents, and send the experience data in the form of multi-agents to the central training server through the information transmission channel between the simulation environment and the central training server; if not, send the experience data of the agents included in the simulation environment to the central training server through the information transmission channel between the simulation environment and the central training server.

[0222] S705: If the simulation environment is finished running, close the information transmission channel between the simulation environment and the central training server.

[0223] S706, if the simulation environment has not yet finished running, receive the predicted operation strategy information sent by the central training server, so that the simulation environment executes the predicted operation strategy corresponding to the predicted operation strategy information, and returns to execute the step of obtaining the experience data of the intelligent agent included in the simulation environment.

[0224] In one possible implementation, if the simulation environment controls multiple agents, converting the experience data of the agents included in the simulation environment into experience data in the form of a multi-agent comprises: if the simulation environment controls multiple agents, determining whether the multiple agents controlled by the simulation environment are asynchronous agents; if the multiple agents controlled by the simulation environment are asynchronous agents, converting the experience data of the agents included in the simulation environment into experience data in the form of a multi-agent.

[0225] In one possible implementation, Figure 8 A flow chart of environment destruction provided by an embodiment of the present invention, such as Figure 8 As shown, after any simulation environment is started, the steps for destroying the simulation environment include:

[0226] S801, determining whether the simulation environment has been actively run to completion.

[0227] S802: If the active operation of the simulation environment ends, the simulation environment is deleted from the environment server where the simulation environment is located.

[0228] S803: If the simulation environment has not been actively run to completion, obtain the running information of the simulation environment.

[0229] S804: Based on the operation information, determine whether the information transmission channel between the simulation environment and the central training server is closed.

[0230] S805: If the information transmission channel between the simulation environment and the central training server is closed, the simulation environment is deleted from the environment server where the simulation environment is located.

[0231] S806: If the information transmission channel between the simulation environment and the central training server is not closed, send an environment destruction request to the central training server through the information transmission channel between the simulation environment and the central training server.

[0232] S807, receiving an environment destruction instruction sent by the central training server, and deleting the simulation environment from the environment server where the simulation environment is located according to the environment destruction instruction.

[0233] The environment destruction instruction is determined by the central training server according to the environment destruction request.

[0234] The embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of the associated agents, the experience data can be shared between the agents, thereby improving the accuracy of the reinforcement learning model obtained by training. In addition, by adopting the method provided by the embodiment of the present invention, by decoupling the experience pool, the reinforcement learning model and the simulation environment, any existing offline strategy reinforcement learning algorithm can be replaced at any time, and the experience pool can collect sample data of each step in real time, so that the use of sample data is more timely. In addition, the training environment of the reinforcement learning model can support multi-process and distributed coexistence, which can fully utilize the performance of a single server and perform horizontal performance expansion, thereby improving the utilization rate of server resources. In addition, the training of the reinforcement learning model is not performed in parallel, and the number of samples can be greatly increased while ensuring the quality of the sample data, thereby improving the convergence speed of the algorithm. The central training server can use the CPU or the GPU for reinforcement learning model training, which can give full play to the hardware performance and improve the training speed.

[0235] Fig. 9 A structural diagram of a distributed multi-agent reinforcement learning model training provided by an embodiment of the present invention, such as Fig. 9 As shown, the environment server can include a multi-agent adaptation module. When the simulation environment controls multiple agents, the environment server can load the multi-agent adaptation model to create a specified multi-agent in the simulation environment and complete the adaptation of dynamic functions. Multiple simulation environments can be run in each environment server. For example, environment server 1 runs environment 1, environment 2 and environment 3, environment server 2 runs environment 4, environment 5 and environment 6, and environment server N runs environment 7, environment 8 and environment 9.

[0236] The central training server includes a distributed simulation environment management module and a multi-agent reinforcement learning module, wherein the distributed simulation environment management module may include an experience collection submodule, a multi-server environment management submodule and a strategy distribution submodule; the multi-agent reinforcement learning module may include a multi-agent experience state mixing submodule, multiple preset experience pools and an environment control action queue submodule.

[0237] The environment server can be used to send the experience data of each agent included in any simulation environment to the central training server, so that the central training server can perform the steps of reinforcement learning model training. The environment server can also be used to run the simulation environment and destroy the simulation environment.

[0238] The multi-server environment management submodule can be used to obtain configuration information of each of the environment servers; select an environment server to be configured based on the configuration information; create an SSH connection between the central training server and the environment server to be configured based on the configuration information of the environment server to be configured; send a simulation environment startup instruction to the environment server to be configured through the SSH connection, so that the environment server to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the central training server returns the transmission port information corresponding to the simulation environment; based on the transmission port information, create an information transmission channel between the central training server and the simulation environment, and update the number of simulation environments running in the environment server to be configured; if the number of simulation environments running in the environment server to be configured does not reach the limited number of environments corresponding to the environment server to be configured, return to execute the step of sending the simulation environment startup instruction to the environment server to be configured through the SSH connection; otherwise, stop creating the simulation environment for the environment server to be configured, and return to execute the step of selecting the environment server to be configured based on the configuration information for the remaining environment servers until the number of simulation environments running in each environment server reaches the limited number of environments corresponding to the environment server.

[0239] The experience collection submodule can be used to receive the experience data of each intelligent agent included in any simulation environment sent by the environment server through the information transmission channel between each simulation environment in the environment server and the central training server.

[0240] The strategy distribution submodule can be used to determine the simulation environment corresponding to the prediction operation strategy based on the environment identifier carried by each prediction operation strategy in the prediction operation strategy information; distribute the prediction operation strategy to the simulation environment corresponding to the environment identifier in the environment server through the information transmission channel between the simulation environment and the central training server, so that the simulation environment in the environment server executes the prediction operation strategy.

[0241] The multi-agent reinforcement learning module can be used to mix the experience data of related agents when the data volume of the experience data is not less than a first preset data volume, and store the mixed experience data in a preset experience pool; when the data volume in the preset experience pool reaches a second preset data volume, obtain the mixed experience data from the preset experience pool as sample data; specifically, an experience adoption algorithm or a random adoption algorithm can be used to obtain the mixed experience data from the preset experience pool as sample data, wherein each agent corresponds to a preset experience pool, for example, there is an "agent A experience pool" for agent A, there is an "agent B experience pool" for agent B, and there is an "agent C experience pool" for agent C. For each agent, the experience data of other agents associated with the agent and the mixed data of the experience data of the agent are stored in the preset experience pool corresponding to the agent; then the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain the output prediction operation strategy information; specifically, the reinforcement learning model to be trained can be Fig. 9 The multi-agent reinforcement learning algorithm model in the , and then continuously train and update the multi-agent reinforcement learning algorithm model to obtain the predicted operation strategy, and store the predicted operation strategy in Fig. 9 The environment control queue shown in the figure waits to be sent to the environment server; then, the predicted operation strategy information is sent to the environment server so that the corresponding simulation environment in the environment server executes the corresponding predicted operation strategy; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server.

[0242] The following is a specific example provided by an embodiment of the present invention:

[0243] 1. Application scenario: In the process of optimizing road traffic lights, the optimization of traffic light signals at multiple intersections is completed, so that it can automatically generate the optimal signal light timing under various conditions, so that the traffic efficiency through multiple intersections is the highest. That is, the traffic simulation environment, through traffic simulation software such as sumo, the roads at three intersections are simulated, and the number of vehicle queues, the travel time at the intersection, and the travel time on a section of the road can be obtained through the interface. The timing of the signal light can be modified through the interface. For example, if the signal light is changed to 30 seconds for north-south and 40 seconds for east-west, the signal light will be in the next 30+40=70 seconds. Set the north-south straight and left turn signal lights to 30 seconds green light in turn, and the east-west direction is red; then set the north-south straight and left turn to 40 seconds red light, and the east-west straight and left turn to 30 seconds green light.

[0244] 2. Operation strategy: Control the timing of each street light, for example: north-south traffic is released for 30 seconds, and the east-west light is red; east-west traffic is released for 40 seconds, and the north-south light is red.

[0245] 3. Rewards: Rewards can include local rewards and global rewards. Among them, the local reward is: the travel time of a single intersection (the lower the time, the better), and the global reward is: the travel time of a road (the lower the time, the better). Local rewards and global rewards can be combined: for example, directly adding them; or weighted addition, giving higher weights to global rewards, and by continuously optimizing rewards, optimizing the travel time to the minimum, and keeping it unchanged for a period of time, the entire training process converges and ends. Here are two examples of reward targets. In actual design, there will be more complex reward and penalty parameters.

[0246] 4. Intelligent agent: Each intersection can be designed as an intelligent agent, which receives the status given by the simulation environment and then issues a signal timing strategy.

[0247] 5. Multi-agent: The three intersections are respectively called Agent 1 (Intersection 1), Agent 2 (Intersection 2) and Agent 3 (Intersection 3). Agent 1, Agent 2 and Agent 3 need to continuously interact with the simulation environment to complete the reinforcement learning model training.

[0248] 6. Multi-agent reinforcement learning model: The MA-SAC (Multi-Agent SAC) algorithm can be used. The MA-SAC algorithm is a multi-agent learning algorithm designed based on the single-agent reinforcement learning algorithm SAC.

[0249] 7. Distributed parallel environment: A peer-to-peer simulation environment can be used, that is, the initial information of all simulation environments is the same, so the configuration of the process is the same when it is started, and the road network and traffic use the same initial data.

[0250] 8. Running servers: 1 GPU server, 3 CPU servers, environment server A, environment server B, environment server C.

[0251] 9. Training process: You can first start and initialize a MA-SAC multi-agent reinforcement learning algorithm on the GPU server, and then start the distributed simulation environment management module, create 4, 3, and 3 simulation environment processes on environment server A, environment server B, and environment server C respectively, a total of 10 simulation environment processes in parallel, and each simulation environment runs 100 rounds.

[0252] At intersection 1 of simulation environment process 1 of environment server A (i.e., agent 1 in simulation environment 1, hereinafter referred to as A-1-1), the state information State can be returned first, and a reward Reward and its corresponding operation strategy Action can be calculated to obtain a trajectory [S, R, A]. However, there is no other information at this time, so this information [S, R, A] can be stored first. Then, intersection 2 of simulation environment process 1 (i.e., agent 2 in simulation environment 1, hereinafter referred to as A-1-2) and intersection 1 of simulation environment process 2 (i.e., agent 1 in simulation environment 2, hereinafter referred to as A-2-1) are returned. Until the same simulation environment process and the same intersection have sent back 5 data, the calculation of a new trajectory matrix can be started. For the state of A-1-1, its previous state is represented as [S ′ ,R ′ ,A ′ ] and the new state [S,R,A], becoming [S,R,A,S ′ ,R ′ ,A ′]. Then merge the intersection status of A-1-2 associated with A-1-1 into A-1-1 and express it in the same format. This information can be put into the preset experience pool corresponding to A-1-1, that is, an array, and the number of arrays is currently 1. When the number of arrays is greater than 64, 64 samples can be randomly sampled from it (all selected for the first selection, and sampling can be performed for subsequent selections) and merged into new matrix data. Input to the MA-SAC algorithm for reinforcement learning model training. Among them, S can be interpreted as the traffic flow in the east, south, west, and north [10,20,10,15] + the number of vehicles queuing in the east, south, west, and north [0,0,1,0]. The reward R can be set as: the traffic flow is average, but the queue is low, and the reward is 10. Operation strategy A: The forecast operation strategy issued is that the east-west and north-south travel time is [30,30], that is, the east-west direction of intersection 1 in simulation environment 1 is green for 30 seconds, and the north-south direction is red for 30 seconds; and the east-west direction of intersection 1 in simulation environment 1 is red for 30 seconds, and the north-south direction is green for 30 seconds. The traffic state after intersection 1 and intersection 2 in simulation environment 1 are merged: [10,20,10,15,30,23,13,12], the first 4 are the traffic of intersection 1, and the last 4 are the traffic of intersection 2. The number of queued vehicles is also merged in a similar way.

[0253] After training, the reinforcement learning model training can predict new prediction operation strategies. For example, after training, the algorithm believes that [30,40] can obtain higher rewards. In this state, a new state is generated: [30,40], and then sent to simulation environment 1.

[0254] This process repeats itself until the reinforcement learning model to be trained has an evaluation of all states and can output the action it believes to have the highest value, so that the reward of the entire simulation environment always remains high.

[0255] After each simulation environment is trained for 100 rounds, the training of the reinforcement learning model is completed. During the training, the reinforcement learning model obtained after each training or every 10 trainings can be automatically saved, and some models with better results can be selected.

[0256] Corresponding to the above-mentioned reinforcement learning model training method for an intelligent agent, an embodiment of the present invention further provides a reinforcement learning model training device for an intelligent agent. The reinforcement learning model training device for an intelligent agent provided by an embodiment of the present invention is introduced below. Fig.10A structural schematic diagram of a reinforcement learning model training device for an agent of a central training server in a reinforcement learning model training system provided by an embodiment of the present invention, wherein the system further comprises at least one environment server, each of which runs at least one simulation environment, each of which comprises at least one agent, and the total number of agents is greater than 1. Fig.10 As shown, the device comprises:

[0257] The experience data receiving module 1001 is used to receive the experience data of each agent included in any simulation environment sent by the environment server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located;

[0258] An experience data mixing module 1002, configured to mix the experience data of the associated agents and store the mixed experience data in a preset experience pool if the data amount of the experience data is not less than a first preset data amount;

[0259] The model training module 1003 is used to obtain mixed experience data from the preset experience pool as sample data when the amount of data in the preset experience pool reaches a second preset data amount, and trigger the training of the reinforcement learning model to be trained based on the sample data to obtain output prediction operation strategy information; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0260] The strategy information sending module 1004 is used to send the predicted operation strategy information to the environment server, so that: the corresponding simulation environment in the environment server executes the corresponding predicted operation strategy, and sends the status information of each simulation environment to the central training server after executing the predicted operation strategy;

[0261] The end condition determination module 1005 is used to receive the status information of each simulation environment sent by the environment server, and determine whether the preset model training end condition is met based on the status information of each simulation environment; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server.

[0262] The device provided by the embodiment of the present invention is adopted, by receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server; when the data amount of the experience data is not less than the first preset data amount, the experience data of the associated intelligent agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data amount in the preset experience pool reaches a second preset data amount, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server, so that: the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy, and after executing the prediction operation strategy, the status information of each simulation environment is sent to the central training server; the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; if the preset model training end condition is not met, the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned. That is, the embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of related agents, the experience data can be shared between agents, thereby improving the accuracy of the trained reinforcement learning model.

[0263] Corresponding to the above-mentioned reinforcement learning model training method for an intelligent agent, an embodiment of the present invention further provides a reinforcement learning model training device for an intelligent agent. The reinforcement learning model training device for an intelligent agent provided by an embodiment of the present invention is introduced below. Fig.11 A structural schematic diagram of a reinforcement learning model training device for an agent of any environment server in a reinforcement learning model training system provided by an embodiment of the present invention, wherein the system comprises a central training server and at least one environment server, each of which runs at least one simulation environment, each of which comprises at least one agent, and the total number of agents is greater than 1. Fig.11 As shown, the device comprises:

[0264] The experience data sending module 1101 is used to send the experience data of each agent included in any simulation environment to the central training server, so that the central training server performs the following steps: when the amount of data in the preset experience pool reaches a second preset data amount, obtain the mixed experience data from the preset experience pool as sample data, and trigger the training of the reinforcement learning model to be trained based on the sample data to obtain the output prediction operation strategy information; send the prediction operation strategy information to the environment server; receive the status information of each simulation environment sent by the environment server, and determine whether the preset model training end condition is met based on the status information of each simulation environment; if the preset model training end condition is met, determine the current reinforcement learning model to be trained as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, return to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server;

[0265] The strategy information receiving module 1102 is used to receive the predicted operation strategy information sent by the central training server, and enable the corresponding simulation environment in the environment server to execute the corresponding predicted operation strategy in the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy; wherein the experience data of each intelligent agent includes: the status information of the intelligent agent, the reward information determined by the environment server based on the status information of the intelligent agent, and the operation strategy of the simulation environment where the intelligent agent is located.

[0266] The embodiment of the present invention proposes a new and efficient reinforcement learning model training framework that supports multiple agents and multiple simulation environments, and by mixing the experience data of related agents, the experience data can be shared between agents, thereby improving the accuracy of the trained reinforcement learning model.

[0267] The embodiment of the present invention also provides a server, such as Fig.12 As shown, it includes a processor 1201, a communication interface 1202, a memory 1203 and a communication bus 1204, wherein the processor 1201, the communication interface 1202, and the memory 1203 communicate with each other through the communication bus 1204.

[0268] Memory 1203, used for storing computer programs;

[0269] The processor 1201 is used to implement the steps of the reinforcement learning model training method of any of the intelligent agents when executing the program stored in the memory 1203.

[0270] The communication bus mentioned in the above server can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0271] The communication interface is used for communication between the above server and other devices.

[0272] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0273] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0274] In another embodiment provided by the present invention, a computer-readable storage medium is also provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the reinforcement learning model training method of any of the above-mentioned intelligent agents are implemented.

[0275] In another embodiment provided by the present invention, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute the reinforcement learning model training method of any intelligent agent in the above-mentioned embodiments.

[0276] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center to another website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk Solid State Disk (SSD)), etc.

[0277] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0278] Each embodiment in this specification is described in a related manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system, server, computer-readable storage medium, and computer program product embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0279] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.

Claims

1. A reinforcement learning model training method for an intelligent agent, It is characterized in that A central training server applied to a reinforcement learning model training system, the system further comprising at least one environment server, each of the environment servers running at least one simulation environment, each simulation environment comprising at least one intelligent agent, the total number of intelligent agents being greater than 1, the method comprising: Receive the experience data of each agent included in any simulation environment sent by the environment server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located; When the amount of the experience data is not less than a first preset amount of data, the experience data of the associated agents are mixed, and the mixed experience data is stored in a preset experience pool; When the amount of data in the preset experience pool reaches a second preset amount of data, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server; Sending the prediction operation strategy information to the environment server so that: the corresponding simulation environment in the environment server executes the corresponding prediction operation strategy, and after executing the prediction operation strategy, sends the status information of each simulation environment to the central training server; Receive status information of each simulation environment sent by the environment server, and determine whether a preset model training end condition is met based on the status information of each simulation environment; If the preset model training end condition is reached, the current reinforcement learning model to be trained is determined as the target reinforcement learning model training obtained by training; If the preset model training end condition is not met, return to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server; The step of mixing the experience data of the associated agents when the amount of the experience data is not less than a first preset amount of data and storing the mixed experience data in a preset experience pool includes: Acquire the association relationship between each intelligent agent from the environment server; When the amount of the experience data is not less than the first preset data amount, for each agent, the experience data of the agent associated with the agent and the experience data of the agent are mixed according to the association relationship to obtain mixed experience data, and stored in the preset experience pool corresponding to the agent.

2. The method according to claim 1, It is characterized in that The determining whether a preset model training end condition is reached based on the status information of each simulation environment includes: Based on the status information of each simulation environment, determining whether each simulation environment in the environment server has completed a preset number of operations; If each simulation environment in the environment server has completed the preset number of runs, it is determined that the preset model training end condition has been reached.

3. The method according to claim 1, It is characterized in that Before receiving the experience data of each agent included in any simulation environment sent by the environment server, the method further includes: Obtaining configuration information of each of the environment servers; Selecting a server for the environment to be configured based on the configuration information; Based on the configuration information of the environment server to be configured, an SSH connection is created between the central training server and the environment server to be configured; Sending a simulation environment startup instruction to the environment server to be configured through an SSH connection, so that the environment server to be configured executes a simulation environment according to the environment startup instruction, and after the simulation environment is started, the central training server returns the transmission port information corresponding to the simulation environment; Based on the transmission port information, an information transmission channel is created between the central training server and the simulation environment, and the number of simulation environments running in the environment server to be configured is updated; If the number of simulation environments running in the environment server to be configured does not reach the limited number of environments corresponding to the environment server to be configured, return to execute the step of sending a simulation environment startup instruction to the environment server to be configured through the SSH connection; otherwise, stop creating a simulation environment for the environment server to be configured, and return to execute the step of selecting the environment server to be configured based on the configuration information for the remaining environment servers until the number of simulation environments running in each environment server reaches the limited number of environments corresponding to the environment server.

4. The method according to claim 3, It is characterized in that The receiving of the experience data of each agent included in any simulation environment sent by the environment server comprises: The experience data of each intelligent agent included in any simulation environment sent by the environment server is received through the information transmission channel between each simulation environment in the environment server and the central training server.

5. The method according to claim 3, It is characterized in that The prediction operation strategy corresponding to each simulation environment carries the environment identifier of the simulation environment; The sending the prediction operation strategy information to the environment server includes: Determine the simulation environment corresponding to each prediction operation strategy based on the environment identifier carried by each prediction operation strategy in the prediction operation strategy information; The prediction operation strategy is distributed to the simulation environment corresponding to the environment identifier in the environment server through the information transmission channel between the simulation environment and the central training server, so that the simulation environment in the environment server executes the prediction operation strategy.

6. The method according to claim 1, It is characterized in that Before sending the prediction operation strategy information to the environment server, the method further includes: When the amount of data in the preset experience pool does not reach the second preset data amount, preset operation strategy information is sent to the environment server, so that the corresponding prediction operation strategy in the preset operation strategy information is executed in the corresponding simulation environment in the environment server, and after executing the prediction operation strategy, the status information of each simulation environment is sent to the central training server, and the step of receiving the experience data of each intelligent agent included in any simulation environment sent by the environment server is returned.

7. A reinforcement learning model training method for an intelligent agent, It is characterized in that Applied to any environment server in a reinforcement learning model training system, the system includes a central training server and at least one environment server, each of the environment servers runs at least one simulation environment, each simulation environment includes at least one intelligent agent, and the total number of intelligent agents is greater than 1, the method includes: Sending the experience data of each agent included in any simulation environment to the central training server so that the central training server performs the following steps: When the amount of data in the preset experience pool reaches a second preset amount of data, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server; the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met; if the preset model training end condition is met, the current reinforcement learning model to be trained is determined as the target reinforcement learning model obtained by training; if the preset model training end condition is not met, The model training end condition returns to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server; wherein the prediction operation strategy information includes the prediction operation strategy of the corresponding simulation environment in the environment server; the mixed experience data is the association relationship between the agents obtained by the environment server, and when the data amount of the experience data is not less than the first preset data amount, for each agent, according to the association relationship, the experience data of the agent associated with the agent and the experience data of the agent are mixed to obtain the mixed experience data, and the mixed experience data is stored in the preset experience pool corresponding to the agent; Receive the predicted operation strategy information sent by the central training server, and make the corresponding simulation environment in the environment server execute the predicted operation strategy corresponding to the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy; wherein the experience data of each intelligent agent includes: the status information of the intelligent agent, the reward information determined by the environment server based on the status information of the intelligent agent, and the operation strategy of the simulation environment where the intelligent agent is located.

8. The method according to claim 7, It is characterized in that Before sending the experience data of each agent included in any simulation environment to the central training server, the method further includes: Receiving a simulation environment startup instruction sent by the central training server through an SSH connection; wherein the SSH connection is a connection between the central training server and the environment server created based on the configuration information of the environment server; Start a simulation environment according to the simulation environment start instruction; After the simulation environment is started, the transmission port information corresponding to the simulation environment is returned to the central training server, so that the central training server performs the following steps: based on the transmission port information, an information transmission channel is created between the central training server and the simulation environment, and the number of simulation environments running in the environment server is updated; if the number of simulation environments running in the environment server does not reach the limited number of environments corresponding to the environment server, a step of sending a simulation environment startup instruction to the environment server through an SSH connection is performed; otherwise, the creation of the simulation environment for the environment server is stopped.

9. The method according to claim 8, It is characterized in that The sending of the experience data of each agent included in any simulation environment to the central training server includes: Through the information transmission channel between each simulation environment in the environment server and the central training server, the experience data of each intelligent agent included in any simulation environment is sent to the central training server.

10. The method according to claim 8, It is characterized in that The prediction operation strategy corresponding to each simulation environment carries the environment identifier of the simulation environment; The receiving the prediction operation strategy information sent by the central training server, and causing the corresponding simulation environment in the environment server to execute the prediction operation strategy corresponding to the prediction operation strategy information, includes: Based on the environment identifier carried by each predicted operation strategy in the predicted operation strategy information, the predicted operation strategy carrying the environment identifier in the predicted operation strategy information is received through the information transmission channel between the simulation environment corresponding to the environment identifier in the environment server and the central training server, so that the simulation environment executes the predicted operation strategy.

11. The method according to claim 8, It is characterized in that After any simulation environment is started, it also includes: Initialize the simulation environment; After issuing the preset operation strategy, obtaining the experience data of the intelligent agent included in the simulation environment; Determining whether the simulation environment controls a plurality of agents based on experience data of agents included in the simulation environment; If the simulation environment controls multiple agents, convert the experience data of the agents included in the simulation environment into experience data in the form of multiple agents, and send the experience data in the form of multiple agents to the central training server through the information transmission channel between the simulation environment and the central training server; if not, send the experience data of the agents included in the simulation environment to the central training server through the information transmission channel between the simulation environment and the central training server; If the simulation environment ends, the information transmission channel between the simulation environment and the central training server is closed; If the simulation environment has not yet finished running, receive the predicted operation strategy information sent by the central training server so that the simulation environment executes the predicted operation strategy corresponding to the predicted operation strategy information, and return to execute the step of obtaining the experience data of the intelligent agent included in the simulation environment.

12. The method according to claim 11, It is characterized in that If the simulation environment controls multiple agents, converting the experience data of the agents included in the simulation environment into experience data in a multi-agent form includes: If the simulation environment controls multiple agents, determining whether the multiple agents controlled by the simulation environment are asynchronous agents; If the multiple agents controlled by the simulation environment are asynchronous agents, the experience data of the agents included in the simulation environment are converted into experience data in a multi-agent form.

13. The method according to claim 8, It is characterized in that After any simulation environment is started, it also includes: Determine whether the simulation environment has ended actively; If the simulation environment ends its active operation, the simulation environment is deleted from the environment server where the simulation environment is located; If the simulation environment does not automatically end, obtain the operation information of the simulation environment; Based on the operation information, determining whether the information transmission channel between the simulation environment and the central training server is closed; If the information transmission channel between the simulation environment and the central training server is closed, deleting the simulation environment from the environment server where the simulation environment is located; If the information transmission channel between the simulation environment and the central training server is not closed, sending an environment destruction request to the central training server through the information transmission channel between the simulation environment and the central training server; Receive an environment destruction instruction sent by the central training server, and delete the simulation environment from the environment server where the simulation environment is located according to the environment destruction instruction; wherein the environment destruction instruction is determined by the central training server according to the environment destruction request.

14. A reinforcement learning model training system, It is characterized in that include: A central training server and at least one environment server, each of the environment servers runs at least one simulation environment, each simulation environment includes at least one intelligent agent, and the total number of intelligent agents is greater than 1; The environment server is used to send the experience data of each agent included in any simulation environment to the central training server; wherein the experience data of each agent includes: the state information of the agent, the reward information determined by the environment server based on the state information of the agent, and the operation strategy of the simulation environment where the agent is located; The central training server is used to receive the experience data of each intelligent agent included in any simulation environment sent by the environment server; when the data volume of the experience data is not less than a first preset data volume, the experience data of the associated intelligent agents are mixed, and the mixed experience data is stored in a preset experience pool; when the data volume in the preset experience pool reaches a second preset data volume, the mixed experience data is obtained from the preset experience pool as sample data, and the training of the reinforcement learning model to be trained is triggered based on the sample data to obtain output prediction operation strategy information; the prediction operation strategy information is sent to the environment server, the status information of each simulation environment sent by the environment server is received, and based on the status information of each simulation environment, it is determined whether the preset model training end condition is met, and if the preset model training end condition is met, the current reinforcement learning model to be trained is terminated. Determine the training of the target reinforcement learning model obtained by training; if the preset model training end condition is not reached, return to the step of receiving the experience data of each agent included in any simulation environment sent by the environment server; wherein the predicted operation strategy information includes the predicted operation strategy of the corresponding simulation environment in the environment server; when the data amount of the experience data is not less than the first preset data amount, the experience data of the associated agents are mixed, and the mixed experience data is stored in the preset experience pool, including: obtaining the association relationship between each agent from the environment server; when the data amount of the experience data is not less than the first preset data amount, for each agent, according to the association relationship, the experience data of the agent associated with the agent and the experience data of the agent are mixed to obtain mixed experience data, and stored in the preset experience pool corresponding to the agent; The environment server is also used to receive the predicted operation strategy information sent by the central training server, and enable the corresponding simulation environment in the environment server to execute the corresponding predicted operation strategy in the predicted operation strategy information, and send the status information of each simulation environment to the central training server after executing the predicted operation strategy.

Citation Information

Patent Citations

  • Reinforced learning implementation method and device based on public information and storage medium

    CN110796266A

  • Power distribution network auxiliary decision-making method and system fusing deep reinforcement learning and expert experience

    CN113159341A