Intelligent agent training method, device, computer equipment and storage medium

By combining evolutionary algorithms and deep reinforcement learning, the initial reinforcement learning agent is trained using the empirical actions of sample agents in the evolutionary population, the problems of sparse rewards and premature convergence in deep reinforcement learning are solved, and the learning efficiency and effect are improved.

CN113919482BActive Publication Date: 2025-05-06SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111106047.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-05-06
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

In deep reinforcement learning, there are problems of sparse rewards, premature convergence to local optimality, and hyperparameter sensitivity and fragile convergence characteristics.

Method used

Combining evolutionary algorithms and deep reinforcement learning, the initial reinforcement learning agent is trained to improve learning efficiency and effect through the empirical actions of sample agents interacting with the environment in the evolutionary population.

Benefits of technology

The learning efficiency and effectiveness of deep reinforcement learning are improved, premature convergence and hyperparameter sensitivity problems are avoided, and the optimization ability for high-dimensional problems is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113919482B_ABST
    Figure CN113919482B_ABST
Patent Text Reader

Abstract

The present application relates to an agent training method, device, computer equipment and storage medium. The method includes: obtaining multiple empirical action data, the empirical action data is the empirical action of multiple target sample agents in the evolutionary population and the environment interactive learning; based on the multiple empirical action data, obtaining the reward information of the action data output by the initial reinforcement learning agent; according to the reward information and the preset loss function, the network parameters of the initial reinforcement learning agent are updated; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, then the update of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent. The present application combines evolutionary algorithms with deep reinforcement learning, which can improve the learning efficiency and effect of deep reinforcement learning, thereby better controlling the reinforcement agent to complete continuous control tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent agent collaborative control technology, and in particular to an intelligent agent training method, apparatus, computer equipment and storage medium. Background Art

[0002] Deep reinforcement learning is a new algorithm that combines deep learning and reinforcement learning to achieve direct mapping from perception to action. It inputs perceptual information (such as vision) and then directly outputs actions through a deep neural network without any hard-coding in between. Deep reinforcement learning combines the advantages of deep neural networks and reinforcement learning, and can effectively solve the perception and decision-making problems of agents in high-dimensional complex problems. It is a cutting-edge research direction in the field of general artificial intelligence and has broad application prospects.

[0003] The key to deep reinforcement learning is to obtain training samples through an agent that continuously interacts with the environment, thereby training a deep policy network. In the deep policy network, the agent receives data representing the current state of the environment, and responds to the received state of the environment by executing actions from the continuous action space in an attempt to perform the corresponding task in the environment. Summary of the invention

[0004] Based on this, it is necessary to provide an intelligent agent training method, device, computer equipment and storage medium that can solve the reward sparsity problem in response to the above technical problems.

[0005] In a first aspect, a method for training an intelligent agent is provided, the method comprising:

[0006] Acquire multiple experience action data, where the experience action data is the experience actions of multiple target sample agents in the evolutionary population interacting and learning with the environment;

[0007] Based on multiple empirical action data, obtain reward information of the action data output by the initial reinforcement learning agent;

[0008] Update the network parameters of the initial reinforcement learning agent based on the reward information and the preset loss function;

[0009] If the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0010] In one embodiment, before acquiring a plurality of experience motion data, the method includes:

[0011] The first sample intelligent agent in the evolutionary population is reproduced by an evolutionary strategy to obtain a second sample intelligent agent; the first sample intelligent agent is a sample intelligent agent in the evolutionary population that meets a preset fitness condition;

[0012] The second sample intelligent agent is reproduced by an evolutionary algorithm to obtain a third sample intelligent agent;

[0013] Determine a plurality of target sample agents according to the fitness of the first sample agent, the second sample agent, and the third sample agent;

[0014] The experience action data of each target sample agent's interactive learning with the environment is stored in a circular replay buffer.

[0015] In one embodiment, performing reproduction processing on a first sample intelligent agent in an evolutionary population by using an evolution strategy to obtain a second sample intelligent agent includes:

[0016] Recombining the first sample agent to obtain a first offspring;

[0017] Perform mutation treatment on the first offspring to obtain the second offspring;

[0018] According to the first fitness function preset by the evolution strategy, the second offspring that meets the fitness condition in the second offspring is used as the second sample intelligent agent.

[0019] In one embodiment, performing breeding processing on the second sample intelligent agent through an evolutionary algorithm to obtain a third sample intelligent agent includes:

[0020] Perform cross processing on each sample agent group in the second sample agent to obtain a third offspring; each sample agent group includes any two sample agents in the second sample agent;

[0021] The third offspring is subjected to mutation processing to obtain the fourth offspring;

[0022] According to the second fitness function preset by the evolutionary algorithm, the fourth offspring that meets the fitness condition in the fourth offspring is used as the third sample intelligent agent.

[0023] In one embodiment, the experience action data of each target sample agent interactively learning with the environment is stored in a circular replay buffer, including:

[0024] The experience action data of each target sample agent's interactive learning with the environment are stored in a circular replay buffer in the form of experience tuples; each experience tuple includes: the current environment state, the sample action in the continuous action space executed by the target sample agent in response to the current environment state, the reward obtained by the target sample agent after executing the sample action, and the next environment state of the target sample agent's interactive environment.

[0025] In one embodiment, obtaining a plurality of experience action data includes:

[0026] Get a preset number of experience tuples from the circular replay buffer;

[0027] The experience action data in a preset number of experience tuples are determined as a plurality of experience action data.

[0028] In one embodiment, obtaining reward information of action data output by an initial reinforcement learning agent based on a plurality of experience action data includes:

[0029] For each experience tuple in the preset number of experience tuples, obtaining a first reward of the initial reinforcement learning agent according to the action data output by the initial reinforcement learning agent in response to the current environment state in each experience tuple;

[0030] Obtaining a second reward for the initial reinforcement learning agent according to the action data output by the initial deep learning agent in response to the next environment state in each experience tuple;

[0031] Reward information is determined based on the first reward and the second reward.

[0032] In one embodiment, the loss function includes a mean square error loss function and a gradient strategy loss function;

[0033] According to the reward information and the preset loss function, the network parameters of the initial reinforcement learning agent are updated, including:

[0034] According to the reward information and the mean square error loss function, the reward parameters are updated through the back propagation of the gradient neural network;

[0035] According to the updated reward parameters and gradient policy loss function, the network parameters of the initial reinforcement learning agent are updated through the back propagation of the gradient neural network.

[0036] In one embodiment, the method further comprises:

[0037] According to the preset synchronization cycle, the network parameters of the reinforcement learning agent are copied to all sample agents in the evolving population, and the network parameters of the reinforcement learning agent are used to update the network parameters of all sample agents in the evolving population.

[0038] In a second aspect, an intelligent agent training device is provided, the device comprising:

[0039] An experience acquisition module is used to acquire multiple experience action data, where the experience action data is the experience actions of multiple target sample agents in the evolutionary population interacting with the environment;

[0040] An action training module is used to obtain reward information of action data output by the initial reinforcement learning agent based on multiple empirical action data;

[0041] The parameter updating module is used to update the network parameters of the initial reinforcement learning agent according to the reward information and the preset loss function; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0042] In a third aspect, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0043] Acquire multiple experience action data, where the experience action data is the experience actions of multiple target sample agents in the evolutionary population interacting and learning with the environment;

[0044] Based on multiple empirical action data, obtain reward information of the action data output by the initial reinforcement learning agent;

[0045] Update the network parameters of the initial reinforcement learning agent based on the reward information and the preset loss function;

[0046] If the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0047] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0048] Acquire multiple experience action data, where the experience action data is the experience actions of multiple target sample agents in the evolutionary population interacting and learning with the environment;

[0049] Based on multiple empirical action data, obtain reward information of the action data output by the initial reinforcement learning agent;

[0050] Update the network parameters of the initial reinforcement learning agent based on the reward information and the preset loss function;

[0051] If the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0052] The above-mentioned agent training method, device, computer equipment and storage medium obtain multiple experience action data, which are experience actions of multiple target sample agents in the evolutionary population and the environment interactive learning; based on multiple experience action data, the reward information of the action data output by the initial reinforcement learning agent is obtained; according to the reward information and the preset loss function, the network parameters of the initial reinforcement learning agent are updated; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the update of the network parameters of the initial reinforcement learning agent is terminated to obtain the trained reinforcement learning agent. That is, this application combines evolutionary algorithms with deep reinforcement learning, and the two algorithms complement each other. The evolution of the evolutionary algorithm is used as the experience pool of deep reinforcement learning, and the experience actions of the sample agents in the evolutionary population and the environment are interactively learned, and the frequency of interaction between the agent and the environment is increased to explore and learn more experience actions in the environment. Further, the initial reinforcement learning agent is trained by the experience actions of multiple target sample agents in the evolutionary population and the environment interactive learning, which can improve the learning efficiency and effect of deep reinforcement learning, thereby better controlling the deep reinforcement learning agent to complete continuous control tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is an application environment diagram of an intelligent agent training method in one embodiment;

[0054] Figure 2 is a flow chart of an agent training method in one embodiment;

[0055] Figure 3 A schematic diagram of a process for creating a circular replay buffer in one embodiment;

[0056] Figure 4 A schematic diagram of an optimization process of an evolving population in one embodiment;

[0057] Figure 5 A schematic diagram of an optimization process of an evolving population in another embodiment;

[0058] Figure 6 A schematic diagram of a process for obtaining experience action data in another embodiment;

[0059] Figure 7 is a flow chart of an agent training method in another embodiment;

[0060] Figure 8 is a flow chart of an agent training method in another embodiment;

[0061] Fig. 9 is a schematic diagram of an agent training system in one embodiment;

[0062] Fig.10is a flow chart of an agent training method in another embodiment;

[0063] Fig.11 is a structural block diagram of an intelligent agent training device in one embodiment;

[0064] Fig.12 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0066] Deep reinforcement learning (DRL) is a new algorithm that combines deep learning and reinforcement learning to achieve direct mapping from perception to action. It inputs perceptual information (such as vision) and then directly outputs actions through a deep neural network without any hard-coding in between. Deep reinforcement learning combines the advantages of deep neural networks and reinforcement learning, and can effectively solve the perception and decision-making problems of intelligent agents in high-dimensional complex problems. It is a cutting-edge research direction in the field of general artificial intelligence and has broad application prospects.

[0067] However, deep reinforcement learning also has certain defects, which mainly manifest in three aspects: sparse rewards for short-term action-related tasks, premature convergence to local optimality in high-dimensional action and state spaces, and hyperparameter sensitivity and fragile convergence characteristics.

[0068] Based on this, this application combines evolutionary strategies and evolutionary reinforcement learning, which includes evolutionary algorithms and deep reinforcement learning. That is, this application is a hybrid algorithm that combines evolutionary algorithms, evolutionary strategies, and deep reinforcement learning, combining the experience generated by the population-based method of the evolutionary algorithm to train the reinforcement learning agent, and regularly transfers the behavior strategy learned by the reinforcement learning agent into the evolutionary population to inject gradient information into the evolutionary algorithm. At the same time, the evolutionary strategy is used to perform black-box optimization on large-scale parameters of high-dimensional problems, avoiding the evolutionary algorithm from suffering from high sample complexity.

[0069] The intelligent agent training method provided in the present application can be applied to computer devices, including but not limited to various personal computers, laptops, smart phones, tablet computers, portable wearable devices, mobile terminals, etc.

[0070] Among them, the internal structure of the computer device is as follows Figure 1As shown, the processor in the internal structure is used to obtain multiple empirical action data, and according to the multiple empirical action data, train the initial reinforcement learning agent, and update the network parameters of the initial reinforcement learning agent until the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters. The memory in the internal structure includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database is used to store the empirical action data of the interactive learning of multiple target sample agents in the evolutionary population and the environment. The network interface is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an agent training method is implemented.

[0071] In one embodiment, Figure 2 As shown, a method for training an intelligent agent is provided, which is applied to Figure 1 The computer device in the example is used to illustrate, including the following steps:

[0072] Step 210: Acquire a plurality of experience action data, where the experience action data is the experience actions of a plurality of target sample agents in the evolutionary population interactively learning with the environment.

[0073] In order to interact with the environment, the agent receives data representing the current state of the environment, and executes actions from the continuous action space in response to the received data in an attempt to perform tasks in the environment. Therefore, the experience action data obtained in step 210 includes the environmental state of the target sample agent's interaction environment, and the action output by the target sample agent in response to the environmental state.

[0074] In the present application, the interaction environment of the sample agents in the evolving population and the initial reinforcement learning agent is the same. The sample agents in the evolving population interact with the environment to learn efficient actions, and then the initial reinforcement learning agent is trained based on the empirical actions learned by the sample agents.

[0075] In some embodiments, the interactive environment of the above-mentioned agents (including sample agents in the evolutionary population and initial reinforcement learning agents to be trained in deep reinforcement learning) is a simulation environment and the agents are implemented as one or more computer programs that interact with the simulation environment. As an example, the simulation environment can be a video game and the agent can be a simulated user playing the video game. As another example, the simulation environment can be a robot or motion simulation environment, such as a driving simulation or a flight simulation, and the agent is a simulated vehicle that navigates through motion simulation to perform a certain navigation task (e.g., navigating to a specific point in the environment without violating safety constraints). In these embodiments, the action output by the agent can be a point in the space of possible control inputs for controlling the simulated user or simulated vehicle.

[0076] In some embodiments, the interaction environment of the above-mentioned agents (including sample agents in the evolutionary population and initial reinforcement learning agents to be trained in deep reinforcement learning) is a real-world environment and the agents are mechanical agents that interact with the real-world environment. As an example, the agent can be a robot that interacts with the environment to complete a specific task (e.g., moving to a specific location or interacting with objects in the environment in a desirable way, such as performing reaching and / or grabbing and / or placing actions). As another example, the agent can be an autonomous or semi-autonomous vehicle that navigates through the environment. In these embodiments, the action output by the agent can be a point in the space of possible control inputs for controlling the robot or autonomous vehicle.

[0077] In some cases, a low-dimensional feature vector is used to characterize the state of the environment, and the values ​​of different dimensions of the low-dimensional feature vector can have a range of variation. For example, the state of the environment can include information identifying the current position (e.g., angle) and optionally the motion (e.g., angular velocity) of the joints of the agent. The state of the environment can also include information identifying the position of objects in the environment, the distance from the agent to these objects, or both.

[0078] In some other cases, the state of the environment is represented using high-dimensional pixel input from one or more images representing the state of the environment (e.g., images of a simulated environment and / or images captured by the agent's sensors as the agent interacts with the real-world environment).

[0079] It should be noted that, based on deep reinforcement learning, this application is provided with a loop replay buffer, and the experiential actions learned by the target sample agent in the evolving population through interaction with the environment are stored in the loop replay buffer. The loop replay buffer is the core mechanism for information to flow from the evolving population to the reinforcement learning agent. In contrast to the standard evolutionary algorithm, which extracts fitness indicators from these experiences and immediately ignores them, in the evolutionary reinforcement learning algorithm, the experiential actions learned by the sample agents in the population are retained in the loop replay buffer, and the reinforcement learning agent uses a powerful gradient-based method to repeatedly learn from the loop replay buffer. This mechanism can maximize the extraction of information from each individual experience, thereby improving sampling efficiency.

[0080] In a possible implementation, the implementation process of the above step 210 is as follows: the evolutionary population includes multiple sample agents, each sample agent uses its own network parameters to interact with the environment, and selects an output action based on the current state of the environment. Based on the learning effect of each sample agent, a target sample agent is selected from the multiple sample agents. Further, the experience actions learned by the interaction between the multiple target sample agents and the environment are stored in an empty loop replay buffer. When the experience actions stored in the replay buffer are greater than a preset experience threshold, multiple experience action data are obtained from the loop replay buffer, and the initial reinforcement learning agent is started for training.

[0081] In some embodiments, the sample agents in the evolving population adopt an Actor-Critic system, and each sample agent is randomly set with network parameters. For example, the parameters of the actor neural network are pre-set to θ π , the parameters of the evaluator neural network are pre-set to θ Q .

[0082] Among them, the Actor neural network takes the environment state s in which the agent interacts as input, fits the strategy π, and outputs the response action a=π(s), that is, directly selects the output action a of the agent based on the current state s. The Critic neural network takes the state s and action a as input, and outputs the cumulative reward Q=(s,a), which is used to evaluate the effect of taking action a under the environment state s. In other words, the Actor neural network in the sample agent is used to select actions, and the Critic neural network is used to estimate action rewards to evaluate the quality of the sample agent's output actions. The two work together to continuously improve the decision-making effect.

[0083] Step 220: Based on the multiple empirical action data, obtain reward information of the action data output by the initial reinforcement learning agent.

[0084] The initial reinforcement learning agent uses a dual Actor-Critic network system, and the initial reinforcement learning agent has pre-set network parameters. Similarly, the parameters of the Actor neural network are pre-set as θ π , the parameters of the Critic neural network are pre-set to θ Q At the same time, the parameters of the target Actor neural network are initialized to θ π′ , initialize the parameters of the target Critic neural network to θ Q′ .

[0085] It should be noted that the training of the initial reinforcement learning agent, that is, the action output by the initial reinforcement learning agent for different environmental states, updates the parameters θ of the actor's neural network according to the reward information of the action π , and the parameters of the evaluator neural network are θ Q .

[0086] In a possible implementation, the implementation process of step 220 may be: the Actor neural network is configured according to the parameter θ π , according to the current state of the environment, select the action output by the initial reinforcement learning agent, and the critic neural network follows the parameter θ Q , score the action output by the initial reinforcement learning agent to obtain the current reward information; the target Actor neural network is based on the parameter θ π′ , according to the next environment state, the next action output by the initial reinforcement learning agent is selected, and the target critic neural network is adjusted according to the parameter θ Q′ , score the action output by the initial reinforcement learning agent to obtain the next reward information; and then obtain the reward information of the action data output by the initial reinforcement learning agent based on the current reward information and the next reward information.

[0087] Among them, the initial reinforcement learning agent outputs actions in a continuous action space. Therefore, the current environment state input to the Actor neural network and the next environment state input to the target Actor neural network are continuous environment states.

[0088] Further, in some embodiments, the reward information of the action data output by the initial reinforcement learning agent is a cumulative reward value, which is not limited in this embodiment of the present application.

[0089] Step 230: Update the network parameters of the initial reinforcement learning agent according to the reward information and the preset loss function; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, then end the update of the network parameters of the initial reinforcement learning agent to obtain a trained reinforcement learning agent.

[0090] Among them, the network parameters of the initial reinforcement learning agent are updated, that is, the parameters θ of the Actor neural network are updated π , and update the parameters θ of the Critic neural network Q , so that the updated network parameters are the same as the target network parameters, or the updated network parameters and the target network parameters meet the preset update conditions, which is not limited in this application.

[0091] In a possible implementation, the implementation process of step 230 may be: according to the accumulated error between the current reward parameter and the next reward parameter, and the preset loss function, the parameter θ of the Critic neural network is adjusted. Q , and then adjust the parameters θ of the Actor neural network according to the updated parameters of the Critic neural network π to update.

[0092] It should be noted that the network parameter update in deep reinforcement learning is updated using a gradient strategy. Therefore, if the updated network parameters of the initial reinforcement learning agent are different from the target network parameters, multiple experience action data are obtained from the replay buffer again, and the above steps 210-230 are executed in a loop until the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters. The update of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0093] In the embodiment of the present application, a plurality of experience action data are obtained, and the experience action data are the experience actions of multiple target sample agents in the evolutionary population and the environment for interactive learning; based on the plurality of experience action data, the reward information of the action data output by the initial reinforcement learning agent is obtained; according to the reward information and the preset loss function, the network parameters of the initial reinforcement learning agent are updated; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the update of the network parameters of the initial reinforcement learning agent is terminated, and the reinforcement learning agent that has been trained is obtained. That is, the present application combines evolutionary algorithms with deep reinforcement learning, and the two algorithms complement each other, using the evolution of evolutionary algorithms as the experience pool of deep reinforcement learning, and learning experience actions through interaction between sample agents in the evolutionary population and the environment, increasing the frequency of interaction between agents and the environment, so as to explore and learn more experience actions in the environment. Further, by training the initial reinforcement learning agent through the experience actions of multiple target sample agents in the evolutionary population and the environment for interactive learning, the learning efficiency and effect of deep reinforcement learning can be improved, so as to better control the deep reinforcement learning agent to complete the continuous control task.

[0094] In view of the three defects of deep reinforcement learning, the problem of sparse rewards for short-term tasks related to actions, hyperparameter sensitivity and fragile convergence characteristics, this application adopts evolutionary algorithms (EA) to make improvements, but evolutionary algorithms cannot solve the problem of premature convergence to local optimality in high-dimensional action and state space. Furthermore, regarding the problem of dimensionality reduction, this application adopts evolutionary strategies to make improvements.

[0095] Based on this technical concept, the present application reproduces the initial sample agents in the evolving population through evolutionary algorithms and evolutionary strategies, so as to use the reproduced sample agents to explore and learn in an interactive environment to obtain more experience actions.

[0096] Among them, the evolutionary algorithm is essentially a random local search optimization algorithm, which is a population-based method that can generate, maintain and optimize a population of candidate solutions. In the specific implementation, the population is initialized first and the candidate solutions are evaluated. Then a cycle begins. During this period, a better parent is selected in each iteration, and the sub-population is used to generate new candidate solutions with random mutation operators. Each solution in the sub-population will be evaluated. Finally, the initial population will be partially or completely replaced by the newly generated population. The iterative cycle will continue until the stopping criteria are met.

[0097] Evolutionary algorithms have three main advantages: First, as a class of black-box optimization methods, evolutionary algorithms are indifferent to reward distribution (sparse or dense), do not require back-propagation of gradients, and tolerate potentially arbitrarily long time frames. An evolutionary algorithm usually uses a fitness metric that integrates returns throughout the iterative episode to ensure long-term robustness; second, the population-based characteristics of evolutionary algorithms also have the advantage of supporting diverse exploration, especially when combined with explicit diversity maintenance techniques. Trial-and-error learning algorithms such as evolutionary algorithms allow agents to creatively discover better behaviors without being restricted by the designer's assumptions about how solutions occur and how to optimize each solution, because the population at the core of such algorithms naturally covers a wide range of extended behaviors; third, evolutionary algorithms can effectively solve the fragile convergence characteristics and robustness problems in deep reinforcement learning. The inherent redundancy in the population also promotes robustness and stable convergence characteristics, especially when combined with the "elite selection" technique, which aims to improve the quality of solutions and convergence speed (reaching the global optimum) at acceptable memory and computational costs.

[0098] However, evolutionary algorithms often suffer from high sample complexity and often have difficulty solving high-dimensional problems that require optimizing a large number of parameters. The main reason is that evolutionary algorithms cannot take advantage of powerful gradient descent methods, which are at the heart of more efficient deep reinforcement learning methods. Therefore, combining evolutionary algorithms with deep reinforcement learning can combine the advantages of both and overcome the above-mentioned shortcomings of deep reinforcement learning.

[0099] Furthermore, as an optimization technique, evolution strategies can perform comparable to standard reinforcement learning techniques on reinforcement learning benchmarks (such as Atari / MuJoCo), while also overcoming many of the inconveniences of reinforcement learning. As a special form of controller design, evolution strategies obtain multiple candidate models by perturbing the internal parameters of the model in deep network learning, and finally update the internal parameters according to the reward value to obtain better ones. The evolution strategy (ES) algorithm is summarized as the following equation (1):

[0100]

[0101] Among them, F is the evaluation function (objective function), θ is the internal parameter of the model, and θ is sampled using distribution P. ψ is the internal parameter of P distribution.

[0102] The algorithm equation of the evolution strategy explains: If you want to optimize the internal parameters of the value evaluation function F, you can first sample some internal parameters θ on the right side of the equation, calculate the evaluation value F(θ) under these θ, and calculate the second term logp ψ (θ), that is, in which direction ψ should be updated to obtain the desired θ.

[0103] Among them, the combination of evolutionary strategies and uncertainty handling techniques can reduce the cost of internal algorithms, thus providing an effective solution for dealing with high-dimensional optimization problems. Some evolutionary strategy mechanisms increase the robustness of the search. In addition, direct strategy search (such as random search) does not suffer from time credit allocation problems or sparse rewards, which is very important for avoiding high deviations in sampling environments. In addition, evolutionary strategies ignore the information contained in each state transition and reward.

[0104] Based on the advantages of evolutionary algorithms and evolutionary strategies, this application uses the evolutionary population as a sample data set for deep reinforcement learning. Evolutionary reinforcement learning inherits the fitness method of the evolutionary algorithm and handles temporary credit allocation by consolidating the return value of the entire training phase. The selection operator of evolutionary reinforcement learning operates based on this fitness, exerting selection pressure on regions of the strategy space, resulting in higher returns across the training phase. In addition, evolutionary reinforcement learning inherits the population-based method of the evolutionary algorithm, resulting in redundancy, thereby stabilizing the convergence characteristics and making the learning process more robust.

[0105] In one possible implementation, evolutionary reinforcement learning uses a two-layer loop learning, where the same set of data (experience) generated by the evolving population is used by the reinforcement learning agent. The recycling of the same data can maximize the information extracted from individual experience, thereby improving sample efficiency.

[0106] In summary, this application uses evolutionary algorithms and evolutionary strategies to enrich sample agents, so that sample agents can interact with the environment and explore and learn more efficient experiential actions.

[0107] Next, the learning process of the sample agent in the evolving population and the process of creating a circular replay buffer based on the learned empirical action data are explained.

[0108] In one embodiment, Figure 3 As shown, before obtaining a plurality of experience action data from the circular replay buffer (corresponding to the above step 210), the implementation process of creating a circular replay buffer includes the following steps:

[0109] Step 310: The first sample intelligent agent in the evolutionary population is reproduced by using an evolutionary strategy to obtain a second sample intelligent agent.

[0110] The first sample agent is a sample agent that meets the preset fitness condition in the evolutionary population. Multiple sample agents are pre-added in the evolutionary population, and the sample agents can interact with the environment to learn corresponding experience actions in the environment. However, the exploration ability of the sample agent is limited. In order to obtain more experience actions, it is necessary to reproduce the multiple sample agents pre-added so that more sample agents can interact with the environment and learn more experience actions.

[0111] As an example, the preset fitness condition is that the fitness is less than a preset fitness threshold.

[0112] In a possible implementation, the implementation process of the above step 310 may be: according to the fitness score of each sample agent interacting with the environment, a first sample agent is selected from multiple sample agents in the evolution population. The first agent is reproduced by an evolution strategy, and the reproduced offspring sample agents interact with the environment. Further, the fitness of each offspring sample agent is obtained, and the offspring sample agent whose fitness is less than a preset fitness threshold is used as the second sample agent.

[0113] As an example, if there are 20 first sample agents and 40 sample agents are obtained after breeding, the fitness of these 40 sample agents is obtained. If the fitness of 30 sample agents among the 40 sample agents after breeding is less than the preset fitness threshold, the number of second sample agents obtained is 30.

[0114] Step 320: The second sample intelligent agent is reproduced by an evolutionary algorithm to obtain a third sample intelligent agent.

[0115] In one possible implementation, the implementation process of step 320 may be: breeding the second sample agent through an evolutionary algorithm; similarly, the offspring sample agent after breeding interacts with the environment to obtain the fitness of each offspring sample agent, and the offspring sample agent whose fitness is less than a preset fitness threshold is used as the third sample agent.

[0116] It should be noted that the fitness conditions of the evolution strategy and the evolution algorithm can be the same or different. That is, the values ​​set for the fitness thresholds can be the same or different, and this application does not impose any restrictions on this.

[0117] Step 330: Determine multiple target sample agents based on the fitness of the first sample agent, the second sample agent, and the third sample agent.

[0118] It should be noted that the target sample agent is a sample agent that interacts well with the environment and can learn effective actions from the environment. The fitness of the target agent is greater than the preset fitness threshold.

[0119] Therefore, in a possible implementation, the implementation process of the above step 330 can be: according to the multiple sample agents included in the evolutionary population and the preset fitness condition, the first sample agent that meets the fitness condition is reproduced to obtain the second sample agent, and the other sample agents that do not meet the fitness condition (i.e., the fitness is greater than the preset fitness threshold) are used as target agents. Similarly, the second sample agent is reproduced, and the one that meets the fitness condition is determined as the third sample agent, and the other sample agents that do not meet the fitness condition (i.e., the fitness is greater than the preset fitness threshold) are determined as target sample agents.

[0120] It should be noted that, for the third sample agent, in some embodiments, the third sample agent can be eliminated, and the target sample agent can be extracted through the first sample agent and the second sample agent. In other embodiments, the above steps 310-320 are looped, and the third sample agent is further propagated to obtain the target sample agent. This embodiment of the application does not limit this.

[0121] Step 340: Store the experiential action data of each target sample agent's interactive learning with the environment into a circular replay buffer.

[0122] Among them, the loop replay buffer is initially empty and has an experience threshold set. When the action experience data in the loop replay buffer reaches the preset experience threshold, the experience action data stored in the loop replay buffer will be used to train the initial reinforcement learning agent.

[0123] In a possible implementation, the implementation process of the above step 340 may be: storing the experience action data of each target sample agent's interactive learning with the environment in the form of experience tuples in a circular replay buffer.

[0124] Each experience tuple includes: the current environment state, the sample action in the continuous action space executed by the target sample agent in response to the current environment state, the reward obtained by the target sample agent after executing the sample action, and the next environment state of the target sample agent's interaction environment.

[0125] As an example, an experience tuple can be represented as: (s i ,a i ,r i ,s i +1).

[0126] Where i represents the time step, s i is the state of the target sample agent at the current time point, a i is the action performed by the target sample agent in response to the current environment state, r i is the reward information obtained after the target sample agent performs action ai, s i +1 is the target sample agent performing action a i The next state after.

[0127] In this embodiment, the sample agents in the evolving population are reproduced through evolutionary strategies and evolutionary algorithms, and the target sample agents are selected from them according to their fitness. The experiential actions of the target sample agents' interactive learning with the environment are stored in a circular replay buffer, providing a high-quality training sample data set for the reinforcement learning agent.

[0128] Next, through Figure 4 and Figure 5 The corresponding embodiments explain the implementation process of the evolutionary strategy and the evolutionary algorithm to optimize the sample agents in the population.

[0129] Based on the above Figure 3 In one embodiment, as shown in Figure 4As shown, the implementation process of performing breeding processing on the first sample agent in the evolutionary population through the evolution strategy to obtain the second sample agent (the above step 310) includes the following steps:

[0130] Step 410: Recombining the first sample agent to obtain the first offspring.

[0131] In this step, at least two first sample agents are selected from a plurality of first samples as parent generations, and the parent generations are recombined to obtain first offspring generations.

[0132] In other words, the parent can be a subset of the first sample agent or a true subset thereof. The parent can be randomly selected or selected from the first sample agent based on a preset selection method, and the present application does not limit this. As an example, the preset selection method can be to select multiple first sample agents whose fitness is lower than the fitness threshold as the parent based on a ranking from high to low based on fitness.

[0133] In a possible implementation, the recombination process includes: discrete recombination, median recombination and hybrid recombination. In the above step 410, any one of the recombination process modes may be selected, and the present application does not impose any limitation thereto.

[0134] Among them, discrete recombination is to randomly select two parent individuals in the parent generation, and then randomly exchange their components to form the components of the new offspring individuals, thereby obtaining the first offspring; median recombination is also to randomly select two parent individuals in the parent generation, and then use the average value of each component of the parent individuals as the component of the new offspring individuals, thereby obtaining the first offspring; the characteristic of mixed recombination lies in the selection of parent individuals. Mixed recombination is to randomly select a fixed parent individual first, and then randomly select the second parent individual from the parent population for each component of the offspring individual, that is, the second parent individual is variable. As for the specific combination of the two parent individuals, both the discrete method and the median method can be used, and even the 1 / 2 in the median recombination can be changed to any weight between [0, 1].

[0135] In addition, for the first generation, the fitness score of each first generation is obtained, and the first K first generations with high fitness scores are used as candidate solutions, which are not affected by the mutation process. The experience actions learned after the candidate solution interacts with the environment are directly stored in the loop replay buffer. Among them, K>0.

[0136] For the other first children except the candidate solution in the first child, the following step 420 is performed.

[0137] Step 420: Perform mutation processing on the first offspring to obtain the second offspring.

[0138] Among them, the mutation process is to add a Gaussian distribution change with zero mean and standard deviation to each component of the first offspring to produce the second offspring. The standard deviation is the degree of variation, which does not always remain unchanged. At the beginning of the evolution strategy algorithm, the degree of variation is relatively large, and when it is close to convergence, the degree of variation will begin to decrease.

[0139] It should be noted that, in order to increase the diversity of mutations, the present application predefines an Ornstein-Uhlenbeck noise generator O. Furthermore, through the noise processor O, a small amount of noise is added to the components of each offspring in the second offspring.

[0140] The above steps 410-420 are executed cyclically, and evolution and selection are repeated until convergence is reached. There are two types of selection in the evolutionary strategy: (μ+λ) selection is to deterministically select μ individuals from μ parent individuals and λ second offspring to form a new population of the next generation; (μ, λ) selection is to deterministically select μ individuals from λ second offspring (λ>μ) to form the next generation population, and each individual survives only one generation and is then replaced by a new individual. (μ+λ) selection can ensure the survival of the best individuals, so that the evolution process of the population shows a monotonically increasing trend. (μ+λ) selection retains old individuals, which is sometimes a local optimal solution.

[0141] In other words, for the new population composed of the second offspring, the above steps 410-420 are executed in a loop to optimize the sample agents in the population until the maximum number of iterations is reached and multiple second offspring are obtained.

[0142] Step 430: According to the first fitness function preset by the evolution strategy, the second offspring that meets the fitness condition in the second offspring is used as the second sample intelligent agent.

[0143] In a possible implementation, according to the first fitness function preset by the evolution strategy, N second offspring with low fitness scores are selected from the second offspring, and the N second offspring are used as second sample agents. In addition, the other second offspring except the N second offspring in the second offspring are used as target sample agents. Where N>0.

[0144] In this embodiment, the sample agents in the evolutionary population are optimized through evolutionary strategies, so that the sample agents can better interact with the environment and learn rich and effective experiential actions.

[0145] Based on the above Figure 3 In one embodiment, as shown in Figure 5 As shown, the implementation process of performing breeding processing on the second sample agent through the evolutionary algorithm to obtain the third sample agent (the above step 320) includes the following steps:

[0146] Step 510: Perform cross processing on each sample agent group in the second sample agent to obtain a third offspring; each sample agent group includes any two sample agents in the second sample agent.

[0147] The sample agent group is obtained by randomly combining two of the second sample agents.

[0148] In this step, for two second sample agents in each sample agent group, part of the binary code values ​​of the two second sample agents are exchanged, that is, the crossover process is completed, and each sample agent group obtains a third offspring after the crossover process.

[0149] Similarly, for the third generation, the fitness scores of each third generation are obtained, and the first M third generations with high fitness scores are used as the population elite (this population refers to the new population composed of all third generations), which is not affected by the mutation process. The experience actions learned by the population elite after interacting with the environment are directly stored in the loop replay buffer. Among them, M>0.

[0150] For the other third generations except the population elites in the third generation, the following step 520 is performed.

[0151] Step 520: Perform mutation processing on the third offspring to obtain the fourth offspring.

[0152] There is a predefined random number generator r()∈[0,1), and the mutation process is to change the values ​​of certain positions of the binary code of the third child. According to the random number generated by the random number generator, the third child code number is mutated, for example, 1 is changed to 0, or 0 is changed to 1.

[0153] Step 530: According to the second fitness function preset by the evolutionary algorithm, the fourth offspring that meets the fitness condition in the fourth offspring is used as the third sample intelligent agent.

[0154] In a possible implementation, according to the second fitness function preset by the evolutionary algorithm, L fourth offspring with the lowest fitness scores are selected from the fourth offspring, and the L fourth offspring are used as the third sample intelligent agents. In addition, the fourth offspring other than the L fourth offspring in the fourth offspring are used as the target sample intelligent agents. Wherein, L>0.

[0155] In this embodiment, the sample agents in the evolutionary population are optimized by using an evolutionary algorithm, so that the sample agents can better interact with the environment and learn rich and effective experiential actions.

[0156] Based on any of the above embodiments, in optimizing the evolutionary population through evolutionary strategies and evolutionary algorithms, the experiential actions learned by the target agent through interaction with the environment are cached in a loop replay buffer, and the experiential action data can be obtained from the loop replay buffer to train the initial reinforcement learning agent.

[0157] In one embodiment, Figure 6 As shown, the implementation process of obtaining multiple experience action data (the above step 210) includes the following steps:

[0158] Step 610: Obtain a preset number of experience tuples from the circular replay buffer.

[0159] It should be noted that in deep reinforcement learning, the training of the reinforcement learning agent is updated using a gradient strategy. By repeatedly updating the network parameters, the network parameters of the initial reinforcement learning agent are made the same as the target network parameters.

[0160] Therefore, when the experience tuple data in the loop replay buffer reaches a preset experience threshold, a preset number of experience tuples are obtained from the loop replay buffer in batches, and the initial reinforcement learning agent is trained based on the preset number of experience tuples obtained each time.

[0161] The preset number may be any real number, such as 20, 30, etc., and the sample actions included in the experience tuples obtained based on the preset number are continuous in the action space.

[0162] Step 620: Determine the experience action data in a preset number of experience tuples as a plurality of experience action data.

[0163] That is, the experience action data in a preset number of experience tuples are used as the experience action data for initial reinforcement learning agent training.

[0164] It should be noted that the loop replay buffer set in this application is the core mechanism between the evolving population and the reinforcement agent. Through the interaction between the sample agent in the evolving population and the environment, the rich, effective and high-frequency actions learned by the target sample agent in the evolving population are stored in the loop replay buffer. Then, a preset number of experience tuples are obtained from the loop replay buffer to train the initial reinforcement learning agent.

[0165] In this embodiment, a preset number of experience tuples are obtained in batches from the loop replay buffer, and the preset number of experience tuples are used to train the initial reinforcement learning agent. Since the loop replay buffer stores the high-frequency actions learned by the target sample agent in the evolutionary population, the training effect and performance of the reinforcement learning agent can be improved by using the experience tuples in the loop replay buffer to train the initial reinforcement learning agent.

[0166] Based on any of the above embodiments, the experience tuples in the loop replay buffer are used to train the initial reinforcement agent. It should be noted that evolutionary reinforcement learning is instantiated through deep deterministic policy gradient (DDPG). Therefore, the initial reinforcement learning agent in this application adopts the Actor-Critic (actor neural network-evaluator neural network) dual network system in DDPG.

[0167] That is, the initial reinforcement learning agent includes four neural networks: Actor neural network, target Actor neural network, Critic neural network and target Critic neural network. In the initialization process, the network parameters of each neural network are pre-set.

[0168] Based on this, in one embodiment, Figure 7 As shown, the implementation process of obtaining the reward information of the action data output by the initial reinforcement learning agent (the above step 220) based on multiple experience action data includes the following steps:

[0169] Step 710: For each experience tuple in a preset number of experience tuples, obtain a first reward for the initial reinforcement learning agent based on the action data output by the initial reinforcement learning agent in response to the current environment state in each experience tuple.

[0170] Among them, each experience tuple includes: the current environment state, the sample action in the continuous action space executed by the target sample agent in response to the current environment state, the reward obtained by the target sample agent after executing the sample action, and the next environment state of the target sample agent's interactive environment.

[0171] Therefore, in a possible implementation, the implementation process of step 710 may be: the Actor neural network responds to the current environment state in each experience tuple and selects the action data output by the initial reinforcement learning agent based on its own initialized network parameters. The Critic neural network scores the output action based on the action data output by the initial reinforcement learning agent to obtain the first reward of the initial reinforcement learning agent.

[0172] Among them, the first reward is the reward value obtained by the initial reinforcement learning agent outputting the current action.

[0173] Step 720: Obtain a second reward for the initial reinforcement learning agent based on the action data output by the initial deep learning agent in response to the next environment state in each experience tuple.

[0174] In one possible implementation, the target Actor neural network responds to the next environment state in each experience tuple and selects the action data output by the initial reinforcement learning agent based on its own initialized network parameters. The target Critic neural network scores the output action based on the action data output by the initial reinforcement learning agent to obtain the second reward of the initial reinforcement learning agent.

[0175] Among them, the second reward is the reward value obtained by the initial reinforcement learning agent outputting the next action.

[0176] Step 730: Determine reward information based on the first reward and the second reward.

[0177] In one possible implementation, the implementation process of step 730 may be: determining reward information according to the accumulated error between the first reward and the second reward corresponding to each experience tuple in each experience tuple. In another possible implementation, the implementation process of step 730 may also be: determining reward information according to the error between the accumulated value of the first reward and the accumulated value of the second reward corresponding to each experience tuple in each experience tuple. This application makes a limitation on this.

[0178] As an example, if the acquired experience tuple is represented as (s i ,a i ,r i ,s i +1), the reward information can be determined by the following formula (2):

[0179] y i =r i +γQ′(s i +1,π′(s i +1|θ π′ )|θ Q′ ) (2)

[0180] Among them, y i represents the cumulative error between the first reward and the second reward, r i represents the first reward, s i +1 represents the next environment state, Q′ represents the reward function of the target Critic neural network, θ Q′ represents the target network parameters of the target Critic neural network, π′ represents the target Actor neural network, and θ π′ Represents the target network parameters of the target Actor's neural network.

[0181] In this embodiment, the initial reinforcement learning agent is trained by the DDPG algorithm to obtain reward information of the action data output by the initial reinforcement learning agent in response to the current environment state and the next environment state in each experience tuple. In this way, the network parameters of the initial reinforcement learning agent can be updated according to the reward information to obtain a trained reinforcement learning agent.

[0182] Furthermore, according to the above-mentioned embodiment, the reward information of the action data output by the initial reinforcement learning agent and the preset loss function are obtained, and the network parameters of the initial reinforcement learning agent can be updated.

[0183] In some possible implementations, the loss functions used when training a reinforcement learning agent include: a mean square error loss function and a gradient strategy loss function.

[0184] Based on this, in one embodiment, Figure 8 As shown, the implementation process of updating the network parameters of the initial reinforcement learning agent (the above step 230) according to the reward information and the preset loss function includes the following steps:

[0185] Step 810: According to the reward information and the mean square error loss function, the reward parameters are updated through back propagation of the gradient neural network.

[0186] In a possible implementation, the mean square error loss function can be expressed by the following formula (3):

[0187]

[0188] Among them, L represents the updated value of the reward parameter, T is the number of experience tuples, and y i is the reward information of the action data output by the initial reinforcement learning agent, Q is the reward function of the Critic neural network, and s i is the current environment state, a i is the action output by the initial reinforcement learning agent, θ Q are the network parameters of the Critic neural network.

[0189] It should be noted that updating the reward parameters is actually updating the network parameters of the Critic neural network, so that the Critic neural network can output rewards with higher accuracy when evaluating the actions output by the initial reinforcement learning agent.

[0190] Step 820: Based on the updated reward parameters and gradient strategy loss function, the network parameters of the initial deep learning agent are updated through back propagation of the gradient neural network.

[0191] In one possible implementation, the gradient strategy loss function can be expressed by the following formula (4):

[0192]

[0193] in, represents the policy gradient, which is obtained by simply sampling different Actor neural networks and averaging to estimate the policy gradient, π represents the policy function of the Actor neural network, and θ π represents the network parameters of the Actor neural network, Q represents the reward function of the Critic neural network, and θ Q Represents the network parameters of the Critic neural network.

[0194] It should be noted that the initial reinforcement learning agent is trained by repeatedly obtaining a preset number of experience tuples from the loop replay buffer to continuously update the network parameters of the initial deep learning agent. The update target of the network parameters is shown in the following formulas (5) and (6):

[0195] θ π′ ←τθ π +(1-τ)θ π′ (5)

[0196] θ Q′ ←τθ Q +(1-τ)θ Q′ (6)

[0197] Among them, τ is the update coefficient, θ π is the network parameter of the Actor neural network, θ π′ is the target network parameter of the target Actor neural network; θ Q is the network parameter of the Critic neural network, θ Q′ is the target network parameter of the target Critic neural network.

[0198] That is, updating the network parameters of the initial reinforcement learning agent includes: updating the network parameters of the Actor neural network so that the network parameters of the updated Actor neural network are the same as the target network parameters of the target Actor neural network; updating the network parameters of the Critic neural network so that the network parameters of the updated Critic neural network are the same as the target network parameters of the target Critic neural network.

[0199] In this embodiment, the reward parameters are updated through the back propagation of the gradient neural network through the reward information and the mean square error loss function, and then the network parameters of the initial reinforcement learning agent are updated through the back propagation of the gradient neural network according to the updated reward parameters and the gradient strategy loss function. In this way, the network parameters of the initial reinforcement learning agent are updated through the back propagation of the gradient neural network, which improves the training efficiency of the reinforcement learning agent.

[0200] Regarding the above Figure 2-8 The agent training method shown in the figure can be found in the following table. Fig. 9 Schematic diagram of the training system shown.

[0201] For the sample agents in the evolving population, on the one hand, they are optimized through evolutionary strategies, including parent selection, mutation recombination, offspring generation, and offspring fitness evaluation; on the other hand, they are optimized through evolutionary algorithms, including parent selection, crossover mutation, obtaining a new population composed of offspring and merging them into the evolving population.

[0202] The sample agents in the evolving population interact with the environment and learn experiential actions. According to the fitness scores of the sample agents, the experiential actions learned by the target sample agents with high fitness scores are stored in the circular replay buffer in the form of experiential tuples.

[0203] When training a reinforcement learning (RL) agent, a preset number of experience tuples are obtained in batches from the loop replay buffer, and the initial reinforcement learning agent is trained through the experience action data of these experience tuples. Then, the network parameters of the Actor neural network and the Critic neural network are continuously updated through the back propagation of the gradient neural network until the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, then the update of the network parameters of the initial reinforcement learning agent is terminated, and the trained reinforcement learning agent is obtained.

[0204] In addition, based on the agent training method shown in the above embodiment, in the process of training the initial reinforcement learning agent, it is also necessary to synchronously update the sample agents in the evolutionary population.

[0205] In one possible implementation, the network parameters of the reinforcement learning agent are copied to all sample agents in the evolving population according to a preset synchronization period, and the network parameters of the reinforcement learning agent are used to update the network parameters of all sample agents in the evolving population.

[0206] As an example, pop represents the evolution population, ω represents the preset synchronization period, and mod is the remainder operation. If the pop of the evolution population πIf modω=0, the network parameters of the Actor neural network in the reinforcement learning agent are copied to the population, and the worst policy function π in the evolving population is updated so that the network parameters of the Actor neural network of the sample agent in the evolving population are the same as those of the Actor neural network of the reinforcement learning agent.

[0207] In this embodiment, the network parameters of the Actor neural network of the reinforcement learning agent are periodically copied to the evolving sample agent population. The frequency of synchronization controls the flow of information from the reinforcement learning agent to the evolving population. This is the core mechanism that enables the evolutionary framework to directly utilize information obtained through gradient descent. The process of permeating the strategy learned by the reinforcement learning agent into the evolving population also helps stabilize learning and make it less resistant to deception.

[0208] Furthermore, if the policy learned from the RL agent is good, it will be selected in the evolving population and survive, with its impact extending to future generations. If the policy learned from the RL agent is bad, it will simply be unselected and discarded in the evolving population. This mechanism ensures that the information flow from the RL agent to the evolving population is constructive, rather than destructive. This is especially important for domains with sparse rewards and deceptive local minima, to which gradient-based methods can be extremely vulnerable.

[0209] Based on any of the above embodiments, the present application also provides another intelligent agent method, such as Fig.10 As shown, this method is applied to Figure 1 Taking the mobile device in the example as an example, the agent training method includes the following steps:

[0210] Step 1002: Performing breeding processing on the first sample intelligent agent in the evolution population through an evolution strategy to obtain a second sample intelligent agent;

[0211] Step 1004: breeding the second sample intelligent agent through an evolutionary algorithm to obtain a third sample intelligent agent;

[0212] Step 1006: Determine a plurality of target sample agents according to the fitness of the first sample agent, the second sample agent, and the third sample agent;

[0213] Step 1008: storing the experience action data of each target sample agent's interactive learning with the environment in the form of experience tuples in a circular replay buffer;

[0214] Step 1010: Obtain a preset number of experience tuples from the circular replay buffer;

[0215] Step 1012: for each experience tuple in the preset number of experience tuples, obtaining a first reward of the initial reinforcement learning agent according to the action data output by the initial reinforcement learning agent in response to the current environment state in each experience tuple;

[0216] Step 1014: Obtain a second reward for the initial reinforcement learning agent according to the action data output by the initial deep learning agent in response to the next environment state in each experience tuple;

[0217] Step 1016: Determine reward information according to the first reward and the second reward;

[0218] Step 1018: Update the reward parameters through back propagation of the gradient neural network according to the reward information and the mean square error loss function;

[0219] Step 1020: Update the network parameters of the initial deep learning agent through back propagation of the gradient neural network according to the updated reward parameters and gradient strategy loss function;

[0220] Step 1022: According to a preset synchronization period, the network parameters of the reinforcement learning agent are copied to all sample agents in the evolving population. The network parameters of the reinforcement learning agent are used to update the network parameters of all sample agents in the evolving population.

[0221] Above Fig.10 The specific implementation process and beneficial effects of the steps shown can be found in the above embodiments and will not be described in detail here.

[0222] It should be understood that although Figure 2-10 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-10 At least part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the steps or stages in other steps.

[0223] In one embodiment, Fig.11 As shown, an intelligent agent training device is provided, the device 1100 includes: an experience acquisition module 1110, an action training module 1120 and a parameter updating module 1130, wherein:

[0224] The experience acquisition module 1110 is used to acquire a plurality of experience action data, where the experience action data is the experience actions of a plurality of target sample agents in the evolutionary population interacting and learning with the environment;

[0225] An action training module 1120 is used to obtain reward information of action data output by an initial reinforcement learning agent based on a plurality of empirical action data;

[0226] The parameter updating module 1130 is used to update the network parameters of the initial reinforcement learning agent according to the reward information and the preset loss function; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0227] In one embodiment, the apparatus 1100 further includes:

[0228] The first sample optimization module is used to perform breeding processing on the first sample intelligent agent in the evolutionary population through an evolutionary strategy to obtain a second sample intelligent agent; the first sample intelligent agent is a sample intelligent agent in the evolutionary population that meets a preset fitness condition;

[0229] The second sample optimization module is used to perform breeding processing on the second sample intelligent agent through an evolutionary algorithm to obtain a third sample intelligent agent;

[0230] A determination module, used to determine a plurality of target sample agents according to the fitness of the first sample agent, the second sample agent and the third sample agent;

[0231] The experience storage module is used to store the experience action data of each target sample agent's interactive learning with the environment into a circular replay buffer.

[0232] In one embodiment, the first sample optimization module includes:

[0233] A recombination unit, used for performing recombination processing on the first sample intelligent agent to obtain a first offspring;

[0234] A mutation unit, used to perform mutation processing on the first offspring to obtain a second offspring;

[0235] The first screening unit is used to select the second offspring that meets the fitness condition in the second offspring as the second sample intelligent agent according to the first fitness function preset by the evolution strategy.

[0236] In one embodiment, the second sample optimization module includes:

[0237] A crossover unit, used for performing crossover processing on each sample agent group in the second sample agent to obtain a third offspring; each sample agent group includes any two sample agents in the second sample agent;

[0238] A mutation unit, used for performing mutation processing on the third offspring to obtain a fourth offspring;

[0239] The second screening unit is used to select the fourth offspring that meets the fitness condition in the fourth offspring as the third sample intelligent agent according to the second fitness function preset by the evolutionary algorithm.

[0240] In one embodiment, the experience storage module is specifically used to:

[0241] The experience action data of each target sample agent's interactive learning with the environment are stored in a circular replay buffer in the form of experience tuples; each experience tuple includes: the current environment state, the sample action in the continuous action space executed by the target sample agent in response to the current environment state, the reward obtained by the target sample agent after executing the sample action, and the next environment state of the target sample agent's interactive environment.

[0242] In one embodiment, the experience acquisition module 1110 includes:

[0243] an acquisition unit, used for acquiring a preset number of experience tuples from the circular replay buffer;

[0244] The first determining unit determines the experience action data in a preset number of experience tuples as a plurality of experience action data.

[0245] In one embodiment, the action training module 1120 includes:

[0246] A first action training unit is used to obtain a first reward of the initial reinforcement learning agent according to action data output by the initial reinforcement learning agent in response to the current environment state in each experience tuple of a preset number of experience tuples;

[0247] A second action training unit, configured to obtain a second reward for the initial reinforcement learning agent according to action data output by the initial deep learning agent in response to the next environment state in each experience tuple;

[0248] The second determining unit is used to determine reward information according to the first reward and the second reward.

[0249] In one embodiment, the loss function includes a mean square error loss function and a gradient strategy loss function;

[0250] The parameter updating module 1130 includes:

[0251] A reward parameter updating unit, used to update the reward parameters through back propagation of the gradient neural network according to the reward information and the mean square error loss function;

[0252] The network parameter updating unit is used to update the network parameters of the initial reinforcement learning agent through the back propagation of the gradient neural network according to the updated reward parameters and gradient strategy loss function.

[0253] In one embodiment, the apparatus 1100 further includes:

[0254] The synchronization module is used to copy the network parameters of the reinforcement learning agent to all sample agents in the evolutionary population according to a preset synchronization period. The network parameters of the reinforcement learning agent are used to update the network parameters of all sample agents in the evolutionary population.

[0255] The specific definition of the intelligent agent training device can be found in the definition of the intelligent agent training method above, which will not be repeated here. Each module in the above-mentioned intelligent agent training device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0256] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Fig.12 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, an intelligent agent training method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0257] Those skilled in the art will understand that Fig.12The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0258] In one embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:

[0259] Acquire multiple experience action data, where the experience action data is the experience actions of multiple target sample agents in the evolutionary population interacting and learning with the environment;

[0260] Based on multiple empirical action data, obtain reward information of the action data output by the initial reinforcement learning agent;

[0261] Update the network parameters of the initial reinforcement learning agent based on the reward information and the preset loss function;

[0262] If the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0263] When the computer device provided in this embodiment implements the above steps, its implementation principle and technical effects are similar to those of the above method embodiments, which will not be repeated here.

[0264] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:

[0265] Acquire multiple experience action data, where the experience action data is the experience actions of multiple target sample agents in the evolutionary population interacting and learning with the environment;

[0266] Based on multiple empirical action data, obtain reward information of the action data output by the initial reinforcement learning agent;

[0267] Update the network parameters of the initial reinforcement learning agent based on the reward information and the preset loss function;

[0268] If the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

[0269] When the computer-readable storage medium provided in this embodiment implements the above steps, its implementation principle and technical effects are similar to those of the above method embodiments, and will not be repeated here.

[0270] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0271] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0272] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for training an intelligent agent, characterized in that: The method comprises: Acquire a plurality of experience action data, wherein the experience action data is the experience action of multiple target sample agents in the evolutionary population interacting with the environment; the experience action data includes the environmental state of the target sample agent interaction environment, and the action output by the target sample agent in response to the environmental state; the action output by the agent is a point in the space of the control input for controlling the robot or autonomous vehicle; Based on the multiple empirical action data, obtaining reward information of the action data output by the initial reinforcement learning agent; According to the reward information and a preset loss function, updating the network parameters of the initial reinforcement learning agent; If the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

2. The method according to claim 1, characterized in that Before acquiring a plurality of experience action data, the method includes: The first sample intelligent agent in the evolutionary population is reproduced by an evolutionary strategy to obtain a second sample intelligent agent; the first sample intelligent agent is a sample intelligent agent in the evolutionary population that meets a preset fitness condition; Performing breeding processing on the second sample intelligent agent through an evolutionary algorithm to obtain a third sample intelligent agent; Determine a plurality of target sample agents according to the fitness of the first sample agent, the second sample agent and the third sample agent; The experiential action data of the interactive learning between each target sample agent and the environment is stored in a circular replay buffer.

3. The method according to claim 2, characterized in that The method of performing reproduction processing on the first sample intelligent agent in the evolutionary population by using an evolution strategy to obtain a second sample intelligent agent includes: Performing a recombination process on the first sample agent to obtain a first offspring; Performing mutation treatment on the first progeny to obtain a second progeny; According to the first fitness function preset by the evolution strategy, the second offspring in the second offspring that meets the fitness condition is used as the second sample intelligent agent.

4. The method according to claim 2, characterized in that: The step of performing breeding processing on the second sample agent by using an evolutionary algorithm to obtain a third sample agent includes: Performing cross processing on each sample agent group in the second sample agent to obtain a third offspring; each sample agent group includes any two sample agents in the second sample agent; performing mutation processing on the third offspring to obtain a fourth offspring; According to the second fitness function preset by the evolutionary algorithm, the fourth offspring that meets the fitness condition in the fourth offspring is used as the third sample intelligent agent.

5. The method according to any one of claims 2 to 4, characterized in that: The step of storing the experience action data of the interaction learning between each target sample agent and the environment into a circular replay buffer includes: The experience action data of the interaction learning between each target sample agent and the environment are stored in the loop replay buffer in the form of experience tuples; each experience tuple includes: the current environment state, the sample action of the continuous action space executed by the target sample agent in response to the current environment state, the reward obtained by the target sample agent after executing the sample action, and the next environment state of the interaction environment of the target sample agent.

6. The method according to claim 5, characterized in that The obtaining of a plurality of experience action data comprises: Obtaining a preset number of experience tuples from the loop replay buffer; The experience action data in the preset number of experience tuples are determined as the multiple experience action data.

7. The method according to claim 6, characterized in that The step of obtaining reward information of the action data output by the initial reinforcement learning agent based on the plurality of experience action data includes: For each experience tuple in the preset number of experience tuples, obtaining a first reward of the initial reinforcement learning agent according to action data output by the initial reinforcement learning agent in response to the current environment state in each of the experience tuples; Obtaining a second reward for the initial reinforcement learning agent according to the action data output by the initial reinforcement learning agent in response to the next environment state in each of the experience tuples; The reward information is determined according to the first reward and the second reward.

8. The method according to any one of claims 1 to 4, characterized in that: The loss function includes a mean square error loss function and a gradient strategy loss function; The updating of the network parameters of the initial reinforcement learning agent according to the reward information and the preset loss function includes: According to the reward information and the mean square error loss function, the reward parameter is updated through back propagation of a gradient neural network; According to the updated reward parameters and the gradient strategy loss function, the network parameters of the initial reinforcement learning agent are updated through back propagation of the gradient neural network.

9. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: According to a preset synchronization cycle, the network parameters of the reinforcement learning agent are copied to all sample agents in the evolving population, and the network parameters of the reinforcement learning agent are used to update the network parameters of all sample agents in the evolving population.

10. An intelligent agent training device, characterized in that: The device comprises: An experience acquisition module is used to acquire a plurality of experience action data, wherein the experience action data is the experience action of a plurality of target sample agents in the evolutionary population interacting with the environment; the experience action data includes the environmental state of the target sample agent's interaction environment, and the action output by the target sample agent in response to the environmental state; the action output by the agent is a point in the space of the control input for controlling the robot or autonomous vehicle; An action training module, used for obtaining reward information of action data output by the initial reinforcement learning agent based on the plurality of empirical action data; A parameter updating module is used to update the network parameters of the initial reinforcement learning agent according to the reward information and a preset loss function; if the updated network parameters of the initial reinforcement learning agent are the same as the target network parameters, the updating of the network parameters of the initial reinforcement learning agent is terminated to obtain a trained reinforcement learning agent.

11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Anti-unmanned aerial vehicle task allocation method based on reinforcement learning

    CN112507622A

  • Evolutionary algorithm driven by multi-model online self-adaptive preferential technology based on reinforcement learning

    CN117273125A

  • Method and device for learning a strategy and for implementing the strategy

    US20220027743A1