Labyrinth environment exploration method based on curiosity and curriculum type reinforcement learning
By introducing a method based on curiosity and curriculum-based reinforcement learning in deep reinforcement learning, the problem of exploration-utilization dilemma of agents in sparse reward environments is solved, and the exploration ability and training efficiency of agents in maze environments is improved.
Patent Information
- Application Number
- CN202510102244.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
Deep reinforcement learning faces exploration-utilization dilemma in sparse reward environments, resulting in agents having few samples acquired and model training in maze environments, and they cannot even converge or fall into local optimal solutions.
A maze environment exploration method based on curiosity and curriculum-based reinforcement learning is proposed. By constructing a curriculum learning framework for teacher agents and student agents, combining A2C algorithms and curiosity modules, the exploration ability and sample utilization efficiency of agents are optimized.
It improves the environmental exploration ability and sample utilization efficiency of the agent in sparse reward scenarios, promotes more effective training of the agent in the maze environment, and improves the convergence speed and average reward.
Smart Images

Figure CN120022605A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep reinforcement learning, and in particular to a maze environment exploration method based on curiosity and curriculum-based reinforcement learning. Background Art
[0002] The nature of reinforcement learning based on feedback and evaluation determines that it needs to obtain samples through continuous exploration and discover better strategies from a large amount of sample information, or prove that the current strategy is the best. The introduction of deep neural networks makes reinforcement learning's demand for a large number of samples more obvious. With the continuous expansion of application fields, the bottleneck caused by the high dependence of traditional deep reinforcement learning algorithms on the number of samples, that is, the number of rewards, has gradually emerged. In recent years, Roguelike games have become increasingly popular. Such games often contain elements such as mazes. Some researchers have tried to apply deep reinforcement learning algorithms to game exploration with maze elements. There are two major difficulties in such tasks. First, the reward signal is sparse or difficult to define, which makes the number of samples available to the agent scarce; second, the action space is relatively complex, which leads to slow model training or even failure to converge or fall into a local optimal solution. How to solve the above problems, that is, to explore how to enable computers to use deep reinforcement learning to cope with sparse reward environments, is an important challenge facing deep reinforcement learning.
[0003] For environments with sparse rewards such as mazes, the key to solving the problem is how to balance the exploration-utilization dilemma to improve sample efficiency. In this regard, there are currently two mainstream approaches: one is to optimize the exploration side and enhance the ability of the agent to obtain samples; the other is to optimize the utilization side and improve the model's utilization rate of unit samples. Among them, the ability of the agent to obtain samples is mainly achieved by shaping intrinsic rewards. The classic reinforcement learning algorithm emphasizes obtaining reward signals from external environmental feedback. This reward signal is called extrinsic reward. The agent takes maximizing extrinsic rewards as the training goal. Extrinsic rewards are also the standard for evaluating the pros and cons of action strategies. The development of intrinsic motivation theory has spawned an intrinsic reward mechanism. A large number of experiments have shown that for the policy gradient method, superimposing intrinsic rewards on the basis of extrinsic rewards is the most significant solution to enhance the effect of exploration ability. The shaping of intrinsic rewards includes state prediction, namely curiosity, information entropy maximization, pseudo counting, action difference drive, parameter space noise and other methods. Among them, the most mainstream method is curiosity-driven exploration based on state prediction methods.
[0004] Improving sample utilization can be achieved through curriculum learning, experience replay, round reverse update, and importance sampling. Among them, curriculum learning has long been considered a key component of many machine learning problems. By imitating the effective learning sequence in human education, the original task is decomposed into a series of subtasks, or precursor tasks, when it is difficult to complete the original task. The agent can start learning from simple tasks and then gradually learn more complex tasks. The above training strategy with increasing complexity and difficulty is called curriculum learning.
[0005] However, both of the above coping strategies have certain defects. On the one hand, curiosity-driven exploration based solely on state prediction has limitations. For example, there is often no lack of diverse states, but it is impractical to explore such a large state space. On the other hand, the effects of most course learning methods are not significant enough. Specifically, the course learning methods cannot explicitly improve the exploration ability of the intelligent agent, and manual course learning relies too much on human prior knowledge, which is costly and inflexible, and does not consider the feedback of the model itself during the training process. In addition, samples that humans think are easy are not necessarily easy for the model, that is, the decision boundaries of human and machine models are inconsistent. Summary of the invention
[0006] In view of the above-mentioned deficiencies of the prior art, the present invention proposes a maze environment exploration method based on curiosity and curriculum-based reinforcement learning, aiming to improve the agent's ability to explore the environment and the efficiency of sample utilization in sparse reward scenarios, thereby helping the agent to train more effectively in a maze environment.
[0007] The present invention proposes a maze environment exploration method based on curiosity and curriculum-based reinforcement learning, which comprises the following steps:
[0008] Step 1: Construct a maze game environment, including: several sets of walls, several mechanisms, a pyramid, and an agent for performing maze exploration tasks;
[0009] Step 2: Initialize the environment parameters of the maze game environment, build and train a curriculum learning framework based on curiosity and curriculum reinforcement learning; the curriculum learning framework based on curiosity and curriculum reinforcement learning includes: a teacher agent and a student agent;
[0010] Step 3: Construct a curiosity module, including: a feature extractor, a forward dynamics model, and an inverse dynamics model; the curiosity module is used to generate intrinsic rewards for the student agent;
[0011] Step 4: Based on the curriculum learning framework and curiosity module of curiosity and curriculum-based reinforcement learning, the A2C algorithm is used to train the student agent to obtain the trained student agent, and the trained student agent is used to perform the maze exploration task;
[0012] The step 1 further comprises:
[0013] Step 1.1: Construct a maze consisting of several groups of walls, randomly set several mechanisms and a pyramid made of multiple cubes at any position in the maze, and each pyramid has a target cube on the top;
[0014] Step 1.2: Create an agent G for performing maze exploration tasks, and use the center of the maze as the starting position of agent G;
[0015] The maze exploration task is as follows: the agent G starts from the starting position and starts to look for a mechanism. When the agent G finds a mechanism, it needs to touch the mechanism. At this time, a pyramid appears randomly at any position in the maze. The agent G continues to move in the maze to find the pyramid. When the agent G finds the pyramid, it pushes it down to get the target block set on the top of the pyramid, thereby completing a maze exploration task.
[0016] Step 1.3: Construct a state space S containing the states of agent G at different time steps;
[0017] The state is: For any time step t, the state s of agent G t Including: the speed of agent G, the position of agent G and the label information obtained by the sensor; the label information includes: walls, cubes forming the pyramid, target blocks, triggered mechanisms and untriggered switches;
[0018] Step 1.4: Construct an action space A containing the actions of agent G at different time steps;
[0019] The action is: For any time step t, the action a of agent G t For stationary 1 , walk forward 2 , walk backwards 3 , turn left 4 Or turn right 5 an action in t Denoted as a t ={a l |l∈{1,2,3,4,5}}; where l is the index;
[0020] Step 1.5: Construct the reward function of agent G;
[0021] The reward function of the agent G is: At any time step t, the reward R of the agent G t It is expressed as:
[0022]
[0023] in represents the external reward of agent G at time step t; represents the curiosity-based intrinsic reward obtained by agent G at time step t
[0024] The step 2.1 further comprises:
[0025] Step 2.1: Initialize the environment parameters of the maze game environment, and discretize the initialized environment parameters to generate several subtasks and construct a subtask set;
[0026] The specific content of step 2.1 is: taking the initialized environmental parameters as a subtask, the environmental parameters include: maze size p 1 、Number of wall groups p 2 and the number of organs p 3 ; Based on the initialized environmental parameters, several groups of different environmental parameters are predefined as several subtasks, and the maze size in all subtasks is required to be less than or equal to the maze size in the initialized environmental parameters; a subtask set is constructed using all subtasks; wherein the subtask is recorded as: h = f (p j |j=1,2,3); where h represents the subtask; f represents the mapping function from the environment parameter to the subtask; p represents the environment parameter; p j represents the jth environmental parameter;
[0027] Step 2.2: Use the agent G used to perform the maze exploration task as the student agent, create a teacher agent, and define the actions of the teacher agent The actions of the teacher agent To do this: In the subtask set, select a subtask for the student agent based on the task importance of each subtask;
[0028] Step 2.3: Construct a learning efficiency measure to generate the learning efficiency of each subtask for the student agent and calculate the cumulative reward of the teacher agent;
[0029] Step 2.4: Construct a task selector to update the action value of the teacher agent according to the accumulated reward of the teacher agent, and select subtasks for the student agent based on the action value of the teacher agent;
[0030] The specific contents of step 2.3 are as follows: defining the learning efficiency, and using optimistic initialization to set the initial learning efficiency for each subtask in the subtask set; at the same time, maintaining a queue for each subtask to store the most recent K training scores of the subtask and their corresponding training start time;
[0031] The learning efficiency is: for any subtask, the difference between the training score of the student agent completing the subtask this time and the training score of the student agent completing the subtask last time in a unit of time;
[0032] For any action of the teacher agent The subtask selected by the teacher agent is assigned to the student agent for learning. The student agent generates a training score for the subtask and records the student agent's training score on the subtask and its corresponding training start time in the queue of the subtask. At the same time, the action value of the teacher agent is generated.
[0033] The training scores saved in the queue of each subtask are used to calculate the learning efficiency of the student agent on the subtask, and the cumulative reward of the teacher agent is calculated by fitting the training scores in the queue of each subtask based on the least squares method;
[0034] The cumulative reward of the teacher agent is expressed as:
[0035]
[0036] in represents the cumulative reward of the teacher agent at time step t; K is the queue capacity, which is a hyperparameter to be set; k represents the index value of the training score in the queue; t k Indicates the training start time corresponding to the kth training score; Represents the mean of the training start time corresponding to all training scores in the queue; Indicates that the student agent starts training the subtask at time step t The training score obtained when
[0037] The specific content of step 2.4 is: using the accumulated rewards of the teacher agent to update the action value of the teacher agent, and adaptively adjusting the learning rate according to the updated action value of the teacher agent; establishing a Boltzmann distribution based on the action value of the teacher agent, the teacher agent selects subtasks for the student agent from the subtask set according to the defined sampling probability, and arranges the selected subtasks to the student agent for multi-process parallel training;
[0038] The updating method of the action value of the teacher agent is expressed as:
[0039]
[0040] in represents the updated action value of the teacher agent; α tea represents the learning rate;
[0041] The adaptive method of the learning rate is expressed as:
[0042]
[0043] Among them, They are adjustment factors and are hyperparameters to be set;
[0044]
[0045] in Indicates that the teacher agent takes action The probability of; exp is the exponential function; τ is the temperature parameter to be set; N is the number of subtasks; n is the subtask index; h n Represents the subtask with index n in the subtask set;
[0046] The specific content of step 3 is: when the subtask selected by the task selector is assigned to the student agent for learning, observe the state s of the student agent at time step t t and the state s at time step t+1 t+1 And input feature extractor Perform feature extraction in and get the state s t The corresponding eigenvector and status t+1 The corresponding eigenvector And input the inverse dynamics model to get the predicted action of the student agent Leveraging Predictive Actions The loss function of the inverse dynamics model is constructed based on the actual actions of the student agent, and the optimal inverse dynamics model is obtained by minimizing the loss function.
[0047] The loss function of the inverse dynamics model is:
[0048]
[0049] Where I represents the inverse dynamics model; θ I represents the model parameters of the inverse dynamics model; θ E represents the model parameters of the feature extractor; represents the loss function of the inverse dynamics model; a tis the action of the student agent at time step t; g(·) represents the learning function of the inverse dynamics model; H(·) represents the cross entropy loss function;
[0050] The student agent state s t The corresponding eigenvector and the student agent’s action a at time step t t Input the forward dynamics model for feature prediction and generate the feature vector of the predicted state at time step t+1 Using the feature vector of the predicted state and status t The corresponding eigenvector Constructing a loss function of the forward dynamics model, and obtaining the optimal forward dynamics model by minimizing the loss function;
[0051] The loss function of the forward dynamics model is expressed as:
[0052]
[0053] where F represents the forward dynamics model; θ F represents the model parameters of the forward dynamics model; represents the loss function of the forward dynamics model; f(·) represents the learning function of the forward dynamics; represents the L2 norm;
[0054] Use the feature vector of the student agent's predicted state at time step t+1 and the state s at time step t+1 t+1 Generate the curiosity-based intrinsic reward obtained by the student agent at time step t It is expressed as:
[0055]
[0056] Where η is the standard factor;
[0057] The step 4 further comprises:
[0058] Step 4.1: The teacher agent randomly selects a subtask from the subtask set as the teacher agent’s action And assign this subtask to the student agent for learning;
[0059] Step 4.2: Set the network parameters of the policy network to ω π , the network parameter of the value network is ω v ;
[0060] Step 4.3: Use the policy network to control the student agent to interact with the maze game environment. After completing the learning of the current subtask, the student agent generates a trajectory, which is a sequence of all states, actions, and rewards of the student agent in the process of completing the subtask;
[0061] The policy network is represented as a t ~π(·|s t ;ω v ), which is used to calculate the current state s of the student agent. t To determine the action a that the student agent should take t ;
[0062] The trajectory is recorded as: 1 ,a 1 ,r 1 ,s 2 ,a 2 ,r 2 ,……,s n ,a n ,r n}; where n is the total number of steps required for the student agent to complete the current subtask; s 1 、a 1 and r 1 are the state, action and reward of the student agent at time step 1; s 2 、a 2 and r 2 are the state, action and reward of the student agent at time step 2; s n 、a n and r n are the state, action and reward of the student agent at time step n respectively;
[0063] Step 4.4: For each state of the student agent, the student agent observes the predicted trajectory of the next m steps starting from the current state, including the state estimate, action estimate, and reward estimate of the student agent in the next m steps;
[0064] For the current state s of the student agent t , will be changed from the current state s t The predicted trajectory of the next m steps is recorded as: t ,s t+1 ,a t+1 ,r t+1 ,……,s t+m-1 ,a t+m-1 ,r t+m-1 ,s t+m ,a t+m}, where m is the number of multi-step bootstrap steps to be set; r t is the current statet Corresponding rewards; t+1 ,a t+1 ,r t+1 are the state estimation, action estimation, and reward estimation for time step 1 starting from the current state; s t+m-1 ,a t+m-1 ,r t+m-1 are the state estimation, action estimation and reward estimation of time step m-1 starting from the current state; s t+m ,a t+m They are the state estimation and action estimation of time step m starting from the current state respectively;
[0065] Step 4.5: Use the value network to calculate all state-action pairs (s t ,a t )’s state action value;
[0066] Step 4.6: Based on the multi-step bootstrapping method, construct the objective function of the temporal difference TD target through the intercepted m-step discounted reward, and calculate the TD target of the student agent;
[0067] Step 4.7: Calculate the TD error using the TD target of the student agent and the state-action value of each state-action pair, and construct the loss function;
[0068] Step 4.8: Use the loss function to perform gradient descent on the network parameters of the policy network and the network parameters of the value network, and update the policy network and the value network;
[0069] Step 4.9: Use the learning efficiency measurer to generate the learning efficiency of the student agent on the current subtask and calculate the cumulative reward of the teacher agent;
[0070] Step 4.10: According to the result of the learning efficiency measurer, use the task selector to select the next subtask for the student agent to learn, and repeat steps 4.3-4.9 until the learning efficiency of the student agent reaches the set threshold or reaches the set maximum number of iterations, then stop training, obtain the trained student agent, and use it to perform the maze exploration task;
[0071] The objective function is expressed as:
[0072]
[0073] in is the TD target, which is used to represent the estimated value of the state action value; e represents the time step index; γ is the reward discount factor; γ e and γ m They represent γ to the eth power and γ to the mth power respectively; r t+eRepresents the reward obtained by the student agent at time step t+e, and the reward value is the reward obtained by the student agent from the external environment and intrinsic rewards generated by the curiosity module The sum of V(s t+m ,a t+m ;ω v ) represents the state-action pair (s) of the student agent at time step t+m t+m ,a t+m ) is the state action value of the student agent in state s t+m Take action a t+m Expected return;
[0074] The loss function is expressed as:
[0075]
[0076] Where L(ω v ) is the loss function; δ t represents the TD error at time step t; V(s t ,a t ;ω v ) represents the state-action pair (s t ,a t )’s state action value; where t∈[1,2,3,…,nm];
[0077] The updated network parameters of the strategy network are expressed as:
[0078]
[0079] where ω π,next Represents the updated network parameters of the policy network; α π represents the learning rate of the policy network; represents the gradient of the policy network;
[0080] The updated network parameters of the value network are expressed as:
[0081]
[0082] where ω v,next Represents the updated network parameters of the value network; α v represents the learning rate of the value network; Represents the gradient of the value network.
[0083] The beneficial effects of adopting the above technical solution are:
[0084] The method of the present invention constructs a nested reinforcement learning framework based on the curriculum learning idea, in which the curriculum learning idea is expressed as a teacher-student curriculum learning framework. The inner loop of the teacher-student curriculum learning framework is used as the student end, which is responsible for completing the actual maze exploration task and can use most reinforcement learning algorithms; the outer loop is used as the teacher end, which is responsible for monitoring and evaluating the student's training progress and determining which task the student should train on. The teacher-student curriculum learning framework can automatically formulate the task sequence rather than manually, which can reduce the need for prior domain knowledge in traditional curriculum learning.
[0085] In the method of the present invention, the teacher end can adaptively adjust the learning rate. In the teacher end, a lower learning rate is adopted for subtasks with a higher learning effectiveness change rate; a higher learning rate is adopted for subtasks with a lower learning effectiveness change rate; thereby avoiding problems such as gradient explosion and improving training stability.
[0086] The method of the present invention adopts the optimization learning efficiency based on the fixed-length queue, maintains a queue with a capacity of K for each subtask, and stores the total scores of the most recent K subtasks and their corresponding training start times in the queue. The relationship between the total score of the task and the time is processed using linear regression, and the absolute value of the slope is used as the teacher's reward.
[0087] The method of the present invention uses the A2C algorithm to train the student end. In the A2C algorithm, the single-step bootstrapping in the value network is changed to multi-step bootstrapping, the value distribution of the intelligent agent is compressed to the value of the mth step, and the objective function of the intelligent agent is constructed by intercepting the m-step discounted reward, thereby improving the learning efficiency of the intelligent agent.
[0088] The method of the present invention uses curiosity as an intrinsic reward to effectively improve the exploration ability of the agent in a sparse reward environment. The value of curiosity is defined as the prediction error of the state prediction model for the next state, and a self-supervised inverse dynamics model is introduced to exclude features that are irrelevant to the action of the agent. A curiosity module is added to the training of the student agent, and the intrinsic reward based on curiosity is superimposed on the original external reward, so that the agent can continuously explore unknown states. By integrating the idea of course learning with curiosity, it is possible to simultaneously improve the sample utilization rate and the ability to explore unknown environments, and further improve the efficiency of curiosity in exploring large spaces.
[0089] In summary, the method of the present invention can enable the intelligent agent to achieve effective exploration of the environment and full utilization of samples in a maze environment with sparse rewards, thereby improving the convergence speed and average reward of the intelligent agent in the maze treasure hunt game. BRIEF DESCRIPTION OF THE DRAWINGS
[0090] Figure 1 is a flow chart of a maze environment exploration method based on curiosity and curriculum-based reinforcement learning in this embodiment;
[0091] Figure 2 is a schematic diagram of the student agent training process in this embodiment;
[0092] Figure 3 Schematic diagram of the A2C algorithm in this implementation. DETAILED DESCRIPTION
[0093] In order to facilitate the understanding of the present application, the specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thoroughly understood.
[0094] Since the maze treasure hunt game contains many elements such as variable environmental parameters and random map generation, this implementation method takes the maze treasure hunt game built based on Unity3D as a representative of the maze environment, and designs the state input, action input and reward function of the maze treasure hunt environment; by setting different environmental parameters, a number of subtasks of different difficulties are generated for the original target task; then combined with the teacher-student course learning framework, and optimized the framework by adjusting the learning rate and introducing fixed-length queues, the problem of manual course learning's dependence on expert knowledge and the efficiency bottleneck of curiosity-driven exploration for large state spaces are solved; in the A2C algorithm, the single-step bootstrap in the value network is changed to multi-step bootstrap, which improves the learning efficiency; on the basis of curiosity-driven exploration based on state prediction, a curiosity module is designed, so as to add intrinsic rewards to the training of intelligent agents using the A2C algorithm, and the feasibility of the method is verified through experiments.
[0095] The present embodiment is a maze environment exploration method based on curiosity and curriculum-based reinforcement learning, such as Figure 1 As shown, the method comprises the following steps:
[0096] Step 1: Construct a maze game environment, including: several sets of walls, several mechanisms, a pyramid, and an agent for performing maze exploration tasks.
[0097] Step 1.1: Construct a maze consisting of several groups of walls, randomly set several mechanisms and a pyramid made of multiple cubes at any position in the maze, and there is a target block on the top of each pyramid.
[0098] In this embodiment, under initial conditions, all pyramids in the maze are in a hidden state. When the agent starts to perform the maze exploration task, each time a mechanism is triggered, a pyramid will randomly appear in the maze.
[0099] Step 1.2: Create an agent G for performing maze exploration tasks, and use the center of the maze as the starting position of agent G.
[0100] The maze exploration task is as follows: agent G starts from the starting position and begins to look for a mechanism. When agent G finds a mechanism, it needs to touch the mechanism. At this time, a pyramid appears randomly at any position in the maze. Agent G continues to move in the maze to find the pyramid. When agent G finds the pyramid, it pushes it down to get the target block set on the top of the pyramid, thereby completing a maze exploration task.
[0101] Step 1.3: Construct a state space S containing the states of agent G at different time steps.
[0102] The state is: For any time step t, the state s of agent G t Including: the speed of the intelligent agent G, the position of the intelligent agent G and the label information obtained by the sensor; the label information includes: walls, cubes forming the pyramid, target blocks, triggered mechanisms and untriggered switches.
[0103] In this embodiment, the agent uses three Unity ML-Agent ray sensor components RayPerception Sensor 3D to obtain a stereoscopic perspective of top, middle and bottom. Each of the top, middle and bottom layers has seven rays, totaling 21 rays. The label information obtained by detecting each ray includes walls, cubes that make up the pyramid, target blocks, closed switches, i.e., untriggered switches, and open switches, i.e., triggered switches. The relative positions of the above objects to the agent G are obtained. If they cannot be detected, the relative position is (-∞,-∞,-∞).
[0104] Step 1.4: Construct an action space A containing the actions of agent G at different time steps.
[0105] The action is: For any time step t, the action a of agent G t For stationary 1 , walk forward 2 , walk backwards 3 , turn left 4 Or turn right 5 an action in t Denoted as a t ={a l |l∈{1,2,3,4,5}}; where l is the index.
[0106] In this embodiment, the model output for subsequent agent training is a discrete action, which contains five values: static a 1 , walk forward a2 , walk backwards 3 , turn left 4 , turn right 5 The agent can only perform one action at a time, for example, it cannot back up and turn left at the same time, but fewer outputs will greatly reduce the complexity of the neural network and reduce training time.
[0107] Step 1.5: Construct the reward function of agent G.
[0108] The reward function of the agent G is: At any time step t, the reward R of the agent G t It is expressed as:
[0109]
[0110] in represents the external reward of agent G at time step t; represents the curiosity-based intrinsic reward obtained by agent G at time step t.
[0111] In this embodiment, when the agent reaches a dead end, there will be no optional actions in the next moment, and agent G will receive a large negative reward, that is, except for reaching the end point, that is, touching the target block, agent G cannot obtain effective external rewards from the external environment.
[0112] In this embodiment, the Unity3D engine is used to create a game scene, which consists of a closed platform surrounded by a wall. The starting position of the agent is the center of the scene, that is, the coordinate zero point (0,0,0). The platform is composed of a maze composed of multiple sets of walls. The maze is scattered with several organs that appear randomly at any position in the maze, as well as a pyramid composed of multiple cubes. For the original task environment, the area occupied by the agent unit is the standard unit 1, the area of the entire environment is 400×400 standard units, and the maximum moving speed of the agent G is 5 units per second. The task of the agent G is to find a target block in this rather large environment. The block is located at the top of a pyramid, so if you want to touch this block, you must knock down the pyramid. The pyramid and the target block are not there at the beginning, and you must touch the organ to appear at a random location. Therefore, if the agent G wants to complete the task, it needs to go through the following steps: find the organ, touch the organ, find the pyramid, knock down the pyramid, and touch the target block. The complexity of the steps and the huge environment cause the reward sparse problem, which makes it difficult for the normal algorithm to converge.
[0113] Step 2: Initialize the environmental parameters of the maze game environment, and construct and train a curriculum learning framework based on curiosity and curriculum-based reinforcement learning; the curriculum learning framework based on curiosity and curriculum-based reinforcement learning includes: a teacher agent and a student agent.
[0114] Step 2.1: Initialize the environmental parameters of the maze game environment, and discretize the initialized environmental parameters to generate several subtasks and construct a subtask set.
[0115] The specific content of step 2.1 is: taking the initialized environmental parameters as a subtask, the environmental parameters include: maze size p 1 、Number of wall groups p 2 and the number of organs p 3 ; Based on the initialized environmental parameters, several groups of different environmental parameters are predefined as several subtasks, and the maze size in all subtasks is required to be less than or equal to the maze size in the initialized environmental parameters; a subtask set is constructed using all subtasks; wherein the subtask is recorded as: h = f (p j |j=1,2,3); where h represents the subtask; f represents the mapping function from the environment parameter to the subtask; p j represents the jth environmental parameter.
[0116] In this embodiment, the environment parameters of the maze game environment are initialized, and the parameters when the environment is created are used as the original task. The environment parameters are discretized to divide several subtasks, specifically: based on the initialized environment parameters, by changing different environment parameters, three other subtasks of different difficulty are predefined, and the original task is also one of the subtasks. Each subtask consists of a maze size p 1 、Number of wall groups p 2 、Number of agencies p 3 These three environmental parameters jointly determine the subtask h = f(p j |j=0,1,2,3). The task division based on the discretization of environmental parameters, that is, the environmental parameter settings of tasks of different difficulty levels are shown in Table 1.
[0117] Table 1 Environmental parameter settings for tasks of different difficulty
[0118]
[0119]
[0120] Step 2.2: Use the agent G used to perform the maze exploration task as the student agent, create a teacher agent, and define the actions of the teacher agent
[0121] The actions of the teacher agent To: In the subtask set, select a subtask for the student agent based on the task importance of each subtask.
[0122] Step 2.3: Construct a learning efficiency measure to generate the learning efficiency of each subtask for the student agent and calculate the cumulative reward of the teacher agent.
[0123] The specific content of step 2.3 is: defining the learning efficiency, and using optimistic initialization to set the initial learning efficiency for each subtask in the subtask set; at the same time, maintaining a queue for each subtask to store the most recent K training scores of the subtask and their corresponding training start times.
[0124] The learning efficiency is defined as: for any subtask, the difference between the training score of the student agent for completing the subtask this time and the training score of the student agent for completing the subtask last time within a unit of time.
[0125] In this embodiment, the student agent should be trained on the task in which he has made the fastest progress. The teacher randomly selects a subtask from the subtask set. In order to avoid forgetting, the student agent should also be trained on tasks with a negative slope of the learning curve. That is, the student agent should prioritize training tasks with a larger absolute value of learning efficiency. The learning efficiency is defined as the cumulative reward change for the student agent's completion of the subtask this time relative to the last time it completed the same subtask per unit time. The learning efficiency is initialized optimistically, that is, the initial learning efficiency of each subtask is the maximum value of 1. The reason for this setting is that the teacher agent will randomly select the first subtask, and after the training of the subtask is completed, its corresponding learning efficiency will be less than the initial value, so that in the early training, the teacher will have a greater probability of evenly training each subtask. The reward function of the teacher agent is shown in the following formula:
[0126]
[0127] in represents the cumulative reward of the teacher agent at time step t; Represents the teacher agent's state The observed value of , which is the score obtained by the student agent when it starts training subtask h at time step t, that is, the cumulative reward; represents the score obtained by the student agent in the last training subtask h; t′ h represents the time when the student agent last trained subtask h; subtask h comes from the previous action of the teacher
[0128] In order to make the measurement of learning efficiency more stable, such as Figure 2As shown, this implementation adopts a fixed-length queue to optimize the calculation of learning efficiency. A queue is maintained for each subtask, and the queue is used as a FIFO (First Input First Output) buffer to store the most recent K scores of each subtask and its corresponding training start time. If the queue is full, that is, K training scores have been stored, the score and time that entered the queue earliest are removed to ensure that there are always only the most recent K data points in the queue. Linear regression is used to fit the relationship between the score and the time on all the data in the queue, and the absolute value of the slope of the fitted line is used as the reward.
[0129] For any action of the teacher agent The subtask selected by the teacher agent is assigned to the student agent for learning. The student agent generates a training score for the subtask and records the student agent's training score on the subtask and its corresponding training start time in the queue of the subtask. At the same time, the action value of the teacher agent is generated.
[0130] This implementation uses the Teacher-Student course learning framework, in which the teacher side, i.e., the input of the teacher agent, is the state And actions The action It means that at time step t, a task is selected from the subtask set according to the task importance to be given to the student agent for training, and the cumulative reward expectation of this task is output. The teacher side then judges the task importance based on the cumulative reward expectation. Since the iterative update of the teacher side is a multi-armed bandit problem, although the teacher can theoretically observe other aspects of the student state, such as network weights, optimizers, etc., the only effect on the training process is the cumulative reward. Therefore, this implementation method chooses to expose only the cumulative reward and does not involve the state The switch, therefore Can be ignored, thus simplifying the teacher end to:
[0131]
[0132] in Represents the action value, that is, the student agent is in The absolute value of the learning efficiency of training on the corresponding subtask. The larger the value, the higher the expected cumulative reward of the corresponding subtask output by the teacher, and the more important the subtask is. The teacher agent will tend to choose The student side learns the subtask itself according to the teacher's arrangement.
[0133] The training scores saved in the queue of each subtask are used to calculate the learning efficiency of the student agent on the subtask, and the cumulative reward of the teacher agent is calculated by fitting the training scores in the queue of each subtask based on the least squares method;
[0134] The cumulative reward of the teacher agent is expressed as:
[0135]
[0136] in represents the cumulative reward of the teacher agent at time step t; K is the queue capacity, which is a hyperparameter to be set; k represents the index value of the training score in the queue; t k Indicates the training start time corresponding to the kth training score; Represents the mean of the training start time corresponding to all training scores in the queue; Indicates that the student agent starts training the subtask at time step t The score obtained when , which is the kth training score in the queue.
[0137] In this implementation, for any time step t, the state of the teacher agent is Represents all the information of the student agent that the teacher agent contacts, such as network weights, optimizers, and the student agent's performance in subtasks The cumulative reward obtained after completing the training, etc., but only the cumulative reward is used for the training process, so this implementation only observes the cumulative reward
[0138] Step 2.4: Construct a task selector to update the action value of the teacher agent according to the accumulated rewards of the teacher agent, and the teacher agent selects subtasks for the student agent based on the action value of the teacher agent.
[0139] The specific content of step 2.4 is: using the accumulated rewards of the teacher agent to update the action value of the teacher agent, and adaptively adjusting the learning rate according to the updated action value of the teacher agent; establishing a Boltzmann distribution based on the action value of the teacher agent, and the teacher agent selects subtasks for the student agent from the subtask set according to the defined sampling probability, and arranges the selected subtasks to the student agent for multi-process parallel training.
[0140] The updating method of the action value of the teacher agent is expressed as:
[0141]
[0142] in represents the updated action value of the teacher agent; αtea Represents the learning rate.
[0143] The adaptive method of the learning rate is expressed as:
[0144]
[0145] Among them, They are adjustment factors and are hyperparameters to be set.
[0146] In this embodiment, the absolute value of learning efficiency is used as the action value Q in the teacher end. tea In order to make the training iteration more stable, this embodiment proposes adaptive adjustment of the learning rate. As the training progresses, a lower learning rate is adopted for subtasks with high absolute learning efficiency, and a higher learning rate is adopted for subtasks with low learning efficiency.
[0147] The sampling probability is expressed as:
[0148]
[0149] in represents the subtask selected by the teacher agent for the student agent at time step t, that is, the action of the teacher agent at time step t; Indicates that the teacher agent takes action The probability of; exp is the exponential function; τ is the temperature parameter to be set; N is the number of subtasks, in this embodiment, N = 4; n is the subtask index, h n Represents the subtask with index n in the subtask set.
[0150] The sampling method for subtasks provided by the present invention is based on Q tea The Boltzmann distribution is established to sample the subtasks, and multi-process parallel training is performed after sampling from the subtask set several times.
[0151] Step 3: Construct a curiosity module, including: a feature extractor, a forward dynamics model, and an inverse dynamics model; the curiosity module is used to generate intrinsic rewards for the student agent.
[0152] In this embodiment, the reward r received by the student agent during the training process is t Equal to the rewards obtained from the external environment With intrinsic rewards The sum of the intrinsic rewards of the student agent Can be generated by the curiosity module. The curiosity module is a system that helps generate curiosity rewards. It consists of three modules: feature extractor, forward dynamics model, and inverse dynamics model.
[0153] The specific content of step 3 is: when the subtask selected by the task selector is assigned to the student agent for learning, observe the state s of the student agent at time step t t and the state s at time step t+1 t+1 And input feature extractor Perform feature extraction in and get the state s t The corresponding eigenvector and status t+1 The corresponding eigenvector And input the inverse dynamics model to get the predicted action of the student agent Leveraging Predictive Actions The loss function of the inverse dynamics model is constructed based on the actual actions of the student agent, and the optimal inverse dynamics model is obtained by minimizing the loss function.
[0154] The loss function of the inverse dynamics model is:
[0155]
[0156] Where I represents the inverse dynamics model; θ I represents the model parameters of the inverse dynamics model; θ E represents the model parameters of the feature extractor; represents the loss function of the inverse dynamics model; a t is the action of the student agent at time step t; g(·) represents the learning function of the inverse dynamics model; H(·) represents the cross entropy loss function.
[0157] In this embodiment, the task selector selects a subtask, and for this subtask, obtains the state s of the student agent at time step t t and the state s at time step t+1 t+1 , since the curiosity module only wants to obtain state information related to the action performed by the agent and ignores irrelevant information. Therefore, self-supervision is used here to learn the feature representation of the current state and the next state, thereby removing state features that are irrelevant to the predicted action in the feature space. To achieve this self-supervision, the agent needs to train the feature extractor through the inverse dynamics model The agent's actions are predicted through the feature vectors of the current and future states, making them as close to the real actions as possible.
[0158] The student agent state s t The corresponding eigenvector and the student agent’s action a at time step t t Input the forward dynamics model for feature prediction and generate the feature vector of the predicted state at time step t+1 Using the feature vector of the predicted state and status t The corresponding eigenvector A loss function of the forward dynamics model is constructed, and the optimal forward dynamics model is obtained by minimizing the loss function.
[0159] The loss function of the forward dynamics model is expressed as:
[0160]
[0161] where F represents the forward dynamics model; θ F represents the model parameters of the forward dynamics model; represents the loss function of the forward dynamics model; f(·) represents the learning function of the forward dynamics; represents the L2 norm.
[0162] Use the feature vector of the student agent's predicted state at time step t+1 and the state s at time step t+1 t+1 Generate the curiosity-based intrinsic reward obtained by the student agent at time step t It is expressed as:
[0163]
[0164] Where η is the standard factor.
[0165] In this embodiment, the forward dynamics model predicts the feature vector of the next state based on the feature vector of the current state and the behavior of the agent. Finally, the prediction error in the forward dynamics model is provided to the student agent as an intrinsic reward to stimulate its curiosity. Specifically, the forward dynamics model and a t As input, it outputs the feature vector of the next predicted state, the feature vector of the predicted state and the eigenvector of the actual state, The bigger the difference, the bigger the reward, which encourages risk-taking. The intrinsic reward based on curiosity is the difference between the predicted feature vector of the next state and the actual feature vector of the next state.
[0166] Step 4: Based on the curriculum learning framework and curiosity module of curiosity and curriculum-based reinforcement learning, the A2C algorithm is used to train the student agent to obtain the trained student agent, and the trained student agent is used to perform the maze exploration task.
[0167] In this embodiment, if Figure 3As shown in the figure, the student side uses the A2C algorithm to train the student agent. The algorithm uses the famous AC architecture, combines a value network and a policy network, uses the policy network as an Actor to generate actions, and uses the value network as a Critic to estimate the state value function or state-action value function. Finally, the value network and the policy network are trained simultaneously through the policy gradient algorithm. In addition, the A2C algorithm replaces the Critic with its desired form, that is, the advantage function represented by the state value function, which can suppress the instability problem caused by large variance.
[0168] Step 4.1: The teacher agent randomly selects a subtask from the subtask set as the teacher agent’s action And assign this subtask to the student agent for learning.
[0169] Step 4.2: Set the network parameters of the policy network to ω π , the network parameter of the value network is ω v .
[0170] Step 4.3: Use the policy network to control the student agent to interact with the maze game environment. After completing the learning of the current subtask, the student agent generates a trajectory, which is a sequence of all states, actions, and rewards of the student agent in the process of completing the subtask.
[0171] The policy network is represented as a t ~π(·|s t ;ω v ), used to calculate the current state s of the student agent t To determine the action a that the student agent should take t .
[0172] The trajectory is recorded as: 1 ,a 1 ,r 1 ,s 2 ,a 2 ,r 2 ,……,s n ,a n ,r n}; where n is the total number of steps required for the student agent to complete the current subtask; s 1 、a 1 and r 1 are the state, action and reward of the student agent at time step 1; s 2 、a 2 and r 2 are the state, action and reward of the student agent at time step 2; s n 、an and r n are the state, action and reward of the student agent at time step n respectively.
[0173] Step 4.4: For each state of the student agent, the student agent observes the predicted trajectory for the next m steps starting from the current state, including the state estimate, action estimate, and reward estimate of the student agent in the next m steps.
[0174] For the current state s of the student agent t , will be changed from the current state s t The predicted trajectory of the next m steps is recorded as: t ,s t+1 ,a t+1 ,r t+1 ,……,s t+m-1 ,a t+m-1 ,r t+m-1 ,s t+m ,a t+m}, where m is the number of multi-step bootstrap steps to be set; r t is the current state t Corresponding rewards; t+1 ,a t+1 ,r t+1 are the state estimation, action estimation, and reward estimation for time step 1 starting from the current state; s t+m-1 ,a t+m-1 ,r t+m-1 are the state estimation, action estimation and reward estimation of time step m-1 starting from the current state; s t+m ,a t+m are the state estimation and action estimation for time step m starting from the current state, respectively.
[0175] In this embodiment, a policy network is used to make decisions. t ~π(·|s t ;ω v ) controls the agent to interact with the environment, completes a subtask, and obtains a complete trajectory. For all states s t , the observation trajectory needs to be observed, where t∈[1,2,3,…,nm].
[0176] Step 4.5: Use the value network to calculate all state-action pairs (s t ,a t )’s state action value.
[0177] Step 4.6: Based on the multi-step bootstrapping method, construct the objective function of the temporal difference TD target through the intercepted m-step discounted reward, and calculate the TD target of the student agent.
[0178] The objective function is expressed as:
[0179]
[0180] in is the TD target, which is used to represent the estimated value of the state action value; e represents the time step index; γ is the reward discount factor; γ e and γ m They represent γ to the eth power and γ to the mth power respectively; r t+e Represents the reward obtained by the student agent at time step t+e, and the reward value is the reward obtained by the student agent from the external environment and intrinsic rewards generated by the curiosity module The sum of V(s t+m ,a t+m ;ω v ) represents the state-action pair (s) of the student agent at time step t+m t+m ,a t+m ) is the state action value of the student agent in state s t+m Take action a t+m expectations of return.
[0181] In this embodiment, the value network in the original A2C model uses the reward r obtained at the current moment t and the next state estimate s t+1 Calculate the temporal difference target (TD target). When the network deviation is large, the target value deviation obtained by this method is also large, resulting in a slow convergence of the algorithm. This implementation replaces the single-step bootstrap in the temporal difference method of agent training with multi-step bootstrap. After the zeroth step TD(0) of the temporal difference, multi-step sampling and multi-step bootstrap are performed to compress the value distribution of the agent to the mth step, that is, the state estimate s after m action-state transitions. t+m , and construct the objective function by intercepting the m-step discounted reward, thereby improving the efficiency of the agent's learning.
[0182] Step 4.7: Calculate the TD error using the TD target of the student agent and the state-action value of each state-action pair, and construct the loss function.
[0183] The loss function is expressed as:
[0184]
[0185] Where L(ω v ) is the loss function; δ t represents the TD error at time step t; V(s t,a t ;ω v ) represents the state-action pair (s t ,a t )’s state action value; where t∈[1,2,3,…,nm].
[0186] Step 4.8: Use the loss function to perform gradient descent on the network parameters of the policy network and the network parameters of the value network respectively, and update the policy network and the value network.
[0187] The updated network parameters of the strategy network are expressed as:
[0188]
[0189] where ω π,next Represents the updated network parameters of the policy network; α π represents the learning rate of the policy network; represents the gradient of the policy network.
[0190] The updated network parameters of the value network are expressed as:
[0191]
[0192] where ω v,next Represents the updated network parameters of the value network; α v represents the learning rate of the value network; Represents the gradient of the value network.
[0193] Step 4.9: Use the learning efficiency measurer to generate the learning efficiency of the student agent on the current subtask and calculate the cumulative reward of the teacher agent.
[0194] Step 4.10: Based on the results of the learning efficiency meter, use the task selector to select the next subtask for the student agent to learn, and repeat steps 4.3-4.9 until the learning efficiency of the student agent reaches the set threshold or reaches the set maximum number of iterations. Stop training and obtain a trained student agent to perform the maze exploration task.
[0195] This implementation method builds a maze game environment and creates an intelligent agent based on the Unity3D game engine and its machine learning toolkit Unity ML-Agent, uses C# to write game logic scripts, and connects the Unity game environment with the algorithm provided by the present invention based on the distributed framework Ray.
[0196] This embodiment improves on the original A2C model and implements the code in Pycharm. The original algorithm converges slowly in the maze treasure hunt game and is difficult to effectively complete the original task, resulting in very low average rewards and final returns. The maze environment exploration method based on curiosity and curriculum reinforcement learning in this embodiment enables the agent to gradually improve the training efficiency after training on subtasks and successfully complete the original task within 5000 steps. At the same time, as shown in Table 2, under the influence of the curiosity mechanism, the agent continuously explores and learns, improving the average cumulative reward and final return obtained in the original task training.
[0197] Table 2 Comparison of indicators when the original algorithm and the method in the present invention are trained for 4000 rounds on the original task of maze treasure hunt
[0198]
[0199] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present invention.
Claims
1. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning, characterized in that: The method comprises the following steps: Step 1: Construct a maze game environment, including: several sets of walls, several mechanisms, a pyramid, and an agent for performing maze exploration tasks; Step 2: Initialize the environment parameters of the maze game environment, build and train a curriculum learning framework based on curiosity and curriculum reinforcement learning; the curriculum learning framework based on curiosity and curriculum reinforcement learning includes: a teacher agent and a student agent; Step 3: Construct a curiosity module, including: a feature extractor, a forward dynamics model, and an inverse dynamics model; the curiosity module is used to generate intrinsic rewards for the student agent; Step 4: Based on the curriculum learning framework and curiosity module of curiosity and curriculum-based reinforcement learning, the A2C algorithm is used to train the student agent to obtain the trained student agent, and the trained student agent is used to perform the maze exploration task.
2. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 1, characterized in that: The step 1 further comprises: Step 1.1: Construct a maze consisting of several groups of walls, randomly set several mechanisms and a pyramid made of multiple cubes at any position in the maze, and each pyramid has a target cube on the top; Step 1.2: Create an agent G for performing maze exploration tasks, and use the center of the maze as the starting position of agent G; Step 1.3: Construct a state space S containing the states of agent G at different time steps; The state is: For any time step t, the state s of agent G t Including: the speed of agent G, the position of agent G and the label information obtained by the sensor; the label information includes: walls, cubes forming the pyramid, target blocks, triggered mechanisms and untriggered switches; Step 1.4: Construct an action space A containing the actions of agent G at different time steps; The action is: For any time step t, the action a of agent G t is one of the following actions: stay still a1, move forward a2, move backward a3, turn left a4, or turn right a5; t Denoted as a t ={a l |l∈{1,2,3,4,5}}; where l is the index; Step 1.5: Construct the reward function of agent G.
3. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 2, characterized in that: The maze exploration task is as follows: agent G starts from the starting position and begins to look for a mechanism. When agent G finds a mechanism, it needs to touch the mechanism. At this time, a pyramid appears randomly at any position in the maze. Agent G continues to move in the maze to find the pyramid. When agent G finds the pyramid, it pushes it down to get the target block set on the top of the pyramid, thereby completing a maze exploration task.
4. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 3, characterized in that: The reward function of agent G in step 1.5 is: At any time step t, the reward R of agent G t It is expressed as: in represents the external reward of agent G at time step t; represents the curiosity-based intrinsic reward obtained by agent G at time step t.
5. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 4, characterized in that: The step 2.1 further comprises: Step 2.1: Initialize the environment parameters of the maze game environment, and discretize the initialized environment parameters to generate several subtasks and construct a subtask set; The specific content of step 2.1 is: taking the initialized environmental parameters as a subtask, the environmental parameters include: maze size p1, wall group number p2 and trap number p3; pre-defining several groups of different environmental parameters based on the initialized environmental parameters as several subtasks, and requiring that the maze size in all subtasks should be less than or equal to the maze size in the initialized environmental parameters; constructing a subtask set using all subtasks; wherein the subtask is recorded as: h = f(p j |j=1,2,3); where h represents the subtask; f represents the mapping function from the environment parameter to the subtask; p represents the environment parameter; p j represents the jth environmental parameter; Step 2.2: Use the agent G used to perform the maze exploration task as the student agent, create a teacher agent, and define the actions of the teacher agent The actions of the teacher agent To do this: In the subtask set, select a subtask for the student agent based on the task importance of each subtask; Step 2.3: Construct a learning efficiency measure to generate the learning efficiency of each subtask for the student agent and calculate the cumulative reward of the teacher agent; Step 2.4: Construct a task selector to update the action value of the teacher agent according to the accumulated rewards of the teacher agent, and the teacher agent selects subtasks for the student agent based on the action value of the teacher agent.
6. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 5, characterized in that: The specific contents of step 2.3 are as follows: defining the learning efficiency, and using optimistic initialization to set the initial learning efficiency for each subtask in the subtask set; at the same time, maintaining a queue for each subtask to store the most recent K training scores of the subtask and their corresponding training start time; The learning efficiency is: for any subtask, the difference between the training score of the student agent completing the subtask this time and the training score of the student agent completing the subtask last time in a unit of time; For any action of the teacher agent The subtask selected by the teacher agent is assigned to the student agent for learning. The student agent generates a training score for the subtask and records the student agent's training score on the subtask and its corresponding training start time in the queue of the subtask. At the same time, the action value of the teacher agent is generated. The training scores saved in the queue of each subtask are used to calculate the learning efficiency of the student agent on the subtask, and the cumulative reward of the teacher agent is calculated by fitting the training scores in the queue of each subtask based on the least squares method; The cumulative reward of the teacher agent is expressed as: in represents the cumulative reward of the teacher agent at time step t; K is the queue capacity, which is a hyperparameter to be set; k represents the index value of the training score in the queue; t k Indicates the training start time corresponding to the kth training score; Represents the mean of the training start time corresponding to all training scores in the queue; Indicates that the student agent starts training the subtask at time step t The training score obtained when .
7. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 6, characterized in that: The specific content of step 2.4 is: using the accumulated rewards of the teacher agent to update the action value of the teacher agent, and adaptively adjusting the learning rate according to the updated action value of the teacher agent; establishing a Boltzmann distribution based on the action value of the teacher agent, the teacher agent selects subtasks for the student agent from the subtask set according to the defined sampling probability, and arranges the selected subtasks to the student agent for multi-process parallel training; The updating method of the action value of the teacher agent is expressed as: in represents the updated action value of the teacher agent; α tea represents the learning rate; The adaptive method of the learning rate is expressed as: Among them, σ and θ are adjustment factors and are hyperparameters to be set; in Indicates that the teacher agent takes action The probability of; exp is the exponential function; τ is the temperature parameter to be set; N is the number of subtasks; n is the subtask index; h n Represents the subtask with index n in the subtask set.
8. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 7, characterized in that: The specific content of step 3 is: when the subtask selected by the task selector is assigned to the student agent for learning, observe the state s of the student agent at time step t t and the state s at time step t+1 t+1 And input feature extractor Perform feature extraction in and get the state s t The corresponding eigenvector and status t+1 The corresponding eigenvector And input the inverse dynamics model to get the predicted action of the student agent Leveraging Predictive Actions The loss function of the inverse dynamics model is constructed based on the actual actions of the student agent, and the optimal inverse dynamics model is obtained by minimizing the loss function. The loss function of the inverse dynamics model is: Where I represents the inverse dynamics model; θ I represents the model parameters of the inverse dynamics model; θ E represents the model parameters of the feature extractor; represents the loss function of the inverse dynamics model; a t is the action of the student agent at time step t; g(·) represents the learning function of the inverse dynamics model; H(·) represents the cross entropy loss function; The student agent state s t The corresponding eigenvector and the student agent’s action a at time step t t Input the forward dynamics model for feature prediction and generate the feature vector of the predicted state at time step t+1 Using the feature vector of the predicted state and status t The corresponding eigenvector Constructing a loss function of the forward dynamics model, and obtaining the optimal forward dynamics model by minimizing the loss function; The loss function of the forward dynamics model is expressed as: where F represents the forward dynamics model; θ F represents the model parameters of the forward dynamics model; represents the loss function of the forward dynamics model; f(·) represents the learning function of the forward dynamics; represents the L2 norm; Use the feature vector of the student agent's predicted state at time step t+1 and the state s at time step t+1 t+1 Generate the curiosity-based intrinsic reward obtained by the student agent at time step t It is expressed as: Where η is the standard factor.
9. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 8, characterized in that: The step 4 further comprises: Step 4.1: The teacher agent randomly selects a subtask from the subtask set as the teacher agent’s action And assign this subtask to the student agent for learning; Step 4.2: Set the network parameters of the policy network to ω π , the network parameter of the value network is ω v ; Step 4.3: Use the policy network to control the student agent to interact with the maze game environment. After completing the learning of the current subtask, the student agent generates a trajectory, which is a sequence of all states, actions, and rewards of the student agent in the process of completing the subtask; The policy network is represented as a t ~π(·|s t ;ω v ), which is used to calculate the current state s of the student agent. t To determine the action a that the student agent should take t ; The trajectory is recorded as: {s1, a1, r1, s2, a2, r2, ..., s n ,a n ,r n }; n is the total number of steps required for the student agent to complete the current subtask; s1, a1 and r1 are the state, action and reward of the student agent at time step 1; s2, a2 and r2 are the state, action and reward of the student agent at time step 2; s n 、a n and r n are the state, action and reward of the student agent at time step n respectively; Step 4.4: For each state of the student agent, the student agent observes the predicted trajectory of the next m steps starting from the current state, including the state estimate, action estimate, and reward estimate of the student agent in the next m steps; For the current state s of the student agent t , will be changed from the current state s t The predicted trajectory of the next m steps is recorded as: t ,s t+1 ,a t+1 ,r t+1 ,……,s t+m-1 ,a t+m-1 ,r t+m-1 ,s t+m ,a t+m }, where m is the number of multi-step bootstrap steps to be set; r t is the current state t Corresponding rewards; t+1 ,a t+1 ,r t+1 are the state estimation, action estimation, and reward estimation for time step 1 starting from the current state; s t+m-1 ,a t+m-1 ,r t+m-1 are the state estimation, action estimation and reward estimation of time step m-1 starting from the current state; s t+m ,a t+m They are the state estimation and action estimation of time step m starting from the current state respectively; Step 4.5: Use the value network to calculate all state-action pairs (s t ,a t )’s state action value; Step 4.6: Based on the multi-step bootstrapping method, construct the objective function of the temporal difference TD target through the intercepted m-step discounted reward, and calculate the TD target of the student agent; Step 4.7: Calculate the TD error using the TD target of the student agent and the state-action value of each state-action pair, and construct the loss function; Step 4.8: Use the loss function to perform gradient descent on the network parameters of the policy network and the network parameters of the value network, and update the policy network and the value network; Step 4.9: Use the learning efficiency measurer to generate the learning efficiency of the student agent on the current subtask and calculate the cumulative reward of the teacher agent; Step 4.10: Based on the results of the learning efficiency meter, use the task selector to select the next subtask for the student agent to learn, and repeat steps 4.3-4.9 until the learning efficiency of the student agent reaches the set threshold or reaches the set maximum number of iterations. Stop training and obtain a trained student agent to perform the maze exploration task.
10. A maze environment exploration method based on curiosity and curriculum-based reinforcement learning according to claim 9, characterized in that: The objective function is expressed as: in is the TD target, which is used to represent the estimated value of the state action value; e represents the time step index; γ is the reward discount factor; γ e and γ m They represent γ to the eth power and γ to the mth power respectively; r t+e Represents the reward obtained by the student agent at time step t+e, and the reward value is the reward obtained by the student agent from the external environment and intrinsic rewards generated by the curiosity module The sum of V(s t+m ,a t+m ;ω v ) represents the state-action pair (s) of the student agent at time step t+m t+m ,a t+m ) is the state action value of the student agent in state s t+m Take action a t+m Expected return; The loss function is expressed as: Where L(ω v ) is the loss function; δ t represents the TD error at time step t; V(s t ,a t ;ω v ) represents the state-action pair (s t ,a t )’s state action value; where t∈[1,2,3,…,nm]; The updated network parameters of the strategy network are expressed as: where ω π,next Represents the updated network parameters of the policy network; α π represents the learning rate of the policy network; represents the gradient of the policy network; The updated network parameters of the value network are expressed as: where ω v,next Represents the updated network parameters of the value network; α v represents the learning rate of the value network; Represents the gradient of the value network.
Citation Information
Cited By
Super-lens inverse design optimization method based on reinforcement learning
CN120597582A