A Traceable Reinforcement Learning Agent Training Method
The record-based reinforcement learning method addresses exploration limitations by archiving agent states and using state mapping and self-imitation learning to ensure comprehensive environmental coverage and effective path learning.
Patent Information
- Application Number
- CN202210096139.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-26
AI Technical Summary
The existing reinforcement learning methods have problems in the complex industrial field that cannot explore all spaces, are difficult to learn effective paths in the sparse reward environment, agents deviate from good state, and lack backtracking mechanisms.
The state of the agent is recorded by archive method, and the agent is guided back to the archived state through the target, and combined with the state mapping to cell, self-imitation learning and multi-strategy selection, optimize the exploration path.
It realizes the exploration of all spaces in complex environments and learns effective paths, enhances the learning ability in reward sparse scenarios, and improves the exploration efficiency and path learning effect of the agent.
Smart Images

Figure CN114511096B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence machine learning, and more specifically, to a method for training a backtrackable reinforcement learning agent. Background Art
[0002] As one of the current machine learning algorithms in artificial intelligence, reinforcement learning has a wide range of application fields, including robot control, video games, drones, etc. Current reinforcement learning methods all use agents to explore based on policies, learn and update algorithm models according to the explored data, and then continue to explore data in the direction of algorithm update, as Figure 1 shown; its deficiencies mainly include: 1. Since it will keep exploring forward along the direction of algorithm update, it is impossible to explore all spaces; 2. In an environment with sparse rewards, it is difficult for agents to learn effective paths; 3. If an agent gets a very high reward at the beginning, it is very likely to deviate further and further from good states; 4. In a deviated state, there is no effective mechanism to return to good states that have been explored before. The above reasons have led to the fact that reinforcement learning still cannot be well promoted and implemented in some complex industrial fields.
[0003] Based on the above, this application expects to propose a method for training a backtrackable reinforcement learning agent to improve the above deficiencies. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for training a backtrackable reinforcement learning agent, which first uses an archiving method to record all states reached by the agent, and guides the agent to be able to return to the states in the archive in a targeted manner; after the agent returns to any state, it starts exploring again, and theoretically can explore all spaces in the environment, having obvious advantages over other reinforcement learning algorithms.
[0005] Embodiments of the present invention are implemented through the following technical solutions:
[0006] This application provides a method for training a backtrackable reinforcement learning agent, including the following steps:
[0007] S1. Create a dictionary with a preset length, which is used to save the states of the agent, the Cells mapped by the states, actions, rewards, and done data;
[0008] S2. Conduct data exploration. First, select a Cell from the dictionary, and use the selected Cell as the target, return the selected target, use the target as a new starting point, select a new target for exploration, and record all states and actions encountered during the return stage and the exploration stage. Map all states to Cells, and update all states, all Cells, and actions to the dictionary;
[0009] S3. Obtain all Cell and behavior data collected in data exploration, learn based on the reinforcement learning algorithm, and update the relevant parameters of the learning algorithm.
[0010] Further, in S2, selecting a Cell from the dictionary and using the selected Cell as the target specifically means that each Cell mapped by a state has a corresponding weight. Based on the weight as the probability, sample the Cells in the dictionary, and the sampled Cell is used as the target returned by the agent.
[0011] In S2, the specific process of returning the selected target is as follows:
[0012] A. Select a behavior based on the behavior selection strategy in the current return stage.
[0013] B. Transmit the behavior and the state of the agent when executing this behavior to the running environment, and obtain the first new state, the first reward, and the first done data returned by the running environment.
[0014] C. Compare whether the first new Cell mapped by the first new state is equal to the target Cell. If they are equal, the agent has returned to the selected target state, and the return stage ends. Otherwise, record the obtained first new Cell and return to A to continue execution until the agent reaches the target Cell.
[0015] Further, in the stage of returning the selected target in S2, the target is divided into the final target and sub-targets. The final target is the selected target, and the sub-targets are the sub-states that guide the achievement of the final target. The sub-states are the states corresponding to any non-target Cell in the path to the final target.
[0016] Further, in S2, the specific process of selecting a new target for exploration is as follows: a. Select the exploration strategy of the agent. Use the threshold P as an adjustable parameter, P ∈ (0, 1), and set a random function. When the number obtained by the random function is greater than P, then use the strategy of selecting behaviors based on the PPO model in the exploration stage. Otherwise, adopt the strategy of randomly sampling behaviors from the action space.
[0017] b. Selection of the new target Cell. First, preferentially select the Cells in the unexplored area of the current state as the new target. Second, select the Cells in all areas as the new target. Finally, select the Cells based on the weight from the dictionary as the new target.
[0018] c. Select a behavior based on the strategy in a, and input the current state and the behavior into the running environment, and obtain the second new state, the second reward, and the second done data returned by the running environment.
[0019] d. Compare whether the second new Cell of the second new state mapping is equal to the new target Cell. If they are equal, the agent has returned to the selected new target state, and the return phase ends. Otherwise, record the obtained second new Cell and return to step c to continue execution until the agent reaches the new target Cell.
[0020] Further, in a complex operating environment, when an agent needs to perform multiple actions to obtain a reward, multiple states can be mapped to one Cell.
[0021] Further, step S3 also includes obtaining the data obtained by self-imitation learning and using it together with all the Cell and behavior data collected by exploration for the training of the reinforcement learning algorithm.
[0022] Further, before the self-imitation learning in step S3, it also includes selecting a better path for self-imitation learning. First, randomly select a target Cell based on the weight of the Cell as a probability, then obtain the complete path corresponding to the target Cell, and the paths corresponding to each non-target Cell in the complete path.
[0023] Further, the self-imitation learning in step S3 is specifically to perform imitation learning according to the complete path and the paths corresponding to the non-target Cells. First, select the mapping Cell of the penultimate state in the complete path as the final target, and use the mappings of all the states before the mapping Cell of the penultimate state as sub-targets; then use the paths corresponding to the non-target Cells as the training objects for training.
[0024] Further, the learning based on the reinforcement learning algorithm in step S3 is specifically to use the PPO model, A3C model or DQN model for training and learning.
[0025] Further, in one cycle, multiple agents can respectively perform the operations of selecting a Cell from the dictionary and using the selected Cell as the target, returning the selected target, using the target as a new starting point and selecting a new target for exploration, and performing the self-imitation learning at the same time; and record the new paths obtained by all agents and update the file.
[0026] The technical solution of the embodiment of the present invention has at least the following advantages and beneficial effects:
[0027] This application first uses an archiving method to record all the states that the agent has reached, guiding the agent to return to the states in the archive in a goal-oriented manner; after the agent returns to any state, it starts exploring again, and theoretically can explore all the spaces in the environment, which is an advantage over other reinforcement learning algorithms; at the same time, the state can be mapped to a low-dimensional Cell, that is, multiple states are mapped to one Cell, enabling the algorithm to learn effective paths even in scenarios with sparse rewards; additionally, self-imitation learning training is also adopted, and the dominant paths are selected according to the weight information in the dictionary for self-learning, enhancing the learning of the dominant paths. Brief Description of the Drawings
[0028] Figure 1 It is a schematic flow chart of an existing reinforcement learning method for exploration based on policies;
[0029] Figure 2 It is a schematic flow chart of the reinforcement learning method of the present invention;
[0030] Figure 3 It is a schematic flow chart of the process for the present invention to select the said goal of return;
[0031] Figure 4 It is a process case diagram of the present invention for splitting the goal and executing to obtain rewards;
[0032] Figure 5 It is a schematic flow chart of the exploration stage of the present invention;
[0033] Figure 6 It is a schematic flow chart of the specific method for the present invention to select a new goal Cell;
[0034] Figure 7 It is a schematic flow chart of the process for the present invention to update the dictionary;
[0035] Figure 8 It is a process case diagram of the self-imitation learning of the present invention. Detailed Implementation Modes
[0036] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.
[0037] From Figure 1It can be seen that existing reinforcement learning methods still have the following deficiencies: 1. Since they will keep exploring forward along the algorithm update direction, they cannot explore all spaces; 2. In an environment with sparse rewards, it is very difficult for the agent to learn an effective path; 3. If the agent receives a very high reward at the beginning, it is very likely to deviate further and further from the good state; 4. In the deviated state, there is no effective mechanism to return to the good state that has been explored before. To improve the above deficiencies, the present application proposes the following methods to improve the existing technical methods.
[0038] The present application provides a method for training a backtrackable reinforcement learning agent, as Figure 2 shown, including the following steps:
[0039] S1. Create a dictionary with a preset length, which is used to save the state of the agent, the Cell mapped by the state, the action, the reward, and the done data.
[0040] In this embodiment, the length of the dictionary is related to the data classification to be stored. For example, the preset length of the dictionary that only saves the state of the agent, the Cell mapped by the state, the action, and the reward is shorter than the length of the dictionary that saves the state of the agent, the Cell mapped by the state, the action, the reward, and the done data. It can be known that the preset length of the dictionary is determined according to the requirements of the specific running environment, and its adaptability is relatively wide. The above-mentioned Cell is the cell array data mapped by the state. When the running environment is relatively simple, the state can also not be converted into a Cell, but only the state is stored and used.
[0041] S2. Conduct data exploration. First, select a Cell from the dictionary, and use the selected Cell as the target, return the selected target, use the target as a new starting point, select a new target for exploration, and record all the states and actions encountered in the return stage and the exploration stage. Map all the states to Cells, and update all the states, all the Cells, and the actions to the dictionary. It should be noted that the exploration strategy in the exploration stage can adopt the PPO model, the A3C model, or the random sampling method.
[0042] In S2, selecting a Cell from the dictionary and using the selected Cell as the target specifically means that each Cell mapped by the state has a corresponding weight, and the Cells in the dictionary are sampled based on the weight as the probability, and the sampled Cell is used as the target for the agent to return.
[0043] A specific presentation of this embodiment is as follows. The state S of the agent at step t t , the state S tThe mapped Cell has a weight w. Based on this weight as a probability, sample the Cells in the dictionary, and the sampled Cell is the target returned by the agent. The setting of the weight includes the following strategies: (1) The highest score strategy, where the Cell with the highest score has a weight of 1 and the others are 0. (2) The number of actions inside the Cell strategy, that is, the reciprocal of the number of actions of a Cell is the strategy weight of the Cell. (3) The score decay strategy, that is, in the descending order of the scores of the Cells, the weights of the Cells decay. Of course, it is not limited to the above strategies. In our solution, it also supports combining multiple strategies and selecting according to different operating environments. It should be understood that the environments in the specific implementation manners of this application all refer to the game operating environments, and the application of this application is not limited to the game operating environments.
[0044] In S2, the selected target returned is as Figure 3 shown as follows:
[0045] A. Select an action based on the action selection strategy in the current return stage. The specific strategy here can be the PPO model, the A3C model, or the random strategy. It should be noted that the selection of the specific strategy has little impact on the technical effects that the application solution can bring. However, considering that in the game operating environment, the actions of people may be early or delayed, etc., and there is no such precise judgment by machines. In order to make the actions of the agent closer to the human situation, a certain probability can be set to continue to adopt the previous action when selecting an action. The optimal value of this probability is 0.2, so as to simulate the human operation situation.
[0046] B. Transmit the action and the state of the agent when executing this action to the operating environment, and obtain the first new state, the first reward, and the first done data returned by the operating environment.
[0047] C. Compare whether the first new Cell mapped by the first new state is equal to the target Cell. If they are equal, the agent has returned to the selected target state, and the return stage ends. Otherwise, record the obtained first new Cell and return to A to continue execution until the agent reaches the target Cell.
[0048] In the return stage, to enable the agent to better return to the selected target state, the target can be divided into a final target and sub-goals. The final target is the selected target, and the sub-goals are the sub-states that guide the achievement of the final target. The sub-states are the states corresponding to any non-target Cell in the path to the final target. That is, in the initial state, to reach the final target, there will be multiple sub-goals in the middle, which means that it is necessary to go through multiple sub-states to reach the final target. In this case, to make the sub-goals more robust, we set a path window with a default size of 10. As long as any Cell corresponding to a sub-state within the path window is reached, it is considered that the sub-goal is achieved. When the sub-goal is achieved, the window where the next position of the sub-goal in the path is located is selected as the new sub-goal. Specific examples are as follows Figure 4 as shown
[0049] In a complex operating environment, when the agent needs to perform multiple actions to obtain rewards, multiple states can be mapped to one Cell. This processing can also not be performed in some simple or non-dimension-reducing operating environments
[0050] The selection of a new target for exploration in S2 is as follows Figure 5 as shown, specifically including: a. Select the agent's exploration strategy, use the threshold P as an adjustable parameter, P ∈ (0, 1), and set a random function. When the number obtained by the random function is greater than P, the strategy of using the PPO model to select actions is used in the exploration stage; otherwise, the strategy of randomly sampling actions from the action space is adopted
[0051] b. Selection of the new target Cell. First, preferentially select the Cell in the unexplored area of the current state as the new target. Second, select the Cell in all areas as the new target. Finally, select the Cell based on weights from the dictionary as the new target. As shown Figure 6 where both P1 and P2 are adjustable parameters between 0 and 1
[0052] c. Select actions based on the strategy in a, input the current state and actions into the operating environment, and obtain the second new state, the second reward, and the second done data returned by the operating environment. Similarly, here it is still necessary to simulate the delay or advance of human actions, and the probability of sampling the previous action with an optimal value of 0.2 can also be set
[0053] d. Compare whether the second new Cell mapped by the second new state is equal to the new target Cell. If they are equal, the agent has returned to the selected new target state, and the return stage ends. Otherwise, record the obtained second new Cell and return to c to continue execution until the agent reaches the new target Cell
[0054] as Figure 7As shown, all states, all Cells, and actions are updated into the dictionary. Specifically, during the return phase and the exploration phase, path information of {{S i ,a i ,S i+1 ,R i},{S i+1 ,a i+1 ,S i+2 ,R i+1},…} is generated. This path information, along with the new Cells mapped by the states in the path, is updated into the dictionary. The update needs to follow the strategy to ensure that the new state and new Cells are updated into the dictionary. If the cell corresponding to the state already exists in the dictionary, it is recorded as the old Cell. Then, it is compared whether the score of the new Cell is higher than that of the old cell, or the path to the new Cell is shorter than that of the old cell. If so, it means the new Cell is better and needs to be updated into the dictionary.
[0055] S3. Obtain all Cell and behavior data collected during data exploration, or add the data obtained through self-imitation learning, and perform learning based on the reinforcement learning algorithm, and update the relevant parameters of the learning algorithm.
[0056] Before self-imitation learning in S3, it also includes selecting a better path for self-imitation learning. First, randomly select a target Cell based on the weight of the Cell as the probability, and then obtain the complete path corresponding to the target Cell, such as [[S0,S1,S2,…,S n-1 ,[a0,a1,a2,…,a n-1 , where S i represents the state of the agent at the i-th moment in this path, and a i represents the action taken by the agent at the i-th moment in this path; and the paths corresponding to each non-target Cell in the complete path, such as [[C0,C1,C2,…,C m-1 ,[b0,b1,b2,…,b m-1 , where C j represents the j-th cell corresponding to the state of the agent in this path, and b j represents the number of actions taken by the agent in C j .
[0057] The self-imitation learning in S3 is specifically to perform imitation learning according to the complete path and the paths corresponding to non-target Cells. First, select the second-to-last state S n-2Taking the mapped Cell as the final goal, and taking the mappings of all states before the mapped Cell in the penultimate state as sub-goals; then using the path corresponding to the non-goal Cell as the training object for training; where the path window size is set to 10, simulating the steps in the training operation environment, comparing whether there is a Cell within the path window size at the current position of the path corresponding to the non-goal Cell that reaches the sub-goal. If so, give a reward, otherwise continue; specific examples are as Figure 8 shown.
[0058] The specific learning based on the reinforcement learning algorithm in S3 is to use the PPO model, A3C model or DQN model for training and learning. In this embodiment, the PPO model is used, and its loss function is as follows:
[0059] L(θ) = L PG (θ) + ω VF L VF (θ) + ω ENT L ENT (θ) + ω L2 L L2 + ω SIL L SIL (θ)
[0060] where L PG (θ) is the policy gradient loss function based on the PPO model; L VF (θ) is the value loss function based on the PPO model; L ENT (θ) is the entropy loss function; L L2 is the L2-norm regularization function; L SIL (θ) is the loss function of self-imitation learning, that is, the loss function calculated from the data set generated during the self-imitation learning process, defined as:
[0061] L SIL (θ) = L SIL_PG (θ) + ω SIL_VF L SIL_VF (θ) + ω SIL_ENT L SIL_ENT (θ)
[0062] L SIL_PG (θ) = E s,a,R∈D [-logπ θ (a|s)·max(0, R - V θold (s))]
[0063] L SIL_VF (θ) = E s,a,r∈D [max(0, R - V θ (s))2]
[0064] LSIL_ENT \((\theta)\) is a regularization function of entropy.
[0065] In one cycle, multiple agents can respectively perform the operations of selecting a Cell from the dictionary and using the selected Cell as the target, returning the selected target, using the target as a new starting point and selecting a new target for exploration, and performing the self-imitation learning at the same time; and record the new paths obtained by all agents and update the archive.
[0066] Through the method of this application, all possible states that the agent may have reached can be recorded, and the agent can be guided in the form of a target to return to the state in the archive; after the agent returns to any state, exploration starts again, and theoretically all spaces in the environment can be explored, which is an advantage compared with other reinforcement learning algorithms; at the same time, the state can be mapped to a low-dimensional Cell, that is, multiple states are mapped to one Cell, so that the algorithm can also learn effective paths in scenarios with sparse rewards; in addition, self-imitation learning training is also adopted, and the dominant path is selected for self-learning according to the weight information in the dictionary, enhancing the learning of the dominant path.
[0067] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training a backtraceable reinforcement learning agent, characterized in that, It includes the following steps: S1. Create a dictionary of a preset length, which is used to store the state of the agent, the Cell mapped by the state, the action, the reward, and the done data; S2. Conduct data exploration. First, select a Cell from the dictionary and use the selected Cell as the target. Return the selected target. Use the target as the new starting point, select a new target for exploration, and record all the states and actions encountered during the return phase and the exploration phase. Map all the states to Cells, and update all the states, all the Cells, and the actions into the dictionary; Among them, selecting a Cell from the dictionary and using the selected Cell as the target specifically means: Each Cell mapped by a state has a corresponding weight. Sample the Cells in the dictionary based on the weight as the probability, and the sampled Cell is used as the target for the agent to return; The specific process of "return the selected target" is as follows: A. Select an action based on the action selection strategy in the current return phase; B. Pass the action and the state of the agent when executing this action to the running environment, and obtain the first new state, the first reward, and the first done data returned by the running environment; C. Compare whether the first new Cell mapped by the first new state is equal to the target Cell. If they are equal, the agent has returned to the selected target state, and the return phase ends. Otherwise, record the obtained first new Cell and return to A to continue execution until the agent reaches the target Cell; S3. Obtain all the Cell and action data collected during data exploration, learn based on the reinforcement learning algorithm, and update the relevant parameters of the learning algorithm.
2. The method for training a backtraceable reinforcement learning agent according to claim 1, wherein, In S2, when returning the selected target, divide the target into the final target and the sub-target. The final target is the selected target, and the sub-target is the sub-state that guides the achievement of the final target. The sub-state is the state corresponding to any non-target Cell in the path to the final target.
3. The trainable method of a backtrackable reinforcement learning agent according to claim 1, characterized in that, The specific process of "select a new target for exploration" in S2 is as follows: a. Select the agent exploration strategy. Use the threshold P as an adjustable parameter, where P ∈ (0, 1), and set a random function. When the number obtained by the random function is greater than P, use the strategy of selecting actions based on the PPO model during the exploration phase. Otherwise, adopt the strategy of randomly sampling actions from the action space; b. Selection of the new target Cell. First, preferentially select the Cells in the unexplored area of the current state as the new target. Secondly, select the Cells in all areas as the new target. Finally, select the Cells based on the weight from the dictionary as the new target; c. Select an action based on the strategy in a, and input the current state and the action into the running environment, and obtain the second new state, the second reward, and the second done data returned by the running environment; d. Compare whether the second new Cell of the second new state mapping is equal to the new target Cell. If they are equal, the agent has returned to the selected new target state, and the return phase ends. Otherwise, record the obtained second new Cell and return to c to continue execution until the agent reaches the new target Cell.
4. The trainable method of a backtrackable reinforcement learning agent according to claim 3, wherein In a complex operating environment, when an agent needs to perform multiple actions to obtain a reward, multiple states can be mapped to one Cell.
5. The method for training a backtrackable reinforcement learning agent according to claim 2, wherein S3 further includes obtaining data from self-imitation learning and using it together with all the Cell and behavior data collected through exploration for the training of the reinforcement learning algorithm.
6. The trainable method of the backtrackable reinforcement learning agent according to claim 5, wherein, Before self-imitation learning in S3, it also includes selecting a better path for self-imitation learning. First, randomly select a target Cell based on the weight of the Cell as the probability, then obtain the complete path corresponding to the target Cell, and the paths corresponding to each non-target Cell in the complete path.
7. The method for training a backtrackable reinforcement learning agent according to claim 6, wherein Self-imitation learning in S3 specifically means performing imitation learning according to the complete path and the paths corresponding to non-target Cells. First, select the mapping Cell of the second-to-last state in the complete path as the final target, and use the mappings of all states before the mapping Cell of the second-to-last state as sub-targets; then use the paths corresponding to the non-target Cells as the training objects for training.
8. The method for training a backtrackable reinforcement learning agent according to claim 7, wherein, Learning based on the reinforcement learning algorithm in S3 specifically means using the PPO model, A3C model, or DQN model for training and learning.
9. The method for training a traceable reinforcement learning agent according to any one of claims 5-8, characterized in that, In one cycle, multiple agents can respectively perform the operations of selecting a Cell from the dictionary and using the selected Cell as the target, returning the selected target, using the target as a new starting point and selecting a new target for exploration, and performing the self-imitation learning at the same time. And record the new paths obtained by all agents and update the archive.
Citation Information
Patent Citations
Experience playback sampling reinforcement learning method and system based on confidence upper bound thought
CN112734014A