Robot strategy model training methods, equipment and storage media

By acquiring an initial policy model through imitation learning and combining it with online reinforcement learning and human intervention information, the limitations of imitation learning and reinforcement learning are overcome, enabling efficient, safe, and autonomous exploration in robot training.

CN120297357BActive Publication Date: 2025-10-28BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510394937.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-10-28
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

Imitation learning in existing technologies relies heavily on expert demonstration data, making it difficult to surpass expert levels and exhibiting low robustness. In contrast, reinforcement learning is costly to train and carries the risk of destructive actions, making it difficult to meet the actual needs of robot training.

Method used

After acquiring an initial policy model through imitation learning based on demonstration data, online reinforcement learning is combined with the introduction of human intervention information. Through parameter updates in online and offline data pools, the robot's autonomous exploration capabilities and safety are improved.

Benefits of technology

It shortened the robot training time, improved the robustness and adaptability of the model, ensured the safety and reliability of the training process, and realized the transformation from demonstration data to autonomous exploration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297357B_ABST
    Figure CN120297357B_ABST
Patent Text Reader

Abstract

This application provides a method, device, and storage medium for training a robot's policy model. The method includes: in an offline learning phase, firstly, collecting expert demonstration data, allowing the robot to learn by imitation based on the expert demonstration data, obtaining an initial policy model and an offline data pool. In the online learning phase, the robot can perform reinforcement learning based on the expert policy, determining the training data for the current time step based on human intervention information, adding the training data to the online data pool, and updating the model parameters based on the offline and online data pools. After all tasks are completed, the obtained initial policy model is used as the target policy model. This application combines autonomous model exploration with human intervention in the robot's online reinforcement learning process, further improving the robustness of the robot's policy network and achieving an efficient transformation of the robot's policy from demonstration to autonomy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and more specifically, to a method, apparatus, and storage medium for training a strategy model for a robot. Background Art

[0002] Imitation learning and reinforcement learning are two main end-to-end learning paradigms for robot operation training, and they have significant advantages in complex operation scenarios such as robot grasping, assembly, and handling.

[0003] However, imitation learning relies heavily on expert demonstration data, making it difficult to surpass expert levels and handle unknown states, resulting in low robustness. Reinforcement learning, on the other hand, has high training costs and may produce destructive actions in actual robot operation scenarios, making the risks during the learning process difficult to control. Therefore, both imitation learning and reinforcement learning have certain limitations and are insufficient to meet the requirements of robot training. Summary of the Invention

[0004] The purpose of this application is to address the shortcomings of the prior art by providing a method, device, and storage medium for training a robot's strategy model, thereby solving the problem that imitation learning and reinforcement learning in the prior art have certain limitations and are difficult to meet the requirements of robot training.

[0005] To achieve the above objectives, the technical solution adopted in this application is as follows:

[0006] Firstly, this application provides a method for training a robot's policy model, the method comprising:

[0007] Imitation learning is performed based on demonstration data to obtain the robot's initial policy model and offline data pool;

[0008] During the current time step of executing the current task, the training data for the current time step is determined based on the human intervention information. The training data is added to the online data pool, and the parameters of the initial policy model are updated based on the offline data pool and the online data pool to perform online reinforcement learning on the robot's initial policy model. The human intervention information includes: the current action and the number of human interventions. The number of human interventions indicates the number of times the robot has been manually taken over. The current action is either a human-instructed intervention action or a target action generated by the initial policy model.

[0009] The initial policy model obtained after all tasks are completed is used as the target policy model to be used by the robot.

[0010] Optionally, if the current action is the intervention action, then the step of performing online reinforcement learning on the initial strategy model based on the human intervention information, determining the training data for the current time step, and adding the training data to the online data pool includes:

[0011] After the robot performs the intervention action, the number of manual interventions is incremented by 1, and the current state of the robot is obtained.

[0012] The current reward value is determined based on the number of manual interventions, the current state, and the current action.

[0013] The current action, the current state, and the current reward value are used as training data for the current time step, and the training data is added to the online data pool.

[0014] Optionally, determining the current reward value based on the number of manual interventions, the current state, and the current action includes:

[0015] The first reward value is determined based on the number of manual interventions.

[0016] Determine the second reward value based on the current state and the current action;

[0017] The current reward value is determined based on the first reward value and the second reward value.

[0018] Optionally, determining the first reward value based on the number of manual interventions includes:

[0019] If the number of manual interventions indicates that the current time step is the first consecutive manual intervention, then the first reward value is determined to be a preset negative reward value;

[0020] If the number of manual interventions indicates that no manual intervention has been performed at the current time step, then the first reward value is determined to be a preset positive reward value.

[0021] Optionally, if the current action is the target action generated by the initial policy model, then the step of performing online reinforcement learning on the initial policy model based on human intervention information, determining the training data for the current time step, and adding the training data to the online data pool includes:

[0022] After the robot performs the target action, the current state of the robot is obtained;

[0023] The current reward value is determined based on the number of manual interventions, the current state, and the current action.

[0024] The current action, the current state, and the current reward value are used as training data for the current time step, and the training data is added to the online data pool.

[0025] Optionally, updating the parameters of the initial strategy model based on the offline data pool and the online data pool includes:

[0026] If the amount of data in the online data pool is greater than or equal to the amount of data in the offline data pool, then the parameters of the initial strategy model are updated based on the data in the offline data pool and the data in the online data pool.

[0027] Determine whether the current task is completed. If so, set the number of manual interventions to the initial value and proceed to the next task or end all tasks. Otherwise, proceed to the next time step.

[0028] Optionally, updating the parameters of the initial strategy model based on the data in the offline data pool and the data in the online data pool includes:

[0029] Data is sampled from the offline data pool and the online data pool respectively to obtain an offline data set and an online data set;

[0030] The network parameters and value function parameters of the initial policy model are updated based on the offline data set and the online data set to obtain the updated policy model.

[0031] Optionally, the step of performing robot imitation learning based on demonstration data to obtain an initial policy model and an offline data pool includes:

[0032] The demonstration data is labeled with a reward value, and the demonstration data is added to the offline data pool;

[0033] The state data in the demonstration data is input into the robot's original strategy model, and the original strategy model generates the predicted action corresponding to the state data.

[0034] Determine the reward value of the predicted action, and add the state data, the predicted action, and the reward value to the offline data pool;

[0035] Based on the predicted action and the action data corresponding to the state data, a loss value is calculated, and the original policy model is iteratively updated based on the loss value. When a preset iteration termination condition is met, the updated original policy model is used as the initial policy model.

[0036] Secondly, this application provides a strategy model training device for a robot, the device comprising:

[0037] The offline learning module is used for imitation learning based on demonstration data to obtain the robot's initial policy model and offline data pool;

[0038] The online learning module is used to perform online reinforcement learning on the initial policy model during the current time step of executing the current task. It determines the training data for the current time step based on human intervention information, adds the training data to the online data pool, and updates the parameters of the initial policy model based on the offline data pool and the online data pool. The human intervention information includes the current action and the number of human interventions. The number of human interventions indicates the number of times the robot has been manually taken over. The current action is either a human-instructed intervention action or a target action generated by the initial policy model.

[0039] The generation module is used as the target policy model to be used after all tasks are completed, based on the initial policy model obtained.

[0040] Optionally, if the current action is the intervention action, then the online learning module is specifically used for:

[0041] After the robot performs the intervention action, the number of manual interventions is incremented by 1, and the current state of the robot is obtained.

[0042] The current reward value is determined based on the number of manual interventions, the current state, and the current action.

[0043] The current action, the current state, and the current reward value are used as training data for the current time step, and the training data is added to the online data pool.

[0044] Optionally, the online learning module is specifically used for:

[0045] The first reward value is determined based on the number of manual interventions.

[0046] Determine the second reward value based on the current state and the current action;

[0047] The current reward value is determined based on the first reward value and the second reward value.

[0048] Optionally, the online learning module is specifically used for:

[0049] If the number of manual interventions indicates that the current time step is the first consecutive manual intervention, then the first reward value is determined to be a preset negative reward value;

[0050] If the number of manual interventions indicates that no manual intervention has been performed at the current time step, then the first reward value is determined to be a preset positive reward value.

[0051] Optionally, if the current action is the target action generated by the initial policy model, then the online learning module is specifically used for:

[0052] After the robot performs the target action, the current state of the robot is obtained;

[0053] The current reward value is determined based on the number of manual interventions, the current state, and the current action.

[0054] The current action, the current state, and the current reward value are used as training data for the current time step, and the training data is added to the online data pool.

[0055] Optionally, the online learning module is specifically used for:

[0056] If the amount of data in the online data pool is greater than or equal to the amount of data in the offline data pool, then the parameters of the initial strategy model are updated based on the data in the offline data pool and the data in the online data pool.

[0057] Determine whether the current task is completed. If so, set the number of manual interventions to the initial value and proceed to the next task or end all tasks. Otherwise, proceed to the next time step.

[0058] Optionally, the online learning module is specifically used for:

[0059] Data is sampled from the offline data pool and the online data pool respectively to obtain an offline data set and an online data set;

[0060] The network parameters and value function parameters of the initial policy model are updated based on the offline data set and the online data set to obtain the updated policy model.

[0061] Optionally, the offline learning module is specifically used for:

[0062] The demonstration data is labeled with a reward value, and the demonstration data is added to the offline data pool;

[0063] The state data in the demonstration data is input into the original policy model, and the original policy model generates the predicted action corresponding to the state data.

[0064] Determine the reward value of the predicted action, and add the state data, the predicted action, and the reward value to the offline data pool;

[0065] Based on the predicted action and the action data corresponding to the state data, a loss value is calculated, and the original policy model is iteratively updated based on the loss value. When a preset iteration termination condition is met, the updated original policy model is used as the initial policy model.

[0066] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of a robot strategy model training method as described in any one of the first aspects.

[0067] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of a robot strategy model training method as described in any one of the first aspects.

[0068] The beneficial effects of this application are as follows: By using imitation learning, the robot acquires expert-level decision-making capabilities. Further online reinforcement learning shortens the robot's training time and enhances its adaptability to unexpected situations during autonomous exploration, thereby improving the model's robustness. Furthermore, the introduction of human intervention during the online learning phase allows for control over risky actions taken by the robot during training, improving the safety and reliability of the reinforcement learning process. In addition, determining the reward value in the training data based on the number of human interventions avoids frequent triggers of human intervention, thus mitigating learning difficulties. This approach ensures safety while reducing the model's over-reliance on human operation, achieving a transition from demonstration data to autonomous exploration.

[0069] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0070] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 A schematic diagram of the process of imitation learning (left) and reinforcement learning (right) in the prior art is shown;

[0072] Figure 2 A flowchart of a strategy model training method for a robot provided in an embodiment of this application is shown;

[0073] Figure 3 This application provides a flowchart of a method for obtaining training data according to an embodiment of the present application.

[0074] Figure 4 This application provides a flowchart for determining the current reward value according to an embodiment of the present application.

[0075] Figure 5 This application provides a flowchart for determining a first reward value according to an embodiment of the present application.

[0076] Figure 6 This invention provides a flowchart of another method for obtaining training data, as illustrated in an embodiment of this application.

[0077] Figure 7 This document illustrates a flowchart of a model parameter update method provided in an embodiment of this application.

[0078] Figure 8 This document illustrates another flowchart of a model parameter update method provided in an embodiment of this application.

[0079] Figure 9 This document illustrates a flowchart of an imitation learning process provided in an embodiment of this application.

[0080] Figure 10 The diagram shows an overall flowchart of a robot strategy model training method provided in an embodiment of this application;

[0081] Figure 11 A schematic diagram of the structure of a strategy model training device for a robot provided in an embodiment of this application is shown;

[0082] Figure 12 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0083] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0084] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0085] Imitation learning and reinforcement learning are two commonly used end-to-end learning paradigms in robot training scenarios. For example... Figure 1 The diagram shows flowcharts of imitation learning and reinforcement learning methods in existing technologies. The left side shows the flowchart for imitation learning, and the right side shows the flowchart for reinforcement learning.

[0086] In this process, imitation learning first obtains diverse and high-quality "state-action" data through remote operation demonstrations by human operators in real environments, and then fits the expert demonstration data through supervised learning to quickly obtain decision-making strategies that approximate the level of experts.

[0087] Reinforcement learning focuses on the robot's autonomous exploration capabilities within the "state-action-reward" framework. It records various interaction data through a continuously updated online experience replay pool, repeatedly extracts small batches of samples to update the Q-network using the Bellman equation, and then optimizes the robot's policy network based on the new Q-network, forming a self-improving closed-loop training process.

[0088] However, imitation learning is highly dependent on expert demonstration data, requiring a large amount of high-quality expert demonstration data. Once the quantity or quality of demonstration data is limited, model performance is often significantly constrained. Furthermore, the quality of imitation learning strategies is usually limited by the upper limit of expert level, preventing further breakthroughs. Especially in practical applications, when encountering environmental states or task requirements not present in the expert demonstration data, imitation learning strategies often lack adaptive adjustment capabilities and are insufficiently resistant to interference.

[0089] Reinforcement learning requires a large amount of interactive data to explore the environment and collect reward signals. Especially in high-dimensional continuous control scenarios such as robotic arms, the training process is time-consuming and may cause equipment wear and tear and high experimental costs. Furthermore, in actual robot operation, the "explore-exploit" strategy may produce destructive actions, and without robust safety mechanisms or degradation measures, there is a problem of uncontrollable risks.

[0090] In summary, existing methods for training robots based on imitation learning or reinforcement learning have certain limitations and cannot meet the actual needs of robot training.

[0091] Based on this, this application proposes a method for training a robot's policy model, comprising two stages: offline learning and online learning. In the offline learning stage, expert demonstration data is first collected, allowing the robot to learn by imitation based on this data, enabling it to achieve expert-level performance. In the online learning stage, the robot can perform reinforcement learning based on the expert policy, thereby improving training efficiency. Furthermore, during the robot's online reinforcement learning process, autonomous model exploration and human expert intervention are combined to further enhance the robustness of the robot's policy network and achieve an efficient transformation of the robot's policy from demonstration to autonomy.

[0092] First, the scenario for robot training in this application will be described. Robot training can be task-driven, with the training environment including multiple objects corresponding to the task, such as randomly placed cups or cabinets with open doors. The tasks the robot completes during training could be, for example, "moving a cup from position A to position B" or "closing the open door of cabinet C." After receiving a task, the robot can perform online reinforcement learning under task-driven conditions. When the robot training is complete, the final robot policy model can be used as a usable policy model.

[0093] Next, combine Figure 2 This paper describes the strategy model training method for the robot according to this application. The execution subject of this method can be a robot with data processing capabilities or an electronic device that is connected to the robot for communication, such as... Figure 2 As shown, the method includes:

[0094] S201. Based on the demonstration data, perform imitation learning to obtain the robot's initial policy model and offline data pool.

[0095] Among them, demonstration data can be expert demonstration data, that is, demonstration data of experts or high-level operators on a specific task. Demonstration data can include robot sensor data, action sequences, environmental conditions, etc.

[0096] After obtaining the demonstration data, it can be labeled to mark key actions and decision points. Simultaneously, the data should be preprocessed, such as through data cleaning and normalization, to ensure data quality and consistency.

[0097] Optionally, by training the imitation learning algorithm using preprocessed demonstration data, an initial policy model can be obtained. The initial policy model can output corresponding actions or decisions based on the input environmental state.

[0098] The initial strategy model can be deployed on the robot to predict its actions over multiple future time steps based on the input robot state data. The robot can be a humanoid robot, a robotic arm, or other similar device; this application does not impose any restrictions.

[0099] Optionally, the offline data pool includes demonstration data and data generated during training. The offline data pool can be a database, file system, etc. As one possible implementation, the offline data pool stores multiple data triples, each of which includes: state, action, and reward value. The state includes the environment state and the robot's state, the action can be an action performed by the robot, and the reward value can be the reward obtained after performing the action.

[0100] For example, after the imitation learning phase is completed, the reward corresponding to the last action of each trajectory can be marked as 1, and the rewards of all intermediate actions can be marked as 0, thereby obtaining the reward value corresponding to each action.

[0101] S202. During the current time step of the current task, the training data for the current time step is determined based on the information from human intervention. The training data is added to the online data pool, and the parameters of the initial policy model are updated based on the offline data pool and the online data pool to perform online reinforcement learning on the robot's initial policy model.

[0102] The human intervention information includes: the current action and the number of human interventions. The number of human interventions indicates the number of times the robot has been taken over by humans. The current action is either the intervention action instructed by humans or the target action generated by the initial strategy model.

[0103] Optionally, during online reinforcement learning, the robot can perform multiple online reinforcement learning tasks. For each current task, the robot can break down the current task into multiple time steps, and perform an action at each time step.

[0104] At the start of the current time step, historical state data and sensor data of the robot can be acquired and input into the initial policy model. The initial policy model predicts the target action. If the target action is risky, such as being destructive or potentially causing task failure, the human operator can intervene and modify the target action to the intervention action. After the robot executes the action of the current time step, the reward value and the updated state of the current time step can be obtained. The reward value, the current action, and the updated state of the current time step are used as training data.

[0105] Optionally, the human intervention information includes the current action and the number of human interventions. The current action refers to the intervention action instructed by the human or the target action generated by the initial policy model in the current time step. The number of human interventions records the number of times the robot has been taken over by a human, reflecting the frequency with which the model requires human intervention during task execution.

[0106] An online data pool can store training data generated during the online reinforcement learning phase, including interaction experiences at the current time step. This data can reflect the robot's performance and problems in actual operation. Initially, the online data pool can be empty.

[0107] Optionally, by combining the data obtained from the imitation learning phase in the offline data pool with the training data accumulated in real time in the online data pool, the parameters of the initial policy model can be comprehensively updated. As one implementation method, the model parameters can be adjusted through optimization algorithms such as gradient descent, enabling the model to fully utilize historical knowledge while adapting to the specific needs of the current task and environmental changes, thereby improving the accuracy and robustness of decision-making.

[0108] It's worth noting that after each task, the model's predictive ability can be evaluated based on the number of human interventions. If there are zero human interventions after a task, the model's predictive ability is good, achieving relatively accurate predictions without human intervention. If there are multiple human interventions after a task, the model's predictive ability needs improvement, and human intervention may be necessary during subsequent training.

[0109] S203. The initial policy model obtained after all tasks are completed shall be used as the target policy model to be used by the robot.

[0110] In the first implementation, before the online reinforcement learning stage, the initial policy model can be copied and saved. After completing the pre-set N tasks, that is, after reaching the preset number of training rounds, the parameters of the pre-saved initial policy model can be updated based on the final obtained initial policy model to obtain the target policy model.

[0111] In the second implementation, the number of tasks can be variable. It can be determined whether the currently executed task should be the last task based on the task completion rate of the current task. If the task completion rate remains stable in multiple iterations, or if the reward value obtained remains stable in multiple iterations and reaches the expected level, then the initial policy model obtained after the current task ends can be used as the target policy model.

[0112] In the third implementation, the convergence of the current initial policy model can also be judged. If the initial policy model has converged, the training can be terminated.

[0113] Optionally, after all tasks are completed, the number of human interventions can be reassessed. If the number of human interventions is 0, it means that the model can complete the task without human intervention. In this case, the initial strategy model can be used as the final target strategy model.

[0114] After obtaining the target policy model, the robot's historical data from at least one time step in the past can be used as input to the target policy model, which can then predict and output the action sequence for multiple future time steps.

[0115] In this embodiment, the robot first acquires expert-level decision-making capabilities through imitation learning. Based on this, online reinforcement learning is then performed, shortening the robot's training time and enhancing its adaptability to unexpected situations during autonomous exploration, thereby improving the model's robustness. Furthermore, human intervention is introduced during the online learning phase to control risky actions taken by the robot during training, improving the safety and reliability of the reinforcement learning process. In addition, determining the reward value in the training data based on the number of human interventions avoids frequent triggers of human intervention, thus mitigating learning difficulties. This ensures safety while reducing the model's over-reliance on human operation, achieving a transition from demonstration data to autonomous exploration.

[0116] Optionally, if the current action is an intervention action, then the initial policy model is subjected to online reinforcement learning based on the human intervention information to determine the training data for the current time step, and the training data is added to the online data pool, as follows: Figure 3 As shown, it includes:

[0117] S301. After the robot performs the intervention action, increment the number of human interventions by 1 and obtain the current state of the robot.

[0118] The robot's current state includes both the current environmental state and the robot's own state. The environmental state includes information about surrounding obstacles, task progress, and terrain, while the robot's own state includes its pose information.

[0119] Optionally, if human intervention is required during the robot's task execution, a human operator will intervene and perform an action, called an intervention action. Intervention actions are used to correct the robot's erroneous behavior, respond to unexpected situations, or complete complex tasks.

[0120] The number of human interventions is used to record the number of times the robot has been taken over by a human. Each time a human intervenes and performs an action, the counter increments by 1. Recording the number of human interventions helps assess the robot's autonomy and reliability during task execution.

[0121] S302. Determine the current reward value based on the number of manual interventions, the current status, and the current action.

[0122] Optionally, the current reward value can reflect the quality of the agent's performance of a specific action in a specific state, thereby guiding the agent to learn the optimal strategy.

[0123] It should be understood that if there are too many human interventions, it indicates that the robot is performing poorly in certain situations. In this case, the current reward value can be appropriately reduced to encourage the robot to reduce human intervention. If the current state indicates that the robot is in an unfavorable or dangerous situation, the current reward value can be appropriately reduced to encourage the robot to avoid such a state. If the current action is reasonable and effective, the current reward value can be appropriately increased to encourage the robot to continue performing such an action; conversely, if the current action is unreasonable or dangerous, the current reward value can be appropriately reduced.

[0124] S303. Use the current action, current state, and current reward value as training data for the current time step, and add the training data to the online data pool.

[0125] After obtaining the current reward value, the current action, current state, and current reward value can be stored as a data triple in the online data pool. The storage format can be (s... t ,a t ,r t ), where s t Indicates the current state, a t Indicates the current action, r t This represents the current reward value.

[0126] Furthermore, the process of determining the current reward value based on the number of manual interventions, the current state, and the current action, as described above, is as follows: Figure 4 As shown, it includes:

[0127] S401. Determine the first reward value based on the number of manual interventions.

[0128] S402. Determine the second reward value based on the current state and the current action.

[0129] S403. Determine the current reward value based on the first reward value and the second reward value.

[0130] As an optional implementation, corresponding evaluation functions can be set for the number of manual interventions, the current state, and the current action, respectively. The number of manual interventions is calculated using the intervention number evaluation function to obtain the first reward value, and the current state and current action are calculated using the state evaluation function and the action evaluation function to obtain the second reward value.

[0131] The weighting coefficients of the first reward value and the second reward value can be set based on actual needs. For example, if the impact of the number of manual interventions on the reward value is to be emphasized, the first reward value can be set to a larger weighting coefficient.

[0132] For example, the reward function R(s) t ,a t ,I) can be expressed as the following formula (1).

[0133] R(s t ,a t ,I)=w1·f(s t )+w2·g(a t )+w3·h(I) (1)

[0134] Where I represents the number of interventions, f(s) t ) is the state evaluation function, g(a) t ) is the action evaluation function, h(I) is the intervention number evaluation function, and w1, w2 and w3 are weight coefficients.

[0135] Optionally, the process of determining the first reward value based on the number of manual interventions described above, such as... Figure 5 As shown, it includes:

[0136] S501. If the number of manual interventions indicates that the current time step is the first consecutive manual intervention, then the first reward value is determined to be a preset negative reward value.

[0137] The first consecutive manual intervention can be determined based on the number of manual interventions in consecutive time steps. If I t =I t-1 +1, and I t-1 =I t-2 If so, the current time step can be determined as the first continuous manual intervention, and the first reward value can be set to a preset negative reward value, such as -1.

[0138] S502. If the number of manual interventions indicates that no manual intervention has been performed at the current time step, then the first reward value is determined to be a preset positive reward value.

[0139] If no manual intervention is performed at the current time step, the first reward value can be set to a preset positive reward value, such as 1.

[0140] It is worth noting that when judging the reward value of the last time step of the task, if the last time step is judged to be over by a human, the reward value can be a positive reward value of 1; if the last time step is judged to be over by the model, the reward value can be 0.

[0141] Optionally, if the current action is the target action generated by the initial policy model, then the initial policy model is subjected to online reinforcement learning based on human intervention information to determine the training data for the current time step, and the training data is added to the online data pool, as follows: Figure 6 As shown, it includes:

[0142] S601. After the robot performs the target action, obtain the robot's current state.

[0143] The robot's current state includes both the current environmental state and the robot's own state. The environmental state includes information about surrounding obstacles, task progress, and terrain, while the robot's own state includes its pose information.

[0144] Optionally, if no human intervention is required during the robot's task execution, the target action for the current time step can be generated by the initial strategy model, and the robot's current state can be obtained after the robot executes the target action.

[0145] S602. Determine the current reward value based on the number of manual interventions, the current status, and the current action.

[0146] S603. Use the current action, current state, and current reward value as training data for the current time step, and add the training data to the online data pool.

[0147] In the first implementation, the number of manual interventions in the previous time step can be used as the input parameter for calculation. The current reward value is determined based on the number of manual interventions, the current state, and the current action.

[0148] In the second implementation, if there is no human intervention at the current time step, the current reward value can be determined based on the current state and the current action.

[0149] The specific implementation process of steps S602-S603 can be referred to steps S302-S303 above, and will not be repeated here.

[0150] Optionally, the above process of updating the parameters of the initial strategy model based on the offline data pool and the online data pool, such as... Figure 7 As shown, it includes:

[0151] S701. If the amount of data in the online data pool is greater than or equal to the amount of data in the offline data pool, then the parameters of the initial strategy model are updated based on the data in the offline data pool and the data in the online data pool.

[0152] Optionally, the amount of data in the online data pool can be adjusted. Data volume of offline data pool If a comparison is made, Then determine whether the task is completed. If the task is not completed, proceed to the next time step. If the task is completed, the task environment can be initialized, and online reinforcement learning can be performed again based on the current task.

[0153] like The parameters of the initial strategy model can then be updated based on the data in the offline data pool and the data in the online data pool.

[0154] S702. Determine whether the current task is completed. If so, set the number of manual interventions to the initial value and proceed to the next task or end all tasks. Otherwise, proceed to the next time step.

[0155] Optionally, when each task is completed and the next task is about to begin, the number of manual interventions can be set to an initial value, such as 0.

[0156] If the current task is the last task, all tasks can be terminated. If the current task is not the last task, the next task can be started. If the current task is not completed, the next time step can be started, and step S202 above can be re-executed.

[0157] Optionally, the steps described above for updating the parameters of the initial strategy model based on data from the offline data pool and the online data pool to obtain the target strategy model are as follows: Figure 8 As shown, it includes:

[0158] S801. Data is sampled from the offline data pool and the online data pool respectively to obtain the offline data set and the online data set.

[0159] Optionally, a portion of the samples can be randomly selected from the offline data pool to obtain an offline dataset, and a portion of the samples can be randomly selected from the online data pool to obtain an online dataset.

[0160] The amount of data sampled in the offline data pool and the online data pool can be the same. For example, the same number of small batch samples can be sampled, including batch state S, batch action A, batch reward R, and batch next state S. ′ .

[0161] In this embodiment of the application, by independently extracting data samples from these two different data sources, historical data and real-time data can be fully utilized, enabling the model to learn from different data distributions and improving the model's generalization ability and adaptability.

[0162] S802. Update the network parameters and value function parameters of the initial policy model based on the offline and online datasets to obtain the updated policy model.

[0163] Optionally, offline and online datasets can be integrated to form a comprehensive dataset. A loss function can be used to calculate the loss value of the data in the comprehensive dataset, then gradient calculation can be performed, and the parameter values ​​of the model can be adjusted according to the calculated gradient to minimize the loss function.

[0164] Optionally, the parameters of the value function can be optimized based on the Bellman equation, as shown in equation (2) below.

[0165]

[0166] Where ||·|| is the root mean square error, α is a hyperparameter that can take the value 0.1, and γ is a hyperparameter that can take the value 0.99.

[0167] The following describes the steps for obtaining the initial policy model and offline data pool through imitation learning based on the demonstration data. In the first implementation method, imitation learning can be performed through behavior cloning, such as... Figure 9 As shown, the method includes:

[0168] S901. Label the demonstration data with reward values ​​and add the demonstration data to the offline data pool.

[0169] Optionally, the demonstration data can be obtained by splitting the trajectory sequence, and the reward value of the action at the last time step in the trajectory sequence can be marked as 1, and the reward value of the actions at the remaining time steps can be marked as 0.

[0170] S902. Input the state data from the demonstration data into the robot's original strategy model, and the original strategy model generates the predicted action corresponding to the state data.

[0171] The original policy model can be a supervised learning model deployed on the robot, such as a linear model, a multilayer perceptron, or other complex neural network architecture.

[0172] S903. Determine the reward value of the predicted action and add the state data, predicted action, and reward value to the offline data pool.

[0173] If the predicted action is the action of the last time step to complete the task, the reward value can be set to 1; otherwise, the reward value is set to 0.

[0174] S904. Calculate the loss value based on the predicted action and the action data corresponding to the state data, and iteratively update the original policy model based on the loss value. When the preset iteration termination condition is met, the updated original policy model is used as the initial policy model.

[0175] Among them, the motion data corresponding to the state data can be motion data demonstrated by experts.

[0176] Optionally, the mean squared error function can be used as the loss function to calculate the difference between the predicted action and the action data, and thus obtain the loss value.

[0177] The iteration termination condition can be either the convergence of the original policy model or the completion of a preset number of training iterations.

[0178] It should be noted that the above steps S901-S904 are one implementation method of imitation learning given in this application. It should be understood that the initial policy model and offline data pool can also be obtained through generative adversarial imitation learning. The specific method is not limited here. The labeling process of the reward value is the same as the above steps S901-S904.

[0179] Next, combine Figure 10 This paper describes the overall process of the strategy model training method for the robot in this application.

[0180] Reference Figure 10 After the offline learning phase, an initial strategy model and an offline data pool are built. Before the online learning phase begins, an online data pool can be built first, and the number of manual interventions can be initialized to 0.

[0181] Within each time step, the robot can make a judgment on the current action. If the current action is a human intervention action, the number of human interventions is incremented by 1. After the current action is executed, the current reward value of the current action is calculated and the training data is stored in the online data pool.

[0182] If the current action is predicted by the model, the number of manual interventions is not incremented by 1. After the current action is executed, the current reward value of the current action is calculated, and the training data is stored in the online data pool.

[0183] At each time step, the amount of data in the online data pool and the offline data pool is determined. If the amount of data in the online data pool is D... on The amount of data D that is greater than or equal to the offline data pool off If the data volume in the online data pool is less than that in the offline data pool, then a small batch of samples can be obtained from both the online and offline data pools, and the parameters of the initial strategy model can be updated based on the small batch of samples. If the data volume in the online data pool is less than that in the offline data pool, then the model parameters are not updated.

[0184] At the end of each time step, it is necessary to check whether the task is completed. If not, proceed to the next time step and regenerate the target action. If the task is completed, the number of manual interventions can be initialized to 0, and the next task can begin.

[0185] Once all tasks are completed, the number of human interventions can be assessed. If the number of human interventions is zero, the final initial policy model can be used as the target policy model. If the number of human interventions is not zero, training can continue until the model can complete the tasks autonomously without human intervention.

[0186] Based on the same inventive concept, this application also provides a robot strategy model training device corresponding to the robot strategy model training method. Since the principle of the device in this application is similar to the robot strategy model training method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0187] Figure 11 A schematic diagram of the structure of a strategy model training device for a robot provided in an embodiment of this application is shown.

[0188] The offline learning module 1101 is used for imitation learning based on demonstration data to obtain the robot's initial policy model and offline data pool.

[0189] The online learning module 1102 is used to determine the training data for the current time step based on human intervention information during the execution of the current task, add the training data to the online data pool, and update the parameters of the initial policy model based on the offline data pool and the online data pool to perform online reinforcement learning on the initial policy model. The human intervention information includes: the current action and the number of human interventions. The number of human interventions indicates the number of times the robot has been taken over by humans. The current action is the intervention action indicated by humans or the target action generated by the initial policy model.

[0190] The generation module 1103 is used as the initial policy model obtained after all tasks are completed as the target policy model to be used by the robot.

[0191] Optionally, the online learning module 1102 is specifically used for:

[0192] After the robot performs the intervention action, increment the number of human interventions by 1 and obtain the robot's current state;

[0193] The current reward value is determined based on the number of manual interventions, the current status, and the current action.

[0194] The current action, current state, and current reward value are used as training data for the current time step, and the training data is added to the online data pool.

[0195] Optionally, the online learning module 1102 is specifically used for:

[0196] The first reward value is determined based on the number of manual interventions.

[0197] Determine the second reward value based on the current state and the current action;

[0198] The current reward value is determined based on the first reward value and the second reward value.

[0199] Optionally, the online learning module 1102 is specifically used for:

[0200] If the number of manual interventions indicates that the current time step is the first consecutive manual intervention, then the first reward value is determined to be the preset negative reward value;

[0201] If the number of manual interventions indicates that no manual intervention has been performed at the current time step, then the first reward value is determined to be a preset positive reward value.

[0202] Optionally, the online learning module 1102 is specifically used for:

[0203] After the robot performs the target action, obtain the robot's current state;

[0204] The current reward value is determined based on the number of manual interventions, the current status, and the current action.

[0205] The current action, current state, and current reward value are used as training data for the current time step, and the training data is added to the online data pool.

[0206] Optionally, the online learning module 1102 is specifically used for:

[0207] If the amount of data in the online data pool is greater than or equal to the amount of data in the offline data pool, the parameters of the initial strategy model are updated based on the data in the offline data pool and the data in the online data pool.

[0208] Determine if the current task is completed. If so, set the number of manual interventions to the initial value and proceed to the next task or end all tasks. Otherwise, proceed to the next time step.

[0209] Optionally, the online learning module 1102 is specifically used for:

[0210] Data is sampled from the offline data pool and the online data pool respectively to obtain the offline data set and the online data set;

[0211] The network parameters and value function parameters of the initial policy model are updated based on the offline and online datasets to obtain the updated policy model.

[0212] Optionally, the offline learning module 1101 is specifically used for:

[0213] The demonstration data is labeled with reward values ​​and added to the offline data pool;

[0214] The state data from the demonstration data is input into the robot's original strategy model, and the original strategy model generates the predicted action corresponding to the state data.

[0215] Determine the reward value for the predicted action, and add the state data, predicted action, and reward value to the offline data pool;

[0216] Based on the predicted action and the corresponding action data of the state data, the loss value is calculated, and the original policy model is iteratively updated based on the loss value. When the preset iteration termination condition is met, the updated original policy model is used as the initial policy model.

[0217] Figure 12 This illustration shows a schematic diagram of an electronic device provided in an embodiment of this application, including: a processor 1201, a storage medium 1202, and a bus 1203. The storage medium 1202 stores machine-readable instructions executable by the processor 1201. When the electronic device runs a robot strategy model training method as described in the embodiment, the processor 1201 communicates with the storage medium 1202 via the bus 1203. The processor 1201 executes the machine-readable instructions. The preamble of the method item of the processor 1201 executes the steps in the robot strategy model training method described above.

[0218] This application also provides a computer-readable storage medium storing a computer program that is executed by a processor, which performs the steps in the above-described robot strategy model training method.

[0219] In this embodiment, the computer program, when run by the processor, can also execute other machine-readable instructions to perform other methods as described in the embodiments. For details on the specific execution steps and principles, please refer to the description of the embodiments, which will not be repeated here.

[0220] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0221] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0222] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0223] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0224] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0225] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A method for training a robot's strategy model, characterized in that, include: Imitation learning is performed based on demonstration data to obtain the robot's initial policy model and offline data pool; During the current time step of executing the current task, the training data for the current time step is determined based on the human intervention information. The training data is added to the online data pool, and the parameters of the initial policy model are updated based on the offline data pool and the online data pool to perform online reinforcement learning on the robot's initial policy model. The human intervention information includes: the current action and the number of human interventions. The number of human interventions indicates the number of times the robot has been taken over by a human. The current action is either a human-instructed intervention action or a target action generated by the initial policy model. The initial strategy model obtained after all tasks are completed is used as the target strategy model to be used by the robot. The step of determining the training data for the current time step based on human intervention information includes: After the robot performs the current action, obtain the robot's current state; The current reward value is determined based on the number of manual interventions, the current state, and the current action; the current action, the current state, and the current reward value are used as training data for the current time step.

2. The method according to claim 1, characterized in that, If the current action is the intervention action, after the robot performs the intervention action, the following steps are also included: Increment the number of manual interventions by 1.

3. The method according to claim 1, characterized in that, The process of determining the current reward value based on the number of manual interventions, the current state, and the current action includes: The first reward value is determined based on the number of manual interventions. Determine the second reward value based on the current state and the current action; The current reward value is determined based on the first reward value and the second reward value.

4. The method according to claim 3, characterized in that, The step of determining the first reward value based on the number of manual interventions includes: If the number of manual interventions indicates that the current time step is the first consecutive manual intervention, then the first reward value is determined to be a preset negative reward value; If the number of manual interventions indicates that no manual intervention has been performed at the current time step, then the first reward value is determined to be a preset positive reward value.

5. The method according to claim 1, characterized in that, After updating the parameters of the initial strategy model based on the offline data pool and the online data pool, the method further includes: Determine whether the current task is completed. If so, set the number of manual interventions to the initial value and proceed to the next task or end all tasks. Otherwise, proceed to the next time step.

6. The method according to claim 1, characterized in that, The step of updating the parameters of the initial strategy model based on the offline data pool and the online data pool includes: If the amount of data in the online data pool is greater than or equal to the amount of data in the offline data pool, then the parameters of the initial strategy model are updated based on the data in the offline data pool and the data in the online data pool.

7. The method according to claim 6, characterized in that, The step of updating the parameters of the initial strategy model based on the data in the offline data pool and the data in the online data pool includes: Data is sampled from the offline data pool and the online data pool respectively to obtain an offline data set and an online data set; The network parameters and value function parameters of the initial policy model are updated based on the offline data set and the online data set to obtain the updated policy model.

8. The method according to claim 1, characterized in that, The imitation learning based on demonstration data, which yields the robot's initial policy model and offline data pool, includes: The demonstration data is labeled with a reward value, and the demonstration data is added to the offline data pool; The state data in the demonstration data is input into the robot's original strategy model, and the original strategy model generates the predicted action corresponding to the state data. Determine the reward value of the predicted action, and add the state data, the predicted action, and the reward value to the offline data pool; Based on the predicted action and the action data corresponding to the state data, a loss value is calculated, and the original policy model is iteratively updated based on the loss value. When a preset iteration termination condition is met, the updated original policy model is used as the initial policy model.

9. An electronic device, characterized in that, include: The device includes a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of a strategy model training method for a robot as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the strategy model training method for a robot as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-level human intelligence enhanced automatic driving vehicle decision control method and system

    CN117227758A

  • Mechanical arm control method based on simulation and variable parameter two-stage reinforcement learning

    CN118357922A