Reinforcement learning policy training method and system

CN121912377BActive Publication Date: 2026-09-08BEIJING HUMANOID ROBOTICS INNOVATION CENTER CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610043146.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-09-08
Estimated Expiration
2046-01-13

AI Technical Summary

Technical Problem

[0005]本申请的目的在于,针对上述现有技术中的不足,提供一种强化学习策略训练方法以及系统,以解决现有技术中训练得到的强化学习策略性能较差、训练过程更新不稳定且收敛速度较慢的问题

Benefits of technology

[0008] The beneficial effects of this application are as follows: By setting up asynchronous action and learning ends in the reinforcement learning policy training system, and deploying a first policy network in the action end and a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network in the learning end, the policy is executed through the action end, triggering human intervention and collecting real interaction data. The learnable state network, intervention experience pool, online experience pool, value network, and expert network are updated through the learning end. This achieves a closed-loop training structure of real-machine online reinforcement learning and human intervention constraints. It can model the update of the policy network as a state-dependent constrained reinforcement learning problem and learn from human interventions that may not be optimal and contain noise. During the learning process, the constraint strength is adaptively adjusted through the learnable state network. This allows for the relaxation of constraints in highly uncertain states where human intervention fluctuates greatly and may be suboptimal, avoiding suboptimal interventions that limit the policy upper limit and encouraging the policy to rely more on the reinforcement learning objective for autonomous improvement. In states where human intervention is consistent and reliable with low uncertainty, the constraints are tightened, encouraging the policy to conform to human guidance, thereby improving sample efficiency, accelerating convergence speed, and achieving adaptive compatibility and robust learning with suboptimal human intervention data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121912377B_ABST
    Figure CN121912377B_ABST
Patent Text Reader

Abstract

The application provides a reinforcement learning strategy training method and system, wherein the method comprises: generating a plurality of action experience data; updating an intervention experience pool and an online experience pool according to each action experience data; updating network parameters of a learnable state network according to the intervention experience pool and the online experience pool, and updating network parameters of a second strategy network based on the updated learnable state network, a value network and an expert network; obtaining the network parameters of the second strategy network in a preset interaction cycle, and updating network parameters of a first strategy network according to the network parameters of the second strategy network; taking the second strategy network at the time of stopping updating as a decision network of a mechanical arm, and deploying the decision network to the mechanical arm for action control. The application can improve sample efficiency, accelerate convergence speed, and realize adaptive compatibility and robust learning of suboptimal human intervention data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of robotic arm control technology, and more specifically, to a reinforcement learning strategy training method and system. Background Technology

[0002] In vision-based real-world robotic arm operations, online reinforcement learning is typically used as the core framework, with human intervention introduced to improve the safety and efficiency of real-machine training while adapting to the needs of real-machine training.

[0003] In existing technologies, real-machine reinforcement learning training methods based on human intervention typically include: reinforcement learning training methods that directly utilize correct human intervention actions of robotic arms, and reinforcement learning training methods that treat human intervention as preferences or feedback signals.

[0004] However, existing real-machine reinforcement learning training methods based on human intervention can only learn from correct or optimal robotic arm intervention actions. When human intervention is biased, hesitant, or erroneous, the reinforcement learning strategy trained has poor performance, unstable updates during training, and slow convergence speed. Summary of the Invention

[0005] The purpose of this application is to provide a reinforcement learning policy training method and system to address the shortcomings of the prior art, thereby solving the problems of poor performance, unstable updates during training, and slow convergence speed of the reinforcement learning policy trained in the prior art.

[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, one embodiment of this application provides a reinforcement learning policy training method applied to a reinforcement learning policy training system. The system includes: an action end and a learning end. The action end is deployed with a first policy network, and the learning end is deployed with a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network. The first policy network and the second policy network have the same network structure. The method includes: The action terminal makes action decisions and executes actions based on the first policy network, generating multiple action experience data; The learning terminal acquires multiple action experience data from the action terminal in real time, and updates the intervention experience pool and the online experience pool based on each action experience data, or updates the online experience pool. The learning end updates the network parameters of the learnable state network based on the intervention experience pool and the online experience pool, and updates the network parameters of the second policy network based on the updated learnable state network, the value network and the expert network. The action terminal acquires the network parameters of the second strategy network according to a preset interaction cycle, and updates the network parameters of the first strategy network according to the network parameters of the second strategy network. The second strategy network at the point of no update is used as the decision network for the robotic arm, and the decision network is deployed to the robotic arm for motion control.

[0007] Secondly, another embodiment of this application provides a reinforcement learning policy training system, the system comprising: an action end and a learning end, the action end being deployed with a first policy network, and the learning end being deployed with a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network, wherein the first policy network and the second policy network have the same network structure.

[0008] The beneficial effects of this application are as follows: By setting up asynchronous action and learning ends in the reinforcement learning policy training system, and deploying a first policy network in the action end and a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network in the learning end, the policy is executed through the action end, triggering human intervention and collecting real interaction data. The learnable state network, intervention experience pool, online experience pool, value network, and expert network are updated through the learning end. This achieves a closed-loop training structure of real-machine online reinforcement learning and human intervention constraints. It can model the update of the policy network as a state-dependent constrained reinforcement learning problem and learn from human interventions that may not be optimal and contain noise. During the learning process, the constraint strength is adaptively adjusted through the learnable state network. This allows for the relaxation of constraints in highly uncertain states where human intervention fluctuates greatly and may be suboptimal, avoiding suboptimal interventions that limit the policy upper limit and encouraging the policy to rely more on the reinforcement learning objective for autonomous improvement. In states where human intervention is consistent and reliable with low uncertainty, the constraints are tightened, encouraging the policy to conform to human guidance, thereby improving sample efficiency, accelerating convergence speed, and achieving adaptive compatibility and robust learning with suboptimal human intervention data.

[0009] Furthermore, the training process can appropriately change the task objectives, provide states in various scenarios, actively construct and introduce diverse training data, and force the policy model to learn to extract more generalized state representations and decision-making logic, thereby systematically enhancing its adaptability and robustness in unknown scenarios.

[0010] Simultaneously, during training, it not only supports real-time human intervention in data processing but also possesses a certain degree of online autonomous exploration capability. When data is limited or of insufficient quality, it can effectively compensate for performance degradation caused by insufficient data. Furthermore, it reduces time consumption, equipment wear and tear, and security risks associated with frequent trial and error in real-world environments. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A schematic diagram of a reinforcement learning strategy training system provided in an embodiment of this application; Figure 2 A flowchart illustrating a reinforcement learning strategy training method provided in an embodiment of this application; Figure 3 Another flowchart illustrating the reinforcement learning strategy training method provided in this application embodiment; Figure 4 This is a flowchart illustrating the process of updating network parameters of a learnable state network in the reinforcement learning strategy training method provided in this embodiment of the application. Figure 5 This is another flowchart illustrating the process of updating the network parameters of the learnable state network in the reinforcement learning strategy training method provided in this application embodiment. Figure 6 This is a flowchart illustrating the process of updating the network parameters of the second policy network in the reinforcement learning policy training method provided in this application embodiment. Figure 7 This is another flowchart illustrating the process of updating the network parameters of the second policy network in the reinforcement learning policy training method provided in this application embodiment. Figure 8 A flowchart illustrating the generation of multiple action experience data in the reinforcement learning strategy training method provided in this application embodiment; Figure 9 This is another flowchart illustrating the process of updating the network parameters of the second policy network in the reinforcement learning policy training method provided in this application embodiment. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0014] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0015] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0016] In existing technologies, real-machine reinforcement learning training methods based on human intervention typically include: reinforcement learning training methods that directly utilize correct human intervention actions of robotic arms, and reinforcement learning training methods that treat human intervention as preferences or feedback signals.

[0017] However, existing real-machine reinforcement learning training methods based on human intervention can only learn from correct or optimal robotic arm intervention actions. When human intervention is biased, hesitant, or erroneous, the reinforcement learning strategy trained has poor performance, unstable updates during training, and slow convergence speed.

[0018] Based on the aforementioned problems, this application proposes a reinforcement learning policy training method applied to a reinforcement learning policy training system. The system establishes asynchronous action and learning ends. A first policy network is deployed on the action end, while a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network are deployed on the learning end. The action end executes the policy, triggers human intervention, and collects real-world interaction data. The learning end updates the learnable state network, intervention experience pool, online experience pool, value network, and expert network, achieving a closed-loop training structure of real-machine online reinforcement learning and human intervention constraints. This allows learning from potentially suboptimal and noisy human interventions, accelerating convergence and enhancing robustness to human interference. Furthermore, it combines offline learning with online autonomous exploration.

[0019] First, the reinforcement learning strategy training system provided in the embodiments of this application will be described in detail.

[0020] Figure 1 A schematic diagram of a reinforcement learning policy training system provided in an embodiment of this application is shown below. Figure 1 As shown, the reinforcement learning policy training system provided in this application embodiment includes: an actor end (Actor end) and a learner end (Learner end), wherein the actor end is deployed with a first policy network. The learning end is equipped with a second-strategy network. Learnable state networks Intervention experience pool, online experience pool, value network and expert network First Strategy Network With the second strategy network They have the same network structure.

[0021] It is worth noting that the motion end can be deployed on the control side of the current robotic arm, for example, in the control system of the current robotic arm, while the learning end can be deployed in electronic devices, such as in a cloud server or a local server. The motion end controls the current robotic arm to perform actions and generates motion experience data. Through the interaction between the motion end and the learning end, the reinforcement learning strategy is trained on the learning end. After training, the decision network of the current robotic arm is obtained, which can then be applied to the current robotic arm to perform tasks, or applied to other robotic arms with the same configuration to perform tasks, thereby achieving control of the robotic arm.

[0022] In this embodiment, the training process of the reinforcement learning strategy is illustrated using a robotic arm as an example. In addition to a robotic arm, the application object can also be other mechanical devices with autonomous operation capabilities, such as robots.

[0023] Specifically, the learning mechanism includes a reinforcement learning objective link and an intervention constraint objective link. The reinforcement learning objective link uses real-world interaction data for value assessment and policy iteration. The intervention constraint objective link provides usable, but potentially suboptimal, human intervention information to the policy learning process.

[0024] It is understandable that human intervention or human intervention data refers to the data generated when a human operator temporarily takes over control of a robotic arm and personally operates the robotic arm to complete some of its actions during the autonomous execution of a task, when the arm's behavior deviates from expectations or faces high risks.

[0025] Human intervention includes ideal human intervention and suboptimal human intervention. Ideal human intervention is considered the optimal solution. Suboptimal human intervention, or suboptimal human intervention data, refers to intervention data that, although derived from human operator actions, is of low quality or even erroneous due to human factors. Suboptimal human intervention is not the globally optimal solution and may be of poor quality or even incorrect due to factors such as action deviations, inappropriate timing, decision-making errors, or inconsistent styles.

[0026] The reinforcement learning objective link includes a second policy network, a learnable state network, an online experience pool, and a value network. The intervention constraint objective link includes an expert network and an intervention experience pool.

[0027] The reinforcement learning objective link and the intervention constraint objective link work together in the second policy network to realize a closed-loop training structure of real-machine online reinforcement learning and human intervention constraints.

[0028] Among them, the first strategy network and the second strategy network are used to generate the execution strategy of the robotic arm.

[0029] Among them, the learnable state network can be a Lagrange multiplier network. For state-based Learnable Lagrange multiplier networks, specifically, refer to learnable Lagrange multiplier networks that adaptively adjust the trade-off between reinforcement learning objectives and human intervention / imitation objectives in the state space. Their parameters are updated based on the degree of constraint violation using the dual gradient ascent method.

[0030] The intervention experience pool stores human action data. Human intervention action sequences refer to the interaction trajectories generated when a human operator takes over or corrects the robotic arm during its operation. These sequences can be obtained at the moment of human intervention before the system detects policy deviations, increased risk, or task failure. The intervention experience pool is used to train and update the expert network.

[0031] The online experience pool stores complete interaction trajectories from the robot's autonomous exploration process, including states, actions, rewards, and next states. It is also used for calculating reinforcement learning objectives.

[0032] Among them, the value network serves as a value assessment component, used to estimate the long-term reward of the current state or state-action pair.

[0033] In this context, the expert network is used to represent human behavior patterns. The goal of the expert network is to fit the distribution of human intervention behavior, that is, to predict the actions that a human operator may take under a given state. It is obtained by fitting human offline teaching and online intervention data. It does not assume that human actions are optimal, but exists as a reference strategy to model human intervention strategies and provide constraint signals or reference signals for imitation learning.

[0034] The reinforcement learning strategy training method provided in this application will be described in detail below with reference to several embodiments.

[0035] Figure 2 A flowchart illustrating a reinforcement learning strategy training method provided in this application embodiment is shown below. Figure 2 As shown, the execution entity of this method can be the aforementioned reinforcement learning policy training system, and the method includes: S201. The action end makes action decisions and executes actions based on the first policy network, generating multiple action experience data.

[0036] It is understandable that the motion end is deployed on the control side of the current robotic arm.

[0037] Optionally, by time For example, the action end can obtain the current environment image and the current pose of the robot arm through the control side of the robot arm. and compare the current environment image with the current pose. Combined into the current state and through the first policy network Regarding the current state Make motion decisions and generate the current motion of the robotic arm. .in, .

[0038] Among them, the current pose Let be the 3D position and Euler angles relative to the base coordinate system of the robotic arm.

[0039] Optionally, the current action is executed, action experience data for the current action is generated, and the action is repeated to generate multiple action experience data.

[0040] For example, the current action Action experience data can be ( , , , , ).in, This is the current state. For the current action, This is the status signal for the current action. This serves as a reward signal for the current action. This is the task signal for the current task in which the current action takes place.

[0041] in, The current task then ends. The task is not finished yet; when the task ends, When the task is not finished, . The current action is a human intervention action. The current action is not a human intervention action.

[0042] Specifically, the first policy network can infer from the environmental image and the current pose to generate the current action, and then generate action experience data for the current action based on the current action.

[0043] The network parameters of the first strategy network can be the same as those of the second strategy network, or the network parameters of the first strategy network can be different from those of the second strategy network.

[0044] For example, in the initial stage of reinforcement learning training, the network parameters of the first policy network are the same as those of the second policy network. In the middle stage of reinforcement learning training, the network parameters of the first policy network and the second policy network may be the same or different.

[0045] In one example, after the first policy network generates the current action, it executes the current action through a robotic arm, obtains the execution result of the current action, and determines the task signal of the current task, the reward signal of the current action, and the state signal of the current action based on the execution result of the current action, thereby generating the action experience data of the current action.

[0046] In another example, after the first policy network generates the current action, the reinforcement learning policy training system judges the current action through preset boundary conditions to determine whether the current action needs to be replaced or adjusted to obtain the target action. The target action is then executed by a robotic arm, and the execution result of the target action is obtained. Based on the execution result of the target action, the task signal of the current task, the reward signal of the target action, and the state signal of the target action are determined, thereby generating motion experience data of the target action. The motion experience data of the target action is then used as the motion experience data of the current action.

[0047] S202. The learning end acquires multiple action experience data from the action end in real time, and updates the intervention experience pool and the online experience pool based on each action experience data, or updates the online experience pool.

[0048] Optionally, the learning end acquires multiple action experience data from the action end in real time, and makes judgments based on each action experience data, thereby updating the intervention experience pool and the online experience pool, or updating the online experience pool.

[0049] The online experience pool stores action experience data from real interactions, while the intervention experience pool stores human intervention data from the action experience data of real interactions.

[0050] For example, multiple action experience data can be updated to an online experience pool, and the multiple action experience data can be judged to find human intervention actions in the multiple action experience data, and the human intervention actions can be updated to the intervention experience pool.

[0051] S203. The learning end updates the network parameters of the learnable state network based on the intervention experience pool and the online experience pool, and updates the network parameters of the second policy network based on the updated learnable state network, value network and expert network.

[0052] It is understandable that after the intervention experience pool and the online experience pool are updated, the learning end can sample data from the intervention experience pool and the online experience pool to perform multi-objective joint optimization, thereby updating the network parameters of the second policy network.

[0053] In the process of updating the network parameters of the second policy network, the action experience data in the online experience pool can be used as input data for the reinforcement learning objective link to maximize long-term returns.

[0054] In the process of updating the network parameters of the second policy network, human intervention data in the intervention experience pool serves as input data for the intervention constraint target link, acting as a soft constraint for the second policy network's learning. The human intervention data in the intervention experience pool can be considered a reference signal, which carries uncertainty, may be suboptimal, but is still valuable.

[0055] Optionally, during the update of the network parameters of the second policy network, the learnable state network can be updated after the human intervention data in the intervention experience pool is explicitly modeled by the expert network. The weights between the reinforcement learning objective and the intervention constraint objective can be dynamically adjusted through the updated learnable state network. Based on the suboptimal intervention, the network parameters of the second policy network can be updated, allowing the second policy network to benefit from the human intervention data while getting rid of the dependence on the perfection of the human intervention data. This enables safer, more efficient, and robust policy learning in a real, noisy, and high-cost physical environment.

[0056] Among them, the reinforcement learning objective refers to the objective of the reinforcement learning objective link in the learning end. The reinforcement learning objective can be represented by the second policy network and the value network. The intervention constraint objective refers to the objective of the intervention constraint objective link. The intervention constraint objective can be represented by the second policy network and the expert network.

[0057] For example, continue to refer to Figure 1 As shown, first training data can be sampled from the intervention experience pool, and second training data can be sampled from the online experience pool. The first training data and the second training data are then combined to form the current training dataset.

[0058] In one example, the network parameters of the learnable state network can be updated based on the current training dataset, and the network parameters of the second policy network can be updated based on the current training dataset, the updated learnable state network, the value network, and the expert network.

[0059] In another example, the expert network can be updated based on the first training data, the value network can be updated based on the second training data, and the network parameters of the learnable state network can be updated based on the current training dataset. Thus, the network parameters of the second policy network are updated based on the current training dataset, the updated learnable state network, the updated value network, and the updated expert network.

[0060] S204. The action terminal obtains the network parameters of the second strategy network according to the preset interaction cycle, and updates the network parameters of the first strategy network according to the network parameters of the second strategy network.

[0061] Optionally, the action terminal can obtain the network parameters of the second strategy network according to a preset interaction period, and update the network parameters of the first strategy network based on the network parameters of the second strategy network.

[0062] Optionally, the action terminal can make action decisions and execute actions based on the updated first policy network, that is, repeat steps S201-S204.

[0063] The execution order of S201-S204 in this application can be sequential, for example, executing S201-S204 in turn. Alternatively, it can be non-sequential or asynchronous. This can be adjusted according to the actual situation in practical applications. This embodiment only illustrates sequential execution as an example.

[0064] S205. Use the second policy network when updates stop as the decision network of the robotic arm, and deploy the decision network to the robotic arm for motion control.

[0065] Optionally, the second policy network when updates stop can be used as the decision network of the robotic arm, and the decision network can be deployed to the robotic arm for motion control.

[0066] Optionally, the second policy control network at the end of training can be used as the decision network of the robotic arm, and the decision network can be deployed to other robotic arms with the same configuration as the current robotic arm to achieve control of the robotic arm.

[0067] In this embodiment, by setting up asynchronous action and learning ends in the reinforcement learning policy training system, and deploying a first policy network in the action end and a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network in the learning end, the policy is executed through the action end, triggering human intervention and collecting real interaction data. The learnable state network, intervention experience pool, online experience pool, value network, and expert network are updated through the learning end. This achieves a closed-loop training structure of real-machine online reinforcement learning and human intervention constraints. The policy network update can be modeled as a state-dependent constrained reinforcement learning problem, and it learns from human interventions that may not be optimal and contain noise. During the learning process, the constraint strength is adaptively adjusted through the learnable state network. Under high uncertainty states where human intervention fluctuates greatly and may be suboptimal, the constraints are relaxed to avoid limiting the policy upper limit due to suboptimal interventions, encouraging the policy to rely more on the reinforcement learning objective for autonomous improvement. Under low uncertainty states where human intervention is consistent and reliable, the constraints are tightened to encourage the policy to conform to human guidance, thereby improving sample efficiency, accelerating convergence speed, and achieving adaptive compatibility and robust learning with suboptimal human intervention data.

[0068] Furthermore, the training process can appropriately change the task objectives, provide states in various scenarios, actively construct and introduce diverse training data, and force the policy model to learn to extract more generalized state representations and decision-making logic, thereby systematically enhancing its adaptability and robustness in unknown scenarios.

[0069] Simultaneously, during training, it not only supports real-time human intervention in data processing but also possesses a certain degree of online autonomous exploration capability. When data is limited or of insufficient quality, it can effectively compensate for performance degradation caused by insufficient data. Furthermore, it reduces time consumption, equipment wear and tear, and security risks associated with frequent trial and error in real-world environments.

[0070] In one possible implementation, Figure 3 This is another flowchart illustrating the reinforcement learning strategy training method provided in an embodiment of this application, referring to... Figure 3 As shown, in S202 above, the intervention experience pool and the online experience pool are updated based on the experience data of each action, or the online experience pool is updated, including: S301. Add all action experience data to the online experience pool to update the online experience pool.

[0071] Optionally, all action experience data can be added to the online experience pool to update the online experience pool.

[0072] S302. Determine whether there is at least one human action data in each action experience data. If so, add each human action data to the intervention experience pool to update the intervention experience pool.

[0073] Optionally, the motion experience data is traversed, and for the current motion experience data that is traversed, it is determined whether the current motion experience data is human motion data.

[0074] For example, for current motion experience data, it can be determined whether the current motion experience data is human motion data based on the state signals in the current motion experience data.

[0075] For example: if the status signal in the current action experience data Then it can be determined that the current action experience data is human action data, if the state signal in the current action experience data If so, it can be determined that the current action experience data is not human action data.

[0076] Optionally, if the current action experience data is human action data, the current action experience data is added to the intervention experience pool to update the intervention experience pool.

[0077] Optionally, if the current motion experience data is not human motion data, then the next motion experience data is traversed.

[0078] By updating the intervention experience pool and the online experience pool using data from various actions, or by updating the online experience pool, we can separate ordinary interaction data from human intervention data. This allows us to model long-term rewards and human behavioral patterns under dynamic environmental conditions, preventing low-quality or suboptimal interventions from polluting the overall experience pool and improving learning stability.

[0079] In one possible implementation, before updating the network parameters of the learnable state network based on the intervention experience pool and the online experience pool in step S203 above, the learning end can also be initialized, specifically including: Initialize the second policy network 2. Value function network Expert Network Lagrange multiplier network The spatial dimension of robotic arm movement Initialize the intervention experience playback pool Online reinforcement learning experience replay pool Initialize training duration .

[0080] For example, the motion space dimension of a robotic arm Training duration It can last between 40 and 60 minutes.

[0081] Optionally, load the initial intervention data into the intervention experience playback pool. middle.

[0082] In one possible implementation, Figure 4 This is a flowchart illustrating the process of updating network parameters of the learnable state network in the reinforcement learning policy training method provided in this application embodiment, with reference to... Figure 4 As shown, in S203 above, the network parameters of the learnable state network are updated based on the intervention experience pool and the online experience pool, including: S401. When the current amount of data in the online experience pool is greater than the preset first quantity threshold, determine the current amount of newly added data corresponding to the intervention experience pool.

[0083] Optionally, the current amount of data in the online experience pool can be determined, i.e., the current sample size. | Whether it is greater than the preset first quantity threshold, if the current sample size| If the number is greater than the first threshold, then the intervention experience pool is determined. The corresponding amount of newly added data.

[0084] Among them, the intervention experience pool The corresponding current new data volume refers to the number of new action experience data added to the intervention experience pool.

[0085] The preset first quantity threshold can be a positive integer, such as 100.

[0086] By determining the current amount of data in the online experience pool, we can ensure the minimum data support for basic training in reinforcement learning, thereby preventing early overfitting or gradient explosion and improving learning stability.

[0087] S402. Based on the current amount of newly added data, determine whether to update the expert network.

[0088] Optionally, the intervention experience pool The system determines the amount of newly added data. If the amount of newly added data exceeds the preset second threshold, the system will update the expert network.

[0089] For example, the second quantity threshold can be a positive integer, and the second quantity threshold is less than the first quantity threshold, such as 50.

[0090] By analyzing the current volume of newly added data, it determines whether to update the expert network. It can dynamically decide whether to enable the imitation learning branch based on the real-time data stream, achieving on-demand activation of the imitation learning branch, controlling the update frequency of the expert network, avoiding redundant computation, reducing GPU load fluctuations caused by frequent backpropagation, and improving training efficiency. It also supports modeling non-stationary intervention behaviors.

[0091] S403. If so, update the expert network based on the intervention experience pool.

[0092] Optionally, taking a single update as an example, if the amount of newly added data exceeds a preset second threshold, then the data will be retrieved from the intervention experience pool. Sampling is performed, and the expert network is calculated according to the following formula. loss function and update the expert network. Network parameters :

[0093] in, To extract from the intervention experience pool The "state-action" pairs obtained from the sampling, For state, For action.

[0094] It is understandable that step S403 can be repeated multiple times, for example, 50 times.

[0095] Optionally, if the amount of newly added data is less than or equal to a preset second threshold, the expert network will not be updated. Proceed directly to step S404.

[0096] S404. Sample the first training data from the intervention experience pool and sample the second training data from the online experience pool, and combine the first training data and the second training data into the current training dataset.

[0097] Optionally, from the intervention experience pool The first training data was obtained by sampling from the online experience pool. The second training data is obtained by sampling from the first training data, and the first training data and the second training data are combined to form the current training dataset. The sampling can be performed according to a preset sampling algorithm to obtain the first training data and the second training data.

[0098] S405. Update the network parameters of the learnable state network based on the expert network and the current training dataset.

[0099] Optionally, if the expert network is updated in S403 above, the network parameters of the learnable state network can be updated based on the updated expert network and the current training dataset.

[0100] Optionally, if the expert network is not updated in S403 above, the network parameters of the learnable state network can be updated based on the expert network and the current training dataset.

[0101] For example, the network parameters of the learnable state network can be updated based on the deviation between the expert network and the second policy network on the current training dataset.

[0102] By using an expert network to update the network parameters of a learnable state network, it is possible to ensure that the update of the learnable state network is based on the latest and most accurate estimates of human behavior, thereby achieving accurate state-based adaptive constraints.

[0103] In one possible implementation, Figure 5 This is another flowchart illustrating the process of updating the network parameters of the learnable state network in the reinforcement learning policy training method provided in this application embodiment, referring to... Figure 5 As shown, in S405 above, the network parameters of the learnable state network are updated based on the expert network and the current training dataset, including: S501. Obtain the current training data in the current training dataset.

[0104] Optionally, the current training data in the current training dataset can be obtained, wherein the current training data includes the state. .

[0105] S502. Determine the distance between the expert network and the second policy network under the current training data.

[0106] Optionally, an expert network can be identified. With the second strategy network 2. The current state of the training data The distance below .

[0107] Specifically, distance The following formula can be used to calculate it:

[0108] S503. Determine the variance of the expert network under the current training data.

[0109] Optionally, the expert network can be calculated. The current state of the training data variance .

[0110] S504. Based on the distance and variance, calculate the loss result of the learnable state network corresponding to the current training data.

[0111] Alternatively, a learnable state network can be calculated based on the distance and variance. The current state of the training data Corresponding loss results For details, please refer to the following formula:

[0112] in, This represents the spatial dimension of the robotic arm's movements.

[0113] S505. Update the network parameters of the learnable state network based on the loss results corresponding to each training data in the current training dataset.

[0114] In one example, after obtaining the loss result of the learnable state network corresponding to the current training data in the current training dataset, the gradient of the loss result corresponding to the current training data can be calculated, and the network parameters of the learnable state network can be updated.

[0115] In another example, after obtaining the loss results of each training data in the current training dataset for the learnable state network, the gradient can be calculated and the network parameters of the learnable state network can be updated after fusing the loss results of each training data.

[0116] By determining the distance between the expert network and the second policy network under the current training data, and determining the variance of the expert network under the current training data, and calculating the loss result of the learnable state network under the current training data based on the distance and variance, the network parameters of the learnable state network are updated according to the loss result. The network parameters of the learnable state network can be updated through dual gradient ascent. Under the low uncertainty state of consistent and reliable human intervention, the network parameters of the learnable state network are automatically increased to strengthen the imitation. Under the high uncertainty state of large fluctuations in human intervention and possible suboptimal results, the network parameters of the learnable state network are automatically decreased to release the exploration space. This achieves adaptive compatibility and robust learning for suboptimal human intervention, solves the problem of blindly trusting intervention signals in existing technologies, and improves training efficiency and final policy performance.

[0117] In one possible implementation, Figure 6 This is a flowchart illustrating the process of updating the network parameters of the second policy network in the reinforcement learning policy training method provided in this application embodiment, with reference to... Figure 6 As shown, in S203 above, the network parameters of the second policy network are updated based on the updated learnable state network, value network, and expert network, including: S601. Update the value network based on the current training dataset.

[0118] Optionally, the value network can be updated based on the current training dataset.

[0119] The current training dataset can be obtained by referring to the aforementioned step S404, which will not be elaborated here.

[0120] For example, the value network can be calculated based on the following formula. loss function And compute gradient update value network Network parameters :

[0121] in, To start from the current training dataset The "state-action-next state" pairs obtained from the sampling, The real value of the reward obtained from the reward model. For the objective value function network, the following is obtained based on the soft update of the value network: , , For the target policy network, the following is obtained through soft updates of the second policy network: .

[0122] S602. Update the network parameters of the second policy network based on the updated value network, learnable state network, expert network, and the current training dataset.

[0123] Optionally, the network parameters of the second policy network can be updated based on the updated value network, learnable state network, expert network, and current training dataset, with the goal of maximizing long-term returns while conditionally approximating human intervention behavior without blindly obeying it.

[0124] In one possible implementation, Figure 7 This is another flowchart illustrating the process of updating the network parameters of the second policy network in the reinforcement learning policy training method provided in this application embodiment, referring to... Figure 7 As shown, in S602 above, the network parameters of the second policy network are updated based on the updated value network, the learnable state network, the expert network, and the current training dataset, including: S701. Determine the current reinforcement learning objective based on the updated value network and the current training dataset.

[0125] For example, the state of the current training data in the current training dataset. For example, the network can be updated based on the value. and status Determine the current reinforcement learning objective .

[0126] S702. Based on the expert network and the current training dataset, determine the current imitation learning target.

[0127] For example, the state of the current training data in the current training dataset. For example, it can be based on expert networks and status Determine the current imitation learning objective .

[0128] S703. Based on the current reinforcement learning objective, the current imitation learning objective, the learnable state network, and the current training dataset, calculate the current loss of the second policy network.

[0129] Optionally, the current loss of the second policy network can be obtained by adjusting the balance between the current reinforcement learning objective and the current imitation learning objective through a learnable state network.

[0130] For example, the loss function of the second policy network can be calculated using the following formula. :

[0131] It is understandable that the learnable state network adjusts the balance between the current reinforcement learning objective and the current imitation learning objective in the second policy network. 2. With expert network Distance between Exceeding the limit back, The loss of the current imitation learning target increases, and the second policy network... 2. More consideration should be given to the current situation. Approaching Experts Strategy If the second strategy network 2. With expert network Distance between Meet the restrictions ,but Gradually reduce to 0, second policy network 2. More dependent on its own value network estimation update of the second policy network 2 Network parameters This allows for an adaptive adjustment of the distance between human intervention strategies and learning strategies, thereby leveraging potentially suboptimal and noisy human interventions to accelerate learning.

[0132] S704. Update the network parameters of the second policy network based on the current loss of the second policy network.

[0133] Optionally, the second policy network can be updated by calculating gradients based on the current loss of the second policy network. 2 Network parameters .

[0134] In one possible implementation, before updating the network parameters of the learnable state network based on the intervention experience pool and the online experience pool in step S203 above, the method further includes: Initial intervention data is generated through the action terminal and added to the intervention experience pool.

[0135] Optionally, the initial intervention data can be generated offline, specifically including: Step 001: Set the external intervention signal I=1, set the mode to fully human intervention mode, and initialize at the specified time. Number of successfully initialized trajectories .

[0136] Step 002: Reset the task scene.

[0137] Step 003: Obtain the RGB image of the environment where the robotic arm is located from the camera, and obtain the current pose from the robotic arm. The current state is obtained by combining the results. .

[0138] Step 004: The human operator controls the movement of the robotic arm via the isomorphic arm.

[0139] Step 005: Determine the pose after movement at the action end. According to the current pose pose at the previous moment Calculate the relative movement of the robotic arm and save it as a human intervention action. .

[0140] in, This represents the relative motion of the gripper end effector of the robotic arm in 6D space. Specifically, the first three dimensions are... , , The translation increment in the direction, with the latter three dimensions expressed in Euler angles. , , Rotation increments in the (roll, pitch, yaw) directions.

[0141] Step 006: The human operator determines whether the task has been successfully completed and updates the task signal for the current task. .

[0142] Step 007: The action end determines the task signal based on the current task. Assign a reward signal if A reward signal will be sent upon completion of the task. Otherwise, a reward signal. .

[0143] Step 008: The action end transfers a piece of action experience data ( , , , , Place it into the experience replay pool.

[0144] Step 009: If the task is successfully completed, Proceed to the next step; otherwise, return to step 003.

[0145] Step 010, if If the content in the experience replay pool is used as the initial intervention data, the offline data collection task ends; otherwise, return to step 002.

[0146] Optionally, the initial intervention data can be added to the intervention experience pool.

[0147] In one possible implementation, Figure 8 This is a flowchart illustrating the generation of multiple action experience data in the reinforcement learning strategy training method provided in this application embodiment, with reference to... Figure 8 As shown, in S201 above, action decisions are made and actions are executed based on the first policy network, generating multiple action experience data, including: S801. Obtain the current state of the action terminal, and generate the current action based on the current state through the first policy network.

[0148] Optionally, the current environment image of the current robotic arm's location and the current pose of the current robotic arm can be obtained. and compare the current environment image with the current pose. Combined into the current state .

[0149] Optionally, through the first policy network Regarding the current state Make motion decisions and generate the current motion of the robotic arm. .

[0150] S802. Based on the current action, determine whether to enter the takeover state.

[0151] In one example, upon obtaining the current action... Then, the current action can be adjusted according to preset boundary conditions. Make a judgment to determine whether to enter the takeover state.

[0152] In another example, after obtaining the current action... Then, the current action can be... The message is pushed to the user interaction device corresponding to the current robotic arm, and the system responds to the user's actions on the user interaction device to determine whether to enter the takeover state. The user interaction device can be a homogeneous arm, such as a teach pendant or a teleoperated device.

[0153] The takeover state refers to the temporary interruption of the automatic control of the current robotic arm when the task strategy being executed autonomously deviates, faces high-risk operations, or tends to fail, and the current robotic arm is manually controlled to complete the corrective action.

[0154] S803, if so, then in response to the user's operation data, generate action experience data.

[0155] Optionally, if it is determined that the takeover state has been entered, the robot arm's post-movement pose can be determined in response to the user's operation on the user interaction device. Based on the pose of the robotic arm after movement and the current forward pose of the robotic arm Calculate the relative movement of the robotic arm to obtain the actual motion, and then use the actual motion to analyze the current action. Perform a replacement to obtain the new current action. Among these, users can be human operators.

[0156] Optionally, upon obtaining the new current action Then, using the pre-trained reward model, it is determined whether the current task has ended, and based on whether the current task has ended, the reward signal of the current action, the task signal of the current task, and the state signal of the current action are assigned values ​​to obtain action experience data.

[0157] S804. If not, generate action experience data based on the current action.

[0158] Optionally, if it is determined not to enter the takeover state, then based on the current action, the reward model obtained through pre-training is used to determine whether the current task has ended, and based on whether the current task has ended, the reward signal of the current action, the task signal of the current task, and the state signal of the current action are assigned values ​​to generate action experience data.

[0159] In one possible implementation, Figure 9 This is another flowchart illustrating the process of updating the network parameters of the second policy network in the reinforcement learning policy training method provided in this application embodiment, referring to... Figure 9 As shown, in the above S803, in response to the user's operation data, action experience data is generated, including: S901, responding to user operation data, determines the pose after the operation.

[0160] Optionally, the pose of the robotic arm after movement is determined in response to the user's operation on the user interaction device. and the current pose of the robotic arm after movement. As the pose after the operation.

[0161] S902. Generate the actual action based on the current pose and the pose after the operation.

[0162] Optionally, based on the pose after the operation and current pose The relative movement of the robotic arm is calculated to obtain the actual motion.

[0163] S903. Determine the status signal, reward signal, and task signal corresponding to the actual action.

[0164] Optionally, after obtaining the actual action, the state signal of the actual action can be... The value is assigned to 1, and the pre-trained reward model is used to determine whether the current task has ended. Based on whether the current task has ended, the reward signal for the actual action and the task signal for the current task are assigned values ​​to obtain the reward signal corresponding to the actual action. and mission signals .

[0165] For example, if the current task ends, then... , If the current task is not finished, then... , .

[0166] S904. Generate action experience data based on the actual action, the status signal corresponding to the actual action, the reward signal corresponding to the actual action, and the task signal corresponding to the actual action.

[0167] Optionally, the current action can be replaced by the actual action, and the current state of the current action can be used as the current state corresponding to the actual action. Based on the actual action, the current state corresponding to the actual action, the state signal corresponding to the actual action, the reward signal corresponding to the actual action, and the task signal corresponding to the actual action, an action experience data can be generated.

[0168] In one possible implementation, before S201 above, which involves making action decisions and executing actions based on the first policy network to generate multiple action experience data, the following further step is included: Initialize the first policy network and load the reward model.

[0169] Optionally, before making action decisions based on the first policy network, the action end can also initialize the first decision network, load the reward model, establish a connection with the learning end, obtain the network parameters of the second policy network, and assign the network parameters of the second policy network to the first policy network.

[0170] Optionally, an experience cache pool can also be set in the action terminal. After obtaining multiple action experience data, the multiple action experience data can be cached in the experience cache pool so that the action experience data can be sent to the learning terminal when the preset interaction period is reached or in real time.

[0171] In one possible implementation, the reward model The training process of the reward model is illustrated by example. This can be obtained through offline training, specifically including: Step 1: Load the initial intervention data into the training dataset of the reward model to initialize the reward model. , The network parameters are for the reward model.

[0172] Step 2: Traverse the training dataset of the reward model and perform preprocessing label steps. Specifically, when traversing the current training data, if... Then the label Otherwise, the label .

[0173] Step 3: Calculate the predicted loss, which can be found using the following formula:

[0174] in, This refers to the current state-current label pair sampled from the training dataset of the reward model.

[0175] Step 4: The network optimizer moves towards... Lowering the direction updates the reward model Network parameters .

[0176] Step 5: Repeat the update 500 times and then save the reward model. .

[0177] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A reinforcement learning strategy training method, characterized in that, An application is made to a reinforcement learning policy training system, the system comprising: an action end and a learning end, wherein the action end is deployed with a first policy network, and the learning end is deployed with a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network, wherein the first policy network and the second policy network have the same network structure, and the learnable state network is a learnable Lagrange multiplier network that adaptively adjusts the trade-off between the reinforcement learning objective and the human intervention imitation objective in the state space, the method comprising: The action terminal makes action decisions and executes actions based on the first policy network, generating multiple action experience data; The learning terminal acquires multiple action experience data from the action terminal in real time, and updates the intervention experience pool and the online experience pool based on each action experience data, or updates the online experience pool. The learning end updates the network parameters of the learnable state network based on the intervention experience pool and the online experience pool, and updates the network parameters of the second policy network based on the updated learnable state network, the value network and the expert network. The action terminal acquires the network parameters of the second strategy network according to a preset interaction cycle, and updates the network parameters of the first strategy network according to the network parameters of the second strategy network. The second strategy network at the point of no update is used as the decision network for the robotic arm, and the decision network is deployed to the robotic arm for motion control.

2. The reinforcement learning strategy training method according to claim 1, characterized in that, The step of updating the intervention experience pool and the online experience pool based on the action experience data, or updating the online experience pool, includes: Add all action experience data to the online experience pool to update the online experience pool; Determine whether at least one human action data exists in each of the action experience data. If so, add each of the human action data to the intervention experience pool to update the intervention experience pool.

3. The reinforcement learning strategy training method according to claim 1, characterized in that, The step of updating the network parameters of the learnable state network based on the intervention experience pool and the online experience pool includes: When the current amount of data in the online experience pool is greater than a preset first quantity threshold, the current amount of newly added data corresponding to the intervention experience pool is determined. Based on the current amount of newly added data, determine whether to update the expert network; If so, then update the expert network according to the intervention experience pool; First training data is sampled from the intervention experience pool, and second training data is sampled from the online experience pool. The first training data and the second training data are then combined to form the current training dataset. The network parameters of the learnable state network are updated based on the expert network and the current training dataset.

4. The reinforcement learning strategy training method according to claim 3, characterized in that, The step of updating the network parameters of the learnable state network based on the expert network and the current training dataset includes: Obtain the current training data from the current training dataset; Determine the distance between the expert network and the second policy network under the current training data; Determine the variance of the expert network under the current training data; Based on the distance and the variance, the loss result of the learnable state network corresponding to the current training data is calculated; The network parameters of the learnable state network are updated based on the loss results corresponding to each training data in the current training dataset.

5. The reinforcement learning strategy training method according to claim 1, characterized in that, The step of updating the network parameters of the second policy network based on the updated learnable state network, the value network, and the expert network includes: Update the value network based on the current training dataset; The network parameters of the second policy network are updated based on the updated value network, the learnable state network, the expert network, and the current training dataset.

6. The reinforcement learning strategy training method according to claim 5, characterized in that, The step of updating the network parameters of the second policy network based on the updated value network, the learnable state network, the expert network, and the current training dataset includes: Based on the updated value network and the current training dataset, the current reinforcement learning objective is determined; Based on the expert network and the current training dataset, determine the current imitation learning target; The current loss of the second policy network is calculated based on the current reinforcement learning objective, the current imitation learning objective, the learnable state network, and the current training dataset. Update the network parameters of the second policy network based on the current loss of the second policy network.

7. The reinforcement learning strategy training method according to claim 1, characterized in that, Before updating the network parameters of the learnable state network based on the intervention experience pool and the online experience pool, the method further includes: Initial intervention data is generated through the action terminal and added to the intervention experience pool.

8. The reinforcement learning strategy training method according to claim 1, characterized in that, The process of making action decisions and executing actions based on the first policy network, generating multiple action experience data, includes: Obtain the current state of the action terminal, and generate the current action based on the current state through the first policy network; Based on the current action, determine whether to enter the takeover state; If so, then in response to the user's operation data, a set of action experience data is generated; If not, then generate action experience data based on the current action.

9. The reinforcement learning strategy training method according to claim 8, characterized in that, The process of generating action experience data in response to user operation data includes: In response to user action data, determine the current pose and the previous pose; Generate the target action based on the previous pose and the current pose; Determine the target signal, target state, and target reward corresponding to the target action; Based on the target action, the target signal corresponding to the target action, the target state, and the target reward, generate action experience data.

10. A reinforcement learning strategy training system, characterized in that, include: The system comprises an action terminal and a learning terminal. The action terminal is equipped with a first policy network, and the learning terminal is equipped with a second policy network, a learnable state network, an intervention experience pool, an online experience pool, a value network, and an expert network. The first policy network and the second policy network have the same network structure. The system is used to execute the steps of the reinforcement learning policy training method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Robot walking control method and system based on deep reinforcement learning and medium

    CN111580385A