A compensation method for a robot imitation learning training process

CN121028677BActive Publication Date: 2026-09-25JIANGSU YUNMU ZHIZAO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511122318.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2026-09-25
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

[0004]本发明所要解决的技术问题是:针对人形机器人运动控制中因机器人系统的强非线性使得策略网络难以训练和部署等问题,本发明提供了一种在机器人模仿学习过程中进行补偿的方法

Benefits of technology

[0044]1、本发明提出了一种基于模型的机器人关节角补偿方法,该方法能显著降低模仿学习中训练和部署的难度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121028677B_ABST
    Figure CN121028677B_ABST
Patent Text Reader

Abstract

The application discloses a kind of compensation methods of robot imitation learning training process, include the following steps: step S1: real data acquisition;The motion trajectory, mechanical characteristics and behavior intention of human demonstrator are accurately recorded by multi-modal sensing system, step S2: robot data redirection;Step S3: parallel training environment construction;Step S4: imitative learning discriminator initialization;Step S5.1: dynamics compensation system construction;Step S5.2: robot system simplification;Step S6: imitative learning strategy optimization;Step S7: real machine deployment.The application proposes a kind of robot joint angle compensation method based on model, which can significantly reduce the difficulty of training and deployment in imitation learning, can improve the accuracy of robot action and the integrity of task, effectively fuses visual information and tactile information, lays a key technical foundation for future intelligent service robot autonomous operation and decision system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot control, and specifically relates to a compensation method for the robot's imitation learning and training process. Background Technology

[0002] Imitation learning, as an important branch of reinforcement learning, has become an effective method for solving robot motion control problems by using expert demonstration data to guide agent policy optimization. Humanoid robots can accelerate convergence and avoid getting trapped in poor local optima through imitation learning. Common imitation learning methods typically collect data from real humans and redirect it to the robot's joint movements. By following the joint movement metrics, the robot imitates the movement style and behavior taught by humans. However, in imitation learning, following specific joint movement targets can easily reduce the robot's generalization and robustness, leading to an inability to adapt to changes in the external environment (e.g., from Earth to space, where gravitational acceleration decreases, rendering imitation learning deployment ineffective). Furthermore, the highly nonlinear dynamics of robots during training not only increase the difficulty of policy training but also create a "simulation-reality gap" during actual deployment.

[0003] To address the aforementioned issues in robot imitation learning, such as training difficulty, generalization performance, and deployment effectiveness, this invention proposes a model-based compensation method (external controller). This method approximates the robot's actuator as a linear system and compensates for the nonlinear components through a dynamic model, thereby reducing the difficulty of training and deployment and improving the robot's performance. Summary of the Invention

[0004] The technical problem to be solved by this invention is: in the motion control of humanoid robots, the strong nonlinearity of the robot system makes it difficult to train and deploy the policy network. This invention provides a method for compensation during the robot's imitation learning process.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a compensation method for the robot imitation learning training process, comprising the following steps:

[0006] Step S1: Real-person data collection;

[0007] The motion trajectory, mechanical characteristics, and behavioral intentions of human demonstrators are accurately recorded using multimodal sensing systems (such as optical motion capture, inertial measurement units, and force feedback devices).

[0008] Step S2: Robot data redirection;

[0009] Based on collected real-human data, and combined with biomechanical parameter calibration and motion redirection algorithms, human movements are mapped to the robot's joint space, forming multi-dimensional demonstration data encompassing kinematics (position / velocity / acceleration) and dynamics (torque / power). Through adaptive motion scaling and domain randomization enhancement techniques, the problem of motion distortion caused by differences in human-robot morphology is effectively solved.

[0010] Step S3: Constructing the parallel training environment;

[0011] Construct an environment for parallel training of multiple robots, including the construction of core layers such as the physical simulation layer, data pipeline layer, and policy training layer;

[0012] Step S4: Initialize the learning discriminator;

[0013] A multi-layered network architecture is constructed, and expert data is preprocessed to allow the policy network to learn by imitating expert demonstrations.

[0014] Step S5:

[0015] Step S5.1: Construction of the dynamic compensation system;

[0016] The robot system can be modeled as follows:

[0017]

[0018] At this point, the system's dynamic compensator can be expressed as:

[0019]

[0020] The compensator includes G(q) with only static compensation term retained, and Coriolis force compensation.

[0021] The compensation items can be expanded as needed, as follows:

[0022]

[0023] Among them, J T F ext It can be a real end reaction force or a planned end reaction force (such as the ground reaction force acting on the foot). For system dissipation forces (such as friction between joints), f is the friction coefficient, and compensators with various hybrid compensation methods;

[0024] The corresponding terms of external compensation can be extended as constant parameters, such as extending the gravity term G(q):

[0025] G(q)→G(q,M,g);

[0026] This allows operators to modify the gravitational acceleration in the current environment, as well as the mass of the replaced part, and the external force F. ext To expand:

[0027] F ext →F ext (F z ,μ)

[0028] →F ext (F z ,c (i) ,φ (i) );

[0029] The adhesion coefficient μ is used to correlate the vertical force and the normal force F. z To adapt to different paved road conditions, or to model the soil so that the robot operator can adjust the soil cohesion in real time. ( i ) and shear angle φ (i) Enables operation in complex terrains such as mud, sand, and vegetation.

[0030] Step S5.2: Simplify the robot system;

[0031] According to step S5.1, the constructed nonlinear component is used as the feedback linearization component to construct a feedback system:

[0032]

[0033] Among them, K p (q d -q) is the integral term of the PD controller, K p The integral coefficient is... K is the differential term of the PD controller. d is the differential coefficient. q is a nonlinear term in the dynamics. d , This is the output of the policy network. Because... The processing methods used are difficult to observe and are relatively small in real-world environments, therefore they are either ignored or learned through imitation of learning styles.

[0034] Furthermore, compensation for nonlinear components can also be achieved using optimization techniques, as follows:

[0035]

[0036] Where, f(q,τ) act ) is the constraint function for the system state and output.

[0037] At this point, the robot joint actuator retains only linear characteristics, and the model simplifies to:

[0038]

[0039] Step S6: Optimize the imitation learning strategy;

[0040] The goal of imitation learning is to continuously optimize the parameters of the policy network using a series of data collected from real humans, ultimately achieving a robot movement that is similar to or even identical to that of a real human. During policy optimization, an imitation learning discriminator rewards the robot's actions, and the reward results are used to adjust the actuator parameters. Linear actuators, combined with compensation for nonlinear components, generate the final actuator torque to control the robot's movements. In a simulation environment, the robot interacts with the environment, and the reward mechanism of imitation learning continuously updates the parameters of the policy network, enabling rapid training of the robot's imitation learning.

[0041] Step S7: Deploy on real devices;

[0042] The designed nonlinear compensator and the collaboratively trained policy network are used as the controller for the robot's physical machine.

[0043] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:

[0044] 1. This invention proposes a model-based method for robot joint angle compensation, which can significantly reduce the difficulty of training and deployment in imitation learning.

[0045] 2. This invention uses visual and tactile information to understand and analyze the object grasped by the robot, which can improve the accuracy of the robot's movements and the completeness of the task.

[0046] 3. This invention effectively integrates visual and tactile information, laying a key technical foundation for the autonomous operation and decision-making system of future intelligent service robots. Attached Figure Description

[0047] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0048] Figure 1 This is a flowchart of the robot imitation learning training process using the model compensation method.

[0049] Figure 2 This is a flowchart illustrating the deployment process of robot imitation learning using the model compensation method. Detailed Implementation

[0050] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] Example 1

[0052] The following describes in detail a compensation method for the robot imitation learning training process provided by the present invention, including the following steps:

[0053] Step S1: Real-person data collection;

[0054] The motion trajectory, mechanical characteristics, and behavioral intentions of human demonstrators are accurately recorded using multimodal sensing systems (such as optical motion capture, inertial measurement units, and force feedback devices).

[0055] Step S2: Robot data redirection;

[0056] Based on collected real-human data, and combined with biomechanical parameter calibration and motion redirection algorithms, human movements are mapped to the robot's joint space, forming multi-dimensional demonstration data encompassing kinematics (position / velocity / acceleration) and dynamics (torque / power). Through adaptive motion scaling and domain randomization enhancement techniques, the problem of motion distortion caused by differences in human-robot morphology is effectively solved.

[0057] Step S3: Constructing the parallel training environment;

[0058] Construct an environment for parallel training of multiple robots, including the construction of core layers such as the physical simulation layer, data pipeline layer, and policy training layer;

[0059] Step S4: Initialize the learning discriminator;

[0060] A multi-layered network architecture is constructed, and expert data is preprocessed to allow the policy network to learn by imitating expert demonstrations.

[0061] Step S5:

[0062] Step S5.1: Construction of the dynamic compensation system;

[0063] The robot system can be modeled as follows:

[0064]

[0065] At this point, the system's dynamic compensator can be expressed as:

[0066]

[0067] The compensator includes G(q) with only static compensation term retained, and Coriolis force compensation.

[0068] The compensation items can be expanded as needed, as follows:

[0069]

[0070] Among them, J T F ext It can be a real end reaction force or a planned end reaction force (such as the ground reaction force acting on the foot). For system dissipation forces (such as friction between joints), f is the friction coefficient, and compensators with various hybrid compensation methods;

[0071] The corresponding terms of external compensation can be extended as constant parameters, such as extending the gravity term G(q):

[0072] G(q)→G(q,M,g);

[0073] This allows operators to modify the gravitational acceleration in the current environment, as well as the mass of the replaced part, and the external force F. ext To expand:

[0074] F ext →F ext (F z ,μ)

[0075] →F ext (F z ,c (i) ,φ (i) );

[0076] The adhesion coefficient μ is used to correlate the vertical force and the normal force F. z To adapt to different paved road conditions, or to model the soil so that the robot operator can adjust the soil cohesion in real time. ( i ) and shear angle φ (i) Enables operation in complex terrains such as mud, sand, and vegetation.

[0077] Step S5.2: Simplify the robot system;

[0078] According to step S5.1, the constructed nonlinear component is used as the feedback linearization component to construct a feedback system:

[0079]

[0080] Among them, K p (q d -q) is the integral term of the PD controller, K p The integral coefficient is... K is the differential term of the PD controller. d is the differential coefficient. q is a nonlinear term in the dynamics. d , This is the output of the policy network. Because... The processing methods used are difficult to observe and are relatively small in real-world environments, therefore they are either ignored or learned through imitation of learning styles.

[0081] Furthermore, compensation for nonlinear components can also be achieved using optimization techniques, as follows:

[0082]

[0083] Where, f(q,τ) act ) is the constraint function for the system state and output.

[0084] At this point, the robot joint actuator retains only linear characteristics, and the model simplifies to:

[0085]

[0086] Step S6: Optimize the imitation learning strategy;

[0087] The goal of imitation learning is to continuously optimize the parameters of the policy network using a series of data collected from real humans, ultimately achieving a robot movement that is similar to or even identical to that of a real human. During policy optimization, an imitation learning discriminator rewards the robot's actions, and the reward results are used to adjust the actuator parameters. Linear actuators, combined with compensation for nonlinear components, generate the final actuator torque to control the robot's movements. In a simulation environment, the robot interacts with the environment, and the reward mechanism of imitation learning continuously updates the parameters of the policy network, enabling rapid training of the robot's imitation learning.

[0088] Step S7: Deploy on real devices;

[0089] The designed nonlinear compensator and the collaboratively trained policy network are used as the controller for the robot's physical machine.

[0090] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A compensation method for a robot's imitation learning and training process, characterized in that, Includes the following steps: Step S1: Real-person data collection; The system uses a multimodal sensing system to accurately record the motion trajectory, mechanical characteristics, and behavioral intentions of human demonstrators. Step S2: Robot data redirection; Based on collected real-person data, and combined with biomechanical parameter calibration and motion redirection algorithms, human movements are mapped to the robot joint space to form multi-dimensional demonstration data containing kinematics and dynamics; through adaptive motion scaling and domain randomization enhancement technology, the motion distortion problem caused by human-machine morphology differences is effectively solved. Step S3: Constructing the parallel training environment; Construct an environment for parallel training of multiple robots, including the construction of the core layers of physical simulation layer, data pipeline layer and policy training layer; Step S4: Initialize the learning discriminator; A multi-layered network architecture is built, and the expert data is preprocessed to allow the policy network to learn by imitating the expert demonstrations. Step S5: Step S5.1: Construction of the dynamic compensation system; The robot system is modeled, and the model is represented as follows: ; At this point, the system's dynamic compensator is expressed as: ; This compensator includes those that retain only the static compensation term. Coriolis force compensation , The compensation items are expanded as follows: ; in, For actual or planned end reaction forces, For the system dissipation force, The coefficient of friction; The corresponding terms of external compensation are extended to constant parameters, including the gravity term. To expand: ; This allows operators to modify the gravitational acceleration in the current environment, as well as the mass of the replaced parts. External forces To expand: ; With adhesion coefficient Related vertical and normal forces To adapt to different pavement conditions, or to model the soil so that robot operators can adjust soil cohesion in real time. and shear angle Enables operation in complex terrains with mud, sand, and vegetation; Step S5.2: Simplify the robot system; Based on step S5.1, the constructed nonlinear components are used as feedback linearization components to construct a feedback system: ; in, For the proportional term of the PD controller, This is the proportionality coefficient. For the differential term of the PD controller, These are the differential coefficients; This represents the nonlinear term of the dynamics; , This is the output of the policy network; because The processing methods are difficult to observe in real-world environments, so they are either ignored or learned by imitating learning styles; Furthermore, the compensation for the nonlinear components is achieved using an optimization method, as follows: ; in, Generate functions to constrain the system state and output; At this point, the robot joint actuator retains only linear characteristics, and the model simplifies to: ; Step S6: Optimize the imitation learning strategy; The goal of imitation learning is to continuously optimize the parameters of the policy network through a series of data collected from real people, so that the robot's actions are similar to or even identical to those of real people. During the policy optimization process, the imitation learning discriminator rewards the robot's actions, and the reward results are used to adjust the actuator parameters. The linear actuator, combined with the compensation of the nonlinear part, generates the final actuator torque to control the robot's actions. In the simulation environment, the robot interacts with the environment, and the reward mechanism of imitation learning continuously updates the parameters of the policy network to achieve rapid training of the robot's imitation learning. Step S7: Deploy on real devices; The designed nonlinear compensator and the collaboratively trained policy network are used as the controller for the robot's physical machine.

Citation Information

Patent Citations

  • Mechanical arm tail end variable load dynamics self-diagnosis and compensation method in quick change mode

    CN119328746A

  • Humanoid robot imitation learning method and device, computer equipment and storage medium

    CN120116218A