Action control strategy training method and device, computer equipment and storage medium

By rewarding and optimizing the robot's motion control strategy at each time interval, combined with joint position correction, the problem of low accuracy in performing high-difficulty movements by the robot was solved, resulting in more efficient training and a lower failure rate.

CN121809578APending Publication Date: 2026-04-07SHENZHEN PUDU TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies suffer from low accuracy in training robots to perform complex stunts, making it difficult to complete demanding tasks, especially due to high failure rates caused by replication errors.

Method used

By rewarding the simulated action subset in each time period, adjusting the motion control strategy, and optimizing it at the end of each round, combined with the correction of joint position vectors, the motion control strategy is optimized using reinforcement learning algorithms until the preset conditions are met, thereby reducing the gap between the simulation environment and the real robot.

Benefits of technology

It improves the accuracy of robots in performing complex actions, reduces the failure rate, avoids action failures caused by replication errors, and improves sample utilization efficiency and learning convergence speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809578A_ABST
    Figure CN121809578A_ABST
Patent Text Reader

Abstract

The invention relates to an action control strategy training method and device, computer equipment and a storage medium. The method comprises the steps that in each time step, calculation and correction are carried out based on a joint position vector corresponding to the time step output by a current action control strategy, and a target torque vector is determined; the target torque vector is transmitted to a simulation environment so as to perform interaction between the intelligent agent and the simulation environment, simulation action sub-data corresponding to the time step of the current round are obtained, the current round comprises a plurality of continuous time periods, and each time period comprises a plurality of continuous time steps; rewarding the simulation action sub-data set in each time period in the current round so as to adjust the current action control strategy; the interaction and reward of multiple rounds are repeated until the simulation action data meet a preset end condition to obtain a target action control strategy corresponding to the task instruction, and the simulation action data are formed by simulation action sub-data of all time steps in the corresponding rounds. The method can improve the accuracy of action execution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, computer device, and storage medium for training a motion control strategy. Background Technology

[0002] Performing acrobatic maneuvers is an excellent way for multi-jointed robots to demonstrate their mobility, and it is also an important testing method and starting point for positive design when developing robot hardware. Acrobatic maneuvers typically require robots to unleash explosive power in an instant, pushing their physical or mechanical limits to complete highly difficult movements. However, this characteristic also places extremely high demands on the details of the movements. For example, movements such as front flips, back flips, and even consecutive back flips require precise coordination between multiple joints, and the robot must land accurately to avoid falls or unstable landings.

[0003] Currently, robot acrobatic trajectories are typically designed through replication (e.g., recording, writing motion sequences). However, for demanding actions, any replication error makes it difficult to complete these complex movements. Therefore, robots trained in this conventional way often exhibit low accuracy in their actions. Summary of the Invention

[0004] Therefore, it is necessary to provide a training method, device, computer equipment, and storage medium for motion control strategies that can improve the accuracy of robot motion execution, in order to address the above-mentioned technical problems.

[0005] Firstly, this application provides a method for training an action control strategy, the method comprising:

[0006] In each time step, the target torque vector is determined by calculating and correcting the joint position vector corresponding to the time step output by the current motion control strategy.

[0007] The target torque vector is transmitted to the simulation environment to enable the agent to interact with the simulation environment and obtain the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps.

[0008] Rewards are given to the simulated action subset in each time period of the current round to adjust the current action control strategy; the simulated action subset includes simulated action subsets corresponding to multiple consecutive time steps within the same time period of the current round;

[0009] The interaction and reward process is repeated for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination condition, so as to obtain the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

[0010] In some embodiments, the step of calculating and correcting the joint position vector corresponding to the time step output by the current motion control strategy to determine the target torque vector includes:

[0011] Based on the joint position vector corresponding to the time step output by the current motion control strategy, determine the target joint position vector;

[0012] Determine the joint torque vector based on the target joint position vector;

[0013] The target torque vector is determined by correcting the joint torque vector using a pre-modeled and calibrated motor torque-speed curve.

[0014] In some embodiments, rewarding a subset of simulated actions in each time period of the current round to adjust the current action control strategy includes:

[0015] Within each time period of the current round, the expected reward of the simulated action subset is determined based on the reward item corresponding to the simulated action subset.

[0016] The parameters of the current motion control strategy are adjusted based on the expected return.

[0017] In some embodiments, the reward item includes at least one of position difference item, rotation difference item, linear velocity difference item, angular velocity difference item, joint angle difference item, and joint angle difference item.

[0018] In some embodiments, determining the expected return of the simulated action subset based on the reward item corresponding to the simulated action subset includes:

[0019] Based on the reward and penalty items corresponding to the simulated action subset, the expected return of the simulated action subset is determined;

[0020] The penalty items include at least one of the following: gravity mapping penalty item, trajectory deviation penalty item, joint power penalty item, torque amplitude penalty item, joint velocity penalty item, joint acceleration penalty item, joint torque change rate penalty item, motion change rate penalty item, and joint limit penalty item; the joint limit penalty item includes a penalty item determined based on at least one of joint limit constraint, joint maximum power constraint, and joint maximum acceleration constraint.

[0021] In some embodiments, it also includes:

[0022] The target motion control strategy is deployed to several different simulation platforms for verification; the verified motion control strategy is then used as the target motion control strategy.

[0023] If the motion control strategy fails verification on any simulation platform, modify the motion control strategy and continue to verify the modified motion control strategy on the different simulation platforms until all the different simulation platforms have passed the verification.

[0024] In some embodiments, it also includes:

[0025] The target motion control strategy is deployed to the robot and optimized using an inference engine;

[0026] Based on the robot's equipment parameters, the optimized target motion control strategy is limited.

[0027] The robot's operation is controlled by a target motion control strategy that limits the amplitude.

[0028] Specifically, the target joint position vector is obtained by adjusting the joint position vector output by the target motion control strategy based on the feedback data of the operation process, and the robot operation is controlled by the target joint position vector; wherein, the feedback data includes at least one of the feedback data of the joint encoder, the feedback data of the force sensor, and the feedback data of the inertial measurement unit.

[0029] Secondly, this application provides a training device for motion control strategies, the device comprising:

[0030] The determination module is used to calculate and correct the target torque vector at each time step based on the joint position vector corresponding to the time step output by the current motion control strategy.

[0031] An interaction module is used to transmit the target torque vector to the simulation environment so that the agent can interact with the simulation environment and obtain the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps.

[0032] The reward adjustment module is used to reward the simulated action subset in each time period of the current round in order to adjust the current action control strategy; the simulated action subset includes simulated action subsets corresponding to multiple consecutive time steps in the same time period of the current round;

[0033] The training module is used to repeat the interaction and reward for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination condition, so as to obtain the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

[0034] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0035] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0036] The above-mentioned training method, apparatus, computer device, and storage medium for a motion control strategy include: in each time step, calculating and correcting a target torque vector based on the joint position vector corresponding to the time step output by the current motion control strategy; transmitting the target torque vector to a simulation environment for interaction between the agent and the simulation environment to obtain simulation motion sub-data corresponding to the time step of the current round, wherein the current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps; rewarding the simulation motion sub-data set in each time period of the current round to adjust the current motion control strategy; wherein the simulation motion sub-data set includes simulation motion sub-data corresponding to multiple consecutive time steps in the same time period of the current round; repeating the interaction and reward for multiple rounds to iteratively adjust the current motion control strategy until the simulation motion data meets a preset termination condition to obtain the target motion control strategy corresponding to the task instruction, wherein the simulation motion data is formed by the simulation motion sub-data of all time steps in the corresponding round.

[0037] Compared to the traditional approach of uniformly calculating rewards and updating the action control strategy after a complete round (i.e., after the completion of the complete action sequence), this application provides rewards and optimizes the action control strategy at the end of each time period within each round. This improves the credit allocation problem in long-running tasks and enhances sample utilization efficiency. Furthermore, since learning can begin immediately after each segment, convergence is accelerated. Additionally, because replication is not required as in conventional methods (such as recording and writing action sequences), it avoids the inability to perform complex actions due to replication errors, thus improving the accuracy of action execution. Moreover, joint position vectors are corrected at each stage, reducing the gap between the robot model in the simulation environment and the real robot. This lowers the failure rate during action execution and improves the accuracy of action execution. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 This is a diagram illustrating the application environment of a training method for an action control strategy in one embodiment.

[0040] Figure 2 This is a flowchart illustrating the training method for an action control strategy in one embodiment;

[0041] Figure 3 This is a flowchart illustrating the process of determining the target torque vector in one embodiment;

[0042] Figure 4 This is a structural block diagram of a training device for an action control strategy in one embodiment;

[0043] Figure 5 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0045] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0046] In related technologies, previously designed movements are replicated by recording and writing motion sequences, enabling the robot to reproduce the movement. However, if the replication of some movements fails, the robot cannot perform the movement. For movements requiring high skill and explosive power, every detail is critical. Furthermore, achieving a complete high-difficulty movement requires precise coordination of multiple joints; even the slightest error will prevent the entire movement from being completed. Therefore, existing training methods cannot accurately reproduce previously designed acrobatic movements, resulting in low precision in movement execution.

[0047] Therefore, in order to solve the above problems, this application proposes a training method for action control strategies.

[0048] The training method for motion control strategies provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 101 communicates with server 102 via a network. A data storage system can store the data that server 102 needs to process. The data storage system can be integrated onto server 102 or placed on a cloud or other network server. Server 102 receives task instructions (i.e., expected action data) sent by terminal 101. After receiving the task instructions, server 102 controls the agent to interact with the simulation environment according to the current action control strategy to obtain the simulation action data for the current round of simulation training. Simulation training is repeated until the simulation action data meets the preset termination conditions to obtain the target action control strategy corresponding to the task instructions. Terminal 101 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can be smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 102 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides cloud computing services.

[0049] To make the solution in this application easier to understand, some terms are explained below:

[0050] An intelligent agent may include a motion control strategy, a controller, and / or an observation module. The motion control strategy can output a corresponding motion vector (which may be a target joint position vector, target joint velocity, or joint torque, etc.) based on an input state vector (including joint angles, angular velocities, body posture, contact states, etc.). The observation module can convert acquired observation data (e.g., raw physical data from the simulation environment (joint angles, velocities, force sensors, etc.)) into the aforementioned state vector. The controller can output a corresponding torque vector based on the motion vector output by the motion control strategy. Therefore, the intelligent agent sends the motion vector output by the motion control strategy (or the joint torque vector obtained after conversion by the controller) to the simulation environment, thereby completing the interaction.

[0051] The simulation environment refers to the environment in which the intelligent agent is trained. The simulation environment can include a physics engine, scene and task models, an observation module, and / or a reward calculation module. The physics engine is used to obtain torque vectors and integrate them according to the rigid body dynamics equations, performing calculations such as collision, contact, and friction to update parameters such as the position, velocity, and contact force of all links. The scene / task model can include parameters such as the mass, inertia, joint constraints, collision geometry, ground friction, and external force field of the multi-rigid-body system. These parameters are used for calculations by the physics engine (the scene / task model can be understood as including the robot model and the scene model; the robot model includes its own parameters such as link geometry, mass, inertia, joint constraints, and motor pose; the scene model includes external environmental parameters such as the ground / obstacles / external forces / friction coefficients). The observation module is used to select the required quantities from the raw data (joint angles, angular velocities, contact forces, etc.) output by the physics engine for normalization, trimming, noise addition, and delay processing to simulate real sensors. Finally, the obtained data is assembled into a state vector for the motion control strategy. The reward calculation module can simultaneously acquire raw data (joint angles, angular velocities, contact forces, target relative pose, etc.) output from the physics engine, calculate the reward corresponding to the raw data, and then transmit the reward and new state data to the training interface of the training framework. The training interface stores the complete experience tuple (which may include, for example, the reward, the current state before the action, the action generated based on the current state, the immediate reward returned by the environment after the action, and the new state data that the environment transitions to after the action) into the agent's experience buffer. The agent then calls its internal reinforcement learning (RL) algorithm to calculate the policy gradient based on the experience tuple and update the parameters of the action control policy. The reinforcement learning algorithm can be Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), etc.

[0052] The above describes an intelligent agent and a simulation environment according to this application. Of course, the above are just examples; the structure and parameters of the intelligent agent and simulation environment can differ in different scenarios, and no limitations are imposed here. Therefore, it can be seen that the interaction process between the intelligent agent and the simulation environment provides sampling data for training robot motion control strategies within the simulation environment.

[0053] The purpose of training the motion control strategy is to enable the real robot to autonomously perform specified actions after deployment, allowing it to successfully complete the corresponding movements. Furthermore, to ensure the trained motion control strategy is well-matched to the real robot to be deployed, the robot description file corresponding to the real robot needs to be imported into the simulation environment before training to build the robot model. The robot description file includes the robot's physical structure, kinematic relationships, and visual attributes, such as high-difficulty movements (e.g., backflips, frontflips, sideflips), manipulations, and movements.

[0054] Furthermore, before training, the motion control strategy is in an initial state, and this initial state motion control strategy is used as the initial motion control strategy. Since the initial motion control strategy has not been sufficiently trained, its output control commands cannot effectively drive the robot to achieve the desired action. Therefore, to improve the control accuracy of the initial motion control strategy, it needs to be deployed to a simulation platform and repeatedly interact with the simulation environment. During this interaction, the motion state of the robot model is continuously collected, and reinforcement learning algorithms are used to update and adjust the real-time current motion control strategy based on the motion state. Training stops when the current motion control strategy can drive the robot to achieve the desired action, and the motion control strategy at the point of successful training is used as the target motion control strategy.

[0055] To make the training framework easier to understand, the following explains the training framework and its relationship with the objects mentioned above:

[0056] The training framework is an end-to-end system for the entire training process, which may include agents, simulation environments, task definitions, and / or training scheduling logic. Task definitions include reward functions, termination conditions, and target specifications; the training scheduling logic includes determining the timing of sampling, buffering, parameter updates, logging, and saving, and controlling the process flow.

[0057] Optionally, to implement the above training framework, a Markov Decision Process (MDP) can be used, i.e., M = (S, A, P, R, γ). Here, M represents the Markov Decision Process; S represents the set of all possible states (e.g., robot joint angles, velocities, body postures, contact states, etc.); A represents the set of all possible actions (e.g., motor torque, desired joint angular velocity, etc.); P(s′|s, a) represents the state transition probability, indicating the transition to a new state s′ after performing action a in state s; R(s, a) is the reward function, used to measure the immediate benefit of the robot performing action a in state s, such as maintaining a stable gait, forward speed, energy consumption penalty, etc.; γ∈(0,1) represents the discount factor, used to weigh the importance of current and future rewards.

[0058] Further, π(a|s) represents the action control policy. The agent learns π(a|s) in the MDP to maximize the cumulative expected reward. Optionally, training of the action control policy stops when the cumulative expected reward is maximized. At this point, the action control policy corresponding to maximizing the cumulative expected reward is the target action control policy. It should be noted that in practice, training of the action control policy generally stops when a preset termination condition (such as a convergence criterion or resource limit) is reached.

[0059] The foregoing has described the training objectives, training environment, and training framework of the motion control strategy in this application. The following sections will further illustrate how this application trains the motion control strategy using different embodiments.

[0060] In one exemplary embodiment, such as Figure 2 As shown, a method for training a motion control strategy is provided, which includes the following steps S110 to S410. Wherein:

[0061] Step S110: In each time step, the target torque vector is determined by calculating and correcting the joint position vector corresponding to the time step output by the current motion control strategy.

[0062] Step S210: Transmit the target torque vector to the simulation environment to enable interaction between the agent and the simulation environment, and obtain the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps.

[0063] Understandably, the training framework needs to acquire task instructions for the desired actions. These task instructions may carry or include desired action data, which refers to the desired action states that the robot should achieve, represented by target quantities in the state space (e.g., velocity, pose, trajectory points). Task instructions can be imported into the training framework via wireless transmission (including page input) or wired transmission. Training start instructions can be acquired through fixed buttons, natural language, or a human-computer interaction interface. For example, a training start instruction can be acquired in response to a training start button press, or through natural language input, such as manually entering "Start Training" via voice, or inputting the training start instruction via text in a user interface. Of course, in some cases, task instructions can also serve as the training start instruction, meaning the training framework begins training after acquiring the task instructions. Once the training framework acquires the task instructions, it determines the reward calculation module and the observation module in the simulation environment based on these instructions.

[0064] The training process, from start to finish, is divided into multiple rounds, each round representing the imitation of a desired action. As the number of training rounds increases, the agent continuously optimizes its action control strategy through reinforcement learning algorithms, reducing state deviations (such as the difference between actual and desired speeds) until a preset termination condition is met, at which point training ceases. The simulated action data in each round consists of several sub-data points corresponding to consecutive time steps. These sub-data points include control actions (such as joint torques and target positions), robot states (such as the position of the robot's center of mass in the world coordinate system, robot orientation, linear velocity, angular velocity, joint angles, joint angular velocities, and foot contact forces), and more.

[0065] Optionally, at each time step, the current motion control strategy outputs the desired joint position vector based on the observation vector containing the body state vector and the task command. The joint position vector and the current joint state (including position and velocity) are input to a proportional-derivative (PD) controller for calculation to output an initial torque vector. Considering the gap between the real robot and the simulated robot, a motor model is introduced during training to narrow the gap. Therefore, the initial torque vector output by the PD control is also input to the motor model for optimization and correction to output the target torque vector. Here, the time step determines the frequency of system information updates and environmental responses, and is defined as the minimum numerical integration interval used to advance the physical state in the simulation environment. For example, if the time step is 10 ms, the physical state change is calculated once every 10 ms.

[0066] For example, by introducing a motor model, the actual state of the motor can be simulated to obtain a more accurate target torque vector. For instance, the damping characteristics of the motor can be simulated, such as calculating the torque compensation value corresponding to the damping loss of the motor based on the current joint speed, and performing reverse damping correction on the initial torque vector to match the energy loss characteristics of the actual motor during operation. The response delay characteristics of the motor can also be simulated, such as adding delay compensation that conforms to the response curve of the actual motor to the initial torque vector based on the time step interval parameter, so that the timing of the torque output matches the response rhythm of the actual motor.

[0067] Furthermore, the total time segment corresponding to the simulated motion data in each round is divided into several consecutive time segments, each of which includes multiple consecutive time steps. Since each time step corresponds to a simulated motion sub-data, each time segment corresponds to multiple consecutive simulated motion sub-data.

[0068] Step S310: Reward the simulated action subset in each time period of the current round to adjust the current action control strategy. The simulated action subset includes simulated action subsets corresponding to multiple consecutive time steps within the same time period of the current round.

[0069] Step S410: Repeat the interaction and reward for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination condition to obtain the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

[0070] Understandably, the number of time steps in each time period can be the same or different; there is no restriction here. At the end of each time period, a reward is given and the motion control strategy is adjusted. The next time period is then simulated based on the adjusted motion control strategy. Each time period corresponds to a subset of simulated motion data. When the simulated motion data meets the preset termination conditions, the resulting motion control strategy is the target motion control strategy, and this target motion control strategy can drive the robot to perform the target action corresponding to the task instruction.

[0071] Optionally, the set of simulation action sub-data for multiple consecutive time steps corresponding to each time period can be used as the simulation action sub-data set.

[0072] Optionally, the set of simulation action sub-data for multiple consecutive time steps corresponding to each time period is called the simulation action sub-data subset. This simulation action sub-data subset is a proper subset of the simulation action dataset, that is, in addition to including the simulation action sub-data subset, the simulation action dataset also includes simulation action sub-data corresponding to several other time steps.

[0073] For example, if a time period corresponds to 50ms and a time step of 10ms, then the simulation action subset corresponding to the time period can be composed of simulation action subsets corresponding to the 5 time steps within the current time period. The simulation action subset corresponding to the time period can also include, in addition to the simulation action subsets corresponding to the 5 time steps within the current time period, several simulation action subsets corresponding to the time steps within the previous d time periods, where d is a positive integer, such as 1, 2, 3, 4, 5, etc.

[0074] Optionally, the preset termination condition can be that the performance of the simulated action data meets the standard, such as the average cumulative reward of N consecutive evaluation rounds being greater than or equal to the preset average expected return, the proportion of successfully completing the task in M ​​tests being greater than or equal to the preset task completion proportion threshold, the state tracking error being less than or equal to the preset state tracking error threshold, or the average reward change in the historical i rounds being less than or equal to the preset average reward change threshold, etc.

[0075] In this embodiment, in order to enable the motion control strategy to drive the robot to achieve the target action (including highly difficult acrobatic actions) more accurately, the following two methods are used:

[0076] The first approach, compared to the traditional method of uniformly calculating rewards and updating the action control strategy after a complete round (i.e., after the complete action sequence ends), rewards are given and the action control strategy is optimized and adjusted at the end of each time period within each round. This improves the credit allocation problem in long-duration tasks, increases sample utilization efficiency, and accelerates convergence because learning can begin at the end of each segment. Furthermore, since a complete action in a round includes multiple sub-actions, even if the landing fails, the successful experiences of the previous sub-actions can still be used to update the strategy, greatly effectively utilizing some successful trajectories for training. Additionally, since a complete high-difficulty action includes multiple stages, each with different characteristics, this approach can use differentiated rewards for different stages, improving the accuracy of action execution while supporting complex multimodal skills. Moreover, since there is no need for replication like conventional methods (such as recording and writing action sequences), it avoids the phenomenon of being unable to complete high-difficulty actions due to replication errors, further improving the accuracy of action execution.

[0077] The second approach involves calculating and correcting the joint position vectors at each stage. This is because it addresses the difference between the robot model in the simulation environment and the real robot. In a real robot, each joint has a maximum torque; if the torque is too large, it will not only damage the robot structure but also prevent the robot from performing the intended action. Therefore, it is necessary to consider the optimal torque for each joint, optimize and correct the torque vector input to the simulation environment, and then transmit the optimized and corrected target torque vector back to the simulation environment for simulation. This reduces the failure rate of action execution and improves the accuracy of action execution.

[0078] In some embodiments, the trajectory data of the desired action can be calculated using trajectory optimization (TO). Different methods can be used for dynamic modeling of different actions. Taking the front somersault of a quadruped robot as an example, since this action is strictly symmetrical from left to right, the robot can be simplified to dynamic modeling in the sagittal plane, thereby halving the degrees of freedom, making the calculation faster while maintaining the calculation accuracy.

[0079] Furthermore, impact dynamics is used to model the state transitions between adjacent dynamic stages, defining the connection conditions between stages. All dynamic stages are sequentially assembled to form a complete motion trajectory. Continuous-time dynamic equations (typically ordinary differential equations) are discretized into finite-dimensional discrete dynamic constraints using numerical transcription methods (such as the trapezoidal rule or collocation method). Based on this, a nonlinear programming (NLP) problem is constructed, incorporating a cost function with task objectives and physical / kinematic constraints. Motion priors are introduced to impose state constraints (such as end-effector pose, centroid trajectory, or contact configuration) on keyframes of each dynamic stage. This nonlinear programming problem is solved using a nonlinear optimization solver (such as IPOPT, SNOPT, or the CasaADi backend) to obtain the optimal discrete state-control sequence. The continuous-time motion trajectory is then reconstructed through interpolation or integration. The continuous motion trajectory includes the position, orientation (rotation), linear velocity, and angular velocity of all links in the world coordinate system, as well as the joint position and joint velocity of all joints.

[0080] This embodiment uses data derived from trajectory optimization as a reference trajectory to replicate various acrobatic maneuvers. Trajectory optimization can solve for the optimal open-loop maneuver in a deterministic model (difficult but deterministic); while reinforcement learning can learn robust closed-loop tracking in uncertain environments (flexible but requires guidance). The combination of these two approaches ensures both high-quality and feasible maneuvers, while also achieving robustness and adaptability on real-world testing.

[0081] In some embodiments, such as Figure 3 As shown, each time step in step S110 above includes the following steps S111 to S113.

[0082] Step S111: Determine the target joint position vector based on the joint position vector corresponding to the time step output by the current motion control strategy.

[0083] Understandably, the joint position vectors output by the current motion control strategy are initial joint position vectors and need to be scaled to obtain the target joint position vectors. This is because the neural network of the motion control strategy is sensitive to the numerical range during training. If the actual joint angles are directly output, the gradients become unstable due to large differences in joint scale. Therefore, the parameter range of the strategy network output is a normalized vector. Then, through scaling coefficients and offsets, the normalized output is transformed into the physically feasible target joint position corresponding to each joint.

[0084] Step S112: Determine the joint torque vector based on the target joint position vector.

[0085] Step S113: Based on the joint torque vector, the torque-speed curve is corrected to determine the target torque vector.

[0086] Understandably, the controller calculates parameters such as the target joint position and outputs an initial torque vector. This initial torque vector, along with the joint angular velocity, is then corrected using a pre-modeled and calibrated motor torque-speed curve. The torque-speed curve represents the maximum output torque the actuator can provide at different motor speeds. Correction via the torque-speed curve includes the following two scenarios: 1) When the joint torque output by the PD controller is greater than the maximum output torque the motor can actually provide at the current joint angular velocity, the joint torque is corrected to the maximum output torque at or below that at the corresponding angular velocity. This prevents excessive torque from hindering the completion of the target action and causing damage to the machine structure when the motion control strategy is deployed to the actual robot; 2) When the joint torque output by the PD controller is less than the maximum output torque the motor can actually provide at the current joint angular velocity, the torque-speed curve can be used, combined with the motor's actual operating parameters, to correct the torque value to better match the motor's actual output capability, resulting in higher accuracy of the output torque while ensuring the stability of the motion execution.

[0087] For example, if the maximum output torque is 50 N·m, and the joint torque output by the PD controller is 52 N·m, then 52 N·m needs to be corrected to 50 N·m or less to avoid damage or inability to complete the action after deployment to the actual robot.

[0088] Alternatively, the torque-speed curve of the motor can be represented by the following formula:

[0089] ;

[0090] Where T(ω) represents the maximum output torque of the motor at the joint angular velocity ω. This indicates the maximum static torque of the motor. This indicates the maximum no-load speed of the motor.

[0091] It is understandable that the duration of motor use has a significant impact on damping. As the duration of motor use increases, wear occurs in the internal friction components, and the performance of the lubricating medium deteriorates, directly leading to a gradual increase in the motor's damping coefficient and a corresponding increase in damping losses. Furthermore, changes in motor temperature also affect the motor's damping coefficient by altering the viscosity of the lubricating medium and the degree of thermal expansion and contraction of components. Therefore, in other embodiments, the damping coefficient C of the actual motor can be pre-calibrated. m(t) The dynamic change curve of (t) with running time (or motor temperature) is used to calculate the torque compensation value corresponding to the damping loss based on the current joint speed and real-time damping coefficient, and the initial torque vector is corrected in reverse. The specific calculation formula is as follows:

[0092] ;

[0093] Where, τ damp (t) is the torque compensation vector corresponding to the motor damping characteristics at time step t; C m (t) represents the real-time damping coefficient of the motor at time step t; act τ is the actual velocity vector of the robot joint at the current time step; init τ is the initial torque vector output by the PD controller. temp This is the intermediate moment vector after damping correction. The intermediate moment vector can be further corrected by considering other influencing factors to obtain the target moment vector. If other influencing factors are not considered, the intermediate moment vector can be directly used as the target moment vector.

[0094] In some embodiments, considering the various environmental disturbances that occur during the actual robot's actions, such as frictional losses between joints and displacement of the link's center of mass, a domain randomization mechanism needs to be introduced to enhance the motion control strategy's ability to migrate between high and low altitudes in real-world environments. The randomization parameters include, but are not limited to: robot-specific parameters (e.g., joint friction coefficient, link mass, link center of mass), environmental parameters (e.g., ground friction coefficient, terrain morphology, external disturbance forces), and sensor parameters (e.g., noise amplitude, latency, packet loss rate).

[0095] Furthermore, to express this mathematically, domain randomization can be viewed as a parameterized perturbation of the environment transition probabilities during training:

[0096] ;

[0097] in, Indicates the probability of environmental transition; Represents a set of environmental parameters. The parameter distribution is represented by a uniform or Gaussian distribution. This method ensures that the agent's learned policy π(a|s) maintains high performance in diverse environments, thereby improving the ability to transfer stunt actions from simulation to reality.

[0098] In some embodiments, step S310 includes the following steps:

[0099] Within each time period of the current round, the expected reward of the simulated action subset is determined based on the reward item corresponding to the simulated action subset.

[0100] Adjust the parameters of the current motion control strategy based on the expected return.

[0101] Optionally, each time step's simulation action sub-data corresponds to multiple reward items. The target reward value for each time step's simulation action sub-data can be determined by weighted summation of the reward values ​​of the corresponding reward items. Based on the target reward values ​​of all simulation action sub-data in the simulation action sub-data set, the expected return of the simulation action sub-data set can be determined. The number of reward items corresponding to each time step's simulation action sub-data can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or more; there is no limitation here.

[0102] For example, assuming the reward items corresponding to the simulated action sub-data include reward item A, reward item B, reward item C, and reward item D, then the reward value of reward item A can be determined as A', the reward value of reward item B as B', the reward value of reward item C as C', and the reward value of reward item D as D'. The weight corresponding to reward item A is A1, the weight corresponding to reward item B is B1, the weight corresponding to reward item C is C1, and the weight corresponding to reward item D is D1. Then, the target reward value R of the simulated action sub-data can be calculated using the following formula: .

[0103] As an example, the formula for calculating the expected return of a subset of simulated actions can be the following:

[0104] ;

[0105] Where J(π) represents the expected reward of the action control policy; π represents the action control policy; E τ-π τ represents the expected value; τ represents the trajectory over a time period; τ∼π indicates that the trajectory is generated by policy π; t represents the time step; S t Indicates state; a t Indicates an action; R (S) t a t ) indicates that in state S t Perform action a t The reward after the period; γ represents the discount factor; T represents the total number of time steps in the time period.

[0106] Optionally, the reward items include at least one of the following: position difference item, rotation difference item, linear velocity difference item, angular velocity difference item, joint angle difference item, and joint angle difference item.

[0107] In some embodiments, in addition to the traditional motion stability reward, a trajectory tracking-based reward can be additionally designed to constrain the robot to replicate the target action obtained from trajectory optimization. The difference terms used in calculating the reward term may include at least one of the following: link position difference term, link rotation difference term, link linear velocity difference term, link angular velocity difference term, joint angle difference term, and joint angular velocity difference term.

[0108] The link position difference term refers to the difference between the actual spatial position of the link in the robot model's motion during the simulation environment and the target spatial position of the corresponding link in the target trajectory obtained through trajectory optimization. This can be calculated based on coordinate values ​​in a three-dimensional coordinate system (the square of the difference can be used to represent the degree of difference), and is used to measure the accuracy of the link's position replication in space during the simulation environment. The link rotation difference term refers to the difference between the actual attitude angle of the link during the robot model's motion in the simulation environment and the target attitude angle of the corresponding link in the target trajectory. This can be calculated based on Euler angles or quaternions, and is used to measure the accuracy of the link's spatial attitude replication in the simulation environment. The link linear velocity difference term refers to the difference between the actual linear velocity of the link during the robot model's motion in the simulation environment and the target linear velocity of the corresponding link in the target trajectory. Linear velocity refers to the moving speed of a specified feature point on the link (generally the end point or center of mass of the link), and can be calculated using the square of the velocity difference, used to measure the accuracy of the velocity replication during the link's movement in the simulation environment. The link angular velocity difference term refers to the difference between the actual angular velocity of a link during the robot model's motion in the simulation environment and the target angular velocity of the corresponding link in the target trajectory. Angular velocity is the speed at which a link rotates around its own axis of rotation, and can be calculated as the square of the angular velocity difference. It is used to measure the accuracy of velocity replication during link rotation in the simulation environment. The joint angle difference term refers to the difference between the actual angle of each joint during the robot model's motion in the simulation environment and the target angle of the corresponding joint in the target trajectory. Joints are the connecting parts between the robot's links, and this term is used to measure the accuracy of joint angle replication in the simulation environment. The joint angular velocity difference term refers to the difference between the actual angular velocity of each joint during the robot model's motion in the simulation environment and the target angular velocity of the corresponding joint in the target trajectory. It is used to measure the accuracy of velocity replication during joint rotation in the simulation environment.

[0109] In some embodiments, the expected reward of the simulated action subset is determined based on the reward item corresponding to the simulated action subset, including:

[0110] Based on the reward and penalty items corresponding to the simulated action subset, the expected return of the simulated action subset is determined.

[0111] The penalty items include at least one of the following: gravity mapping penalty, trajectory deviation penalty, joint power penalty, torque amplitude penalty, joint velocity penalty, joint acceleration penalty, joint torque change rate penalty, motion change rate penalty, and joint limit penalty. The joint limit penalty includes a penalty item determined based on at least one of joint limit constraints, joint maximum power constraints, and joint maximum acceleration constraints. All penalty items determined based on at least one of these constraints belong to the aforementioned joint limit penalty items. Therefore, the number of penalty items determined by combining these three constraints can be multiple and is not limited here.

[0112] Optionally, the robot's gravity mapping direction differs for different acrobatic maneuvers. Therefore, prior knowledge can be used to specify constraints on the gravity mapping direction for different maneuvers. The gravity mapping direction refers to associating and mapping the direction of gravity with the robot's maneuver. The gravity mapping penalty term may include at least one of the following penalty terms:

[0113] Robot tipping over penalty: During a backflip or frontflip, as the body rotates around the x-axis, This represents the penalty imposed on the robot's gravity mapping in the y-direction at time step t, i.e., the robot's rollover penalty.

[0114] Robot flipping penalty: During a side somersault, primarily during the rotation around the y-axis, This represents the gravity mapping of the robot in the x-direction at time step t, i.e., the robot's right-side-over penalty.

[0115] Trajectory away from penalty: ,in, To avoid the penalty for the trajectory, Let t represent the positions of all links in the world coordinate system at time step t of the robot. This penalty is applied to the positions of all links in the world coordinate system at time step t in the trajectory optimization data. The penalty increases significantly when the robot's trajectory deviates greatly from the optimized data.

[0116] Optionally, when the robot performs complex stunts, in order to ensure that the performance of the motors and other hardware can cover the requirements of extreme actions such as somersaults, and at the same time avoid excessive energy consumption and unstable control, this application also introduces energy and motion smoothness constraints, including at least one of the following penalty terms: joint power penalty, torque amplitude penalty, joint velocity penalty, joint acceleration penalty, joint torque change rate penalty, and motion change rate penalty. The calculation method for each penalty term is to accumulate the penalty values ​​of the corresponding joints.

[0117] The joint power penalty term constrains the power output of the joint motors in the robot model within the simulation environment. When the actual power output of a joint exceeds a threshold, a penalty is triggered to prevent the motor from exceeding its hardware performance limits and to control overall energy consumption. The torque amplitude penalty term constrains the torque output of the robot model's joints within the simulation environment. When the amplitude of the actual torque output exceeds the upper limit, a penalty is triggered to prevent excessive torque from causing hardware damage to the actual robot and to ensure motion stability. The joint speed penalty term constrains the rotational speed of the robot model's joints within the simulation environment. When the actual rotational speed exceeds the upper limit, a penalty is triggered to prevent control instability caused by excessively fast joint rotation and to reduce unnecessary energy consumption. The joint acceleration penalty term constrains the rotational acceleration of the robot model's joints within the simulation environment. When the actual rotational acceleration exceeds the upper limit, a penalty is triggered to prevent overly drastic changes in the joint's rotational state and to ensure smooth motion. The joint torque change rate penalty term constrains the rate of change of joint torque in the robot model within the simulation environment. When the rate of torque change exceeds the upper limit, a penalty is triggered to ensure smoother output changes in joint torque, avoid abrupt torque changes, and improve control stability. The motion change rate penalty term constrains the overall rhythm of the robot's motion within the simulation environment. When the robot's motion transitions from one state to another too quickly, a penalty is triggered to ensure the overall smoothness and stability of complex acrobatic movements such as somersaults, avoiding control problems caused by abrupt motion transitions.

[0118] Optionally, to prevent the motion control strategy from outputting excessive motion commands that could damage the hardware, the robot's motion range can be further constrained by adding at least one of the following joint limit penalties during training:

[0119] ;

[0120] Where, q i This represents the velocity of the i-th joint. This represents the acceleration of the i-th joint. This represents the torque of the i-th joint. In other words, if the action output by the control strategy does not meet the above constraints, a joint limit penalty is added to the reward function; the greater the deviation, the heavier the penalty.

[0121] Furthermore, in addition to penalizing joint limitation as described above, penalties are also imposed on joint power and / or joint acceleration. The penalties for joint power and joint acceleration are as follows:

[0122] ;

[0123] in, This represents the maximum power of the i-th joint. This represents the maximum acceleration of the i-th joint. In other words, if the action output by the control strategy does not meet the above constraints, a penalty for joint power and joint acceleration is added to the reward function; the greater the deviation, the heavier the penalty.

[0124] Considering that high-difficulty acrobatic maneuvers (such as somersaults, jumps followed by turns) are highly nonlinear, multi-stage coupled, and highly dynamically sensitive, it is often difficult to achieve stable convergence if the strategy is directly driven from scratch by sparse or general reward functions, and it is extremely sensitive to reward structure and hyperparameters.

[0125] To address this, this embodiment introduces a pre-calculated reference motion trajectory (e.g., generated through offline trajectory optimization, motion capture data, or physical simulation) and transforms the control objective into a high-precision tracking task of this trajectory. Based on this, only a few simple, physically meaningful basic reward and penalty terms (such as joint position tracking error, posture stability, energy consumption suppression, and motion smoothness) need to be designed to effectively guide the policy to learn complex acrobatic behaviors. Since the reward function no longer needs to explicitly encode the stage-specific semantics or fine-grained dynamic constraints of the acrobatic actions, its design complexity and parameter tuning burden are significantly reduced, which is beneficial for the deployment and transfer of the policy on real robots.

[0126] In some embodiments, verification may also be performed using different simulation platforms.

[0127] The target motion control strategy is deployed to several different simulation platforms for verification; the verified motion control strategy is then used as the target motion control strategy.

[0128] If the motion control strategy fails verification on any simulation platform, modify the strategy and continue to verify it on several different simulation platforms until all of them pass verification.

[0129] Optionally, the Mujoco simulation platform's physics engine offers higher speed and accuracy, facilitating rapid verification of dynamic consistency. By importing models such as URDF or MJCF that are consistent with real robots and connecting the trained target motion control strategy to Mujoco's control interface, the stability and trajectory tracking performance of the target motion control strategy under a superior physics engine simulation platform can be verified.

[0130] Alternatively, due to its deep integration with ROS, the Gazebo simulation platform is suitable for testing in software stacks that closely resemble real-world robot applications. By loading the robot model (URDF / SDF format) into Gazebo and mapping the policy output to joint control commands, evaluation and verification can be performed in more complex simulation environments with sensor noise, communication delays, and other factors.

[0131] Understandably, after verification on multiple simulation platforms, it is shown that the motion control strategy meets the requirements of multiple simulation platforms, which is the final target motion control strategy.

[0132] It should be noted that if the verification fails on any simulation platform, the reward, penalty, or other parameters in the motion control strategy need to be modified, and the verification should be performed again on multiple simulation platforms until all simulation platforms pass the verification.

[0133] In some embodiments, considering the inherent limitations of a single simulation platform in terms of physical modeling accuracy and sensor and actuator abstraction granularity, and the significant differences between different simulation platforms (such as MuJoCo and Gazebo) in dynamics engine implementation, contact force models, friction handling, and numerical integration methods, relying solely on a single system for training is insufficient to fully cover the system-level uncertainties faced by real robots. To effectively narrow the sim-to-real gap, this solution adopts a multi-simulation platform joint training strategy: that is, training the strategy simultaneously or alternately in multiple heterogeneous simulation environments (such as MuJoCo based on rigid body precise modeling and Gazebo based on the ROS ecosystem and supporting complex sensor simulation), enabling it to learn robust control behavior under diverse virtual physical conditions, thereby significantly improving the generalization ability and deployment success rate of the strategy on real hardware.

[0134] In some embodiments, it also includes:

[0135] The target action control strategy is deployed and trained on several different simulation platforms. The action control strategy trained on the last simulation platform is then used as the target action control strategy.

[0136] Understandably, training on several different simulation platforms can be achieved through cross-training and synchronous training. This allows for the early detection of potential failure modes of the strategy under different physical modeling and simulation conditions, while also enabling learning and training under diverse virtual physical conditions.

[0137] Cross-training refers to alternating training across multiple simulation platforms. For example, if training only needs to be performed on two simulation platforms, the training should be completed on the first platform, then on the second, then back to the first platform, and so on, until at least the training stopping condition is met. The target action control policy trained on one simulation platform becomes the initial action control policy for training on the next platform.

[0138] Synchronous training involves training on multiple simulation platforms using the same motion control strategy and sampling data from different platforms in parallel to jointly update the motion control strategy. Specifically, this application simultaneously collects multiple corresponding time-segment datasets of simulated actions from different simulation platforms, updates the current motion control strategy across multiple platforms based on these datasets, and continues simulation according to the updated strategy. The number of simulated action datasets is the same as the number of simulation platforms. The entire time period of the simulated action data in each training round for each simulation platform is divided into an equal number of time segments, each time segment is ordered chronologically, and time segments from the same round and sequence across different simulation platforms are considered as corresponding time segments.

[0139] In some embodiments, following the target action control strategy obtained above, the following further steps are included:

[0140] The target motion control strategy is deployed to the robot and optimized using an inference engine.

[0141] Based on the robot's equipment parameters, the optimized target motion control strategy is limited.

[0142] The robot's operation is controlled by a target motion control strategy that limits the amplitude.

[0143] Understandably, the motion control strategy, in the form of neural network parameters, undergoes serialization, graph optimization, and hardware acceleration by a high-performance inference engine (such as PyTorchJIT, ONNX Runtime, or TensorRT) to generate a lightweight inference model. This model is optimized for deployment on embedded computing platforms (such as NVIDIA Jetson, Intel NUC, or custom robot controllers) to achieve a certain end-to-end inference latency, meeting the stringent real-time requirements of highly dynamic robots. The latency can range from 0.1 to 3 ms, but other values ​​such as 4, 5, 6, 7, 8, 9, and 10 ms are also possible and not limited here.

[0144] After deployment, the raw motion signals output by the motion control strategy (usually normalized joint target values) need to be converted into control commands acceptable to the underlying drive system by the motion decoding module.

[0145] Optionally, depending on the type of robot actuator (such as a servo motor, frameless torque motor, or hydraulic cylinder) and the drive capability, the system can support the following three control modes:

[0146] Position control mode: Outputs the target joint angle, suitable for high reduction ratio servo systems;

[0147] Speed ​​control mode: Outputs target angular velocity, suitable for medium and low impedance drive scenarios;

[0148] Torque control mode: Outputs target joint torque, suitable for high bandwidth, low transmission ratio direct drive or SEA systems.

[0149] The control mode selection is automatically matched by the Hardware Abstraction Layer (HAL) to ensure that the target motion control strategy and the actuator work together.

[0150] Furthermore, after deployment, in order to ensure the safe operation of the real robot, it is necessary to limit and verify the output of the motion control strategy before it is output to the actuator. The parameters for limiting and verifying the safety include joint angle, speed, torque, etc.

[0151] Optionally, if the motion control strategy output continuously exceeds the limit, sensor data is abnormal, or communication is interrupted, the safety monitoring module will immediately trigger degrade control (such as switching to impedance holding) or emergency stop, cut off the motor enable, and prevent mechanical overload or damage.

[0152] During actual robot operation, the robot's inertial measurement unit (IMU), joint encoders, and foot / joint force sensors collect the robot's body state in real time (including body posture, angular velocity, linear acceleration, joint position / velocity, contact force, etc.) and input it into the motion control strategy, forming a closed-loop control circuit. This closed-loop mechanism enables the strategy to dynamically respond to external disturbances (such as pushing and terrain changes) and internal uncertainties (such as friction changes and battery voltage fluctuations), significantly improving the robustness of the real robot and the success rate of tasks.

[0153] The advantages of this application include:

[0154] 1. By establishing and optimizing strategies in stages, the credit allocation problem for long-duration tasks is improved. Since learning can begin after each stage, sample utilization efficiency is also enhanced. Furthermore, even if a landing fails, the action experience from previous time periods can still be reused, thus increasing the utilization rate of some successful trajectories. Different rewards can be used for each stage, improving the accuracy of action execution while supporting complex multimodal skills. Simultaneously, since there is no need for replication like conventional methods (such as recording and writing action sequences), the inability to complete high-difficulty actions due to replication errors is avoided, further improving the accuracy of action execution.

[0155] 2. This application takes into account the gap between simulation and reality. Since the torque output directly may exceed the hardware's tolerance, the gap between the simulated robot and the real robot can be narrowed by torque vector calculation and correction, so as to avoid excessive torque damaging the hardware and avoid motion failure.

[0156] 3. High-difficulty acrobatic maneuvers (such as somersaults) are highly nonlinear, making it difficult for traditional sparse rewards to achieve stable convergence. This application reduces the complexity of reward function design by introducing basic reward or penalty terms such as trajectory tracking reward functions and energy constraints during training (eliminating the need to explicitly encode the semantic constraints of acrobatic maneuvers), balancing multiple objectives such as action accuracy, energy consumption, and hardware security, thus making policy learning more efficient.

[0157] 4. This application further introduces an optimal open-loop motion reference trajectory generated based on the trajectory optimization (TO) method, and embeds it as a motion prior into the reinforcement learning training framework. On this basis, through end-to-end policy learning, the agent can reproduce complex acrobatic movements (such as somersaults, jumps, and twists) with high precision.

[0158] This hybrid paradigm effectively combines the global optimality and feasibility guarantee of trajectory optimization with the closed-loop robustness and environmental adaptability of reinforcement learning, thereby significantly improving the deployment success rate and operational robustness of the strategy on real robot platforms while ensuring high-quality execution of actions.

[0159] 5. This application has high applicability and can be applied to the replication of motion trajectories obtained by motion capture or trajectory optimization for any robot (e.g., humanoid robot, fixed-base collaborative arm, quadruped robot, etc.).

[0160] To make it easier to understand, this application also provides a more specific embodiment:

[0161] In each time step, the target joint position vector is determined based on the joint position vector corresponding to the time step output by the current motion control strategy; the joint torque vector is determined based on the target joint position vector; and the target torque vector is determined by correcting the joint torque vector through the torque-speed curve.

[0162] The target torque vector is transmitted to the simulation environment to enable interaction between the agent and the simulation environment, thereby obtaining the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps.

[0163] Within each time period of the current round, the expected reward of the simulated action subset is determined based on the reward item corresponding to the simulated action subset.

[0164] The parameters of the current motion control strategy are adjusted according to the expected return; the simulation motion subset includes simulation motion subsets corresponding to multiple consecutive time steps within the same time period of the current round.

[0165] Repeated interactions and rewards for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination conditions, thereby obtaining the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

[0166] The target motion control strategy is deployed to several different simulation platforms for verification; the verified motion control strategy is then used as the target motion control strategy.

[0167] If the motion control strategy fails verification on any simulation platform, modify the strategy and continue to verify it on several different simulation platforms until all of them pass verification.

[0168] The aforementioned target action control strategy is deployed to the robot and optimized using an inference engine.

[0169] Based on the robot's equipment parameters, the optimized target motion control strategy is limited.

[0170] The robot's operation is controlled by a target motion control strategy that limits the amplitude.

[0171] Specifically, the target joint position vector is obtained by adjusting the joint position vector output by the target motion control strategy based on the feedback data of the operation process, and the robot operation is controlled by the target joint position vector; wherein, the feedback data includes at least one of the feedback data of the joint encoder, the feedback data of the force sensor and the feedback data of the inertial measurement unit.

[0172] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0173] Based on the same inventive concept, such as Figure 4 As described above, another embodiment of this application also provides a training device 500 for implementing the motion control strategy involved above, the training device 500 comprising:

[0174] The determination module 510 is used to calculate and correct the target torque vector at each time step based on the joint position vector corresponding to the time step output by the current motion control strategy.

[0175] The interaction module 520 is used to transmit the target torque vector to the simulation environment for interaction between the agent and the simulation environment, and to obtain the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps.

[0176] The reward adjustment module 530 is used to reward the simulated action subset in each time period of the current round in order to adjust the current action control strategy; the simulated action subset includes the simulated action subset data corresponding to multiple consecutive time steps in the same time period of the current round.

[0177] The training module 540 is used to repeat the interaction and reward for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination condition, so as to obtain the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

[0178] Each module in the training device for the aforementioned motion control strategy can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0179] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data for the training method of the aforementioned motion control strategy. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a training method for a motion control strategy.

[0180] In an exemplary embodiment, a computer device is provided, which may be a terminal. The computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a training method for an action control strategy. The display unit of the computer device is used to form a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0181] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0182] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0183] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described above.

[0184] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0185] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0186] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0187] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A training method for an action control strategy, characterized in that, The method includes: In each time step, the target torque vector is determined by calculating and correcting the joint position vector corresponding to the time step output by the current motion control strategy. The target torque vector is transmitted to the simulation environment to enable the agent to interact with the simulation environment and obtain the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps. Rewards are given to the simulated action subset in each time period of the current round to adjust the current action control strategy; the simulated action subset includes simulated action subsets corresponding to multiple consecutive time steps within the same time period of the current round; The interaction and reward process is repeated for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination condition, so as to obtain the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

2. The method according to claim 1, characterized in that, The joint position vector corresponding to the time step output by the current motion control strategy is calculated and corrected to determine the target torque vector, including: Based on the joint position vector corresponding to the time step output by the current motion control strategy, determine the target joint position vector; Determine the joint torque vector based on the target joint position vector; The target torque vector is determined by correcting the joint torque vector using a pre-modeled and calibrated motor torque-speed curve.

3. The method according to claim 1, characterized in that, The step of rewarding a subset of simulated actions within each time period of the current round to adjust the current action control strategy includes: Within each time period of the current round, the expected reward of the simulated action subset is determined based on the reward item corresponding to the simulated action subset. The parameters of the current motion control strategy are adjusted based on the expected return.

4. The method according to claim 3, characterized in that, The reward items include at least one of the following: position difference item, rotation difference item, linear velocity difference item, angular velocity difference item, joint angle difference item, and joint angle difference item.

5. The method according to claim 3, characterized in that, The determination of the expected return of the simulated action subset based on the reward item corresponding to the simulated action subset includes: Based on the reward and penalty items corresponding to the simulated action subset, the expected return of the simulated action subset is determined; The penalty items include at least one of the following: gravity mapping penalty item, trajectory deviation penalty item, joint power penalty item, torque amplitude penalty item, joint velocity penalty item, joint acceleration penalty item, joint torque change rate penalty item, motion change rate penalty item, and joint limit penalty item; the joint limit penalty item includes a penalty item determined based on at least one of joint limit constraint, joint maximum power constraint, and joint maximum acceleration constraint.

6. The method according to claim 1, characterized in that, Also includes: The target motion control strategy was deployed to several different simulation platforms for verification. The action control strategy that passes the verification will be used as the target action control strategy. If the motion control strategy fails verification on any simulation platform, modify the motion control strategy and continue to verify the modified motion control strategy on the different simulation platforms until all the different simulation platforms have passed the verification.

7. The method according to claim 1, characterized in that, Also includes: The target motion control strategy is deployed to the robot and optimized using an inference engine; Based on the robot's equipment parameters, the optimized target motion control strategy is limited. The robot's operation is controlled by a target motion control strategy that limits the amplitude. Specifically, the target joint position vector is obtained by adjusting the joint position vector output by the target motion control strategy based on the feedback data of the operation process, and the robot operation is controlled by the target joint position vector; wherein, the feedback data includes at least one of the feedback data of the joint encoder, the feedback data of the force sensor, and the feedback data of the inertial measurement unit.

8. A training device for a motion control strategy, characterized in that, The device includes: The determination module is used to calculate and correct the target torque vector at each time step based on the joint position vector corresponding to the time step output by the current motion control strategy. An interaction module is used to transmit the target torque vector to the simulation environment so that the agent can interact with the simulation environment and obtain the simulation action sub-data corresponding to the time step of the current round. The current round includes multiple consecutive time periods, and each time period includes multiple consecutive time steps. The reward adjustment module is used to reward the simulated action subset in each time period of the current round in order to adjust the current action control strategy; the simulated action subset includes simulated action subsets corresponding to multiple consecutive time steps in the same time period of the current round; The training module is used to repeat the interaction and reward for multiple rounds to iteratively adjust the current action control strategy until the simulated action data meets the preset termination condition, so as to obtain the target action control strategy corresponding to the task instruction. The simulated action data is formed by the simulated action sub-data of all time steps in the corresponding round.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • A strategy network training method and device, a robot control method and a robot

    CN122287768A