Robot motion control method and equipment

By adopting a multi-task training method based on reinforcement learning algorithms in humanoid robots, the problems of poor path planning effectiveness and insufficient motion stability in the prior art are solved, and more efficient and stable robot motion control is achieved.

CN120206522APending Publication Date: 2025-06-27人形机器人(上海)有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510435007.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the motion path planning of humanoid robots is poor, and the stability of the robot cannot be guaranteed.

Method used

A robot motion control method based on reinforcement learning algorithm is adopted to carry out path planning, gait control and balance control tasks together, and a motion control model is obtained through multi-task training.

Benefits of technology

It improves the efficiency and stability of robot motion control, allowing robots to efficiently and reasonably plan paths and maintain stable balance in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120206522A_ABST
    Figure CN120206522A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a robot motion control method and device.The method comprises the steps that the current state of a robot is obtained, the current state comprises the current position and the target position of the robot, the current state is input into a motion control model, and the current action needing to be executed by the robot is obtained; the motion control model is obtained by performing multi-task training on the reinforcement learning model, the multi-task training comprises path planning task training and at least one of gait control task training and balance control task training, and the robot is controlled to execute the current action. And if the robot reaches the target position after completing the current action, completing the current motion task. According to the method provided by the embodiment of the invention, the robot can improve the stability on the basis of efficiently and reasonably planning the path.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202510073211.6, and the filing date of the original application is January 17, 2025. The entire content of the original application is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the technical field of humanoid robots, and in particular to a robot motion control method and device. Background Art

[0003] With the rapid development of technology, humanoid robots have shown great application potential in many fields. In order to improve the performance of humanoid robots, it is necessary to focus on improving the motion control level of humanoid robots.

[0004] In related technologies, during the motion of a humanoid robot, path planning is usually carried out by means of geometric algorithms or search algorithms.

[0005] However, in the process of implementing this application, the inventors found that there are at least the following problems in the prior art: the effectiveness of the above path planning method is poor, and the stability of the robot during motion cannot be guaranteed. Summary of the Invention

[0006] Embodiments of this application provide a robot motion control method and device to improve the efficiency and stability of motion control.

[0007] In a first aspect, embodiments of this application provide a robot motion control method, including:

[0008] Obtain the current state of the robot; the current state includes the current position and the target position of the robot;

[0009] Input the current state into a motion control model to obtain the current action that the robot needs to execute; the motion control model is obtained by multi-task training of a reinforcement learning model; the multi-task training includes path planning task training, and at least one of gait control task training and balance control task training; the gait control task training and the balance control task training are carried out on the basis of the path planning task training;

[0010] Control the robot to execute the current action, and if the robot reaches the target position after executing the current action, the current motion task is completed.

[0011] In a possible design, the current state further includes: the grid attribute of the grid to which the current position belongs, joint angles, joint angular velocities, stride, stride frequency, external force action, body tilt angle, body angular velocity, and sole pressure value.

[0012] In a possible design, the method further includes:

[0013] Construct a first reinforcement learning model corresponding to the path planning task based on an initial policy function and an initial value function; the first reinforcement learning model includes a first reward function; the first reward function is related to whether a collision occurs with an obstacle and whether the distance to the end point is shortened;

[0014] Train the first reinforcement learning model to obtain a trained first reinforcement learning model;

[0015] Construct a second reinforcement learning model corresponding to the gait control task based on the trained first reinforcement learning model; the second reinforcement learning model includes a second reward function; the second reward function is related to the stride uniformity and the step frequency uniformity of walking;

[0016] Train the second reinforcement learning model to obtain a trained second reinforcement learning model;

[0017] Determine the motion control model according to the trained second reinforcement learning model.

[0018] In a possible design, the training of the first reinforcement learning model includes:

[0019] Update the value function based on the following expression:

[0020] ,

[0021] where Q is the value function, is the state at the current moment, is the state at the next moment, is the action at the current moment, is the learning rate, is the discount factor, is the reward obtained after executing the action at the current moment and represents the action at the next moment.

[0022] In a possible design, the training of the second reinforcement learning model includes:

[0023] Update the policy based on the following expression:

[0024]

[0025] where, are the parameters of the policy network, is the policy network, is the advantage function, , is the action-value function, is the state-value function.

[0026] In a possible design, determining the motion control model according to the trained second reinforcement learning model includes:

[0027] Constructing a third reinforcement learning model corresponding to the balance control task training based on the trained second reinforcement learning model; the third reinforcement learning model includes a third reward function; the third reward function is related to whether falling when being subjected to an external force and the balance recovery speed when being subjected to an external force;

[0028] Training the third reinforcement learning model to obtain the trained third reinforcement learning model;

[0029] Determining the motion control model according to the trained third reinforcement learning model.

[0030] In a possible design, determining the motion control model according to the balance policy function and the balance value function includes:

[0031] Constructing a fourth reinforcement learning model based on the trained third reinforcement learning model; the fourth reward function corresponding to the fourth reinforcement learning model is determined according to the first reward function, the second reward function, and the third reward function;

[0032] Training the fourth reinforcement learning model to obtain the motion control model.

[0033] In a possible design, training the fourth reinforcement learning model to obtain the motion control model includes:

[0034] Modeling the environment where the robot is located to obtain an environment model; the environment model includes a starting position and a target position;

[0035] Placing the robot at the starting position to obtain an initial environment state;

[0036] For each time step, determining the current action to be executed according to the current state, and after executing the current action, obtaining the next state and a reward value; the reward value is determined according to the first reward function;

[0037] Storing the current state, the current action, the next state, and the reward value into an experience replay buffer;

[0038] Randomly extracting a preset number of training data from the experience replay buffer;

[0039] Based on the training data, perform gradient descent training on the first reinforcement learning model to obtain the motion control model.

[0040] In a possible design, the fourth reinforcement learning model is a deep Q-network; the training of the fourth reinforcement learning model to obtain the motion control model includes:

[0041] For each time step, generate a random number M, where M is greater than or equal to 0 and less than or equal to 1;

[0042] If M is less than , then randomly select an action from the action space and execute it;

[0043] If M is greater than or equal to , then select the action corresponding to the maximum Q value among the Q values output by the Q-network in the current state as the target action, and execute the target action; is a probability value between 0 and 1, representing the degree of random exploration; the of the current iteration round is obtained by multiplying the of the previous iteration round by the decay factor; the decay factor is greater than 0 and less than 1.

[0044] In a second aspect, an embodiment of the present application provides a robot motion control device, including:

[0045] An acquisition module, configured to acquire the current state of the robot; the current state includes the current position and the target position of the robot;

[0046] An input module, configured to input the current state into the motion control model to obtain the current action that the robot needs to execute; the motion control model is obtained by performing multi-task training on a reinforcement learning model; the multi-task training includes path planning task training, and at least one of gait control task training and balance control task training; the gait control task training and the balance control task training are performed on the basis of the path planning task training;

[0047] A control module, configured to control the robot to execute the current action, and if the robot reaches the target position after executing the current action, the current motion task is completed.

[0048] In a third aspect, an embodiment of the present application provides a robot motion control device, including: at least one processor and a memory;

[0049] The memory stores computer execution instructions;

[0050] The at least one processor executes the computer-executable instructions stored in the memory, such that the at least one processor executes the method as described in the first aspect above and various possible designs of the first aspect.

[0051] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method as described in the first aspect above and various possible designs of the first aspect.

[0052] In a fifth aspect, an embodiment of the present application provides a computer program product including a computer program, which, when executed by a processor, implements the method as described in the first aspect above and various possible designs of the first aspect.

[0053] The robot motion control method and device provided in this embodiment, the method includes obtaining the current state of the robot, the current state including the current position and the target position of the robot, inputting the current state into a motion control model to obtain the current action that the robot needs to execute, the motion control model being obtained by performing multi-task training on a reinforcement learning model, the multi-task training including path planning task training and at least one of gait control task training and balance control task training, controlling the robot to execute the current action, and if the robot reaches the target position after executing the current action, the current motion task is completed. The robot motion control method provided in the embodiments of the present application, by combining gait control training or balance control training with path planning training based on a reinforcement learning algorithm, trains to obtain a motion control model, enabling the robot to improve stability on the basis of being able to efficiently and reasonably plan a path. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] To more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0055] Figure 1 It is a schematic flowchart of the robot motion control method provided in the embodiments of the present application;

[0056] Figure 2 It is a schematic flowchart of the training process of the motion control model provided in the embodiments of the present application Figure 1 ;

[0057] Figure 3 It is a schematic flowchart of the training process of the motion control model provided in the embodiments of the present application Figure 2 ;

[0058] Figure 4 This is a schematic structural diagram of the robot motion control device provided by the embodiments of the present application;

[0059] Figure 5 This is a schematic hardware structure diagram of the robot motion control device provided by the embodiments of the present application. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0061] It should be noted that the robot motion control method provided by the present application can be used in the field of humanoid robot technology, and can also be used in any field other than the field of humanoid robots. The application field of the robot motion control method provided by the present application is not limited.

[0062] With the rapid development of technology, humanoid robots have shown great application potential in many fields, such as industrial manufacturing, medical care, home services, and rescue. However, achieving efficient, stable, and intelligent motion control of humanoid robots in complex environments has always been a major challenge in the field of robot technology.

[0063] In the related art, path planning usually adopts the method based on geometric algorithms or search algorithms. For example, the algorithm can be used to find the path from the starting point to the ending point by calculating the heuristic cost of nodes in a known map environment. However, although this method can achieve certain effects in simple and static environments, it has obvious limitations for complex and changeable actual scenarios. When there are dynamic obstacles in the environment or the map information is inaccurate, traditional path planning algorithms are difficult to quickly and effectively adjust the path, which easily causes the robot to get into trouble or collide, and affects the stable balance of the robot during walking.

[0064] To solve the above technical problems, the inventors of the present application have found that a reinforcement learning algorithm can be used to enable the robot to continuously try mistakes in the environment and learn the optimal behavior strategy according to the reward signal feedback by the environment to cope with situations such as obstacles in the environment or inaccurate map information. And, through the reinforcement learning algorithm, gait control training or balance control training is carried out together with path planning training, so that the robot can improve stability on the basis of being able to efficiently and reasonably plan the path.

[0065] The technical solution of the present application will be described in detail below with specific embodiments. These several specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0066] Figure 1 It is a schematic flowchart of the robot motion control method provided by the embodiment of the present application. As Figure 1 shown, the method includes:

[0067] 101. Obtain the current state of the robot; the current state includes the current position and the target position of the robot.

[0068] The execution subject of this embodiment may be a humanoid robot, a control module in the humanoid robot, etc. The control module may be a hardware, software, or a combination of hardware and software component.

[0069] In some embodiments, the current state may further include at least one of the following: the grid attribute of the grid to which the current position belongs, joint angles, joint angular velocities, stride, stride frequency, external force action, body tilt angle, body angular velocity, sole pressure value.

[0070] Specifically, during the movement of the robot, information such as the position, posture, joint state, and sensor data of the robot are obtained in real time, and these information are fed back as the current state to the motion control module of the robot (for performing at least one of path planning, gait control, and balance control).

[0071] 102. Input the current state into the motion control model to obtain the current action that the robot needs to execute; the motion control model is obtained by multi-task training of a reinforcement learning model; the multi-task training includes path planning task training, and at least one of gait control task training and balance control task training; the gait control task training and the balance control task training are carried out on the basis of the path planning task training.

[0072] Specifically, the motion control model is used to determine the current action that needs to be executed corresponding to the current state based on the trained policy function.

[0073] In this embodiment, during the process of training the motion control model, only two tasks can be trained. Specifically, first, the path planning task is trained, and then the gait control task or the balance control task is trained, that is, the training is completed. It is also possible to train three tasks. Specifically, considering the different degrees of influence of the three-task training on the stability of the robot, the training can be carried out in descending order according to the degree of influence from high to low. Exemplarily, the path planning task can be trained first, then the gait control task, and finally the balance control task.

[0074] 103. Control the robot to execute the current action. If the robot reaches the target position after executing the current action, the current motion task is completed.

[0075] Specifically, after the robot executes the current action, it can judge whether it has reached the target position. If it has reached, the task is completed. If it has not reached, it returns to repeat steps 101 to 103 until it reaches the target position or receives a stop instruction.

[0076] The robot motion control method provided in this embodiment, through the reinforcement learning algorithm, combines the gait control training or the balance control training with the path planning training, and trains to obtain a motion control model, so that the robot can improve stability on the basis of being able to efficiently and reasonably plan the path.

[0077] Figure 2 It is a flow chart of the training process of the motion control model provided in the embodiment of the present application Figure 1 As Figure 2 shown, the method includes:

[0078] 201. Construct a first reinforcement learning model corresponding to the path planning task training based on the initial policy function and the initial value function; the first reinforcement learning model includes a first reward function; the first reward function is related to whether a collision occurs with an obstacle and whether the distance to the end point is shortened.

[0079] 202. Train the first reinforcement learning model to obtain the trained first reinforcement learning model.

[0080] Specifically, first, parameters related to robot motion control can be defined, including but not limited to the learning rate, the discount factor, environmental parameters (such as the grid size of path planning, the number of joints for gait control and related physical parameters, the number of sensors for balance control and related thresholds, etc.). Initialize the value function and the policy function in the first reinforcement learning model to obtain the initial policy function and the initial value function.

[0081] For training on path planning tasks, the environment can first be set up: The environment where the robot is located is discretized or parametrically modeled. For example, its activity space is divided into small grid regions, each with different attributes (such as whether there are obstacles, the flatness of the ground, etc.), which serves as the environment for reinforcement learning. The position of the robot in this environment and the target end position constitute the state space. Secondly, the actions can be defined: The actions that the robot can take can include moving a certain step length in different directions (such as moving forward, backward, left, or right a short distance), and multiple step length actions constitute the action space. Finally, the rewards can be designed: When the robot moves towards the end point and does not collide with obstacles, a certain positive reward is given; if it collides with an obstacle, a large negative reward is given; when it successfully reaches the end point, a high positive reward is given. By continuously trying and making mistakes in this environment (that is, continuously trying different action sequences), the robot can learn the optimal action strategy from the starting point to the end point, thereby achieving more efficient and safer path planning, and obtaining the principle of the optimal control strategy through continuous sampling (trial and error).

[0082] Exemplarily, the training process of the first reinforcement learning model corresponding to the path planning task training may include the following steps:

[0083] Step 2021, define the state representation of the first reinforcement learning model: Let the position of the robot in the two-dimensional plane be , and the target position be . The environment is discretized into a grid of , and the state is represented by , where represents the grid where the robot is located (such as whether there is an obstacle, which can be represented by 0 for no obstacle and 1 for an obstacle).

[0084] Step 2022, define the action space of the first reinforcement learning model: Define the action set , where represents moving forward a distance of one grid cell , represents moving backward a distance of one grid cell , represents moving left a distance of one grid cell , represents moving right a distance of one grid cell .

[0085] Step 2023, calculate the reward function of the first reinforcement learning model: Let the current position of the robot be , the target position be , and the distance function be .

[0086] If the robot does not encounter an obstacle after moving and is closer to the target, then the reward . Among them, is the distance after moving, is the distance before moving, is a positive reward less than a preset value, and the preset value can be less than or equal to 0.2. For example can be 0.1.

[0087] If an obstacle is encountered, then the reward = -10.

[0088] If the target position is reached, then the reward = 100.

[0089] Step 2024, calculate the state transition probability of the first reinforcement learning model: Assume that the environment is deterministic (the actions of the robot can be accurately executed), then the state transition probability is:

[0090] If the action is legal (it will not cause the robot to go beyond the environment boundary or encounter an obstacle), then is the state after updating the position according to the action . .

[0091] If the action is illegal, then , .

[0092] Step 2025, calculate the value function update process of the first reinforcement learning model: Use the Q-learning algorithm to update the value function , and the update formula is:

[0093] (1)

[0094] Among them, Q is the value function, is the state at the current moment, is the state at the next moment, is the action at the current moment, is the learning rate, is the discount factor, is the reward obtained after executing the action at the current moment, represents the action at the next moment.

[0095] Repeat steps 2023 to 2025 until the robot reaches the target position or the maximum number of iterations is reached, to obtain the optimal path planning strategy corresponding to the maximum reward value from the starting point to the ending point, and obtain the trained first reinforcement learning model.

[0096] 203. Construct a second reinforcement learning model corresponding to the gait control task based on the trained first reinforcement learning model; the second reinforcement learning model includes a second reward function; the second reward function is related to the stride uniformity and step frequency uniformity of walking.

[0097] Specifically, after obtaining the trained first reinforcement learning model, the policy function and value function of the trained first reinforcement learning model can be used as the initial policy function and initial value function of the second reinforcement learning model, and a reward function specifically for the gait control task can be designed for the second reinforcement learning model.

[0098] 204. Train the second reinforcement learning model to obtain the trained second reinforcement learning model.

[0099] Specifically, for the gait control task, first, the environment and state representation can be defined: the parameters such as the angles, angular velocities, and torques of each joint of the robot, as well as the body posture (such as the inclination angle, center of gravity position, etc.) are used as the state to describe the current motion state of the robot, and the entire motion scene (such as the terrain condition, whether there is external force interference, etc.) is used as the environmental information. Secondly, the action setting can be carried out: the actions can be operations such as torque adjustment applied to each joint and change of joint angle, and these actions will affect the gait of the robot. Finally, the reward mechanism can be designed: if the gait of the robot can be kept stable and conform to the expected motion pattern (such as uniform stride and stable rhythm during walking), a positive reward is given; if unstable situations such as gait imbalance and falling occur, a negative reward is given. Using the reinforcement learning algorithm, the robot can learn the optimal action strategy for achieving a stable gait in different environments (such as flat ground, rough terrain, etc.) through continuous experiments, that is, by continuously sampling (trying the influence of different actions on the gait) to optimize the gait control based on the dynamic model.

[0100] Exemplarily, the training process of training the second reinforcement learning model corresponding to the gait control task can include the following steps:

[0101] Step 2041. Define the state representation: Assume the robot has joints, and use X=(X1,X2,…,X n ) to represent the joint angles, to represent the joint angular velocities, F=(F1, F2,…, F j) represents the external forces acting on the robot body (such as gravity, friction, etc.), where both n and j are positive integers greater than 1. represents the state.

[0102] Step 2042, Define the action space: The action represents the torque adjustment applied to n joints. is the torque applied to the i-th joint.

[0103] Step 2043, Calculate the reward value based on the reward function: Let be a function to measure gait stability, which can be defined according to, for example, the step deviation and body sway degree when the robot is walking.

[0104] The reward , if the robot does not fall within a time step and the gait is more stable (the value of G is smaller), then a higher reward is obtained.

[0105] Step 2044, Define the dynamics model (state transition): According to the dynamics equation of the robot , where is a function based on the physical structure and mechanical principles of the robot, describing the change of the state over time. For example, for a simple single-joint robot, the following expression can be adopted:

[0106] (2)

[0107] where, is the moment of inertia, is the mass, is the acceleration due to gravity, is the link length, is the damping coefficient, represents the joint angular velocity, is the first derivative of

[0108] Step 2045, Policy update (policy gradient method): The policy gradient algorithm can be used to update the policy , and the gradient update formula is:

[0109] (3)

[0110] where, are the parameters of the policy network, is the policy network, is the advantage function, , is the action-value function, is the state-value function.

[0111] Repeat steps 2043 to 2045 until the stop condition is met (e.g., reaching the iteration number threshold). During the process of the robot moving along the planned path, continuously adjust the gait control strategy to enable the robot to walk stably.

[0112] 205. Construct a third reinforcement learning model corresponding to the balance control task based on the trained second reinforcement learning model; the third reinforcement learning model includes a third reward function. The third reward function is related to whether falling when subjected to an external force and the speed of restoring balance when subjected to an external force.

[0113] Specifically, after obtaining the trained second reinforcement learning model, the policy function and value function of the trained second reinforcement learning model can be used as the initial policy function and initial value function of the third reinforcement learning model, and a reward function specifically for the balance control task can be designed for the third reinforcement learning model.

[0114] 206. Train the third reinforcement learning model to obtain the trained third reinforcement learning model.

[0115] Specifically, for the training of the balance control task, first define the environment and state: Combine factors such as the robot's posture (e.g., body tilt angle, angular velocity, etc.), the contact state between the soles of the feet and the ground (e.g., pressure distribution, etc.), and external disturbance forces (if any) as state information to describe the current balance status of the robot. The surrounding terrain conditions, the existence of sudden external forces, etc. constitute the environment. Secondly, define actions: The actions can be fine-tuning different parts of the robot's body (such as leg joints, waist, etc.) to adjust the center of gravity, change the supporting force, etc., with the aim of maintaining balance. Finally, reward settings can be made: When the robot can maintain balance under external disturbances (e.g., still stand firm after being slightly pushed), give a positive reward; if it loses balance and falls, give a negative reward. Let the robot continuously try and make mistakes (try different balance adjustment actions) in such an environment through the reinforcement learning algorithm, learn the optimal action strategy for maintaining balance in various situations, and achieve the transition from continuously sampling and trying to make mistakes in the simulated environment to effectively performing balance control in the real scenario (i.e., adopt the sim-to-real method, first train the balance control strategy in the simulated environment and then apply it to the actual scenario).

[0116] Exemplarily, the training process of the third reinforcement learning model corresponding to the balance control task may include the following steps:

[0117] Step 2061. Define the state representation: Let the tilt angle of the robot's body be , the angular velocity be , and the measured value of the sole pressure sensor be Assume the robot has a sole pressure sensor), the state is .

[0118] Step 2062, define the action space: The action represents an operation to finely adjust a key part of the robot's body, such as adjusting the leg joint angle, changing the waist torque, etc., to adjust the center of gravity and support force.

[0119] Step 2063, calculate the reward value based on the reward function: Let be a function to measure the balance state, which can be defined according to whether the tilt angle is within the stable range, whether the sole pressure distribution is uniform, etc. Among them, is the tilt angle of the robot's body, is the angular velocity of the robot, and P is the measured value of the robot's sole pressure sensor.

[0120] The reward is , if the robot can recover balance faster after being disturbed externally (that is, the value of is smaller), then a higher reward is obtained.

[0121] Step 2064, calculate the state transition probability (in the presence of external disturbances): Assume that the external disturbance is a random variable, and the state transition probability depends on the physical structure of the robot, the control action and the external disturbance d. For example, according to the rigid body kinematics and dynamics equations of the robot, can be established, where α t is the tilt angle of the robot's body at the current moment, α t+1 is the predicted value of the tilt angle of the robot's body at the next moment, ω t is the angular velocity of the robot at the current moment, is the time interval, is a function obtained according to physical principles.

[0122] Step 2065, perform policy update: The asynchronous advantage actor-critic (A3C) algorithm can be used to update the policy and value function in multiple parallel threads. For the policy update part, the gradient of the policy can be updated according to the advantage function, and the value function can be updated simultaneously to better estimate the state value. The specific formula involves the gradient update of the policy network and the value network , such as the policy network gradient update formula: (where, are the parameters of the policy network), r is the reward, is the discount factor, V() is the state-value function, S t is the state at the current moment, S t+1 is the state at the next moment.

[0123] Steps 2063 to 2065 are repeatedly executed until a stop condition is met (such as reaching an iteration count threshold). During the process of the robot moving along the planned path, external disturbances and its own state are continuously monitored, and actions are adjusted according to the balance control strategy to cope with various terrain changes and external disturbances and maintain balance.

[0124] 207. Determine the motion control model according to the trained third reinforcement learning model.

[0125] The robot motion control method provided in this embodiment, through the reinforcement learning algorithm, conducts gait control training or balance control training together with path planning training to train and obtain a motion control model, enabling the robot to improve stability on the basis of being able to efficiently and reasonably plan a path.

[0126] In some embodiments, as Figure 3 shown, on the basis of the above embodiments, for example, on the basis of the above Figure 2 shown embodiments, step 207 may include:

[0127] 2071. Construct a fourth reinforcement learning model based on the trained third reinforcement learning model; the fourth reward function corresponding to the fourth reinforcement learning model is determined according to the first reward function, the second reward function, and the third reward function.

[0128] Specifically, after obtaining the trained third reinforcement learning model, the policy function and value function of the trained third reinforcement learning model can be used as the initial policy function and initial value function of the fourth reinforcement learning model, and the fourth reward function corresponding to the fourth reinforcement learning model is determined according to the first reward function, the second reward function, and the third reward function.

[0129] 2072. Train the fourth reinforcement learning model to obtain the motion control model.

[0130] In some embodiments, training the fourth reinforcement learning model to obtain the motion control model may include: modeling the environment where the robot is located to obtain an environment model; the environment model includes a starting position and a target position; placing the robot at the starting position to obtain an initial environmental state; for each time step, determining a current action to be executed according to the current state, and after executing the current action, obtaining the next state and a reward value; the reward value is determined according to the first reward function; storing the current state, the current action, the next state, and the reward value into an experience replay buffer; randomly extracting a preset number of training data from the experience replay buffer; and performing gradient descent training on the first reinforcement learning model according to the training data to obtain the motion control model.

[0131] Specifically, after obtaining the trained third reinforcement learning model, considering the mutual influence among path planning, gait control, and balance control, in order to further balance the three, a comprehensive reward function, i.e., the fourth reward function, can be designed to take into account the reward factors of path planning, gait control, and balance control.

[0132] Exemplarily, for the design of the fourth reward function: the fourth reward function can be defined as . Wherein, R{path} is the reward related to path planning. For example, for every meter the robot approaches the target position, R{path} increases by 10 points. R{gait} is the reward related to gait control. Each step of a normal and stable gait is rewarded 2 points, and points are deducted if gait abnormalities occur. R{balance} is the balance control reward. The robot is rewarded 3 points per second for maintaining upright balance during movement, and points are deducted if the tilt angle exceeds a certain threshold. w1, w2, and w3 are weights, which are allocated according to the task priorities. For example, w1 = 0.4, w2 = 0.3, and w3 = 0.3.

[0133] During the reinforcement learning training process, the robot takes actions according to the environmental state, and calculates the reward value according to this reward function after each action. The agent will learn how to ensure good gait and balance while obtaining a high path planning reward. Specifically, it may include the following steps:

[0134] Environment construction: Construct a simulation environment for the humanoid robot, including terrain such as flat ground, slopes, obstacles, etc., a starting position, and a target position.

[0135] Agent Setup: Define the humanoid robot as an agent. Its action space includes discrete or continuous actions such as moving forward, backward, turning left, turning right, adjusting stride, and changing step frequency. The state space covers the robot's position information (x, y coordinates), speed, attitude angles (pitch angle, yaw angle, roll angle), surrounding environment information sensed by sensors (such as distance to obstacles), and current gait parameters (stride, step frequency, etc.).

[0136] Reinforcement Learning Algorithm Selection: The fourth reinforcement learning model is determined based on the third reinforcement learning model. Therefore, the fourth reinforcement learning model can adopt algorithms such as Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) in the same way as the third reinforcement learning model. Taking DQN as an example, initialize the parameters of the Q-network based on the trained third reinforcement learning model, including the weights and biases of the neural network.

[0137] The training loop process can specifically include the following steps:

[0138] For each training episode:

[0139] Reset the environment: Place the robot at the starting position and initialize the environmental state s0.

[0140] For each time step t:

[0141] According to the current state , the agent selects an action through the Q-network (in DQN) or the policy network (in PPO). . For example, select an action according to the magnitude of the Q value, or sample an action according to the action probability distribution output by the policy network.

[0142] Execute the action After that, the robot generates a new state in the environment , and obtains a reward , calculated by the fourth reward function . For example, if the robot moves forward a certain distance towards the target position and maintains balance and a good gait, it will receive a positive reward; if the robot falls or deviates from the path plan, it will receive a negative reward.

[0143] Store the data in the experience replay buffer. In DQN, the experience replay buffer is used to store a certain number of historical experiences for subsequent random sampling for network training.

[0144] Randomly sample a batch of data from the experience replay buffer . For the DQN algorithm, calculate the target Q value: , where is the discount factor, which is used to weigh the importance of the current reward and future rewards. are the parameters of the target Q-network (the target network is used in DQN to stabilize the training process). Then, according to the loss function perform gradient descent training on the Q-network to update the network parameters θ1. For algorithms based on policy gradients such as PPO, the gradients are calculated according to the policy gradient theorem and the parameters of the policy network are updated to maximize the expected cumulative reward.

[0145] When the robot reaches the target position or exceeds the maximum number of time steps, one training episode ends.

[0146] To improve the performance of the model, training evaluation and optimization can be carried out. Specifically, the performance of the agent can be evaluated every certain number of training episodes. For example, calculate metrics such as the average number of steps, average reward, and number of falls when the robot reaches the target position in multiple test training episodes. Adjust the training parameters according to the evaluation results, such as the learning rate, discount factor, weights w1, w2, w3 of the reward function, etc. If it is found that the robot pays too much attention to path planning and ignores balance, resulting in frequent falls, the weight w3 of the balance reward can be appropriately increased.

[0147] Repeat the above training loop until the performance of the agent reaches a satisfactory level, that is, the robot can efficiently and stably reach the target position from the starting position in various complex environments, while maintaining good gait and balance control.

[0148] To clearly illustrate the training process, the following is an example of the training process for the Q-network:

[0149] First, collect interaction data. Let the robot act in the environment according to the current policy, which selects actions based on the output of the Q-network. For example, the robot is in the initial state and selects the action with the highest Q-value according to the Q-value estimates of different actions by the Q-network to execute, such as moving forward in a straight line for a certain distance.

[0150] Record the state , action , reward and the next state for each interaction, and store these experience data in the experience replay buffer.

[0151] Sample from the experience replay: Randomly sample a batch of experience data from the experience replay buffer, denoted as , where N is the number of sampled samples.

[0152] Secondly, calculate the target Q-value. For each sampled experience data, use the target Q-network to calculate the target Q-value: The target value is used to guide the training of the Q-network. Among them, is the discount factor, which is used to balance the importance of the current reward and future rewards. are the weight parameters of the target Q-network, S i+1 is the next state, and a' is the optimal action among all possible actions that can be taken in the next state S i+1 . If the current state is a terminal state, then . Among them, the target Q-network has the same structure as the trained Q-network. The target Q-network is used to provide a relatively stable estimate of the target Q-value. The trained Q-network is used to collect experience data.

[0153] Again, calculate the loss function and update the Q-network. Use the mean squared error loss function to calculate the error between the predicted Q-value of the current Q-network and the target Q-value: . Among them, is the target Q-value, is the predicted Q-value.

[0154] Calculate the gradient of the loss function with respect to the Q-network parameters through the backpropagation algorithm , and use an optimization algorithm (such as Adam) to update the parameters , and the update formula is: where is the learning rate. During the update process, the weight parameters in the target Q-network remain unchanged.

[0155] Once again, perform policy update. According to the updated Q-network, use the greedy policy or other policy improvement methods to update the policy. For example, in the greedy policy, with probability 1 - select the action with the largest Q-value, and with probability randomly select an action. As the training progresses, gradually decrease the value of so that the robot more often selects the optimal action based on the estimate of the Q-network. It should be noted that the greedy policy is a simple and commonly used method for balancing between exploring new actions and exploiting known optimal actions. Among them, is a probability value between 0 and 1, representing the degree of random exploration.

[0156] The specific process of policy update can include: performing action selection: at each time step, when the robot needs to select an action, first generate a random number M uniformly distributed in the interval [0, 1]. Then compare the random number M with : if M is less than , indicating exploration, the robot will randomly select an action from the action space for execution. For example, if the robot's action space contains 5 actions (move forward, turn left, turn right, adjust stride, adjust balance torque), then one will be randomly selected from the 5 actions with equal probability, regardless of the Q-value evaluation of these actions by the current Q-network. If M is greater than or equal to , indicating exploitation, the robot will select the action with the largest Q-value output by the Q-network in the current state for execution. That is, by inputting the current state s into the already updated Q-network , obtaining the Q-values corresponding to different actions , and then selecting the action with the largest Q-value as the action to be executed.

[0157] In some embodiments, the value can be adjusted during training: at the initial stage of training, a relatively large value is usually set, such as = 0.5 or even higher, to ensure that the robot has a greater probability of exploring different actions in the environment, trying various possible path planning methods, gait control adjustments, and balance control operations, so as to collect more data under different circumstances and help the Q-network better learn the value of different state-action pairs. As the training progresses, that is, after multiple rounds of interaction with the environment and updates of the Q-network, the value will gradually decrease. For example, the in the current iteration round is obtained by multiplying the in the previous iteration round by a decay factor; the decay factor is greater than 0 and less than 1. Also, for example, according to a certain decay rule, like every several training episodes, let be multiplied by a decay factor less than 1, such as 0.99, to make it gradually become smaller. When becomes very small, such as approaching 0.05 or smaller, the robot will rely more on the optimal actions estimated by the Q-network to make decisions, that is, more exploitation, because at this time the Q-network has undergone a large amount of data learning and parameter updates, and its estimation of the action value in each state is relatively more accurate, so the probability that the robot selects the optimal action to achieve path planning, maintain a good gait, and balance control is greater.

[0158] Finally, continuous training and optimization can be carried out. Continuously repeat the above steps, allowing the robot to continuously interact in the environment and collect experience data, updating the Q-network and the policy. After multiple iterations, the Q-network of the robot will gradually learn the value of each action in different states, and the policy will also be continuously optimized, so as to perform better and better in path planning, gait control, and balance control.

[0159] Figure 4This is a schematic structural diagram of the robot motion control device provided by the embodiments of the present application. As Figure 4 shown, the robot motion control device 40 includes: an acquisition module 401 and an input module 402.

[0160] The acquisition module 401 is used to acquire the current state of the robot; the current state includes the current position and the target position of the robot.

[0161] The input module 402 is used to input the current state into the motion control model to obtain the current action that the robot needs to execute; the motion control model is obtained by performing multi-task training on the reinforcement learning model; the multi-task training includes path planning task training, and at least one of gait control task training and balance control task training; the gait control task training and the balance control task training are performed on the basis of the path planning task training.

[0162] The control module 403 is used to control the robot to execute the current action, and if the robot reaches the target position after executing the current action, the current motion task is completed.

[0163] The robot motion control device provided by the embodiments of the present application, by combining gait control training or balance control training with path planning training based on the reinforcement learning algorithm, trains to obtain a motion control model, so that the robot can improve stability on the basis of being able to efficiently and reasonably plan the path.

[0164] The robot motion control device provided by the embodiments of the present application can be used to execute the above method embodiments, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0165] In some embodiments, the current state further includes: the grid attribute of the grid to which the current position belongs, joint angle, joint angular velocity, stride, stride frequency, external force action, body tilt angle, body angular velocity, and sole pressure value.

[0166] In some embodiments, the device 40 further includes a training module (not shown), and the training module is used to: construct a first reinforcement learning model corresponding to the path planning task training based on the initial policy function and the initial value function; train the first reinforcement learning model to obtain a trained first reinforcement learning model; construct a second reinforcement learning model corresponding to the gait control task based on the trained first reinforcement learning model; train the second reinforcement learning model to obtain a trained second reinforcement learning model; determine the motion control model according to the trained second reinforcement learning model.

[0167] In some embodiments, the training module is used to update the value function based on the following expression:

[0168] ,

[0169] where Q is the value function, is the state at the current moment, is the state at the next moment, is the action at the current moment, is the learning rate, is the discount factor, is the reward obtained after executing the action at the current moment , represents the action at the next moment.

[0170] In some embodiments, the training module is used to update the policy based on the following expression:

[0171]

[0172] where, are the parameters of the policy network, is the policy network, is the advantage function, , is the action-value function, is the state-value function.

[0173] In some embodiments, the first reinforcement learning model includes a first reward function; the first reward function is related to whether a collision occurs with an obstacle and whether the distance to the end point is shortened; the second reinforcement learning model includes a second reward function; the second reward function is related to the uniformity of the walking stride and the uniformity of the walking frequency.

[0174] In some embodiments, the training module is specifically configured to: construct a third reinforcement learning model corresponding to the balance control task training based on the trained second reinforcement learning model; the third reinforcement learning model includes a third reward function; the third reward function is related to whether a fall occurs when an external force is applied and the balance recovery speed when an external force is applied; train the third reinforcement learning model to obtain the trained third reinforcement learning model; determine the motion control model according to the trained third reinforcement learning model.

[0175] In some embodiments, the training module is specifically configured to: construct a fourth reinforcement learning model based on the trained third reinforcement learning model; the fourth reward function corresponding to the fourth reinforcement learning model is determined according to the first reward function, the second reward function, and the third reward function; train the fourth reinforcement learning model to obtain the motion control model.

[0176] In some embodiments, the training module is specifically configured to: model the environment where the robot is located to obtain an environment model; the environment model includes a starting position and a target position; place the robot at the starting position to obtain an initial environmental state; for each time step, determine a current action to be executed according to the current state, and after executing the current action, obtain the next state and a reward value; the reward value is determined according to the first reward function; store the current state, the current action, the next state, and the reward value into an experience replay buffer; randomly extract a preset number of training data from the experience replay buffer; and perform gradient descent training on the first reinforcement learning model according to the training data to obtain the motion control model.

[0177] In some embodiments, the training module is specifically configured to: for each time step, generate a random number M, where M is greater than or equal to 0 and less than or equal to 1; if M is less than , then randomly select an action from the action space and execute it; if M is greater than or equal to , then select the action corresponding to the maximum Q value in the Q values output by the Q-network in the current state as the target action, and execute the target action; is a probability value between 0 and 1, indicating the degree of random exploration; in the current iteration round is obtained by multiplying in the previous iteration round by a decay factor; the decay factor is greater than 0 and less than 1.

[0178] Figure 5 FIG.

[0179] Device 50 may include one or more of the following components: a processing component 501, a memory 502, a power component 503, a multimedia component 504, an audio component 505, an input / output (I / O) interface 506, a sensor component 507, and a communication component 508.

[0180] The processing component 501 generally controls the overall operation of the device 50, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 501 may include one or more processors 509 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 501 may include one or more modules to facilitate the interaction between the processing component 501 and other components. For example, the processing component 501 may include a multimedia module to facilitate the interaction between the multimedia component 504 and the processing component 501.

[0181] The memory 502 is configured to store various types of data to support the operation of the device 50. Examples of such data include instructions for any application or method operating on the device 50, contact data, phone book data, messages, pictures, videos, and the like. The memory 502 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0182] The power supply component 503 provides power to various components of the device 50. The power supply component 503 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 50.

[0183] The multimedia component 504 includes a screen that provides an output interface between the device 50 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 504 includes a front camera and / or a rear camera. When the device 50 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0184] The audio component 505 is configured to output and / or input audio signals. For example, the audio component 505 includes a microphone (MIC) that is configured to receive external audio signals when the device 50 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 502 or transmitted via the communication component 508. In some embodiments, the audio component 505 further includes a speaker for outputting audio signals.

[0185] The I / O interface 506 provides an interface between the processing component 501 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power-on button, and a lock button.

[0186] The sensor assembly 507 includes one or more sensors for providing a status assessment of various aspects of the device 50. For example, the sensor assembly 507 can detect the on / off state of the device 50, the relative positioning of components, such as the display and keypad of the device 50. The sensor assembly 507 can also detect a change in the position of the device 50 or a component of the device 50, the presence or absence of user contact with the device 50, the orientation or acceleration / deceleration of the device 50, and the temperature change of the device 50. The sensor assembly 507 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 507 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 507 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0187] The communication component 508 is configured to facilitate communication between the device 50 and other devices in a wired or wireless manner. The device 50 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 508 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 508 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0188] In an exemplary embodiment, the device 50 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0189] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions, such as the memory 502 including instructions, is also provided. The above instructions can be executed by the processor 509 of the device 50 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0190] The above-mentioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0191] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0192] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical discs.

[0193] The embodiments of the present application also provide a computer program product, including a computer program, which, when executed by a processor, implements the robot motion control method executed by the above robot motion control device.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A robot motion control method, characterized in that: include: Get the current state of the robot; The current state includes the current position and target position of the robot; The current state is input into the motion control model to obtain the current action that the robot needs to perform; the motion control model is obtained by performing multi-task training and comprehensive training on the reinforcement learning model; the multi-task training includes path planning task training, gait control task training and balance control task training; the gait control task training is performed on the basis of the path planning task training, and the balance control task training is performed on the basis of the gait control task training; the comprehensive training is performed after the balance control task training; the reward function of the comprehensive training includes the reward factor of the multi-task training; The robot is controlled to execute the current action. If the robot reaches the target position after completing the current action, the current motion task is completed.

2. The method according to claim 1, characterized in that The current state also includes: grid attributes of the grid to which the current position belongs, joint angles, joint angular velocities, stride length, step frequency, external force, body inclination angle, body angular velocity, and sole pressure value.

3. The method according to claim 1 or 2, characterized in that: The method further comprises: A first reinforcement learning model corresponding to the path planning task training is constructed based on the initial strategy function and the initial value function; the first reinforcement learning model includes a first reward function; the first reward function is related to whether a collision with an obstacle occurs and whether the distance to the end point is shortened; Training the first reinforcement learning model to obtain a trained first reinforcement learning model; Constructing a second reinforcement learning model corresponding to the gait control task based on the trained first reinforcement learning model; the second reinforcement learning model includes a second reward function; the second reward function is related to the uniformity of walking stride and the uniformity of step frequency; Training the second reinforcement learning model to obtain a trained second reinforcement learning model; The motion control model is determined according to the trained second reinforcement learning model.

4. The method according to claim 3, characterized in that The training of the first reinforcement learning model includes: The value function is updated based on the following expression of the gradient update formula: , Where Q is the value function, is the current state, For the state of the next moment, is the action at the current moment, is the learning rate, is the discount factor, Is the action that executes the current moment After the reward, Indicates the action at the next moment.

5. The method according to claim 3, characterized in that: The training of the second reinforcement learning model comprises: The policy update is based on the following expression: in, are the parameters of the policy network, is a strategic network, is the advantage function, , is the action-value function, is the state-value function.

6. The method according to claim 3, characterized in that Determining the motion control model according to the trained second reinforcement learning model includes: Constructing a third reinforcement learning model corresponding to the balance control task training based on the trained second reinforcement learning model; the third reinforcement learning model includes a third reward function; the third reward function is related to whether to fall down when subjected to external force, and the speed of restoring balance when subjected to external force; Training the third reinforcement learning model to obtain a trained third reinforcement learning model; The motion control model is determined according to the trained third reinforcement learning model.

7. The method according to claim 6, characterized in that Determining the motion control model according to the trained third reinforcement learning model includes: Constructing a fourth reinforcement learning model based on the trained third reinforcement learning model; a fourth reward function corresponding to the fourth reinforcement learning model is determined according to the first reward function, the second reward function and the third reward function; The fourth reinforcement learning model is trained to obtain the motion control model.

8. The method according to claim 7, characterized in that The step of training the fourth reinforcement learning model to obtain the motion control model includes: Modeling the environment in which the robot is located to obtain an environment model; the environment model includes a starting position and a target position; Placing the robot at the starting position to obtain an initial environment state; For each time step, according to the current state, determine the current action to be performed, and after performing the current action, obtain the next state and reward value; the reward value is determined according to the fourth reward function; Storing the current state, the current action, the next state, and the reward value in an experience replay buffer; Randomly extracting a preset amount of training data from the experience replay buffer; According to the training data, gradient descent training is performed on the fourth reinforcement learning model to obtain the motion control model.

9. The method according to claim 8, characterized in that The fourth reinforcement learning model is a deep Q network; and the training of the fourth reinforcement learning model includes: For each time step, generate a random number M, M is greater than or equal to 0 and M is less than or equal to 1; If M is less than , then randomly select an action from the action space to execute; If M is greater than or equal to , then select the action corresponding to the maximum Q value among the Q values ​​output by the Q network in the current state as the target action, and execute the target action; is a probability value between 0 and 1, indicating the degree of random exploration; The previous iteration round The attenuation factor is greater than 0 and less than 1.

10. A robot motion control device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the robot motion control method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Biped robot gait control method and control device

    CN113467235A

  • Biped robot gait control method and device, storage medium and equipment

    CN117572877A

  • Biped robot multi-mode adaptive gait generation method based on reinforcement learning

    CN118672291A

  • Efficient reinforcement learning based on merging of trained learners

    US20200193333A1

  • Ai-based control for robotics systems and applications

    US20240100694A1